Go + Redis Distributed Lock

Contents

Interview notes on Go + Redis distributed locks. Covers: why a distributed lock is needed, atomic locking with SET NX EX, Lua-script unlock, expiry and watchdog renewal, master-replica failover and Redlock.

Why you need a distributed lock

sync.Mutex only protects goroutines inside one process: the lock state lives in memory that other processes can’t see. Once a service runs on multiple instances, the same logic executes in several processes (possibly on several machines) with no shared memory, so you need a lock every instance recognizes.

Redis is shared storage that every instance can reach. Modeling the lock as a Redis key with an expiry is the most common approach.

Typical use cases:

  • Scheduled jobs: any instance may trigger the same job, but only one may run it.
  • Inventory deduction / flash sales: multiple instances decrement the same product’s stock.
  • Deduplication: message consumption and reconciliation jobs must process each batch of data only once.

A distributed lock needs three properties:

  • Mutual exclusion: only one client holds it at any time.
  • No deadlock: the lock releases automatically after its holder crashes (via the expiry).
  • Identifiability: when unlocking, you can prove “this lock is mine” (via a unique value).

The pitfalls around the last two are covered below.

Locking: SET NX EX

The original approach: SETNX

SETNX (SET if Not eXists) is a natural fit for lock acquisition: it writes the key and returns 1 when the key doesn’t exist, and does nothing and returns 0 when it does. A return of 1 means you got the lock.

Since Redis 2.6.12 SETNX is marked deprecated; the docs recommend SET with the NX option instead. It’s still the right place to start the story.

The two-step trap

The beginner approach is SETNX to lock, then a separate EXPIRE for the expiry:

// Wrong: two separate commands; a crash in between deadlocks everything
ok, _ := rdb.SetNX(ctx, "lock:order", "token").Result() // set only, no expiry
if ok {
	rdb.Expire(ctx, "lock:order", 30*time.Second) // expiry set separately
	// if the process crashes between the two commands:
	// the key has no expiry and never releases; other instances can never acquire the lock
}

SETNX and EXPIRE are two independent commands. If the process crashes after SETNX but before EXPIRE, the key stays in Redis with no expiry. The no-deadlock property is gone.

The atomic way: SET NX EX

Since Redis 2.6.12, SET supports the NX and EX options, doing “write only if absent + set expiry” in one step. go-redis exposes this as SetNX; the third argument is the expiry:

// Lock: atomic, writes only if the key is absent, with a 30s expiry
ok, err := rdb.SetNX(ctx, "lock:order", "token-1", 30*time.Second).Result()

The underlying command is SET lock:order token-1 NX EX 30. ok == true means the lock is yours; false means someone else holds it.

Verified behavior (Redis 7.4):

Client 1 lock: ok=true
Client 2 lock: ok=false (already held)
Lock TTL: 30s

The value must be unique

The value can’t be a constant. Each lock contender uses its own unique identifier (random string, UUID, process ID + goroutine ID) so unlocking can prove “this lock is mine”. Why, see the unlock section.

Unlocking: a Lua script with an ownership check

The direct-DEL trap

Releasing with a bare DEL is the most common mistake:

// Wrong: may delete someone else's lock
rdb.Del(ctx, "lock:order")

Scenario: client A acquires the lock, its work takes longer than 30 seconds, and the lock expires and gets deleted by Redis. Client B acquires the lock. A finishes and its DEL removes B’s lock. B is still in its critical section, C acquires the lock — mutual exclusion is broken and two clients run the critical section at once.

Verified:

Client 2 bare DEL: 1 (deleted client 1's lock)

Check the value before deleting

Check the value is yours before deleting. But “GET-compare + DEL” is two commands; the lock can change hands in between, reopening the window. The right way is to put both steps in a Lua script, which Redis executes atomically:

if redis.call("get", KEYS[1]) == ARGV[1] then
    return redis.call("del", KEYS[1])
else
    return 0
end

Calling it from go-redis:

// Unlock script: delete only if the value matches, return 1; otherwise return 0
const unlockScript = `
if redis.call("get", KEYS[1]) == ARGV[1] then
    return redis.call("del", KEYS[1])
else
    return 0
end`

// Unlock with someone else's value: returns 0, lock stays
n, _ := rdb.Eval(ctx, unlockScript, []string{"lock:order"}, "token-2").Int()
// Unlock with your own value: returns 1, lock deleted
n, _ = rdb.Eval(ctx, unlockScript, []string{"lock:order"}, "token-1").Int()

Verified behavior:

Wrong value: 0 (lock still there)
Right value: 1 (lock deleted)

This is why the value must be unique: the script uses it to tell “is this lock mine?”. With a hardcoded value, any client could unlock someone else’s lock with the same value.

Expiry and the watchdog

How long should the expiry be

The lock must have an expiry, otherwise a crashed holder leaves it forever. But the expiry itself cuts both ways:

  • Too short: the lock expires before the work is done, another client acquires it, and two clients are in the critical section at once.
  • Too long: after the holder crashes, other clients wait a long time to acquire the lock.

There’s no perfect fixed value. When execution time varies a lot, a fixed expiry forces a choice between “expires too early” and “wait too long”.

The watchdog: automatic renewal

The fix is to let the business decide how long the lock lives: while the holder is alive it keeps renewing; when it dies nobody renews and the lock expires naturally. This is the watchdog mechanism.

Redisson (the most common Java Redis client) has a built-in watchdog. From its source (RedissonLock / RenewalTask):

  • Locking without a leaseTime uses a default expiry of 30 seconds (lockWatchdogTimeout = 30 * 1000).
  • After a successful lock, a background task renews every lockWatchdogTimeout / 3 = 10 seconds, resetting the expiry back to 30 seconds.
  • The renewal script first checks the lock is still held by you (hexists); if ownership changed, renewal stops.
  • On unlock or holder crash the task stops, and the lock expires within at most 30 seconds.

Result: the lock lives as long as the work, and a crashed holder’s lock still releases within 30 seconds.

The Go ecosystem today

Go has no Redisson-style all-in-one. go-redis only provides primitives, so renewal is on you: after acquiring the lock, start a goroutine that periodically runs a Lua renewal script (check the value matches, then PEXPIRE), and stop it on unlock.

redsync (Go’s Redlock implementation) has no automatic renewal by default: the expiry defaults to 8 seconds, and locks expire once work exceeds that. Check whether you need renewal before using it; if you do, implement it yourself.

Lock loss on failover, and Redlock

The single-instance risk

Everything above assumes a single Redis instance. With master-replica replication there’s a classic lock-loss scenario:

  1. Client A acquires the lock on the master.
  2. The master crashes before the write is replicated to the replica.
  3. The replica is promoted to master, without the lock key.
  4. Client B acquires the same lock on the new master.

A and B both believe they hold the lock; mutual exclusion is broken. Redis replication is asynchronous by default, and this window is real.

The official docs (the Distributed locks page) call this a SAFETY VIOLATION outright. The root cause is asynchronous replication: the new master’s state is a lagging snapshot of the old master from some earlier moment. Mutual exclusion is fundamentally “all participants agreeing on who holds the lock”; async replication can’t guarantee that agreement, and consistency requires a consensus protocol (Raft/Paxos), which Redis doesn’t provide. The lock-loss window of “master-replica + SET NX EX” is structural; no config change removes it.

The Redlock algorithm

Redlock’s idea: don’t rely on a single node. Try to acquire the lock on N independent Redis instances (the docs suggest 5) and consider it acquired only if more than half (N/2 + 1) succeed. One node going down doesn’t matter; a minority losing the lock doesn’t break mutual exclusion.

Procedure:

  1. Record the current time, send SET key value NX PX ttl to all nodes.
  2. Measure the total time the acquisition took; the lock is acquired only if a majority succeeded and the total time is less than the lock’s expiry.
  3. Once acquired, the lock’s effective lifetime = expiry - acquisition time.
  4. On failure (not enough nodes or over time), send the unlock script to all nodes to clean up.

The debate: is Redlock safe?

Redlock has been controversial since it appeared. Two articles to read:

  • Martin Kleppmann, “How to do distributed locking”: his core argument is that a distributed lock’s expiry can’t prevent the “holder paused by GC, lock expired, resumes and keeps writing” scenario; the lock service must hand out monotonically increasing fencing tokens and the storage side must reject writes with stale tokens. Redlock has no such mechanism and relies on clock assumptions, so it isn’t safe.
  • antirez (Redis’ author), “Is Redlock safe?”: a point-by-point rebuttal. Fencing tokens require the storage side to do linearizable checks — at which point you might as well use the token scheme directly. Redlock only needs node clocks to run at roughly comparable rates (e.g. counting 5 seconds within a 10% error), not absolute time agreement; monotonic clocks are enough.

The debate is really about different system-model assumptions, with no consensus. For an interview, being able to state both sides is enough.

Pragmatic choices

  • The lock is only for efficiency (avoiding duplicate computation or notifications), and losing it just costs a bit more work: a single Redis instance with SET NX EX is the first choice — simple, fast, already deployed; Redlock’s complexity isn’t worth it. The lock’s “mutual exclusion with an expiry date” doesn’t matter here.
  • The lock is for correctness (inventory deduction, file writes): don’t rely on the lock as the backstop. Let the resource side defend concurrency itself (database unique constraints, version CAS, idempotency keys), or switch to a strongly consistent lock like etcd/ZooKeeper; the lock is just the first gate.
  • With master-replica replication, even without Redlock, know that the loss window exists; don’t treat the lock’s mutual exclusion as an absolute guarantee.

Choose by the cost of losing the lock: duplicated work → a Redis lock is fine; data corruption → add resource-side protection or move to a strongly consistent lock.

FAQ

How long should the lock expiry be?

There’s no standard answer. If execution time is stable, set it slightly above the worst case; if it varies a lot, use a renewal mechanism instead of betting on a fixed value.

Why unlock with a Lua script?

Splitting “compare value + delete” into two commands leaves a window: after the compare and before the delete, the lock can expire and change hands, and you delete someone else’s lock. A Lua script runs atomically in Redis; no other command can slip between the compare and the delete.

What happens when watchdog renewal fails?

Renewal failure means Redis is unreachable or the lock changed hands. The lock releases itself after the expiry; the business code has to cope with “the lock may be gone” — this is the difference between distributed and local locks: a local lock’s mutual exclusion is absolute while held, a distributed lock’s has an expiry date.

Edit this page

Contents