Go Mutex

Contents

Interview notes on Go mutexes. Covers: basic usage, non-reentrancy and the consequences of misuse, starvation mode, TryLock, RWMutex, the memory model, common patterns and pitfalls.

Basic usage

var mu sync.Mutex // zero value is usable, no need for new

mu.Lock()   // lock; blocks if someone else holds it
// critical section: only one goroutine can be in here at a time
mu.Unlock() // unlock

Key properties:

  • The zero value is a usable, unlocked mutex — declare it and use it, no initialization needed.
  • Locking and unlocking don’t have to be done by the same goroutine: sync.Mutex is not associated with an owner. Goroutine A can lock it and goroutine B can unlock it. Many languages have “whoever locks must unlock” semantics; Go doesn’t.
  • Since the lock isn’t tied to an owner, there’s no way to tell “do I hold this lock” — hence no reentrancy. This is the root of the deadlock case below.
  • Keep the critical section short: the mutex protects reads and writes of shared data; don’t leave extra work inside the lock.

The correct pattern is to unlock with defer, guaranteeing release on every return path:

mu.Lock()
defer mu.Unlock() // unlocked when the function returns, including on panic

The cost of defer: when locking and unlocking a lot inside a loop, defer has a small overhead and you can switch to manual Unlock, but in most cases it’s not worth optimizing.

Non-reentrancy

Locking an already-held mutex again from the same goroutine blocks forever — deadlock:

var mu sync.Mutex
mu.Lock()
mu.Lock() // blocks forever: the mutex doesn't know the lock is held by itself

This is a design trade-off: tracking the owner requires extra state and a check (every Lock/Unlock would have to inspect the goroutine identity), and Go chose not to do it. The consequence: inside a critical section you can’t touch code that needs the same lock, including indirect calls — when calling a function while holding the lock, make sure that function doesn’t lock the same mutex again.

A similar deadlock exists with RWMutex recursive read locks; see the RWMutex section.

Misuse is a fatal error, not a panic

Operation Result
Unlock an unlocked Mutex fatal error: sync: unlock of unlocked mutex
RUnlock an unlocked RWMutex fatal error: sync: RUnlock of unlocked RWMutex
Call Unlock while only holding a read lock fatal error: sync: Unlock of unlocked RWMutex
Lock again while already holding it Blocks forever, deadlock (all goroutines are asleep - deadlock!)

fatal error is different from panic: a panic can be caught with recover, while a fatal error prints the stack and crashes the process — it can’t be caught. Lock pairing errors are unrecoverable at runtime, so you have to get the pairing right when writing the code.

Copying an in-use mutex is another class of error. The documentation states A Mutex must not be copied after first use. A copy carries the lock state along; the two copies each manage their own state and the lock effectively stops protecting anything. go vet’s copylocks check catches it:

mu2 := mu // go vet error: assignment copies lock value to mu2: sync.Mutex

If a struct embeds a mutex, always pass it by pointer, never by value — a value copy carries the lock state along.

Starvation mode

The mutex has two operating modes: normal mode and starvation mode, introduced in Go 1.9. The source code, in the state-machine comment in the sync package, has the full description.

In normal mode, blocked waiters queue in FIFO order. But when the lock is released, the woken waiter doesn’t get it directly — it races against newly arriving goroutines. The newcomer is already running on a CPU and has the advantage, so the waiter often loses and queues back at the front of the queue to keep waiting. A single goroutine can acquire the lock many times in a row even with a long queue behind it — throughput is good, but in extreme cases a waiter may never get the lock and starve.

So there’s a switch mechanism: after a waiter has failed to acquire the lock for more than 1ms of contention, the mutex switches to starvation mode, where the rules are reversed:

  • When the lock is released, it’s handed off directly to the front waiter; newcomers don’t compete.
  • A newly arriving goroutine sees the starving flag, doesn’t spin or try to acquire, and queues at the tail.
  • A waiter that gets the lock switches back to normal mode if it’s the last one in the queue, or if its wait was under 1ms.

Plainly: normal mode is “queue up, but latecomers can cut in line”; starvation mode is “strict first-come, first-served”. What it prevents is tail latency: in extreme cases a waiter goes a long time without the lock and individual request latency spikes.

About spinning: in normal mode, when the lock is held and the machine is multi-core, a new goroutine spins for a few iterations (polling the lock state) instead of parking — for short critical sections, spinning is much cheaper than parking and waking a thread. Starvation mode forbids spinning, because the lock is destined for the front waiter.

Implementation

The source lives in the sync package (since Go 1.26 the implementation moved to src/internal/sync/mutex.go). Mutex has only two fields:

type Mutex struct {
	state int32 // bitfield; one field stores four pieces of information
	sema  uint32 // semaphore for the waiter queue
}

The bits of state:

  • Bit 0, locked: whether the lock is held.
  • Bit 1, woken: whether a goroutine is already on its way to acquire (spinning or woken).
  • Bit 2, starving: whether the mutex is in starvation mode.
  • The remaining upper bits: the waiter count.

What is CAS?

Question: what is CAS, and why is it used here?

Answer: CAS (compare-and-swap) is an atomic instruction. It does three things:

  1. Reads the current value in memory.
  2. Compares it with the expected value.
  3. If equal, writes the new value; if not, does nothing.

The whole “compare + swap” is indivisible — no other goroutine can interleave. In Go it maps to the sync/atomic functions:

// If the value at addr equals old, replace it with new and return true
// Otherwise do nothing and return false
atomic.CompareAndSwapInt32(addr, old, new)

The mutex uses it for three reasons:

  • Multiple goroutines modifying state at the same time must be atomic. If two goroutines both read “unlocked” and both write “locked”, the lock is broken.
  • state can’t be protected by another lock — that’s chicken-and-egg: the lock itself needs a lock-free atomic operation to bootstrap.
  • Locking is the hottest path. CAS completes in one CPU instruction, an order of magnitude faster than “take a lock, then check”.

So: CAS is a lock-free atomic update, and the mutex uses it to acquire in a single instruction when there’s no contention.

Lock fast path

With no contention it’s a single CAS: CAS(&state, 0, locked). Nobody holds the lock, and one atomic operation acquires it. In the uncontended case, the entire cost of locking is this one instruction.

Lock slow path

Under contention it goes to lockSlow, four steps:

  1. On a multi-core machine with a spare P (processor, the logical processor in GMP scheduling), spin up to 4 times (active_spin = 4). Spinning is cheaper than parking for short critical sections.
  2. If spinning still doesn’t acquire, increment the waiter count.
  3. Park itself via sema and sleep in the waiter queue (the next section covers sema). On Linux this is built on futex (the kernel’s sleep/wake primitive).
  4. After being woken, compete again. A waiter that loses queues back at the front of the queue (source comment: queue at the front of the queue) — having waited once, it keeps reinserting at the front when it loses again.

What sema does

Question: what does sema do?

Answer: sema is the anchor of the sleep/wake primitive. It’s just a uint32 and its value is meaningless (always 0); what matters is its memory address. The runtime uses the address to attach waiters of the same lock to the same queue; different locks have different addresses, so the queues are isolated.

Two key operations:

  • When Lock can’t acquire: runtime_SemacquireMutex(&m.sema, ...) — the current goroutine is attached to the waiter queue for that sema address and sleeps.
  • When Unlock needs to wake someone: runtime_Semrelease(&m.sema, ...) — wakes one goroutine from the queue.

Don’t treat it as a counting semaphore. The runtime source comment says it verbatim: don’t think of these as semaphores, think of them as a way to implement sleep and wakeup, with every sleep paired with a wakeup. Counting is the waiter field’s job in state; sema only handles queuing and waking. It’s the same kind of thing as futex — futex is also a kernel sleep/wake primitive, and the runtime’s sema is built on it on Linux.

Unlock

Fast path: AddInt32(&state, -locked) clears the locked bit. If nothing else needs handling, return.

The slow path looks at the waiters; three cases where it does NOT wake anyone:

  • No waiters.
  • The woken flag is already set.
  • The lock was re-acquired by someone else.

The woken bit exists for exactly this: avoiding wasted wakeups. Wake a waiter and it runs over only to lose to a newcomer — that wakeup was wasted. When a wake is warranted, decrement the waiter count, set the woken bit, and semrelease wakes the front of the queue.

How starvation mode is implemented

When Unlock runs with the starving bit set, it hands off directly:

  • It does not set the locked bit.
  • It yields the time slice (Gosched) so the front waiter runs immediately.
  • The woken waiter sets the locked bit itself.

The switch back to normal mode is decided when acquiring the lock: the waiter is the last one in the queue, or its wait was under 1ms. The source comment explains why the check lives here: starvation mode is so inefficient that once two goroutines switch into it they can lock-step indefinitely (lock-step: marching in step — A holds the lock while B waits, B holds it while A waits, alternating forever), so exit as soon as possible.

TryLock also reads the status bits: if the lock is held, or the mutex is in starvation mode, it returns false immediately without queuing.

TryLock

The non-blocking lock attempt added in Go 1.18:

if mu.TryLock() { // returns true on success and holds the lock
	defer mu.Unlock()
	// ...
} else {
	// lock is held; don't wait, go here
}

Semantics: success is equivalent to Lock; failure establishes no synchronization relationship (in the memory model, a failed TryLock can be treated as if nothing happened). RWMutex’s TryRLock is the same.

The official docs are cold water on TryLock: correct uses are rare, and using it is often a sign of a deeper problem. The typical anti-pattern is a “retry if I can’t get it” loop — exactly the breeding ground for starvation. Use it for one-shot decisions like “if I get the lock, do the work; otherwise take another path”, not as an optimistic lock.

What is an optimistic lock?

Question: what is an optimistic lock?

Answer: two schools. A pessimistic lock (a mutex is one): assume someone will contend — lock first, then do the work, release when done; nobody else touches the shared data the whole time. An optimistic lock is the opposite: assume no contention — do the work without locking, and at commit time verify “is the value I read still the current value”; if yes, write; if not, someone changed it in the meantime, retry or give up.

The most common optimistic lock in Go is CAS (compare-and-swap) from sync/atomic. Using a simulated debit as an example:

var balance int32 = 100

for {
	old := atomic.LoadInt32(&balance)   // read a snapshot
	next := old - 30                    // do the work (pure computation, no touching shared variables)
	if atomic.CompareAndSwapInt32(&balance, old, next) {
		// CAS succeeded: nobody changed it in between, commit complete
		return
	}
	// CAS failed: the value is no longer old, retry
}

CompareAndSwapInt32(&balance, old, next) atomically performs “compare + swap”: if balance still equals old, replace it with next and return true; otherwise do nothing and return false. With no conflicts the whole thing is lock-free and concurrent; on conflict, loop and retry or give up. Database optimistic locking is the same idea: add a version column, UPDATE ... SET amount = ?, version = version+1 WHERE id = ? AND version = ? — zero affected rows means conflict.

TryLock is not an optimistic lock. It’s still acquiring a lock: you get the lock before doing the work, and conflict happens before the work; with an optimistic lock the conflict happens after the work, and no lock is ever taken. Using TryLock as an optimistic lock (retry when you can’t get it) is using a locking mechanism to do a validation job — you get neither the mutual exclusion nor CAS’s atomic commit, and you just manufacture starvation.

So: with a low conflict rate, use an optimistic lock and save the locking overhead; with a high conflict rate, use a pessimistic lock — retrying costs more than the lock.

RWMutex

Reader-writer lock: multiple goroutines can hold the read lock at the same time, the write lock is exclusive, and reads and writes exclude each other.

var rw sync.RWMutex

rw.RLock() // read lock: concurrent
// read shared data
rw.RUnlock()

rw.Lock() // write lock: exclusive
// write shared data
rw.Unlock()

Two hard rules:

  • Writer-preference: when a writer is waiting, new RLock calls all block until the writer has acquired and released the lock. This is explicit behavior in the source comments, with the goal of guaranteeing the writer eventually gets the lock instead of starving under a steady stream of readers. Implementation-wise, when a writer is waiting, readerCount is flipped internally as a marker.
  • No recursive read locks: RLock inside RLock can deadlock if a writer slips in between — the second RLock queues behind the writer, and the writer waits for the first reader to release. Without a writer, nested read locks are fine (the read count just increments), but the moment a writer shows up it hangs. So don’t write nested RLock code; don’t bet on “there’s no writer this time”.

Also don’t try to upgrade a read lock to a write lock (calling Lock while holding RLock deadlocks); downgrading a write lock to a read lock is likewise not allowed.

When it pays off: RWMutex makes sense for read-heavy, write-light workloads with a large enough critical section. If the critical section just reads a few fields, RWMutex’s atomic counting overhead costs more than Mutex — use plain Mutex. The deciding standard is benchmarks; don’t assume a reader-writer lock is faster by default.

RWMutex’s internals are worth a mention: an embedded Mutex handles writer mutual exclusion, readerCount tracks the reader count, and two semaphores wait for readers to drain and for the writer to release. Readers just atomically increment; the writer waits for the count to hit zero — so the reader path is light and the writer path is heavy, which is why it suits read-heavy workloads.

Memory model

The basics of happens-before are in the memory model section of the go-channel notes; here I’ll only cover the mutex rules. First, one term: synchronized before.

synchronized before is a relationship established directly between synchronization operations. When a Lock observes a previous Unlock, that Unlock is said to be synchronized before that Lock. happens-before is its superset: program order plus synchronized before, then the transitive closure. Ordinary reads and writes are not synchronization operations, so they can’t participate in synchronized before — they connect to synchronization edges through program order. Interviews don’t require a strict distinction between the two terms: the lock rules are phrased in terms of synchronized before only because Lock/Unlock themselves are synchronization operations.

The official rules (the Locks section of go.dev/ref/mem):

  • For any sync.Mutex or sync.RWMutex variable l and n < m, the nth call to l.Unlock() is synchronized before the return from the mth call to l.Lock().

In plain language: every unlock establishes a synchronization edge for later locks. When you acquire the lock, you see all writes made before the previous unlock — regardless of who wrote them. This is what “the lock provides visibility, not just mutual exclusion” means: shared data read inside the critical section is never stale.

The official example, verified with real code (-race is silent, and it always prints hello, world):

var l sync.Mutex
var a string

f := func() {
	a = "hello, world"
	l.Unlock() // 1st Unlock
}

l.Lock()
go f()
l.Lock()      // 2nd Lock: sees f's write to a
fmt.Println(a) // always "hello, world"

The chain: the write to a precedes Unlock in f by program order, the 1st Unlock is synchronized before the return of the 2nd Lock, and the print comes after the 2nd Lock. So there’s no data race and the result is deterministic.

RWMutex has a companion rule: for each RLock call, there exists some n such that the nth Unlock is synchronized before the return of that RLock, and the RUnlock corresponding to that RLock is synchronized before the return of the (n+1)th Lock. The effect: a read lock sees every write completed before the most recent write lock was released, and a write lock sees every write completed during the read-lock period — both sides of the reader-writer lock are covered by synchronization edges.

A TryLock addendum: success equals Lock; failure is nothing — in the memory model, a failed TryLock doesn’t even guarantee the lock was unlocked at the time, and it may be considered to always return false. So after a failed TryLock, just give up; don’t infer anything about the lock state from the failure.

Last, sync.Once — internally it’s a mutex plus a flag, and the memory model has a dedicated rule for it: the completion of f in once.Do(f) is synchronized before the return of any once.Do(f) call. Everything f initializes is visible to all callers once Do returns — this is what makes singleton initialization correct.

Common patterns

Protecting shared data, the standard shape:

type Counter struct {
	mu sync.Mutex
	n  int
}

func (c *Counter) Inc() {
	c.mu.Lock()
	defer c.mu.Unlock()
	c.n++
}

Note that sync.Mutex here is an embedded field, and a value-receiver method copies the struct — so such methods must use a pointer receiver, otherwise every call copies a struct carrying lock state; vet warns and the behavior is wrong.

sync.Once for singleton initialization, replacing double-checked locking — which doesn’t hold in Go. go.dev/ref/mem explicitly gives the double-check example reading an empty string, because observing the done flag isn’t the same as observing the write to a:

var once sync.Once
var cfg *Config

func GetConfig() *Config {
	once.Do(func() { cfg = loadConfig() })
	return cfg
}

Deadlock prevention with multiple locks: when goroutines acquire several locks in different orders, circular waiting is possible. The rule is a global, consistent lock order — either always acquire from smallest to largest, or follow an agreed layering. Also: don’t call functions that might lock from inside a lock — indirect calls are a sneaky source of deadlock.

Choosing between channel and mutex (echoing the go-channel notes’ conclusion): use channels for transferring data ownership and cooperative scheduling; use mutexes for critical sections protecting shared data structures. go.dev’s words on the sync package: apart from Once and WaitGroup, the sync package types are mostly intended for low-level library use; for higher-level synchronization, prefer channels and communication.

Pitfall checklist

  • Forgetting to Unlock, or an early return path that skips it → fix all at once with defer.
  • defer Unlock in a huge loop: defer runs when the function returns; a loop running millions of times accumulates a large backlog of pending defers. Wrap the loop body in an anonymous function and lock/unlock inside it.
  • Copying a mutex: passing structs by value and reassigning slices/maps both copy; vet’s copylocks catches it.
  • Reentrancy: locking the same mutex again directly or indirectly from inside a critical section deadlocks. Go locks are not reentrant — a feature, not a bug.
  • Inconsistent lock ordering: circular waiting in multi-lock scenarios.
  • Lock scope too large: locking the whole function also locks unrelated slow operations, killing concurrency. The lock should only wrap reads and writes of shared data.

Interview questions

Is Go’s Mutex reentrant?

No. Mutex doesn’t record an owner and has no reentrancy counter; locking an already-held mutex again from the same goroutine blocks forever — deadlock. This is a design trade-off: tracking the owner requires extra state and checks. Don’t touch code that needs the same lock inside a critical section, including indirect calls.

What happens when you Unlock an unlocked Mutex?

fatal error: sync: unlock of unlocked mutex, and the program crashes immediately. It’s a fatal error, not a panic, so recover can’t catch it. Lock pairing errors are unrecoverable at runtime; get the pairing right when writing the code.

Why does starvation mode exist?

In normal mode waiters queue FIFO, but newly arriving goroutines can cut in and win the lock, so a waiter may go a long time without it. After waiting more than 1ms the mutex switches to starvation mode: the lock is handed directly to the front waiter and newcomers queue, guaranteeing first-come, first-served. It prevents tail latency: individual request latency spiking. A waiter that gets the lock switches back to normal mode if it’s the last one in the queue or waited under 1ms.

What is RWMutex’s writer-preference?

When a writer is waiting, new RLock calls all block until the writer has acquired and released the lock. This guarantees the writer doesn’t starve under a steady stream of readers. The side effect: no recursive read locks — RLock inside RLock deadlocks if a writer slips in between.

How is Mutex implemented?

Two fields: state (a bitfield: locked/woken/starving plus a waiter count) and sema (the anchor of the sleep/wake primitive). The Lock fast path is CAS(0 → locked), a single instruction; under contention it takes the slow path: spin up to 4 times on a multi-core machine with a P available, then increment the waiter count and park via sema (futex-based on Linux). Unlock is symmetric: clear the locked bit, wake one waiter when there are waiters, with the woken bit preventing duplicate wakeups. In starvation mode Unlock hands off directly to the front waiter and yields the time slice.

Does a failed TryLock have synchronization effects?

No. A successful TryLock is equivalent to Lock; a failure is nothing in the memory model — it doesn’t even guarantee the lock was unlocked at the time. Don’t use TryLock as an optimistic lock: it’s a locking mechanism, and the conflict happens before the work; an optimistic lock’s conflict happens at commit time.

When should you use Mutex, and when channel?

Use Mutex for critical sections protecting shared data structures; use channel for transferring data ownership and cooperative scheduling (task distribution, shutdown notifications). Mutex handles “don’t touch simultaneously”; channel handles “who gets what”.

Edit this page

Contents