Go Goroutine

Contents

Goroutines are the core of Go concurrency interviews, but this series already covers most of the ground in separate posts: the scheduler in Go GMP Model, channel semantics and the memory model in Go channel, locks in Go Mutex, context in Go Context, and Go sync.Once. This post fills in what the series doesn’t cover in detail: goroutines vs threads, panics across goroutines, race detection, WaitGroup, and leak diagnosis, ending with a topic map.

1. Goroutines vs threads

A goroutine is a lightweight coroutine managed by the Go runtime. It runs on top of OS threads, with many goroutines multiplexed onto a few threads. The official FAQ puts it this way: multiplex independently executing functions onto a set of threads; when a coroutine blocks, the runtime automatically moves other coroutines on the same OS thread to a different, runnable thread.

The differences from threads come down to three things:

  • Stack size. A goroutine starts with a 2KB stack (since Go 1.4, stackMin), allocated on the heap and grown automatically when needed; a thread stack is measured in megabytes (8MB by default on Linux), and creating a thread reserves a large chunk of virtual memory.
  • Switch cost. Goroutine switches happen in user space without entering the kernel; thread switches go through the kernel scheduler, with context save/restore and user/kernel mode transitions.
  • Scale. The official FAQ says creating hundreds of thousands of goroutines in one address space is practical; threads would exhaust system resources at a much smaller number.

So “goroutines are lightweight” comes down to: small stacks, user-space switching, and the runtime managing memory and scheduling itself.

Don’t read “lightweight” as “free”. Goroutines have scheduler overhead, and more goroutines means more stacks for the GC to scan; higher concurrency is not always better, which is why worker pools are a common pattern.

2. Panic and recover

Two rules, and a lot of people get them wrong:

  • recover only works inside a defer function; calling it directly returns nil.
  • A panic in one goroutine can only be recovered by that goroutine’s own recover. If a goroutine panics without a recover, the whole process crashes; a recover in another goroutine cannot save it.

The classic mistake: recover in main, panic in a child goroutine, expecting it to be caught. It crashes. The correct pattern is a recover in the goroutine’s own defer:

go func() {
    defer func() {
        if r := recover(); r != nil {
            log.Printf("goroutine panic: %v", r)
        }
    }()
    // do work
}()

For library functions that spawn goroutines, the standard practice is to start the goroutine inside the function with its own recover, converting the panic into an error return so the caller’s process doesn’t get taken down.

Also worth distinguishing panic from fatal error: panics can be recovered, while fatal errors like sync: unlock of unlocked mutex crash the runtime directly and cannot be recovered (see Go Mutex).

3. Data races and race detection

Multiple goroutines reading and writing the same variable at the same time is a data race. The result is reading stale or torn values, or even a crash. Go’s detection tool is the race detector:

go build -race ./...
go test -race ./...

A binary built with -race reports and exits when it detects a race at runtime. Mentioning “I run the race detector” in an interview is a plus.

Common race triggers: goroutine closures capturing outer variables (loop variables), concurrent reads/writes on a shared map, and reading/writing shared fields without a lock. Prevention is locks and atomics; for selection guidance see Go Mutex. For a single counter, sync/atomic is lighter than a lock:

var counter atomic.Int64
counter.Add(1)

4. WaitGroup

WaitGroup waits for a group of goroutines to finish. There is really one pitfall: Add must be called before the goroutines start, and from the main goroutine — don’t call Add inside a child goroutine, or Wait can return before the counter is incremented. Done should be deferred:

var wg sync.WaitGroup
for _, task := range tasks {
    wg.Add(1)
    go func(t Task) {
        defer wg.Done()
        t.Run()
    }(task)
}
wg.Wait()

Note that the loop variable must be passed into the closure as a parameter; referencing the outer variable directly makes every goroutine read the last value (before Go 1.22).

Since Go 1.25 there is wg.Go, which combines Add, starting the goroutine, and Done into one call. Internally it’s Add(1) + go func() + deferred Done(), so the manual pattern above becomes:

var wg sync.WaitGroup
for _, task := range tasks {
    wg.Go(func() { task.Run() })
}
wg.Wait()

Dropping the manual Add/Done also avoids the counting-mismatch pitfall. One detail: Go has the same call-timing requirement as Add — it must be called before Wait (or after the previous Wait returns); the reuse rules haven’t changed.

5. Goroutine leaks and diagnosis

A goroutine that never reaches its exit condition sits there holding its stack and memory, and too many of them slow down the GC. Typical scenarios:

  • Sending on a channel with no receiver, blocking forever.
  • Receiving from a channel that nobody sends on or closes.
  • A select waiting on a case that never fires, without a default.
  • A time.Ticker that is never Stopped, leaking the timer.
  • An infinite loop with no exit condition.

How to diagnose:

  • Log runtime.NumGoroutine() and watch whether the count keeps climbing.
  • The pprof goroutine profile: import net/http/pprof, then go tool pprof http://localhost:6060/debug/pprof/goroutine. It shows every goroutine’s stack and count, and it’s obvious which channel or which line each one is stuck on.

To make a goroutine cancellable, the standard approach is a context or a done channel combined with select. A context’s cancellation propagates down the call chain, cancelling child contexts too; context.WithTimeout adds automatic timeout cancellation (see Go Context). The core principle: every goroutine must have an answer for how it exits — if it doesn’t, that’s a leak waiting to happen.

6. Topic map

Topic Where
Goroutines vs threads, why they’re lightweight This post
GMP scheduling, work stealing, handoff, preemption Go GMP Model
Channel close semantics, hchan, select, memory model Go channel
Panics across goroutines, recover rules This post
Data races, race detection, WaitGroup This post
Mutex/RWMutex internals, starvation mode Go Mutex
Context cancellation, timeout, values Go Context
sync.Once Go sync.Once
Goroutine leak scenarios and diagnosis This post
Edit this page

Contents