Go context: Timeouts, Cancellation, and Why Goroutines Keep Running

Key takeaways

context only signals; it never stops code by force. This post walks through how deadlines nest, how cancellation reaches (or fails to reach) blocking calls, how r.Context() and http.Client.Timeout interact, and the leaks that come from forgetting all of that.

A context.Context answers two questions for a piece of work: how long may this run and has someone upstream already given up on it. The standard library threads it through net/http, database/sql, os/exec and net, so once a request context exists, most I/O can observe it.

The one thing to internalize before any API detail: context never stops code by force. It closes a channel (ctx.Done()) and sets an error (ctx.Err()). Whether anything actually stops depends on the code below you checking that channel. Nearly every “the timeout didn’t work” bug comes from forgetting this.

The three constructors, and what cancel() is for

ctx, cancel := context.WithCancel(parent)                 // cancel manually
ctx, cancel := context.WithTimeout(parent, 800*time.Millisecond) // now + d
ctx, cancel := context.WithDeadline(parent, t)            // absolute time

WithTimeout(parent, d) is literally WithDeadline(parent, time.Now().Add(d)), so the choice is only about which is easier to read. A per-call budget reads better as a timeout; “finish before the batch window closes at 02:00” reads better as a deadline.

All three return a cancel function, and you must call it even if the work finishes successfully. A derived context registers itself with its parent so the parent can cancel it later; with timeouts it also holds a timer. Calling cancel releases both. If you drop it, go vet catches the obvious cases:

ctx, _ = context.WithTimeout(ctx, time.Second)
the cancel function returned by context.WithTimeout should be called, not discarded, to avoid a context leak

defer cancel() right after the constructor is the default. The one place it bites is inside a long loop: defers only run when the function returns, so a loop that creates a timeout per iteration piles up live contexts. Move the loop body into its own function, or call cancel() explicitly at the end of each iteration.

How cancellation and deadlines propagate

Contexts form a tree. Cancelling a node cancels everything derived from it, never its parent or siblings. Deadlines follow the same rule: a child can only be cut shorter.

outer, c1 := context.WithTimeout(context.Background(), 100*time.Millisecond)
defer c1()
inner, c2 := context.WithTimeout(outer, 5*time.Second)
defer c2()

d1, _ := outer.Deadline()
d2, _ := inner.Deadline()
fmt.Println(d1.Equal(d2)) // true: the 5s request is ignored
<-inner.Done()
fmt.Println(inner.Err())  // context deadline exceeded

This is useful in layered services: the HTTP handler sets an overall budget, and each downstream call can set a tighter one without worrying that it accidentally grants more time than the caller has.

ctx.Err() tells you why it ended: context.Canceled after cancel() or a client disconnect, context.DeadlineExceeded when the clock ran out. Check with errors.Is, because libraries wrap it:

if errors.Is(err, context.DeadlineExceeded) {
    // our budget ran out: maybe retry, maybe 504
} else if errors.Is(err, context.Canceled) {
    // the caller left: log at debug level, don't page anyone
}

That distinction matters for alerting. A spike in Canceled usually means clients are hanging up (or a load balancer timeout is shorter than yours); a spike in DeadlineExceeded means your own dependencies are slow.

Making blocking code actually observe cancellation

If you own the blocking loop, select on ctx.Done() alongside the real work:

func poll(ctx context.Context, interval time.Duration) error {
    t := time.NewTicker(interval)
    defer t.Stop()
    for {
        select {
        case <-ctx.Done():
            return ctx.Err()
        case <-t.C:
            if err := checkOnce(ctx); err != nil {
                return err
            }
        }
    }
}

For library calls, use the Context variants: QueryContext, ExecContext, PingContext, exec.CommandContext, http.NewRequestWithContext, (*net.Dialer).DialContext. A db.Query without context will happily wait on a locked row until the database gives up, no matter what your handler’s deadline says.

The failure I have hit most often is not the blocking call itself but the channel send after it. A worker computes a result and does results <- v on an unbuffered channel, but the caller already returned because its context expired. Nobody will ever receive, so the worker goroutine sits there forever, holding whatever it allocated. It shows up weeks later as a slowly climbing goroutine count in pprof. The fix is either a buffered channel sized to the number of senders or a select that also watches ctx.Done() on the send.

Fan-out: stop the siblings when one fails

When a request needs several upstream calls in parallel, one failure should cancel the rest. WithCancelCause (Go 1.20+) lets you record which error triggered the cancel, which makes the siblings’ errors far more readable than a bare “context canceled”:

func fanOut(ctx context.Context) error {
    ctx, cancel := context.WithCancelCause(ctx)
    defer cancel(nil)

    errs := make(chan error, 2) // buffered: senders never block after we return
    go func() { errs <- call(ctx, "A", 500*time.Millisecond, false) }()
    go func() { errs <- call(ctx, "B", 50*time.Millisecond, true) }()

    var first error
    for i := 0; i < 2; i++ {
        if err := <-errs; err != nil && first == nil {
            first = err
            cancel(err) // siblings see context.Cause(ctx) == err
        }
    }
    return first
}

func call(ctx context.Context, name string, d time.Duration, fail bool) error {
    t := time.NewTimer(d)
    defer t.Stop()
    select {
    case <-t.C:
        if fail {
            return fmt.Errorf("%s: upstream failed", name)
        }
        return nil
    case <-ctx.Done():
        return fmt.Errorf("%s: %w", name, context.Cause(ctx))
    }
}

Running it, B fails at 50ms, and A returns immediately with A: B: upstream failed instead of waiting its full 500ms. In real code, golang.org/x/sync/errgroup with errgroup.WithContext does the same bookkeeping (cancel on first error, wait for all) and is what I reach for once there are more than two calls.

A note on time.After inside select: before Go 1.23 each call created a timer that was not garbage-collected until it fired, so a hot loop with time.After(time.Minute) could accumulate a lot of pending timers. Since 1.23 unreferenced timers are collected, but time.NewTimer plus Stop is still the clearer choice when the timer is per-iteration.

HTTP clients: request context vs Client.Timeout

There are two knobs on the client side, and they overlap:

client := &http.Client{Timeout: 10 * time.Second} // create once, reuse

func proxy(w http.ResponseWriter, r *http.Request) {
    ctx, cancel := context.WithTimeout(r.Context(), 2*time.Second)
    defer cancel()

    req, err := http.NewRequestWithContext(ctx, http.MethodGet, upstreamURL, nil)
    if err != nil {
        http.Error(w, err.Error(), http.StatusInternalServerError)
        return
    }
    resp, err := client.Do(req)
    if err != nil {
        if errors.Is(err, context.DeadlineExceeded) {
            http.Error(w, "upstream timeout", http.StatusGatewayTimeout)
            return
        }
        http.Error(w, err.Error(), http.StatusBadGateway)
        return
    }
    defer resp.Body.Close()
    w.Header().Set("Content-Type", resp.Header.Get("Content-Type"))
    _, _ = io.Copy(w, resp.Body)
}
  • Client.Timeout is a hard ceiling for the whole exchange, including reading the response body. It is a safety net for every request that client makes.
  • The request context carries the caller’s situation: this request’s budget, and cancellation when the incoming client disconnects.

Whichever fires first wins. Against a server that takes 2 seconds, a 100ms context produces Get "...": context deadline exceeded, and a 100ms Client.Timeout produces Get "...": context deadline exceeded (Client.Timeout exceeded while awaiting headers). On current Go versions both satisfy errors.Is(err, context.DeadlineExceeded), and both close the connection, so the upstream handler’s own r.Context() is cancelled too.

Two related details: build the http.Client once instead of per request (a new client per call is fine functionally with the default transport, but a new Transport per call throws away connection pooling), and do not rely on http.DefaultClient for outbound calls in a server, since it has no timeout at all.

HTTP servers: what r.Context() means

On the server, r.Context() is cancelled when the client’s connection closes, when an HTTP/2 stream is reset, or when ServeHTTP returns. That last one surprises people:

func handler(w http.ResponseWriter, r *http.Request) {
    go sendAuditEvent(r.Context(), r.URL.Path) // cancelled almost immediately
    w.WriteHeader(http.StatusAccepted)
}

As soon as the handler returns, the goroutine’s context is done and its outbound call fails with context canceled. For work that should outlive the response, Go 1.21 added context.WithoutCancel, which keeps the values (trace ID, user) but drops cancellation. Always give the detached work its own bound:

bg, cancel := context.WithTimeout(context.WithoutCancel(r.Context()), 5*time.Second)
go func() {
    defer cancel()
    sendAuditEvent(bg, r.URL.Path)
}()

http.Server also has ReadTimeout, ReadHeaderTimeout, WriteTimeout and IdleTimeout, which protect the server from slow clients at the connection level. They are complementary to per-request contexts, not a replacement. ReadHeaderTimeout in particular is worth setting on anything exposed to the internet, since the zero value means a client can trickle headers indefinitely.

Graceful shutdown uses a separate context: srv.Shutdown(ctx) stops accepting connections and waits for in-flight handlers until ctx expires. Request contexts are not cancelled by Shutdown itself; if you want handlers to stop long-running work on shutdown, set BaseContext on the server to a context you cancel yourself after Shutdown times out.

Context values: narrow on purpose

context.WithValue is for request-scoped data that crosses API boundaries without being part of every signature: trace IDs, the authenticated principal, a request-scoped logger. Use an unexported key type so packages cannot collide:

type ctxKey struct{}

func WithUser(ctx context.Context, u *User) context.Context {
    return context.WithValue(ctx, ctxKey{}, u)
}

func UserFrom(ctx context.Context) (*User, bool) {
    u, ok := ctx.Value(ctxKey{}).(*User)
    return u, ok
}

What goes wrong in practice is using values as a hidden parameter bag. A function that reads a page size or feature flag from ctx.Value has an invisible dependency; tests pass a bare context.Background() and get the zero value, and nothing tells them why. If a function needs something to do its job, make it an argument. Relatedly, do not store a context in a struct field for later use. Pass it as the first parameter of each call, so each operation gets the context of whoever invoked it.

Hooks that run on cancellation

context.AfterFunc(ctx, f) (Go 1.21) runs f in its own goroutine once ctx is done, and returns a stop function to unregister it. It is handy for bridging to APIs that have their own cancel mechanism, such as closing a connection or calling a vendor SDK’s Abort() when the request context ends, without dedicating a goroutine to wait on ctx.Done() yourself.

Diagnosing common symptoms

  • Timeout never fires: some call in the path does not take the context. Grab a goroutine dump (/debug/pprof/goroutine?debug=2) during the hang and look at what the stuck goroutine is blocked in.
  • Goroutine count grows over time: look for sends on channels nobody reads after an early return, and for defer cancel() inside loops.
  • Lots of context canceled in logs under normal traffic: clients or a proxy are giving up first. Compare your server’s budget with the load balancer’s idle and request timeouts.
  • Tests hang instead of failing: give tests a context with a timeout (t.Context() in Go 1.24+ is cancelled when the test ends, which pairs well with an explicit WithTimeout).