My Go API Returned 503 After 100 ms. The Handler Kept Running

—

от автора

I was looking at what happens after a Go HTTP handler times out. The client receives a 503 response, so the request appears finished from the outside. But the wrapped handler can still be running inside the process.

I wanted to separate two events that are easy to treat as one: sending a timeout response and stopping the work that produced it. In Go, the first can happen without the second. A context tells code that its result is no longer needed. It does not forcibly interrupt a goroutine that never checks the signal.

I wrote a small example with two handlers. They have the same HTTP timeout. One ignores request cancellation; the other waits either for permission to finish or for the request context to end. A channel holds the first handler open until the client has received its timeout response, so the central observation does not depend on an accurately timed sleep.

A small experiment with one request

This is the complete program. Save it as main.go and run it with go run main.go.

package mainimport (    "errors"    "fmt"    "io"    "net/http"    "net/http/httptest"    "sync/atomic"    "time")func run(label string, cooperate bool) error {    var active atomic.Int64    started := make(chan struct{})    finished := make(chan struct{})    release := make(chan struct{})    writeResult := make(chan error, 1)    handler := http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {        active.Add(1)        defer close(finished)        defer active.Add(-1)        close(started)        if cooperate {            select {            case <-release:                // The work was allowed to finish.            case <-r.Context().Done():                return            }        } else {            // Wait until main explicitly releases this handler.            <-release        }        _, err := w.Write([]byte("ok"))        writeResult <- err    })    server := httptest.NewServer(        http.TimeoutHandler(handler, 100*time.Millisecond, "timed out"),    )    defer server.Close()    type result struct {        status int        err    error    }    results := make(chan result, 1)    go func() {        response, err := http.Get(server.URL)        if err != nil {            results <- result{err: err}            return        }        _, _ = io.Copy(io.Discard, response.Body)        _ = response.Body.Close()        results <- result{status: response.StatusCode}    }()    <-started // Inspect only after the handler has started.    response := <-results    if response.err != nil {        close(release)        <-finished        return response.err    }    // The client has its 503. The first handler is still blocked.    if cooperate {        <-finished    }    fmt.Printf("%s: status=%d active-after-response=%d\n",        label, response.status, active.Load())    close(release)    <-finished    if !cooperate {        fmt.Printf("%s: late-write-rejected=%t\n",            label, errors.Is(<-writeResult, http.ErrHandlerTimeout))    }    return nil}func main() {    if err := run("ignores context", false); err != nil {        panic(err)    }    if err := run("checks context", true); err != nil {        panic(err)    }}

The important value is the state after the response. In the first case the client has received a 503, while active-after-response is still 1. The handler is blocked on release. The program then releases it, and its late attempt to write through the wrapper should return http.ErrHandlerTimeout.

In the second case the client still gets a 503. The program waits for the handler to return before reading the counter. It should read 0 because the handler observes r.Context().Done() and returns. Both cases are explicit about when the counter is inspected.

This is a state demonstration, not a latency measurement. The channel controls when the handler may finish, while the 100-millisecond timeout controls when the wrapper sends its response. There is one request per case.

What the timeout actually does

http.TimeoutHandler gives the wrapped handler a request context with a deadline. It runs the wrapped handler and waits for either completion or that deadline. If the deadline wins, the wrapper sends 503 to the client. The wrapped handler can still be executing. When it eventually attempts to write through the wrapper, that write returns http.ErrHandlerTimeout.

The key detail is that context cancellation is a signal, not a goroutine termination mechanism. A channel receive that does not also watch ctx.Done() cannot receive that signal. Neither can a CPU-heavy loop that never checks it, a call to a library that ignores context, or a background goroutine started with an unrelated context.

I think this distinction gets lost because the user-facing part behaves correctly. The browser stops waiting. The access log records a timeout. A dashboard shows a bounded response time. None of those observations says how long the server continued using resources for that request.

Consider a handler that starts slow work and times out after 100 milliseconds. The response is gone, but work that ignores cancellation may continue. Under repeated requests, those pieces overlap. That is how an API can return timeout responses quickly while the server continues accumulating work it no longer needs.

This example uses a blocked goroutine, so it does not imply high CPU usage. Replace the channel receive with a computation, a blocked database call, a remote request, or a lock wait and the resource picture changes. The common property is that the work outlives the response.

The easy fix is only the first boundary

The context-aware branch in the demo is deliberately simple:

select {case <-workFinished:    // Use the result.case <-r.Context().Done():    return}

In a real service, I would not stop at that select in the HTTP handler. I would pass the same request context through the service method and into every operation that is supposed to end with the request.

func (s *Service) LoadAccount(    ctx context.Context,    id int64,) (Account, error) {    return s.repository.LoadAccount(ctx, id)}func (r *Repository) LoadAccount(    ctx context.Context,    id int64,) (Account, error) {    var account Account    err := r.db.QueryRowContext(        ctx,        "SELECT id, name FROM accounts WHERE id = $1",        id,    ).Scan(&account.ID, &account.Name)    return account, err}

Those fragments show the boundary that matters: the database operation receives the request context instead of context.Background(). Whether an in-flight query stops promptly also depends on the database driver and server behavior. I would measure that separately, rather than assume that accepting a context argument proves prompt cancellation all the way down.

The same rule applies to outbound HTTP. Construct the downstream request with the incoming context, or with a shorter child deadline when that call needs its own budget. Otherwise the upstream request can time out while an outbound call keeps waiting.

There is a second trap here. Returning from the handler after selecting on ctx.Done() does not stop a child goroutine if that goroutine continues working independently. The goroutine itself must receive cancellation, finish, and release what it owns. A channel send to a handler that has already returned can also leave a worker blocked unless the send is designed with cancellation in mind.

Three time limits that answer different questions

I used to treat timeout settings as though they formed one general limit. They do not.

An HTTP client timeout limits how long that client is willing to wait. A server-side TimeoutHandler limits the response path and cancels the context it passes to the handler. A database or outbound-request timeout limits a particular dependency operation, provided that operation honors its context or its own timeout mechanism.

These limits should fit together, but setting one does not automatically enforce the others. If a client gives up after 100 milliseconds and the backend continues an 800-millisecond operation, the client timeout protected the client. It did not necessarily protect the backend. If a server sends a 503 after 100 milliseconds while its handler ignores cancellation, it bounded the response. It did not bound the handler’s work.

For a real endpoint I would inspect at least two timelines: when the client receives its final response, and when all work started on behalf of that request actually stops. A histogram of HTTP response times captures the first. An active-operation gauge, cancellation counter, or trace of the downstream operation can help reveal the second. I would also check what happens when a client disconnects before a response, because that uses the request context too.

What this example does not prove

The demo has no database, downstream network dependency, or production workload. It isolates the behavior of these two handlers under a controlled timeout. It does not tell me how many abandoned operations exist in any particular service, which database drivers stop promptly, or how much CPU a real timeout wastes.

It also does not mean every task should die when the caller leaves. Some tasks are intentionally durable: a payment already accepted, an audit event, or a job handed to a queue may need to complete independently. Those tasks need an explicit owner and lifecycle. Accidentally detaching them by replacing r.Context() with context.Background() inside a request path is a different design decision, often made without realizing it.

The distinction is still useful. A 503 after the deadline tells me when the client stopped waiting. It tells me nothing, by itself, about when the work stopped. I now want both moments whenever I evaluate a request timeout.

ссылка на оригинал статьи https://habr.com/ru/articles/1088260/