Errors a client can actually act on

An error response is a set of instructions for a program that cannot ask you a follow-up question.

The idea

When your API says no, something on the other end has to decide what to do next — usually a piece of code at 3am, not a person. That code can only ask three things: is this my fault, should I try again, and what exactly do I change?

Every field in an error body exists to answer one of those three questions. If a response answers none of them, the client is left with two bad moves: hammer you with retries, or throw the work away. Both look like your outage.

The same bug, three error designs

One batch of four invoices. Invoice 3 has a negative amount. Watch what a simulated client can do with each response.

POST /v1/invoices/batch   idempotency-key: bk_77a
  [0] acme        amount_cents:   1200
  [1] borden      amount_cents:    850
  [2] cinder      amount_cents:  -1200   <- invalid
  [3] dovetail    amount_cents:    430
design a · opaque 500 idle
clientapi

waiting for the request.

requests sent
0
confirmed of 4
0
duplicate rows
0
design b · flat 400 idle
clientapi

waiting for the request.

requests sent
0
confirmed of 4
0
duplicate rows
0
design c · per-item results idle
clientapi

waiting for the request.

requests sent
0
confirmed of 4
0
duplicate rows
0

How it works

Design the body backwards from the client's decision tree. Three questions, three parts of the response — and each part is machine-readable before it is human-readable.

question              answered by                          in design c
whose fault?          status class (4xx caller / 5xx us)    422 -> mine
should i retry?       explicit "retryable" + Retry-After    false -> no
what do i fix?        stable "code" + "pointer" + detail    /items/2/amount_cents

The status code is the cheap signal every generic HTTP client already reads; the body is where the specifics live. Never let them disagree. A validation failure returned as 500 tells the client's retry layer "transient, back off and try again" while the truth is "this will fail identically forever".

Why the batch arithmetic matters

Here is what the simulation actually counted, with the server writing valid rows before it hits the bad one:

design a — opaque 500, client retries with backoff
  attempts            = 1 initial + 5 retries      = 6
  rows written        = 6 attempts x 3 valid items = 18
  intended rows       =                              3
  duplicate rows      = 18 - 3                     = 15
  confirmed to client =                              0

design b — flat 400, client correctly stops
  requests = 1   rows = 0   duplicates = 0   humans paged = 1

design c — per-item results
  call 1: 3 x 201, 1 x 422 (retryable:false)
  call 2: corrected item [2] -> 201
  requests = 2   rows = 4   duplicates = 0   humans paged = 0

Design a is worse than doing nothing: it silently manufactured fifteen duplicate invoices while reporting total failure. That gap between "what the client believes" and "what the database holds" is the real cost of a vague error.

Field-level detail that a client can apply

A pointer is only useful if it addresses the request the client actually sent. Use a JSON Pointer (RFC 6901) into the request body, keep the array index, and pair it with a stable code that you promise never to repurpose. Prose is for humans and may change freely; code is an API surface.

The two kinds of "not found"

You still need to debug it, so keep the distinction where it belongs: log reason=forbidden against the request_id you echoed in the body, and let support join the two.

When to use it

situationshapetrade-off
bulk endpoint, items are independent207 (or 200) with a per-item results arraycallers must read the body, not just the status — document it loudly
bulk endpoint that must be all-or-nothing422 + a list of errors with ; nothing writtenone bad row blocks a good batch; only choose it when the items are genuinely one transaction
single resource, bad input400 / 422 + code + pointer, retryable:falsecheap and obvious; nothing to trade
you are overloaded or throttling429/503 + Retry-After, retryable:truehonest beats a slow success every time
genuinely your bug500 + request_id, no detailkeep it opaque, but make it rare — 5xx is a promise that retrying is reasonable

Note that 207 Multi-Status comes from WebDAV; plenty of teams use a plain 200 with the same results array. Either is defensible — being consistent and documented is what matters.

Watch out for

Worked example

You're reviewing a payroll import endpoint: POST /v1/timesheets/batch, up to 500 rows. Today it validates rows in a loop, and on the first bad one it throws — the framework maps that to 500, after the earlier rows have already been committed. The customer's integration retries on 5xx, six times, with .

In review you'd name the three symptoms concretely: the status lies about whose fault it is, the retry is guaranteed to fail identically, and because the writes were not atomic each attempt re-created the same valid rows — six attempts over three good rows is fifteen duplicates, and the client still believes nothing was saved.

The change is small. Process each row independently and return 207 with one result per index: 201 plus the created id for successes, and for row 47 a 422 carrying code: "rate_out_of_range", pointer: "/items/47/hourly_rate", retryable: false. Keep the envelope-level 4xx for things that make the whole call unacceptable — malformed JSON, missing auth, more than 500 items — and reserve 500 for the cases where retrying really is the right advice. Add an Idempotency-Key on the batch so even a mistaken replay can't duplicate. Now the client's recovery is two lines: keep the successes, fix row 47, resend one item.

Check yourself

Your rate limiter trips halfway through a 200-item batch. What does the client most need in the response?

A tenant asks for an invoice that exists but belongs to a different tenant. What do you return?