An error response is a set of instructions for a program that cannot ask you a follow-up question.
When your API says no, something on the other end has to decide what to do next — usually a piece of code at 3am, not a person. That code can only ask three things: is this my fault, should I try again, and what exactly do I change?
Every field in an error body exists to answer one of those three questions. If a response answers none of them, the client is left with two bad moves: hammer you with retries, or throw the work away. Both look like your outage.
One batch of four invoices. Invoice 3 has a negative amount. Watch what a simulated client can do with each response.
POST /v1/invoices/batch idempotency-key: bk_77a
[0] acme amount_cents: 1200
[1] borden amount_cents: 850
[2] cinder amount_cents: -1200 <- invalid
[3] dovetail amount_cents: 430
500 Internal Server Error
{ "error": "Internal server error",
"request_id": "req_8f21" }
waiting for the request.
400 Bad Request
{ "title": "Invalid batch",
"detail": "One or more invoices are invalid." }
waiting for the request.
207 Multi-Status
{ "results": [
{ "index": 0, "status": 201, "id": "inv_a1f" },
{ "index": 1, "status": 201, "id": "inv_b2c" },
{ "index": 2, "status": 422, "error": {
"code": "amount_must_be_positive",
"pointer": "/items/2/amount_cents",
"detail": "must be > 0, got -1200",
"retryable": false } },
{ "index": 3, "status": 201, "id": "inv_d4e" } ],
"request_id": "req_3c07" }
waiting for the request.
Design the body backwards from the client's decision tree. Three questions, three parts of the response — and each part is machine-readable before it is human-readable.
question answered by in design c whose fault? status class (4xx caller / 5xx us) 422 -> mine should i retry? explicit "retryable" + Retry-After false -> no what do i fix? stable "code" + "pointer" + detail /items/2/amount_cents
The status code is the cheap signal every generic HTTP client already reads; the body is where the specifics live. Never let them disagree. A validation failure returned as 500 tells the client's retry layer "transient, back off and try again" while the truth is "this will fail identically forever".
Here is what the simulation actually counted, with the server writing valid rows before it hits the bad one:
design a — opaque 500, client retries with backoff attempts = 1 initial + 5 retries = 6 rows written = 6 attempts x 3 valid items = 18 intended rows = 3 duplicate rows = 18 - 3 = 15 confirmed to client = 0 design b — flat 400, client correctly stops requests = 1 rows = 0 duplicates = 0 humans paged = 1 design c — per-item results call 1: 3 x 201, 1 x 422 (retryable:false) call 2: corrected item [2] -> 201 requests = 2 rows = 4 duplicates = 0 humans paged = 0
Design a is worse than doing nothing: it silently manufactured fifteen duplicate invoices while reporting total failure. That gap between "what the client believes" and "what the database holds" is the real cost of a vague error.
A pointer is only useful if it addresses the request the client actually sent. Use a JSON Pointer (RFC 6901) into the request body, keep the array index, and pair it with a stable code that you promise never to repurpose. Prose is for humans and may change freely; code is an API surface.
unknown_endpoint.invoice_not_found.You still need to debug it, so keep the distinction where it belongs: log reason=forbidden against the request_id you echoed in the body, and let support join the two.
| situation | shape | trade-off |
|---|---|---|
| bulk endpoint, items are independent | 207 (or 200) with a per-item results array | callers must read the body, not just the status — document it loudly |
| bulk endpoint that must be all-or-nothing | 422 + a list of errors with ; nothing written | one bad row blocks a good batch; only choose it when the items are genuinely one transaction |
| single resource, bad input | 400 / 422 + code + pointer, retryable:false | cheap and obvious; nothing to trade |
| you are overloaded or throttling | 429/503 + Retry-After, retryable:true | honest beats a slow success every time |
| genuinely your bug | 500 + request_id, no detail | keep it opaque, but make it rare — 5xx is a promise that retrying is reasonable |
Note that 207 Multi-Status comes from WebDAV; plenty of teams use a plain 200 with the same results array. Either is defensible — being consistent and documented is what matters.
400 during your own outage stops clients retrying at the exact moment they should back off."field": "amount" when the body had items[2].amount_cents forces a human to guess the mapping. Point at the exact path the caller sent.429s are not the same thing: "slow down for 2 seconds" and "your monthly quota is gone until the 1st" need very different client behaviour. Ship an explicit boolean plus Retry-After.detail string, and your next copy edit will be their outage. Give them a stable code so the prose stays free.You're reviewing a payroll import endpoint: POST /v1/timesheets/batch, up to 500 rows. Today it validates rows in a loop, and on the first bad one it throws — the framework maps that to 500, after the earlier rows have already been committed. The customer's integration retries on 5xx, six times, with .
In review you'd name the three symptoms concretely: the status lies about whose fault it is, the retry is guaranteed to fail identically, and because the writes were not atomic each attempt re-created the same valid rows — six attempts over three good rows is fifteen duplicates, and the client still believes nothing was saved.
The change is small. Process each row independently and return 207 with one result per index: 201 plus the created id for successes, and for row 47 a 422 carrying code: "rate_out_of_range", pointer: "/items/47/hourly_rate", retryable: false. Keep the envelope-level 4xx for things that make the whole call unacceptable — malformed JSON, missing auth, more than 500 items — and reserve 500 for the cases where retrying really is the right advice. Add an Idempotency-Key on the batch so even a mistaken replay can't duplicate. Now the client's recovery is two lines: keep the successes, fix row 47, resend one item.
Your rate limiter trips halfway through a 200-item batch. What does the client most need in the response?
A tenant asks for an invoice that exists but belongs to a different tenant. What do you return?