Two writers, one resource
Two people open the same record, both press save, and the second save quietly erases the first — unless the write is forced to say which version it was based on.
The idea
A read is a photograph, not a reservation. When client a and client b both read version 1 and then both write, the server has no way of knowing that b's write was composed before a's landed. It applies it anyway, and a's change disappears with a cheerful 200 OK.
The fix is small: the server hands out a version label with every read (an ETag), and the client hands it back with every write (If-Match). If the label has moved on, the server refuses with 412 instead of guessing. Nothing is locked — conflicts are detected, not prevented, and the caller decides what to do.
Watch the collision, then replay it
Same two clients, same timing. Only the precondition changes.
Both clients are about to read version 1 of incident 4821. Press play, or step through one message at a time.
How it works
- Give every representation a validator.
ETag: "v1"— a row version counter, a monotonic revision, or a strong hash of the canonical body. It must change on every state change and be identical across replicas. - Return it on read, and have the client keep it next to the copy it is editing.
- Require it on unsafe writes.
PUT,PATCHandDELETEcarryIf-Match: "v1". Missing header on a contended resource? Answer428 Precondition Requiredrather than writing blind. - Compare and write atomically, in one statement or one transaction — never read-then-check-then-write.
- Answer a mismatch with 412, having written nothing, and include the current
ETag(ideally the current representation) so the retry costs one round trip. - Resolve on the client: auto-merge disjoint fields, ask a human when the same field moved twice.
The exchange from the demo, in full:
GET /incidents/4821
200 OK
ETag: "v1"
{ "status": "investigating", "owner": "unassigned" }
# client a writes first, quoting what it read
PUT /incidents/4821 If-Match: "v1"
{ "status": "investigating", "owner": "dana" }
200 OK ETag: "v2"
# client b writes second, still holding "v1"
PUT /incidents/4821 If-Match: "v1"
412 Precondition Failed ETag: "v2" <- nothing was written
# b re-reads, merges intent, retries against the version it just saw
GET /incidents/4821 -> 200, ETag: "v2", owner = "dana"
PUT /incidents/4821 If-Match: "v2"
{ "status": "mitigated", "owner": "dana" }
200 OK ETag: "v3"
Underneath, the precondition and the write are the same statement. This is the part people get wrong:
UPDATE incidents
SET status = 'mitigated', owner = 'dana',
version = version + 1
WHERE id = 4821
AND version = 1; -- the version the client quoted in If-Match
rows affected = 1 -> 200 OK, ETag: "v2"
rows affected = 0 -> 412 Precondition Failed, no side effects
Same idea with no HTTP in sight: an expected_version field in a gRPC request or an event-store append, answered with FAILED_PRECONDITION / 409 Conflict. The header is a convention; the compare-and-set is the mechanism.
When to use it
- Human-speed edits, low to moderate , an HTTP surface ETag + If-Match → 412 Stateless and cheap: no lock to hold, leak or time out. The cost is that every caller must handle 412 and have a merge story.
- Non-HTTP surface, or you want the conflict visible in the version field / expected_version → 409 Identical semantics, less protocol ceremony, and easier to log and test. You own the error mapping and the docs.
- Fields are genuinely independent and the newest value is the truth last-writer-wins, written down Device telemetry, presence, cursor position, cache warmers. Honest only if you state it in the API contract — silent LWW is a bug wearing a policy costume.
- The change is a delta, not a replacement commutative op (increment, add-to-set) Two concurrent "+1"s both count and there is nothing to resolve. You give up "I saw exactly this state before acting".
- Hot row, or work that must never be done twice pessimistic lock / lease with Seat allocation, payment capture, ledger posting. Correct under real contention — but you now own , lock waits, and the operator who closed their laptop mid-edit.
- Long, overlapping editing sessions on rich text CRDT or operational transform Genuinely concurrent editing with no rejections. A large, permanent engineering commitment; do not reach here to avoid writing a merge dialog.
The trade-off in one line: optimistic concurrency does not prevent conflicts, it surfaces them. If nobody owns the resolution — a merge screen, a retry rule, a documented policy — you have moved the lost update from the server to the user.
Watch out for
- Thinking PATCH is the fix. Sending only the changed fields shrinks the — a's
ownerand b'sstatusstop colliding — but two PATCHes to the same field still lose one. Patch shape narrows the window; only a precondition closes it. - An ETag that is not stable, or is weak.
If-Matchuses strong comparison, soW/"v1"can never match and those writes 412 forever. a serialised body is just as risky: reorder a JSON key, include alast_seen_atthat ticks, or let one replica gzip differently, and the validator changes when the resource did not. A version column is boring and correct. - A 412 with no way forward. Returning bare "precondition failed" makes the client guess. Return the current
ETagand, where it is cheap, the current representation or a field-level diff, so the retry is one round trip. And a 412 must be truly side-effect free — check inside the write transaction, not before it. - Read-check-then-write in two statements.
SELECT version, compare in application code, thenUPDATEis the very same lost update one layer down: another writer slips between the two statements. UseUPDATE … WHERE id = ? AND version = ?and treat zero rows affected as the conflict. - Blind retry on 412. A generic retry wrapper that resends the same body — or a client that drops back to a PUT without
If-Matchafter a rejection — converts a detected conflict straight back into a lost update. Retry only after re-reading and re-applying the user's intent, and consider428so a client that forgets the header cannot clobber quietly.
Worked example
An incident tool. Support reports that assignments "randomly unassign themselves" during busy pages. You pull the access log for one such incident and find two PUT /incidents/4821 requests 310 ms apart from different sessions, both preceded by a GET that returned the same ETag: "v1". Neither request carried If-Match, and the API takes a full-document PUT — so the later request, composed from a copy that was already stale, wrote owner: "unassigned" back over dana. No error was raised anywhere, which is exactly why it took three weeks to report.
The fix has three parts, and saying all three is what makes the answer senior. Server: derive the ETag from an existing version column, require If-Match on PUT/PATCH/DELETE, answer 428 when it is absent and 412 (plus the current ETag and body) when it is stale, with the compare folded into the UPDATE … AND version = ?. Client: on 412, refetch, auto-merge fields the user did not touch, and show "this incident changed while you were editing — dana was assigned" for any field that moved twice. Rollout: preconditions cannot be flipped on for existing callers overnight, so log missing-If-Match writes first, then enforce per client version.
If the interviewer pushes on the hot path — say twenty responders hammering one incident — that is where you concede that optimistic concurrency degrades: every writer keeps losing the race. Then you narrow the resource (per-field endpoints), make the operation commutative, or take a short lease. Not because ETags are wrong, but because the retry cost has overtaken the lock cost.
Check yourself
A client sends PUT with If-Match: "v7", but the resource is now at "v9". What should the server answer?
Which of these is the most honest fit for a documented last-writer-wins policy?