Two writers, one resource

Two people open the same record, both press save, and the second save quietly erases the first — unless the write is forced to say which version it was based on.

The idea

A read is a photograph, not a reservation. When client a and client b both read version 1 and then both write, the server has no way of knowing that b's write was composed before a's landed. It applies it anyway, and a's change disappears with a cheerful 200 OK.

The fix is small: the server hands out a version label with every read (an ETag), and the client hands it back with every write (If-Match). If the label has moved on, the server refuses with 412 instead of guessing. Nothing is locked — conflicts are detected, not prevented, and the caller decides what to do.

Watch the collision, then replay it

Same two clients, same timing. Only the precondition changes.

client aagent's browser tab no copy
api/incidents/4821 ETag "v1"
statusinvestigating
ownerunassigned
client bon-call mobile app no copy
step 0 / 10

Both clients are about to read version 1 of incident 4821. Press play, or step through one message at a time.

server versionv1
writes accepted0
412 conflicts0
lost updates0

How it works

  1. Give every representation a validator. ETag: "v1" — a row version counter, a monotonic revision, or a strong hash of the canonical body. It must change on every state change and be identical across replicas.
  2. Return it on read, and have the client keep it next to the copy it is editing.
  3. Require it on unsafe writes. PUT, PATCH and DELETE carry If-Match: "v1". Missing header on a contended resource? Answer 428 Precondition Required rather than writing blind.
  4. Compare and write atomically, in one statement or one transaction — never read-then-check-then-write.
  5. Answer a mismatch with 412, having written nothing, and include the current ETag (ideally the current representation) so the retry costs one round trip.
  6. Resolve on the client: auto-merge disjoint fields, ask a human when the same field moved twice.

The exchange from the demo, in full:

GET /incidents/4821
200 OK
ETag: "v1"
{ "status": "investigating", "owner": "unassigned" }

# client a writes first, quoting what it read
PUT /incidents/4821          If-Match: "v1"
{ "status": "investigating", "owner": "dana" }
200 OK   ETag: "v2"

# client b writes second, still holding "v1"
PUT /incidents/4821          If-Match: "v1"
412 Precondition Failed      ETag: "v2"     <- nothing was written

# b re-reads, merges intent, retries against the version it just saw
GET /incidents/4821          -> 200, ETag: "v2", owner = "dana"
PUT /incidents/4821          If-Match: "v2"
{ "status": "mitigated", "owner": "dana" }
200 OK   ETag: "v3"

Underneath, the precondition and the write are the same statement. This is the part people get wrong:

UPDATE incidents
   SET status = 'mitigated', owner = 'dana',
       version = version + 1
 WHERE id = 4821
   AND version = 1;      -- the version the client quoted in If-Match

rows affected = 1  ->  200 OK,  ETag: "v2"
rows affected = 0  ->  412 Precondition Failed, no side effects

Same idea with no HTTP in sight: an expected_version field in a gRPC request or an event-store append, answered with FAILED_PRECONDITION / 409 Conflict. The header is a convention; the compare-and-set is the mechanism.

When to use it

The trade-off in one line: optimistic concurrency does not prevent conflicts, it surfaces them. If nobody owns the resolution — a merge screen, a retry rule, a documented policy — you have moved the lost update from the server to the user.

Watch out for

Worked example

An incident tool. Support reports that assignments "randomly unassign themselves" during busy pages. You pull the access log for one such incident and find two PUT /incidents/4821 requests 310 ms apart from different sessions, both preceded by a GET that returned the same ETag: "v1". Neither request carried If-Match, and the API takes a full-document PUT — so the later request, composed from a copy that was already stale, wrote owner: "unassigned" back over dana. No error was raised anywhere, which is exactly why it took three weeks to report.

The fix has three parts, and saying all three is what makes the answer senior. Server: derive the ETag from an existing version column, require If-Match on PUT/PATCH/DELETE, answer 428 when it is absent and 412 (plus the current ETag and body) when it is stale, with the compare folded into the UPDATE … AND version = ?. Client: on 412, refetch, auto-merge fields the user did not touch, and show "this incident changed while you were editing — dana was assigned" for any field that moved twice. Rollout: preconditions cannot be flipped on for existing callers overnight, so log missing-If-Match writes first, then enforce per client version.

If the interviewer pushes on the hot path — say twenty responders hammering one incident — that is where you concede that optimistic concurrency degrades: every writer keeps losing the race. Then you narrow the resource (per-field endpoints), make the operation commutative, or take a short lease. Not because ETags are wrong, but because the retry cost has overtaken the lock cost.

Check yourself

A client sends PUT with If-Match: "v7", but the resource is now at "v9". What should the server answer?

Which of these is the most honest fit for a documented last-writer-wins policy?