One key, and nobody dares rotate it

A credential is not a login. It is something you issue, scope, age and revoke, and whether anyone ever replaces it comes down to one question: can two of them be valid at once?

The idea

Pick the credential by asking who the caller is. An API key says this machine belongs to the customer, and it is the honest answer for a server the customer runs and controls. OAuth says this application is acting for a person who agreed to it, and it is the only honest answer when somebody else's product calls you on your customer's behalf. Offering both is fine. Leaving the reader of your documentation to guess which one their case is, is not.

Then treat whatever you issue as a thing with a life: issued, scoped, used, aged, replaced, revoked. Most credential incidents are not clever attacks. They are a key pasted somewhere public that could not be replaced quickly, because replacing it meant taking the customer's integration down. That is a design problem, not a discipline problem, and it has one fix.

One leak, three ways to live with it

Same integration, same key, same leak. The only thing that changes is how many keys are allowed to be valid at once.

design
response

With one key you have to pick one of these two. Neither is good.

scopes
expiry
day 8
A 21 day credential timeline for one integration, described by the caption below

Each bar is a key. Where two bars overlap in time, both keys are valid. The tinted band is the exposure window, from the leak to the moment revocation actually takes effect. The strip underneath is the integration's own traffic, at 40 requests a minute.

day 0.0

exposure0 h
requests rejected0
stageissued

if the leak lands on a live keyexposedrejected
one key, revoke on detection
one key, schedule the swap
two keys, confirm idle, then revoke

How it works

A key is a row, not a string. Design the row first and most of the hard questions answer themselves.

  1. Issue it to an owner that cannot resign. The owner is a service account or a team, never a person's login. Record who created it separately, for the audit trail. A key owned by a person is revoked by offboarding, at two in the morning, by someone who has never heard of the integration.
  2. Scope narrow by default. A read integration gets read. The token carries resource scopes, but the screen where a customer picks them has to speak their language: not orders:write but "create and change orders, including issuing refunds". A security reviewer approving a scope they cannot read is approving nothing.
  3. Store the hash, keep the signposts. The secret is hashed at rest and shown exactly once. What you keep in clear is a visible prefix and the last_four, so a value found in a public repository can be matched to an account in seconds without you ever having stored it.
  4. Give it an end, and say so early. An expires_at, a warning carried in the response for the whole run up to the deadline, a written grace window, and an error the client can tell apart from an ordinary rejection.
  5. Allow two. This is the one that changes behaviour. If an account may hold two valid keys, rotation stops being a cutover and becomes an ordinary deploy.
  6. Show last_used_at per key. It turns "I think nothing uses the old one" into something the customer can look at before they press revoke. Rotation runs on proof, not courage.
  7. Revoke on the read path, not on a cache refresh. The check hits the store the next request makes, so revocation lands on the next request rather than in up to five minutes.

The row

key_id        key_7Qd2fR             stable, safe to log, safe to show in a UI
prefix        prk_live_              visible, tells a scanner whose key it is
last_four     9c31                   enough to identify a leaked value
secret_hash   argon2id(...)          the only stored form of the secret
owner         svc-orders-sync        a service account, not a person
created_by    dana@acme.example      audit only, never the owner
scopes        ["orders:read"]        narrow by default
created_at    2026-03-02T09:14Z
expires_at    2026-03-23T09:14Z      null is a decision, not a default
last_used_at  2026-03-14T22:07Z      the field that makes rotation safe
revoked_at    null

The secret itself is returned once, at creation, and never again. "Support can read it back to you" and "the secret is hashed" cannot both be true.

The arithmetic behind the timeline

one integration, 40 requests per minute, key leaked on day 8

one key, revoke the moment you detect it
  leak to detection                           6 h
  exposure                                    6 h
  integration down until a new key ships      9 h  = 540 min
  rejected requests               540 x 40  = 21,600

one key, wait for the customer's change window
  leak to detection                           6 h
  wait for the window                        72 h
  exposure                           6 + 72  = 78 h
  the swap itself                             4 min
  rejected requests                 4 x 40  = 160

two keys may be valid at once
  leak to detection                           6 h
  key b issued at once, customer deploys      4 h
  last_used_at on key a quiet for             1 h   then revoke
  exposure                       6 + 4 + 1   = 11 h
  rejected requests                           0

The overlap does not win on exposure alone. Revoking on detection is faster: six hours against eleven. It wins because it is the only column where you are not paying for that speed with somebody else's production traffic, and because being cheap is what makes it get used again next month.

an expiry that arrives with only one key in service

  integration fails, someone notices, generates, deploys    90 min
  rejected requests                            90 x 40  = 3,600
  and again at the next expiry, and the one after that

with a second key allowed, the same expiry costs 0, because
key c is deployed before key b's deadline and key b is
revoked once its last_used_at has gone quiet

When to use it

The trade-off in one line: an API key is cheap to adopt and expensive to retire. OAuth is expensive to adopt and cheap to retire, one application at a time.

Watch out for

Worked example

You are asked to design credentials for a payments-adjacent API. Two callers: the customer's own order-sync server, and a partner analytics product that hundreds of your customers install. Say the split first, because it is the part most answers skip. The order-sync server acts as the customer, so it gets an API key scoped to orders:read. The partner acts for the customer, so it gets OAuth: the customer grants the partner's application a per-customer access token with the same narrow scope, and can withdraw that grant without touching their own key. One leaked partner secret then compromises one grant, not every customer who installed the product.

Now the incident, which is where the design is really tested. A customer's key turns up in a public repository. The and last four match in seconds, so detection is six hours, not days. Because the account may hold two valid keys, the order of operations is written down and boring: issue key b, tell the customer, watch last_used_at on key a stop moving once they deploy, then revoke key a. Eleven hours of exposure and no rejected requests. Under the one-key design the same incident is a choice between revoking immediately and taking their integration down for nine hours, or waiting three days for their change window with the leaked key live the whole time. The timeline above is that choice, in numbers.

When the interviewer asks what happens at the instant of revocation, be exact. The check is on the read path, so a request that has already been authorized finishes and returns normally, and the very next request gets 401 credential_revoked. Nothing that is mid-flight is killed, and nothing that is cached keeps a dead key alive. If you have a streaming endpoint, name it as the exception and say how it re-checks, because that connection passed the gate once and would otherwise outlive the revocation by hours.

Then say what you would change if the scope had been account:admin. Revoking the leaked value would not end the incident, because an admin key can mint more keys. You would also have to list every credential created during the exposure window and revoke the ones the customer does not recognise. That is the argument for narrow scopes, made concrete: it is not about least privilege as a slogan, it is about whether revocation is sufficient.

Check yourself

A partner builds a product that reads your customers' orders, and hundreds of your customers install it. What should the partner's calls carry?

A customer tells you they cannot rotate their key, because rotating means downtime. You get one change. Which one?