One key, and nobody dares rotate it
A credential is not a login. It is something you issue, scope, age and revoke, and whether anyone ever replaces it comes down to one question: can two of them be valid at once?
The idea
Pick the credential by asking who the caller is. An API key says this machine belongs to the customer, and it is the honest answer for a server the customer runs and controls. OAuth says this application is acting for a person who agreed to it, and it is the only honest answer when somebody else's product calls you on your customer's behalf. Offering both is fine. Leaving the reader of your documentation to guess which one their case is, is not.
Then treat whatever you issue as a thing with a life: issued, scoped, used, aged, replaced, revoked. Most credential incidents are not clever attacks. They are a key pasted somewhere public that could not be replaced quickly, because replacing it meant taking the customer's integration down. That is a design problem, not a discipline problem, and it has one fix.
One leak, three ways to live with it
Same integration, same key, same leak. The only thing that changes is how many keys are allowed to be valid at once.
With one key you have to pick one of these two. Neither is good.
Each bar is a key. Where two bars overlap in time, both keys are valid. The tinted band is the exposure window, from the leak to the moment revocation actually takes effect. The strip underneath is the integration's own traffic, at 40 requests a minute.
How it works
A key is a row, not a string. Design the row first and most of the hard questions answer themselves.
- Issue it to an owner that cannot resign. The owner is a service account or a team, never a person's login. Record who created it separately, for the audit trail. A key owned by a person is revoked by offboarding, at two in the morning, by someone who has never heard of the integration.
- Scope narrow by default. A read integration gets read. The token carries resource scopes, but the screen where a customer picks them has to speak their language: not
orders:writebut "create and change orders, including issuing refunds". A security reviewer approving a scope they cannot read is approving nothing. - Store the hash, keep the signposts. The secret is hashed at rest and shown exactly once. What you keep in clear is a visible
prefixand thelast_four, so a value found in a public repository can be matched to an account in seconds without you ever having stored it. - Give it an end, and say so early. An
expires_at, a warning carried in the response for the whole run up to the deadline, a written grace window, and an error the client can tell apart from an ordinary rejection. - Allow two. This is the one that changes behaviour. If an account may hold two valid keys, rotation stops being a cutover and becomes an ordinary deploy.
- Show
last_used_atper key. It turns "I think nothing uses the old one" into something the customer can look at before they press revoke. Rotation runs on proof, not courage. - Revoke on the read path, not on a cache refresh. The check hits the store the next request makes, so revocation lands on the next request rather than in up to five minutes.
The row
key_id key_7Qd2fR stable, safe to log, safe to show in a UI
prefix prk_live_ visible, tells a scanner whose key it is
last_four 9c31 enough to identify a leaked value
secret_hash argon2id(...) the only stored form of the secret
owner svc-orders-sync a service account, not a person
created_by dana@acme.example audit only, never the owner
scopes ["orders:read"] narrow by default
created_at 2026-03-02T09:14Z
expires_at 2026-03-23T09:14Z null is a decision, not a default
last_used_at 2026-03-14T22:07Z the field that makes rotation safe
revoked_at null
The secret itself is returned once, at creation, and never again. "Support can read it back to you" and "the secret is hashed" cannot both be true.
The arithmetic behind the timeline
one integration, 40 requests per minute, key leaked on day 8
one key, revoke the moment you detect it
leak to detection 6 h
exposure 6 h
integration down until a new key ships 9 h = 540 min
rejected requests 540 x 40 = 21,600
one key, wait for the customer's change window
leak to detection 6 h
wait for the window 72 h
exposure 6 + 72 = 78 h
the swap itself 4 min
rejected requests 4 x 40 = 160
two keys may be valid at once
leak to detection 6 h
key b issued at once, customer deploys 4 h
last_used_at on key a quiet for 1 h then revoke
exposure 6 + 4 + 1 = 11 h
rejected requests 0
The overlap does not win on exposure alone. Revoking on detection is faster: six hours against eleven. It wins because it is the only column where you are not paying for that speed with somebody else's production traffic, and because being cheap is what makes it get used again next month.
an expiry that arrives with only one key in service
integration fails, someone notices, generates, deploys 90 min
rejected requests 90 x 40 = 3,600
and again at the next expiry, and the one after that
with a second key allowed, the same expiry costs 0, because
key c is deployed before key b's deadline and key b is
revoked once its last_used_at has gone quiet
When to use it
- A server your customer runs, calling you as themselves long-lived API key, scoped, rotatable The cheapest thing for a customer to adopt: paste a string into a config. The cost is that the secret then sits at rest in their environment indefinitely, so the whole lifecycle above becomes yours to build and theirs to operate.
- Somebody else's product, acting for your customer OAuth authorization code with PKCE The only design where the customer can see what they granted, and revoke that one application without revoking anything else. You take on consent screens, refresh tokens and per-user revocation in exchange.
- A machine-to-machine caller you want inside OAuth anyway client credentials grant, short-lived token An API key with more ceremony. You get expiry for free because the token is minted per hour, and you pay a token endpoint round trip plus a caching story on the client.
- One of your own services calling another mTLS or a workload identity token Minutes, not months, and no secret to leak because the identity comes from the platform. No self-service story, so it is useless for customers.
- CI jobs, scripts, one-off backfills a token minted per run Nothing survives the job, so nothing can leak from it later. Needs an identity provider the runner can prove itself to, which small customers will not have.
- Your own web or mobile client a session, not an API key A key shipped in a mobile binary is a public string with a long life. This row is here as a mistake to avoid, not a trade-off to weigh.
The trade-off in one line: an API key is cheap to adopt and expensive to retire. OAuth is expensive to adopt and cheap to retire, one application at a time.
Watch out for
- A key owned by a person. It gets created by whoever was on the ticket, and the audit trail records them as the owner. Eighteen months later they leave, offboarding revokes everything in their name, and a nightly settlement job stops at 02:00 with nobody on the incident who knows what
key_7Qd2fRwas for. Own keys with a service account or a team, and keep the human only ascreated_by. - One key per account. Everything else on this page follows from this. If the account may hold exactly one valid key, then rotating it is a cutover with downtime on the customer's side, so it is scheduled, so it is deferred, so it never happens. A key with no expiry and no second slot is a key that will be five years old on the day it leaks.
- Storing the secret so support can read it back. Once the value is recoverable, your database is now the interesting target, and any admin who can view an account can act as it. Hash at rest, show it once at creation, and keep
prefixpluslast_fourso both the support conversation and the leak scanner still work without the value. - Revocation that is , and no answer about requests in flight. If the auth check reads a cache with a five minute lifetime, then "revoke" means "some time in the next five minutes", and an incident call will ask you which. Be able to say it precisely: a request that has already passed the check completes, the next request is rejected, and a long-lived stream or websocket needs its own re-check because it passed the gate once and never comes back.
- An expiry the client only discovers as a 401. An email to the address on the account reaches the person who signed up, not the service that will break. Carry the warning in the response for the whole run up to the deadline, publish the grace window, and return a distinguishable code,
credential_expiredrather than a bare 401, so the client's retry logic can stop and page a human instead of hammering you with the same dead secret for an hour.
Worked example
You are asked to design credentials for a payments-adjacent API. Two callers: the customer's own order-sync server, and a partner analytics product that hundreds of your customers install. Say the split first, because it is the part most answers skip. The order-sync server acts as the customer, so it gets an API key scoped to orders:read. The partner acts for the customer, so it gets OAuth: the customer grants the partner's application a per-customer access token with the same narrow scope, and can withdraw that grant without touching their own key. One leaked partner secret then compromises one grant, not every customer who installed the product.
Now the incident, which is where the design is really tested. A customer's key turns up in a public repository. The and last four match in seconds, so detection is six hours, not days. Because the account may hold two valid keys, the order of operations is written down and boring: issue key b, tell the customer, watch last_used_at on key a stop moving once they deploy, then revoke key a. Eleven hours of exposure and no rejected requests. Under the one-key design the same incident is a choice between revoking immediately and taking their integration down for nine hours, or waiting three days for their change window with the leaked key live the whole time. The timeline above is that choice, in numbers.
When the interviewer asks what happens at the instant of revocation, be exact. The check is on the read path, so a request that has already been authorized finishes and returns normally, and the very next request gets 401 credential_revoked. Nothing that is mid-flight is killed, and nothing that is cached keeps a dead key alive. If you have a streaming endpoint, name it as the exception and say how it re-checks, because that connection passed the gate once and would otherwise outlive the revocation by hours.
Then say what you would change if the scope had been account:admin. Revoking the leaked value would not end the incident, because an admin key can mint more keys. You would also have to list every credential created during the exposure window and revoke the ones the customer does not recognise. That is the argument for narrow scopes, made concrete: it is not about least privilege as a slogan, it is about whether revocation is sufficient.
Check yourself
A partner builds a product that reads your customers' orders, and hundreds of your customers install it. What should the partner's calls carry?
A customer tells you they cannot rotate their key, because rotating means downtime. You get one change. Which one?