Code RoomDNS TTL cache split
MediumPrep Room Coding #3002

DNS TTL cache split

On-callNetworking & APIsMid–Senior~35 min

Twenty minutes after a planned DNS change repointing api.example.com to a new ingress IP, you get partial outage reports: roughly 40% of users can't reach the API while 60% are fine, and the split doesn't correlate with region. Your new ingress is healthy and serving the 60%. Some clients resolve the old IP (now decommissioned, connections refused) and some the new IP. The old record had a TTL of 3600s. A handful of large enterprise customers behind corporate resolvers are heavily represented in the failures. Triage and stabilize.

What a strong answer looks like

Stop the bleeding first (mitigate), then form hypotheses from real signals. Separate root cause from symptom, communicate status as you go, and close with what prevents a repeat.

0:00 of about 35 min
Which questions mattered is sealed until you submit. Telling you now would just be handing over the edge cases.