Code RoomWarm pool depleted by deploy
HardPrep Room Coding #2843

Warm pool depleted by deploy

On-callReliability & on-callSenior–Staff~30 min

To absorb your nightly dinner-rush spike you keep a warm pool of 40 pre-booted, pre-warmed instances that the ASG promotes into service in seconds. It has worked for months. Tonight you shipped a routine deploy at 17:30, and then at 18:00 the dinner ramp hit and scale-up was slow again — 4-minute cold boots, 503s, the exact symptom the warm pool was meant to prevent. The warm pool shows 0 available instances. CPU and the app are otherwise healthy. What happened, how do you mitigate right now, and how do you prevent it?

What a strong answer looks like

Stop the bleeding first (mitigate), then form hypotheses from real signals. Separate root cause from symptom, communicate status as you go, and close with what prevents a repeat.

0:00 of about 30 min
Which questions mattered is sealed until you submit. Telling you now would just be handing over the edge cases.