Code RoomUpstream accept queue full
MediumPrep Room Coding #3006

Upstream accept queue full

On-callNetworking & APIsMid–Senior~35 min

Under a traffic surge (a marketing email went out 10 minutes ago), your reverse proxy starts returning 502 Bad Gateway for ~6% of requests, climbing. The proxy's own CPU is moderate (~60%). Upstream app pods show no crashes and their own request latency is normal for the requests that succeed, but their accept queue / connection backlog is growing and some new connections to them are being refused or slow to accept. The proxy logs show 'upstream prematurely closed connection' and connect timeouts to upstreams. Autoscaling on the app tier hasn't added pods yet. Triage and mitigate.

What a strong answer looks like

Stop the bleeding first (mitigate), then form hypotheses from real signals. Separate root cause from symptom, communicate status as you go, and close with what prevents a repeat.

0:00 of about 35 min
Which questions mattered is sealed until you submit. Telling you now would just be handing over the edge cases.