Code RoomConnection pool saturation
HardPrep Room Coding #2572

Connection pool saturation

On-callReliability & on-callSenior–Staff~40 min

Your monolith talks to 6 microservices through a shared HTTP client connection pool (max 200). At 19:40 the whole app's p99 goes to 30s and error rate spikes across ALL endpoints, even ones unrelated to search. Dashboards: the recommendations service (one of the 6) has p99 of 28s after a deploy at 19:35; the shared pool shows 200/200 connections in use and a growing wait queue; thread pool on the monolith is saturated. Other 5 downstreams are healthy. How do you triage and mitigate?

What a strong answer looks like

Stop the bleeding first (mitigate), then form hypotheses from real signals. Separate root cause from symptom, communicate status as you go, and close with what prevents a repeat.

0:00 of about 40 min
Which questions mattered is sealed until you submit. Telling you now would just be handing over the edge cases.