Code RoomDependency latency spike
MediumPrep Room Coding #2905

Dependency latency spike

On-callML systemsCode quality & reviewMid–Senior~30 min

Your real-time inference service enriches each request by calling a third-party 'identity-risk' API synchronously to build one feature before scoring. At 12:40 your service's p99 latency triples and error rate hits 4%. Tracing shows the time and errors are all in the outbound call to the third-party API, which is returning slow responses and intermittent 503s; their status page later confirms a partial outage. Your own model and feature store are healthy. Triage and mitigate.

What a strong answer looks like

Stop the bleeding first (mitigate), then form hypotheses from real signals. Separate root cause from symptom, communicate status as you go, and close with what prevents a repeat.

0:00 of about 30 min
Which questions mattered is sealed until you submit. Telling you now would just be handing over the edge cases.