Autoscaler sees stale metrics
During an evening ramp your service saturates and p99 blows out, but the autoscaler never reacts — the fleet stays flat at its current size for 12 minutes. The HPA/scaler shows `desired = current` the whole time and its last-seen metric value is a low, stale number. Meanwhile your own monitoring graphs (a different pipeline) clearly show CPU and RPS climbing. There was no deploy to the service; an hour earlier, though, the metrics/observability platform had a brief incident. How do you triage and mitigate?
What a strong answer looks like
Stop the bleeding first (mitigate), then form hypotheses from real signals. Separate root cause from symptom, communicate status as you go, and close with what prevents a repeat.
0:00 of about 30 min
Which questions mattered is sealed until you submit. Telling you now would just be handing over the edge cases.
Run or narrate your approach, then ask the coach.