Gray failure disk await
Customers complain a specific availability zone 'feels slow' and some uploads time out, but every dashboard is green: the service's health checks pass, the load balancer marks all targets healthy, and synthetic monitors from your monitoring VPC show normal latency. Dashboards: error rate is 0.3% (within SLO), but support tickets are clustered to clients in one AZ; one storage node in that AZ shows disk await climbing slowly over 6 hours; its health-check endpoint (a simple /healthz that returns 200) is fine. Nothing was deployed. How do you triage and mitigate?
What a strong answer looks like
Stop the bleeding first (mitigate), then form hypotheses from real signals. Separate root cause from symptom, communicate status as you go, and close with what prevents a repeat.
0:00 of about 40 min
Which questions mattered is sealed until you submit. Telling you now would just be handing over the edge cases.
Run or narrate your approach, then ask the coach.