Upstream quota limit per worker
Your service calls a third-party payments provider whose API allows 100 requests/second across your whole account. You autoscaled your worker fleet to clear a backlog, and now a growing fraction of provider calls return HTTP 429 with `Retry-After`. The more workers you add, the higher the 429 rate climbs and the slower the backlog drains. Your own infra is healthy — CPU, memory, and queues on your side are fine. Dashboards show provider 429s rising in lockstep with your worker count. How do you triage and mitigate?
What a strong answer looks like
Stop the bleeding first (mitigate), then form hypotheses from real signals. Separate root cause from symptom, communicate status as you go, and close with what prevents a repeat.
0:00 of about 25 min
Which questions mattered is sealed until you submit. Telling you now would just be handing over the edge cases.
Run or narrate your approach, then ask the coach.