Kubernetes node disk full from debug logs
At 20:30, multiple unrelated services on the same Kubernetes node start failing health checks and pods are getting evicted with 'DiskPressure'. The node's root/ephemeral disk is at 100%. A change this afternoon bumped one service's log level to DEBUG, and it's been emitting verbose JSON logs at high volume; container logs aren't being rotated/shipped fast enough. How do you triage, restore the node, and prevent recurrence?
What a strong answer looks like
Stop the bleeding first (mitigate), then form hypotheses from real signals. Separate root cause from symptom, communicate status as you go, and close with what prevents a repeat.
0:00 of about 30 min
Which questions mattered is sealed until you submit. Telling you now would just be handing over the edge cases.
Run or narrate your approach, then ask the coach.