First Kubernetes Failure: Node Out of Disk
Adam
- One minute read - 130 wordsWhat Happened
Our production Kubernetes cluster experienced a complete outage when all worker nodes ran out of disk space simultaneously. This was caused by:
- Unbounded log growth from a misconfigured application
- No disk space monitoring alerts
- Default emptyDir volumes filling up node storage
The Impact
- All pods became pending as nodes were marked unschedulable
- 4 hours of downtime during peak traffic
- Data loss for applications using emptyDir storage
Lessons Learned
- Monitor disk usage: Set up alerts at 70%, 80%, and 90% capacity
- Log rotation: Configure proper log rotation for all containers
- Resource limits: Set storage limits on emptyDir volumes
- Multi-zone deployment: Distribute workloads across availability zones
The Fix
We implemented:
- Prometheus alerts for node disk usage
- Fluentd with log rotation configuration
- Resource quotas for storage
- Regular storage capacity planning reviews