Your requests and limits are where the cloud bill hides
How right-sizing Kubernetes workloads cut public cloud spend by roughly 40% across thirty-odd clusters.
Kubernetes schedules on requests, not on what your pods actually use. So when every team copies cpu: 1, memory: 2Gi from the last service they wrote, the cluster autoscaler dutifully buys nodes for capacity nobody touches.
The pattern
- Measure real usage over a representative window (p95, not average).
- Set requests close to that, and leave headroom where it matters.
- Be careful with CPU limits: throttling is often worse than the noisy neighbour you were afraid of.
- Automate it, because hand-tuned values drift within a quarter.
resources:
requests:
cpu: 150m
memory: 320Mi
limits:
memory: 512Mi
Tooling beats policy
We wrote a small Go tool that compared requests against observed usage and opened pull requests with suggestions. Engineers accepted most of them, because a PR with numbers is easier to say yes to than a wiki page with rules.