VPA and capacity planning
Stage 7 · Orbital Maneuvering
While the Horizontal Pod Autoscaler adjusts the number of replicas, the Vertical Pod Autoscaler (VPA) optimizes the CPU and memory requests allocated to each individual container.
What the VPA Recommender provides
Diagram SC-04 — in recommendation-only mode, VPA observes real resource consumption and outputs suggestions for human review.
- Continuous monitoring: Tracks historical memory peaks and CPU percentiles.
- Three recommended tiers:
lowerBound: Minimum allocation to prevent CPU starvation.target: Ideal baseline request based on steady-state traffic.upperBound: Maximum limit recommended to absorb unexpected spikes without triggeringOOMKilled.
Operating modes: why updateMode: "Off" is safest
| Mode | Behavior | Risk Level |
|---|---|---|
Off (Apollo standard) | Computes recommendations without mutating Pods | Zero risk; human reviews before applying |
Initial | Assigns values only when new Pods are first created | Low; existing running Pods are not restarted |
Auto / Recreate | Forcibly evicts running Pods to apply new limits | High; disrupts active traffic and causes rolling restarts |
Resolving the HPA vs. VPA conflict
Running HPA and VPA concurrently on the same CPU metric causes a destructive race condition:
-
- CPU rises → HPA scales out additional replicas.
-
- Workload distributes → CPU per Pod drops.
-
- VPA interprets lower CPU as over-provisioning → shrinks CPU requests.
-
- Lower CPU request artificially inflates HPA utilization percentage → HPA scales out again.
Operational rule: Never run HPA and VPA simultaneously in
Automode on CPU. Use VPA inOffmode to determine baseline resource sizing, and configure HPA to manage live horizontal scaling.
Evidence and limits
Diagram SC-06 — HPA can request Pods, but scheduler and node capacity determine whether those replicas can run and become ready.
- 1. Inspect VPA recommendations:
kubectl get vpa search -n apollo-airlines-apps -o yaml
- 2. Compare recommendations against live requests:
kubectl top pods -n apollo-airlines-apps -l app=search
- 3. Check for VPA evictions:
kubectl get events -n apollo-airlines-apps --field-selector reason=EvictedByVPA