Skip to main content

VPA and capacity planning

Stage 7 · Orbital Maneuvering

While the Horizontal Pod Autoscaler adjusts the number of replicas, the Vertical Pod Autoscaler (VPA) optimizes the CPU and memory requests allocated to each individual container.


What the VPA Recommender provides​

Diagram SC-04 — in recommendation-only mode, VPA observes real resource consumption and outputs suggestions for human review.

  • Continuous monitoring: Tracks historical memory peaks and CPU percentiles.
  • Three recommended tiers:
    • lowerBound: Minimum allocation to prevent CPU starvation.
    • target: Ideal baseline request based on steady-state traffic.
    • upperBound: Maximum limit recommended to absorb unexpected spikes without triggering OOMKilled.

Operating modes: why updateMode: "Off" is safest​

ModeBehaviorRisk Level
Off (Apollo standard)Computes recommendations without mutating PodsZero risk; human reviews before applying
InitialAssigns values only when new Pods are first createdLow; existing running Pods are not restarted
Auto / RecreateForcibly evicts running Pods to apply new limitsHigh; disrupts active traffic and causes rolling restarts

Resolving the HPA vs. VPA conflict​

Running HPA and VPA concurrently on the same CPU metric causes a destructive race condition:

    1. CPU rises → HPA scales out additional replicas.
    1. Workload distributes → CPU per Pod drops.
    1. VPA interprets lower CPU as over-provisioning → shrinks CPU requests.
    1. Lower CPU request artificially inflates HPA utilization percentage → HPA scales out again.

Operational rule: Never run HPA and VPA simultaneously in Auto mode on CPU. Use VPA in Off mode to determine baseline resource sizing, and configure HPA to manage live horizontal scaling.


Evidence and limits​

Diagram SC-06 — HPA can request Pods, but scheduler and node capacity determine whether those replicas can run and become ready.

  • 1. Inspect VPA recommendations:
    kubectl get vpa search -n apollo-airlines-apps -o yaml
  • 2. Compare recommendations against live requests:
    kubectl top pods -n apollo-airlines-apps -l app=search
  • 3. Check for VPA evictions:
    kubectl get events -n apollo-airlines-apps --field-selector reason=EvictedByVPA