Skip to main content

Horizontal Pod Autoscaling

Stage 7 · Orbital Maneuvering

When airline ticket sales launch, search traffic spikes tenfold. Manually adjusting replica counts is slow and reactive.

A HorizontalPodAutoscaler (HPA) automatically adjusts Deployment replica counts based on observed CPU utilization or custom metric thresholds.


The HPA controller feedback loop​

Diagram SC-03 — the HPA reads metrics via the metrics API, evaluates target ratios, and writes updated replica targets to the Deployment.

  • 1. Metrics Pipeline: metrics-server aggregates container CPU and memory metrics from node kubelets.
  • 2. Recommendation Algorithm: Evaluates current utilization against target ratio:
    Desired Replicas = ceil( Current Replicas * ( Current Metric / Target Metric ) )
  • 3. Deployment Scale: Writes new replica values to the Deployment controller.

Why resource requests determine autoscaling correctness​

HPA evaluates CPU utilization as a percentage of requested CPU:

Utilization % = ( Actual CPU Usage / Requested CPU ) * 100
Actual CPURequested CPUCalculated UtilizationHPA reaction (Target: 80%)
80m100m80%Stable; no scale
80m500m16%Erroneously scales down
80m50m160%Rapidly scales up

Critical rule: Resource requests are not optional decorations. Without accurate requests, HPA calculations are mathematically meaningless.


Scale-down stabilization window​

Diagram SC-05 — downscale stabilization retains recent recommendations and uses the highest relevant value instead of sleeping for a fixed period.

  • Flapping hazard: Rapidly alternating between adding and removing Pods when traffic fluctuates.
  • Stabilization window (stabilizationWindowSeconds: 300):
    • Remembers the highest recommended replica count over the preceding 5 minutes.
    • Ensures Pods are not prematurely terminated during brief traffic dips.

Evidence and limits​

  • 1. HPA operational status: Inspect current vs. target metric values:
    kubectl get hpa search -n apollo-airlines-apps
  • 2. Detailed evaluation events: Review scaling decisions:
    kubectl describe hpa search -n apollo-airlines-apps
  • 3. Watch real-time replica scaling:
    kubectl get pods -n apollo-airlines-apps -l app=search -w