Horizontal Pod Autoscaling
Stage 7 · Orbital Maneuvering
When airline ticket sales launch, search traffic spikes tenfold. Manually adjusting replica counts is slow and reactive.
A HorizontalPodAutoscaler (HPA) automatically adjusts Deployment replica counts based on observed CPU utilization or custom metric thresholds.
The HPA controller feedback loop
Diagram SC-03 — the HPA reads metrics via the metrics API, evaluates target ratios, and writes updated replica targets to the Deployment.
- 1. Metrics Pipeline:
metrics-serveraggregates container CPU and memory metrics from node kubelets. - 2. Recommendation Algorithm: Evaluates current utilization against target ratio:
Desired Replicas = ceil( Current Replicas * ( Current Metric / Target Metric ) )
- 3. Deployment Scale: Writes new replica values to the Deployment controller.
Why resource requests determine autoscaling correctness
HPA evaluates CPU utilization as a percentage of requested CPU:
Utilization % = ( Actual CPU Usage / Requested CPU ) * 100
| Actual CPU | Requested CPU | Calculated Utilization | HPA reaction (Target: 80%) |
|---|---|---|---|
| 80m | 100m | 80% | Stable; no scale |
| 80m | 500m | 16% | Erroneously scales down |
| 80m | 50m | 160% | Rapidly scales up |
Critical rule: Resource requests are not optional decorations. Without accurate requests, HPA calculations are mathematically meaningless.
Scale-down stabilization window
Diagram SC-05 — downscale stabilization retains recent recommendations and uses the highest relevant value instead of sleeping for a fixed period.
- Flapping hazard: Rapidly alternating between adding and removing Pods when traffic fluctuates.
- Stabilization window (
stabilizationWindowSeconds: 300):- Remembers the highest recommended replica count over the preceding 5 minutes.
- Ensures Pods are not prematurely terminated during brief traffic dips.
Evidence and limits
- 1. HPA operational status: Inspect current vs. target metric values:
kubectl get hpa search -n apollo-airlines-apps
- 2. Detailed evaluation events: Review scaling decisions:
kubectl describe hpa search -n apollo-airlines-apps
- 3. Watch real-time replica scaling:
kubectl get pods -n apollo-airlines-apps -l app=search -w