Skip to main content

Scheduling and placement

Stage 4 · Flight Control

The Kubernetes kube-scheduler assigns unscheduled Pods (spec.nodeName empty) to feasible cluster nodes.

Placement decisions follow a deterministic multi-stage algorithm rather than arbitrary round-robin assignment.


The three-phase scheduling lifecycle​

Diagram RL-04 — the scheduler filters infeasible nodes before scoring feasible ones; binding writes nodeName.

  • 1. Filtering (Predicates):
    • Eliminates all nodes incapable of hosting the Pod (insufficient CPU/RAM, untolerated taints, missing node affinity labels).
  • 2. Scoring (Priorities):
    • Ranks remaining feasible nodes using algorithms such as NodeResourcesLeastAllocated or topology spread weights.
  • 3. Binding:
    • Atomically writes the winning node's name to spec.nodeName, triggering the node's local kubelet to pull images and boot containers.

What a Pending status actually indicates​

A Pending Pod indicates placement constraints have not been met:

Symptom / ReasonUnderlying CauseRemediation
Insufficient cpu/memoryTotal node requests exceed allocatable capacityAdd node capacity or scale down workloads
Untolerated taintNode has taint (e.g. dedicated=search:NoSchedule)Add matching toleration to Pod spec
Node affinity mismatchPod requires label not present on any nodeLabel the target nodes
Unbound PVCStorage provisioner waiting on first consumerInvestigate PVC binding status

Taints, tolerations, and affinity controls​

  • Taints & Tolerations:
    • Taint on Node: Repels Pods (kubectl taint nodes node1 dedicated=db:NoSchedule).
    • Toleration on Pod: Allows the Pod to schedule on tainted nodes.
  • Node Affinity:
    • Attracts Pods to nodes with matching labels (kubernetes.io/os=linux).
    • requiredDuringSchedulingIgnoredDuringExecution: Strict hard requirement.
    • preferredDuringSchedulingIgnoredDuringExecution: Soft preference.
  • Topology Spread Constraints:
    • Distributes replicas evenly across failure domains (racks, availability zones, hostnames) to prevent single-point-of-failure concentration (maxSkew: 1).

Evidence and limits​

  • 1. Scheduling failure events: Inspect why the scheduler rejected nodes:
    kubectl describe pod <pod-name> -n apollo-airlines-apps | grep -A 10 Events:
  • 2. Node allocation breakdown: Review allocatable resources per node:
    kubectl describe node apollo11-worker | grep -A 10 "Allocated resources"
  • 3. Active node taints:
    kubectl get nodes -o custom-columns='NAME:.metadata.name,TAINTS:.spec.taints'