Clouds
Kubernetes·4 min read

Kubernetes Scheduling: Taints, Affinity, and the Pods That Won't Move

A pod that will not schedule is rarely a capacity problem. It is almost always a rule you wrote six months ago that nobody remembers writing.


The usual reaction to Pending is to add a node. I have done it three times, and two of those times the cluster had spare capacity the whole time — the pod was refusing a node that already had room, because of a taint someone added during an incident and never removed.

Scheduling is a negotiation between what a pod asks for and what a node advertises. When it fails, one of those two sides has a rule the other side cannot satisfy, and kubectl describe pod will tell you which.

Read the event before you touch anything

text
Warning  FailedScheduling  0/6 nodes are available:
  3 node(s) didn't match Pod's node affinity/selector.
  2 node(s) had untolerated taint {dedicated: monitoring}.
  1 node(s) didn't have enough pods.

Those three lines are three different problems. The first means your nodeSelector or affinity rules exclude everything. The second means a taint exists and your pod has no matching toleration. The third is the one that actually means "add a node."

Most Pending pods fall into the first two, and adding capacity fixes neither.

Taints push, tolerations allow, selectors pull

The distinction is worth being precise about, because conflating them produces rules that half-work.

A taint is on a node and repels pods that do not tolerate it. It is how you reserve a node pool — for monitoring agents, for GPU work, for a tenant — without editing every workload that must not land there.

A toleration is on a pod and says "I accept that taint." It does not attract the pod to the node; it merely stops the taint from excluding it.

A node selector / affinity term is on the pod and pulls it toward matching nodes.

yaml
# node: reserved for monitoring agents
taints:
  - key: dedicated
    value: monitoring
    effect: NoSchedule

# pod: the only thing allowed there
tolerations:
  - key: dedicated
    operator: Equal
    value: monitoring
    effect: NoSchedule
nodeSelector:
  workload: monitoring

Note that the toleration alone would let any pod with it land on any node. The nodeSelector is what actually directs the agent to the reserved pool. Taints keep things out; selectors put things in; they solve different halves.

Preferences versus requirements

requiredDuringSchedulingIgnoredDuringExecution is a hard gate. If no node matches, the pod stays Pending forever.

preferredDuringSchedulingIgnoredDuringExecution is a nudge with a weight. The scheduler tries to honour it and falls back to anything feasible.

The "Ignored" half means the rule is evaluated at scheduling time and then forgotten — if you later label the node differently, a running pod will not move. That surprises people who expect relabelling to migrate workloads. It does not; you delete the pod and it is re-evaluated.

For anything that genuinely must not run somewhere — a regulatory boundary, a single-tenant node pool — use the required form plus a taint on the node, so the constraint holds even if someone deploys without the selector.

Topology spread is the one people skip

Spreading replicas across zones is usually what you actually wanted from affinity, expressed more honestly:

yaml
topologySpreadConstraints:
  - maxSkew: 1
    topologyKey: topology.kubernetes.io/zone
    whenUnsatisfiable: ScheduleAnyway
    labelSelector:
      matchLabels: { app: api }

ScheduleAnyway rather than DoNotSchedule is deliberate. The strict variant can make a pod unschedulable when one zone is full; the lenient variant degrades your spread instead of blocking a deployment. For most services, degraded spread beats a Pending replica.

Where teams go wrong

Taints added during incidents and never cleaned up — the single most common cause of mystery Pending pods I have diagnosed. Tolerations without selectors, which grant access without directing placement. Required affinity with an impossible term, which converts a preference into an outage. And treating Pending as a capacity signal when the event clearly says otherwise.

If you want the same "read the event, not the symptom" habit applied to crashes, it is the same discipline as triaging a failing pod in five minutes.

Summary

Scheduling failures are rule mismatches, not capacity problems. Read FailedScheduling first — it names the exact constraint. Use taints to exclude, tolerations to permit, selectors and affinity to direct, and topology spread for distribution. Prefer soft constraints where a blocked deployment costs more than imperfect placement, and clean up incident-era taints before they become folklore.

#kubernetes#scheduling#taints#affinity#nodes

SDP Clouds Team

DevOps and cloud engineers writing practical, battle-tested guides on CI/CD, Kubernetes, infrastructure as code, and production operations — every article is based on real incidents and real pipelines, not docs-page rewrites.

More about us →

Related articles