Clouds
Cloud·4 min read

Autoscaling Without the Oscillation

Horizontal autoscalers are simple to enable and easy to misconfigure, and the symptom is never a slow page, it is a sawtooth of latency that no single threshold will explain.


The dashboard showed a perfect sawtooth: latency climbing for four minutes, a scale-out event, a drop, then a scale-in that arrived just in time to lose the cache and start climbing again. Throughput was flat the entire time. The service was not under load; it was under a rule that kept undoing itself.

Autoscaling had been enabled by a checkbox. The parameters had never been reasoned about.

The two knobs that actually matter

Almost every horizontal autoscaler reduces to a target and a stabilization window, and the second one is what teams skip.

The target is the metric threshold you want to hold: 70% CPU, 100 in-flight requests per pod, 500ms p95. The controller compares current value against target and computes a desired replica count.

The stabilization window decides how long the controller waits before acting on a change. Scale-out and scale-in usually have separate windows, and they should: scaling out quickly when latency spikes is responsive, while scaling in quickly after the spike ends is what causes the oscillation.

yaml
behavior:
  scaleUp:
    stabilizationWindowSeconds: 0
    policies:
      - type: Percent
        value: 100
        periodSeconds: 60
  scaleDown:
    stabilizationWindowSeconds: 300
    policies:
      - type: Pods
        value: 1
        periodSeconds: 60

The asymmetry is the point. Scale out when demand arrives; scale in slowly after demand has stayed low for five minutes. A symmetric window on both directions produces the sawtooth, because the controller is equally eager to undo what it just did.

Why CPU is the wrong default

CPU utilization is the default target because it is always available and easy to graph. It is also frequently the wrong signal.

A service blocked on a database connection pool sits at 12% CPU while requests queue for seconds. A service doing heavy JSON serialization sits at 90% CPU while responding in 20ms. In both cases the autoscaler's chosen metric is decoupled from what users experience.

Concurrency-based targets align better for request-serving workloads: target a fixed number of in-flight requests per pod, and replica count becomes a direct function of arrival rate rather than of an intermediate resource. If each pod handles 100 concurrent requests comfortably, adding 1,000 requests per second should add replicas automatically, without anyone picking a CPU number.

The catch is that concurrency requires instrumentation the platform does not provide for free: you export it from the application, so the metric is only as trustworthy as the middleware emitting it.

The cold-start tax

Scale-in is where the cost hides. A new replica has to load the model, warm the JIT, populate the connection pool, and pull the asset bundle, and during that window it either serves slowly or fails health checks while the load balancer still sends it traffic. Capacity therefore arrives seconds late, exactly when the scale-out decision was supposed to help.

The better lever is making the replica ready faster: warm pools of pre-initialized instances, health checks that gate on readiness rather than liveness, connection pools opened at startup rather than on first request, and P99 startup time tracked as a first-class metric alongside latency.

What to watch for

Four signals distinguish a healthy autoscaler from an oscillating one:

Replica count changing more than a few times per hour under roughly constant traffic. Demand that does not vary should not produce varying capacity.

Scale-in events shortly after scale-out events. This is the sawtooth, and it almost always means the stabilization window is too short on the down side.

Latency p95 rising while CPU stays flat. The autoscaler is watching the wrong metric.

Cost per request trending up while throughput stays flat. Scaling is adding capacity the workload cannot use: the same signal right-sizing EC2 before it bites your bill tracks per instance.

The rollout worth doing

Turn the autoscaler on with a deliberately conservative target, record replica count over a full business cycle, then tighten. The mistake is an aggressive target on day one, discovering the interaction between scale-in windows and cache warmth in production.

promql
# should be smooth; spiky means the window is too short
changes(kube_deployment_status_replicas_available[15m])

# should track load; divergence means the target is wrong
kube_pod_container_resource_requests_cpu_cores
  / on() group_left label_replace(kube_pod_info, "x", "1", "x", "x")

Two charts, checked weekly, catch nearly every misconfiguration before it becomes a bill.

Summary

Autoscaling oscillates when scale-in is as eager as scale-out; asymmetric stabilization windows fix most sawtooths on their own. Replace CPU with a concurrency or latency target whenever the workload is request-serving, since a service waiting on a database is idle by CPU's definition. Track replica-count churn and cost-per-request as first-class signals, and treat cold-start time as part of the scaling budget rather than an afterthought.

#cloud#autoscaling#scaling#performance#devops

SDP Clouds Team

DevOps and cloud engineers writing practical, battle-tested guides on CI/CD, Kubernetes, infrastructure as code, and production operations: every article is based on real incidents and real pipelines, not docs-page rewrites.

More about us →

Related articles