Clouds
IaC·4 min read

Terraform Drift Is a Process Problem, Not a Tool Problem

Drift detection will tell you what changed. It will not tell you why someone changed it, or who decides whether it stays. The tooling is the easy half.


Every drift conversation I have been part of starts the same way. Someone runs terraform plan after three quiet months, sees forty-seven resources to change, and concludes that the team needs a better drift detection tool.

The tooling is not the problem. The missing thing is a rule about who is allowed to touch infrastructure outside Terraform, and what happens when they do — because the drift is almost always somebody doing their job.

What drift actually is

Drift is any difference between what your state file claims and what exists in the cloud. That definition hides the important part: nearly all of it came from a human with a good reason.

Separating them matters, because each needs a different response:

  • Emergency edits — a console change during an incident, a security group narrowed for a support ticket, a scale-up at 3 AM. Correct at the time, never written back.
  • Permanent escapes — records someone will always add by hand because Terraform should not own them.
  • Managed rot — the resource changed underneath you: an AMI id rotated, a version bumped, a provider default moved.

Only the third is genuinely nobody's fault.

The three checks answer the wrong question

You have three ways to notice, and they all report the same thing:

  • terraform plan against a refreshed state — interactive, depends on someone running it.
  • A scheduled terraform plan in CI, failing a Slack channel — automatic, but the failure becomes wallpaper fast.
  • A dedicated drift job that compares a fresh read to state — the same output, more plumbing.

All three answer what changed. None of them answer why, or should it still be this way, or who owns the decision. That gap is where drift accumulates: the alert fires, nobody is assigned, and the next run looks about the same as the last one.

Detection is the easy half

Once a nightly plan is running and posting somewhere visible, stop counting resources in sync. That number goes up and down with every intentional change and tells you nothing.

The number I watch is the age of the oldest unreviewed drift. A resource that drifted three days ago and was acknowledged is fine — someone looked at it and made a call. A resource that drifted forty days ago means the loop does not close, and no amount of tooling will close it for you.

The rule that makes it stick

Two constraints I have seen actually hold:

Break-glass changes get a deadline, not a prohibition. You cannot stop people fixing production, and trying to will just push the changes somewhere you cannot see them. Instead, require that every console change is resolved within a fixed window — adopt it into state or revert it. Twenty-four hours is a comfortable default. The point is that "someone decides" is written down rather than assumed.

Everything Terraform can reach, Terraform owns. If a resource is editable in the console, make that an explicit choice today. Where a field genuinely must float, an ignore_changes entry in code records the decision and survives review; silent drift records nothing.

hcl
resource "aws_instance" "web" {
  # launch_template owns the AMI; re-reading it just creates false drift
  lifecycle {
    ignore_changes = [ami]
  }
}

Where this sits next to state hygiene

Drift is only one way state and reality can disagree. If the state file itself is wrong — a resource deleted outside Terraform, a half-finished migration, a partial apply — no amount of drift reporting helps you, and the repair looks nothing like a plan. That is state surgery, and it is a different discipline from reading a refresh diff.

Drift handling is what you do once you trust the state file and the world moved anyway. Worth being precise about that order: teams that reach for drift tooling while their state is still unreliable just get a louder report of a problem they already had. And once the adopt-or-revert rule exists, it stops being something anyone remembers — the same plan output can feed a policy gate on the PR without a second pipeline running somewhere else.

Summary

Drift detection is a solved problem — a scheduled plan and somewhere visible to post it. The unsolved half is organisational: a deadline for emergency edits, an explicit decision about what Terraform owns, and one metric that measures how stale the backlog is rather than how green the dashboard looks. Fix the rule first and the tooling stops mattering, because there is far less left for it to find.

#terraform#drift#infrastructure-as-code#process#devops

SDP Clouds Team

DevOps and cloud engineers writing practical, battle-tested guides on CI/CD, Kubernetes, infrastructure as code, and production operations — every article is based on real incidents and real pipelines, not docs-page rewrites.

More about us →

Related articles