Terraform State Is the Whole Ballgame
What actually lives in a Terraform state file, why the backend choice matters more than your module code, and how to recover from a stuck lock.
Everybody reads the .tf files. Almost nobody reads the state file until the day something is stuck. That day is when you discover the state is the actual product of Terraform and your HCL is just a request.
What is actually in there
State is a JSON mapping between your configuration and real cloud resources. Each managed resource gets an entry with its ID, its attributes as Terraform last observed them, and dependency edges. For an AWS estate it routinely runs to megabytes and contains things you did not intend to publish: database endpoints, connection strings, and any secret you passed through a variable without sensitive = true.
Two consequences follow. First, state is a source of truth, not a cache — if you delete it, Terraform loses track of everything and will happily try to create duplicates. Second, state is a secret store, which is why committing it to git alongside your HCL is a category error, not a shortcut.
The backend is not a detail
Local state works for one person on one machine. The moment a second engineer, a CI runner, or a second workstation exists, you need a remote backend with locking. For AWS the standard answer is unchanged:
terraform {
backend "s3" {
bucket = "acme-tfstate"
key = "platform/network/terraform.tfstate"
region = "eu-west-1"
dynamodb_table = "acme-tfstate-lock"
encrypt = true
}
}
The DynamoDB table is the part people skip. Without it, two apply runs read the same state, both compute a plan, and the second one overwrites the first. You do not find out until a resource exists twice, or half of one change and none of another.
Use one state file per blast radius. A single enormous state means a typo in an unrelated module can lock the entire estate behind the same lock.
When the lock will not go away
The failure mode that generates the most panic: an apply is interrupted, the lock record survives, and every subsequent run returns Error acquiring the state lock.
Do not reach for force-unlock first. Ask what was holding it. If a pipeline is genuinely still running, breaking the lock lets two applies race — which is the exact problem locking exists to prevent. Check whether the run is still live, and if it is not, clear it deliberately:
terraform force-unlock jib0oq-4h9k2p-7c1m3x
The ID comes from the error message. Treat the command as an outage tool, not a housekeeping step. A stale lock you clear often means an interrupted process you have not accounted for.
Recovering from bad state
Sometimes the state is wrong rather than locked — a resource was deleted in the console, or an entry points at something that no longer exists. The safe moves are narrow and surgical:
terraform state list
terraform state show aws_s3_bucket.logs
terraform state rm aws_s3_bucket.logs
terraform import aws_s3_bucket.logs acme-logs
state rm detaches without destroying. import re-attaches an existing object. Between them you can repair almost anything without touching real infrastructure. What you should not do is hand-edit terraform.tfstate: the serial number, the resource IDs, and the dependency graph all have to stay consistent, and one wrong character turns a recoverable problem into a rebuild.
Before any of this, copy the state file somewhere. A state mv executed against a shared backend has no undo.
Secrets deserve their own answer
If a secret has already reached state, marking the variable sensitive = true hides it from plan output but does not remove it from the file. Rotate the credential, then decide whether it belongs in state at all — that is really a question about your secrets management habits. The durable pattern is to keep secrets in a manager and reference them by ARN, so the value never lands in Terraform's hands.
Summary
State is the real artifact: protect it with a remote backend and locking, keep one file per blast radius, repair it with state rm and import rather than hand edits, and assume anything you put in it is readable. Your modules can be perfect and still lose you an afternoon if the state behind them is a single unencrypted file on someone's laptop.
SDP Clouds Team
DevOps and cloud engineers writing practical, battle-tested guides on CI/CD, Kubernetes, infrastructure as code, and production operations — every article is based on real incidents and real pipelines, not docs-page rewrites.
More about us →