Policy as Code with OPA and Conftest: Catching Bad Plans Before Apply
Terraform plans run to thousands of lines and nobody reads them all. Turn 'is this plan acceptable?' into a test that fails the PR, not an argument after it lands.
The plan was 4,100 lines and approving it took nine seconds.
Nothing in it looked alarming, because nothing in a 4,100-line diff looks like anything. It removed a database subnet group, opened port 22 to the world, and rewrote an S3 bucket policy. Two of those were intended. The third was not, and we learned about it from an alert instead of from the pull request.
"Read the plan more carefully" has been the advice since 2016 and it has never once scaled. Review is not a control. A test is a control.
What actually gets evaluated
Policy-as-code takes the machine-readable plan, hands it to an engine, and asks questions in a real language: is any security group open to the world on an admin port? does this destroy a resource whose name starts with prod-? does everything carry an owner tag?
The engine is usually OPA, which means writing Rego. For Terraform there are three entry points people genuinely use, and the second is the one that matters:
terraform plan -out=tfplan
terraform show -json tfplan > plan.json
conftest test plan.json # fails the job, prints the violated rule
conftest is where I start teams. It reads plain YAML policies, prints output a human can act on, and exits non-zero on failure — so it drops into an existing CI step with no new infrastructure and no new dashboard.
Four rules that pay for themselves
Write policies that catch mistakes you have actually made. If you have never made them, you are writing compliance theatre.
- No
0.0.0.0/0on 22, 3389, or your database port. - No destroys of anything named
prod-*without an explicit label on the PR. - Every resource tagged with an owner. Cheapest quality win in the whole stack.
- No new resources in the
sharedaccount without going through the module.
Then let it run for two weeks before you make it blocking. The first thing you will discover is that half your existing infrastructure violates rule 3, and blocking on that day one turns the team against the practice permanently.
The gate that becomes a nuisance
The failure mode is not that policies are too strict. It is that they are too noisy, and the team learns to work around them.
Symptoms: a policy-exceptions.tf file that grows faster than the policies; approvals granted by whoever is least tired; a Slack channel called #waive-it. Every one of those means the rules were written before the workflows were understood.
Three things keep it honest:
- Judge only what this PR introduces. Compare the current plan against the proposed one rather than the whole estate. Legacy debt stops being a blocker the day you adopt the practice.
- Keep the rules short enough to memorise. Six rules a developer can recite beat forty they have to look up. More than eight and people stop reading the failure messages.
- Make the message teach. "Rule
no-open-sshviolated" produces a fight. "Port 22 is open to 0.0.0.0/0 onweb-sg— if this is intentional, add the resource toallowed-openwith a comment" produces a fix.
What it will not catch
It will not catch a plan that is well-formed and wrong. If someone deletes the correct subnet group by mistake, every rule passes and the outage happens anyway. That is what environment separation and state discipline are for — policy checks the shape of a change, not the intent behind it.
It also does not replace human review of the interesting parts. The value is that review time gets spent on the twenty lines that are genuinely ambiguous rather than the four thousand that are generated.
And it is not a substitute for locking: policy runs before apply, but state is still the whole ballgame when two runs race.
Summary
Plans are too large to review by eye, so stop asking eyes to do it. Emit the plan as JSON, run conftest over it with four rules drawn from your own incident history, let it report without blocking for two weeks, then make it blocking on new changes only. The goal is not compliance — it is that the mistake gets caught while it is still a diff.
SDP Clouds Team
DevOps and cloud engineers writing practical, battle-tested guides on CI/CD, Kubernetes, infrastructure as code, and production operations — every article is based on real incidents and real pipelines, not docs-page rewrites.
More about us →