GitOps Failure Modes: Drift, Sync Loops, and Breaking Glass
What actually goes wrong once Argo CD is running — out-of-band changes, sync loops fighting your controller, and how to break glass without breaking GitOps.
The first month of GitOps is the honeymoon. Everything in git matches the cluster, argocd app diff is empty, and you feel untouchable.
Then someone fixes something directly during an incident, a sync loop starts fighting you, and the diff shows forty files. None of this means GitOps failed — these failure modes each have a known shape.
Drift is not an incident, it's Tuesday
Drift means the cluster no longer matches git. It has exactly three causes:
- Someone changed it out of band —
kubectl edit, a console click, an emergency fix at 3 AM. - Something outside your controller changed it — the HPA scaled replicas, a mutating webhook injected annotations, a cloud controller rewrote a field, the API server defaulted a value you didn't set.
- A previous sync partially applied and never finished.
The second category is the one that surprises people. Fields your manifest never specified get set by defaults, and depending on your diff configuration you'll see a permanent phantom diff — or an auto-sync that reverts the controller's own change on a loop.
Ignore differences for fields you deliberately don't own. Annotations like kubectl.kubernetes.io/last-applied-configuration, HPA-owned spec.replicas, and injected sidecar state are not yours to reconcile:
ignoreDifferences:
- group: apps
kind: Deployment
jsonPathReplicas: /spec/replicas # owned by HPA
- kind: MutatingWebhookConfiguration
jsonPath: /webhooks/*/clientConfig/caBundle
Do this once, properly, and half your phantom diffs disappear permanently.
Detecting drift before your users do
Auto-sync answers "make it match." It does not answer "tell me someone changed it." Different features, and the second one needs enabling.
Turn on diff status in every PR and every Slack channel. Argo CD can report drift even for apps configured to wait for manual sync — worth having everywhere. The goal is an out-of-band change producing a notification within minutes, so the question is "why did you change this?" while the context is fresh.
Pair it with a periodic resync interval: drift caught at the next sync is a Monday-morning ticket, drift caught in three weeks is a postmortem.
The sync loop: when the controller fights you
The classic pattern: sync succeeds, drift reappears within seconds, sync reverts it, repeat. It looks like a bug. It's almost always a disagreement about ownership.
Typical culprits: a controller writing a field your manifest also specifies (HPA and replicas is the canonical one), a webhook injecting defaults into every apply, or two writers — someone's CI still kubectl apply-ing alongside Argo. That last one is worth panicking about: git is not the source of truth and every sync is a coin flip.
The fix is ownership, not retries: stop writing the field if another controller owns it, ignore the difference if it's injected, and kill the second writer immediately if one exists. Two rules prevent most of it — see the Argo CD production setup for the long version: exactly one writer ever applies changes, and every resource has exactly one owner.
Breaking glass without breaking GitOps
Incidents happen and sometimes you must change prod in thirty seconds. Do it while keeping the exception temporary, visible, and reversible.
Freeze, don't fight. Pause auto-sync on the affected app first. A controller reverting your emergency fix mid-incident is how small outages become large ones — the most common GitOps escalation I've seen.
Then:
- Make the change directly. You are now knowingly in drift.
- Announce it — a line in the incident channel with what, why, and which app.
- Fix git in the same incident, not tomorrow: the branch, the PR, the merge — or the next sync turns your fix into an outage.
- Re-enable auto-sync and confirm the diff is empty.
Step 3 is the discipline. A direct change with no git follow-up isn't breaking glass; it's undocumented infrastructure, and it gets reverted the moment sync resumes.
The audit trail you'll want afterwards
Three questions should be answerable from records alone after any break-glass: what changed, who changed it, and does git reflect it now. If any answer is "I'd have to ask around," the process is the problem — which is what a blameless postmortem is built to surface.
Also check whether the emergency change touched anything promotion-related: a hotfix applied straight to production shows up as unexplained drift on your staging app the next morning.
Summary
Drift is normal; undetected drift is the failure. Ignore fields you don't own, surface every diff as a notification, and treat a self-fighting controller as an ownership dispute. When you must break glass: freeze sync first, change prod, fix git in the same incident, unfreeze. GitOps doesn't forbid emergency changes — it makes undocumented ones expensive.
SDP Clouds Team
DevOps and cloud engineers writing practical, battle-tested guides on CI/CD, Kubernetes, infrastructure as code, and production operations — every article is based on real incidents and real pipelines, not docs-page rewrites.
More about us →