SDP Clouds
GitOps·4 min read

GitOps Promotion: Dev to Staging to Prod Without Clicking Buttons

The promotion path is where most GitOps setups quietly break — environments, image tags, and the promotion model that stops staging drifting from production.


Getting Argo CD to sync one environment is the easy part. The setup that survives six months is the one where a change moves from dev to staging to production without anyone kubectl apply-ing anything — and where staging is genuinely the same artifact you're about to ship.

Most teams get the controller right and the promotion path wrong. Here's what I've settled on.

Environments are separate apps, not separate clusters

The mistake I see most often: one repo, one kustomization.yaml, one set of values, and a comment saying "change this for staging." That's not multi-environment, that's a foot-gun with an audience.

Split by values, not by repo. One application repo, three directories:

code
deploy/
  base/          # deployment, service, ingress — no env-specific anything
  overlays/
    dev/         # kustomization + replicas: 1, small resources
    staging/     # kustomization + replicas: 2, prod-like resources
    prod/        # kustomization + replicas: 4, real limits

Then three Argo Application resources pointing at the three overlays, each syncing to its own namespace (or cluster, if you've earned that complexity). Same manifest, different patches. When you fix a securityContext in base, all three get it.

The image tag is the promotion

This is the decision that makes or breaks the pipeline: what exactly is being promoted?

Not "the latest commit to main." Not "whatever :latest resolves to right now." The promotion unit is an immutable image digest, tagged with the git SHA that built it.

yaml
# staging overlay
images:
  - name: app
    newName: registry.example.com/app
    newTag: 1.14.2-ge3a9c1b

Promotion is then a one-line diff to the tag in the staging overlay, reviewed and merged like any other change. Promote to production = the same tag, in the prod overlay. The binary never changes between environments, so "it worked in staging" means something.

latest breaks all of this. If you're still using it, your supply chain posture has a bigger problem than promotion.

The promotion model: pull, don't push

Two patterns work. Pick one deliberately.

Git-only promotion. A human edits the overlay, opens a PR, merges. Argo syncs within seconds. Simple, auditable, and every promotion has an approver. Slower by design — which is usually correct for production.

Automated promotion with a gate. CI opens the PR after tests pass; a manual approval or an automated health check merges it. Faster, and it works well once the health signal is trustworthy.

What doesn't work is CI pushing directly to prod and Argo also syncing — you've built two writers to the same state, which is precisely what GitOps exists to prevent. The rule from the Argo CD production setup still holds: the controller is the only thing that applies to the cluster.

Sync waves for the boring dependencies

Promotion isn't just the app. It's the migration job, the config, the certificate, the app itself. Get the order wrong and staging looks broken for ninety seconds on every deploy.

Argo's sync waves handle it: annotations on resources that sort them.

yaml
metadata:
  annotations:
    argocd.argoproj.io/sync-wave: "-1"   # migration, before the app

Negative waves run first, positive later. Database migrations at -1, application at 0, anything that depends on it at 1. Combined with a sync hook that runs and fails loudly when the migration fails, this turns "did it apply in the right order?" from a source of incidents into a sorted list.

Health checks that gate the promotion

The promotion gate is only as good as the health signal behind it. A Deployment that reports Healthy the instant pods exist tells you nothing about whether the app works.

Add a proper health.lua or use the built-in checks for:

  • Deployment ready replicas — not just created.
  • An HTTP health endpoint that verifies a real dependency, not just "the process started."
  • Migration status — did it complete, and is the schema version what the new code expects?

If the gate passes and the rollout then fails, your gate is decorative. If the gate correctly blocks, you've prevented a bad promotion rather than discovered one in the morning.

The drift you can't see

The last failure mode: promotion works, but something outside Git changes the cluster — someone scales a replica count manually, a kubectl edit fixes a resource limit during an incident.

Argo will eventually notice the drift, but the promotion is now testing a configuration that doesn't match what production actually ran.

Turn on drift detection explicitly, and treat any out-of-band change as a bug to be reverted and re-applied through the overlay. Postmortems should record these — the pattern usually reveals a missing runbook or an alert that fires too slowly to be handled through the normal path.

Summary

One repo, three overlays, immutable image tags as the promotion unit, Git as the only writer, and health checks that actually gate the rollout. Get those five right and promotion stops being a series of tense Slack messages and becomes a diff someone approves in ninety seconds.

#gitops#argo-cd#kubernetes#cd#environments

SDP Clouds Team

DevOps and cloud engineers writing practical, battle-tested guides on CI/CD, Kubernetes, infrastructure as code, and production operations — every article is based on real incidents and real pipelines, not docs-page rewrites.

More about us →

Related articles