Clouds
GitOps·4 min read

GitOps at Fleet Scale: Fifty Clusters Without Fifty Tabs

One cluster in Git is a repository. Fifty clusters in Git is a system with a topology, and the branching strategy you choose decides whether it stays manageable.


At four clusters I had a folder per environment and it was fine. At eighteen the folders had folders, and two clusters were running an image tag that existed nowhere in the repository because someone had patched a deployment directly "just for today." At fifty, the question stopped being how to deploy and became how to describe the fleet at all.

The moment your clusters outnumber your team's memory, the repository stops being a place to store manifests and becomes a place to store intent.

Three shapes, and what each one costs

The options are not really about YAML layout. They are about where a change is expressed once versus fifty times.

text
per-cluster     clusters/prod-eu-1/deployment.yaml   # N copies of everything
per-environment overlays/prod/kustomization.yaml     # N/env, drifts per region
per-role        base/ + roles/db + roles/edge        # compose capabilities

Per-cluster is honest and unscalable. Every cluster is explicit, which makes auditing trivial and makes a one-line change a fifty-line diff.

Per-environment is what most teams reach for next, until they have three regions and a fix must land in one. Overlays compose by replacement, so a patch changing three fields silently discards the other forty.

Per-role inverts the question: instead of listing clusters, list the capabilities a cluster needs — database, ingress, mesh, observability — and compose them. Adding a cluster becomes a declaration of what it is for rather than a copy of what it contains.

I do not recommend starting with role composition at four clusters. I do recommend that every folder-per-cluster decision come with an explicit answer to "what happens when this change must reach forty-nine of these."

App-of-apps, and the blast radius problem

The standard pattern is a root application that points at child applications, one per cluster or per team. It gives you a single place to see the fleet and a single sync that brings everything forward.

It also means a bad edit at the root deploys to everything at once.

yaml
# the root app: one sync, fifty clusters
source:
  repoURL: https://git.example.com/fleet
  path: clusters
destination:
  name: in-cluster          # each child targets its own
syncPolicy:
  automated: { prune: true, selfHeal: true }

prune: true with selfHeal: true at the root is the combination that turns a typo into a fleet-wide deletion that the controller then corrects into. Guard it the way you would guard a production apply: progressive sync waves, so clusters roll in an explicit order, and a synchronisation window that refuses to run while a human is editing.

Rollouts that stop at a region

Progressive delivery across a fleet needs an ordering you chose, not one Git happened to produce.

Three tiers — canary cluster, region, fleet — each gated on the previous one holding for a stated period. Argo Rollouts or a parameter matrix on argocd app set both work; the mechanism matters less than the rule that a failure at tier one stops tier two.

The rule I would insist on: the gate must fail closed without a human. A gate that requires someone to notice and act works only during business hours.

Drift is the fleet's real failure mode

Cluster drift — a hotfixed deployment, a manually scaled replica count — is what GitOps exists to eliminate, and selfHeal handles it. Fleet drift is subtler: two clusters with the same declared intent and different behaviour, because a Helm chart resolved a different dependency version on one of them.

Neither shows up as a sync failure. The controller reports everything synced, which is correct — the declared state matches. It simply was not the whole story.

The answer is verification beyond sync status: a policy check against what is actually running, compared periodically to intent rather than trusting the controller's summary.

What I would tell myself at cluster eighteen

Declare intent at the level where the change is true — role, environment, or cluster — and never higher, because over-generalising a change is how a regional fix reaches every region. Make ordering explicit and the failure path automatic. And treat "is this in Git?" as a weaker claim than "is this what I meant?": a controller will faithfully deploy the wrong sentence to fifty clusters.

For the gating mechanics underneath, the environment promotion post covers the same logic at application scale.

Summary

At fleet scale the repository is a topology, not a storage location. Choose a layout by answering how one change reaches N clusters, gate rollouts in an order you defined with a path that fails closed automatically, and verify live state rather than trusting sync status. The goal is that describing a new cluster should take a sentence, not a copied folder.

#gitops#argocd#kubernetes#multi-cluster#platform-engineering

SDP Clouds Team

DevOps and cloud engineers writing practical, battle-tested guides on CI/CD, Kubernetes, infrastructure as code, and production operations — every article is based on real incidents and real pipelines, not docs-page rewrites.

More about us →

Related articles