Peter Wiggers
Clay: Kubernetes & platform

Making Flux fast with 1,200 tenants in one repository

What I learned tuning Flux's kustomize-controller for a monorepo with 1,200 tenants, including a /tmp trick that made it into the official Flux docs.

3 min read

I love Flux. It does one job, GitOps, and it does it well. But at one point we were pushing it harder than most setups do.

We ran a single-tenant SaaS: every customer got their own set of about 20 Kubernetes objects. Around 1,200 tenants lived on one cluster, and our GitOps repository had one directory per tenant. One GitRepository, 1,200 Kustomizations.

It worked. It just wasn’t fast.

The problem: every commit reconciles everything

A full reconciliation of all tenants took about 13 minutes. That would be fine if it only happened occasionally. But changing the config of one tenant triggered a reconciliation of all of them. Push a one-line fix for a customer, and you could wait up to 13 minutes to see it applied.

I asked about it in a GitHub discussion, and Stefan Prodan, one of Flux’s maintainers, explained why. Flux can’t decide what changed by looking at files, because with Kustomize a change in a shared base can affect every overlay. And even if nothing changed in Git, something may have drifted in the cluster. So on every new revision, Flux runs a server-side apply dry-run for every Kustomization and only applies what actually differs.

That’s correct behaviour. So the job was to make each reconciliation as cheap as possible.

What actually helped

1. More parallelism, a lot more. By default the kustomize-controller reconciles only a few objects at a time. Each reconciliation runs in a goroutine, which is lightweight, so you can go much higher than the number of CPU cores. We went to --concurrent=50.

2. Enough CPU, without throttling. At first, raising concurrency made things worse: individual reconciliations went from 0.5 seconds to about 20. Two bottlenecks were hiding behind each other. My test cluster’s Kubernetes control plane was too small for the load (on GKE it scales with cluster size), and a CPU limit of 4 cores was throttling the controller. On a production-sized control plane, with 16 cores available, the controller used all of its CPU during a run, which is exactly what you want.

3. Put /tmp in memory. The kustomize-controller builds every overlay in /tmp. On a normal emptyDir, that’s disk I/O, and disks get throttled. Mounting /tmp as a RAM disk removed that bottleneck completely. Others have since run into the same thing on AWS, where gp2 volumes ran out of burst credits.

4. A longer interval. With 1,200 objects, there’s no need to check for drift every few minutes. We set interval: 60m on the tenant Kustomizations. New commits still trigger a reconciliation immediately.

If you bootstrap Flux, you can patch the controller in your flux-system/kustomization.yaml:

apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
  - gotk-components.yaml
  - gotk-sync.yaml
patches:
  - target:
      kind: Deployment
      name: kustomize-controller
    patch: |
      - op: add
        path: /spec/template/spec/containers/0/args/-
        value: --concurrent=50
      - op: replace
        path: /spec/template/spec/volumes/0
        value:
          name: temp
          emptyDir:
            medium: Memory

Raise the CPU and memory limits in the same patch, to match whatever your cluster can give it.

Now in the official docs

Stefan liked the RAM disk tip enough to add it to Flux’s recommendations. Today it’s part of the official vertical scaling guide, next to the concurrency and resource settings. It’s a small thing, but seeing a fix from your own production problem end up in the docs of a CNCF project is a nice feeling.

The real fix for monorepos, since Flux 2.7

Tuning makes a full run cheaper. It doesn’t stop the full run from happening. Since Flux 2.7, you can do that too.

The new ArtifactGenerator API (enable it with --components-extra=source-watcher) splits one repository into separate artifacts, one per directory. A tenant’s Kustomization then only reconciles when its own artifact changes:

apiVersion: source.extensions.fluxcd.io/v1beta1
kind: ArtifactGenerator
metadata:
  name: tenants
  namespace: flux-system
spec:
  sources:
    - alias: repo
      kind: GitRepository
      name: flux-system
  artifacts:
    - name: tenant-acme
      originRevision: "@repo"
      copy:
        - from: "@repo/tenants/acme/**"
          to: "@artifact/"

The tenant’s Kustomization points at kind: ExternalArtifact, name: tenant-acme instead of the GitRepository. With 1,200 tenants you’d generate that list instead of writing it by hand, but the effect is exactly what I was asking for back then: change one tenant, reconcile one tenant.

The takeaway

When a controller is slow, it’s tempting to give it more replicas or more cores and hope. It’s more useful to find out what it’s actually waiting for. In our case that was the API server, CPU throttling and disk I/O, in that order. Fix them one at a time and measure in between.