SDP Clouds
CI/CD·4 min read

CI Build Times: Caching, Parallelism, and the Fixes That Cut Minutes

A four-minute pipeline people trust beats a forty-minute one they skip. Layered caching, path-based triggering, and parallel tests — with the order that matters.


The pipeline got to nineteen minutes. Nobody complained out loud, so nobody fixed it — people stopped running it before lunch, started batching changes, and batch size quietly grew until a single merge broke four things at once.

Slow CI is a correctness problem disguised as a convenience problem. The fix is layered: cut what's redundant, then what's serial, then what you didn't need to run.

Layer one: stop re-downloading the world

On most pipelines the majority of wall-clock time isn't your code — it's dependency resolution running on every job, in every run.

code
- uses: actions/setup-node@v4
  with:
    node-version: 20
    cache: npm

Two details decide whether this helps:

  • The cache key must include the lockfile. Hash package-lock.json, not a static string. A static key gives a stale cache that either misses constantly or restores the wrong versions.
  • Caches are read-only after creation by default. A lockfile that changes daily will miss if you key on it alone. Fall back to a keyed restore (lockfile hash) plus a broader base key so you still get warmth on miss.

The same applies to base images, Go module downloads, and anything fetched at build time.

Layer two: don't run jobs that have nothing to do

A pipeline that runs the full suite for a one-line docs change trains people to ignore it. Path-based triggering is the fix:

yaml
on:
  pull_request:
    paths:
      - "src/**"
      - "package.json"
      - ".github/workflows/ci.yml"

Be careful what you exclude: lockfiles, workflow definitions, and anything affecting the build image must always trigger. The classic bug is excluding the workflow file itself, so a pipeline change silently doesn't run.

For a monorepo, this compounds: path filters per package turn one big build into "only the packages that changed."

Layer three: run tests in parallel

Serial tests are the most common single source of wasted minutes, and the fix is usually one line.

Most runners have multiple cores doing nothing:

bash
# jest / vitest
npx jest --maxWorkers=4
# go
go test -p 4 ./...
# pytest
pytest -n auto

Two caveats that make parallel tests flaky, which is worse than slow:

  • Shared fixtures. Tests writing to the same temp file, port, or database table will collide. Give each worker its own namespace.
  • Random ports and isolated state. If a test binds :3000, four workers will fight over it.

If tests share state, fix that first. Parallelising broken tests just breaks them four times faster.

Layer four: the jobs that shouldn't be on the critical path

Some work is genuinely required before merge: linting, type-checking, unit tests, the security gate. Other work is important but not blocking — and treating them identically is what turns four minutes into forty.

Split them:

  • Required checks — run on every PR, kept aggressively fast. This is your budget and it should be visible to the team.
  • Scheduled or post-merge — full integration suites, dependency audits, image scans that take minutes, performance smoke tests.

The split that mattered most for us was moving image building and scanning out of the lint job — an artifact that doesn't exist yet can't be scanned in the same job that's still compiling. That single restructure took the required path from nineteen minutes to six. It's the same principle behind gating security checks so they don't crawl: each gate, one job, one artifact, one owner.

Order matters more than technique

Do them in this order, because each one changes what the next is worth:

  1. Measure. Get per-step timings before changing anything. The slow step is rarely the one people guess.
  2. Cache dependencies — usually the biggest single win, and the least invasive.
  3. Parallelise tests — big win, medium risk of flakiness.
  4. Trigger only what changed — structural, and it compounds with the others.
  5. Split blocking from non-blocking — the one that requires a conversation about what "done" means.

Skipping to step 5 first usually just moves the wait somewhere less visible.

Also worth knowing: pipeline structure and the contract it should keep determines how much of this is even possible. If your pipeline is one monolithic job, caching and parallelism are both fighting the shape of the thing — split it first.

Summary

Measure before you optimise, then work outward: cache what you download, skip what you don't need, parallelise what's serial, and get non-blocking work off the required path. A pipeline people actually run before every change is worth more than a thorough one they batch around — and the fastest build is the one whose steps you've looked at rather than assumed.

#ci-cd#build-times#caching#github-actions#devops

SDP Clouds Team

DevOps and cloud engineers writing practical, battle-tested guides on CI/CD, Kubernetes, infrastructure as code, and production operations — every article is based on real incidents and real pipelines, not docs-page rewrites.

More about us →

Related articles