SDP Clouds
AI·4 min read

AI Evals for People Who Don't Do Machine Learning

A golden set, a rubric, and a regression run are enough to know whether your prompt change helped or hurt — you don't need a research team to stop shipping AI on vibes.


You changed a prompt. The output looks better. You tried it on five examples you already knew the answer to, felt good, and merged it. Two weeks later someone reports that it started confidently summarising the wrong document, and nobody can say when that started.

This is not a model problem. It is a testing problem, and it is the same testing problem every other change in your system has — you just haven't built the harness yet. The good news is that an eval good enough to catch real regressions takes an afternoon, not a research budget.

Start with ten cases you already know the answer to

Every team that ships an AI feature has already been shown the failure modes by users. Those complaints are your test set. Pull out ten to twenty real inputs: the document that got misread, the query that returned nothing, the summary that invented a deadline, the two phrasings that produced different answers for the same question.

Write down what the correct output should look like. Not the ideal output — the acceptable output. Sometimes that's a bullet list, sometimes it's an ID, sometimes it's "refuses and asks a clarifying question."

Store them in the repo as fixtures, alongside the prompt. A prompt with no fixtures next to it is a prompt nobody can safely change.

json
{
  "input": "summarise the Q3 incident channel",
  "expect": {
    "must_mention": ["root cause", "customer impact"],
    "must_not": ["invented headcount", "specific dollar figures"],
    "acceptable_style": ["bullets", "short paragraph"]
  }
}

Keep the set small and stable. Fifteen cases you trust beat three hundred you generated once and never reviewed.

Grade with a rubric, not a vibe

The instinct is to read the outputs and decide. That works until you're tired, and it doesn't work at all when two people disagree — which is the actual situation an eval exists to resolve.

Turn your judgement into three or four questions with yes/no answers:

  1. Did it answer the question that was asked?
  2. Is every claim supported by the input?
  3. Is it the right shape and length?
  4. Would I ship this to a customer as-is?

Grade each output against the rubric rather than against your mood. Disagreement between two graders on the same case is a signal your rubric is vague — fix the rubric, not the grader.

The judge model is a colleague, not an oracle

You can use a model to grade outputs, and for a lot of mechanical checks — was the JSON valid, did it include the required key, is it under 200 words — it is cheap and consistent. Use it there.

Where it gets dangerous is using a model to grade another model on quality. Models are generous to each other, they share blind spots, and they will happily rate a fluent wrong answer above a clumsy right one. If you use an LLM judge for subjective criteria, calibrate it: take thirty cases you graded by hand, see where it disagrees with you, and treat those disagreements as the places it is unreliable.

Wire it into CI like any other test

The whole point is catching regressions before they ship. A nightly job that nobody reads is a dashboard, not a gate.

  • Run the fast, deterministic checks (structure, required fields, refusal behaviour) on every prompt change, in CI.
  • Run the slower, model-graded checks on merge to main, or nightly if they cost real money.
  • Fail the build when a case that used to pass now fails. That is the definition of a regression.
  • Track a score over time so you can see whether the last three prompt edits netted out positive.

You will also want a canary: one request against production traffic patterns per deploy, logged, so you notice distribution drift that your fixtures don't cover.

What I would skip until it hurts

Not now: a custom evaluation platform, a preference-model trained on your data, or a benchmark suite for capabilities your product doesn't use. Evaluation tooling has a habit of becoming its own project. The fifteen fixtures and the four-question rubric will outlive all of it.

Also skip the accuracy percentage in your sprint report. A single number invites the question "is 87% good?" and the honest answer is always "compared to what, on which cases, graded how?" Keep the case list and the pass/fail per case visible instead.

Summary

Fixtures from real failures, a rubric two people can agree on, a judge used only where it's calibrated, and a gate in CI that fails on regressions — that is a working eval. It costs an afternoon to build and it is the difference between shipping an AI feature you can change safely and one you're afraid to touch. See also what survived a year of AI in our pipeline.

#ai#evals#testing#devops

SDP Clouds Team

DevOps and cloud engineers writing practical, battle-tested guides on CI/CD, Kubernetes, infrastructure as code, and production operations — every article is based on real incidents and real pipelines, not docs-page rewrites.

More about us →

Related articles