← All posts

Engineering · AI

Evals: How to Test a Feature That Never Gives the Same Answer Twice

Mihajlo Petrović5 min read

You cannot unit-test an LLM feature, so most teams ship prompt changes on vibes and break things silently. How to build a golden set, assert cheaply before reaching for a judge model, and wire it all into CI.

Every engineer hits the same wall on their first LLM feature. You write the prompt, it works on your handful of examples, you ship it. A week later you tweak the wording to fix one complaint — and quietly break three things that used to work. You have no way of knowing, because there's no test suite. There can't be: the output is different every time.

This is the problem evals solve, and it's the difference between shipping an AI feature and maintaining one.


Why Normal Tests Don't Apply

expect(output).toBe("...") requires a single correct answer. LLM features have a space of acceptable answers: different wording, different ordering, different level of detail, all fine.

So evals don't assert equality. They assert properties — and they report a score rather than pass/fail on a single run. A suite at 94% that drops to 81% after a prompt change is the signal. The absolute number matters much less than the direction.


Step 1: The Golden Set

Collect 30-100 real cases. Not invented ones — real user inputs if you have them, realistic ones if you don't. Each case is the input plus what a good answer must contain.

The composition matters more than the size:

  • Common cases (~60%) — the boring path, most of your traffic
  • Edge cases (~30%) — empty input, wrong language, hostile input, ambiguity, missing data
  • Known failures (~10%) — every bug you've ever fixed, kept forever as a regression guard

Thirty well-chosen cases beat a thousand generated ones. And when a user reports a problem, the fix isn't only the prompt change — it's adding that case to the set.


Step 2: Assert Cheaply First

Reach for the expensive tooling last. A surprising share of real quality issues is caught by assertions that cost nothing to run:

  • Shape. Valid JSON? Required fields present? Enum values in range? (Structured output modes turn this from a check into a guarantee.)
  • Must contain / must not contain. A refund answer that never mentions the 14-day window is wrong regardless of how well written it is. A support answer that says "as an AI language model" is wrong, period.
  • Grounding. Every cited id/order number/section appears in the source you passed in. This catches a large fraction of hallucinations mechanically.
  • Budgets. Length, latency, tokens. Quality regressions and cost regressions both show up here.

These are deterministic, fast, and cheap enough to run on every commit.

Then: the model as judge

For "is this answer actually good", grade with a model. Two things make it work, and their absence is why people dismiss the technique:

A specific rubric. Not "rate 1-10". Ask a few binary questions: Does it answer the question asked? Is every claim supported by the provided context? Does it follow the required tone? Does it avoid advice outside the documented policy? Binary questions are far more stable across runs than a numeric score.

Validation of the judge. Hand-label 20 cases yourself, then check whether the judge agrees with you. If it doesn't, your judge is broken and every number it produces is noise. People skip this step and then trust the output — that's how you get a dashboard that measures nothing.

Use a strong model for judging. It's a harder task than the one you're grading, and this is not where to save money.


Step 3: Wire It Into the Workflow

An eval suite you run manually is an eval suite you stop running.

Locally, a fast subset (10-15 cases, deterministic checks only) before every prompt change. Seconds, not minutes.

In CI, the full suite on any change to prompts, retrieval, model version, or tool definitions. Fail the build on a meaningful drop, not on noise — run each case a few times and compare against a threshold, because a single run at temperature will wobble on its own.

In production, the things you can only measure with real users: thumbs up/down, escalation-to-human rate, retry rate, abandonment. Every thumbs-down is a candidate for the golden set. This closes the loop — offline evals tell you if a change is safe to ship, online metrics tell you what to fix next.


The Payoff Nobody Mentions

The obvious win is catching regressions. The bigger one is that evals make model upgrades a 20-minute decision.

A new model comes out. Without evals, adopting it is a leap of faith followed by weeks of anecdotes. With evals, you run the suite against both, look at the score and the cost, and decide before lunch. Same for prompt refactors, retrieval changes, cheaper models for sub-tasks, or dropping a chunk of prompt that may no longer be earning its place.

In a regulated environment this stops being a nice-to-have. If you have to demonstrate that a model change was tested before it reached customers, an eval suite with recorded results is that evidence — a point I go into further in the post on AI in regulated FinTech.


Start Smaller Than You Think

You do not need a platform. A JSON file of cases, a script that runs them, and a printed score is a complete eval system, and it's a weekend's work.

Twenty cases and a pass rate you actually look at will do more for your feature's quality than any amount of prompt tinkering by vibes. The teams that ship reliable LLM features aren't the ones with the cleverest prompts — they're the ones who can tell, in a minute, whether today's change made things better or worse.

  • #evals
  • #llm
  • #testing
  • #ci
  • #ai

Written by

Mihajlo Petrović

Software engineer in Belgrade. Builds his own products and the AI automations that keep them running.

Have a task that repeats every week?

Tell me about it. If it can be automated well, I will show you how. If it cannot, I will say that too.

Tell me what to automate