The failure: prompt change, no regression net
Someone shortens the system prompt to save tokens. Tone improves. The bot now invents SKUs that are not in the catalogue. There is no eval.spec.ts, so CI stays green. Customers see the SKUs on Monday.
12 September 2026 · OpenCrevo
An LLM evaluation framework that lives in a Colab notebook will not stop a bad prompt merge. You need one fixture, one spec, and a threshold that fails GitHub Actions the same way a unit test does. Pair it with RAG evaluation metrics (/blog/rag-evaluation-metrics-that-catch-failures/) if you retrieve documents. Score the baseline first with the QA maturity assessment (/qa-maturity/).
Someone shortens the system prompt to save tokens. Tone improves. The bot now invents SKUs that are not in the catalogue. There is no eval.spec.ts, so CI stays green. Customers see the SKUs on Monday.
[
{
"id": "sku-grounded",
"input": "What is the SKU for the travel mug?",
"must_include": ["SKU-MUG-14"],
"must_not_include": ["SKU-MUG-99"]
}
]import { test, expect } from "@playwright/test";
import cases from "./fixtures/llm-eval.json";
import { complete } from "../src/llm";
for (const c of cases) {
test(`eval ${c.id}`, async () => {
const out = await complete(c.input);
for (const s of c.must_include) expect(out).toContain(s);
for (const s of c.must_not_include) expect(out).not.toContain(s);
});
}- name: LLM eval
env:
EVAL_MODEL: gpt-4.1-mini
run: npx playwright test tests/eval.spec.ts --project=evalBook a free quality audit. We'll review your AI system, identify the highest-risk failure modes, and map a quality roadmap tailored to your stack.