Skip to main content
LLM evaluation frameworkCI/CDAI quality engineering

LLM Evaluation Framework You Can Run in CI Tonight

12 September 2026 · OpenCrevo

An LLM evaluation framework that lives in a Colab notebook will not stop a bad prompt merge. You need one fixture, one spec, and a threshold that fails GitHub Actions the same way a unit test does. Pair it with RAG evaluation metrics (/blog/rag-evaluation-metrics-that-catch-failures/) if you retrieve documents. Score the baseline first with the QA maturity assessment (/qa-maturity/).

The failure: prompt change, no regression net

Someone shortens the system prompt to save tokens. Tone improves. The bot now invents SKUs that are not in the catalogue. There is no eval.spec.ts, so CI stays green. Customers see the SKUs on Monday.

Minimal LLM evaluation framework

[
  {
    "id": "sku-grounded",
    "input": "What is the SKU for the travel mug?",
    "must_include": ["SKU-MUG-14"],
    "must_not_include": ["SKU-MUG-99"]
  }
]
import { test, expect } from "@playwright/test";
import cases from "./fixtures/llm-eval.json";
import { complete } from "../src/llm";

for (const c of cases) {
  test(`eval ${c.id}`, async () => {
    const out = await complete(c.input);
    for (const s of c.must_include) expect(out).toContain(s);
    for (const s of c.must_not_include) expect(out).not.toContain(s);
  });
}
- name: LLM eval
  env:
    EVAL_MODEL: gpt-4.1-mini
  run: npx playwright test tests/eval.spec.ts --project=eval
  • Pin the eval model. Do not let the eval model float with production.
  • Store fixtures in git. A spreadsheet export is not a framework.
  • Fail closed: missing API key should fail the job, not skip it.
  • When the suite grows, fold it into AI quality engineering services so evals stay on the same production line as UI tests.

Pass, fail, cleanup

  • Pass: Deleting SKU-MUG-14 from the prompt or catalogue fails eval sku-grounded on the PR.
  • Fail: The job is allowed to skip when the key is missing, or evals run only on main after merge.
  • Cleanup: Remove notebook screenshots from the wiki once the spec is green in CI.
START YOUR QUALITY JOURNEY

Your next chapter starts with a conversation.

Book a free quality audit. We'll review your AI system, identify the highest-risk failure modes, and map a quality roadmap tailored to your stack.