How to evaluate an LLM app
LLM apps fail in quiet ways. An evaluation set turns 'it seems better' into a number you can compare between versions and check before every release.
- 1
Collect real test cases
Start with questions real users ask, each with the answer or behavior you expect. A few dozen good cases beat thousands of synthetic ones.
- 2
Run the app on every case
Run your full app, prompt, retrieval and model, on each case, exactly as users would hit it.
- 3
Score each answer
Use exact checks where you can, and an LLM judge with a clear rubric where you can't. Spot-check the judge against human labels.
- 4
Compare versions
Run the same cases on every change. Now a new prompt or model is a measured difference, not a feeling.
- 5
Gate releases
Set minimums and run evals in CI. A version that drops below them is blocked before it reaches users.
Collect real test cases
Start with questions real users ask, each with the answer or behavior you expect. A few dozen good cases beat thousands of synthetic ones.
Run the app on every case
Run your full app, prompt, retrieval and model, on each case, exactly as users would hit it.
Score each answer
Use exact checks where you can, and an LLM judge with a clear rubric where you can't. Spot-check the judge against human labels.
Compare versions
Run the same cases on every change. Now a new prompt or model is a measured difference, not a feeling.
Gate releases
Set minimums and run evals in CI. A version that drops below them is blocked before it reaches users.
In short
- Look at failures, not just averages: each failed case usually points to a specific fix.
- Track quality, cost and latency together. A better answer that takes three times longer may not be better.
- Keep adding cases from real failures in production, so the set keeps up with how people use the app.
- The cases and scores in the animation are illustrative.