Juan Pablo García
All writing
Explainer5 steps · scroll to play

How to evaluate an LLM app

LLM apps fail in quiet ways. An evaluation set turns 'it seems better' into a number you can compare between versions and check before every release.

  1. 1

    Collect real test cases

    Start with questions real users ask, each with the answer or behavior you expect. A few dozen good cases beat thousands of synthetic ones.

  2. 2

    Run the app on every case

    Run your full app, prompt, retrieval and model, on each case, exactly as users would hit it.

  3. 3

    Score each answer

    Use exact checks where you can, and an LLM judge with a clear rubric where you can't. Spot-check the judge against human labels.

  4. 4

    Compare versions

    Run the same cases on every change. Now a new prompt or model is a measured difference, not a feeling.

  5. 5

    Gate releases

    Set minimums and run evals in CI. A version that drops below them is blocked before it reaches users.

Collect real test cases

Start with questions real users ask, each with the answer or behavior you expect. A few dozen good cases beat thousands of synthetic ones.

Run the app on every case

Run your full app, prompt, retrieval and model, on each case, exactly as users would hit it.

Score each answer

Use exact checks where you can, and an LLM judge with a clear rubric where you can't. Spot-check the judge against human labels.

Compare versions

Run the same cases on every change. Now a new prompt or model is a measured difference, not a feeling.

Gate releases

Set minimums and run evals in CI. A version that drops below them is blocked before it reaches users.

In short

  • Look at failures, not just averages: each failed case usually points to a specific fix.
  • Track quality, cost and latency together. A better answer that takes three times longer may not be better.
  • Keep adding cases from real failures in production, so the set keeps up with how people use the app.
  • The cases and scores in the animation are illustrative.

Have an AI feature to build? Let's talk for 15 minutes.