AI 4 min read
LLM evaluation
Also known as: LLM evals, LLM eval, model evaluation, language model evaluation
Definition
LLM evaluation is the systematic testing of a language model's outputs on a fixed set of inputs against agreed acceptance criteria, using automatic checks, model-based grading and human review. It decides whether a model, prompt or version is good enough to ship.
Cite this entry
Text
"LLM evaluation". Order Group, Software glossary, 10 October 2026. https://ordergroup.co/glossary/llm-evaluation/
HTML
<a href="https://ordergroup.co/glossary/llm-evaluation/">LLM evaluation</a> - Order Group
How LLM evaluation works
An evaluation, often shortened to eval, has four parts. The test set is a list of inputs that represent real use, including the hard and unusual cases. The success criteria say what a good output is for this task, specifically enough to be measured. Graders score each output against the criteria. A run report compares the scores with the previous version.
Model makers describe the same structure in their documentation. Anthropic's guide asks for success criteria that are specific and measurable, and for test cases that mirror the real distribution of tasks, edge cases included. It recommends automated grading where possible, by code or by a language model, prefers more automatically graded cases over fewer cases graded by hand, and advises using a different model as the grader than the one that produced the output. OpenAI's evals guide builds evals from a dataset of inputs and testing criteria in the same way.
Model-based grading, also called LLM-as-a-judge, lets one model score another model's outputs against a rubric. Zheng et al. (2023) found that strong judge models such as GPT-4 reached over 80% agreement with human preferences, the same level of agreement as between humans. They also named the judges' biases: position (favoring the answer shown first), verbosity (favoring longer answers) and self-enhancement (favoring answers written by the judge model itself). A judge is therefore calibrated on a sample that people have scored.
Public benchmarks measure general ability on someone else's tasks. They help to shortlist models, but they do not say whether a model simplifies your documents correctly or answers your customers' questions. That needs a test set built from your own data.
Evaluation also continues after release. Before deployment it runs offline on the fixed test set. In production the system samples real outputs, collects feedback from users, and turns every confirmed failure into a new test case.
What LLM evaluation means for your software
For a buyer, the evaluation set is part of the deliverable, much like automated tests are part of ordinary software. Requirements worth writing into the specification:
- The test set is versioned in the repository, built from real inputs (anonymized where they contain personal data) and covers each type of document or question the system handles.
- Acceptance criteria are written down before the build and are measurable: rule checks, thresholds for scores, and the share of outputs a human reviewer accepts.
- Every change of model, quantization, system prompt or inference server version reruns the full evaluation and is compared with the current baseline. A regression blocks the release.
- Automatic metrics filter quickly but do not decide alone. A readability score can improve while the meaning gets lost.
- Human review uses a written rubric, and for a specialized audience, people from that audience take part.
- Outputs are not deterministic, so each case is run several times and results are reported with their spread; a temperature of 0 reduces the variation but does not remove it.
- Production logging keeps enough to rebuild a failed case, within the retention and anonymization rules of the system.
| Method | Good for | Weak at | Cost |
|---|---|---|---|
| Code-based checks | Format, length, required fields, forbidden words | Meaning and tone | Lowest; runs on every change |
| Text metrics (readability, overlap with a reference) | Trends between versions | Whether the content is correct | Low |
| Model-based grading | Rubric scoring at scale | Bias toward position, length and its own outputs; needs calibration | Medium |
| Human review | Meaning, harm, fitness for the audience | Scale and consistency between reviewers | Highest |
Rules and standards
The NIST AI Risk Management Framework (AI RMF 1.0, January 2023) has four functions: Govern, Map, Measure and Manage. Measure covers testing, evaluation, verification and validation, before deployment and regularly in operation. NIST's Generative AI Profile (NIST AI 600-1, July 2024) adds risks specific to generative models, including confabulation, the term NIST uses for hallucinations, together with suggested actions for testing them. In the EU, the AI Act (Regulation 2024/1689) requires high-risk AI systems to be tested against prior defined metrics and probabilistic thresholds, in any event before they are placed on the market or put into service (Article 9(8)). Article 15 requires an appropriate level of accuracy maintained throughout the lifecycle, with the accuracy levels and metrics declared in the instructions for use. After the 2026 amendment (Regulation (EU) 2026/1744), these obligations apply to the stand-alone high-risk systems listed in Annex III from December 2, 2027. Most text features in products are not high-risk, but these requirements are a reasonable model for an acceptance procedure.
From our projects
In Generator ETR, the app we are building for PSONI, a Polish association for people with intellectual disabilities, to simplify documents into easy-to-read Polish, the first evaluation pipeline was delivered in June 2026. It included the first system prompt for Gemma 4, which instructs the model to follow the easy-to-read rules, a Python script that scores the difficulty of each output with the Gunning Fog Index, and an aggregated Markdown report from runs on synthetic samples. Gunning Fog was designed for English, so on Polish text it is useful as a trend between versions rather than as a verdict.
The first round of model fixes, described in the easy-to-read entry, targeted failures that a single readability score would not catch. For reading aloud, the text normalization step is checked against a fixed sample text that covers numerals, ordinal numbers, decades, Roman numerals, ranges and negative numbers, and the known limits are listed next to the fixes.
Sources
- Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 - NIST
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1 - NIST
- Define success criteria and build evaluations - Anthropic
- Evals guide - OpenAI
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena - Zheng et al., arXiv
- Regulation (EU) 2024/1689 (AI Act) - EUR-Lex
- Regulation (EU) 2026/1744 amending Regulation (EU) 2024/1689 (Digital Omnibus on AI) - EUR-Lex
FAQ
-
Large enough to cover every type of input and every failure you already know about. Many teams start with a few dozen well-chosen cases and add one with each confirmed production bug. More automatically graded cases give a steadier signal than a handful graded by hand.
-
Only to make a shortlist. Benchmarks measure general tasks, and the final choice needs your own test set and acceptance criteria.
-
Yes, with a written rubric and calibration against human scores on a sample. Judge models favor longer answers, the first answer shown and their own outputs, so the rubric and the order of answers should account for that.
-
A sample of real outputs, user feedback and every reported failure, which becomes a new test case. Any change of model, prompt, inference server or its configuration reruns the full set before release.
Building a system that depends on LLM evaluation?
See how we build software for this domain, with case studies and the stack we use.