Evaluators
Evaluators turn raw experiment outputs into comparable scores. Instead of eyeballing which variant's output looks better, you get a number per item and an average per variant — so you can pick a winner with confidence.
Why Evaluators
Without evaluators, comparing variants means reading their outputs side by side and forming a subjective impression. That works for two or three items, but falls apart fast:
- Scale: Can you really judge 100 outputs fairly?
- Consistency: Your opinion drifts as you get tired.
- Reproducibility: Your reviewer sees different nuances than you.
An evaluator applies the same scoring rule to every item, every time. The comparison matrix shows who scored highest, where variants disagree, and which items are hard for every model.
How They Work
When an experiment run finishes an item, each configured evaluator inspects the output and produces a score between 0.0 and 1.0 plus a label:
| Label | Meaning |
|---|---|
| PASS | Output meets quality threshold (typically ≥ 0.8) |
| PARTIAL | Partially correct (typically 0.5 – 0.8) |
| FAIL | Below threshold |
| SKIP | Evaluator can't score this item (e.g. missing ground truth) |
Scores appear as columns in the experiment results matrix — one column per evaluator, one row per dataset item. The best score in each row is highlighted, and variant averages are shown at the bottom.
Available Evaluators
Grounding
Checks whether the values in the extracted output actually appear in the source document. Useful for extraction experiments where hallucinations are the main risk.
A grounding score of 1.0 means every extracted value was found verbatim (or with minor normalization) in the input. Lower scores indicate the model invented values that aren't supported by the source.
No ground truth needed
Grounding doesn't require an expected output in your dataset — it only needs the source document. That makes it cheap to add to any extraction experiment.
Use when: You're testing extraction models and want to detect hallucination. Default for extraction experiments.
Exact Match
Compares the output to the expectedOutput you've defined on each dataset item, field by field. Score = fraction of expected fields that match exactly.
A score of 1.0 means every expected field is present with the expected value. 0.5 means half the fields match. SKIP means the dataset item doesn't have an expectedOutput to compare against.
Use when: Your dataset has gold-standard outputs and you want strict accuracy measurement.
LLM-as-Judge
Sends each result to a language model together with a rubric you write, and records the score the model returns. Use it for qualities a string comparison cannot measure: answer quality, factual accuracy, tone, whether a summary covers the key points.
For every result, the judge receives your rubric followed by the item's input, the variant's output and — when the dataset item has one — its expectedOutput, each wrapped in its own XML tag (<input>, <output>, <expected>). It must answer with JSON of the form { "score": 0–1, "label": "…", "explanation": "…" }. The label is whatever your rubric asks for; the default rubric asks for PASS, PARTIAL or FAIL. A score outside 0–1 is clipped into range.
An LLM-as-Judge evaluator is saved as a reusable evaluator template. Create one on the Evaluators page of the console, or from the New evaluator button next to the evaluator picker in the experiment form. Its configuration:
| Field | Required | Default | Meaning |
|---|---|---|---|
rubricPrompt | Yes | — | Instructions for the judge: what to assess and how to score it |
model | Yes | — | The judging model: a platform model name (e.g. gpt-5.1) or provider/model for one of your BYOK providers |
temperature | No | 0 | Sampling temperature. Keep 0 for reproducible scores |
maxFieldTokens | No | 4000 | Approximate token budget per field. Longer input, output or expected values are shortened in the middle, with a marker telling the judge text was left out |
timeoutSeconds | No | 60 | Time limit for one judge call |
Choosing a model: the judge only needs to read and compare, so a smaller, cheaper model is often enough for clear-cut rubrics (format, completeness, presence of a fact). Use a stronger model for rubrics that need reasoning or domain knowledge. Keep the judge model fixed across runs you want to compare: a different judge gives different scores for the same output.
Cost: every result scored by an LLM-as-Judge evaluator is one extra model call on the judging model. An experiment with 100 items and 2 variants makes 200 judge calls per run, for each judge evaluator. The calls are billed like any other request on that model; with a BYOK model they run on your own provider account.
The result is SKIP when the configuration lacks rubricPrompt or model, when the result has no output, or when the judge's reply is not valid JSON. It is ERROR when the judge call fails or runs past timeoutSeconds.
Use when: You need to score qualities that need judgment rather than an exact comparison, including outputs with no single correct answer.
Evaluator Templates
An evaluator template is a named, reusable scorer that belongs to your organization. Define it once on the Evaluators page in the console, then reference it from any number of experiments and use it in annotation queues. Editing a template changes it for every experiment that uses it from then on; each score keeps a snapshot of the config it was produced with.
A template has one of two types:
| Type | Console name | Who scores |
|---|---|---|
LLM_JUDGE | LLM as Judge | A model, automatically, during experiment runs |
HUMAN | Human review | A person, in an annotation queue or through the human-scores API |
Both types need a config object. The console doesn't let you change the type of an existing template — the config of one type means nothing to the other. An LLM_JUDGE template's config is the same fields as the LLM-as-Judge table above. To keep the judge's labels to a fixed set, add "dataType": "CATEGORICAL" and a categories list (see below) — a label outside the list is then not stored as the score's label, and the judge's original answer is kept in details.rawLabel.
Human review templates
A Human review template is the rubric a reviewer scores against. In the console you choose between:
- Categorical labels — your own list of options, for example Pass / Partial / Fail. The reviewer picks one.
- True / False — a binary judgement.
Via the API the config can also be NUMERIC with a minValue / maxValue range, or have no dataType at all ("config": {}) — a note-only rubric that the queue shows under Annotation as a free-text field. A human score is always between 0.0 and 1.0, so a numeric range can only narrow that interval: keep minValue and maxValue inside it.
The server enforces the rubric for categorical and numeric templates: a human score whose label is not one of the active categories, or whose score is outside the numeric range, is rejected. True / False templates don't restrict the label on the server — the console only offers True and False. See Human scores.
Human review templates score nothing on their own. The experiment form lists them next to the LLM judges, but an experiment run produces no score for them — pick LLM as Judge templates or built-in evaluators there.
Config shape
The typed part of the config is the same for both template types:
config
{
"dataType": "CATEGORICAL",
"categories": [
{ "value": "PASS", "label": "Pass", "numericValue": 1.0 },
{ "value": "FAIL", "label": "Fail", "numericValue": 0.0 },
{ "value": "UNCLEAR", "label": "Unclear", "archived": true }
]
}
| Field | Description |
|---|---|
dataType | CATEGORICAL, NUMERIC or BOOLEAN. Optional. |
categories | CATEGORICAL only, at least one. value is what a score stores as its label and must be unique; label is the display text; numericValue is optional. |
archived | On a category: no new scores may use it, existing scores keep it. |
minValue, maxValue | NUMERIC only, both optional; minValue must not exceed maxValue. For human scores, keep both within 0.0–1.0. |
Treat a category's value as permanent. Rename its label or archive it, but don't change the value — existing scores store it.
Managing templates via the API
Request
// POST /v1/evaluator-templates
{
"name": "Support answer quality",
"description": "Is the answer correct and polite?",
"type": "LLM_JUDGE",
"config": {
"rubricPrompt": "You grade answers of a customer support agent. Score 1 if the answer is correct, complete and polite, 0.5 if it is correct but incomplete, 0 otherwise. Label PASS, PARTIAL or FAIL.",
"model": "gpt-4.1",
"temperature": 0
}
}
Response
{
"id": "3f6c2a9e-…",
"name": "Support answer quality",
"description": "Is the answer correct and polite?",
"type": "LLM_JUDGE",
"config": {
"rubricPrompt": "You grade answers of …",
"model": "gpt-4.1",
"temperature": 0
},
"createdAt": "2026-09-27T10:00:00Z",
"lastModifiedAt": "2026-09-27T10:00:00Z"
}
PUT /v1/evaluator-templates/{id} replaces the template — send every field, not only the ones you change. An unknown id answers 404. Names are unique within the organization: a duplicate answers 409. An LLM_JUDGE template without rubricPrompt or model, or a config that breaks the shape above, answers 400.
Reading templates needs the Viewer role; creating, updating and deleting them needs Member.
| Method | Endpoint | Description |
|---|---|---|
GET | /v1/evaluator-templates | List templates, by name |
POST | /v1/evaluator-templates | Create a template |
GET | /v1/evaluator-templates/{id} | Get a template |
PUT | /v1/evaluator-templates/{id} | Replace a template |
DELETE | /v1/evaluator-templates/{id} | Delete a template |
From the CLI: backbone evaluators list, get, create, update and delete; create takes --type and --config as JSON or a file path.
Selecting Evaluators
When you create an experiment, the form includes an Evaluators multi-select where you choose which ones to run. It lists the built-in code evaluators (Grounding, Exact Match) and your organization's evaluator templates, such as your LLM-as-Judge rubrics. For extraction experiments, Grounding is selected by default — you can add others or deselect Grounding as needed. New evaluator above the picker creates a template without leaving the form.
Paths on this page are relative to https://api.2kw.ai; send your API key as Authorization: Bearer sk_your_api_key. You can also set evaluators via the API when creating or updating an experiment, by including them in the metadata field:
Request
// POST /v1/experiments
{
"name": "Invoice extraction — GPT-4.1 vs GPT-5.1",
"type": "extraction",
"datasetVersionId": "dsv_789",
"metadata": {
"evaluators": [
{ "evaluatorId": "grounding" },
{ "evaluatorId": "exact_match" },
{ "evaluatorId": "3f6c2a1e-8b4d-4c1a-9e2f-7a5b0d9c1e42" }
]
}
}
An evaluatorId is either a built-in evaluator ID (grounding, exact_match) or the ID of an evaluator template (see Evaluator Templates for how to create one). A config next to it replaces the template's config for this experiment only; scores produced that way carry overrodeTemplate: true.
Request excerpt
"evaluators": [
{ "evaluatorId": "grounding" },
{ "evaluatorId": "3f6c2a9e-…" }
]
When you trigger a run, every configured evaluator scores every result — you don't need to opt in per-run.
Reading the Matrix
On the experiment results page, evaluator scores appear as short labels below duration, tokens, and cost in each cell:
Row 1 │ GPT-4.1 │ GPT-5.1
│ dur 2.3s │ dur 1.5s
│ cost $0.0066 │ cost $0.0088
│ GRD 1.00 ◄─ │ GRD 0.87
│ EXA 0.75 │ EXA 1.00 ◄─
GRD,EXA— evaluator IDs abbreviated- Green highlight — best score in the row (higher is better)
- Amber row marker — variants produced different outputs (regardless of score)
- Avg row (bottom) — variant-level averages across all items
What Happens Next
Evaluator scores are stored alongside each result, so you can:
- Compare variants objectively in the matrix — higher average wins
- Find hard items — rows where every variant scores low are candidates for prompt tuning or better training data
- Track regressions — re-running an experiment tells you if changes helped or hurt (relative to the previous run's scores)
API
List available evaluators
Returns the built-in evaluators registered on your instance. Use this to build selection UIs or validate evaluator IDs before submission. Your organization's own evaluators are templates, listed at GET /v1/evaluator-templates.
Request
GET /v1/evaluators
Response
[
{ "id": "grounding", "displayName": "Grounding", "type": "CODE" },
{ "id": "exact_match", "displayName": "Exact Match", "type": "CODE" },
{ "id": "llm_judge", "displayName": "LLM as Judge", "type": "LLM_JUDGE" }
]
type says how an evaluator scores: CODE runs a fixed rule, LLM_JUDGE asks a model with a rubric, and HUMAN marks templates that reviewers score by hand. An LLM_JUDGE evaluator needs a rubric and a model, so select it through an evaluator template rather than by the llm_judge ID alone. The order of the list is not guaranteed.
Read scores
Run results include a scores array — one entry per evaluator that ran against that result:
Response excerpt
{
"id": "res_abc",
"output": { "total": 154.7, "vendor": "Acme" },
"scores": [
{
"evaluatorId": "grounding",
"evaluatorName": "Grounding",
"score": 0.87,
"label": "PASS",
"source": "AUTO",
"details": { /* per-field breakdown */ }
}
]
}
The details field shape depends on the evaluator — grounding returns per-field scoring, exact match returns matched vs. mismatched field lists. LLM-as-Judge returns the judging model and the judge's full judgment, and lists under truncations any field that was shortened to fit maxFieldTokens. A score from a configured evaluator also keeps a copy of that configuration under details.evaluatorConfig, so it stays traceable to its rubric after the template changes. source says where a score came from: AUTO (code evaluator), LLM_JUDGE or HUMAN (a reviewer; see Human scores).
Coming Soon
More evaluators are on the roadmap:
- Semantic Similarity — embedding-based comparison for freeform text outputs
- JSON Schema Valid — validate output structure against a schema
- Latency / Cost — score against budgets and flag outliers
- Webhook — call your own scoring service
The evaluator interface is pluggable, so custom evaluators can be added without changing the experiment engine.