Experiments
Run controlled experiments to compare AI configurations — models, schemas, agent versions, instructions and sampling parameters — against the same test data. Find the best setup before deploying to production.
How It Works
An experiment compares multiple variants (different configurations) against a dataset (versioned test data). Each variant runs against every item in the dataset, producing results you can compare side by side.
The workflow is: create dataset → create experiment → add variants → run → compare results.
Supported experiment types
The current release supports two experiment types: extraction (document → structured JSON) and agent (a dataset item → an agent's answer). To compare models, instructions or sampling settings, use an agent experiment whose variants override them. Variants of any other task type are refused with 400.
Core Concepts
Variants
A variant is one specific configuration you want to test. For extraction experiments, a variant specifies a model, a schema, and optionally a schema version (without one it uses the latest). For agent experiments, it names an agent, either a label to follow or one version to pin, and optional overrides of the model, the instructions, the iteration limit and the sampling options:
Agent variant configuration
{
"agentId": "agt_123",
"version": { "label": "production" },
"overrides": {
"model": "gpt-5-mini",
"instructions": "Answer in one sentence.",
"maxIterations": 8,
"options": { "temperature": 0.2, "top_p": 1, "max_output_tokens": 2048 }
},
"timeoutSeconds": 120
}
version carries exactly one of label and versionId. The experiment type determines what configuration fields are available. All configuration is stored as JSON — the experiment engine is type-agnostic.
Prompts are not a unit of variation
A variant does not reference a managed prompt, and the extraction prompt cannot be overridden per variant. To vary what an extraction does, vary the schema version and its extraction settings; to vary what an agent does, override its instructions or compare agent versions.
Runs
When you start an experiment, the engine creates one run per variant. Each run processes every item in the linked dataset through the variant's configuration and stores the results. Runs execute asynchronously — you can poll for progress.
Runs pin their configuration
Every run stores two copies of its variant's configuration when it starts:
- the submitted configuration, exactly as the variant had it — including open references such as a schema version left at "latest" or an agent label like
production; - the resolved configuration, in which "latest" is replaced by the schema version's id and a label by the agent version it points to at that moment.
The run executes the resolved copy, so every item of a run uses the same schema version and agent version, even if a label moves while the run is in progress. Editing the variant afterwards changes neither copy. Runs started before this behaviour was introduced have no stored configuration.
Results
Each run produces a result per dataset item, containing the AI output, duration, token usage, and estimated cost. Configured evaluators also score each result — grounding scores measure how well extracted values match the source document, and exact-match scores compare against ground-truth expectedOutput when provided.
Creating an Experiment
Create an experiment within your organization, optionally linking it to a dataset version. Paths on this page are relative to https://api.2kw.ai; send your API key as Authorization: Bearer sk_your_api_key.
Request
// POST /v1/experiments
{
"name": "Invoice Model Comparison",
"description": "Compare GPT-4 vs Claude on invoices",
"type": "extraction",
"datasetVersionId": "dsv_789"
}
Response
{
"id": "exp_456",
"name": "Invoice Model Comparison",
"status": "DRAFT",
"type": "extraction",
"datasetVersionId": "dsv_789",
"createdAt": "2026-04-08T10:10:00Z"
}
Adding Variants
Add one or more variants to compare. Each variant specifies a taskType and a JSON configuration:
Request
// POST /v1/experiments/{experimentId}/variants
{
"name": "GPT-4.1",
"taskType": "extraction",
"configuration": {
"schemaId": "sch_789",
"model": "gpt-4.1"
}
}
Response
{
"id": "var_101",
"experimentId": "exp_456",
"name": "GPT-4.1",
"taskType": "extraction",
"configuration": {
"schemaId": "sch_789",
"model": "gpt-4.1"
},
"sortOrder": 1,
"version": "0"
}
A configuration is validated when it is saved, not only when a run starts: an unknown task type, a schema or schema version that does not exist in your organization, a schema without any version, or a model your organization has not configured is refused with 400 and a message naming the problem.
Building variants in the console
The console's variant dialog covers the whole life of a variant:
- Add starts from the experiment type's defaults. Start from… copies an existing variant instead: search for any experiment of your organization, then pick one of its variants of the same task type.
- Edit (the pencil on a variant's row) changes a saved variant in place. Duplicate (the copy icon) opens a new variant pre-filled with a copy of that row.
- The Form tab shows the type's guided fields; the JSON tab edits the same configuration as raw JSON, including keys the form does not know. While the JSON does not parse to an object, Save and the Form tab are disabled and your text is kept.
- Changes shows a diff of the configuration against the variant you started from — the saved variant when editing, the copied one otherwise. Key order never counts as a change.
- An edit that changes nothing cannot be saved. Fields the form declares are dropped when empty; any other key is saved exactly as written.
Updating a variant
PUT …/variants/{variantId} takes the full variant, including the version you read. If someone saved the variant after you read it, the update is refused with 409 and nothing is overwritten; reload the variant and apply your change again. Without version, the last write wins. A variant's taskType cannot change after it is created.
Request
// PUT /v1/experiments/{experimentId}/variants/{variantId}
{
"name": "GPT-4.1, pinned schema",
"taskType": "extraction",
"configuration": {
"schemaId": "sch_789",
"schemaVersionId": "scv_12",
"model": "gpt-4.1"
},
"sortOrder": 1,
"version": "0"
}
Returns the updated variant with its new version. A stale version returns 409, a changed taskType or an invalid configuration returns 400, and a variant that does not belong to the experiment in the path returns 404.
Running an Experiment
Start the experiment to execute all variants against the dataset:
Request
// POST /v1/experiments/{experimentId}/runs
Returns 202 Accepted. Each variant gets its own run that executes asynchronously.
Response
{
"id": "run_201",
"experimentId": "exp_456",
"variantId": "var_101",
"status": "PENDING",
"itemsTotal": 0,
"itemsCompleted": 0,
"itemsFailed": 0
}
Checking Progress
Poll the runs endpoint to check status. Runs transition through: PENDING → RUNNING → COMPLETED (or FAILED).
GET /v1/experiments/{experimentId}/runs
The experiment's own status is derived from its runs — there is no separate field that can get out of sync. As runs change state, the experiment reflects the aggregate (see the lifecycle table below).
Viewing Results
Retrieve results for a specific run:
GET /v1/experiments/{experimentId}/runs/{runId}/results
Response
{
"content": [
{
"id": "res_301",
"runId": "run_201",
"datasetItemId": "dsi_101",
"output": {
"invoiceNumber": "INV-2026-001",
"vendor": "Acme Corp",
"total": 1250.00
},
"durationMs": 2340,
"inputTokens": 156,
"outputTokens": 89,
"estimatedCost": 0.004200,
"error": null
}
]
}
When a run fails on an item (e.g. the model rejects the request or the
provider returns an error), output is null and error carries a
JSON object with type and message describing the failure. The run's
itemsFailed counter increments for that result.
Comparing Results
The experiment's results page bundles a variant × item comparison matrix — the primary way to read an experiment:
- Metric-only cells — each cell shows the duration, total tokens, and cost for that variant on that item. Per-row winners (fastest, fewest tokens, cheapest) are highlighted so you can scan a dataset for performance patterns at a glance.
- Divergence detection — rows where variants produced different outputs are flagged with an amber indicator. A summary banner at the top counts how many rows diverge; rows where all variants agree stay neutral.
- Outlier highlighting — within a divergent row, cells that don't match the row's majority answer get a subtle amber tint. If there's no clear majority, every cell is flagged.
- Aggregate row — averaged duration, averaged token count, and total cost per variant across the whole dataset, with overall winners highlighted the same way.
- Side-by-side diff — clicking a row opens a diff view comparing the outputs of two variants. With two variants the pair is fixed; with three or more, dropdowns let you pick which pair to compare.
- Configuration diff — above the outputs, the same panel diffs the two variants' configurations as their runs executed them. A Resolved / Submitted toggle switches between the pinned copies and the configurations as written; a run from before configurations were stored says so.
- Edited since this run — a variant whose configuration changed after its latest run is marked in the panel and in the matrix header, so a comparison never silently describes a variant that no longer exists in that form. New variant from this run opens the variant dialog pre-filled with that run's submitted configuration.
All of this is served by one public endpoint, so your own tooling can render the same views without joining per-run results itself:
GET /v1/experiments/{experimentId}/comparison?page=0&size=20
The server picks the latest completed run of each variant and joins its results to the items of the experiment's dataset version. The response has three parts:
items— a page of dataset items, driven by the dataset version rather than by the runs. Each row carriesresultsByVariantId, a map from variant ID to that run's result for the item (output, duration, tokens, cost, error and evaluatorscores). A variant whose latest run has no result for the item is absent from the map. Soft-deleted dataset items are left out. Page with the standardpage,sizeandsortquery parameters.aggregates— one entry per variant that has a completed run, computed over all results of that run, not only the current page, so the summary row stays the same while you page. Each entry names therunIdit was computed from and itsconfigurationandresolvedConfiguration. A variant whose run produced no results still gets an entry withitemCount: 0.evaluatorIds— every evaluator that scored at least one result in these runs, including one that skipped every item. Use it for a stable set of score columns.
An experiment with no completed run, or with no dataset version, returns an empty page and empty aggregates.
Response
{
"items": {
"content": [
{
"datasetItemId": "dsi_101",
"resultsByVariantId": {
"var_101": {
"id": "res_301",
"runId": "run_201",
"datasetItemId": "dsi_101",
"output": { "total": 1250.00 },
"durationMs": 2340,
"inputTokens": 156,
"outputTokens": 89,
"estimatedCost": 0.004200,
"error": null,
"scores": [ /* ... */ ]
},
"var_102": { /* ... */ }
}
}
],
"number": 0,
"size": 20,
"totalElements": 48,
"totalPages": 3
},
"aggregates": [
{
"variantId": "var_101",
"variantName": "GPT-4.1",
"avgDurationMs": 2210.5,
"avgTokens": 241.0,
"totalCost": 0.201600,
"itemCount": 48,
"avgEvaluatorScores": { "ev_exact": 0.92 },
"skippedEvaluatorCounts": {},
"runId": "run_201",
"runCompletedAt": "2026-04-08T10:24:00Z",
"configuration": { "schemaId": "sch_789", "model": "gpt-4.1" },
"resolvedConfiguration": { /* ... */ }
}
],
"evaluatorIds": ["ev_exact"]
}
Use GET …/runs/{runId}/results instead when you need every result of one specific run, including a run that is not its variant's latest.
Regression Detection
Iterating on a pipeline means tweaking one thing and hoping nothing else quietly broke — a prompt change that improves invoices might break receipts, a model upgrade that helps accuracy might make latency regress. Regression detection compares any run against a reference and tells you exactly which dataset items improved, which regressed, and by how much.
Score-driven, modality-agnostic
Regression detection operates on evaluator scores, not raw outputs. As long as your variants have evaluators attached, it works the same for extraction and agent experiments.
Baselines
The reference a run is compared against is called the baseline. Baselines are scoped per variant — each variant compares its own runs against its own reference — because re-running a single variant (see POST …/runs?variantId=…) is a first-class operation, and forcing a cross-variant baseline would fight that workflow.
Every regression comparison resolves a baseline in this order:
- Explicit — a
baselineRunIdquery parameter on the request. Overrides everything. - Marked baseline — if any completed run on the same variant has been flagged via
PUT …/runs/{runId}/baseline, it's used. - Prior run (auto) — the most recent older completed run on the same variant. Filtered strictly by
createdAtso looking back from an older run never picks a newer one. - None — if nothing qualifies, no report is produced. The UI shows "No prior run" in place of counters.
The UI's regression panel makes the resolved source explicit (the small vs … label on each row) so users never have to guess what a delta number is measured against.
Mark a run as its variant's baseline:
PUT /v1/experiments/{experimentId}/runs/{runId}/baseline
Clearing any prior baseline on the same variant is atomic — the partial-unique index on experiment_run(variant_id) WHERE baseline = true guarantees only one baseline per variant at any time.
Response
{
"id": "run_201",
"variantId": "var_101",
"status": "COMPLETED",
"baseline": true
}
The Report
Fetch a regression report for any completed run:
GET /v1/experiments/{experimentId}/runs/{runId}/regression
Optional query params:
baselineRunId— compare against an explicit run instead of the resolved defaultthreshold— the absolute score delta above which an item is considered IMPROVED or REGRESSED. Default0.05.
Response
{
"runId": "run_current",
"baselineRunId": "run_baseline",
"baselineRunCreatedAt": "2026-04-14T09:14:00Z",
"baselineSource": "MARKED_BASELINE",
"threshold": 0.05,
"summary": {
"comparedItems": 12,
"improved": 6,
"regressed": 3,
"unchanged": 3,
"gtChanged": 1,
"baselineMean": 0.725,
"currentMean": 0.810,
"meanDelta": 0.085,
"netDelta": 1.020
},
"regressed": [
{
"datasetItemId": "item-5",
"baselineScore": 0.90,
"currentScore": 0.40,
"delta": -0.50,
"classification": "REGRESSED",
"reScoredFromBaseline": false
}
],
"improved": [ /* ... */ ],
"gtChanged": [
{
"datasetItemId": "item-9",
"baselineScore": 1.00,
"currentScore": 0.50,
"delta": -0.50,
"classification": "GT_CHANGED",
"reScoredFromBaseline": false
}
]
}
For each dataset item that exists in both runs, each side's score is the mean over the evaluators that scored the item in both runs with a numeric, non-SKIP score, and delta = currentScore − baselineScore. Each item gets one classification:
| Condition | Classification |
|---|---|
| Ground truth changed since the baseline and could not be re-scored | GT_CHANGED |
delta > +threshold | IMPROVED |
delta < −threshold | REGRESSED |
| otherwise | UNCHANGED |
Items that exist in only one of the two runs (e.g., the dataset version changed) are excluded from the comparison — there's no sensible numeric delta to report.
Ground-truth changes
Every stored score records the expectedOutput it was computed against. When an item's expectedOutput was edited after the baseline run, the baseline and current scores were measured against different ground truth, and their delta would mix model drift with the edit. For each evaluator that depends on expectedOutput, the report then re-scores the baseline run's stored output against the current ground truth before computing the delta:
- Re-scoring succeeds — the item is classified as usual and its
reScoredFromBaselineistrue. The delta compares both runs against the same ground truth. - Re-scoring fails (the evaluator is unknown, the baseline output is missing, or the evaluator errors or returns
SKIP) — the item is classifiedGT_CHANGEDand listed ingtChanged. Its delta is shown for reference only.
GT_CHANGED items are counted in summary.gtChanged and left out of every other number: comparedItems, improved, regressed, unchanged, the means and netDelta. An edited ground truth therefore never shows up as a regression or an improvement.
Reading the UI
The experiment detail page surfaces regression inline in the Runs table: every completed non-baseline row shows three counters (N improved · N regressed · N unchanged) plus a "Details ›" affordance. Clicking opens a side panel with:
- A verdict bar — stacked proportional segments (emerald / destructive / muted) so the shape of the change is visible at a glance.
- A hero Mean Δ — the single most telling number, toned by sign.
- Supporting metrics — Net Δ, baseline mean, current mean.
- Regressed and Improved sections — each item on its own line with a dumbbell plot (baseline dot → current dot on a 0–1 axis) and the signed delta. A refresh icon marks an item whose baseline was re-scored against the current ground truth.
- A Ground truth changed section, when any item is
GT_CHANGED— listed below the others, with its delta for reference only.
Clicking any item row deep-links into the experiment's Results page with the dataset item pre-selected, so you can see the actual input and per-variant outputs side-by-side in the existing diff view.
Experiment Lifecycle
The experiment's status field is derived from its runs at read
time, so it can't drift out of sync with the runs themselves:
| Status | Meaning |
|---|---|
DRAFT | No runs yet |
RUNNING | At least one run is PENDING or RUNNING |
COMPLETED | All runs finished successfully |
PARTIAL_SUCCESS | At least one run succeeded and at least one failed |
FAILED | All runs failed |
API Reference
| Method | Endpoint | Description |
|---|---|---|
POST | /v1/experiments | Create an experiment |
GET | /v1/experiments | List experiments (paginated) |
GET | /v1/experiments/{id} | Get an experiment |
PUT | /v1/experiments/{id} | Update an experiment |
DELETE | /v1/experiments/{id} | Delete an experiment |
POST | .../experiments/{id}/variants | Add a variant |
PUT | .../experiments/{id}/variants/{variantId} | Update a variant |
DELETE | .../experiments/{id}/variants/{variantId} | Delete a variant |
GET | .../experiments/{id}/variants | List variants |
POST | .../experiments/{id}/runs | Start experiment runs |
GET | .../experiments/{id}/runs | List runs (paginated) |
GET | .../experiments/{id}/runs/{runId} | Get run details |
GET | .../experiments/{id}/runs/{runId}/results | List run results (paginated) |
GET | .../experiments/{id}/comparison | Get the comparison matrix: latest completed run per variant (paginated items, unpaged aggregates) |
PUT | .../experiments/{id}/runs/{runId}/baseline | Mark run as its variant's baseline |
DELETE | .../experiments/{id}/runs/{runId}/baseline | Clear the baseline flag |
GET | .../experiments/{id}/runs/{runId}/regression | Get regression report for a run |