Data Extraction
Define a schema once, then extract structured data from any text —consistently, every time.
Snippet setup
The Python and TypeScript snippets on this page assume BASE_URL = "https://api.2kw.ai" and your API key. Python: import requests, api_key = "sk_your_api_key" and headers = {"Authorization": f"Bearer {api_key}"}. TypeScript: const BASE_URL = "https://api.2kw.ai" and const apiKey = "sk_your_api_key".
How It Works
2kw.ai takes a standard JSON Schema definition and uses it to extract structured output from unstructured text. You define the fields, types, and constraints — 2kw.ai returns validated, schema-conformant JSON.
The workflow is: define schema → commit version → send text → receive structured output.
Schemas use JSON Schema
If you've used JSON Schema before, you already know the format. If not —it's just a way to describe what fields you expect, their types, and which ones are required.
Schema Management
Before you can extract anything, you need a schema. Schemas are organization-scoped and support full version control.
Creating a Schema
A schema takes two calls: create the schema with a name and an optional description, then commit its JSON Schema definition as the first version.
POST /v1/schemas
curl -X POST https://api.2kw.ai/v1/schemas \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk_your_api_key" \
-d '{
"name": "invoice",
"description": "Supplier invoices"
}'
The response carries the schema's id. The JSON Schema is not part of this call: the API reference lists a content field on the schema, but a value sent there is not stored.
POST /v1/schemas/{schemaId}/versions
{
"jsonSchema": {
"type": "object",
"properties": {
"company": { "type": "string" },
"invoice_number": { "type": "string" },
"total": { "type": "number" },
"line_items": {
"type": "array",
"items": {
"type": "object",
"properties": {
"description": { "type": "string" },
"quantity": { "type": "integer" },
"price": { "type": "number" }
}
}
}
},
"required": ["company", "invoice_number"]
},
"changeDescription": "First version"
}
| Field | Type | Required | Description |
|---|---|---|---|
jsonSchema | object | Yes | The JSON Schema definition. One that is not a valid JSON Schema is refused with 400 |
changeDescription | string | No | What changed in this version |
The response (201) is the new version with its id, versionNumber and jsonSchema, and the latest label now points to it. Every later change is another call to this endpoint.
Version Control
Every time you change your schema, you create a new version. Versions are immutable —this gives you a full audit trail of every change.
You can:
- Commit new versions with a change description
- Deactivate versions you don't want used anymore
- Re-activate a historical version (creates a new version as a copy, preserving the audit trail)
- Pin extractions to a specific version for reproducibility
Each new version auto-increments the version number and updates the latest label.
Re-activation creates a copy
Re-activating version 3 doesn't revert —it creates version N+1 with the same content. The original stays untouched.
Labels
Labels are named pointers to specific schema versions. They decouple your application code from version numbers, enabling controlled rollouts.
How Labels Work
Every schema gets a latest label automatically, updated on each new version. You create custom labels for deployment stages:
| Label | Points to | Purpose |
|---|---|---|
latest | v5 (auto) | Always the newest version |
production | v3 | What your API consumers use |
staging | v5 | What you're testing |
Your application resolves schemas by ID + label —when you're ready to promote, just move the label pointer.
Managing Labels
POST /v1/schemas/{schemaId}/labels
Create a label:
{
"name": "production",
"schemaVersionId": "version-uuid"
}
Update a label to point to a different version:
PUT /v1/schemas/{schemaId}/labels/{labelName}
{
"schemaVersionId": "new-version-uuid"
}
Label names must be lowercase alphanumeric with hyphens (e.g., production, staging, v2-rollback). The system latest label cannot be deleted.
Resolving Schemas
GET /v1/schemas/{schemaId}/resolve
Resolve a schema by ID and optional label to get the JSON Schema definition for a specific version:
Request
curl "https://api.2kw.ai/v1/schemas/{schemaId}/resolve?label=production" \
-H "Authorization: Bearer sk_your_api_key"
| Parameter | Type | Required | Description |
|---|---|---|---|
label | string | No | Label to resolve (defaults to latest) |
Response:
{
"schemaId": "uuid",
"schemaName": "invoice-schema",
"versionId": "uuid",
"versionNumber": 3,
"jsonSchema": {
"type": "object",
"properties": {
"company": { "type": "string" },
"total": { "type": "number" }
}
},
"label": "production"
}
Validate Before Committing
POST /v1/schemas/{schemaId}/validate
Check if a schema is valid before you commit it. Returns structured errors and warnings:
{
"jsonSchema": {
"type": "object",
"properties": {
"name": { "type": "string" }
}
}
}
The response has three fields: valid, errors and warnings. Errors mean the schema document itself cannot be used. The same checks run when you commit a version, so POST /v1/schemas/{schemaId}/versions answers 400 for a document that fails them. The typical errors are:
JSON Schema must be a JSON objectJSON Schema must contain at least one schema keyword (e.g., type, properties, $schema)Invalid JSON Schema: <detail from the validator>
Check the JSON syntax, make sure the document is an object with at least one JSON Schema keyword, and validate it against the JSON Schema 2020-12 specification.
Errors are about the schema, never about extraction output. An extraction whose output does not match its schema does not fail: there is no failed status and no retry. The mismatches are recorded on the result instead, see Schema Validation.
Schema-safety warnings
Warnings are advisory and never block saving a version: valid stays true. The validate endpoint returns them on request, and the schema editor in the console lints live while you type. Each warning names a stable rule id, a JSON pointer to the exact location in your schema, a message and, where one applies, a suggested fix. The editor offers a one-click fix for the first three rules.
| Rule | Trigger | Why it is risky | Fix |
|---|---|---|---|
required-non-nullable | A required field declares a type that does not include null | The model must emit some value even when the document has none | Allow null ("type": ["string", "null"]) or remove the field from required |
required-enum | A required field has an enum that contains neither "insufficient_evidence" nor null | The model must pick a listed value even when the document contains none | Add "insufficient_evidence" or null to the enum, or remove the field from required |
min-items | minItems greater than 0 on any array | Forces the model to invent entries for documents that have none | Remove minItems and check cardinality after extraction instead, for example with a business rule |
provider-rewrite | The schema uses constraints your provider does not enforce | Constraints look enforced but are ignored at generation time, or the schema is rejected | Informational, no automatic fix. See Provider enforcement below |
A required field that carries an enum is reported as required-enum, not required-non-nullable. A field without a type keyword, or one whose anyOf/oneOf includes a {"type": "null"} branch, is already nullable and is not flagged.
The pointer targets the property for required-non-nullable and required-enum, and the offending keyword for min-items and provider-rewrite. A provider-rewrite warning about prompt-only providers uses the empty pointer, which means the whole schema.
Example. This schema requires two fields, one of them an enum without an escape value:
{
"jsonSchema": {
"type": "object",
"properties": {
"invoice_number": { "type": "string" },
"status": { "type": "string", "enum": ["paid", "unpaid"] }
},
"required": ["invoice_number", "status"]
}
}
With only the platform's default provider in play, the response is:
{
"valid": true,
"errors": [],
"warnings": [
{
"rule": "required-non-nullable",
"pointer": "/properties/invoice_number",
"message": "Required field 'invoice_number' cannot be null. When the document does not contain this value, models fabricate one rather than fail.",
"suggestion": "Allow null via \"type\": [\"string\", \"null\"], or remove 'invoice_number' from \"required\"."
},
{
"rule": "required-enum",
"pointer": "/properties/status",
"message": "Required enum field 'status' forces the model to pick a listed value even when the document contains none — the highest-risk fabrication construct.",
"suggestion": "Add an escape value \"insufficient_evidence\" to the enum, or remove 'status' from \"required\"."
}
]
}
An organization that has an Anthropic or Mistral provider configured also gets a provider-rewrite warning at the empty pointer.
Escape values
Make fields optional and nullable by default. Require a field only when a document without it should count as unprocessable. For enums, include the escape value insufficient_evidence so the model has a legal way to say "not in the document". A null entry in the enum counts as an escape too. An escape value reads as a deliberate absence; a fabricated value looks exactly like a real one.
Provider enforcement
Extraction runs on the provider behind the model you choose, and providers differ in what they enforce while generating:
- Anthropic and Mistral receive the schema as prompt text, without constrained decoding. Required fields, enums and every other constraint are advisory.
- OpenAI, Azure OpenAI and xAI receive the schema as a structured-output
json_schema. Constraint keywords outside the enforced subset are ignored, and Azure OpenAI may reject a schema that uses them. The flagged keywords arepattern,format,minLength,maxLength,minimum,maximum,exclusiveMinimum,exclusiveMaximum,multipleOf,maxItems,uniqueItems,minPropertiesandmaxProperties.minItemsis not enforced either; it is reported by themin-itemsrule, which carries the fix. - Vertex AI and Ollama are not checked, so they produce no
provider-rewritewarnings.
The linter checks your schema against every provider configured for your organization, plus Azure OpenAI, which serves the platform's built-in models.
2kw.ai does not enforce these keywords after extraction either. It does check every output against its schema and records the violations, for enforced and advisory keywords alike, see Schema Validation.
Test Against Sample Text
POST /v1/schemas/{schemaId}/test
Run an extraction against sample text without persisting the result. Great for iterating on your schema:
Request
curl -X POST https://api.2kw.ai/v1/schemas/{schemaId}/test \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk_your_api_key" \
-d '{
"jsonSchema": {
"type": "object",
"properties": {
"name": { "type": "string" },
"email": { "type": "string" }
}
},
"sampleText": "Contact Jane Smith at jane@acme.com",
"model": "gpt-5.1"
}'
The response includes token usage and processing time so you can estimate costs:
{
"success": true,
"extractedData": { "name": "Jane Smith", "email": "jane@acme.com" },
"inputTokens": 85,
"outputTokens": 22,
"processingDurationMs": 890,
"modelUsed": "gpt-5.1"
}
All Schema Endpoints
| Method | Endpoint | What it does |
|---|---|---|
GET | /v1/schemas?search=invoice&page=0&size=20 | List schemas, paginated, optionally filtered by search |
POST | /v1/schemas | Create a schema (name, description) |
GET | /v1/schemas/{id} | Get a schema |
PUT | /v1/schemas/{id} | Update a schema's name or description |
DELETE | /v1/schemas/{id} | Delete a schema and all its versions. Needs the admin role |
POST | /v1/schemas/{schemaId}/versions | Commit a new version (jsonSchema, changeDescription) |
GET | /v1/schemas/{schemaId}/versions?page=0&size=20 | List versions, paginated |
GET | /v1/schemas/{schemaId}/versions/{versionId} | Get one version |
GET | /v1/schemas/{schemaId}/versions/latest | The latest version; 404 while there is none. Deprecated: use resolve |
PUT | /v1/schemas/{schemaId}/versions/{versionId}/activate | Re-activate a historical version as a new copy |
DELETE | /v1/schemas/{schemaId}/versions/{versionId} | Deactivate a version (soft delete); answers 204. Needs the admin role |
GET | /v1/schemas/{schemaId}/labels | List labels |
POST | /v1/schemas/{schemaId}/labels | Create a label (name, schemaVersionId) |
PUT | /v1/schemas/{schemaId}/labels/{labelName} | Point a label at another version (schemaVersionId) |
DELETE | /v1/schemas/{schemaId}/labels/{labelName} | Delete a label; latest cannot be deleted. Needs the admin role |
GET | /v1/schemas/{schemaId}/resolve?label=production | Resolve a label (default latest) to its version |
POST | /v1/schemas/{schemaId}/validate | Validate a JSON Schema without saving it |
POST | /v1/schemas/{schemaId}/test | Extract from sample text without saving the result |
Running Extractions
Synchronous
POST /v1/extractions
The standard extraction endpoint. Send text and, for inputs that fit a synchronous run, get the structured data back in the same response. Larger inputs are handed to background processing automatically; see the responses below the example.
Request body:
| Field | Type | Required | Description |
|---|---|---|---|
schemaId | string | Yes | Schema to use for extraction |
schemaVersionId | string | No | Pin to a specific version (defaults to latest active) |
inputText | string | Conditional | Text to extract data from (required if no inputImages) |
inputImages | array | Conditional | Base64-encoded images to extract from, max 10 (required if no inputText) |
model | string | Yes | Platform model name (e.g., gpt-5.1) or provider/model for BYOK |
Request
curl -X POST https://api.2kw.ai/v1/extractions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk_your_api_key" \
-d '{
"schemaId": "your-schema-id",
"inputText": "Invoice #INV-2024-001 from Acme Corp. Total: $1,250.00",
"model": "gpt-5.1"
}'
Responses. Always read status before you read the result:
| HTTP status | status | Meaning |
|---|---|---|
201 Created | COMPLETED | The extraction ran; result.structuredOutput holds the data |
201 Created | FAILED | The extraction ran and failed; errorMessage says why |
202 Accepted | PENDING | The input was too large for a synchronous run and was queued in the background. The response carries a Location header and Retry-After: 5; poll it as described under Asynchronous |
A failed extraction is still a 201, so checking the HTTP status alone is not enough:
const response = await fetch(`${BASE_URL}/v1/extractions`, {
method: "POST",
headers: {
"Content-Type": "application/json",
"Authorization": `Bearer ${apiKey}`,
},
body: JSON.stringify({ schemaId, inputText, model }),
});
if (!response.ok) {
throw new Error(`Request rejected: ${response.status}`);
}
const data = await response.json();
if (data.status === "COMPLETED") {
console.log(data.result.structuredOutput);
} else if (data.status === "FAILED") {
throw new Error(data.errorMessage ?? "Extraction failed");
} else {
// PENDING: poll GET /v1/extractions/{id} until COMPLETED or FAILED
console.log("Queued, poll", response.headers.get("Location"));
}
Extracting from Images
You can also extract structured data from images (up to 10 per request). Each image must be base64-encoded with its MIME type:
Image extraction
curl -X POST https://api.2kw.ai/v1/extractions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk_your_api_key" \
-d '{
"schemaId": "your-schema-id",
"inputImages": [
{
"data": "iVBORw0KGgoAAAANSUhEUg...",
"mimeType": "image/png"
}
],
"model": "gpt-5.1"
}'
Supported image MIME types: image/png, image/jpeg, image/gif, image/webp. You can combine inputText and inputImages in a single request.
Asynchronous
POST /v1/extractions/async
For large inputs, submit an extraction for background processing. You get back a 202 Accepted with a Location header to poll:
HTTP/1.1 202 Accepted
Location: /v1/extractions/{id}
Retry-After: 5
Poll the Location URL until status changes from PENDING/PROCESSING to COMPLETED or FAILED.
Estimate Tokens First
POST /v1/extractions/estimate
Before running an extraction, you can estimate the cost. Send the same schema and input text you'd use for a real extraction:
Request body:
| Field | Type | Required | Description |
|---|---|---|---|
schemaId | string | Yes | Schema to estimate for |
schemaVersionId | string | No | Pin to a specific version (defaults to latest active) |
inputText | string | Yes | Text to estimate token usage for |
Response:
{
"inputTokens": 4200,
"estimatedOutputTokens": 350,
"strategy": "SINGLE_SHOT"
}
The strategy field tells you which processing path will be used (SINGLE_SHOT, CHUNKED, or ASYNC).
Extraction Strategies
2kw.ai automatically picks the best strategy based on input size:
| Strategy | Token Range | What Happens |
|---|---|---|
| Single-shot | Under 50K tokens | Entire input processed in one LLM call |
| Chunked | 50K - 100K tokens | Input split into overlapping chunks, results merged and validated |
| Async | Over 100K tokens | Background processing with polling |
Chunking details
Chunks are max 8,000 tokens with 200 tokens of overlap. Each chunk receives context about what was already extracted from previous chunks (carry-forward), so the model can extend partial items rather than re-extract them. Results are automatically merged, deduplicated, and validated.
Quality Scoring
Every text extraction result includes quality scores that help you assess how trustworthy the extracted data is. These scores are purely informational — structuredOutput is never modified.
Grounding vs Confidence
2kw.ai distinguishes between two types of quality signals:
Grounding measures whether an extracted value has evidence in the source text. If your schema extracts a sourceFile field and the value "gear.geo" appears in the input, that's grounded. If "phantom.geo" doesn't appear anywhere in the input, it was likely hallucinated by the model. Grounding is deterministic, free (no extra LLM calls), and available on every extraction.
Confidence combines the grounding score with the other evidence 2kw.ai already has about a value: the schema violations recorded for it and the verdicts of your business rules. It is arithmetic over those signals. No model is asked how sure it is, no extra call is made, and the same result always produces the same score. See Composite Confidence for how the number is built.
The scoring.methods array tells you which methods ran: grounding, and confidence when the composite was computed.
Scores never modify output
Grounding and confidence scores are metadata only. structuredOutput always contains the complete extraction result, exactly as the model produced it (after deduplication for chunked extractions). Your application decides how to use the scores — flag low-scoring items, highlight them in the UI, or filter them client-side.
How Grounding Works
For each extracted field, 2kw.ai searches the original input text for the extracted value:
| Basis | Score | Meaning |
|---|---|---|
exact_match | 1.0 | Value found verbatim in the source text |
normalized_match | 0.9 | Found after normalizing separators (underscores, hyphens, spaces) |
partial_match | 0.8 | Substring or filename stem found (e.g. gear matches gear.geo) |
unverifiable | 0.5 | Numbers, booleans, or strings shorter than 3 characters — can't be meaningfully searched |
not_found | 0.0 – 0.1 | String value not found in source text. Key fields score 0.0 (strong hallucination signal), regular fields score 0.1. |
Per-item scores are weighted averages of their fields, with key fields (identifiers like filenames, IDs, names) weighted 3x higher than regular fields. This means a hallucinated filename dominates the item score even if other fields look fine. Unverifiable fields (numbers, booleans, short strings) are excluded from this average — they don't drag the score down or inflate it.
Composite Confidence
Every text extraction that is scored carries a composite confidence per scored field: one number in [0, 1], built from the field's grounding score, its schema violations and its business-rule verdicts. Like the other quality signals it is recorded only. Nothing is blocked, no status changes and nothing is retried.
The scores are not calibrated
resultMetadata.scoring.confidence.calibrated is false, and it is the most important field in this section. The numbers are an ordering, not a probability: 0.7 does not mean "70% likely to be correct". The constants below have not been fitted against labelled outcomes. Use the scores to rank fields against each other within one schema version, derive any threshold from your own observed results, and do not show a confidence number to an end user as a probability.
How the score is computed
The grounding score is the base. Negative evidence attributed to the field either deducts from it or caps it. Nothing raises it: a rule that passed adds to the field's provenance, never to its score.
| Signal attributed to the field | Effect | Amount |
|---|---|---|
| One or more schema violations | caps | 0.2 |
A failed rule with automationBlocker: true | caps | 0.1 |
A failed rule with severity: ERROR | caps | 0.3 |
A failed rule with severity: WARNING | deducts, once per failure | 0.2 |
A failed rule with severity: INFO | deducts, once per failure | 0.05 |
| A rule that passed | no change to the score, adds rule provenance | — |
A rule outcome with verdict SKIPPED or ERROR | nothing | — |
base = grounding score
penalty = sum of the deductions above
cap = the smallest applicable cap, if any
score = clamp(min(round(base − penalty, 2 decimals), cap), 0, 1)
severity and automationBlocker act independently. A blocking WARNING failure deducts 0.2 and caps at 0.1; when several caps apply, the smallest wins. Deductions add up without limit (three failed WARNING rules deduct 0.6) and the result is clamped at 0. Caps beat deductions because the two meet in a min. The SKIPPED and ERROR verdicts do not count: the input was absent or the rule could not run, and neither says anything about the value. Scores are rounded to two decimals, the same grid as the grounding scores.
The basis
confidence.basis names the one cause that decided the number. It is always one of six values:
basis | Meaning |
|---|---|
blocker_capped | A rule marked automationBlocker failed, and its 0.1 cap is the score |
schema_capped | The value violates its schema, and the 0.2 cap is the score |
rule_capped | An ERROR-severity rule failed, and the 0.3 cap is the score |
rule_penalized | Rule failures deducted from the grounding score and no cap applied |
low_information | Grounding could not verify the value, and nothing negative touched it |
grounded | Nothing lowered the grounding score; the composite is the grounding score |
The order of the table is the priority order. A cap claims the basis only when it binds, that is, when it is at or below the score the deductions produced. A field grounded at 0.1 that also has a schema violation reads grounded at 0.1, because the 0.2 cap would have raised it. On an exact tie the cap takes the label. Read the basis as "what decided this number", not as "everything wrong with this field": the violation is still listed in resultMetadata.schemaValidation.
low_information is not a middle score
Grounding checks a value by searching for it in the source text. That cannot work for numbers, booleans, objects, strings shorter than three characters, or arrays with nothing searchable in them. Without a field hint they score 0.5 with grounding basis unverifiable, because nothing was checked. When such a field has no violation and no failed rule, its composite stays 0.5 with basis low_information: this number tells you nothing about this field.
- A
0.5that means "unverified" and a0.5that came out of a real comparison are different claims. Only the basis tells them apart. - Averaging a document's scores without separating out
low_informationmostly measures how many numeric fields the schema has. - Do not infer the basis from a field's type. On array-item fields a field hint is applied first (
system_assignedscores1.0without looking at the document), so read the basis.
Negative evidence still applies: a low_information field that violates its schema is capped like any other, and its basis changes.
Provenance
provenance lists the checks that looked at the value and did not object. It is not a second score.
| Entry | Written when |
|---|---|
grounding | Grounding found the value in the source text (exact_match, normalized_match, partial_match), or the field has the system_assigned hint |
rule | At least one business rule attributed to this field returned PASS |
threshold | A composite threshold is configured for this field and the score met it |
Entries appear in that order. history and human are reserved and not written yet. system_assigned is the one entry that is not a check: the value was trusted because of the field's configuration, not compared against the input.
An empty list is an answer: the composite was computed and nothing vouches for the value. It does not move with the score. A field can score 0.5 with [] (nothing was checked) or 0.2 with ["grounding"] (the text was found, then a schema violation capped it). The list is null only when confidence is null.
Thresholds and flags
Two independent thresholds can be configured per property on the schema version (see Extraction Configuration), and they judge different numbers:
| Config key | Compares | Flag source |
|---|---|---|
validation.properties.{property}.confidenceThreshold | the grounding score | LEGACY_GROUNDING |
confidence.properties.{property}.threshold | the composite score | COMPOSITE |
Both are keyed by property name: a root field's own name, an array property's name (which covers every scored field of every element), or $root for a schema whose root is an array.
A score strictly below its threshold records a flag; a score exactly at the threshold passes. Meeting a composite threshold also adds threshold provenance. One field can raise both flags in the same extraction. A flag is an observation only: nothing is blocked, retried or changed. At most 100 flags are recorded per result. flagCount is the true total and truncated says whether the list was cut, so read the count, not the length of the list.
The composite threshold takes a plain JSON number in [0, 1]. Anything else (a quoted "0.8", a boolean, a number out of range) counts as not configured, without a warning. If a composite threshold never flags anything, check its type first, then the property name, then whether the field was scored at all.
Which field a signal belongs to
Schema violations and rule outcomes both carry a JSON pointer. Each is attributed to the deepest scored field whose pointer is a prefix of it: /vendor/vatId lands on the root field vendor, and /parts/0/price on the item field parts[0].price.
A signal that matches no scored field lowers no score; it stays where it was recorded. The common case is a root field the model omitted or returned as null: it has no score, so the required violation that names it has nothing to lower. Rule outcomes count only when they name one concrete location. An outcome with a wildcard pointer or the empty pointer is not attributed to any field.
Known limitations:
- A signal on an array element as a whole (
/parts/0rather than/parts/0/price) lowers nothing, because only an element's fields are scored. - Flags name fields flatly (
parts[0].price). A root field literally namedparts[0].price, or$root, shares that flag key with the item field. Their scores stay separate everywhere else in the result.
When confidence is null
A null confidence never means "not confident". It means the number was not computed:
- Vision extractions carry no grounding, so they have no scoring at all:
resultMetadata.scoringisnull. - Validation is switched off for the schema version (
validation.enabled: false). That also switches off grounding, so there is nothing to build on. - Nothing was scored. Either there was no source text, or no property of the output could be scored. In the second case the scoring block can still show
methods: ["grounding"]and agrounding.scoreof1.0(the average of nothing) whilefieldsanditemsare empty. Read those two maps, not the presence of the block. - The extraction was scored before composite confidence existed.
- The confidence step could not complete for this extraction. The grounding scores are left as they were. If this happens on every extraction for one schema version, contact support.
The quickest positive check is "confidence" in scoring.methods. Before you gate anything on a score, check that resultMetadata.scoring.confidence is present and treat its absence as "unknown", never as "clean". A missing summary and a summary with flagCount: 0 look nothing alike, and only the second is a statement about your data.
Reading the payload
Confidence sits inside resultMetadata.scoring, next to the grounding scores it is built from:
{
"methods": ["grounding", "confidence"],
"grounding": { "score": 0.95, "basis": null },
"thresholds": { "parts": 0.5 },
"fields": {
"invoiceNumber": {
"grounding": { "score": 1.0, "basis": "exact_match" },
"confidence": { "score": 0.2, "basis": "schema_capped" },
"provenance": ["grounding"]
},
"total": {
"grounding": { "score": 0.5, "basis": "unverifiable" },
"confidence": { "score": 0.5, "basis": "low_information" },
"provenance": []
}
},
"items": {
"parts[0]": {
"grounding": { "score": 0.9, "basis": "normalized_match" },
"fields": {
"price": {
"grounding": { "score": 0.9, "basis": "normalized_match" },
"confidence": { "score": 0.7, "basis": "rule_penalized" },
"provenance": ["grounding"]
}
}
}
},
"confidence": {
"calibrated": false,
"flags": [
{ "field": "invoiceNumber", "confidence": 0.2, "threshold": 0.8, "source": "COMPOSITE" }
],
"flagCount": 1,
"truncated": false
}
}
fieldsholds root-level properties anditemsholds array elements, keyedparts[0](or[0]when the schema's root is an array). A property the model omitted or returned asnullappears in neither.- An item's own
grounding.basisnames the basis of its weakest scored field, the one that pulled the item score down. It readsunverifiablewhen none of the item's fields could be checked. The top-levelgrounding.basisis alwaysnull. flags[].fielduses the same keys:invoiceNumber,parts[0].price,[2].name.flags[].confidenceis the grounding score forLEGACY_GROUNDINGflags and the composite forCOMPOSITEflags. Always read it together withsource. A field that appears twice, once per source, is two signals, not a duplicate.thresholdsis a hint map for the console. It lists only array properties and$root, and flags do not depend on it: a grounding threshold on a root field raises flags without ever appearing here.
Extraction Response
The full extraction response includes structuredOutput, process metadata, and resultMetadata with dedup stats and grounding scores:
{
"id": "extraction-uuid",
"status": "COMPLETED",
"strategy": "CHUNKED",
"result": {
"structuredOutput": {
"company": "ACME GmbH",
"parts": [
{ "sourceFile": "gear.geo", "material": "1.4301", "quantity": 5 },
{ "sourceFile": "shaft.geo", "material": "Steel", "quantity": 2 }
]
},
"metadata": {
"inputTokens": 12500,
"outputTokens": 340,
"cost": 0.135,
"durationMs": 4200,
"model": "gpt-5.1",
"chunksProcessed": 2
},
"resultMetadata": {
"dedup": {
"applied": true,
"itemsRemoved": 1,
"itemsMerged": 0,
"resequenced": true,
"originalItemCount": 3,
"finalItemCount": 2
},
"postValidation": {
"applied": true,
"hallucinatedItemsRemoved": 0,
"requiredFieldsRestored": 0,
"removedItemKeys": []
},
"scoring": {
"methods": ["grounding"],
"grounding": { "score": 0.91 },
"fields": {
"company": {
"grounding": { "score": 1.0, "basis": "exact_match" }
}
},
"items": {
"parts[0]": {
"grounding": { "score": 0.92 },
"fields": {
"sourceFile": { "grounding": { "score": 1.0, "basis": "exact_match" } },
"material": { "grounding": { "score": 1.0, "basis": "exact_match" } },
"quantity": { "grounding": { "score": 0.5, "basis": "unverifiable" } }
}
},
"parts[1]": {
"grounding": { "score": 0.95 },
"fields": {
"sourceFile": { "grounding": { "score": 1.0, "basis": "exact_match" } },
"material": { "grounding": { "score": 0.9, "basis": "normalized_match" } },
"quantity": { "grounding": { "score": 0.5, "basis": "unverifiable" } }
}
}
}
}
}
}
}
The example is shortened to the grounding scores. A current result also carries resultMetadata.schemaValidation (Schema Validation), resultMetadata.rules (Business Rules) and the confidence fields shown under Composite Confidence.
Deduplication
When an extraction uses the chunked strategy, overlapping chunks can produce duplicate items. 2kw.ai automatically detects and merges these before returning results. Items with the same identifier fields (like sourceFile, partNumber, or name) are grouped — gaps between chunks are filled, exact duplicates removed, and sequential numbering fixed.
The resultMetadata.dedup object tells you what happened:
| Field | Description |
|---|---|
applied | Whether dedup was needed |
itemsRemoved | Number of exact duplicates removed |
itemsMerged | Number of partial items merged into one |
resequenced | Whether sequential numbering was fixed |
originalItemCount | Items before dedup |
finalItemCount | Items after dedup |
Dedup just works for most schemas
If your array items have clearly named identifier fields (sourceFile, partId, name, etc.), dedup handles everything automatically. See the Advanced section below for how merging works under the hood and how to override the defaults.
Post-Merge Validation
After deduplication, chunked extractions go through a validation pass that catches two types of issues:
Hallucination removal — When chunks overlap at entity boundaries, models sometimes fabricate items that don't exist in the source (e.g., inventing bbq_0_3_1.geo when only bbq_0_3_0.geo exists). The validator checks each item's literal reference fields (like sourceFile) against all chunk texts. If a filename doesn't appear anywhere in the source, the item is removed.
Required field restoration — When results are merged across chunks, required top-level fields (like date, inquiryNumber) can be silently dropped. The validator checks the schema's required array and inserts null for any missing required field, making the absence explicit rather than silent.
The resultMetadata.postValidation object tells you what happened:
| Field | Description |
|---|---|
applied | Whether post-merge validation ran |
hallucinatedItemsRemoved | Number of items removed because their identifier wasn't found in the source |
requiredFieldsRestored | Number of required fields that were missing and filled with null |
removedItemKeys | The identifier values of removed items (for debugging) |
Hallucination detection uses literal fields
By default, only file-reference fields (sourceFile, fileName, etc.) are checked against the source text. These fields should appear verbatim in the input. Derived values like summaries or computed measurements are never checked — they wouldn't be found via text matching even when correct.
To mark additional fields as literal references, use the literal field hint in your extraction config:
"fieldHints": {
"sourceFile": { "scoring": "literal" },
"documentId": { "scoring": "literal" }
}
Schema Validation
Every extraction output is checked against its schema version before it is scored. The check only records: violations are listed on the result, nothing is blocked, the status does not change and nothing is retried. For chunked extractions only the final merged output is checked, never the individual chunks. If you want non-conforming output to fail hard, enforce it in your own code from this object.
The result is in resultMetadata.schemaValidation:
{
"applied": true,
"violationCount": 2,
"truncated": false,
"violations": [
{
"fieldPath": "/total",
"keyword": "type",
"message": "string found, number expected",
"source": "MODEL",
"property": null
},
{
"fieldPath": "/invoiceNumber",
"keyword": "type",
"message": "null found, string expected",
"source": "RESTORED_NULL",
"property": null
}
]
}
| Field | Description |
|---|---|
fieldPath | RFC 6901 JSON pointer into the output (~ escaped as ~0, / as ~1). The empty string is the document root |
keyword | The JSON Schema keyword that failed: type, required, minimum, … |
property | The field the violation is about, when the keyword names one (in practice required). null otherwise |
message | Display text only, always English, and its wording can change. Match on fieldPath, keyword and property instead |
source | MODEL: the model produced a value that does not satisfy the schema. RESTORED_NULL: see below |
A required violation is anchored on the containing object, not on the missing field. A top-level required field the model left out entirely reports {"fieldPath": "", "keyword": "required", "property": "invoiceNumber"}, so read the field's name from property, never from message.
source: RESTORED_NULL means the field was missing after the chunks were merged and was filled with null by post-merge validation. That is a deliberate absence signal: the input most likely does not contain the value. Treat it as "not found", not as a model error. Only chunked extractions produce it.
At most 100 violations are recorded. violationCount is always the true total, and truncated is true when more were found.
applied: false never means "passed"; a clean output is applied: true with an empty violations list. It means the check did not run, either because validation is switched off for the schema version (validation.enabled: false, which also switches off grounding and confidence scoring) or because it could not run for this extraction.
Absence at a glance:
| Signal | Where | Meaning |
|---|---|---|
Grounding basis default_absent | resultMetadata.scoring.items[...].fields[...].grounding.basis | An array-item field with the default_absent hint came back empty or at its default: a blank string, false, an empty array, or a number that truncates to 0 or 1 (so 0.5 and 1.9 count too). Scored 0.8 as a deliberate absence. Root-level fields are scored without hints and never get this basis |
source: RESTORED_NULL | resultMetadata.schemaValidation | A required field was filled with null after the chunk merge: deliberate absence |
source: MODEL on a null | resultMetadata.schemaValidation | The model itself returned an invalid or null value, with no explanation |
A field that keeps coming back as RESTORED_NULL points at the schema or the input, not at the model: either the value is not in the documents, or the field should not be required.
Context Carry-Forward
When processing chunks sequentially, each chunk after the first receives a summary of what was already extracted from previous chunks. This helps the model:
- Avoid re-extracting items it already found in earlier chunks
- Extend partial items that span chunk boundaries (e.g., adding missing contours to a part that started in the previous chunk)
- Reduce hallucinations by knowing which entities already exist
Carry-forward is automatic and requires no configuration. It works with any schema.
Extraction Configuration
Each schema version can carry an extractionConfig that tunes per-property scoring, dedup, thresholds and business rules. It is stored alongside the JSON Schema but never sent to the model; it controls pipeline behavior only.
The configuration is set on the schema version by support. The API does not accept it when you create a version (POST /v1/schemas/{schemaId}/versions takes only jsonSchema and changeDescription), but you can read it back as extractionConfig on the schema version. To change it, contact support with the schema, the version and the settings you want.
For a schema whose parts property is an array of items with sourceFile, partNumber, material and quantity, a configuration looks like this:
{
"validation": {
"enabled": true,
"properties": {
"parts": {
"confidenceThreshold": 0.5,
"keyFields": ["sourceFile"],
"fieldHints": {
"partNumber": { "scoring": "system_assigned" },
"quantity": { "scoring": "default_absent" }
}
}
}
},
"confidence": {
"properties": {
"parts": { "threshold": 0.6 }
}
}
}
| Config | Description |
|---|---|
validation.enabled | Switches schema validation, grounding and confidence scoring on or off for the version. Defaults to on. It does not affect business rules |
validation.properties.{p}.confidenceThreshold | Grounding score below which a field is flagged. Records a LEGACY_GROUNDING flag (see Composite Confidence) and is shown in the console. Informational, does not remove items |
validation.properties.{p}.keyFields | Override which fields are used as identifiers for dedup and scoring weight. Auto-detected if not set. |
validation.properties.{p}.fieldHints | Override scoring behavior for specific fields. Each hint is an object with a scoring key. |
confidence.properties.{p}.threshold | Composite confidence below which a field is flagged, as a COMPOSITE flag. A JSON number in [0, 1]; anything else is ignored |
rules | The schema version's business rules |
Available field hints:
| Hint | Effect | Use for |
|---|---|---|
system_assigned | Always scores 1.0 | Fields assigned by the system (partNumber, auto-generated IDs) |
default_absent | Scores 0.8 when value is empty, normal scoring otherwise | Fields where empty means "not applicable" (tolerances, surface treatment) |
literal | Enables hallucination detection — value must appear verbatim in source text | Fields that reference entities in the source (filenames, document IDs, reference codes) |
Most fields don't need hints
Grounding scoring works automatically for string fields by searching the source text. You only need hints for fields where the default scoring doesn't apply — typically 2-3 fields per schema.
Business Rules
Business rules are deterministic checks over an extraction's output: a VAT ID that matches a pattern, a total that equals the sum of its line items, a date in ISO form. They need no model call and no calibration data. The same output always gets the same verdict.
Rules are part of the schema version's extraction configuration, in a top-level rules array, and are set by support like the rest of it. They run on every extraction of that version, right after schema validation and before scoring. validation.enabled does not switch them off; each rule has its own enabled flag.
Rules only record. Every outcome, passes included, is written to resultMetadata.rules; nothing is blocked, the status does not change and nothing is retried.
{
"validation": { "enabled": true },
"rules": [
{
"id": "vat-id-format",
"type": "REGEX",
"field": "/supplier/vatId",
"pattern": "^DE[0-9]{9}$",
"severity": "ERROR",
"automationBlocker": true,
"onFail": "FLAG"
}
]
}
Rule fields
Every rule accepts these fields, whatever its type:
| Field | Required | Default | Meaning |
|---|---|---|---|
id | Yes | — | Names the rule in every outcome it produces. Unique within the version's rules |
type | Yes | — | REGEX, RANGE, ENUM, ARITHMETIC, CROSS_FIELD or FORMAT |
severity | No | WARNING | ERROR, WARNING or INFO: how much a failure matters |
automationBlocker | No | false | Whether a failure should stop straight-through processing |
onFail | No | WARN | WARN, FLAG, REJECT or RETRY. Recorded, not enforced (see below) |
enabled | No | true | A disabled rule produces no outcome at all |
failOnMissing | No | false | Report FAIL instead of SKIPPED when the input is absent |
The type-specific fields (field, pattern, min, max, …) sit next to these.
Enum values are case-sensitive: "severity": "error" is not ERROR, so that rule is invalid and reports an ERROR verdict. An explicit null for severity or onFail also makes the rule invalid, while a null for a boolean field falls back to its default. When two rules share an id, the first one is evaluated and every later one reports ERROR with duplicate rule id.
Field pointers
Rules address values with RFC 6901 JSON pointers, the same format as fieldPath in schema validation. The empty pointer is the document root.
A pointer may contain at most one * segment, which expands over an array. The rule then produces one outcome per location that resolves, each with its own concrete pointer. /items/*/price over three items, of which only the first two have a price, produces two outcomes (/items/0/price and /items/1/price) and nothing for the third item. Two wildcards, or a non-empty pointer that does not start with /, is a configuration error.
A pointer into something the document does not have (an absent key, an index out of range, a wildcard over a non-array) is a miss, not an error. When the pointer as a whole resolves to nothing, the rule reports one SKIPPED (or FAIL with failOnMissing) at the pointer as configured, wildcard included. An explicit null at a location that did resolve reports SKIPPED (or FAIL) at that location.
An object key literally named * cannot be addressed, because * is always the wildcard.
Rule types
| Type | Addresses | Type-specific config | Holds when | Outcomes |
|---|---|---|---|---|
REGEX | field | pattern (Java regex) | The string matches the pattern in full | One per resolved location |
RANGE | field | min and/or max (inclusive) | The number lies within the bounds | One per resolved location |
ENUM | field | values (non-empty array) | The value equals one of the values | One per resolved location |
ARITHMETIC | field (no wildcard) | equals: {op, operands}, tolerance | |field − sum of operands| ≤ tolerance | Exactly one |
CROSS_FIELD | left, right (no wildcards) | operator | The relation holds between the two values | Exactly one |
FORMAT | field | format (named format) | The string satisfies the named format | One per resolved location |
Values are read at their JSON type and never converted. A number written as a string, or a string where a number is expected, reports ERROR. Numbers are compared as exact decimals, so 1.50 equals 1.5 and 0.1 + 0.2 is exactly 0.3.
REGEX
{
"id": "iban-shape",
"type": "REGEX",
"field": "/supplier/iban",
"pattern": "^[A-Z]{2}[0-9]{2}[A-Z0-9]{11,30}$"
}
The pattern must match the whole value, not just occur in it, so "AT123 DE123456789 xx" does not satisfy a VAT ID pattern. The syntax is Java's regular expression syntax. Strings only: a number reports ERROR, as does a missing or invalid pattern.
RANGE
{
"id": "vat-rate-plausible",
"type": "RANGE",
"field": "/taxRate",
"min": 0,
"max": 0.27
}
Both bounds are inclusive and at least one is required. A non-numeric bound, or a min above max, reports ERROR. So does a non-numeric value at field.
ENUM
{
"id": "currency-supported",
"type": "ENUM",
"field": "/items/*/currency",
"values": ["EUR", "CHF", "GBP"]
}
Values are compared by their text form, case-sensitively, so "de" does not satisfy ["DE"]. Two numbers compare as decimals, so 1.0 satisfies [1]. Mixed types fall back to text: the string "1" satisfies [1], but the number 1.0 does not satisfy ["1"]. An object or array at field reports ERROR. A missing or empty values array reports ERROR.
ARITHMETIC
{
"id": "total-equals-line-items",
"type": "ARITHMETIC",
"field": "/totals/net",
"equals": { "op": "SUM", "operands": ["/items/*/lineNet"] },
"tolerance": 0.01
}
op is SUM; any other operator reports ERROR. tolerance is an absolute amount, 0 by default, and a difference exactly equal to it passes. Operands may use a wildcard, which is how a line-item sum is written; field may not.
An ARITHMETIC rule always produces one outcome. Its actual text shows both sides and the number of addends, for example 119.05 vs 119 (2 operands). Operands that resolve only partly are summed over what exists, and a null operand is not counted. The rule reports SKIPPED when field is absent or when no operand resolves; the two outcomes look identical, so check the output to see which side was missing.
CROSS_FIELD
{
"id": "net-not-above-gross",
"type": "CROSS_FIELD",
"left": "/totals/net",
"operator": "LTE",
"right": "/totals/gross"
}
operator is EQ, NE, LT, LTE, GT or GTE. Two numbers compare as exact decimals. Two strings can only be compared with EQ and NE; ordering strings reports ERROR. Any other pairing (a number against a string, booleans, objects) reports ERROR. Neither side may use a wildcard. An absent or null side reports SKIPPED, and the outcome names the pointer of the missing side.
FORMAT
{
"id": "invoice-date-iso",
"type": "FORMAT",
"field": "/invoiceDate",
"format": "DATE_ISO"
}
| Format | Accepts | Edges |
|---|---|---|
DATE_ISO | Strict ISO 8601 calendar dates (2026-08-06) | No lenient parsing: 2026-8-6 and 06.08.2026 fail |
EMAIL | One @, a dotted domain, a top-level domain of 2+ letters | A shape check, not a deliverability check. Deliberately loose |
URL | Absolute http and https URLs with a host | International domains must be punycoded (xn--…); a host containing _ is rejected |
IBAN | 15–34 characters, country code, check digits, valid mod-97 checksum | Spaces are removed and letters upper-cased before checking |
An unknown format name reports ERROR. Strings only: a date extracted as a number reports ERROR.
Verdicts
Every outcome has exactly one verdict, and the report counts each:
| Verdict | Meaning | What it tells you |
|---|---|---|
PASS | The rule was evaluated and held | Evidence the value is not wrong in this way |
FAIL | The rule was evaluated and did not hold | A problem in the document |
SKIPPED | The input the rule needed was absent | The document lacks the field |
ERROR | The rule could not be evaluated | A problem in the rule: its configuration, or a value type it will not convert |
Watch failCount for data quality, errorCount for rules that need fixing and skippedCount for coverage: a rule that is always SKIPPED is not checking anything. No verdict ever fails the extraction. See Rule Verdict Precedence for which verdict wins when several apply.
Severity vs automationBlocker
The two are independent. severity says how much a failure matters: ERROR for a defect in the document, WARNING for something that needs human attention, INFO for observation only. automationBlocker says whether a failure should stop straight-through processing. A rule can be INFO and blocking, or ERROR and non-blocking.
Both are copied onto every outcome of the rule, including passes: they describe the rule, the verdict describes what happened. Both also feed into composite confidence.
onFail is recorded, not enforced
onFail states what should happen when a rule fails. Today all four actions behave as WARN: the action is validated and recorded on every outcome, and the extraction result is returned unchanged whatever the verdict.
| Action | Intended behavior |
|---|---|
WARN | Record and continue (the current behavior for all four) |
FLAG | Mark the result for review |
REJECT | Reject the extraction result |
RETRY | Retry the extraction |
If a failing rule should stop your workflow now, enforce it in your own code from resultMetadata.rules.
Reading Rule Outcomes
The report is in resultMetadata.rules:
{
"applied": true,
"passCount": 1,
"failCount": 1,
"skippedCount": 1,
"errorCount": 1,
"truncated": false,
"outcomes": [
{
"ruleId": "gross-equals-net-plus-vat",
"field": "/totals/gross",
"verdict": "FAIL",
"expected": "sum(/totals/net, /totals/vat) ± 0.01",
"actual": "119.05 vs 119 (2 operands)",
"severity": "ERROR",
"automationBlocker": true,
"action": "REJECT"
},
{
"ruleId": "vat-id-format",
"field": "/supplier/vatId",
"verdict": "PASS",
"expected": "^DE[0-9]{9}$",
"actual": "DE123456789",
"severity": "ERROR",
"automationBlocker": true,
"action": "FLAG"
},
{
"ruleId": "invoice-date-iso",
"field": "/invoiceDate",
"verdict": "SKIPPED",
"expected": "value present",
"actual": "absent",
"severity": "WARNING",
"automationBlocker": false,
"action": "WARN"
},
{
"ruleId": "tax-rate-plausible",
"field": "/taxRate",
"verdict": "ERROR",
"expected": "numeric 'min'",
"actual": "non-numeric 'min': ten",
"severity": "WARNING",
"automationBlocker": false,
"action": "WARN"
}
]
}
| Field | Description |
|---|---|
ruleId | The rule's id. Empty when the rule had no id, which is then the reason for its ERROR |
field | The concrete pointer the rule was evaluated at. When nothing resolved, the pointer as configured, wildcard included. Empty when no location applies, such as a duplicate id or an unreadable rule |
verdict | PASS, FAIL, SKIPPED or ERROR |
expected, actual | Display text only, and the wording can change. For an ERROR, actual carries the reason (unknown rule type 'REGEXP', min 10 exceeds max 1, duplicate rule id, wildcard not allowed in 'field', …). Match on ruleId, field and verdict, never on these |
severity, automationBlocker | As configured on the rule |
action | The configured onFail. Recorded, not enforced |
At most 200 outcomes are recorded, because passes are recorded too and a rule multiplies over array items. The four counts are always the true totals, and truncated is true when the list was cut. Read the counts, not the length of the list. expected and actual are cut at 500 characters and marked with a trailing ….
applied: false never means "all rules passed"; that is applied: true with failCount: 0. It means nothing was evaluated: the version has no rules, every rule is disabled, or rule evaluation could not run for this extraction. The three look identical in the result. validation.enabled: false does not cause it.
A rule that is always SKIPPED usually has a pointer that does not match the real shape of the output. Compare it with an output you got back, not with the schema. Set failOnMissing: true if absence is itself the problem, but note that array elements without the field never produce an outcome, so failOnMissing cannot flag them; require the field in the schema instead. A rule that is always ERROR needs fixing: read actual for the reason.
Advanced
Key Field Auto-Detection
When keyFields is not set in your extraction config, 2kw.ai auto-detects identifier fields from your schema using these patterns:
| Pattern | Examples |
|---|---|
Exact name: id, name, key | id, name |
Ends with: id, name, file, key, code, number | sourceFile, partNumber, materialCode |
Starts with: source, file | sourceDocument, fileName |
These fields serve double duty: they're used for dedup grouping (matching items across chunks) and are weighted 3x in grounding scores (a hallucinated identifier dominates the item score).
If auto-detection picks the wrong fields — or misses yours — set keyFields explicitly in your extraction config.
Dedup Merge Algorithm
Understanding how merging works helps when debugging unexpected results in chunked extractions. Here's a complete example before we break down each step.
End-to-end example
A 20-page bill of materials is too large for a single extraction. 2kw.ai splits it into two overlapping chunks. The item housing.geo spans the boundary — Chunk 1 sees the beginning of it (material is mentioned) but cuts off before the quantity. Chunk 2 picks up in the overlap region and sees the quantity, but the material description is already behind it.
Chunk 1 extracts 3 items, the last one incomplete:
{
"parts": [
{ "sourceFile": "gear.geo", "material": "1.4301", "quantity": 5, "partNumber": 1 },
{ "sourceFile": "shaft.geo", "material": "Steel", "quantity": 2, "partNumber": 2 },
{ "sourceFile": "housing.geo", "material": "Aluminum", "quantity": null, "partNumber": 3 }
]
}
Chunk 2 extracts housing.geo again (from the overlap) plus one new item:
{
"parts": [
{ "sourceFile": "housing.geo", "material": null, "quantity": 3, "partNumber": 1 },
{ "sourceFile": "bracket.geo", "material": "1.4301", "quantity": 1, "partNumber": 2 }
]
}
Now 2kw.ai merges the two outputs:
1. Detect key fields. sourceFile matches the *file suffix pattern → used as the dedup key. partNumber matches *number → marked as a sequence field.
2. Group by key. Two items share the key sourceFile: "housing.geo" — one from each chunk. The other three items (gear.geo, shaft.geo, bracket.geo) are unique and pass through unchanged.
3. Merge the group. Both housing.geo items have 3 non-null fields (tie). The one from Chunk 1 was encountered first, so it becomes the base. Its quantity is null → filled with 3 from Chunk 2's item:
Base (Chunk 1): { sourceFile: "housing.geo", material: "Aluminum", quantity: null, partNumber: 3 }
Fill from Chunk 2: quantity: 3 ←── null filled
Result: { sourceFile: "housing.geo", material: "Aluminum", quantity: 3, partNumber: 3 }
4. Resequence. After merge, partNumber values are 1, 2, 3, 2 — duplicate 2. Renumbered to 1, 2, 3, 4.
Final result:
{
"parts": [
{ "sourceFile": "gear.geo", "material": "1.4301", "quantity": 5, "partNumber": 1 },
{ "sourceFile": "shaft.geo", "material": "Steel", "quantity": 2, "partNumber": 2 },
{ "sourceFile": "housing.geo", "material": "Aluminum", "quantity": 3, "partNumber": 3 },
{ "sourceFile": "bracket.geo", "material": "1.4301", "quantity": 1, "partNumber": 4 }
]
}
{
"dedup": {
"applied": true,
"itemsRemoved": 0,
"itemsMerged": 1,
"resequenced": true,
"originalItemCount": 5,
"finalItemCount": 4
}
}
The incomplete housing.geo from Chunk 1 and the incomplete one from Chunk 2 were combined into a single complete item. Neither chunk had all the data, but together they did.
Step 1: Group by key fields
Items are grouped by the normalized values (lowercase, trimmed) of their key fields. Items where all key fields are null cannot be fingerprinted and are kept as-is — they are never merged.
If no key fields are detected at all, 2kw.ai falls back to pairwise similarity matching (items with 80%+ field value overlap are grouped).
Step 2: Merge each group
Within a group, the item with the most non-null fields becomes the base. Then, for each remaining item, any field that is null in the base is filled from the other item. If completeness is equal, the item encountered first (from the earlier chunk) becomes the base.
The end-to-end example above shows the common case: each chunk sees a different part of the item, and the merge fills the gaps.
Edge case: conflicting values. When both chunks have the same field with different non-null values, 2kw.ai uses provenance-aware conflict resolution: it prefers the value from the "authority" chunk — the chunk that contains the item's identifier (e.g., the chunk where sourceFile: "gear.geo" actually appears in the text). The authority chunk is most likely to have seen the item's header and metadata, making its scalar values more reliable.
Chunk 1: { sourceFile: "gear.geo", material: "Steel" } (authority — "gear.geo" appears in chunk 1 text)
Chunk 2: { sourceFile: "gear.geo", material: "1.4301", quantity: 5 }
→ Base: Chunk 2 (more fields), but material overridden from Chunk 1 (authority)
→ Result: { sourceFile: "gear.geo", material: "Steel", quantity: 5 }
If no authority chunk can be determined (the identifier doesn't appear in any chunk text, or both chunks contain it), the base item's value wins (most-complete-item-first, same as before).
Edge case: nested arrays. When both chunks extract array fields (like contours or holes) for the same item, the arrays are concatenated and deduplicated rather than one replacing the other. This is critical for items that span chunk boundaries — Chunk 1 might extract contours 1-10, Chunk 2 might extract contours 8-13, and the merge produces the complete set 1-13 (with duplicates 8-10 removed).
Empty strings block merging
Only null fields are filled during merge. An empty string "" counts as a value and will not be replaced. If your schema has optional fields, use nullable types ("type": ["string", "null"]) rather than defaulting to "" so that null-fill works correctly across chunks.
Step 3: Resequence
After merging, sequential number fields (names ending in number, index, num, order, sequence, position, or named nr/pos) are checked for broken numbering. If gaps or duplicates are found, the field is renumbered starting at 1.
Rule Verdict Precedence
When more than one business rule verdict could apply, these decide:
- A broken rule beats an absent field. A rule's configuration is checked before the document is read. A rule with an invalid pattern, an empty
valuesarray or aminabovemaxreportsERROReven when its field is also absent, so it counts inerrorCount, never inskippedCount. A rule that can never work says so instead of reportingSKIPPEDforever. - A broken rule is one
ERROR, not one per location. A configuration problem on a rule whose pointer covers fifty array items produces a singleERROR, because the problem belongs to the rule. A value-levelERROR, such as a wrong type at one location, is still reported per location, because it belongs to the document. failOnMissingturnsSKIPPEDintoFAILand nothing else. It does not affectERROR. It reaches a pointer that resolved to nothing (oneFAILat the configured pointer) and an explicitnullat a resolved location (oneFAILthere).- A rule that crashes is that rule's
ERROR. It never suppresses the rules around it and never fails the extraction. The exception is a pattern so pathological that evaluation cannot continue at all; the whole report is then recorded asapplied: false.
Re-running Extractions
POST /v1/extractions/{id}/rerun
Need to retry an extraction with the same config? Hit the rerun endpoint. Creates a new extraction —the original stays untouched.
The re-run copies the original's schema version, input text and model, and it runs the same
way the original did: an extraction that was queued asynchronously is queued again and comes
back PENDING, a synchronous one is executed inline and comes back with its result.
Not every extraction can be re-run:
| Response | When |
|---|---|
404 Not Found | No such extraction in your organization |
409 Conflict | The extraction is still PENDING or PROCESSING |
409 Conflict | The extraction was vision-only —images are not stored and cannot be replayed |
409 Conflict | The schema version it ran against has since been deleted |
COMPLETED and FAILED extractions are eligible. Each re-run is billed as a new extraction.
Listing and Filtering
GET /v1/extractions
| Parameter | Type | Description |
|---|---|---|
search | string | Filter by model name |
schemaVersionId | string | Filter by schema version |
status | string | PENDING, PROCESSING, COMPLETED, or FAILED |
page | number | Page number (0-based) |
size | number | Page size (default: 20) |
Use Cases
- Contact extraction —names, emails, phone numbers from unstructured text
- Invoice processing —invoice numbers, dates, amounts, line items from documents
- Resume parsing —skills, experience, education from CVs
- Document analysis —key fields from contracts, reports, forms
- Data entry automation —turn free-text notes into structured database records