Data Extraction

Define a schema once, then extract structured data from any text —consistently, every time.

How It Works

2kw.ai takes a standard JSON Schema definition and uses it to extract structured output from unstructured text. You define the fields, types, and constraints — 2kw.ai returns validated, schema-conformant JSON.

The workflow is: define schema → commit version → send text → receive structured output.

Schema Management

Before you can extract anything, you need a schema. Schemas are organization-scoped and support full version control.

Creating a Schema

A schema takes two calls: create the schema with a name and an optional description, then commit its JSON Schema definition as the first version.

POST /v1/schemas

curl -X POST https://api.2kw.ai/v1/schemas \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer sk_your_api_key" \
  -d '{
    "name": "invoice",
    "description": "Supplier invoices"
  }'

The response carries the schema's id. The JSON Schema is not part of this call: the API reference lists a content field on the schema, but a value sent there is not stored.

POST /v1/schemas/{schemaId}/versions

{
  "jsonSchema": {
    "type": "object",
    "properties": {
      "company": { "type": "string" },
      "invoice_number": { "type": "string" },
      "total": { "type": "number" },
      "line_items": {
        "type": "array",
        "items": {
          "type": "object",
          "properties": {
            "description": { "type": "string" },
            "quantity": { "type": "integer" },
            "price": { "type": "number" }
          }
        }
      }
    },
    "required": ["company", "invoice_number"]
  },
  "changeDescription": "First version"
}
FieldTypeRequiredDescription
jsonSchemaobjectYesThe JSON Schema definition. One that is not a valid JSON Schema is refused with 400
changeDescriptionstringNoWhat changed in this version

The response (201) is the new version with its id, versionNumber and jsonSchema, and the latest label now points to it. Every later change is another call to this endpoint.

Version Control

Every time you change your schema, you create a new version. Versions are immutable —this gives you a full audit trail of every change.

You can:

  • Commit new versions with a change description
  • Deactivate versions you don't want used anymore
  • Re-activate a historical version (creates a new version as a copy, preserving the audit trail)
  • Pin extractions to a specific version for reproducibility

Each new version auto-increments the version number and updates the latest label.

Labels

Labels are named pointers to specific schema versions. They decouple your application code from version numbers, enabling controlled rollouts.

How Labels Work

Every schema gets a latest label automatically, updated on each new version. You create custom labels for deployment stages:

LabelPoints toPurpose
latestv5 (auto)Always the newest version
productionv3What your API consumers use
stagingv5What you're testing

Your application resolves schemas by ID + label —when you're ready to promote, just move the label pointer.

Managing Labels

POST /v1/schemas/{schemaId}/labels

Create a label:

{
  "name": "production",
  "schemaVersionId": "version-uuid"
}

Update a label to point to a different version:

PUT /v1/schemas/{schemaId}/labels/{labelName}

{
  "schemaVersionId": "new-version-uuid"
}

Label names must be lowercase alphanumeric with hyphens (e.g., production, staging, v2-rollback). The system latest label cannot be deleted.

Resolving Schemas

GET /v1/schemas/{schemaId}/resolve

Resolve a schema by ID and optional label to get the JSON Schema definition for a specific version:

Request

curl "https://api.2kw.ai/v1/schemas/{schemaId}/resolve?label=production" \
  -H "Authorization: Bearer sk_your_api_key"
ParameterTypeRequiredDescription
labelstringNoLabel to resolve (defaults to latest)

Response:

{
  "schemaId": "uuid",
  "schemaName": "invoice-schema",
  "versionId": "uuid",
  "versionNumber": 3,
  "jsonSchema": {
    "type": "object",
    "properties": {
      "company": { "type": "string" },
      "total": { "type": "number" }
    }
  },
  "label": "production"
}

Validate Before Committing

POST /v1/schemas/{schemaId}/validate

Check if a schema is valid before you commit it. Returns structured errors and warnings:

{
  "jsonSchema": {
    "type": "object",
    "properties": {
      "name": { "type": "string" }
    }
  }
}

The response has three fields: valid, errors and warnings. Errors mean the schema document itself cannot be used. The same checks run when you commit a version, so POST /v1/schemas/{schemaId}/versions answers 400 for a document that fails them. The typical errors are:

  • JSON Schema must be a JSON object
  • JSON Schema must contain at least one schema keyword (e.g., type, properties, $schema)
  • Invalid JSON Schema: <detail from the validator>

Check the JSON syntax, make sure the document is an object with at least one JSON Schema keyword, and validate it against the JSON Schema 2020-12 specification.

Errors are about the schema, never about extraction output. An extraction whose output does not match its schema does not fail: there is no failed status and no retry. The mismatches are recorded on the result instead, see Schema Validation.

Schema-safety warnings

Warnings are advisory and never block saving a version: valid stays true. The validate endpoint returns them on request, and the schema editor in the console lints live while you type. Each warning names a stable rule id, a JSON pointer to the exact location in your schema, a message and, where one applies, a suggested fix. The editor offers a one-click fix for the first three rules.

RuleTriggerWhy it is riskyFix
required-non-nullableA required field declares a type that does not include nullThe model must emit some value even when the document has noneAllow null ("type": ["string", "null"]) or remove the field from required
required-enumA required field has an enum that contains neither "insufficient_evidence" nor nullThe model must pick a listed value even when the document contains noneAdd "insufficient_evidence" or null to the enum, or remove the field from required
min-itemsminItems greater than 0 on any arrayForces the model to invent entries for documents that have noneRemove minItems and check cardinality after extraction instead, for example with a business rule
provider-rewriteThe schema uses constraints your provider does not enforceConstraints look enforced but are ignored at generation time, or the schema is rejectedInformational, no automatic fix. See Provider enforcement below

A required field that carries an enum is reported as required-enum, not required-non-nullable. A field without a type keyword, or one whose anyOf/oneOf includes a {"type": "null"} branch, is already nullable and is not flagged.

The pointer targets the property for required-non-nullable and required-enum, and the offending keyword for min-items and provider-rewrite. A provider-rewrite warning about prompt-only providers uses the empty pointer, which means the whole schema.

Example. This schema requires two fields, one of them an enum without an escape value:

{
  "jsonSchema": {
    "type": "object",
    "properties": {
      "invoice_number": { "type": "string" },
      "status": { "type": "string", "enum": ["paid", "unpaid"] }
    },
    "required": ["invoice_number", "status"]
  }
}

With only the platform's default provider in play, the response is:

{
  "valid": true,
  "errors": [],
  "warnings": [
    {
      "rule": "required-non-nullable",
      "pointer": "/properties/invoice_number",
      "message": "Required field 'invoice_number' cannot be null. When the document does not contain this value, models fabricate one rather than fail.",
      "suggestion": "Allow null via \"type\": [\"string\", \"null\"], or remove 'invoice_number' from \"required\"."
    },
    {
      "rule": "required-enum",
      "pointer": "/properties/status",
      "message": "Required enum field 'status' forces the model to pick a listed value even when the document contains none — the highest-risk fabrication construct.",
      "suggestion": "Add an escape value \"insufficient_evidence\" to the enum, or remove 'status' from \"required\"."
    }
  ]
}

An organization that has an Anthropic or Mistral provider configured also gets a provider-rewrite warning at the empty pointer.

Escape values

Make fields optional and nullable by default. Require a field only when a document without it should count as unprocessable. For enums, include the escape value insufficient_evidence so the model has a legal way to say "not in the document". A null entry in the enum counts as an escape too. An escape value reads as a deliberate absence; a fabricated value looks exactly like a real one.

Provider enforcement

Extraction runs on the provider behind the model you choose, and providers differ in what they enforce while generating:

  • Anthropic and Mistral receive the schema as prompt text, without constrained decoding. Required fields, enums and every other constraint are advisory.
  • OpenAI, Azure OpenAI and xAI receive the schema as a structured-output json_schema. Constraint keywords outside the enforced subset are ignored, and Azure OpenAI may reject a schema that uses them. The flagged keywords are pattern, format, minLength, maxLength, minimum, maximum, exclusiveMinimum, exclusiveMaximum, multipleOf, maxItems, uniqueItems, minProperties and maxProperties. minItems is not enforced either; it is reported by the min-items rule, which carries the fix.
  • Vertex AI and Ollama are not checked, so they produce no provider-rewrite warnings.

The linter checks your schema against every provider configured for your organization, plus Azure OpenAI, which serves the platform's built-in models.

2kw.ai does not enforce these keywords after extraction either. It does check every output against its schema and records the violations, for enforced and advisory keywords alike, see Schema Validation.

Test Against Sample Text

POST /v1/schemas/{schemaId}/test

Run an extraction against sample text without persisting the result. Great for iterating on your schema:

Request

curl -X POST https://api.2kw.ai/v1/schemas/{schemaId}/test \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer sk_your_api_key" \
  -d '{
    "jsonSchema": {
      "type": "object",
      "properties": {
        "name": { "type": "string" },
        "email": { "type": "string" }
      }
    },
    "sampleText": "Contact Jane Smith at jane@acme.com",
    "model": "gpt-5.1"
  }'

The response includes token usage and processing time so you can estimate costs:

{
  "success": true,
  "extractedData": { "name": "Jane Smith", "email": "jane@acme.com" },
  "inputTokens": 85,
  "outputTokens": 22,
  "processingDurationMs": 890,
  "modelUsed": "gpt-5.1"
}

All Schema Endpoints

MethodEndpointWhat it does
GET/v1/schemas?search=invoice&page=0&size=20List schemas, paginated, optionally filtered by search
POST/v1/schemasCreate a schema (name, description)
GET/v1/schemas/{id}Get a schema
PUT/v1/schemas/{id}Update a schema's name or description
DELETE/v1/schemas/{id}Delete a schema and all its versions. Needs the admin role
POST/v1/schemas/{schemaId}/versionsCommit a new version (jsonSchema, changeDescription)
GET/v1/schemas/{schemaId}/versions?page=0&size=20List versions, paginated
GET/v1/schemas/{schemaId}/versions/{versionId}Get one version
GET/v1/schemas/{schemaId}/versions/latestThe latest version; 404 while there is none. Deprecated: use resolve
PUT/v1/schemas/{schemaId}/versions/{versionId}/activateRe-activate a historical version as a new copy
DELETE/v1/schemas/{schemaId}/versions/{versionId}Deactivate a version (soft delete); answers 204. Needs the admin role
GET/v1/schemas/{schemaId}/labelsList labels
POST/v1/schemas/{schemaId}/labelsCreate a label (name, schemaVersionId)
PUT/v1/schemas/{schemaId}/labels/{labelName}Point a label at another version (schemaVersionId)
DELETE/v1/schemas/{schemaId}/labels/{labelName}Delete a label; latest cannot be deleted. Needs the admin role
GET/v1/schemas/{schemaId}/resolve?label=productionResolve a label (default latest) to its version
POST/v1/schemas/{schemaId}/validateValidate a JSON Schema without saving it
POST/v1/schemas/{schemaId}/testExtract from sample text without saving the result

Running Extractions

Synchronous

POST /v1/extractions

The standard extraction endpoint. Send text and, for inputs that fit a synchronous run, get the structured data back in the same response. Larger inputs are handed to background processing automatically; see the responses below the example.

Request body:

FieldTypeRequiredDescription
schemaIdstringYesSchema to use for extraction
schemaVersionIdstringNoPin to a specific version (defaults to latest active)
inputTextstringConditionalText to extract data from (required if no inputImages)
inputImagesarrayConditionalBase64-encoded images to extract from, max 10 (required if no inputText)
modelstringYesPlatform model name (e.g., gpt-5.1) or provider/model for BYOK

Request

curl -X POST https://api.2kw.ai/v1/extractions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer sk_your_api_key" \
  -d '{
    "schemaId": "your-schema-id",
    "inputText": "Invoice #INV-2024-001 from Acme Corp. Total: $1,250.00",
    "model": "gpt-5.1"
  }'

Responses. Always read status before you read the result:

HTTP statusstatusMeaning
201 CreatedCOMPLETEDThe extraction ran; result.structuredOutput holds the data
201 CreatedFAILEDThe extraction ran and failed; errorMessage says why
202 AcceptedPENDINGThe input was too large for a synchronous run and was queued in the background. The response carries a Location header and Retry-After: 5; poll it as described under Asynchronous

A failed extraction is still a 201, so checking the HTTP status alone is not enough:

const response = await fetch(`${BASE_URL}/v1/extractions`, {
  method: "POST",
  headers: {
    "Content-Type": "application/json",
    "Authorization": `Bearer ${apiKey}`,
  },
  body: JSON.stringify({ schemaId, inputText, model }),
});

if (!response.ok) {
  throw new Error(`Request rejected: ${response.status}`);
}

const data = await response.json();
if (data.status === "COMPLETED") {
  console.log(data.result.structuredOutput);
} else if (data.status === "FAILED") {
  throw new Error(data.errorMessage ?? "Extraction failed");
} else {
  // PENDING: poll GET /v1/extractions/{id} until COMPLETED or FAILED
  console.log("Queued, poll", response.headers.get("Location"));
}

Extracting from Images

You can also extract structured data from images (up to 10 per request). Each image must be base64-encoded with its MIME type:

Image extraction

curl -X POST https://api.2kw.ai/v1/extractions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer sk_your_api_key" \
  -d '{
    "schemaId": "your-schema-id",
    "inputImages": [
      {
        "data": "iVBORw0KGgoAAAANSUhEUg...",
        "mimeType": "image/png"
      }
    ],
    "model": "gpt-5.1"
  }'

Supported image MIME types: image/png, image/jpeg, image/gif, image/webp. You can combine inputText and inputImages in a single request.

Asynchronous

POST /v1/extractions/async

For large inputs, submit an extraction for background processing. You get back a 202 Accepted with a Location header to poll:

HTTP/1.1 202 Accepted
Location: /v1/extractions/{id}
Retry-After: 5

Poll the Location URL until status changes from PENDING/PROCESSING to COMPLETED or FAILED.

Estimate Tokens First

POST /v1/extractions/estimate

Before running an extraction, you can estimate the cost. Send the same schema and input text you'd use for a real extraction:

Request body:

FieldTypeRequiredDescription
schemaIdstringYesSchema to estimate for
schemaVersionIdstringNoPin to a specific version (defaults to latest active)
inputTextstringYesText to estimate token usage for

Response:

{
  "inputTokens": 4200,
  "estimatedOutputTokens": 350,
  "strategy": "SINGLE_SHOT"
}

The strategy field tells you which processing path will be used (SINGLE_SHOT, CHUNKED, or ASYNC).

Extraction Strategies

2kw.ai automatically picks the best strategy based on input size:

StrategyToken RangeWhat Happens
Single-shotUnder 50K tokensEntire input processed in one LLM call
Chunked50K - 100K tokensInput split into overlapping chunks, results merged and validated
AsyncOver 100K tokensBackground processing with polling

Quality Scoring

Every text extraction result includes quality scores that help you assess how trustworthy the extracted data is. These scores are purely informational — structuredOutput is never modified.

Grounding vs Confidence

2kw.ai distinguishes between two types of quality signals:

Grounding measures whether an extracted value has evidence in the source text. If your schema extracts a sourceFile field and the value "gear.geo" appears in the input, that's grounded. If "phantom.geo" doesn't appear anywhere in the input, it was likely hallucinated by the model. Grounding is deterministic, free (no extra LLM calls), and available on every extraction.

Confidence combines the grounding score with the other evidence 2kw.ai already has about a value: the schema violations recorded for it and the verdicts of your business rules. It is arithmetic over those signals. No model is asked how sure it is, no extra call is made, and the same result always produces the same score. See Composite Confidence for how the number is built.

The scoring.methods array tells you which methods ran: grounding, and confidence when the composite was computed.

How Grounding Works

For each extracted field, 2kw.ai searches the original input text for the extracted value:

BasisScoreMeaning
exact_match1.0Value found verbatim in the source text
normalized_match0.9Found after normalizing separators (underscores, hyphens, spaces)
partial_match0.8Substring or filename stem found (e.g. gear matches gear.geo)
unverifiable0.5Numbers, booleans, or strings shorter than 3 characters — can't be meaningfully searched
not_found0.0 – 0.1String value not found in source text. Key fields score 0.0 (strong hallucination signal), regular fields score 0.1.

Per-item scores are weighted averages of their fields, with key fields (identifiers like filenames, IDs, names) weighted 3x higher than regular fields. This means a hallucinated filename dominates the item score even if other fields look fine. Unverifiable fields (numbers, booleans, short strings) are excluded from this average — they don't drag the score down or inflate it.

Composite Confidence

Every text extraction that is scored carries a composite confidence per scored field: one number in [0, 1], built from the field's grounding score, its schema violations and its business-rule verdicts. Like the other quality signals it is recorded only. Nothing is blocked, no status changes and nothing is retried.

The scores are not calibrated

resultMetadata.scoring.confidence.calibrated is false, and it is the most important field in this section. The numbers are an ordering, not a probability: 0.7 does not mean "70% likely to be correct". The constants below have not been fitted against labelled outcomes. Use the scores to rank fields against each other within one schema version, derive any threshold from your own observed results, and do not show a confidence number to an end user as a probability.

How the score is computed

The grounding score is the base. Negative evidence attributed to the field either deducts from it or caps it. Nothing raises it: a rule that passed adds to the field's provenance, never to its score.

Signal attributed to the fieldEffectAmount
One or more schema violationscaps0.2
A failed rule with automationBlocker: truecaps0.1
A failed rule with severity: ERRORcaps0.3
A failed rule with severity: WARNINGdeducts, once per failure0.2
A failed rule with severity: INFOdeducts, once per failure0.05
A rule that passedno change to the score, adds rule provenance—
A rule outcome with verdict SKIPPED or ERRORnothing—
base    = grounding score
penalty = sum of the deductions above
cap     = the smallest applicable cap, if any
score   = clamp(min(round(base − penalty, 2 decimals), cap), 0, 1)

severity and automationBlocker act independently. A blocking WARNING failure deducts 0.2 and caps at 0.1; when several caps apply, the smallest wins. Deductions add up without limit (three failed WARNING rules deduct 0.6) and the result is clamped at 0. Caps beat deductions because the two meet in a min. The SKIPPED and ERROR verdicts do not count: the input was absent or the rule could not run, and neither says anything about the value. Scores are rounded to two decimals, the same grid as the grounding scores.

The basis

confidence.basis names the one cause that decided the number. It is always one of six values:

basisMeaning
blocker_cappedA rule marked automationBlocker failed, and its 0.1 cap is the score
schema_cappedThe value violates its schema, and the 0.2 cap is the score
rule_cappedAn ERROR-severity rule failed, and the 0.3 cap is the score
rule_penalizedRule failures deducted from the grounding score and no cap applied
low_informationGrounding could not verify the value, and nothing negative touched it
groundedNothing lowered the grounding score; the composite is the grounding score

The order of the table is the priority order. A cap claims the basis only when it binds, that is, when it is at or below the score the deductions produced. A field grounded at 0.1 that also has a schema violation reads grounded at 0.1, because the 0.2 cap would have raised it. On an exact tie the cap takes the label. Read the basis as "what decided this number", not as "everything wrong with this field": the violation is still listed in resultMetadata.schemaValidation.

low_information is not a middle score

Grounding checks a value by searching for it in the source text. That cannot work for numbers, booleans, objects, strings shorter than three characters, or arrays with nothing searchable in them. Without a field hint they score 0.5 with grounding basis unverifiable, because nothing was checked. When such a field has no violation and no failed rule, its composite stays 0.5 with basis low_information: this number tells you nothing about this field.

  • A 0.5 that means "unverified" and a 0.5 that came out of a real comparison are different claims. Only the basis tells them apart.
  • Averaging a document's scores without separating out low_information mostly measures how many numeric fields the schema has.
  • Do not infer the basis from a field's type. On array-item fields a field hint is applied first (system_assigned scores 1.0 without looking at the document), so read the basis.

Negative evidence still applies: a low_information field that violates its schema is capped like any other, and its basis changes.

Provenance

provenance lists the checks that looked at the value and did not object. It is not a second score.

EntryWritten when
groundingGrounding found the value in the source text (exact_match, normalized_match, partial_match), or the field has the system_assigned hint
ruleAt least one business rule attributed to this field returned PASS
thresholdA composite threshold is configured for this field and the score met it

Entries appear in that order. history and human are reserved and not written yet. system_assigned is the one entry that is not a check: the value was trusted because of the field's configuration, not compared against the input.

An empty list is an answer: the composite was computed and nothing vouches for the value. It does not move with the score. A field can score 0.5 with [] (nothing was checked) or 0.2 with ["grounding"] (the text was found, then a schema violation capped it). The list is null only when confidence is null.

Thresholds and flags

Two independent thresholds can be configured per property on the schema version (see Extraction Configuration), and they judge different numbers:

Config keyComparesFlag source
validation.properties.{property}.confidenceThresholdthe grounding scoreLEGACY_GROUNDING
confidence.properties.{property}.thresholdthe composite scoreCOMPOSITE

Both are keyed by property name: a root field's own name, an array property's name (which covers every scored field of every element), or $root for a schema whose root is an array.

A score strictly below its threshold records a flag; a score exactly at the threshold passes. Meeting a composite threshold also adds threshold provenance. One field can raise both flags in the same extraction. A flag is an observation only: nothing is blocked, retried or changed. At most 100 flags are recorded per result. flagCount is the true total and truncated says whether the list was cut, so read the count, not the length of the list.

The composite threshold takes a plain JSON number in [0, 1]. Anything else (a quoted "0.8", a boolean, a number out of range) counts as not configured, without a warning. If a composite threshold never flags anything, check its type first, then the property name, then whether the field was scored at all.

Which field a signal belongs to

Schema violations and rule outcomes both carry a JSON pointer. Each is attributed to the deepest scored field whose pointer is a prefix of it: /vendor/vatId lands on the root field vendor, and /parts/0/price on the item field parts[0].price.

A signal that matches no scored field lowers no score; it stays where it was recorded. The common case is a root field the model omitted or returned as null: it has no score, so the required violation that names it has nothing to lower. Rule outcomes count only when they name one concrete location. An outcome with a wildcard pointer or the empty pointer is not attributed to any field.

Known limitations:

  • A signal on an array element as a whole (/parts/0 rather than /parts/0/price) lowers nothing, because only an element's fields are scored.
  • Flags name fields flatly (parts[0].price). A root field literally named parts[0].price, or $root, shares that flag key with the item field. Their scores stay separate everywhere else in the result.

When confidence is null

A null confidence never means "not confident". It means the number was not computed:

  1. Vision extractions carry no grounding, so they have no scoring at all: resultMetadata.scoring is null.
  2. Validation is switched off for the schema version (validation.enabled: false). That also switches off grounding, so there is nothing to build on.
  3. Nothing was scored. Either there was no source text, or no property of the output could be scored. In the second case the scoring block can still show methods: ["grounding"] and a grounding.score of 1.0 (the average of nothing) while fields and items are empty. Read those two maps, not the presence of the block.
  4. The extraction was scored before composite confidence existed.
  5. The confidence step could not complete for this extraction. The grounding scores are left as they were. If this happens on every extraction for one schema version, contact support.

The quickest positive check is "confidence" in scoring.methods. Before you gate anything on a score, check that resultMetadata.scoring.confidence is present and treat its absence as "unknown", never as "clean". A missing summary and a summary with flagCount: 0 look nothing alike, and only the second is a statement about your data.

Reading the payload

Confidence sits inside resultMetadata.scoring, next to the grounding scores it is built from:

{
  "methods": ["grounding", "confidence"],
  "grounding": { "score": 0.95, "basis": null },
  "thresholds": { "parts": 0.5 },
  "fields": {
    "invoiceNumber": {
      "grounding": { "score": 1.0, "basis": "exact_match" },
      "confidence": { "score": 0.2, "basis": "schema_capped" },
      "provenance": ["grounding"]
    },
    "total": {
      "grounding": { "score": 0.5, "basis": "unverifiable" },
      "confidence": { "score": 0.5, "basis": "low_information" },
      "provenance": []
    }
  },
  "items": {
    "parts[0]": {
      "grounding": { "score": 0.9, "basis": "normalized_match" },
      "fields": {
        "price": {
          "grounding": { "score": 0.9, "basis": "normalized_match" },
          "confidence": { "score": 0.7, "basis": "rule_penalized" },
          "provenance": ["grounding"]
        }
      }
    }
  },
  "confidence": {
    "calibrated": false,
    "flags": [
      { "field": "invoiceNumber", "confidence": 0.2, "threshold": 0.8, "source": "COMPOSITE" }
    ],
    "flagCount": 1,
    "truncated": false
  }
}
  • fields holds root-level properties and items holds array elements, keyed parts[0] (or [0] when the schema's root is an array). A property the model omitted or returned as null appears in neither.
  • An item's own grounding.basis names the basis of its weakest scored field, the one that pulled the item score down. It reads unverifiable when none of the item's fields could be checked. The top-level grounding.basis is always null.
  • flags[].field uses the same keys: invoiceNumber, parts[0].price, [2].name.
  • flags[].confidence is the grounding score for LEGACY_GROUNDING flags and the composite for COMPOSITE flags. Always read it together with source. A field that appears twice, once per source, is two signals, not a duplicate.
  • thresholds is a hint map for the console. It lists only array properties and $root, and flags do not depend on it: a grounding threshold on a root field raises flags without ever appearing here.

Extraction Response

The full extraction response includes structuredOutput, process metadata, and resultMetadata with dedup stats and grounding scores:

{
  "id": "extraction-uuid",
  "status": "COMPLETED",
  "strategy": "CHUNKED",
  "result": {
    "structuredOutput": {
      "company": "ACME GmbH",
      "parts": [
        { "sourceFile": "gear.geo", "material": "1.4301", "quantity": 5 },
        { "sourceFile": "shaft.geo", "material": "Steel", "quantity": 2 }
      ]
    },
    "metadata": {
      "inputTokens": 12500,
      "outputTokens": 340,
      "cost": 0.135,
      "durationMs": 4200,
      "model": "gpt-5.1",
      "chunksProcessed": 2
    },
    "resultMetadata": {
      "dedup": {
        "applied": true,
        "itemsRemoved": 1,
        "itemsMerged": 0,
        "resequenced": true,
        "originalItemCount": 3,
        "finalItemCount": 2
      },
      "postValidation": {
        "applied": true,
        "hallucinatedItemsRemoved": 0,
        "requiredFieldsRestored": 0,
        "removedItemKeys": []
      },
      "scoring": {
        "methods": ["grounding"],
        "grounding": { "score": 0.91 },
        "fields": {
          "company": {
            "grounding": { "score": 1.0, "basis": "exact_match" }
          }
        },
        "items": {
          "parts[0]": {
            "grounding": { "score": 0.92 },
            "fields": {
              "sourceFile": { "grounding": { "score": 1.0, "basis": "exact_match" } },
              "material":   { "grounding": { "score": 1.0, "basis": "exact_match" } },
              "quantity":   { "grounding": { "score": 0.5, "basis": "unverifiable" } }
            }
          },
          "parts[1]": {
            "grounding": { "score": 0.95 },
            "fields": {
              "sourceFile": { "grounding": { "score": 1.0, "basis": "exact_match" } },
              "material":   { "grounding": { "score": 0.9, "basis": "normalized_match" } },
              "quantity":   { "grounding": { "score": 0.5, "basis": "unverifiable" } }
            }
          }
        }
      }
    }
  }
}

The example is shortened to the grounding scores. A current result also carries resultMetadata.schemaValidation (Schema Validation), resultMetadata.rules (Business Rules) and the confidence fields shown under Composite Confidence.

Deduplication

When an extraction uses the chunked strategy, overlapping chunks can produce duplicate items. 2kw.ai automatically detects and merges these before returning results. Items with the same identifier fields (like sourceFile, partNumber, or name) are grouped — gaps between chunks are filled, exact duplicates removed, and sequential numbering fixed.

The resultMetadata.dedup object tells you what happened:

FieldDescription
appliedWhether dedup was needed
itemsRemovedNumber of exact duplicates removed
itemsMergedNumber of partial items merged into one
resequencedWhether sequential numbering was fixed
originalItemCountItems before dedup
finalItemCountItems after dedup

Post-Merge Validation

After deduplication, chunked extractions go through a validation pass that catches two types of issues:

Hallucination removal — When chunks overlap at entity boundaries, models sometimes fabricate items that don't exist in the source (e.g., inventing bbq_0_3_1.geo when only bbq_0_3_0.geo exists). The validator checks each item's literal reference fields (like sourceFile) against all chunk texts. If a filename doesn't appear anywhere in the source, the item is removed.

Required field restoration — When results are merged across chunks, required top-level fields (like date, inquiryNumber) can be silently dropped. The validator checks the schema's required array and inserts null for any missing required field, making the absence explicit rather than silent.

The resultMetadata.postValidation object tells you what happened:

FieldDescription
appliedWhether post-merge validation ran
hallucinatedItemsRemovedNumber of items removed because their identifier wasn't found in the source
requiredFieldsRestoredNumber of required fields that were missing and filled with null
removedItemKeysThe identifier values of removed items (for debugging)

Schema Validation

Every extraction output is checked against its schema version before it is scored. The check only records: violations are listed on the result, nothing is blocked, the status does not change and nothing is retried. For chunked extractions only the final merged output is checked, never the individual chunks. If you want non-conforming output to fail hard, enforce it in your own code from this object.

The result is in resultMetadata.schemaValidation:

{
  "applied": true,
  "violationCount": 2,
  "truncated": false,
  "violations": [
    {
      "fieldPath": "/total",
      "keyword": "type",
      "message": "string found, number expected",
      "source": "MODEL",
      "property": null
    },
    {
      "fieldPath": "/invoiceNumber",
      "keyword": "type",
      "message": "null found, string expected",
      "source": "RESTORED_NULL",
      "property": null
    }
  ]
}
FieldDescription
fieldPathRFC 6901 JSON pointer into the output (~ escaped as ~0, / as ~1). The empty string is the document root
keywordThe JSON Schema keyword that failed: type, required, minimum, …
propertyThe field the violation is about, when the keyword names one (in practice required). null otherwise
messageDisplay text only, always English, and its wording can change. Match on fieldPath, keyword and property instead
sourceMODEL: the model produced a value that does not satisfy the schema. RESTORED_NULL: see below

A required violation is anchored on the containing object, not on the missing field. A top-level required field the model left out entirely reports {"fieldPath": "", "keyword": "required", "property": "invoiceNumber"}, so read the field's name from property, never from message.

source: RESTORED_NULL means the field was missing after the chunks were merged and was filled with null by post-merge validation. That is a deliberate absence signal: the input most likely does not contain the value. Treat it as "not found", not as a model error. Only chunked extractions produce it.

At most 100 violations are recorded. violationCount is always the true total, and truncated is true when more were found.

applied: false never means "passed"; a clean output is applied: true with an empty violations list. It means the check did not run, either because validation is switched off for the schema version (validation.enabled: false, which also switches off grounding and confidence scoring) or because it could not run for this extraction.

Absence at a glance:

SignalWhereMeaning
Grounding basis default_absentresultMetadata.scoring.items[...].fields[...].grounding.basisAn array-item field with the default_absent hint came back empty or at its default: a blank string, false, an empty array, or a number that truncates to 0 or 1 (so 0.5 and 1.9 count too). Scored 0.8 as a deliberate absence. Root-level fields are scored without hints and never get this basis
source: RESTORED_NULLresultMetadata.schemaValidationA required field was filled with null after the chunk merge: deliberate absence
source: MODEL on a nullresultMetadata.schemaValidationThe model itself returned an invalid or null value, with no explanation

A field that keeps coming back as RESTORED_NULL points at the schema or the input, not at the model: either the value is not in the documents, or the field should not be required.

Context Carry-Forward

When processing chunks sequentially, each chunk after the first receives a summary of what was already extracted from previous chunks. This helps the model:

  • Avoid re-extracting items it already found in earlier chunks
  • Extend partial items that span chunk boundaries (e.g., adding missing contours to a part that started in the previous chunk)
  • Reduce hallucinations by knowing which entities already exist

Carry-forward is automatic and requires no configuration. It works with any schema.

Extraction Configuration

Each schema version can carry an extractionConfig that tunes per-property scoring, dedup, thresholds and business rules. It is stored alongside the JSON Schema but never sent to the model; it controls pipeline behavior only.

The configuration is set on the schema version by support. The API does not accept it when you create a version (POST /v1/schemas/{schemaId}/versions takes only jsonSchema and changeDescription), but you can read it back as extractionConfig on the schema version. To change it, contact support with the schema, the version and the settings you want.

For a schema whose parts property is an array of items with sourceFile, partNumber, material and quantity, a configuration looks like this:

{
  "validation": {
    "enabled": true,
    "properties": {
      "parts": {
        "confidenceThreshold": 0.5,
        "keyFields": ["sourceFile"],
        "fieldHints": {
          "partNumber": { "scoring": "system_assigned" },
          "quantity": { "scoring": "default_absent" }
        }
      }
    }
  },
  "confidence": {
    "properties": {
      "parts": { "threshold": 0.6 }
    }
  }
}
ConfigDescription
validation.enabledSwitches schema validation, grounding and confidence scoring on or off for the version. Defaults to on. It does not affect business rules
validation.properties.{p}.confidenceThresholdGrounding score below which a field is flagged. Records a LEGACY_GROUNDING flag (see Composite Confidence) and is shown in the console. Informational, does not remove items
validation.properties.{p}.keyFieldsOverride which fields are used as identifiers for dedup and scoring weight. Auto-detected if not set.
validation.properties.{p}.fieldHintsOverride scoring behavior for specific fields. Each hint is an object with a scoring key.
confidence.properties.{p}.thresholdComposite confidence below which a field is flagged, as a COMPOSITE flag. A JSON number in [0, 1]; anything else is ignored
rulesThe schema version's business rules

Available field hints:

HintEffectUse for
system_assignedAlways scores 1.0Fields assigned by the system (partNumber, auto-generated IDs)
default_absentScores 0.8 when value is empty, normal scoring otherwiseFields where empty means "not applicable" (tolerances, surface treatment)
literalEnables hallucination detection — value must appear verbatim in source textFields that reference entities in the source (filenames, document IDs, reference codes)

Business Rules

Business rules are deterministic checks over an extraction's output: a VAT ID that matches a pattern, a total that equals the sum of its line items, a date in ISO form. They need no model call and no calibration data. The same output always gets the same verdict.

Rules are part of the schema version's extraction configuration, in a top-level rules array, and are set by support like the rest of it. They run on every extraction of that version, right after schema validation and before scoring. validation.enabled does not switch them off; each rule has its own enabled flag.

Rules only record. Every outcome, passes included, is written to resultMetadata.rules; nothing is blocked, the status does not change and nothing is retried.

{
  "validation": { "enabled": true },
  "rules": [
    {
      "id": "vat-id-format",
      "type": "REGEX",
      "field": "/supplier/vatId",
      "pattern": "^DE[0-9]{9}$",
      "severity": "ERROR",
      "automationBlocker": true,
      "onFail": "FLAG"
    }
  ]
}

Rule fields

Every rule accepts these fields, whatever its type:

FieldRequiredDefaultMeaning
idYes—Names the rule in every outcome it produces. Unique within the version's rules
typeYes—REGEX, RANGE, ENUM, ARITHMETIC, CROSS_FIELD or FORMAT
severityNoWARNINGERROR, WARNING or INFO: how much a failure matters
automationBlockerNofalseWhether a failure should stop straight-through processing
onFailNoWARNWARN, FLAG, REJECT or RETRY. Recorded, not enforced (see below)
enabledNotrueA disabled rule produces no outcome at all
failOnMissingNofalseReport FAIL instead of SKIPPED when the input is absent

The type-specific fields (field, pattern, min, max, …) sit next to these.

Enum values are case-sensitive: "severity": "error" is not ERROR, so that rule is invalid and reports an ERROR verdict. An explicit null for severity or onFail also makes the rule invalid, while a null for a boolean field falls back to its default. When two rules share an id, the first one is evaluated and every later one reports ERROR with duplicate rule id.

Field pointers

Rules address values with RFC 6901 JSON pointers, the same format as fieldPath in schema validation. The empty pointer is the document root.

A pointer may contain at most one * segment, which expands over an array. The rule then produces one outcome per location that resolves, each with its own concrete pointer. /items/*/price over three items, of which only the first two have a price, produces two outcomes (/items/0/price and /items/1/price) and nothing for the third item. Two wildcards, or a non-empty pointer that does not start with /, is a configuration error.

A pointer into something the document does not have (an absent key, an index out of range, a wildcard over a non-array) is a miss, not an error. When the pointer as a whole resolves to nothing, the rule reports one SKIPPED (or FAIL with failOnMissing) at the pointer as configured, wildcard included. An explicit null at a location that did resolve reports SKIPPED (or FAIL) at that location.

An object key literally named * cannot be addressed, because * is always the wildcard.

Rule types

TypeAddressesType-specific configHolds whenOutcomes
REGEXfieldpattern (Java regex)The string matches the pattern in fullOne per resolved location
RANGEfieldmin and/or max (inclusive)The number lies within the boundsOne per resolved location
ENUMfieldvalues (non-empty array)The value equals one of the valuesOne per resolved location
ARITHMETICfield (no wildcard)equals: {op, operands}, tolerance|field − sum of operands| ≤ toleranceExactly one
CROSS_FIELDleft, right (no wildcards)operatorThe relation holds between the two valuesExactly one
FORMATfieldformat (named format)The string satisfies the named formatOne per resolved location

Values are read at their JSON type and never converted. A number written as a string, or a string where a number is expected, reports ERROR. Numbers are compared as exact decimals, so 1.50 equals 1.5 and 0.1 + 0.2 is exactly 0.3.

REGEX

{
  "id": "iban-shape",
  "type": "REGEX",
  "field": "/supplier/iban",
  "pattern": "^[A-Z]{2}[0-9]{2}[A-Z0-9]{11,30}$"
}

The pattern must match the whole value, not just occur in it, so "AT123 DE123456789 xx" does not satisfy a VAT ID pattern. The syntax is Java's regular expression syntax. Strings only: a number reports ERROR, as does a missing or invalid pattern.

RANGE

{
  "id": "vat-rate-plausible",
  "type": "RANGE",
  "field": "/taxRate",
  "min": 0,
  "max": 0.27
}

Both bounds are inclusive and at least one is required. A non-numeric bound, or a min above max, reports ERROR. So does a non-numeric value at field.

ENUM

{
  "id": "currency-supported",
  "type": "ENUM",
  "field": "/items/*/currency",
  "values": ["EUR", "CHF", "GBP"]
}

Values are compared by their text form, case-sensitively, so "de" does not satisfy ["DE"]. Two numbers compare as decimals, so 1.0 satisfies [1]. Mixed types fall back to text: the string "1" satisfies [1], but the number 1.0 does not satisfy ["1"]. An object or array at field reports ERROR. A missing or empty values array reports ERROR.

ARITHMETIC

{
  "id": "total-equals-line-items",
  "type": "ARITHMETIC",
  "field": "/totals/net",
  "equals": { "op": "SUM", "operands": ["/items/*/lineNet"] },
  "tolerance": 0.01
}

op is SUM; any other operator reports ERROR. tolerance is an absolute amount, 0 by default, and a difference exactly equal to it passes. Operands may use a wildcard, which is how a line-item sum is written; field may not.

An ARITHMETIC rule always produces one outcome. Its actual text shows both sides and the number of addends, for example 119.05 vs 119 (2 operands). Operands that resolve only partly are summed over what exists, and a null operand is not counted. The rule reports SKIPPED when field is absent or when no operand resolves; the two outcomes look identical, so check the output to see which side was missing.

CROSS_FIELD

{
  "id": "net-not-above-gross",
  "type": "CROSS_FIELD",
  "left": "/totals/net",
  "operator": "LTE",
  "right": "/totals/gross"
}

operator is EQ, NE, LT, LTE, GT or GTE. Two numbers compare as exact decimals. Two strings can only be compared with EQ and NE; ordering strings reports ERROR. Any other pairing (a number against a string, booleans, objects) reports ERROR. Neither side may use a wildcard. An absent or null side reports SKIPPED, and the outcome names the pointer of the missing side.

FORMAT

{
  "id": "invoice-date-iso",
  "type": "FORMAT",
  "field": "/invoiceDate",
  "format": "DATE_ISO"
}
FormatAcceptsEdges
DATE_ISOStrict ISO 8601 calendar dates (2026-08-06)No lenient parsing: 2026-8-6 and 06.08.2026 fail
EMAILOne @, a dotted domain, a top-level domain of 2+ lettersA shape check, not a deliverability check. Deliberately loose
URLAbsolute http and https URLs with a hostInternational domains must be punycoded (xn--…); a host containing _ is rejected
IBAN15–34 characters, country code, check digits, valid mod-97 checksumSpaces are removed and letters upper-cased before checking

An unknown format name reports ERROR. Strings only: a date extracted as a number reports ERROR.

Verdicts

Every outcome has exactly one verdict, and the report counts each:

VerdictMeaningWhat it tells you
PASSThe rule was evaluated and heldEvidence the value is not wrong in this way
FAILThe rule was evaluated and did not holdA problem in the document
SKIPPEDThe input the rule needed was absentThe document lacks the field
ERRORThe rule could not be evaluatedA problem in the rule: its configuration, or a value type it will not convert

Watch failCount for data quality, errorCount for rules that need fixing and skippedCount for coverage: a rule that is always SKIPPED is not checking anything. No verdict ever fails the extraction. See Rule Verdict Precedence for which verdict wins when several apply.

Severity vs automationBlocker

The two are independent. severity says how much a failure matters: ERROR for a defect in the document, WARNING for something that needs human attention, INFO for observation only. automationBlocker says whether a failure should stop straight-through processing. A rule can be INFO and blocking, or ERROR and non-blocking.

Both are copied onto every outcome of the rule, including passes: they describe the rule, the verdict describes what happened. Both also feed into composite confidence.

onFail is recorded, not enforced

onFail states what should happen when a rule fails. Today all four actions behave as WARN: the action is validated and recorded on every outcome, and the extraction result is returned unchanged whatever the verdict.

ActionIntended behavior
WARNRecord and continue (the current behavior for all four)
FLAGMark the result for review
REJECTReject the extraction result
RETRYRetry the extraction

If a failing rule should stop your workflow now, enforce it in your own code from resultMetadata.rules.

Reading Rule Outcomes

The report is in resultMetadata.rules:

{
  "applied": true,
  "passCount": 1,
  "failCount": 1,
  "skippedCount": 1,
  "errorCount": 1,
  "truncated": false,
  "outcomes": [
    {
      "ruleId": "gross-equals-net-plus-vat",
      "field": "/totals/gross",
      "verdict": "FAIL",
      "expected": "sum(/totals/net, /totals/vat) ± 0.01",
      "actual": "119.05 vs 119 (2 operands)",
      "severity": "ERROR",
      "automationBlocker": true,
      "action": "REJECT"
    },
    {
      "ruleId": "vat-id-format",
      "field": "/supplier/vatId",
      "verdict": "PASS",
      "expected": "^DE[0-9]{9}$",
      "actual": "DE123456789",
      "severity": "ERROR",
      "automationBlocker": true,
      "action": "FLAG"
    },
    {
      "ruleId": "invoice-date-iso",
      "field": "/invoiceDate",
      "verdict": "SKIPPED",
      "expected": "value present",
      "actual": "absent",
      "severity": "WARNING",
      "automationBlocker": false,
      "action": "WARN"
    },
    {
      "ruleId": "tax-rate-plausible",
      "field": "/taxRate",
      "verdict": "ERROR",
      "expected": "numeric 'min'",
      "actual": "non-numeric 'min': ten",
      "severity": "WARNING",
      "automationBlocker": false,
      "action": "WARN"
    }
  ]
}
FieldDescription
ruleIdThe rule's id. Empty when the rule had no id, which is then the reason for its ERROR
fieldThe concrete pointer the rule was evaluated at. When nothing resolved, the pointer as configured, wildcard included. Empty when no location applies, such as a duplicate id or an unreadable rule
verdictPASS, FAIL, SKIPPED or ERROR
expected, actualDisplay text only, and the wording can change. For an ERROR, actual carries the reason (unknown rule type 'REGEXP', min 10 exceeds max 1, duplicate rule id, wildcard not allowed in 'field', …). Match on ruleId, field and verdict, never on these
severity, automationBlockerAs configured on the rule
actionThe configured onFail. Recorded, not enforced

At most 200 outcomes are recorded, because passes are recorded too and a rule multiplies over array items. The four counts are always the true totals, and truncated is true when the list was cut. Read the counts, not the length of the list. expected and actual are cut at 500 characters and marked with a trailing ….

applied: false never means "all rules passed"; that is applied: true with failCount: 0. It means nothing was evaluated: the version has no rules, every rule is disabled, or rule evaluation could not run for this extraction. The three look identical in the result. validation.enabled: false does not cause it.

A rule that is always SKIPPED usually has a pointer that does not match the real shape of the output. Compare it with an output you got back, not with the schema. Set failOnMissing: true if absence is itself the problem, but note that array elements without the field never produce an outcome, so failOnMissing cannot flag them; require the field in the schema instead. A rule that is always ERROR needs fixing: read actual for the reason.


Advanced

Key Field Auto-Detection

When keyFields is not set in your extraction config, 2kw.ai auto-detects identifier fields from your schema using these patterns:

PatternExamples
Exact name: id, name, keyid, name
Ends with: id, name, file, key, code, numbersourceFile, partNumber, materialCode
Starts with: source, filesourceDocument, fileName

These fields serve double duty: they're used for dedup grouping (matching items across chunks) and are weighted 3x in grounding scores (a hallucinated identifier dominates the item score).

If auto-detection picks the wrong fields — or misses yours — set keyFields explicitly in your extraction config.

Dedup Merge Algorithm

Understanding how merging works helps when debugging unexpected results in chunked extractions. Here's a complete example before we break down each step.

End-to-end example

A 20-page bill of materials is too large for a single extraction. 2kw.ai splits it into two overlapping chunks. The item housing.geo spans the boundary — Chunk 1 sees the beginning of it (material is mentioned) but cuts off before the quantity. Chunk 2 picks up in the overlap region and sees the quantity, but the material description is already behind it.

Chunk 1 extracts 3 items, the last one incomplete:

{
  "parts": [
    { "sourceFile": "gear.geo",    "material": "1.4301",    "quantity": 5,    "partNumber": 1 },
    { "sourceFile": "shaft.geo",   "material": "Steel",     "quantity": 2,    "partNumber": 2 },
    { "sourceFile": "housing.geo", "material": "Aluminum",  "quantity": null,  "partNumber": 3 }
  ]
}

Chunk 2 extracts housing.geo again (from the overlap) plus one new item:

{
  "parts": [
    { "sourceFile": "housing.geo", "material": null,    "quantity": 3,    "partNumber": 1 },
    { "sourceFile": "bracket.geo", "material": "1.4301", "quantity": 1,    "partNumber": 2 }
  ]
}

Now 2kw.ai merges the two outputs:

1. Detect key fields. sourceFile matches the *file suffix pattern → used as the dedup key. partNumber matches *number → marked as a sequence field.

2. Group by key. Two items share the key sourceFile: "housing.geo" — one from each chunk. The other three items (gear.geo, shaft.geo, bracket.geo) are unique and pass through unchanged.

3. Merge the group. Both housing.geo items have 3 non-null fields (tie). The one from Chunk 1 was encountered first, so it becomes the base. Its quantity is null → filled with 3 from Chunk 2's item:

Base (Chunk 1):  { sourceFile: "housing.geo", material: "Aluminum", quantity: null, partNumber: 3 }
Fill from Chunk 2:                                                   quantity: 3  ←── null filled
Result:          { sourceFile: "housing.geo", material: "Aluminum", quantity: 3,    partNumber: 3 }

4. Resequence. After merge, partNumber values are 1, 2, 3, 2 — duplicate 2. Renumbered to 1, 2, 3, 4.

Final result:

{
  "parts": [
    { "sourceFile": "gear.geo",    "material": "1.4301",    "quantity": 5, "partNumber": 1 },
    { "sourceFile": "shaft.geo",   "material": "Steel",     "quantity": 2, "partNumber": 2 },
    { "sourceFile": "housing.geo", "material": "Aluminum",  "quantity": 3, "partNumber": 3 },
    { "sourceFile": "bracket.geo", "material": "1.4301",    "quantity": 1, "partNumber": 4 }
  ]
}
{
  "dedup": {
    "applied": true,
    "itemsRemoved": 0,
    "itemsMerged": 1,
    "resequenced": true,
    "originalItemCount": 5,
    "finalItemCount": 4
  }
}

The incomplete housing.geo from Chunk 1 and the incomplete one from Chunk 2 were combined into a single complete item. Neither chunk had all the data, but together they did.

Step 1: Group by key fields

Items are grouped by the normalized values (lowercase, trimmed) of their key fields. Items where all key fields are null cannot be fingerprinted and are kept as-is — they are never merged.

If no key fields are detected at all, 2kw.ai falls back to pairwise similarity matching (items with 80%+ field value overlap are grouped).

Step 2: Merge each group

Within a group, the item with the most non-null fields becomes the base. Then, for each remaining item, any field that is null in the base is filled from the other item. If completeness is equal, the item encountered first (from the earlier chunk) becomes the base.

The end-to-end example above shows the common case: each chunk sees a different part of the item, and the merge fills the gaps.

Edge case: conflicting values. When both chunks have the same field with different non-null values, 2kw.ai uses provenance-aware conflict resolution: it prefers the value from the "authority" chunk — the chunk that contains the item's identifier (e.g., the chunk where sourceFile: "gear.geo" actually appears in the text). The authority chunk is most likely to have seen the item's header and metadata, making its scalar values more reliable.

Chunk 1: { sourceFile: "gear.geo", material: "Steel" }              (authority — "gear.geo" appears in chunk 1 text)
Chunk 2: { sourceFile: "gear.geo", material: "1.4301", quantity: 5 }

→ Base: Chunk 2 (more fields), but material overridden from Chunk 1 (authority)
→ Result: { sourceFile: "gear.geo", material: "Steel", quantity: 5 }

If no authority chunk can be determined (the identifier doesn't appear in any chunk text, or both chunks contain it), the base item's value wins (most-complete-item-first, same as before).

Edge case: nested arrays. When both chunks extract array fields (like contours or holes) for the same item, the arrays are concatenated and deduplicated rather than one replacing the other. This is critical for items that span chunk boundaries — Chunk 1 might extract contours 1-10, Chunk 2 might extract contours 8-13, and the merge produces the complete set 1-13 (with duplicates 8-10 removed).

Step 3: Resequence

After merging, sequential number fields (names ending in number, index, num, order, sequence, position, or named nr/pos) are checked for broken numbering. If gaps or duplicates are found, the field is renumbered starting at 1.

Rule Verdict Precedence

When more than one business rule verdict could apply, these decide:

  • A broken rule beats an absent field. A rule's configuration is checked before the document is read. A rule with an invalid pattern, an empty values array or a min above max reports ERROR even when its field is also absent, so it counts in errorCount, never in skippedCount. A rule that can never work says so instead of reporting SKIPPED forever.
  • A broken rule is one ERROR, not one per location. A configuration problem on a rule whose pointer covers fifty array items produces a single ERROR, because the problem belongs to the rule. A value-level ERROR, such as a wrong type at one location, is still reported per location, because it belongs to the document.
  • failOnMissing turns SKIPPED into FAIL and nothing else. It does not affect ERROR. It reaches a pointer that resolved to nothing (one FAIL at the configured pointer) and an explicit null at a resolved location (one FAIL there).
  • A rule that crashes is that rule's ERROR. It never suppresses the rules around it and never fails the extraction. The exception is a pattern so pathological that evaluation cannot continue at all; the whole report is then recorded as applied: false.

Re-running Extractions

POST /v1/extractions/{id}/rerun

Need to retry an extraction with the same config? Hit the rerun endpoint. Creates a new extraction —the original stays untouched.

The re-run copies the original's schema version, input text and model, and it runs the same way the original did: an extraction that was queued asynchronously is queued again and comes back PENDING, a synchronous one is executed inline and comes back with its result.

Not every extraction can be re-run:

ResponseWhen
404 Not FoundNo such extraction in your organization
409 ConflictThe extraction is still PENDING or PROCESSING
409 ConflictThe extraction was vision-only —images are not stored and cannot be replayed
409 ConflictThe schema version it ran against has since been deleted

COMPLETED and FAILED extractions are eligible. Each re-run is billed as a new extraction.

Listing and Filtering

GET /v1/extractions

ParameterTypeDescription
searchstringFilter by model name
schemaVersionIdstringFilter by schema version
statusstringPENDING, PROCESSING, COMPLETED, or FAILED
pagenumberPage number (0-based)
sizenumberPage size (default: 20)

Use Cases

  • Contact extraction —names, emails, phone numbers from unstructured text
  • Invoice processing —invoice numbers, dates, amounts, line items from documents
  • Resume parsing —skills, experience, education from CVs
  • Document analysis —key fields from contracts, reports, forms
  • Data entry automation —turn free-text notes into structured database records

Was this page helpful?