Queues

An annotation queue is a worklist for human review. You fill it with traces, spans, dataset items or experiment results, reviewers work through it item by item, and every judgement they make is stored as a human score next to the automated evaluator scores.

Why Queues

Automated evaluators tell you what they were built to measure. They don't tell you what you haven't thought of yet. Reading real production traces is how you find the failure modes worth measuring in the first place — and how you check that an LLM judge agrees with a person.

A queue makes that reading systematic:

  • A defined sample instead of whatever trace you happened to open.
  • Progress you can see — how many items are done, how many are left.
  • Scores you can aggregate — every judgement is filed against a rubric, not buried in a chat thread.

How It Works

A queue holds items. Each item points at one subject:

Subject typeWhat it isScored in the console
TRACEA whole trace from ObservabilityYes, in the queue's review panel
SPANA single span inside a traceAPI only
DATASET_ITEMAn item of a datasetAPI only
RUN_RESULTOne result of an experiment runAPI only (or on the experiment's comparison view)

A subject appears at most once per queue. The same subject can sit in several queues.

Every item moves through four states:

StatusMeaning
PENDINGNobody has scored it yet
IN_PROGRESSAt least one score was recorded for it through the queue
DONEA reviewer marked it done
SKIPPEDA reviewer skipped it, or its subject was deleted (can be set back to PENDING)

The first score recorded against an item moves it from PENDING to IN_PROGRESS automatically. DONE is never set automatically — a reviewer marks an item done when they have finished with it. Queue progress is done / (total − skipped): skipped items don't count against you.

When a subject is deleted — a trace is removed, for example — every queue item pointing at it becomes SKIPPED and loses its assignee. That can move a queue's progress without anyone reviewing anything.

Creating a Queue

In the console, open Queues in the sidebar and create a queue with a name and an optional description. Names are unique among an organization's active queues.

Request

// POST /v1/annotation-queues
{
  "name": "Support agent — weekly review",
  "description": "50 sampled traces per week, split by outcome"
}

Response

{
  "id": "7c1e…",
  "organizationId": "org_123",
  "name": "Support agent — weekly review",
  "description": "50 sampled traces per week, split by outcome",
  "createdBy": "user_42",
  "archivedAt": null,
  "createdAt": "2026-09-27T10:00:00Z",
  "lastModifiedAt": "2026-09-27T10:00:00Z",
  "itemCounts": { "pending": 0, "inProgress": 0, "done": 0, "skipped": 0 }
}

PATCH /v1/annotation-queues/{id} changes the name or description; a field you leave out stays as it is. POST /v1/annotation-queues/{id}/archive archives a queue — archiving twice is harmless. Archived queues drop out of the default list; GET /v1/annotation-queues?archived=true lists them.

Adding Items

There are three ways to fill a queue.

Sample traces. On a queue's page, Sample traces draws a random sample from the 500 most recent traces, optionally filtered to one source service. Choose how to split the sample:

  • No grouping — one random sample of the given size.
  • Outcome (OK / Error) — the given number of traces from each trace status (OK, ERROR, and UNSET when traces without a status are in the pool), so errors are not drowned out by successes.
  • Source service — the given number of traces from each service. A trace that spans several services counts for the first one listed only.

Add to queue from Traces. Select traces in the traces list and choose Add to queue, or open a trace and use the queue button in the span detail: on the root span it queues the whole trace, on any other span it queues just that span. You can pick an existing queue or create one in the same dialog.

Bulk add via the API. Send up to 1000 subjects per call. Subjects already in the queue are skipped, and the response lists only the items that were newly added:

Request

// POST /v1/annotation-queues/{id}/items
{
  "items": [
    { "subjectType": "TRACE", "subjectId": "4bf92f3577b34da6a3ce929d0e0e4736" },
    { "subjectType": "SPAN", "subjectId": "00f067aa0ba902b7" },
    { "subjectType": "DATASET_ITEM", "subjectId": "dsi_901" }
  ]
}

Reviewing Items

On a queue's page, Start annotating opens the next item in a panel beside the item list. The panel shows the trace — span tree and span detail — on the left and the annotation column on the right. The annotation column needs a wide screen (1280 px and up); on narrower screens the panel shows the trace only.

The annotation column has two sections:

  • Annotation — free-text notes, one field per note-only Human review evaluator. Write down what you see before you decide how to score it. If your organization has no note-only evaluator, a single default note field is shown. A note can only be saved once it has text.
  • Failure modes — one scoring widget per Human review evaluator with labels or a range (for example Pass / Partial / Fail, or True / False). Every widget starts empty, and Save stays disabled until you pick a label or a score.

Each widget saves on its own. Saving again replaces your earlier score for that rubric — you can revise a judgement at any time.

The bar above the trace holds the item actions:

  • Assign to me — make yourself the item's assignee. While it is assigned, only the assignee can score it through the queue. This is not a lock: any member can reassign the item.
  • Skip — set the item to SKIPPED.
  • Mark done — set the item to DONE.

The review panel works on TRACE items. SPAN, DATASET_ITEM and RUN_RESULT items appear in the item list, but you score them through the API.

Picking the next item via the API

GET /v1/annotation-queues/{id}/items/next returns the oldest item that is PENDING or IN_PROGRESS and is either unassigned or assigned to you. It answers 204 No Content when there is nothing left for you — an empty queue is a normal state, not an error.

The call does not assign the item. To hold it for yourself, set the assignee:

Request

// PATCH /v1/annotation-queues/{id}/items/{itemId}
{ "assignedTo": "user_42" }

The same call changes the status ({ "status": "DONE" }). Send "assignedTo": "" to unassign. GET /v1/annotation-queues/{id}/items lists items with the optional filters status and assignedTo, paginated with page (from 0) and size (default 50).

To find out whether something is already queued, GET /v1/annotation-queues/by-subject?subjectType=TRACE&subjectId=… returns every queue item that points at that subject, across all queues.

Human Scores

A human score is an evaluation score with source: "HUMAN" and the reviewer's user id as annotatorId. Human scores live in the same store as automated evaluator scores, so they show up next to them — a human score on an experiment result appears in that result's scores array.

Recording a score

Request

// POST /v1/evaluation-scores/human
{
  "subjectType": "TRACE",
  "subjectId": "4bf92f3577b34da6a3ce929d0e0e4736",
  "evaluatorId": "b3d1…",
  "label": "FAIL",
  "comment": "Answered from the wrong knowledge base",
  "annotationQueueItemId": "9a40…"
}
FieldRequiredDescription
evaluatorIdYesThe id of a Human review evaluator template. The legacy id human_quality is still accepted and takes free-form labels.
subjectType, subjectIdYes, except for run resultsWhat you are scoring: TRACE, SPAN, DATASET_ITEM or RUN_RESULT.
runResultIdFor RUN_RESULT onlyThe run result id. Leave subjectType / subjectId out, or set subjectId to the same id. Must be absent for every other subject type.
scoreNoA number from 0.0 to 1.0.
labelNoA category.
commentNoA free-text note. A score with only a comment is a valid annotation.
annotationQueueItemIdNoThe queue item this score belongs to.

One score per reviewer per rubric. Scores are keyed by subject, evaluator and reviewer: posting again for the same three overwrites your earlier score instead of adding a second one. Two reviewers scoring the same subject produce two scores.

The rubric is enforced. A score is rejected with 400 when:

  • evaluatorId is neither human_quality nor a template of type HUMAN,
  • the template has categories and label is not one of them,
  • the template is numeric and score falls outside its minValue / maxValue,
  • score is outside 0.0–1.0, whatever the template says.

True / False templates don't restrict label on the server.

With annotationQueueItemId, the item's subject must be the subject you score, and if the item is assigned, only its assignee may score it through the queue. The first such score moves a PENDING item to IN_PROGRESS.

Reading scores

GET /v1/evaluation-scores/human?subjectType=TRACE&subjectId=… returns the human scores for one subject. Without annotatorId (or with annotatorId=me) you get your own scores; pass another reviewer's user id to read theirs.

Scoring experiment results

The result detail panel of an experiment's comparison view has a human-score widget for each variant's result, so you can score results without putting them in a queue first. Those scores are stored as RUN_RESULT human scores like any other.

Permissions

Queues, queue items and human scores need the Member role or higher; a Viewer is refused with 403. Everything is scoped to the current organization.

API Reference

MethodEndpointDescription
GET/v1/annotation-queuesList active queues (?archived=true for archived ones)
POST/v1/annotation-queuesCreate a queue
GET/v1/annotation-queues/{id}Get a queue with its item counts
PATCH/v1/annotation-queues/{id}Update name or description
POST/v1/annotation-queues/{id}/archiveArchive a queue
GET/v1/annotation-queues/by-subjectQueue items for one subject, across queues
GET/v1/annotation-queues/{id}/itemsList items (status, assignedTo, page, size)
POST/v1/annotation-queues/{id}/itemsBulk-add up to 1000 subjects
GET/v1/annotation-queues/{id}/items/nextNext item for you, or 204
GET/v1/annotation-queues/{id}/items/{itemId}Get one item
PATCH/v1/annotation-queues/{id}/items/{itemId}Set status or assignee
POST/v1/evaluation-scores/humanRecord or revise a human score
GET/v1/evaluation-scores/humanHuman scores for one subject

CLI

backbone queues create -n "Support agent — weekly review"
backbone queues items add --queue <queueId> \
  --items '[{"subjectType":"TRACE","subjectId":"4bf92f3577b34da6a3ce929d0e0e4736"}]'
backbone queues next <queueId>
backbone scores human --subject-type TRACE --subject-id <traceId> \
  --evaluator <templateId> --label FAIL --queue-item <itemId>
backbone queues items update <itemId> --queue <queueId> --status DONE
backbone scores list --subject-type TRACE --subject-id <traceId>

backbone queues --help lists the rest: list, get, update, archive, by-subject, and items list / items get.

What's Next

Was this page helpful?