← Blog

Decision models for robotics data curation

A lot of robot data curation is small multiple-choice calls about text. Does this report back up its verdict? Which slot does this note mean? Is this the episode the curator was looking for? A decision model is built for exactly that. You hand it the text and the options, and it picks one and tells you how sure it is, in well under a second. We've been using Jev from TypeSafe for this. OpenAI just opened its own, the Decision API, so we ran both through every test we have on our robot data, next to the big general models and our own harness.

17 min read engineering

Setup

What we tested, and how

Everything here comes from real recordings of two robots. One is a two-arm robot restocking a shelf, mostly from our hard set, where most attempts go wrong on purpose. The other packs clothes into a bag. The text tests use what our pipeline actually produces: a vision model’s write-up of an episode, a curator’s notes, and the actions a video model detected. Every model got the same questions, the same options and the same answer keys.

One wrinkle. Jev takes a few worked examples with every question, and a lot of our curation leans on them. OpenAI’s Decision API has nowhere to put them. So we asked it two ways: with the examples pasted into the question, which is what Jev sees, and with just the question. The second one is called rules only in the charts.

Speed and the routine stuff

Both are fast, and both get the routine reading right

Both decision models answer a text question in about 0.3 seconds. The fastest general model we tried takes 3.1. That adds up when a dataset throws up thousands of these questions and someone is waiting on the answers. OpenAI’s also takes images, at 0.5 seconds a case, against 4.9 for the fastest vision model.

Time to answer one question. Median seconds; text: every question in the seven reading tests; images: every case in the five image tests. Reading text, OpenAI Decision API, rules only, 0.3 s, OpenAI Decision API, 0.3 s, Jev, 0.3 s, gpt-6-astra, 3.1 s, GPT-6 Luna, 3.2 s, gpt-6.1-sol, 4.1 s, Claude Fable 5.1, 6.0 s, Gemini 3.8 Flash, 6.7 s, Looking at images, OpenAI Decision API, 0.5 s, gpt-6.1-sol, 4.9 s, gpt-6-astra, 5.0 s, GPT-6 Luna, 5.1 s, Gemini ER 2, 5.2 s, Gemini 3.8 Flash, 9.5 s.

On the routine reading they’re both right nearly every time. Turning a curator’s note into fields, reading a phrase like “bottom right of the shelf”, checking whether two reports say the same thing, sorting a written reason into a defect type: OpenAI’s got 95% or better on each with the worked examples, about where Jev and the big models are.

Routine reading, share answered right on each test. A curator's note, into fields: OpenAI Decision API 99.8%, rules only 98%, Jev 99.9%, the general models 99.8%. A note against the curator's own label: OpenAI Decision API 95%, rules only 95%, Jev 95%, the general models 95%. A location phrase, into slot and shelf: OpenAI Decision API 99%, rules only 97%, Jev 98%, the general models 99%–99.6%. Do two reports agree?: OpenAI Decision API 96%, rules only 92%, Jev 95%, the general models 95%–97%. A reason, into a defect type: OpenAI Decision API 97%, rules only 98%, Jev 99%, the general models 96%–99.5%.

Confidence

How much can you let through without checking?

The way we’d actually use a decision model: take the confident answers and send the rest to a person. OpenAI’s guide says to set that cut-off from your own labelled examples, so we did that for every model and asked how many answers each one can let through while keeping mistakes at 1% or less.

Jev lets through 93% of its answers. OpenAI’s Decision API lets through 66%, or 71% without the worked examples. So on the same questions, a person would end up checking about 34% of OpenAI’s answers and 7% of Jev’s.

The big general models do even better here (98% for gpt-6.1-sol) because they’re right more often to begin with. They also take ten times as long per question. Claude Fable 5.1 scores as well, but it wrote the answer key for these tests, so take that with a pinch of salt.

How many answers can go through without a person checking them?. All 2,186 answers on the seven reading tests; each model's confidence cut-off set so that at most 1% of the answers it lets through are wrong. Answers let through (more is better), OpenAI Decision API, 66%, OpenAI Decision API, rules only, 71%, Jev, 93%, Claude Fable 5.1, 98%, gpt-6.1-sol, 98%, gpt-6-astra, 97%, GPT-6 Luna, 88%, Gemini 3.8 Flash, 87%.

Close calls

Where the two decision models split

The tests that actually separate the models need a real call. Does a vision model’s own write-up back up its verdict? Does the count in a report add up? On these, OpenAI’s Decision API mostly said the text doesn’t settle it.

On the report check it caught 25% of the contradictions, and sent back 50% of the good verdicts for no reason. Jev caught 75% and sent back none. Dropping the examples helped a bit (50% caught, 16% sent back), but that’s still fewer catches than any other model.

Does the report back the verdict?. 52 verdicts from the hard set with the vision model's written report: 8 contradicted by it, 44 backed by it. Contradictions caught, % of 8 (more is better), OpenAI Decision API, 25%, OpenAI Decision API, rules only, 50%, Jev, 75%, Claude Fable 5.1, 100%, gpt-6.1-sol, 88%, GPT-6 Luna, 100%, gpt-6-astra, 88%, Gemini 3.8 Flash, 100%, Good verdicts sent back, % of 44 (fewer is better), OpenAI Decision API, 50%, OpenAI Decision API, rules only, 16%, Jev, 0%, Claude Fable 5.1, 5%, gpt-6.1-sol, 11%, GPT-6 Luna, 14%, gpt-6-astra, 18%, Gemini 3.8 Flash, 36%.

The count check is just arithmetic, and one line of code gets every one right. With the examples, OpenAI’s sent back every one of the counts that add up, and caught 13% of the ones that don’t. Jev caught 93% and sent back 18%.

Does the count add up?. 49 reports from the hard set that claim a number of placements: 15 whose box counts disagree, 34 that add up. Wrong counts caught, % of 15 (more is better), OpenAI Decision API, 13%, OpenAI Decision API, rules only, 47%, Jev, 93%, gpt-6-astra, 100%, gpt-6.1-sol, 100%, Claude Fable 5.1, 100%, Gemini 3.8 Flash, 100%, GPT-6 Luna, 93%, Right counts sent back, % of 34 (fewer is better), OpenAI Decision API, 100%, OpenAI Decision API, rules only, 53%, Jev, 18%, gpt-6-astra, 0%, gpt-6.1-sol, 0%, Claude Fable 5.1, 0%, Gemini 3.8 Flash, 0%, GPT-6 Luna, 18%.

We don’t think it’s the model underneath. The Decision API runs on GPT-6 Luna, and plain Luna, asked through the normal API with the same examples, caught all of the contradictions and 93% of the bad counts. Something about how the Decision API reads long questions pushes it to “not enough to tell”. It said that on 73% of its answers across these two tests.

If a tool says “not enough to tell” on clear cases, a person ends up reading them anyway.

Head to head

When they disagree, Jev is usually right

Across all 2,186 answers, the two picked different options 7% of the time. When they did, Jev was right 88% of the time, OpenAI’s 11%, and neither 1%. Most of the splits are in search, the count check and the report check.

When Jev and the OpenAI Decision API disagree, who is right?. The 150 answers, out of 2,186, where the two picked different options, by test. Jev right, OpenAI Decision API right, neither, Search the notes, 55, 7, 62, Does the count add up?, 40, 41, Does the report back the verdict?, 27, 29, Location phrases, 3, 4, 7, Do two reports agree?, 4, 6, Reason into defect type, 4, 4, Note into fields, 1.

Search on its own: asked which notes match something like “a box fell on the bottom shelf”, OpenAI’s found 85% of the notes that do. Jev found 90%, gpt-6-astra 95%.

Search the notes in plain words. Ten searches over the hard set's 50 curator notes, such as “a box fell on the bottom shelf”. Matching notes found, % of 84 (more is better), OpenAI Decision API, 85%, OpenAI Decision API, rules only, 89%, Jev, 90%, gpt-6-astra, 95%, gpt-6.1-sol, 95%, Claude Fable 5.1, 95%, GPT-6 Luna, 95%, Gemini 3.8 Flash, 90%, Non-matching notes picked, % of 416 (fewer is better), OpenAI Decision API, 1%, OpenAI Decision API, rules only, 5%, Jev, 2%, gpt-6-astra, 0%, gpt-6.1-sol, 0%, Claude Fable 5.1, 2%, GPT-6 Luna, 4%, Gemini 3.8 Flash, 0%.

In the app

The job Jev actually does for us

In our app, Jev has one real job on robot recordings. A video model lists every action it sees in an episode (reach, grasp, lift, place), and Jev groups those into the curator’s phases: first small box, second large box, third large box, last small box, back home. For each phase it picks the action where that phase starts.

We gave OpenAI’s Decision API the same requests Jev got for all 50 hard-set episodes. Jev answers all five phases in one go. OpenAI’s guide says questions that depend on an earlier answer should go in separate requests, so we asked it both ways: all at once, and one phase per request with the earlier answers passed along.

Jev came back with all five phases in order on 46% of the episodes. OpenAI’s managed 36% one phase at a time and 36% all at once. An answer with the phases out of order is no use to anyone, and OpenAI’s gave one on 28% and 32% of the episodes, against Jev’s 20%.

Splitting a recording into the curator's phases. Share of all 50 hard-set episodes; each answer is five phase starts picked from the detected actions. dark: all five phases, in order, light: some phases left out, grey: out of order, unusable, Jev, 46%, 34%, 20%, OpenAI, one phase per request, 36%, 36%, 28%, OpenAI, all phases at once, 36%, 32%, 32%.

The curator timed one episode by hand, ep001. There OpenAI’s boundaries were closer: 1.0 seconds off on average, against Jev’s 1.75. One episode isn’t much to go on, though.

Both are steady. Ask the same questions twice and Jev gives the same answer 99.7% of the time, OpenAI’s 99.5%.

Images

Only OpenAI’s can look at pictures

Jev only reads text. OpenAI’s Decision API takes images too, so we also tried it on our image tests, next to the vision models and the Qualia harness. One test went fine. From the first and last frame, it caught 86% of the episodes where the robot didn’t finish, and called 7% of the finished ones unfinished. The vision models caught 90% and called 7%–10% of the finished ones unfinished, and took ten times as long.

The hardest thing we check is a fallen box. On this shelf a fallen box can look a lot like a standing one, even to a person. The hard set has 198 slots we have answers for, and 15% of them hold a fallen box. The Qualia harness caught 90% of the fallen boxes and called 0.6% of the standing ones fallen. Given exactly what the harness sees for each slot, OpenAI’s caught 69%, but it also called 36% of the standing boxes fallen, and someone would have to check each of those.

So the Qualia harness came out on top. The closest any single model got was gpt-6.1-sol, asked once per slot with the harness’s inputs: it caught 86% and called 2% of the standing boxes fallen. And those inputs do a lot of the work. They’re the curator’s examples, reference pictures of the slot and the close-up, which the harness puts together for every slot. Give Sol just the rule and it catches 38% and calls 22% of the standing boxes fallen.

A fallen box, slot by slot. 198 shelf slots at the end of a hard-set episode, 29 with a fallen box; every model sees the harness's own inputs for the slot. Fallen boxes caught, % of 29 (more is better), Qualia harness, 90%, OpenAI Decision API, 69%, gpt-6.1-sol, 86%, Gemini 3.8 Flash, 48%, GPT-6 Luna, 28%, Gemini ER 2, 21%, Standing boxes called fallen, % of 169 (fewer is better), Qualia harness, 0.6%, OpenAI Decision API, 36%, gpt-6.1-sol, 2%, Gemini 3.8 Flash, 17%, GPT-6 Luna, 6%, Gemini ER 2, 6%.

Whole recordings tell the same story. Asked to keep or remove an episode from a sheet of 16 frames, OpenAI’s got an F1 of 0.63. The harness gets 0.97, and gpt-6-astra looking at the same sheet gets 0.75–0.81. On the second robot it found every item left outside the bag, but it also flagged 27% of the good recordings. The harness flagged none.

Whole recordings: keep or remove, and an item left outside the bag. Left: 50 hard-set episodes, one 16-frame sheet each; right: 177 good recordings from the second robot (all found the 5 bad ones). Keep or remove, F1 (higher is better), Qualia harness, 0.97, Gemini 3.8 Flash, 0.91, gpt-6-astra, 0.85, Gemini ER 2, 0.82, gpt-6.1-sol, 0.81, OpenAI Decision API, 0.63, GPT-6 Luna, 0.55, Good recordings flagged, % of 177 (fewer is better), Qualia harness, 0%, Gemini ER 2, 2%, Gemini 3.8 Flash, 6%, gpt-6-astra, 8%, gpt-6.1-sol, 15%, OpenAI Decision API, 27%, GPT-6 Luna, 53%.

All results

The full table

Where a cell has two numbers, the first is the share it caught and the second the share it flagged by mistake, each out of the count in the row’s note. Everything is on the shelf robot’s hard set, except the location phrases, which use both shelf sets, and the bag test, which is the second robot.

The testOpenAI Decision APIOpenAI Decision API, rules onlyJevGPT-6 LunaGpt-6-astraGpt-6.1-solClaude Fable 5.1Gemini 3.8 FlashGemini ER 2Qualia harness
Reading text
Does the report back the verdict? (contradictions caught (of 8) · good verdicts sent back (of 44))25% · 50%50% · 16%75% · 0%100% · 14%88% · 18%88% · 11%100% · 5%100% · 36%——
Does the count add up? (wrong counts caught (of 15) · right counts sent back (of 34))13% · 100%47% · 53%93% · 18%93% · 18%100% · 0%100% · 0%100% · 0%100% · 0%——
Search the notes (matching notes found (of 84) · non-matching picked (of 416))85% · 1%89% · 5%90% · 2%95% · 4%95% · 0%95% · 0%95% · 2%90% · 0%——
A curator’s note, into fields (share right)99.8%98%99.9%99.8%99.8%99.8%99.8%99.8%——
A note against the curator’s own label (share right)95%95%95%95%95%95%95%95%——
A location phrase, into slot and shelf (share right)99%97%98%99%99.6%99.6%99.6%99.6%——
Do two reports agree? (share right)96%92%95%95%96%96%96%97%——
A reason, into a defect type (share right)97%98%99%98%97%97%99.5%96%——
Median seconds per question0.3 s0.3 s0.3 s3.2 s3.1 s4.1 s6.0 s6.7 s——
Looking at images
A fallen box, harness inputs (caught (of 29) · standing called fallen (of 169))69% · 36%——28% · 6%—86% · 2%—48% · 17%21% · 6%90% · 0.6%
A fallen box, rule only (caught (of 29) · standing called fallen (of 169))79% · 83%——62% · 50%45% · 22%38% · 22%—38% · 28%52% · 56%—
Is the task finished? (first and last frame; unfinished caught (of 21) · finished called unfinished (of 29))86% · 7%——90% · 7%90% · 10%90% · 10%—90% · 10%90% · 7%100% · 10%
Keep or remove the recording (one sheet of 16 frames; F1)0.63——0.550.850.81—0.910.820.97
An item left outside the bag (second robot; caught (of 5) · good flagged (of 177))100% · 27%——100% · 53%100% · 8%100% · 15%—100% · 6%100% · 2%100% · 0%
Median seconds per case0.5 s——5.1 s5.0 s4.9 s—9.5 s5.2 s—

Takeaway

So which would we use?

  • A decision model, for the small text questions. There are thousands of them in a dataset. It’s ten times faster than a general model, and when it says it’s sure, you can mostly take its word for it.
  • Jev, for now. It catches more of the close calls, gets the phases in order more often, and lets far more answers through without a person checking them. OpenAI’s is just as fast and just as good on the routine reading.
  • Not a decision model, for pictures. OpenAI’s is the only one of the two that takes images, and it’s quick, but it raises too many false alarms to curate on its own. On whole recordings no single model matched the harness.
  • Code, for counting. One line gets it right every time.

If we could ask OpenAI for one thing, it would be a proper place for worked examples. That’s how what the curator knows ends up in every question.


How we ran it. All OpenAI Decision API answers are from 7 October 2026, on the model OpenAI runs it on (GPT-6 Luna). The text tests come from our earlier decision-model study: questions written on 30 September from real curation text, an answer key written before any model ran (by Claude Fable 5.1), and nine disputed cases settled by a person. Jev was asked again the same week. The general models got the same text, options and examples in one prompt. The phase test reuses the exact requests Jev got on 30 September; the curator’s ep001 phases are one development example. The image tests are scored against the curator’s labels. Requests follow OpenAI’s Decision API guide: text as the input (images as inline base64), each question as plain-text instructions with a list of options. The guide has no field for worked examples, and for the phase test we also asked one phase per request, as it advises for dependent questions. Confidence cut-offs were set per model from the same labelled answers they are scored on, which flatters every model equally. An answer that isn’t one of the options counts as wrong, and “not enough to tell” counts as sending the case back.

Feasibility audit

Tell us how we can help.

0/500
Join the team

Apply to Qualia.

0/500

“It must be confessed, moreover, that perception, and that which depends on it, are inexplicable by mechanical causes, that is, by figures and motions. And, supposing that there were a mechanism so constructed as to think, feel and have perception, we might enter it as into a mill. And this granted, we should only find on visiting it, pieces which push one against another, but never anything by which to explain a perception. This must be sought, therefore, in the simple substance, and not in the composite or in the machine.”

What do you make of this in relation to robotics?
0/500