Decision models for robotics data curation
A lot of robot data curation is small multiple-choice calls about text. Does this report back up its verdict? Which slot does this note mean? Is this the episode the curator was looking for? A decision model is built for exactly that. You hand it the text and the options, and it picks one and tells you how sure it is, in well under a second. We've been using Jev from TypeSafe for this. OpenAI just opened its own, the Decision API, so we ran both through every test we have on our robot data, next to the big general models and our own harness.

Setup
What we tested, and how
Everything here comes from real recordings of two robots. One is a two-arm robot restocking a shelf, mostly from our hard set, where most attempts go wrong on purpose. The other packs clothes into a bag. The text tests use what our pipeline actually produces: a vision model’s write-up of an episode, a curator’s notes, and the actions a video model detected. Every model got the same questions, the same options and the same answer keys.
One wrinkle. Jev takes a few worked examples with every question, and a lot of our curation leans on them. OpenAI’s Decision API has nowhere to put them. So we asked it two ways: with the examples pasted into the question, which is what Jev sees, and with just the question. The second one is called rules only in the charts.
Speed and the routine stuff
Both are fast, and both get the routine reading right
Both decision models answer a text question in about 0.3 seconds. The fastest general model we tried takes 3.1. That adds up when a dataset throws up thousands of these questions and someone is waiting on the answers. OpenAI’s also takes images, at 0.5 seconds a case, against 4.9 for the fastest vision model.
On the routine reading they’re both right nearly every time. Turning a curator’s note into fields, reading a phrase like “bottom right of the shelf”, checking whether two reports say the same thing, sorting a written reason into a defect type: OpenAI’s got 95% or better on each with the worked examples, about where Jev and the big models are.
Confidence
How much can you let through without checking?
The way we’d actually use a decision model: take the confident answers and send the rest to a person. OpenAI’s guide says to set that cut-off from your own labelled examples, so we did that for every model and asked how many answers each one can let through while keeping mistakes at 1% or less.
Jev lets through 93% of its answers. OpenAI’s Decision API lets through 66%, or 71% without the worked examples. So on the same questions, a person would end up checking about 34% of OpenAI’s answers and 7% of Jev’s.
The big general models do even better here (98% for gpt-6.1-sol) because they’re right more often to begin with. They also take ten times as long per question. Claude Fable 5.1 scores as well, but it wrote the answer key for these tests, so take that with a pinch of salt.
Close calls
Where the two decision models split
The tests that actually separate the models need a real call. Does a vision model’s own write-up back up its verdict? Does the count in a report add up? On these, OpenAI’s Decision API mostly said the text doesn’t settle it.
On the report check it caught 25% of the contradictions, and sent back 50% of the good verdicts for no reason. Jev caught 75% and sent back none. Dropping the examples helped a bit (50% caught, 16% sent back), but that’s still fewer catches than any other model.
The count check is just arithmetic, and one line of code gets every one right. With the examples, OpenAI’s sent back every one of the counts that add up, and caught 13% of the ones that don’t. Jev caught 93% and sent back 18%.
We don’t think it’s the model underneath. The Decision API runs on GPT-6 Luna, and plain Luna, asked through the normal API with the same examples, caught all of the contradictions and 93% of the bad counts. Something about how the Decision API reads long questions pushes it to “not enough to tell”. It said that on 73% of its answers across these two tests.
If a tool says “not enough to tell” on clear cases, a person ends up reading them anyway.
Head to head
When they disagree, Jev is usually right
Across all 2,186 answers, the two picked different options 7% of the time. When they did, Jev was right 88% of the time, OpenAI’s 11%, and neither 1%. Most of the splits are in search, the count check and the report check.
Search on its own: asked which notes match something like “a box fell on the bottom shelf”, OpenAI’s found 85% of the notes that do. Jev found 90%, gpt-6-astra 95%.
In the app
The job Jev actually does for us
In our app, Jev has one real job on robot recordings. A video model lists every action it sees in an episode (reach, grasp, lift, place), and Jev groups those into the curator’s phases: first small box, second large box, third large box, last small box, back home. For each phase it picks the action where that phase starts.
We gave OpenAI’s Decision API the same requests Jev got for all 50 hard-set episodes. Jev answers all five phases in one go. OpenAI’s guide says questions that depend on an earlier answer should go in separate requests, so we asked it both ways: all at once, and one phase per request with the earlier answers passed along.
Jev came back with all five phases in order on 46% of the episodes. OpenAI’s managed 36% one phase at a time and 36% all at once. An answer with the phases out of order is no use to anyone, and OpenAI’s gave one on 28% and 32% of the episodes, against Jev’s 20%.
The curator timed one episode by hand, ep001. There OpenAI’s boundaries were closer: 1.0 seconds off on average, against Jev’s 1.75. One episode isn’t much to go on, though.
Both are steady. Ask the same questions twice and Jev gives the same answer 99.7% of the time, OpenAI’s 99.5%.
Images
Only OpenAI’s can look at pictures
Jev only reads text. OpenAI’s Decision API takes images too, so we also tried it on our image tests, next to the vision models and the Qualia harness. One test went fine. From the first and last frame, it caught 86% of the episodes where the robot didn’t finish, and called 7% of the finished ones unfinished. The vision models caught 90% and called 7%–10% of the finished ones unfinished, and took ten times as long.
The hardest thing we check is a fallen box. On this shelf a fallen box can look a lot like a standing one, even to a person. The hard set has 198 slots we have answers for, and 15% of them hold a fallen box. The Qualia harness caught 90% of the fallen boxes and called 0.6% of the standing ones fallen. Given exactly what the harness sees for each slot, OpenAI’s caught 69%, but it also called 36% of the standing boxes fallen, and someone would have to check each of those.
So the Qualia harness came out on top. The closest any single model got was gpt-6.1-sol, asked once per slot with the harness’s inputs: it caught 86% and called 2% of the standing boxes fallen. And those inputs do a lot of the work. They’re the curator’s examples, reference pictures of the slot and the close-up, which the harness puts together for every slot. Give Sol just the rule and it catches 38% and calls 22% of the standing boxes fallen.
Whole recordings tell the same story. Asked to keep or remove an episode from a sheet of 16 frames, OpenAI’s got an F1 of 0.63. The harness gets 0.97, and gpt-6-astra looking at the same sheet gets 0.75–0.81. On the second robot it found every item left outside the bag, but it also flagged 27% of the good recordings. The harness flagged none.
All results
The full table
Where a cell has two numbers, the first is the share it caught and the second the share it flagged by mistake, each out of the count in the row’s note. Everything is on the shelf robot’s hard set, except the location phrases, which use both shelf sets, and the bag test, which is the second robot.
| The test | OpenAI Decision API | OpenAI Decision API, rules only | Jev | GPT-6 Luna | Gpt-6-astra | Gpt-6.1-sol | Claude Fable 5.1 | Gemini 3.8 Flash | Gemini ER 2 | Qualia harness |
|---|---|---|---|---|---|---|---|---|---|---|
| Reading text | ||||||||||
| Does the report back the verdict? (contradictions caught (of 8) · good verdicts sent back (of 44)) | 25% · 50% | 50% · 16% | 75% · 0% | 100% · 14% | 88% · 18% | 88% · 11% | 100% · 5% | 100% · 36% | — | — |
| Does the count add up? (wrong counts caught (of 15) · right counts sent back (of 34)) | 13% · 100% | 47% · 53% | 93% · 18% | 93% · 18% | 100% · 0% | 100% · 0% | 100% · 0% | 100% · 0% | — | — |
| Search the notes (matching notes found (of 84) · non-matching picked (of 416)) | 85% · 1% | 89% · 5% | 90% · 2% | 95% · 4% | 95% · 0% | 95% · 0% | 95% · 2% | 90% · 0% | — | — |
| A curator’s note, into fields (share right) | 99.8% | 98% | 99.9% | 99.8% | 99.8% | 99.8% | 99.8% | 99.8% | — | — |
| A note against the curator’s own label (share right) | 95% | 95% | 95% | 95% | 95% | 95% | 95% | 95% | — | — |
| A location phrase, into slot and shelf (share right) | 99% | 97% | 98% | 99% | 99.6% | 99.6% | 99.6% | 99.6% | — | — |
| Do two reports agree? (share right) | 96% | 92% | 95% | 95% | 96% | 96% | 96% | 97% | — | — |
| A reason, into a defect type (share right) | 97% | 98% | 99% | 98% | 97% | 97% | 99.5% | 96% | — | — |
| Median seconds per question | 0.3 s | 0.3 s | 0.3 s | 3.2 s | 3.1 s | 4.1 s | 6.0 s | 6.7 s | — | — |
| Looking at images | ||||||||||
| A fallen box, harness inputs (caught (of 29) · standing called fallen (of 169)) | 69% · 36% | — | — | 28% · 6% | — | 86% · 2% | — | 48% · 17% | 21% · 6% | 90% · 0.6% |
| A fallen box, rule only (caught (of 29) · standing called fallen (of 169)) | 79% · 83% | — | — | 62% · 50% | 45% · 22% | 38% · 22% | — | 38% · 28% | 52% · 56% | — |
| Is the task finished? (first and last frame; unfinished caught (of 21) · finished called unfinished (of 29)) | 86% · 7% | — | — | 90% · 7% | 90% · 10% | 90% · 10% | — | 90% · 10% | 90% · 7% | 100% · 10% |
| Keep or remove the recording (one sheet of 16 frames; F1) | 0.63 | — | — | 0.55 | 0.85 | 0.81 | — | 0.91 | 0.82 | 0.97 |
| An item left outside the bag (second robot; caught (of 5) · good flagged (of 177)) | 100% · 27% | — | — | 100% · 53% | 100% · 8% | 100% · 15% | — | 100% · 6% | 100% · 2% | 100% · 0% |
| Median seconds per case | 0.5 s | — | — | 5.1 s | 5.0 s | 4.9 s | — | 9.5 s | 5.2 s | — |
Takeaway
So which would we use?
- A decision model, for the small text questions. There are thousands of them in a dataset. It’s ten times faster than a general model, and when it says it’s sure, you can mostly take its word for it.
- Jev, for now. It catches more of the close calls, gets the phases in order more often, and lets far more answers through without a person checking them. OpenAI’s is just as fast and just as good on the routine reading.
- Not a decision model, for pictures. OpenAI’s is the only one of the two that takes images, and it’s quick, but it raises too many false alarms to curate on its own. On whole recordings no single model matched the harness.
- Code, for counting. One line gets it right every time.
If we could ask OpenAI for one thing, it would be a proper place for worked examples. That’s how what the curator knows ends up in every question.
How we ran it. All OpenAI Decision API answers are from 7 October 2026, on the model OpenAI runs it on (GPT-6 Luna). The text tests come from our earlier decision-model study: questions written on 30 September from real curation text, an answer key written before any model ran (by Claude Fable 5.1), and nine disputed cases settled by a person. Jev was asked again the same week. The general models got the same text, options and examples in one prompt. The phase test reuses the exact requests Jev got on 30 September; the curator’s ep001 phases are one development example. The image tests are scored against the curator’s labels. Requests follow OpenAI’s Decision API guide: text as the input (images as inline base64), each question as plain-text instructions with a list of options. The guide has no field for worked examples, and for the phase test we also asked one phase per request, as it advises for dependent questions. Confidence cut-offs were set per model from the same labelled answers they are scored on, which flatters every model equally. An answer that isn’t one of the options counts as wrong, and “not enough to tell” counts as sending the case back.