What does a vision model miss when you just hand it the recording?
The quick way to curate robot data is one API call per recording. Send the video to a vision model and ask whether the episode is good. We built that, gave it every advantage we could, and measured it against the Qualia curation harness on the same 106 recordings, three runs each. We call the one-call approach the straight feed.
Setup · the task
Four boxes, one shelf, a minute of video
The task is simple. A two-arm robot moves four boxes from a container into the empty slots of a two-level shelf. At the end, the shelf holds twelve boxes and the container is empty.
One attempt is called an episode. An episode has two parts: about a minute of video from a fixed camera, and the joint log, where the robot’s motors report their positions thirty times a second.
Hundreds of these episodes, driven by people, teach the robot to do the task on its own. A bad episode teaches it the wrong thing. Curation is the job of finding the bad ones.
Four things go wrong in this data:
- A placed box falls over.
- The task is not finished.
- A joint freezes. The motor log is wrong, but the video looks fine.
- The camera drifts out of step with the joint log.
The task is harder to curate than it looks, and the hardest part is the fallen box.
The camera looks down on the shelf from above. A box standing on the bottom shelf shows only its top, so it looks as if it is lying flat. A box that is twisted or leaning but still on its base counts as standing. A box on its side can look almost the same. Even a person has to look twice.
That is why we built this post around it. The question is hard, but the answer can be checked: after the boxes are placed, has one fallen over, and if so, in which slot?
A curator went through all 106 episodes and marked 58 for removal. The episodes come from two sets. The ordinary set is 56 recordings from a normal collection. The hard set is 50 recordings made on purpose so that most attempts go wrong. Both sets appear in every result below.
The start of an episode from the fixed camera: the shelf at the top with eight boxes, the container with four in front of the robot.
The same kind of episode as sixteen frames, four seconds apart, each stamped with its time. This one image is the second thing we hand the models.
Setup · the two approaches
One API call, or the Qualia curation harness
The straight feed is one call per episode. The model gets the prompt below, plus either the whole video or the sixteen-frame image shown above. Nothing else.
The task enters through the first line only. We ran it two ways. In one, the first line is a placeholder that says nothing about the task. In the other, it is the one sentence a user would type to describe the task. We call that sentence the context.
Gemini Robotics ER 2 was given the video. gpt-6-astra was given the sixteen-frame image.
- Context: A robot performs a task. or Restock the shelf: place the 4 boxes from the container onto the empty slots of the two-level shelf in front of the robot.
- Input: The attached video is the whole episode. or The attached image is a contact sheet of 16 frames sampled evenly over the whole episode, in reading order, each stamped with its time in seconds.
- Ask: You are curating robot teleoperation recordings for training data. Decide whether this episode should be kept. Answer with JSON only.
- Answer:
{"verdict": keep | remove, "defect_type": none | task_not_completed | object_fallen_or_misplaced | robot_or_sensor_fault | other, "where": "…", "when_seconds": …, "confidence": 0 to 1, "why": "…"}
The Qualia curation harness gets the whole episode, video and joint log, plus a brief from the curator. It returns a verdict, and with it the three things a curator needs to act on it:
- the reason
- the place in the scene
- the second in the video
The rest of this post measures how often each approach is right, on the same episodes.
Left, the straight feed: a recording and the curator’s brief in, one model, a one-sentence verdict out. Right, the Qualia curation harness: the whole episode and a brief from the curator in, a verdict with its reason, a place and a second out.
Setup · the results at a glance
What each approach can tell you about an episode
The table is the whole comparison in one place. Every number was measured on the same 106 episodes, against the same curator, three runs per setup. The feeds are shown at their best.
The sections after it explain what is behind each row.
| what you need to know | whole video, straight to a model | sixteen frames, straight to a model | the Qualia curation harness |
|---|---|---|---|
| F1 0 to 1, and 1 is perfect; ordinary set · hard set | .09 · .71 | 0 · .77 | .94 · .97 |
| Task not finished 21 episodes | 92% | 94% | 100% |
| Fallen box 26 episodes; ordinary set · hard set | 0% · 13% | 0% · 24% | 100% · 87% |
| Names the slot the box fell in of the fallen boxes it caught; hard set | 0% | 45% | 100% |
| Times every placement against a person's stopwatch | one time per verdict | one time per verdict | 24 of 24 within 2 s |
| Mistakes per run hard set, 37 bad and 13 good episodes | 2 good removed 15.7 bad missed |
0 good removed 13.7 bad missed |
0 good removed 2 bad missed |
| Time for 5,000 episodes | 28.1 h one at a time 3.5 h eight at a time |
16.3 h one at a time | 5 min, every episode at once |
| What it takes to set up | a prompt per task | a prompt per task | one sentence, one good episode, six clicks |
The straight feed asks one question about a whole recording. Catching a bad episode takes several: what happened, where, when and how many.
Finding 1 .09 → .94
The feed keeps most bad episodes. The Qualia curation harness removes them.
The first question is the simplest. Does each approach remove the episodes the curator removed?
We score this with F1, a number from 0 to 1. It combines two things: how many of the bad episodes were removed, and how many of the removals were right. A 1 means every bad episode was removed and no good one was.
On the ordinary set, the straight feed removed almost nothing. F1 was .09 on video and 0 on frames. On the hard set, where most bad episodes are unfinished tasks, the feed did better: .71 on video and .77 on frames, with context.
The Qualia curation harness scored .94 and .97. It removed every bad episode on the ordinary set, and 35 of the 37 on the hard set.
Same models, same episodes, same three runs. Only the approach differs.
Finding 2 24 of 24
The joint log knows when
The joint log is the robot’s own record of what it did. It holds the moment each gripper closes on a box and the moment it lets go.
That makes it a clock for the whole episode. A person timed 24 placements with a stopwatch. Read from the joint log, all 24 are within two seconds of the person’s times, and 23 within one.
To put those moments on the video, the Qualia curation harness measures the gap between the camera’s clock and the motors’. It does that by matching the motion in the picture to the motion in the log. The same measurement shows when a camera has drifted out of step, one of the four faults in this data.
The joint log says when. The vision models say what happened and where.
Finding 3 21 of 21
Unfinished tasks: the feed guesses, the Qualia curation harness measures
An unfinished task looks finished. The boxes are on the shelf and the arm is at rest. To catch it from a picture, the straight feed has to guess whether all four boxes went in.
Telling it that there were four helps a lot. With context, the feeds caught 92% and 94% of the hard set’s 21 unfinished tasks. Without it, 71% and 43%.
Guessing cuts both ways. The video model also removed the same two good episodes in every run. In one of them, a person’s hand reaches into the frame to straighten a box.
The Qualia curation harness does not guess. It measures whether the task was finished, and it tells three failures apart: the task was not finished, a box ended up where it should not be, or too few boxes were placed. It also says at what second.
On this set it caught all 21 unfinished tasks in every run and removed no good episode. The frames feed removed no good episode either, but only because it removes very little. It missed more than a third of the bad episodes even with context.
With a sentence of context, a model can guess. The Qualia curation harness measures.
Finding 4 33 of 33
Fallen boxes: the right slot, every time
This is the hard question from the start of the post. On the bottom shelf a standing box looks fallen. On the top shelf a fallen box is a few dozen pixels in a frame of a million.
The ordinary set holds 11 fallen boxes. Over three runs, that gave the feeds 33 chances to catch one. They caught none. The Qualia curation harness caught all of them.
On the hard set the feeds did catch some. So we asked a second question: when a feed catches a fallen box, does it say which slot? The video feed named the right slot in 0% of its catches, the frames feed in 45%. The Qualia curation harness named the curator’s slot on every catch, on both sets.
The Qualia curation harness is not perfect here either, and its mistakes are where you would expect. All of them are on the bottom shelf: a standing box called fallen in every run on the ordinary set, and two fallen boxes on the hard set it never caught.
A useful verdict says which slot. Without it, someone has to watch the episode again.
Phase one
Before the higher-level checks: the primal checks
A regular API call answers only the question you ask it. Besides higher-level checks like the fallen boxes and the placement times, the Qualia curation harness runs the primal checks every robotics dataset needs:
- frozen joints
- a camera out of sync with the joint log, found by correlating the motion in the picture with the motion of the joints
- torn or corrupted frames
- recordings that run into the next episode
- episodes far longer or shorter than the rest
- dataset files that disagree with their own index
The primal checks are the first phase of curation. They make a dataset ready for the higher-level checks. Run over this post’s own recordings, they found six torn frames that nobody, us included, had noticed.
One of the six: the frame before, the torn frame, the frame after.
Finding 5 5 min
Five thousand episodes in five minutes
The last question is time. Call a model once per episode and wait for each answer, and the hard set’s 50 episodes take 17 min on video. A thousand episodes take 5.6 h. Five thousand take 28.1 h.
The Qualia curation harness runs every episode at the same time. So the time it needs is the time of the longest single episode, not the sum of all of them: about five minutes, whether the dataset has fifty episodes or five thousand. Even our measured run, kept to a modest number of parallel episodes on purpose, took 12 min for the hard set, while doing far more work per episode than the loop.
That difference is infrastructure, and it is where the work goes as a team’s data grows.
A curation pass is no longer one model call per episode. It is many questions per episode, put to different kinds of models: vision-language models, but also segmentation, depth and small structured-decision models. Some are proprietary. Each has its own limits and its own ways of failing.
Running thousands of episodes at once across all of them, retrying what fails, keeping every answer, and finishing in minutes rather than days is a problem of its own. Qualia builds for exactly that. We work directly with the foundation-model labs, including on our robotics program with Google DeepMind, so that these models can be used at scale, reliably and fast.
You can parallelise your own loop. Then you are building the rest yourself.
Five thousand episodes take five minutes. The model is not the hard part. Running every model on every episode at once is.
A second robot 33% → 95%
On a second robot, the API call missed two thirds of the bad recordings
We ran both approaches on a second robot: two arms packing clothes and shoes into a bag. A curator checked all 202 of its recordings. They removed 7% and kept 88%. The other 5% were a different job: the bag was already packed, and the robot only pulled it to the middle of the workspace. The curator tagged those as a separate task.
The straight feed got the same prompt as on the shelf robot, with and without the task sentence, three runs each. At its best it removed 33% of the bad recordings, and 5% of the good ones with them. Over three runs, the Qualia curation harness removed 95% of the bad recordings on average (all of them in one of three runs) and none of the good ones.
| second robot | whole video, straight to a model | sixteen frames, straight to a model | the Qualia curation harness |
|---|---|---|---|
| Bad recordings removed 7% of the recordings | 31% | 33% | 95% |
| Recordings that run into the next one 57% of the bad ones | 4%, never for that reason | 0% | 100% |
| Nothing packed | 100% | 83% | 100% |
| Something left outside the bag | 0% | 67% | 67%, the rest sent for a look |
| Good recordings removed by mistake | 5% | 4% | 0% |
| Bag moves thrown out as failed packing a different job, tagged by the curator | 0% without the task sentence 73% with it |
0% without the task sentence 67% with it |
0%, every one tagged |
The feed is not blind. When nothing was packed, it usually saw it. It failed in three places, and each one says something about curation.
One recording cannot show where it ends
More than half of the bad recordings run into the next one: their last seconds show the next attempt’s starting table. The feed never noticed. Across all its runs it removed only two of them, and each time it read the next attempt’s items, laid out on the table, as an unfinished task. A single call sees one recording, so it cannot know that the ending belongs to another. The Qualia curation harness checks every recording against the one after it. That is one of the primal checks.
A sentence of context throws out the wrong recordings
Told the task, the feed removed 67% to 73% of the bag moves as failed packing. Without the sentence it kept them all, but then it could not tell them apart from packing either. Context taught it what success looks like, and it failed everything that looked different, including a different job done right. The curator’s brief names the bag move once. The Qualia curation harness tagged every one and removed none.
The hardest misses need a second opinion
The recordings the feed missed most were the subtle ones. In one, a shirt hangs over the bag’s rim. The feed kept it in 10 of its 12 runs. The Qualia curation harness never passed it as clean: it removed it in one of three runs, and in the other two it was unsure and sent it to a person. The next section shows what it takes to see that shirt at all.
The result is that nobody has to watch every recording. The Qualia curation harness passed 84% of them as clean, and none of the bad recordings was ever among them. The rest, about one in six, are the ones to look at: the ones it removed, the bag moves, and a few it flagged for a second look.
The straight feed looks at one recording at a time. Curation needs the dataset, the brief and a second opinion.
Inside the Qualia curation harness
Models catch more when they check each other
Inside the Qualia curation harness, no model works alone. That is how it gets to the hardest recordings on the second robot. Models look at the same recording independently, check each other, and take pointers from each other before an answer stands.
Take the end of each recording. In five, something was still outside the bag: three where nothing had been packed, and two where something was left behind. In one of those two, a shirt hung over the bag’s rim.
First we asked single models, each on its own, whether every item ended up in the bag. Five strong models, each on its own, caught between 40% and 80% of them. gpt-6-astra caught 40%. Not one of them saw the hanging shirt.
Told exactly what counts as outside, astra caught 80%, the hanging shirt among them. But it also flagged two good recordings, once mistaking the robot’s own black gripper for a shoe.
Different models see different things. Given the same instruction, one saw the hanging shirt and missed a shirt left beside the bag. Another did the opposite. Each had false alarms of its own.
In the Qualia curation harness they look independently. Where they agree, the answer stands. Where they disagree, a third model looks closer at exactly the places they name. On the last frame of each recording, together they caught all of them, with no false alarm on the good recordings. Then we ran everything a second time, every call made again, and got the same result. In the full runs on the whole recordings, the hanging shirt came back unsure in two of the three runs, and each time it went to a person.
Fallen boxes on the shelf robot show the same thing. A narrower model’s second opinion, shown with the labelled pictures behind it, took astra from five or six false alarms per run to none, while it caught just as many fallen boxes. Together they did better than either one alone, and better than a fixed rule for combining the two.
Help is not automatic. Told up front that another model had found no leftover box, astra searched harder and reported boxes that were not there. Its false alarms rose from six to nine. Told the same thing after it had answered, it withdrew false alarms instead, down to three.
So the gain is not in having more models. It is in knowing which model to consult on which question, what to show the strong model, and when. That is what we measure, and it is what we build into the Qualia curation harness. One API call, to however strong a model, has no one to consult.
One model alone can only answer from what it sees. In the Qualia curation harness, models check each other and look at what the others found.
Decision models
A decision model checks the words behind every verdict
The vision models write down what they saw before they give a verdict. The curator writes a note with every label. A straight API call never reads any of it again. The Qualia curation harness does, with a decision model: a small model built to answer structured questions about text. Ours is Jev, from TypeSafe.
A vision model can describe one thing and conclude another. In an earlier run on the bag robot, a video model wrote that no item was placed in six recordings, and the next step still marked the packing task complete. The verdict looked fine on its own. The contradiction was in the sentence next to it.
Given 93 real verdicts with their written evidence, the decision model found 94% of the contradictions. A rule that trusts every verdict finds none. We are adding it as a check that sends such a verdict back for another look.
It reads the curator’s notes the same way. Checked against their labels, it found nine episodes where note and label disagree, such as a keep whose note describes a fallen box. Each went back to the curator as a proposal.
Numbers stay in code. On placement counts, a one-line rule was right every time, and the decision model 92% of the time.
Next we are adding a modified decision model, and we are waiting on OpenAI’s new decision model to see how it does on our broader evaluations.
A single API call gives an answer and stops. The Qualia curation harness reads every answer back.
Takeaway
What you would be building
The straight feed is one prompt and a loop. Everything else in this post is what it takes to curate robot data well:
- A clock (24/24). Every grasp and release, timed from the joint log, with the camera’s clock measured against the motors’.
- Measuring, not guessing (21/21). Whether the task was finished, counted rather than guessed from a picture. All 21 unfinished tasks removed in every run, and no good episode with them.
- Where it went wrong (33/33). The fallen box and its slot, on a shelf where standing boxes look fallen. Every fallen box on the ordinary set, with the curator’s slot on every catch, on both sets.
- The primal checks (6). Frozen joints, torn frames, recordings that run into the next one. Six torn frames in our own data that nobody had seen.
- Models that check each other (95%). 95% of the bad recordings on a second robot, where the straight feed found 33%, and one recording in six left to watch.
- Every episode at once (5 min). About five minutes for fifty episodes or five thousand, across every kind of model, with every answer kept.
Each of these is its own piece of work. Each has to be measured against a curator before you can trust it, and measured again when a model changes. That is the work we do, so your team does not have to.
What it asks of you is one sentence describing the task, one good episode, and about six clicks. The Qualia curation harness works out the rest. You confirm it.
Next
Try it on your data
This post is one study on one dataset, and a second robot. We have run the same comparison on other tasks, other robots and other kinds of recordings, with similar results.
Finding what is wrong is only the first step. The work after that is fixing the data: trimming episodes by subtask, adding a subtask at the right moment, creating the right recipes, and doing all of it in one place with your team, in a shared workspace that all of your robotics data streams into. That is what Qualia is built for.
The better test is your own data. Try it. You will probably find something in it you did not know was there. We did, with a lot of ours.
Join the waitlist, or write to us at hello@qualiastudios.dev.
Every run in this post. Three per setup, scored against the same curator’s labels. The second robot: three runs of the Qualia curation harness on each of its four recording sessions, three runs of each straight-feed setup with the same prompts, and two runs of the end-state comparison, all against its curator’s labels.
| Set | Setup | Context | F1, three runs | Precision | Recall | Verdicts flipped |
|---|---|---|---|---|---|---|
| ordinary | whole video | none | .09 ± .00 | .83 | .05 | 2% |
| ordinary | whole video | one sentence | .06 ± .05 | .67 | .03 | 2% |
| ordinary | sixteen frames | none | 0 ± 0 | 0 | 0 | 0% |
| ordinary | sixteen frames | one sentence | 0 ± 0 | 0 | 0 | 0% |
| ordinary | the Qualia curation harness | the brief | .94 ± .03 | .89 | 1.0 | 5% |
| hard | whole video | none | .64 ± 0 | .86 | .51 | 0% |
| hard | whole video | one sentence | .71 ± .01 | .91 | .58 | 2% |
| hard | sixteen frames | none | .47 ± .02 | 1.0 | .31 | 8% |
| hard | sixteen frames | one sentence | .77 ± .03 | 1.0 | .63 | 6% |
| hard | the Qualia curation harness | the brief | .97 ± 0 | 1.0 | .95 | 0% |