From Rollout to Reset: A Graph-Based Harness for Autonomous Long-Horizon Manipulation Evaluation
Scoring a long-horizon rollout on real hardware, and resetting the scene for the next one, without a human in the loop.
Abstract
Robot manipulation policies are improving quickly, and real-robot evaluation remains the standard evidence for that progress. It still relies on a human to reset the scene between rollouts, which consumes operator time and leaves the initial state distribution unspecified, so results reproduce poorly. A recent system, AutoEval, automates both reset and scoring, but only for single-step tasks, because a long-horizon rollout can terminate in combinatorially many configurations that no single learned reset policy covers. We present HALTER, a Harness for Autonomous Long-horizon Task Evaluation and Reset, which restores the scene by planning over a library of learned atomic reset skills, so demonstration cost scales with the size of that library rather than with the number of terminal states. HALTER builds a spatial scene graph online from point clouds and vision foundation models, and an LLM reasons over this graph to score the rollout, plan the reset, and verify that the reset succeeded, without collecting labeled success images for any task.
Three-minute overview — sound on
- 76%
- reset success rate
- 74.7%
- on held-out tasks, against 1.3%
- 72%
- less operator time
- 91%
- reset verification correct
four real tasks · 52% AutoEval · 65% MP skills
no new demonstrations collected
88 min of human effort down to 24
the system checks its own work
Why long-horizon reset is hard
Every real-robot evaluation alternates a rollout with a reset. The reset is still manual: a person walks over, puts the objects back, and starts the next episode. That costs operator time, and because no two hand-placed scenes are identical, it leaves the initial state distribution unspecified.
Automating it works for single-step tasks. A long-horizon rollout breaks the recipe, because it can stop anywhere.
How HALTER works
HALTER answers three questions on every episode: score what the policy accomplished, plan a reset from wherever it stopped, and verify that the reset actually landed. All three need to read the scene state, so the representation has to support both relation tests and planning. A raw image leaves geometry implicit in its pixels; a raw point cloud carries geometry without naming objects. HALTER uses a scene graph instead.
A graph built during the rollout
The scene is partially observable: at the end of an episode a lemon inside a closed drawer looks the same as no lemon at all. The history that resolves the ambiguity is available to any system watching the rollout, but only a system that accumulates state while the rollout runs keeps it.
HALTER samples RGB-D keyframes at 0.5 Hz, segments objects with GroundedSAM, and projects the masks into registered depth to recover per-object point clouds. It reads on from vertical proximity and support, in from volumetric containment.
A hierarchy over atomic skills
Instead of one reset policy per task, HALTER keeps a library of atomic reset skills and plans over it. Twelve skills cover both scenes, at 30 to 50 demonstrations each.
Because the planner composes skills rather than replaying one trajectory, demonstration cost scales with the size of the library, not with the number of states a rollout can stop in. The same library also covers tasks it was never assembled for.
Tasks
Evaluating and resetting four real tasks
Twenty-five episodes per task and method, on the same policy checkpoint under evaluation.
| Task (%) | |||||
|---|---|---|---|---|---|
| Method · Metric | Drawer-Store | Drawer-Retrieve | Pot-Cook | Pan-Cook | Average |
| AutoEval · Scoring acc. | 72 | 64 | 80 | 88 | 76.0 ± 4.3 |
| AutoEval · Verif. acc. | 76 | 80 | 80 | 76 | 78.0 ± 4.1 |
| AutoEval · Reset SR | 52 | 64 | 16 | 76 | 52.0 ± 5.0 |
| MP skills · Reset SR | 68 | 76 | 72 | 44 | 65.0 ± 4.8 |
| HALTER · Scoring acc. | 100 | 96 | 80 | 84 | 90.0 ± 3.0 |
| HALTER · Plan validity | 92 | 84 | 72 | 88 | 84.0 ± 3.7 |
| HALTER · Verif. acc. | 96 | 84 | 88 | 96 | 91.0 ± 2.9 |
| HALTER · Reset SR | 88 | 84 | 48 | 84 | 76.0 ± 4.3 |
The point of automating reset is to spend less human time, so we also measure what the campaign costs a person.
| Method | Interventions | Cycle time (s/ep) | Operator time (min) |
|---|---|---|---|
| Manual reset | 100 | 53 | 88 |
| AutoEval | 46 | 104 | 44 |
| MP skills | 34 | 137 | 29 |
| HALTER | 25 | 121 | 24 |
Composing skills for held-out tasks
Three tasks HALTER had never been evaluated on. They keep the same objects and change the goals: the lemon ends on the table with the drawer closed, the strawberry on the table rather than the plate, the pot on the stove with its lid on and no carrot. These goals produce terminal states that need reset sequences unlike any of the original tasks.
| Held-out task (%) | ||||
|---|---|---|---|---|
| Method · Metric | Store | Retrieve | Pot | Average |
| AutoEval · Scoring acc. | 72 | 72 | 80 | 74.7 ± 5.0 |
| AutoEval · Verif. acc. | 84 | 80 | 84 | 82.7 ± 4.4 |
| AutoEval · Reset SR | 0 | 4 | 0 | 1.3 ± 1.3 |
| HALTER · Scoring acc. | 100 | 100 | 84 | 94.7 ± 2.6 |
| HALTER · Plan validity | 92 | 96 | 80 | 89.3 ± 3.6 |
| HALTER · Verif. acc. | 96 | 92 | 88 | 92.0 ± 3.1 |
| HALTER · Reset SR | 84 | 84 | 56 | 74.7 ± 5.0 |
What the graph buys
Limitations
- The skill library has to cover the reset. A reset may need a skill the rollout never performs, which more demonstrations would supply. Or the rollout may leave the scene physically unrecoverable, as when an object breaks or a liquid spills, which no planner over manipulation skills can undo. In that second case HALTER detects rather than recovers: the post-reset check fails and calls an operator, so an unrecoverable state costs human time instead of silently starting the next rollout from the wrong initial condition.
- A handful of reference cases per task. The reasoner still needs in-context examples written for each task, although it trains no classifier and collects no labeled images.
- Executing a reset is itself a long-horizon problem. The difficulties known from that setting carry over to the skills HALTER composes. Our focus here is on bringing the long-horizon setting into the reset problem itself.
Takeaway
- Problem. Real-robot evaluation still needs a human between every rollout, and a long-horizon rollout can stop in combinatorially many terminal states, which breaks the single-reset-policy recipe.
- Method. Build a spatial scene graph online during the rollout, and let an LLM read that graph to score the episode, plan a reset over a library of atomic skills, and verify the result.
- Result. 76% reset success across four real tasks, 72% less operator time, and 74.7% on held-out tasks that cost no new demonstrations.
BibTeX
@article{jiang2026halter, title={From Rollout to Reset: A Graph-Based Harness for Autonomous Long-Horizon Manipulation Evaluation}, author={Jiang, Jing and Yang, Yue and Jiang, Xinkai and Bertasius, Gedas and Szafir, Daniel J. and Lioutikov, Rudolf}, journal={arXiv preprint arXiv:2609.19413}, year={2026}}