From Rollout to Reset: A Graph-Based Harness for Autonomous Long-Horizon Manipulation Evaluation

Scoring a long-horizon rollout on real hardware, and resetting the scene for the next one, without a human in the loop.

Jing Jiang 1,*
KIT
Yue Yang 2,*
UNC Chapel Hill
KIT
UNC Chapel Hill
UNC Chapel Hill
KIT
1Karlsruhe Institute of Technology, 2Department of Computer Science, University of North Carolina at Chapel Hill, *Equal contribution, †Corresponding author: lioutikov@kit.edu

Abstract

Robot manipulation policies are improving quickly, and real-robot evaluation remains the standard evidence for that progress. It still relies on a human to reset the scene between rollouts, which consumes operator time and leaves the initial state distribution unspecified, so results reproduce poorly. A recent system, AutoEval, automates both reset and scoring, but only for single-step tasks, because a long-horizon rollout can terminate in combinatorially many configurations that no single learned reset policy covers. We present HALTER, a Harness for Autonomous Long-horizon Task Evaluation and Reset, which restores the scene by planning over a library of learned atomic reset skills, so demonstration cost scales with the size of that library rather than with the number of terminal states. HALTER builds a spatial scene graph online from point clouds and vision foundation models, and an LLM reasons over this graph to score the rollout, plan the reset, and verify that the reset succeeded, without collecting labeled success images for any task.

Three-minute overview — sound on

76%
reset success rate

four real tasks · 52% AutoEval · 65% MP skills

74.7%
on held-out tasks, against 1.3%

no new demonstrations collected

72%
less operator time

88 min of human effort down to 24

91%
reset verification correct

the system checks its own work

Why long-horizon reset is hard

Every real-robot evaluation alternates a rollout with a reset. The reset is still manual: a person walks over, puts the objects back, and starts the next episode. That costs operator time, and because no two hand-placed scenes are identical, it leaves the initial state distribution unspecified.

Automating it works for single-step tasks. A long-horizon rollout breaks the recipe, because it can stop anywhere.

Left: a robot rollout alternating with a human resetting the scene. Right: a four-step drawer task shown failing at four different points, each leaving a different terminal state.
A policy can fail at any step, and a task admits many valid step orderings, so the set of terminal states grows combinatorially with the length of the task. This four-step task fails at opening the drawer (1), at removing the lemon (2), in transit (3), and at closing the drawer (4). Both halves of an automated reset break: a single reset policy needs demonstrations that grow with the number of terminal states, and a single-image success checker reads spatial relations unreliably.

How HALTER works

HALTER answers three questions on every episode: score what the policy accomplished, plan a reset from wherever it stopped, and verify that the reset actually landed. All three need to read the scene state, so the representation has to support both relation tests and planning. A raw image leaves geometry implicit in its pixels; a raw point cloud carries geometry without naming objects. HALTER uses a scene graph instead.

HALTER system diagram: multi-view RGB-D keyframes become per-object point clouds and a serialized scene-graph history on the right, which a shared LLM interface reads to score, plan, and verify on the top left, while the planner composes atomic skills into a reset the robot executes on the bottom left.
Overview. During a rollout, HALTER samples multi-view RGB-D keyframes, recovers a point cloud for every object, and serializes the result into a scene-graph history (right). A shared LLM interface reads that history to score task completion, to plan the reset, and to verify it (top left), and the planner composes atomic skills from the library into the plan the robot executes (bottom left).

A graph built during the rollout

The scene is partially observable: at the end of an episode a lemon inside a closed drawer looks the same as no lemon at all. The history that resolves the ambiguity is available to any system watching the rollout, but only a system that accumulates state while the rollout runs keeps it.

HALTER samples RGB-D keyframes at 0.5 Hz, segments objects with GroundedSAM, and projects the masks into registered depth to recover per-object point clouds. It reads on from vertical proximity and support, in from volumetric containment.

A hierarchy over atomic skills

Instead of one reset policy per task, HALTER keeps a library of atomic reset skills and plans over it. Twelve skills cover both scenes, at 30 to 50 demonstrations each.

Because the planner composes skills rather than replaying one trajectory, demonstration cost scales with the size of the library, not with the number of states a rollout can stop in. The same library also covers tasks it was never assembled for.

Tasks

Four long-horizon manipulation tasks in two tabletop scenes: two drawer tasks with a lemon and a strawberry, and two stove tasks cooking a carrot in a pot and in a pan.
Four long-horizon tasks across two tabletop scenes. Drawer-Store opens the lower drawer, places a lemon inside, and closes it. Drawer-Retrieve opens the drawer, takes a strawberry out, and places it on a plate. Pot-Cook and Pan-Cook place a pot or a pan on the stove and cook a carrot in it. All experiments run on a Franka Panda with a Robotiq 2F-85 gripper and three RGB-D cameras.

Evaluating and resetting four real tasks

Twenty-five episodes per task and method, on the same policy checkpoint under evaluation.

Task (%)
Method · Metric Drawer-Store Drawer-Retrieve Pot-Cook Pan-Cook Average
AutoEval · Scoring acc.
72
64
80
88
76.0
± 4.3
AutoEval · Verif. acc.
76
80
80
76
78.0
± 4.1
AutoEval · Reset SR
52
64
16
76
52.0
± 5.0
MP skills · Reset SR
68
76
72
44
65.0
± 4.8
HALTER · Scoring acc.
100
96
80
84
90.0
± 3.0
HALTER · Plan validity
92
84
72
88
84.0
± 3.7
HALTER · Verif. acc.
96
84
88
96
91.0
± 2.9
HALTER · Reset SR
88
84
48
84
76.0
± 4.3
HALTER restores the scene in 76.0% of episodes, 24 points above AutoEval and 11 above a motion-planning reset. Pooled over the 100 episodes, the gap over AutoEval is more than three standard errors. Pot-Cook is the hardest task for every method.

The point of automating reset is to spend less human time, so we also measure what the campaign costs a person.

Method Interventions Cycle time (s/ep) Operator time (min)
Manual reset
100
53
88
AutoEval
46
104
44
MP skills
34
137
29
HALTER
25
121
24
Across 100 episodes, HALTER cuts interventions from 100 to 25 and operator time from 88 minutes to 24, and raises the mean number of episodes between interventions from 1.0 to 4.0. LLM reasoning between reset skills makes each episode slightly longer in wall-clock time, but operator availability rather than wall-clock time bounds how many rollouts a campaign can afford.

Composing skills for held-out tasks

Three tasks HALTER had never been evaluated on. They keep the same objects and change the goals: the lemon ends on the table with the drawer closed, the strawberry on the table rather than the plate, the pot on the stove with its lid on and no carrot. These goals produce terminal states that need reset sequences unlike any of the original tasks.

Held-out task (%)
Method · Metric Store Retrieve Pot Average
AutoEval · Scoring acc.
72
72
80
74.7
± 5.0
AutoEval · Verif. acc.
84
80
84
82.7
± 4.4
AutoEval · Reset SR
0
4
0
1.3
± 1.3
HALTER · Scoring acc.
100
100
84
94.7
± 2.6
HALTER · Plan validity
92
96
80
89.3
± 3.6
HALTER · Verif. acc.
96
92
88
92.0
± 3.1
HALTER · Reset SR
84
84
56
74.7
± 5.0
HALTER reaches 74.7% reset SR without a single new demonstration. AutoEval reaches 1.3%, because it fits one reset policy to the skill sequence of the task it was trained on, and that sequence no longer applies. A library of atomic skills recombines into a new order; a single policy cannot.

What the graph buys

Two panels. Left: grouped bars comparing the full model against images only and rendered point clouds across four metrics. Right: line chart of the four metrics as the graph update rate drops from 0.5 Hz to 0.066 Hz.
Ablation on Drawer-Store, 25 episodes per configuration. (a) Replacing the graph with the images themselves lowers plan validity from 92% to 72%, verification accuracy from 96% to 76%, and reset SR from 88% to 72%. Rendered point clouds lower them further, to 40%, 52%, and 40%. The gain does not come from three-dimensional appearance, but from turning geometry into explicit relations. (b) As the update rate drops, plan validity and reset SR fall while scoring accuracy stays flat: a plan needs the states the rollout passed through, whereas the completion score can often be recovered from the last few keyframes.

Limitations

Takeaway

  • Problem. Real-robot evaluation still needs a human between every rollout, and a long-horizon rollout can stop in combinatorially many terminal states, which breaks the single-reset-policy recipe.
  • Method. Build a spatial scene graph online during the rollout, and let an LLM read that graph to score the episode, plan a reset over a library of atomic skills, and verify the result.
  • Result. 76% reset success across four real tasks, 72% less operator time, and 74.7% on held-out tasks that cost no new demonstrations.

BibTeX

@article{jiang2026halter,
title={From Rollout to Reset: A Graph-Based Harness for Autonomous Long-Horizon Manipulation Evaluation},
author={Jiang, Jing and Yang, Yue and Jiang, Xinkai and Bertasius, Gedas and Szafir, Daniel J. and Lioutikov, Rudolf},
journal={arXiv preprint arXiv:2609.19413},
year={2026}
}