Diffusion Policy
4 tasks · 3 demonstration budgets
Proactive Human-in-the-Loop
Preview a robot’s next move. Correct it before an error. Learn from the difference.
From anticipating errors in XR to learning from corrective actions. The video contains on-screen explanations and no audio.
The video contrasts reactive correction after an error with proactive correction using an XR preview. A human operator takes control of a Franka robot while predicted trajectories continue to update. Rejected predictions and executed corrective action chunks form preference pairs. Side-by-side selected rollouts then show Towel Folding, Cook the Carrot, and Lemon in the Drawer for HG-DAgger with reactive or proactive data and PHIL with proactive data. Playback speed labels are retained from the original video. These selected examples illustrate behavior; the results below summarize the quantitative evaluation.
Five correction episodes
Mean task success after post-training.
4 tasks · 3 demonstration budgets
3 tasks · 3 demonstration budgets
Unweighted means over each policy’s evaluated tasks and demonstration budgets. HG-DAgger uses reactive corrections; PHIL uses proactive corrections. Each condition uses five correction rollout episodes, with multiple interventions allowed.
The idea
Proactive Human-in-the-Loop (PHIL) is a framework for proactive policy correction and preference-based post-training. It lets an operator preview policy-predicted future motion in extended reality (XR) and intervene before an anticipated undesirable action.
During takeover, the policy continues predicting trajectories from the robot’s current state. PHIL pairs rejected predictions with the corrective action chunks subsequently executed from the same physical state over the same horizon. Preference-based post-training encourages the policy to favor these corrections.
Evaluated on real-world manipulation with a diffusion policy and a pretrained vision-language-action policy, PHIL improves mean success rates over HG-DAgger using five correction episodes per condition. A separate analysis finds fewer unnecessary recovery behaviors in successful episodes for policies trained with proactive corrections.
How PHIL works
Turn human intervention into a comparison between what the policy proposed and what the robot should do.
The PHIL framework. Base demonstrations initialize the policy. Proactive correction provides context-aligned preference pairs for post-training.
View full size (new tab)Render predicted end-effector trajectories in the physical workspace, so the operator can inspect upcoming motion.
Take control before an anticipated error. Policy queries and XR previews continue on the same fixed schedule.
Match a rejected prediction with the next executed action chunk, starting at the same query state and spanning the same horizon.
Combine imitation on the original demonstrations with a preference objective that favors the corrective continuation.
The positive chunk contains the next H executed actions and may extend beyond the end of human takeover. The prediction accepted at release is excluded from negative samples. Only complete pairs are used for training.
In the experiments, actions run at 20 Hz with a 30-step prediction horizon and a query every 10 control steps. Both action chunks share the same context; the preference loss encourages a lower native imitation loss for the corrective chunk.
In the workspace
A Meta Quest 3 interface aligns the predicted trajectories with the real robot. Holding the controller trigger enables correction; releasing it returns control to the policy.
Experiments & results
From single-stage manipulation to sequences that require several actions to succeed.
Pick up the banana.
Single stageFold the towel.
Single stagePlace the pot, add the carrot, and close the lid.
3 ordered stagesOpen the drawer, place the lemon, and close it.
3 ordered stagesPHIL matches or exceeds the compared adaptation baselines across the evaluated tasks, budgets, and metrics. PHIL uses proactive corrections; HG-DAgger and Residual Policy use reactive corrections.
Diffusion Policy results. Left to right: success rate for Banana, Towel, Carrot, and Lemon; then task completion score for Carrot and Lemon. The zero-demonstration origins are visual guides, not measured conditions.
View full size (new tab)Cook the Carrot · 10 demonstrations: PHIL reaches 52% success, compared with 24% for HG-DAgger and 4% for the base policy.
PHIL matches or exceeds HG-DAgger across the evaluated X-VLA tasks, budgets, and metrics, reaching 100% full-task success on Lemon in the Drawer with 30 base demonstrations.
| Method | 10 demos | 20 demos | 30 demos | |||
|---|---|---|---|---|---|---|
| SR | TCS | SR | TCS | SR | TCS | |
| Base Policy | 32 | — | 36 | — | 60 | — |
| HG-DAgger | 32 | — | 40 | — | 64 | — |
| PHIL | 64 | — | 72 | — | 72 | — |
| Method | 10 demos | 20 demos | 30 demos | |||
|---|---|---|---|---|---|---|
| SR | TCS | SR | TCS | SR | TCS | |
| Base Policy | 40 | 73.3 | 56 | 77.3 | 56 | 85.3 |
| HG-DAgger | 56 | 78.7 | 48 | 76.0 | 76 | 90.7 |
| PHIL | 56 | 85.3 | 68 | 89.3 | 76 | 92.0 |
| Method | 10 demos | 20 demos | 30 demos | |||
|---|---|---|---|---|---|---|
| SR | TCS | SR | TCS | SR | TCS | |
| Base Policy | 0 | 37.3 | 4 | 46.7 | 0 | 53.3 |
| HG-DAgger | 68 | 72.0 | 56 | 78.7 | 80 | 88.0 |
| PHIL | 92 | 96.0 | 92 | 97.3 | 100 | 100.0 |
All values are percentages over 25 evaluation rollouts per condition. SR: full-task success rate. TCS: normalized task completion score across three ordered stages. TCS is not reported for Towel Folding. Blue shading identifies PHIL.
Proactive collection alone does not consistently improve success under behavior cloning. On the same proactive data, preference post-training matches or improves on HG-DAgger across all four tasks in the lowest-data regime.
Separating the two contributions. Three base demonstrations for Banana and ten for each other task; five correction episodes per condition. Hybrid combines two proactive and three reactive episodes. Mixing improves HG-DAgger here but reduces PHIL performance, showing that the benefit depends on the objective.
View full size (new tab)For both HG-DAgger and PHIL, proactive-trained Diffusion Policies show fewer unnecessary recovery behaviors in every evaluated task–demonstration setting.
A separate evaluation of successful behavior. Each point averages counts over ten successful episodes per condition. Necessary recovery after actual errors is excluded. The horizontal axis is the share of correction frames in the combined training data and decreases to the right.
View full size (new tab)Takeaways
XR previews let people act on a policy’s proposed motion before it is executed.
Rejected predictions and executed corrections provide a context-aligned preference signal.
Five correction episodes improve mean success over HG-DAgger in the evaluated Diffusion Policy and X-VLA settings.