Proactive Human-in-the-Loop

Correct Robots Before They Make MistakesProactive Human-in-the-Loop Intervention via Preference Learning

Karlsruhe Institute of Technology (KIT)

* Corresponding author

Preview a robot’s next move. Correct it before an error. Learn from the difference.

XR trajectory previewHuman correctionPreference learning

PHIL in action

Real-world robot learning · 1 min 50 sec

From anticipating errors in XR to learning from corrective actions. The video contains on-screen explanations and no audio.

Read the video overview

The video contrasts reactive correction after an error with proactive correction using an XR preview. A human operator takes control of a Franka robot while predicted trajectories continue to update. Rejected predictions and executed corrective action chunks form preference pairs. Side-by-side selected rollouts then show Towel Folding, Cook the Carrot, and Lemon in the Drawer for HG-DAgger with reactive or proactive data and PHIL with proactive data. Playback speed labels are retained from the original video. These selected examples illustrate behavior; the results below summarize the quantitative evaluation.

Five correction episodes

Small correction budget.
Stronger policies.

Mean task success after post-training.

Diffusion Policy

44.0%70.3%
HG-DAggerPHIL

4 tasks · 3 demonstration budgets

X-VLA

57.8%76.9%
HG-DAggerPHIL

3 tasks · 3 demonstration budgets

Unweighted means over each policy’s evaluated tasks and demonstration budgets. HG-DAgger uses reactive corrections; PHIL uses proactive corrections. Each condition uses five correction rollout episodes, with multiple interventions allowed.

The idea

A preview changes
when we intervene.

Proactive Human-in-the-Loop (PHIL) is a framework for proactive policy correction and preference-based post-training. It lets an operator preview policy-predicted future motion in extended reality (XR) and intervene before an anticipated undesirable action.

During takeover, the policy continues predicting trajectories from the robot’s current state. PHIL pairs rejected predictions with the corrective action chunks subsequently executed from the same physical state over the same horizon. Preference-based post-training encourages the policy to favor these corrections.

Evaluated on real-world manipulation with a diffusion policy and a pretrained vision-language-action policy, PHIL improves mean success rates over HG-DAgger using five correction episodes per condition. A separate analysis finds fewer unnecessary recovery behaviors in successful episodes for policies trained with proactive corrections.

How PHIL works

See ahead. Step in. Learn a preference.

Turn human intervention into a comparison between what the policy proposed and what the robot should do.

PHIL pipeline: initialize a base policy from demonstrations, preview future actions in XR and collect rejected–corrected action pairs, then post-train with imitation and preference objectives.

The PHIL framework. Base demonstrations initialize the policy. Proactive correction provides context-aligned preference pairs for post-training.

View full size (new tab)
  1. 01

    Preview

    Render predicted end-effector trajectories in the physical workspace, so the operator can inspect upcoming motion.

  2. 02

    Correct

    Take control before an anticipated error. Policy queries and XR previews continue on the same fixed schedule.

  3. 03

    Pair

    Match a rejected prediction with the next executed action chunk, starting at the same query state and spanning the same horizon.

  4. 04

    Post-train

    Combine imitation on the original demonstrations with a preference objective that favors the corrective continuation.

A closer look at preference pairs

The positive chunk contains the next H executed actions and may extend beyond the end of human takeover. The prediction accepted at release is excluded from negative samples. Only complete pairs are used for training.

In the experiments, actions run at 20 Hz with a 30-step prediction horizon and a query every 10 control steps. Both action chunks share the same context; the preference loss encourages a lower native imitation loss for the corrective chunk.

In the workspace

Human guidance,
in context.

A Meta Quest 3 interface aligns the predicted trajectories with the real robot. Holding the controller trigger enables correction; releasing it returns control to the policy.

Sequence of XR views showing trajectory inspection, human takeover, continued policy previews during correction, and return to policy control.
A proactive correction sequence with synchronized robot and trajectory visualization.View full size (new tab)

Experiments & results

Four tasks. Two policy families.

From single-stage manipulation to sequences that require several actions to succeed.

TASK 01

Pick Up Banana

Pick up the banana.

Single stage
TASK 02

Towel Folding

Fold the towel.

Single stage
TASK 03

Cook the Carrot

Place the pot, add the carrot, and close the lid.

3 ordered stages
TASK 04

Lemon in the Drawer

Open the drawer, place the lemon, and close it.

3 ordered stages
A shared evaluation protocol. Each condition uses 25 evaluation rollouts. Diffusion Policy is evaluated on all four tasks; X-VLA on Towel Folding, Cook the Carrot, and Lemon in the Drawer. Base-demonstration budgets are 3/6/9 for Diffusion Policy on Pick Up Banana, and 10/20/30 elsewhere.
01 / DIFFUSION POLICY

Higher success across demonstration budgets.

PHIL matches or exceeds the compared adaptation baselines across the evaluated tasks, budgets, and metrics. PHIL uses proactive corrections; HG-DAgger and Residual Policy use reactive corrections.

Six-panel line plot comparing Base Policy, HG-DAgger, Residual Policy, and PHIL across demonstration budgets. PHIL matches or exceeds adaptation baselines for task success and long-horizon task completion.

Diffusion Policy results. Left to right: success rate for Banana, Towel, Carrot, and Lemon; then task completion score for Carrot and Lemon. The zero-demonstration origins are visual guides, not measured conditions.

View full size (new tab)

Cook the Carrot · 10 demonstrations: PHIL reaches 52% success, compared with 24% for HG-DAgger and 4% for the base policy.

02 / PRETRAINED VLA

The same approach extends to X-VLA.

PHIL matches or exceeds HG-DAgger across the evaluated X-VLA tasks, budgets, and metrics, reaching 100% full-task success on Lemon in the Drawer with 30 base demonstrations.

Towel Folding

X-VLA Towel Folding: full-task success rate and normalized task completion score (%) at three base-demonstration budgets.
Method10 demos20 demos30 demos
SRTCSSRTCSSRTCS
Base Policy32—36—60—
HG-DAgger32—40—64—
PHIL64—72—72—

Cook the Carrot

X-VLA Cook the Carrot: full-task success rate and normalized task completion score (%) at three base-demonstration budgets.
Method10 demos20 demos30 demos
SRTCSSRTCSSRTCS
Base Policy4073.35677.35685.3
HG-DAgger5678.74876.07690.7
PHIL5685.36889.37692.0

Lemon in the Drawer

X-VLA Lemon in the Drawer: full-task success rate and normalized task completion score (%) at three base-demonstration budgets.
Method10 demos20 demos30 demos
SRTCSSRTCSSRTCS
Base Policy037.3446.7053.3
HG-DAgger6872.05678.78088.0
PHIL9296.09297.3100100.0

All values are percentages over 25 evaluation rollouts per condition. SR: full-task success rate. TCS: normalized task completion score across three ordered stages. TCS is not reported for Towel Folding. Blue shading identifies PHIL.

03 / CONTROLLED ABLATIONS

Both collection and post-training matter.

Proactive collection alone does not consistently improve success under behavior cloning. On the same proactive data, preference post-training matches or improves on HG-DAgger across all four tasks in the lowest-data regime.

Four-panel success-rate comparison of reactive, hybrid, and proactive correction data with HG-DAgger and PHIL. Outcomes depend on both the collection mode and the training objective.

Separating the two contributions. Three base demonstrations for Banana and ten for each other task; five correction episodes per condition. Hybrid combines two proactive and three reactive episodes. Mixing improves HG-DAgger here but reduces PHIL performance, showing that the benefit depends on the objective.

View full size (new tab)
04 / BEHAVIOR ANALYSIS

Fewer unnecessary recovery actions.

For both HG-DAgger and PHIL, proactive-trained Diffusion Policies show fewer unnecessary recovery behaviors in every evaluated task–demonstration setting.

Recovery counts for four manipulation tasks: proactive-trained policies exhibit fewer unnecessary recovery behaviors than reactive-trained policies across the evaluated conditions.

A separate evaluation of successful behavior. Each point averages counts over ten successful episodes per condition. Necessary recovery after actual errors is excluded. The horizontal axis is the share of correction frames in the combined training data and decreases to the right.

View full size (new tab)

Takeaways

Better corrections begin
before the mistake.

01

Make future motion visible.

XR previews let people act on a policy’s proposed motion before it is executed.

02

Learn from the contrast.

Rejected predictions and executed corrections provide a context-aligned preference signal.

03

Use a small correction budget.

Five correction episodes improve mean success over HG-DAgger in the evaluated Diffusion Policy and X-VLA settings.

Watch PHIL in action