ICLR 2027 · Anonymous supplementary material

NEEDLEwork: Offline Rewriting of Robot Data with Verified Local Stitches

TL;DR: NEEDLE adds short, verified action bridges between recorded observations, so a policy trained on imperfect demonstrations skips detours, borrows better continuations, and recovers from failed attempts.

How can we achieve better task execution from imperfect demonstrations?

Imperfect demonstrations contain useful behavior alongside repeated grasps, detours, and unsuccessful attempts. A policy that imitates the corpus as recorded inherits those inefficiencies. Yet the best robot behavior may not exist in any single trajectory: one demonstration may do the first half well and another the second. We ask whether the useful pieces can be connected into better training data from offline data alone, without a simulator, new robot time, or privileged state.

Real-robot comparisons, shown at 2× speed. The full evaluation below includes success rates and all 300 trials.

01 / Intuition and method

Why stitch demonstrations?

Consider folding a sweater: one demonstration folds the sleeves reliably but needs several corrective grasps for the bottom fold; another struggles with the sleeves but folds the bottom cleanly. Connecting their useful segments could produce a better demonstration than either alone.

This is the promise of trajectory stitching: connecting useful segments to create better examples than those recorded. An ideal stitch would offer a better continuation and supply actions that physically reach it.

What could stitching offer?

A within-trajectory shortcut (S1) bypasses an inefficient detour inside one demonstration. A cross-trajectory transplant (S2) combines complementary behaviors from two successful demonstrations that were never recorded together. A failure funnel (S3) connects a state on a failed attempt to a successful continuation. The animations illustrate these opportunities; realizing them requires finding connections that are both useful and feasible.

Within a trajectory. The demonstration reaches the target state again after a detour from the source state. A bridge from source to target skips the detour, and the recorded actions from the target complete the task.

Which connections would improve the data?

Connecting two nearby observations may simply add a detour. NEEDLE therefore looks for targets with better recorded continuations. From a successful demonstration, the connection and target continuation must together take fewer steps than the original route. From a failed attempt, the target must offer a successful continuation where none was recorded. Among accepted bridges from the same source, NEEDLE keeps the one with the shortest combined bridge and remaining successful continuation. For successful sources, this also maximizes the steps saved. These criteria identify useful destinations; they do not establish that the robot can reach them.

How should we best extract useful information from existing demonstrations for stitching?

Key insight. Rather than learning a globally accurate dynamics model or task-value function and asking it to reason over long trajectories, NEEDLE reduces stitching in high-dimensional observation spaces to a series of short, locally supervised reachability problems. Short recorded segments already teach both feasibility and duration: they show which observations a sequence of actions connects, and in how many steps, even in an episode that ultimately fails.

How does NEEDLE stitch demonstrations together?

The recordings contain the useful segments, but may contain no actions connecting them. A goal-conditioned inverse dynamics model (IDM) proposes those missing actions from a source observation to a target. Plausible actions can still miss the target, so an action-conditioned verifier, trained on these short recorded segments, assesses whether each prefix reaches it and after how many steps. Its training pairs share one recorded action chunk: the chunk is a positive example for the goal it reaches within the horizon, and a negative example for a goal beyond the horizon or on another trajectory. A negative label says only that this chunk does not reach that goal, not that no action sequence could; finding one that does is the objective of stitching.

From proposal to bridge. For each source–target pair, NEEDLE selects the shortest prefix that passes the verifier’s threshold as a bridge. This requires neither a simulator nor reconstructed images along the bridge.

Eight action sequences are replayed in simulation to illustrate verification. The verifier scores whether each prefix reaches the target; the dashed line marks its acceptance threshold. Three proposals pass, and the shortest accepted prefix becomes the bridge.

How can the policy learn from a bridge?

A bridge supplies new actions but no intermediate images. NEEDLE pairs these actions with recorded observations at and before the source, using recorded actions to lead into the bridge. Varying the starting observation lets the policy learn the new continuation from several preceding situations.

The bridge may end slightly away from its target, so the target’s next actions may not apply there. Training therefore stops at the bridge endpoint; the target continuation is learned separately from its own observations. Because a stitch creates an alternative local continuation rather than an exact synthetic trajectory, the original actions at the bridge source and along the skipped segment stay in the training set, at a controlled sampling weight, so the policy keeps supervision for situations it may still encounter.

Each window pairs a recorded observation with the actions leading into the bridge (cyan), then the bridge actions (green). Sliding the start exposes the connection from several preceding observations. Slots after the bridge carry no training loss. The target continuation (blue) and original source and skipped-segment actions remain separate training examples.

02 / Real-robot evaluation

Task success, completion time, and repeated motion

We compare NEEDLE with diffusion policy (DP) and advantage-weighted regression (AWR) on dish racking and sweater folding. Demonstrations are collected with handheld UMI devices. Policies receive wrist-camera RGB and proprioception relative to each trajectory’s starting pose; there is no shared coordinate frame across demonstrations. We first report task success and all evaluation rollouts, then examine completion time and repeated motion.

Does stitching improve task success?

NEEDLE succeeds in 44 of 50 dish trials and 48 of 50 sweater trials. Each policy is evaluated on the same 50 starting states per task.

Success Across All Trials ↑ Higher Is Better50 trials per policy and task

Compared with the strongest baseline on each task, success rises by 18 percentage points on dish racking and 24 on sweater folding.

All 300 evaluation trials

Select a task and policy. Each grid shows all 50 trials in trial order at 2× speed; cells turn green for success or red for failure and freeze at the end.

Does stitching reduce completion time?

Mean completion time among successes can be misleading, since a policy that solves fewer difficult starts can look faster. We therefore also count, for each time limit, the fraction of all 50 trials completed within it, failures included. NEEDLE has the lowest mean on sweater folding. On dish racking, AWR has a slightly lower mean, while NEEDLE succeeds on more trials.

Time to Complete the Task ↓ Lower Is BetterSuccessful trials only · mean ± standard error

Trials Completed Within a Time Limit ↑ Higher Is BetterAll 50 trials per policy and task · time from the first arm motion to the operator’s success key

Each curve counts trials completed within the time limit; failed trials do not count. NEEDLE reaches higher completion rates earlier on both tasks.

Does the policy repeat less motion, and where do failures stop?

The demonstrations contain hesitations, repeated grasps, and re-attempts. These clips show baselines repeating a motion without completing the task, while NEEDLE completes them. The figure below marks, for every failed trial, the stage the policy never completed.

Points of Failure ↑ Higher Is BetterOf 50 trials per policy, after each stage · a line steps down where trials stopped and ends at the success count

▸ Click a filled point above to watch a trial that failed there.

▸ Click a filled point above to watch a trial that failed there.

The numbers beside a line give the number of trials that policy lost at that stage, shown when the loss is 7 or more. All three policies lift the plate in 47 of 50 dish trials; the difference comes at the rack, where DP and AWR fail 13 and 11 times and NEEDLE 3. On sweaters, AWR most often never gets past the first sleeve, which it folds and lets fall open repeatedly, while DP more often stalls at the second sleeve. Where points overlap, click again to switch policy.

03 / Simulation experiments

How do the bridges change what a policy learns?

We evaluate on three Robomimic tasks with mixed-quality human demonstrations: Can, Square, and two-arm Transport. Each task provides 40 successful demonstrations and 150 failed attempts for methods that use them. Results report the mean and standard error across five training seeds, using each seed’s highest checkpoint success rate on the same 50 fixed evaluation initial conditions. All methods use the same checkpoint-evaluation frequency and fixed evaluation starts.

Does the success gain hold across simulated tasks?

NEEDLE has the highest success rate on all three tasks: on average 8.4 percentage points above DP and 6.9 above the strongest baseline on each task.

(a) Success Rate (%) ↑ Higher Is Bettermean ± standard error over five seeds

(b) Steps Taken on Successful Trials ↓ Lower Is Better

Baselines cover the three usual ways of getting more out of a fixed dataset: adding connections with another stitching method (MBTS), selecting or reweighting its examples (SBR, AWR, SARM), and choosing actions more carefully at run time (IDQL). NEEDLE’s successful trials are also shorter than DP’s on Can and Square.

Do bridges lead to faster completion?

We test whether individual bridges reduce the remaining steps to completion, and whether the trained policy succeeds within a given step budget.

Can, Square, and Transport: accepted bridges shift remaining completion steps lower; NEEDLE’s policy reaches higher success within the evaluation budget than DP and IDQL.
Left of each task: starting from the same recorded moment, follow either the original recorded actions or the bridge for a few steps, then let the same policy finish the task. This is an advantage-style comparison under the same policy: both continuations start from the same source, and we ask which leaves the policy in a state from which success is easier to complete. After a bridge, the task is finished in fewer remaining steps (the blue distribution sits further left). Right of each task: give every policy the same starting states and a time limit, and count how many are solved as the limit grows. NEEDLE’s curve rises earlier and ends higher than DP and IDQL, which means it solves more starts, and solves them sooner.

Can failed demonstrations teach recovery?

A failed trajectory can contain useful intermediate behavior even though it never reaches the goal. A failure funnel connects such a state to a successful continuation, turning a state that had no successful action target into one with explicit supervision toward success.

For each task, visual-feature plots show training starts added by failure-to-success bridges, alongside recovery curves for NEEDLE, DP, and AWR.
Left of each task: a map of the situations the policy is trained in. The gray region is what the successful demonstrations cover; the colored region is what the bridges from failed attempts add. Right of each task: drop each policy into a situation taken from a failed attempt and see whether it recovers and finishes. NEEDLE recovers more often than DP on all three tasks (by 5.1, 3.3, and 3.9 percentage points on Can, Square, and Transport) and more often than AWR, the non-DP baseline with the highest average success across the three tasks on nominal resets.

Does checking the bridges matter?

A useful connection must pass two separate tests. It must be useful: following the target’s recorded actions should give the policy a better route to completion. It must also be feasible: the proposed bridge actions must actually reach the target. Candidate selection estimates usefulness from recorded continuations; the verifier judges feasibility from the proposed actions. We test whether its scores depend on the actions, whether accepted bridges reach their targets in simulation, and whether verification improves policy success.

Does the Verifier Respond to the Actions?

how well it tells the true recorded actions apart from a substitute, with the same source and target (1.0 = perfect, 0.5 = guessing)

Policy Success With and Without Verification (%) ↑ Higher Is Better

same number of bridges per task, dashed line: DP

Do Accepted Bridges Arrive? ↑ Higher Is Betterfraction of bridges that end within a given distance of the target when replayed in the simulator

The verifier separates the true actions from every kind of substitute (AUROC 0.85 to 0.98), including actions taken from the same episode a little later and those from the closest starting arm pose. Accepted bridges are more likely than rejected ones to finish within 5 cm of the target, by 21.9, 34.6, and 61.1 percentage points on Can, Square, and Transport. A policy trained on unchecked proposals, with the same number of bridges, loses success on every task; on Transport it falls from 53.2% to 20.4%. The distance to target is measured only in this simulator diagnostic; the method itself never sees object positions.

What if the demonstrations are already good?

On three DexMimicGen tasks, where the demonstrations are near-optimal, NEEDLE matches DP on Threading and improves by 1.3 and 4.0 percentage points on Coffee and Three Piece Assembly. AWR performs best on Threading.

DexMimicGen Success (%) · Mean ± Standard Error Across Three Training Seeds
Task DP SBR AWR IDQL SARM NEEDLE
Coffee 62.0 ± 1.2 60.7 ± 2.4 52.7 ± 1.8 49.3 ± 1.8 39.3 ± 1.8 63.3 ± 0.7
Three Piece Assembly 51.3 ± 0.7 44.7 ± 1.8 47.3 ± 1.3 43.3 ± 3.5 26.7 ± 1.8 55.3 ± 1.8
Threading 46.0 ± 1.2 40.7 ± 2.4 50.0 ± 1.2 40.0 ± 1.2 34.0 ± 2.0 46.0 ± 1.2

NEEDLE is most valuable when a dataset contains complementary, inefficient, or unsuccessful behavior that can be recombined into better supervision. When the demonstrations are already close to the desired behavior, rewriting has correspondingly less to add, and it does not reduce success relative to standard behavior cloning (DP).

04 / Ablations

Which design choices matter?

Three questions, mirroring Appendix B of the paper: which kinds of connection help, how best to train on stitched data, and whether the bridge actions themselves matter. Bars are success rates on the three Robomimic tasks; DP and NEEDLE are the same numbers as above.

Which kinds of connection help?

NEEDLE uses three kinds of bridge: skipping ahead within a demonstration (S1), connecting two demonstrations (S2), and connecting a failed attempt to a successful demonstration (S3). Each row trains NEEDLE with a subset of them; the full method is S1+S2+S3.

No single kind of connection matches the full method on all three tasks, and which kind helps most depends on the task: cross-demonstration bridges contribute most on Can, while on Transport adding failure bridges to S1+S2 gives the largest gain (45.6% to 53.2%).

What is the best way to train on stitched data?

A bridge gives the policy a second way to act from a recorded observation. Each row below changes one choice about how that is turned into training examples. Under the chart, each change opens into a short animation of what it does to the windows built from one bridge.

Open any of the four to see it applied to the same set of training windows.

Remove Original Source Actions

The source observation normally yields two training examples: one that follows the bridge and one that keeps its own recorded actions. This drops the recorded one, so that observation only ever teaches the bridge.

Sample Original Source Actions More

Keeps both, but draws the recorded one as often as any ordinary example instead of at the reduced weight the method uses by default.

Remove Skipped Segments

Drops every training example whose window starts inside the stretch the bridge skips over, so that part of the demonstration is never trained on.

Append Target’s Remaining Actions

Where a window runs past the end of the bridge, the default marks the remaining slots as not scored. This fills them with the target demonstration’s own next actions instead, which assumes the bridge lands exactly on the target rather than near it.

The default keeps the original actions next to the bridge (at half the sampling weight), keeps the examples from the bypassed segment, and stops the training target where the bridge ends. Removing the original source actions costs 6 to 9 points on every task; removing the bypassed segment costs 3 to 6. Appending the target demonstration’s remaining actions directly after the bridge, as if the connection were exact, lowers success on Square and Transport (42.0% and 45.6% against 47.2% and 53.2%) and is the lowest variant on Transport, consistent with the risk that a bridge’s endpoint differs from the recorded target.

Do the bridge actions themselves matter?

This removes every training example whose window contains bridge actions, while keeping the original demonstrations' examples at the weights NEEDLE gives them. What remains is the recorded data, reweighted by where the bridges were found.

Removing the bridge actions lowers success from 83.6% to 77.2% on Can, from 47.2% to 42.8% on Square, and from 53.2% to 45.6% on Transport, so a large part of NEEDLE’s gain over DP comes from the bridge actions rather than from reweighting the recorded data.

05 / Evaluation details

Protocol and limitations

Limitations

NEEDLE assumes useful connections can be made with short action sequences. It may fail when reaching a target requires long-horizon contact dynamics, substantial object reconfiguration, or feedback unavailable within an open-loop action chunk. Verification is statistical: false positives can introduce harmful training targets. The method requires episode-level success labels and is designed for policies that predict action sequences. Real-robot results cover two tasks and two baselines; simulation uncertainty is reported over five training seeds on Robomimic and three on DexMimicGen.

Real-robot evaluation

Each task is evaluated on 50 starting states. Objects are placed using a reference-image overlay with the arms at their starting pose. Outcomes are marked by hand. The mean completion times run from the first arm motion to the marked completion in the recorded video. Policies see wrist-camera images and relative pose readings only.

Simulation evaluation

Following prior evaluation practice (Chi et al., 2023), we evaluate checkpoints every 25 epochs on the same 50 fixed evaluation initial conditions and report the mean and standard error of the maximum success rate observed across seeds. All methods use the same checkpoint-evaluation frequency and the same fixed evaluation starts.

Robomimic policies train for 300 epochs with five training seeds and a fixed train/validation split; DexMimicGen results use three training seeds. Steps per successful trial are measured at the selected checkpoint. The completion and recovery figures use the evaluation runs from the paper’s analysis, with time limits of 500 steps for Can and Square and 700 for Transport.