Chunked robot policies predict \(H\) actions per call, execute a short prefix, and throw the rest away before replanning, so reactivity is paid for with frequent policy calls. Action Upcycling is a training-free rule that keeps executing the discarded tail while the predicted action velocity stays smooth, reading nothing but the action chunk itself. Across four policies, three simulation benchmarks, and a real robot, it cuts policy calls by 1.2–1.7× with no loss in success rate.
| Action Upcycling | |
|---|---|
| Policy acceleration | 1.2–1.7× fewer policy calls · success rate ≥ baseline |
| Model-agnostic | Compatible with VLA (π0.5, SmolVLA, GR00T) · WAM (FastWAM) |
| Training-free | No fine-tuning · no architecture change |
| Complementary | Stacks with other policy acceleration methods (few-step sampling, streaming decoding) · up to 6.7× |
A chunked policy \(\pi\) predicts \(H\) actions per call, \(A=\pi(o_t)=(a_1,\dots,a_H)\), executes the prefix \(a_1,\dots,a_h\), and discards the tail. Action Upcycling reuses the trustworthy part of the tail, so that the \(i\)-th call executes \(h+\Delta h_i\) actions. With the mean execution length \(\bar h = \frac{1}{N}\sum_i (h+\Delta h_i)\), the upcycling ratio \(r = \bar h / h \in [1, H/h]\) gives the reduction in policy calls: \(r=2\) halves the calls and \(r=1\) recovers the default setting. Everything is computed from the predicted chunk itself: no gradients, no attention maps, no extra samples.
Let \(v_k\) be the velocity induced by action \(a_k\) (\(v_k=a_k\) for relative actions, \(v_k=a_k-a_{k-1}\) for absolute ones). Accumulate the fluctuation from the first tail action onward:
\(c_k\) is non-decreasing, stays small while the predicted motion is steady, and rises once it fluctuates. Empirically, the gap between a discarded action and its replanned version grows with \(c_k\).
Given a threshold \(\tau\), execute the tail while the accumulated fluctuation stays below it:
\(\tau=0\) recovers the standard protocol and \(\tau=\infty\) executes whole chunks. The next policy call is issued when the gate stops instead of on a fixed clock.
\(\Delta h_i\) is non-decreasing in \(\tau\), so the ratio \(r\) grows with \(\tau\) and the threshold for a target \(r\) is found by a simple search. Collect the tail signals \(c_{h+1},\dots,c_H\) of predicted chunks into a pool \(\mathcal C\), which needs no extra rollouts because \(c_k\) depends only on the chunk, and pick the smallest \(\tau\in\mathcal C\) with \(\bar h(\tau)\ge r\,h\).
The pool comes from the policy's own rollouts, without demonstrations or success labels. One scalar \(\tau\) per model and benchmark, fixed offline; no per-task tuning, no run-time adaptation.
We apply Action Upcycling to the released checkpoints of π0.5, SmolVLA, GR00T N1.7, and FastWAM (a world action model) without any fine-tuning, on LIBERO (2,000 trials), LIBERO-Plus (504 perturbed trials), and RoboTwin 2.0 (50 bimanual tasks, clean and randomized settings). The baseline is each model's standard deployment protocol.
| Benchmark | Model | Baseline policy | + Action Upcycling | ||||
|---|---|---|---|---|---|---|---|
| succ. ↑ | calls / ep ↓ | s / ep ↓ | succ. ↑ | calls / ep ↓ | s / ep ↓ | ||
| LIBERO | π0.5 | 96.9 | 32.4 (1×) | 4.41 | 97.9 (+1.0) | 21.9 (1.5×) | 3.00 |
| SmolVLA | 82.6 | 19.3 (1×) | 1.82 | 82.8 (+0.2) | 14.7 (1.3×) | 1.39 | |
| GR00T N1.7 | 96.1 | 22.6 (1×) | 2.59 | 96.6 (+0.5) | 15.0 (1.5×) | 1.72 | |
| FastWAM | 97.4 | 15.7 (1×) | 1.32 | 97.6 (+0.2) | 13.4 (1.2×) | 1.12 | |
| LIBERO-Plus | π0.5 | 83.7 | 38.5 (1×) | 5.27 | 86.5 (+2.8) | 23.9 (1.6×) | 3.27 |
| SmolVLA | 31.8 | 29.4 (1×) | 2.77 | 33.3 (+1.5) | 24.1 (1.2×) | 2.27 | |
| GR00T N1.7 | 82.1 | 36.4 (1×) | 4.18 | 82.3 (+0.2) | 22.7 (1.6×) | 2.61 | |
| FastWAM | 49.4 | 33.4 (1×) | 2.79 | 49.6 (+0.2) | 22.5 (1.5×) | 1.88 | |
| RoboTwin 2.0 | π0.5 | 59.7 | 40.9 (1×) | 6.22 | 60.5 (+0.8) | 29.8 (1.4×) | 4.53 |
| SmolVLA | 34.8 | 50.6 (1×) | 5.77 | 47.4 (+12.6) | 29.7 (1.7×) | 3.39 | |
| FastWAM | 89.5 | 10.6 (1×) | 2.27 | 90.6 (+1.1) | 8.7 (1.2×) | 1.86 | |
Table 1. Success rate (%), policy calls per episode, and inference time per episode. Per-call latency is unchanged (π0.5 136 ms, SmolVLA 94 ms, GR00T 115 ms, FastWAM 84–214 ms), so the time saving equals the call reduction. Action Upcycling reduces policy calls by 1.2–1.7× and matches or improves success in all 11 cells. RoboTwin 2.0 covers all 50 tasks in both clean and randomized settings.
Policy calls per episode. Shorter is better.
| Method | LIBERO | LIBERO-Plus | RoboTwin 2.0 | ||||||
|---|---|---|---|---|---|---|---|---|---|
| succ. | calls / ep | s / ep | succ. | calls / ep | s / ep | succ. | calls / ep | s / ep | |
| Baseline | 96.9 | 32.4 (1×) | 4.41 | 83.7 | 38.5 (1×) | 5.27 | 59.7 | 40.9 (1×) | 6.22 |
| + AAC (CVPR 2026), K = 20 samples | 97.3 | 23.5 (1.38×) | 18.9 | 84.5 | 27.1 (1.42×) | 21.7 | 53.1 | 46.1 (0.89×) | 42.9 |
| + AutoHorizon (ECCV 2026), attention | 97.4 | 22.3 (1.45×) | 3.03 | 85.5 | 24.5 (1.57×) | 3.33 | 56.8 | 28.0 (1.46×) | 4.26 |
| + Action Upcycling (ours) | 97.9 | 21.9 (1.48×) | 3.00 | 86.5 | 23.9 (1.61×) | 3.27 | 60.5 | 29.8 (1.37×) | 4.53 |
Table 2. Success rate (%), policy calls per episode, and inference time per episode (s) of π0.5. Best in bold, second best underlined. AAC draws 20 chunks per call and stops where the samples disagree, so its per-call latency grows six- to seven-fold (802–930 ms) and its time per episode exceeds the baseline; on RoboTwin 2.0 it even calls the policy more often than the baseline (0.89×). AutoHorizon reads the action expert's self-attention; on RoboTwin 2.0 it cuts calls slightly more (1.46×) but drops success below the baseline (56.8 vs. 59.7). Action Upcycling reads only the predicted chunk, adds no latency, and achieves the highest success rate on all three benchmarks.
We deploy Action Upcycling on a YAM 6-DoF arm with a linear gripper, driven at 30 Hz from a top-down camera, a wrist camera, and the joint state. We evaluate two policies fine-tuned on this arm, π0.5 and GR00T N1.6. π0.5 runs eight tasks and GR00T N1.6 runs four, each for 10 episodes under both protocols: 12 task–model cells × 2 protocols × 10 episodes = 240 real-world episodes in total.
Action Upcycling completes tasks faster than the baseline: time per episode drops from 24.9 s to 21.5 s with π0.5 and from 28.0 s to 24.9 s with GR00T N1.6, with fewer policy calls (Tables 3 and 4). Four tasks per page; use the arrows or the chips to browse.
| Task | Success rate ↑ | Calls / ep ↓ | Time / ep (s) ↓ | |||
|---|---|---|---|---|---|---|
| Baseline | Ours | Baseline | Ours | Baseline | Ours | |
| Soccer ball → plate (layout 1) | 10/10 | 10/10 | 26.9 | 26.8 | 14.4 | 17.4 |
| Soccer ball → plate (layout 2) | 10/10 | 10/10 | 36.5 | 20.7 | 19.0 | 12.9 |
| Basketball → basket | 10/10 | 10/10 | 31.3 | 24.9 | 16.5 | 15.8 |
| Baseball → basket | 10/10 | 10/10 | 32.1 | 25.0 | 17.1 | 16.1 |
| Drawer: black ball → drawer | 8/10 | 10/10▲ | 73.8 | 49.2 | 37.3 | 30.9 |
| Drawer: basketball → drawer | 8/10 | 8/10 | 61.2 | 40.7 | 31.6 | 25.5 |
| Drawer: baseball → basket | 7/10 | 9/10▲ | 67.9 | 40.6 | 34.3 | 25.5 |
| Drawer: basketball → pot | 9/10 | 10/10▲ | 56.9 | 44.5 | 28.7 | 27.8 |
| Total / mean | 72/80 | 77/80 | 48.3 | 34.0 | 24.9 | 21.5 |
Table 3. Real-world results with π0.5 on a YAM arm, 10 episodes per task. Success rises from 72/80 to 77/80 while policy calls per episode drop from 48.3 to 34.0 (1.42×). On the pick-and-place tasks the baseline already succeeds in every episode and calls drop by 1.30×; on the long-horizon drawer tasks success improves from 32/40 to 37/40 with 1.48× fewer calls. Time per episode also falls from 24.9 s to 21.5 s.
| Task | Success rate ↑ | Calls / ep ↓ | Time / ep (s) ↓ | |||
|---|---|---|---|---|---|---|
| Baseline | Ours | Baseline | Ours | Baseline | Ours | |
| Soccer ball → plate (layout 1) | 10/10 | 10/10 | 56.8 | 50.3 | 17.8 | 17.7 |
| Soccer ball → plate (layout 2) | 8/10 | 9/10▲ | 102.2 | 81.7 | 31.5 | 28.8 |
| Basketball → basket | 8/10 | 8/10 | 95.6 | 69.9 | 29.6 | 24.3 |
| Baseball → basket | 8/10 | 8/10 | 107.9 | 82.3 | 33.1 | 28.9 |
| Total / mean | 34/40 | 35/40 | 90.6 | 71.0 | 28.0 | 24.9 |
Table 4. Real-world results with GR00T N1.6 on the four pick-and-place tasks, 10 episodes per task. Success 34/40 → 35/40, policy calls per episode 90.6 → 71.0, time per episode 28.0 s → 24.9 s. The short chunk (H = 16) caps the achievable ratio.