Don't Throw Away the Tail: Action Upcycling for Policy Acceleration

Taesung Kwon* Jangho Park* Sunwoo Park Youngmin Kim Seonghyun Jin Youngjun Jun Kyumin Choi Jong Chul Ye
KAIST
*Equal contribution
Paper Code Real-robot videos
Playback

Same policy (π0.5), same checkpoint, same episode index; the only difference is when the policy is re-queried. Left: the standard protocol re-plans every 15 ticks on a fixed clock and discards the remaining 35 actions of each chunk. Right: Action Upcycling keeps executing the tail while its velocity stays smooth. Top-down camera on the left of each frame, wrist camera on the right.

1.2–1.7×fewer policy calls in simulation
11 / 11model × benchmark cells with success ≥ baseline
48.3 → 34.0VLA calls per episode on a real YAM arm (π0.5), 72/80 → 77/80 success
0training, extra samples, or model internals required

Abstract

Chunked robot policies predict \(H\) actions per call, execute a short prefix, and throw the rest away before replanning, so reactivity is paid for with frequent policy calls. Action Upcycling is a training-free rule that keeps executing the discarded tail while the predicted action velocity stays smooth, reading nothing but the action chunk itself. Across four policies, three simulation benchmarks, and a real robot, it cuts policy calls by 1.2–1.7× with no loss in success rate.

Action Upcycling
Policy acceleration1.2–1.7× fewer policy calls · success rate ≥ baseline
Model-agnosticCompatible with VLA (π0.5, SmolVLA, GR00T) · WAM (FastWAM)
Training-freeNo fine-tuning · no architecture change
ComplementaryStacks with other policy acceleration methods (few-step sampling, streaming decoding) · up to 6.7×

Method

A chunked policy \(\pi\) predicts \(H\) actions per call, \(A=\pi(o_t)=(a_1,\dots,a_H)\), executes the prefix \(a_1,\dots,a_h\), and discards the tail. Action Upcycling reuses the trustworthy part of the tail, so that the \(i\)-th call executes \(h+\Delta h_i\) actions. With the mean execution length \(\bar h = \frac{1}{N}\sum_i (h+\Delta h_i)\), the upcycling ratio \(r = \bar h / h \in [1, H/h]\) gives the reduction in policy calls: \(r=2\) halves the calls and \(r=1\) recovers the default setting. Everything is computed from the predicted chunk itself: no gradients, no attention maps, no extra samples.

Speed
1

Velocity fluctuation as a signal

Let \(v_k\) be the velocity induced by action \(a_k\) (\(v_k=a_k\) for relative actions, \(v_k=a_k-a_{k-1}\) for absolute ones). Accumulate the fluctuation from the first tail action onward:

\[c_k=\sum_{j=h+1}^{k}\lVert v_j-v_{j-1}\rVert_2,\quad h<k\le H\]

\(c_k\) is non-decreasing, stays small while the predicted motion is steady, and rises once it fluctuates. Empirically, the gap between a discarded action and its replanned version grows with \(c_k\).

2

Adaptive horizon selection

Given a threshold \(\tau\), execute the tail while the accumulated fluctuation stays below it:

\[h_{\text{exec}}(A;\tau)=\max_{c_k\le\tau} k,\quad h\le k\le H\]

\(\tau=0\) recovers the standard protocol and \(\tau=\infty\) executes whole chunks. The next policy call is issued when the gate stops instead of on a fixed clock.

3

Threshold from a target ratio

\(\Delta h_i\) is non-decreasing in \(\tau\), so the ratio \(r\) grows with \(\tau\) and the threshold for a target \(r\) is found by a simple search. Collect the tail signals \(c_{h+1},\dots,c_H\) of predicted chunks into a pool \(\mathcal C\), which needs no extra rollouts because \(c_k\) depends only on the chunk, and pick the smallest \(\tau\in\mathcal C\) with \(\bar h(\tau)\ge r\,h\).

The pool comes from the policy's own rollouts, without demonstrations or success labels. One scalar \(\tau\) per model and benchmark, fixed offline; no per-task tuning, no run-time adaptation.

Simulation Results

We apply Action Upcycling to the released checkpoints of π0.5, SmolVLA, GR00T N1.7, and FastWAM (a world action model) without any fine-tuning, on LIBERO (2,000 trials), LIBERO-Plus (504 perturbed trials), and RoboTwin 2.0 (50 bimanual tasks, clean and randomized settings). The baseline is each model's standard deployment protocol.

Main result: success rate and policy calls

BenchmarkModelBaseline policy+ Action Upcycling
succ. ↑calls / ep ↓s / ep ↓succ. ↑calls / ep ↓s / ep ↓
LIBEROπ0.596.932.4 (1×)4.4197.9 (+1.0)21.9 (1.5×)3.00
SmolVLA82.619.3 (1×)1.8282.8 (+0.2)14.7 (1.3×)1.39
GR00T N1.796.122.6 (1×)2.5996.6 (+0.5)15.0 (1.5×)1.72
FastWAM97.415.7 (1×)1.3297.6 (+0.2)13.4 (1.2×)1.12
LIBERO-Plusπ0.583.738.5 (1×)5.2786.5 (+2.8)23.9 (1.6×)3.27
SmolVLA31.829.4 (1×)2.7733.3 (+1.5)24.1 (1.2×)2.27
GR00T N1.782.136.4 (1×)4.1882.3 (+0.2)22.7 (1.6×)2.61
FastWAM49.433.4 (1×)2.7949.6 (+0.2)22.5 (1.5×)1.88
RoboTwin 2.0π0.559.740.9 (1×)6.2260.5 (+0.8)29.8 (1.4×)4.53
SmolVLA34.850.6 (1×)5.7747.4 (+12.6)29.7 (1.7×)3.39
FastWAM89.510.6 (1×)2.2790.6 (+1.1)8.7 (1.2×)1.86

Table 1. Success rate (%), policy calls per episode, and inference time per episode. Per-call latency is unchanged (π0.5 136 ms, SmolVLA 94 ms, GR00T 115 ms, FastWAM 84–214 ms), so the time saving equals the call reduction. Action Upcycling reduces policy calls by 1.2–1.7× and matches or improves success in all 11 cells. RoboTwin 2.0 covers all 50 tasks in both clean and randomized settings.

Baseline policy+ Action Upcycling

Policy calls per episode. Shorter is better.

Comparison with adaptive-horizon methods (π0.5)

MethodLIBEROLIBERO-PlusRoboTwin 2.0
succ.calls / eps / epsucc.calls / eps / epsucc.calls / eps / ep
Baseline96.932.4 (1×)4.4183.738.5 (1×)5.2759.740.9 (1×)6.22
+ AAC (CVPR 2026), K = 20 samples97.323.5 (1.38×)18.984.527.1 (1.42×)21.753.146.1 (0.89×)42.9
+ AutoHorizon (ECCV 2026), attention97.422.3 (1.45×)3.0385.524.5 (1.57×)3.3356.828.0 (1.46×)4.26
+ Action Upcycling (ours)97.921.9 (1.48×)3.0086.523.9 (1.61×)3.2760.529.8 (1.37×)4.53

Table 2. Success rate (%), policy calls per episode, and inference time per episode (s) of π0.5. Best in bold, second best underlined. AAC draws 20 chunks per call and stops where the samples disagree, so its per-call latency grows six- to seven-fold (802–930 ms) and its time per episode exceeds the baseline; on RoboTwin 2.0 it even calls the policy more often than the baseline (0.89×). AutoHorizon reads the action expert's self-attention; on RoboTwin 2.0 it cuts calls slightly more (1.46×) but drops success below the baseline (56.8 vs. 59.7). Action Upcycling reads only the predicted chunk, adds no latency, and achieves the highest success rate on all three benchmarks.

Real-World Experiments

We deploy Action Upcycling on a YAM 6-DoF arm with a linear gripper, driven at 30 Hz from a top-down camera, a wrist camera, and the joint state. We evaluate two policies fine-tuned on this arm, π0.5 and GR00T N1.6. π0.5 runs eight tasks and GR00T N1.6 runs four, each for 10 episodes under both protocols: 12 task–model cells × 2 protocols × 10 episodes = 240 real-world episodes in total.

Playback

Baseline comparison

OursBaselinesuccessfailure

More results

Action Upcycling completes tasks faster than the baseline: time per episode drops from 24.9 s to 21.5 s with π0.5 and from 28.0 s to 24.9 s with GR00T N1.6, with fewer policy calls (Tables 3 and 4). Four tasks per page; use the arrows or the chips to browse.

Results

π0.5 (H = 50, h = 15, r = 1.25)

TaskSuccess rate ↑Calls / ep ↓Time / ep (s) ↓
BaselineOursBaselineOursBaselineOurs
Soccer ball → plate (layout 1)10/1010/1026.926.814.417.4
Soccer ball → plate (layout 2)10/1010/1036.520.719.012.9
Basketball → basket10/1010/1031.324.916.515.8
Baseball → basket10/1010/1032.125.017.116.1
Drawer: black ball → drawer8/1010/10▲73.849.237.330.9
Drawer: basketball → drawer8/108/1061.240.731.625.5
Drawer: baseball → basket7/109/10▲67.940.634.325.5
Drawer: basketball → pot9/1010/10▲56.944.528.727.8
Total / mean72/8077/8048.334.024.921.5

Table 3. Real-world results with π0.5 on a YAM arm, 10 episodes per task. Success rises from 72/80 to 77/80 while policy calls per episode drop from 48.3 to 34.0 (1.42×). On the pick-and-place tasks the baseline already succeeds in every episode and calls drop by 1.30×; on the long-horizon drawer tasks success improves from 32/40 to 37/40 with 1.48× fewer calls. Time per episode also falls from 24.9 s to 21.5 s.

GR00T N1.6 (H = 16, h = 9, r = 1.1)

TaskSuccess rate ↑Calls / ep ↓Time / ep (s) ↓
BaselineOursBaselineOursBaselineOurs
Soccer ball → plate (layout 1)10/1010/1056.850.317.817.7
Soccer ball → plate (layout 2)8/109/10▲102.281.731.528.8
Basketball → basket8/108/1095.669.929.624.3
Baseball → basket8/108/10107.982.333.128.9
Total / mean34/4035/4090.671.028.024.9

Table 4. Real-world results with GR00T N1.6 on the four pick-and-place tasks, 10 episodes per task. Success 34/40 → 35/40, policy calls per episode 90.6 → 71.0, time per episode 28.0 s → 24.9 s. The short chunk (H = 16) caps the achievable ratio.