SplineWAM: Adaptive Action Horizons for World Action Models via B-Spline Representations

1Tsinghua University2Shanghai Jiao Tong University3Xiaomi Robotics
4Peking University5CASIA

* Equal contribution † Corresponding authors

A fixed parameter budget. An adaptive action horizon.

Longer horizons through free space. Finer control around contact.

Abstract

World action models usually predict fixed-length action chunks. SplineWAM instead predicts a fixed-size set of B-spline parameters whose knot times determine the motion’s duration and resolution. Video supervision follows the same adaptive time axis. For asynchronous execution, Jacobian-Pullback Real-Time Chunking (JP-RTC) maintains continuity in the actions the robot actually executes. Across two simulation benchmarks and three bimanual real-world tasks, SplineWAM matches or improves task success while executing more control steps per policy call.

LIBERO-Plus zero-shot

+8.2 ppSuccess rate
−22%Policy calls
74.5% vs. 66.3% success · 26.2 vs. 33.6 calls/episode

RoboCasa

+4.4 ppSuccess rate
−26%Policy calls
73.5% vs. 69.1% success · 16.3 vs. 21.9 calls/episode

SplineWAM-Knot vs. action chunking · synchronous evaluation · default execution settings.

Motivation

Free-space motion is smooth and predictable over longer intervals; grasping and insertion require finer resolution and frequent feedback. Fixed action chunks spend unnecessary policy calls on free-space motion yet can delay replanning near contact. For a large WAM backbone, these redundant calls also limit the throughput of shared cloud inference.

Method

Let the action representation decide how far to act and when to replan.

For an interactive introduction to B-splines, see What are B-splines? or Fitting a B-spline from the B-Spline Policy project website.

01 / REPRESENT

Predict a continuous trajectory

A fixed window of cubic B-spline parameters decodes into a variable number of control steps.

02 / ALIGN

See where precision matters

Sample training video frames at fitted knot times, concentrating supervision around complex motion.

03 / EXECUTE

Keep asynchronous actions continuous

JP-RTC corrects spline parameters through the decoder to match the actions already committed.

Video–action alignment

One shared, non-uniform time axis.

Jacobian-Pullback RTC

Match the decoded actions, allowing different spline parameterizations.

Real Robot Experiments

Three bimanual tasks with asynchronous execution. Thirty trials per method, per task.

Table 2: success rate, progress, and executed chunk length on three real robot tasks
MethodArrange BookshelfCharge EarphoneClean Whiteboard
SR (%) ↑PG (%) ↑ChunkSR (%) ↑PG (%) ↑ChunkSR (%) ↑PG (%) ↑Chunk
Action Chunking20.052.232.026.773.332.050.067.532.0
SplineWAM + Naive RTC10.035.639.813.343.346.546.768.353.3
SplineWAM + JP-RTC23.354.438.626.757.545.760.074.252.4

Table 2. SR requires all task nodes; PG is the mean fraction of nodes reached. Chunk is the mean number of executed control steps per policy call, not a speedup in wall-clock time.

Naive RTC executes longer chunks but achieves lower SR and PG than JP-RTC on all three tasks. JP-RTC executes 1.2–1.6× more steps per call than action chunking; the largest success gain is on the highly compressible Clean Whiteboard task (+10.0 pp).

Playback rates below are relative to the supplied recordings.

Lift-2S mobile robot arranging books between a desk and shelf
01

Arrange Bookshelf

Long-horizon mobile manipulation

Move between the desk and shelf to place three books upright, using the Lift-2S mobile platform.

Trial 1 · SplineWAM + JP-RTC · 1× playback
Trial 2 · SplineWAM + JP-RTC · 1× playback
ARX AC One robot erasing a marked whiteboard
02

Clean Whiteboard

Repetitive motion under a visual decision

Identify the remaining marks, erase the board, and return the eraser to its holder. The repetitive strokes make this the most compressible task.

Trial 1 · SplineWAM + JP-RTC · 1× playback
Trial 2 · SplineWAM + JP-RTC · 1× playback
ARX AC One robot connecting a charger and an earphone case
03

Charge Earphone

High-precision alignment

Plug the charger into the wall socket, then insert USB-C into the earphone case. Both insertions require precise alignment.

Trial 1 · SplineWAM + JP-RTC · 1× playback
Trial 2 · SplineWAM + JP-RTC · 1× playback

Adaptive Action Horizons

The plot follows a successful Charge Earphone trajectory with JP-RTC. Horizons lengthen during free-space transport and shorten around gripper closure, contact, and insertion.

Simulation Benchmarks

Synchronous evaluation with the same Fast-WAM backbone, demonstrations, and training budget.

Zero-shot LIBERO-Plus. Models are trained on LIBERO and tested directly under LIBERO-Plus perturbations, without adaptation or fine-tuning.

Table 1: success rate, policy calls per episode, and steps per call on LIBERO-Plus and RoboCasa
Action representationLIBERO-Plus (zero-shot)RoboCasa
Success ↑Calls/ep. ↓Steps/callSuccess ↑Calls/ep. ↓Steps/call
Action Chunking66.3%33.68.0069.1%21.916.00
BEAST66.6%35.58.0065.0%24.616.00
Naive B-Spline w/o video60.0%38.77.9255.0%32.714.26
Naive B-Spline70.7%27.38.3669.5%16.917.98
SplineWAM-Window73.3%27.18.3971.9%17.518.10
SplineWAM-Knot ours74.5%26.28.4173.5%16.318.18

Table 1. Success is episode-weighted; calls count policy-model invocations. Steps/call counts executed control steps. Results are single-run point estimates.

Execution Budget and Fitting Tolerance

When varying the execution budget, SplineWAM matches the action-chunking baseline’s best success rate with fewer policy calls on both RoboCasa and LIBERO-Plus. These are matched-success operating points from the ablation, distinct from the default settings above.

Citation

@article{guo2026splinewam,
  title  = {SplineWAM: Adaptive Action Horizons for World Action
            Models via B-Spline Representations},
  author = {Guo, Jun and Han, Xiaoshen and Li, Qiwei and Sun, Nan
            and Li, Peiyan and Wang, Heyun and Lai, Hang
            and Zhang, Weinan and Li, Xinghang and Liu, Huaping},
  journal = {arXiv preprint arXiv:2609.39873},
  year    = {2026},
  eprint  = {2609.39873},
  archivePrefix = {arXiv},
  url     = {https://arxiv.org/abs/2609.39873}
}

Figure viewer