VISUOMOTOR POLICY LEARNING

Past2Next

Past Action Conditioning Policy
with Data-Augmented Tokenization

Haotian Liu1,*,†Shoukang Yu1,*Haojie Huang1,†Boce Hu1Dian Wang2Robert Platt1

1 Northeastern University2 University of Macau

* Co-first authors   ·   † Corresponding authors

Past actions and their dynamic features narrow next-action prediction from a broad action space to a local feasible region around the extrapolated motion.
Use the recent past to make the next action more predictable.

01 / THE IDEA

Recent motion.
A clearer next step.

What a robot just did helps explain what it should do next.

Discrete action policies face two challenges: representing continuous motion accurately and predicting the right action tokens. A more precise tokenizer alone does not make its tokens easier to predict.

Past2Next addresses both. Recent executed actions and their first- and second-order differences give the policy a local motion context. SO(3) augmentation broadens the tokenizer’s action coverage. Together, these make finer action representations more useful for control.

85.5%

LIBERO-Long success rate

+29.2 percentage points over OAT
66.3%

RoboCasa success rate

+9.5 percentage points over OAT
90%

Prepare Fruit success rate

18 of 20 real-robot evaluation trials

02 / METHOD

Better action tokens.
Better context to predict them.

Two training stages address action reconstruction and next-token predictability.

STAGE 01 / TOKENIZATION

Broaden motion coverage

Train an OAT action tokenizer with SO(3) augmentation. One shared 3D rotation transforms the translations and rotation vectors throughout a chunk, preserving the motion’s shape and leaving gripper commands unchanged.

STAGE 02 / ACTION PREDICTION

Condition on recent motion

Freeze the tokenizer, then train an autoregressive Transformer. Visual observations, seven recent executed actions, and two finite-difference features provide context for predicting the next action tokens.

Stage 1 trains an OAT encoder, finite scalar quantizer and decoder with rotated action chunks. Stage 2 conditions an autoregressive Transformer on two observations, two dynamics features and seven past actions. Eight predicted tokens decode to a 16-step action chunk.
The tokenizer is frozen after Stage 1. At execution time, decoded actions update the history used for the next prediction.
CLOSED-LOOP EXECUTION

Predict, execute, update the history.

The illustrated policy generates eight tokens for a 16-step action chunk. It executes the first eight actions, then uses the latest seven executed actions as context for the next chunk.

16actions predicted
8actions executed
7past actions retained

03 / REAL-ROBOT EXPERIMENTS

Three tasks.
Motion with memory.

A Piper-X arm performs four-stage tabletop tasks across different initial arrangements.

Nut Washer Removal · Three placementsOriginal speed (1×)
CONTACT-RICH MANIPULATION

Nut Washer Removal

Unscrew and remove the nut before lifting the washer. Tight and loose nuts require different numbers of turns, while frames just before and after release can look similar.

  1. Unscrew nut
  2. Place nut
  3. Lift washer
  4. Place washer
15/20

successful evaluation trials
OAT: 6/20 · Diffusion Policy: 8/20

MULTI-OBJECT PICK AND PLACE

Prepare Fruit

Pick the strawberry and banana, then place each onto the plate. Three placements show the policy handling different initial object arrangements.

  1. Pick strawberry
  2. Place strawberry
  3. Pick banana
  4. Place banana
18/20

successful evaluation trials
OAT: 14/20 · Diffusion Policy: 12/20

MULTI-STAGE MANIPULATION

Pen to Drawer

Open the drawer, transfer the pen inside, then close the drawer. The policy must carry the task forward through successive changes in the scene.

  1. Open drawer
  2. Pick pen
  3. Place pen
  4. Close drawer
16/20

successful evaluation trials
OAT: 12/20 · Diffusion Policy: 10/20

Reported task success

Twenty trials per policy and task. Past2Next improves success across all three tasks, with nearly the same per-chunk inference time as OAT.

Successful trials / 20 · Inference per action chunk
PolicyNut & washerFruitDrawerTime
Diffusion Policy8/2012/2010/2040.7 ms
OAT6/2014/2012/2027.2 ms
Past2Next15/2018/2016/2027.4 ms

Videos show demonstration examples across three placements per task. Trial counts are from the paper’s evaluation. Diffusion Policy uses DDIM with ten denoising steps.

04 / SIMULATION EXPERIMENTS

More predictable tokens.
More successful actions.

Evaluated on ten LIBERO-Long tasks and ten RoboCasa tasks against discrete and continuous action policies.

LIBERO-Long

Past2Next reaches 85.5% success, 10.7 percentage points above the strongest reported baseline, A2A.

Success rate (%) ± standard error · higher is better
Diffusion Policy36.6± 0.2
Flow39.0± 1.2
Flow + past actions65.2± 1.4
A2A74.8± 2.8
FAST23.0± 0.5
ACT64.1± 2.2
AR-VLA35.9± 0.5
Dense67.7± 2.9
OAT56.3± 1.0
Past2Next85.5± 0.4

50 demonstrations per task. Long-horizon household tasks require completing several subtasks in sequence.

RoboCasa

Past2Next reaches 66.3% success, 3.7 percentage points above A2A and 9.5 points above OAT.

Success rate (%) ± standard error · higher is better
Diffusion Policy46.6± 1.0
Flow50.3± 1.0
Flow + past actions52.6± 0.9
A2A62.6± 2.0
FAST21.3± 1.1
ACT44.6± 0.9
AR-VLA43.1± 2.6
Dense58.8± 1.5
OAT56.8± 1.2
Past2Next66.3± 0.8

200 demonstrations per task: 50 human demonstrations and 150 generated with MimicGen. Tasks cover kitchen drawers, doors, faucets and appliances.

Evaluation protocol

Policies observe an in-hand camera and a side-view camera. Each method is trained with action chunk sizes of 16 and 32. Results use a chunk size of 32 for OAT on LIBERO-Long and the better chunk size for the remaining settings.

Each checkpoint is evaluated on 50 episodes per task. Values report the best mean success rate and standard error, as reported in Table 2 of the paper.

05 / WHY THE PAST HELPS

A smaller set
of plausible next actions.

Controlled studies separate the effect of conditioning from the tokenizer’s reconstruction accuracy.

SAME TOKENIZER · RICHER CONTEXT

History and dynamics work together.

With the same 1,000-code OAT tokenizer, adding raw past actions and finite-difference features improves LIBERO-Long success.

Observations only56.3%
+ Past actions62.2%
+ Dynamics features63.8%
+ Both68.8%
NEXT-TOKEN UNCERTAINTY

Fewer plausible choices per token.

Past-action conditioning lowers average conditional entropy with the same 1,000-code tokenizer.

Observation-only OAT2.87 nats
Past-action conditioned1.60 nats

The effective branching factor, exp(entropy), falls from about 18 to about 5 candidate tokens.

How does a larger codebook change performance?

A larger codebook provides finer action resolution but also more possible tokens. At a fixed 32-step chunk size, OAT peaks at 1,000 codes. Past2Next without SO(3) augmentation achieves its best result at 5,000 codes in this ablation.

LIBERO-Long success (%) · 32-step chunks · Past2Next without SO(3) augmentation
Codebook sizeOATPast2Next
24029.251.2
51253.558.0
1,00056.368.8
1,92054.663.2
5,00046.973.4

CITATION

Past2Next

If this work is useful for your research, please consider citing it.

@unpublished{liu2026past2next,
  title = {Past2Next: Past Action Conditioning Policy
           with Data-Augmented Tokenization},
  author = {Liu, Haotian and Yu, Shoukang and
            Huang, Haojie and Hu, Boce and
            Wang, Dian and Platt, Robert},
  year = {2026},
  note = {Manuscript}
}