VISUOMOTOR POLICY LEARNING

Heading Flow

Guiding Visuomotor Action Generation
with Predicted Motion Direction

Haotian Liu1,2Wei Li2Xin Wen2Yuan Ma2Xin Li2Peijin Jia2Zhen Zhu2Bailin Li2Kun Zhan2Dian Wang3,*

1 Northeastern University2 Li Auto3 University of Macau

* Corresponding author

An observed robot scene informs a predicted XY heading. This heading shapes a directional Gaussian prior that flow matching transforms into an action chunk.
Predict a direction. Shape the source. Generate the action.

01 / THE IDEA

Coarse intent.
Fine-grained action.

A robot often knows which way to move before it knows every detail of the motion.

Heading Flow makes this motion intent explicit. From visual observations, it predicts a planar heading supervised directly by demonstration actions. The same heading both conditions the action generator and shapes its initial Gaussian distribution.

This directional bias preserves the source distribution’s full support, leaving room for corrective motions. Spatial Motion Query (SMQ) extends the idea to multiple consecutive segments, capturing direction changes within an action chunk.

88.6%

LIBERO-10 success rate

SMQ-DiT · 500 evaluation episodes
+36.7 pp

Under camera perturbations

Heading Flow vs. StarVLA (GR00T)
75%

Pen to Drawer success rate

SMQ-DiT · 15 of 20 real-robot trials

02 / METHOD

Give generation a heading.

A shared geometric cue connects what the robot sees with how action generation begins.

01

Predict motion intent

A lightweight predictor estimates a planar direction and its validity from observations. Supervision comes from the net XY translation in demonstration actions.

02

Shape the starting point

Heading Gaussian shifts the source mean and stretches its covariance along the predicted direction. Nonzero variance across directions preserves corrective alternatives.

03

Generate the details

The same heading conditions the flow model. Visual observations guide the complete action chunk, including vertical movement, orientation, and gripper commands.

SPATIAL MOTION QUERY

One action chunk. More than one direction.

A single heading can hide turns and reversals. SMQ uses learned queries to read spatial observation features and predict a heading for each consecutive segment. In our setup, four queries guide four segments of a 16-step action chunk.

Spatial Motion Query architecture: learned segment queries attend to spatial observation tokens, predict local headings and validity, and guide both a segmented Gaussian source and the DiT velocity model.
Segment-specific predictions guide both the source distribution and the DiT action generator.
Why preserve variation around the predicted direction?

Heading is a partial description of motion. Heading Zero enforces exact alignment of the sampled chunk’s net XY direction. Heading Gaussian instead biases the source while retaining lateral and backward alternatives. When a heading is predicted to be invalid, the policy falls back to the base Gaussian.

Comparison of raw XY source samples and their chunk sums over 128 paired noise draws for Heading Zero, Heading Gaussian, and SMQ-DiT.
Raw XY samples (top) and chunk sums (bottom), across 128 paired noise draws. Green marks the ground-truth direction; dashed magenta marks the predicted global heading.

03 / REAL-ROBOT EXPERIMENTS

Multi-step manipulation.
In the physical world.

A single-arm PiperX robot, wrist and side-view cameras, and delta action control at 15 Hz.

SMQ-DiT / SUCCESSFUL ROLLOUTS

Prepare Fruit

50 training demonstrations
Successful rollout 01Original speed (1×)

Pick up the banana and strawberry and place both onto the plate. These three successful rollouts show SMQ-DiT completing the task across different initial arrangements.

Choose a rollout

UNSEEN-OBJECT GENERALIZATION

Generalization to
an unseen carrot.

We replace the banana with a carrot that was absent from the training demonstrations. Using the same SMQ-DiT checkpoint, the robot picks and places the strawberry and the unseen carrot onto the plate.

Two successful examples of transfer to an unseen carrot.

Example 01SMQ-DiT · Original speed (1×)
Example 02SMQ-DiT · Original speed (1×)
Four real-robot frames: initial drawer arrangement, opening the drawer, positioning the pen above the open drawer, and the final arrangement.
MULTI-STEP MANIPULATION

Pen to Drawer

70 demonstrations

Open the drawer, pick up the pen, place it inside, then push the drawer closed. Demonstrations include positional perturbations of the pen holder.

  1. Open drawer
  2. Pick pen
  3. Place inside
  4. Close drawer

Reported task success

The paper evaluation uses 20 trials per policy and task. The videos above are selected successful examples; the carrot generalization clips are separate from these reported task results.

Real-robot success rates and successful trial counts
PolicyPrepare FruitPen to Drawer
Diffusion Policy40% 8/2050% 10/20
Heading Flow60% 12/2070% 14/20
SMQ-DiT70% 14/2075% 15/20
ON-ROBOT EXECUTION

Smooth transitions between action chunks.

The policy predicts a 16-step chunk and uses an eight-step execution horizon. At each transition, outgoing and incoming commands are blended over three control steps to reduce abrupt changes.

16steps predicted
8steps executed
3steps blended · 0.2 s

04 / SIMULATION EXPERIMENTS

From benchmarks
to robust action.

Heading guidance is evaluated with U-Net and DiT policies, and integrated into VLA action experts.

Local headings improve chunk-level control.

SMQ-DiT reaches 88.6% success, an 11.0 percentage-point gain over the DiT flow-matching baseline.

Ten tasks, 50 initial states per task. Training uses 450 demonstrations and validation uses 50. Values report the best evaluated success rate using one checkpoint across all ten tasks.

Direction guidance under distribution shifts.

With the GR00T-style expert, Heading Flow improves camera-perturbation success by 36.7 percentage points and noise-perturbation success by 25.3 points. Gains vary by perturbation suite.

LIBERO-Plus success rates (%) · Qwen3-VL backbone
Action expert / methodCameraRobotLanguageLightBackgroundNoiseLayoutTotal
StarVLA (PI)64.357.282.894.294.079.678.277.2
+ Heading Flow72.453.278.494.095.883.482.678.6
+ SMQ74.651.286.996.795.690.288.282.2
StarVLA (GR00T)32.950.888.196.285.762.073.667.8
+ Heading Flow69.646.286.395.895.587.382.679.1
+ SMQ70.249.686.494.896.688.486.580.5

Bold values indicate the best result within each action-expert group. Total is the sample-weighted mean, following the StarVLA evaluation calculation. On small screens, scroll the table horizontally to view all suites.

05 / CITATION

Heading Flow

Bibliographic information for this manuscript.

@unpublished{liu2026headingflow,
  title = {Heading Flow: Guiding Visuomotor Action
           Generation with Predicted Motion Direction},
  author = {Liu, Haotian and Li, Wei and Wen, Xin and
            Ma, Yuan and Li, Xin and Jia, Peijin and
            Zhu, Zhen and Li, Bailin and Zhan, Kun
            and Wang, Dian},
  year = {2026},
  note = {Manuscript},
  url = {https://seanliu7081.github.io/heading-flow/}
}