A systematic study of vision-language-action models

VLANeXt Family

From Core Recipes to Emerging Paradigms

Two connected papers. One shared framework for studying and building strong VLA models.

ICML 202601 · Core recipe

VLANeXt: Recipes for Building Strong VLA Models

Xiao-Ming Wu1, Bin Fan2, Kang Liao1, Jian-jian Jiang2, Runze Yang3, Yihang Luo1, Zhonghua Wu4, Wei-Shi Zheng2, Chen Change Loy1,5,*

A unified design study distilling more than 500 experiments into 12 findings and the base VLANeXt model.

Extended paper · 202602 · Emerging paradigms

VLANeXt Family: A Systematic Study of VLA Models from Core Recipes to Emerging Paradigms

Xiao-Ming Wu1, Kang Liao1, Yihang Luo1, Bin Fan2, Jian-jian Jiang2, Runze Yang3, Zhonghua Wu4, Wei-Shi Zheng2, Chen Change Loy1,5,*

Five new variants test the core recipe across model scaling, latent-action pretraining, latent predictive representation learning, and world action modeling.

1 S-Lab, Nanyang Technological University2 Sun Yat-sen University3 Shanghai Jiao Tong University4 SenseTime Research5 ACE Robotics

* Corresponding author: Chen Change Loy

Research film · 2:02

Meet VLANeXt Family

From core recipes to emerging paradigms, with real-world robot demonstrations.

One research trajectory

A strong recipe.
A broader family.

Which VLA design choices matter, and do they remain effective as the field evolves? We study both questions under a unified training and evaluation framework, from a simple RT-2-style baseline to six VLANeXt variants.

500+Core design experiments
12Findings behind the core recipe
4Emerging paradigms
98.4%VLANeXt-L · LIBERO average
83.9%Base · LIBERO-plus average
Ablation trajectory from the RT-2-style baseline to VLANeXt and the S, Base, L, LAM, JEPA, and WAM variants
The core study uses the Spatial suite, moving from LIBERO to LIBERO-plus as performance saturates. The Family results at the bottom are averages across all four LIBERO suites.

01 · Establish the recipe

Make each design choice count

The ICML study examines foundational components, perception essentials, and action modeling perspectives. Their combination yields the original VLANeXt, now the Base model of the family.

Explore the core recipe

02 · Test its generality

Change the scale and the prior

The extended study explores model scaling, latent-action pretraining, latent predictive representation learning, and world action modeling. Controlled ablations examine whether the same recipe remains useful in each setting.

Meet the family

03 · Build on shared foundations

One codebase, six variants

Both papers share a model implementation, training entry point, and evaluation pipeline. Configuration files select the backbone and learning objectives.

Models and configurations

VLANeXt · ICML 2026

The core architecture

Multi-view images, language instructions, and proprioception enter a multimodal backbone. Learnable meta queries softly connect it to a dedicated policy module, which predicts action chunks with flow matching and frequency-domain regularization.

VLANeXt core architecture: multi-view vision, language and proprioception feed the MLLM; soft meta queries condition a diffusion policy with a frequency-domain action loss
The base VLANeXt architecture. A Qwen3-VL-2B backbone and a dedicated policy module combine the effective design choices identified in the core study.
01 · Foundation

A dedicated action pathway

A deeper policy module, action chunking, a strong VLM, and soft layer-wise connections provide the foundation. Flow matching models continuous action distributions.

02 · Perception

The right inputs, in the right place

Third-person and wrist views improve robustness. Proprioception is most effective when conditioned in the VLM; simply adding past images does not improve this setup.

03 · Action modeling

Learn the structure of motion

A frequency-domain loss regularizes action trajectories with little overhead. Auxiliary future-image prediction helps, but is left out of Base because it nearly triples training cost.

Table I · Core design studyFull recipe exploration results

Success rates (%) on the Spatial suite. Each block varies one design aspect along the exploration trajectory. These are Spatial results, distinct from the four-suite benchmark averages below.

Core-recipe ablations · Spatial suite success rate (%) · Family Table I
Design / variantLIBEROLIBERO-plus · unseen perturbations
OriginalCameraRobotLanguageLightBackgroundNoiseLayoutTotal ↑
Foundational components · RT-2-style baseline
Baseline19.8-------< 5.0
Policy module design
Baseline19.8-------< 5.0
Separate Head30.20.810.031.015.424.84.030.116.6
Large Policy Module64.40.512.679.734.632.98.563.134.0
Action chunking horizon
Action Chunk 164.40.512.679.734.632.98.563.134.0
Action Chunk 475.45.328.067.942.550.414.870.440.0
Action Chunk 874.65.626.085.656.855.811.763.943.4
Action learning objective
bin Classification74.65.626.085.656.855.811.763.943.4
VQ-VAE Classification58.83.244.367.242.543.87.148.336.5
Regression85.45.132.390.562.768.67.775.648.4
DDIM80.04.852.680.370.555.49.468.348.3
Flow Matching80.07.234.679.246.957.811.177.445.0
VLM backbone capacity
Paligemma69.81.117.132.122.932.62.824.918.9
LLaMA3.2 + SigLip80.07.234.679.246.957.811.177.445.0
Qwen3VL-2B90.09.642.074.675.068.627.983.653.7
Qwen3VL-4B95.812.266.093.389.081.029.988.664.8
VLM-policy connection
Loose Connection90.09.642.074.675.068.627.983.653.7
Tight Connection90.014.451.781.068.567.825.182.155.4
Soft Connection91.811.858.389.272.974.419.972.556.2
Perception essentials · temporal observation history
Current Frame Image91.811.858.389.272.974.419.972.556.2
Temporal Observation History85.07.268.651.565.862.020.880.850.2
Camera views
Third-person camera view91.811.858.389.272.974.419.972.556.2
Multiview (third-person + wrist)97.664.954.091.897.793.085.590.180.5
Proprioception conditioning
No Proprioception Input97.664.954.091.897.793.085.590.180.5
Proprioception to VLM98.087.262.286.298.393.892.096.687.7
Proprioception to Policy96.262.869.192.392.596.587.288.383.4
Proprioception to VLM & Policy97.677.973.482.390.194.690.688.884.8
Proprioception projector
Linear Projector98.087.262.286.298.393.892.096.687.7
Transformer Projector96.496.858.084.695.298.896.993.388.8
Transformer Projector & MAE97.091.151.189.972.886.978.785.678.9
Action modeling · world modeling perspective
Normal98.087.262.286.298.393.892.096.687.7
World Modelling98.094.480.376.999.798.893.293.290.3
Time-series forecasting perspective
Normal98.087.262.286.298.393.892.096.687.7
Frequency Domain Loss99.095.778.686.999.798.898.096.693.1

“—” or “-” denotes an unreported result. Values follow Table I of the extended paper.

VLANeXt Family · Extended study

Four directions.
A common foundation.

The Family study asks whether the core recipe remains effective across model-capacity priors, action-like priors, semantic priors, and world-dynamics priors.

Four extensions of VLANeXt: model scaling, latent-action pretraining, latent predictive representation learning, and world action modeling
From the core architecture to the Family: (a) model scaling, (b) latent-action pretraining, (c) latent predictive representation learning, and (d) world action modeling.
A · Model-capacity priors

VLANeXt-S / Base / L

Apply the recipe to Qwen3.5-0.8B, Qwen3-VL-2B, and Qwen3-VL-4B backbones. The remaining architecture and training recipe are shared.

B · Action-like priors

VLANeXt-LAM

Learn latent actions from current and future frames with a VQ-VAE. Pretrain on those visual transitions, then fine-tune on labeled robot actions.

C · Semantic priors

VLANeXt-JEPA

Predict future DINOv3 features alongside actions. The best evaluated variant predicts the observation at the end of the current action horizon.

D · World-dynamics priors

VLANeXt-WAM

Replace the VLM with a pretrained Wan video-generation backbone and jointly learn video and actions. The fast connection transfers predictive priors without direct action attention to future video tokens.

Models and configurations

Meet the VLANeXt family

Success rates (%) averaged over Spatial, Object, Goal, and Long. Backbone names follow the released configurations.

VLANeXt family · released configurations and LIBERO average success rates
ModelBackboneFocusLIBERO avg. ↑Configuration
VLANeXt-SQwen3.5-0.8BModel scaling · small96.7YAML ↗
VLANeXt (Base / B)Qwen3-VL-2BCore recipe · ICML 202697.4YAML ↗
VLANeXt-LQwen3-VL-4BModel scaling · large98.4YAML ↗
VLANeXt-LAMQwen3-VL-2BLatent-action pretraining97.6YAML ↗
VLANeXt-JEPAQwen3-VL-2BLatent predictive representation learning97.7YAML ↗
VLANeXt-WAMWan2.2-TI2V-5BWorld action modeling98.2YAML ↗

VLANeXt (Base / B) is the original ICML model. The five additional variants are introduced in the extended paper. Backbone sizes are not total model parameter counts.

Controlled evaluation

Does the recipe
carry over?

Removing core ingredients reduces performance in every reported Family ablation.

Table V · Emerging paradigms

The core recipe across the family

LIBERO four-suite average success rate (%). Parentheses show the change in percentage points relative to each variant’s full recipe.

Core-recipe ablations across emerging paradigms · Family Table V
ModelBackboneFull recipe ↑w/o Chunkingw/o Multi-vieww/o Floww/o Prop.w/o Freq.
VLANeXt-SQwen3.5-0.8B96.788.7 (-8.0)89.9 (-6.8)94.0 (-2.7)95.8 (-0.9)96.6 (-0.1)
VLANeXt-LQwen3-VL-4B98.493.2 (-5.2)94.6 (-3.8)96.6 (-1.8)96.8 (-1.6)97.1 (-1.3)
VLANeXt-LAMQwen3-VL-2B97.687.8 (-9.8)85.8 (-11.8)93.2 (-4.4)94.6 (-3.0)95.3 (-2.3)
VLANeXt-JEPAQwen3-VL-2B97.791.2 (-6.5)93.3 (-4.4)96.7 (-1.0)95.9 (-1.8)95.5 (-2.2)
VLANeXt-WAMWAN2.2-5B98.294.1 (-4.1)89.9 (-8.3)96.0 (-2.2)97.6 (-0.6)98.0 (-0.2)

Full recipe Chunking: action chunking · Flow: flow matching · Prop.: proprioception · Freq.: frequency-domain loss.

Table VI · Latent-action pretrainingThe latent-action constraint matters

VQ-VAE reaches 97.6%, while training the action-only baseline for the same total number of steps reaches 97.0%. Longer training alone does not explain the improvement.

Latent-action pretraining · Family Table VI
ModelLAM constraintPretrain stepsFine-tune stepsSpatialObjectGoalLongAverage ↑
VLANeXt——10k99.099.296.694.897.4
VLANeXt · matched duration——30k98.898.896.294.097.0
VLANeXt-LAMVICReg20k10k95.499.697.095.496.9
VLANeXt-LAMVAE20k10k98.097.895.296.296.8
VLANeXt-LAMVQ-VAE20k10k98.099.895.697.097.6
Table VII · Latent predictive representation learningAlign semantic prediction with the action horizon

Predicting DINOv3 features at the end of the current action horizon reaches 97.7%, compared with 96.6% for the episode-final target and 96.4% for the Emu3.5-token objective.

Latent predictive objectives · Family Table VII
ModelGeneratorObjectiveSpatialObjectGoalLongAverage ↑
VLANeXt—Action only99.099.296.694.897.4
VLANeXtEmu3.5 tokenizer + DiTAction + horizon-final image97.699.895.492.896.4
VLANeXt-JEPADINO encoder + DiTAction + episode-final image98.099.296.493.096.6
VLANeXt-JEPADINO encoder + DiTAction + horizon-final image98.099.897.095.897.7
Table VIII · World action modelingJoint video and action learning improves control

On the same Wan backbone, the action-only baseline reaches 95.6%. Joint video and action learning improves this to 98.2% with the fast connection and 97.6% with the joint connection.

World action modeling variants · Family Table VIII
ModelBackboneObjectiveConnectionSpatialObjectGoalLongAverage ↑
VLANeXtQwen3-VL-2BAction only—99.099.296.694.897.4
VLANeXt-WAMWAN2.2-5BAction only—95.099.696.491.295.6
VLANeXt-WAMWAN2.2-5BAction & videoFast97.899.499.696.098.2
VLANeXt-WAMWAN2.2-5BAction & videoJoint95.299.897.098.297.6
Table IX · Model scalingThe core recipe across model scales

We evaluate the core recipe with different backbone sizes from the Qwen3.5 and Qwen3-VL families under a shared LIBERO training and evaluation protocol.

Backbone scaling · Family Table IX
BackboneSpatialObjectGoalLongAverage ↑
Qwen3.5-0.8B95.499.495.496.496.7
Qwen3.5-2B97.8100.098.097.498.3
Qwen3.5-4B97.499.697.295.297.4
Qwen3-VL-2B99.099.296.694.897.4
Qwen3-VL-4B99.2100.098.096.498.4
Qwen3-VL-8B98.099.497.495.897.7

Tables VI–IX report per-suite success rates (%) and their four-suite average. Bold blue cells mark the best result in each metric column of the displayed comparison, including ties. All table numbers refer to the extended paper.

Core-recipe benchmarks

Strong performance
from the base model

The original VLANeXt reaches 97.4% on LIBERO. On LIBERO-plus, it achieves 83.9% without perturbation-specific training, a gain of 14.3 percentage points over OpenVLA-OFT.

Table II · LIBERO

Four suites of manipulation tasks

Success rate (%) across Spatial, Object, Goal, and Long. Comparisons are those reported in the paper.

LIBERO benchmark · success rate (%) · Family Table II
ModelSpatialObjectGoalLongAverage ↑
Direct policy models
Diffusion Policy78.392.568.350.572.4
Octo78.985.784.651.575.1
MDT78.587.573.564.876.4
Vision-language-action models
TraceVLA84.685.275.154.174.8
OpenVLA84.788.479.253.776.5
SpatialVLA88.289.978.655.578.1
WorldVLA85.689.082.659.079.1
CoT-VLA87.591.687.669.083.9
π0-Fast96.496.888.660.285.5
π090.086.095.073.086.0
NORA92.295.489.474.687.9
SmolVLA93.094.091.077.088.8
UniVLA96.596.895.692.095.2
FLOWER97.599.196.194.996.9
π0.598.898.298.092.496.9
OpenVLA-OFT97.698.497.994.597.1
Our core recipe
VLANeXt (Base)99.099.296.694.897.4

Best result Bold blue cells indicate the best result in each column, including ties.

Table III · LIBERO-plus

Robustness under unseen perturbations

All models are trained on LIBERO and evaluated directly on LIBERO-plus. Seven perturbation categories test visual, physical, and semantic robustness.

LIBERO-plus benchmark · success rate (%) · Family Table III
ModelSuiteCameraRobotLanguageLightBackgroundNoiseLayoutTotal ↑
OpenVLAAverage0.83.523.08.134.815.228.515.6
WorldVLAAverage0.127.941.643.717.110.938.025.0
NORAAverage2.237.065.145.758.612.862.139.0
UniVLAAverage1.846.269.669.081.021.231.942.9
π0Average13.86.058.885.081.479.068.953.6
π0-FastAverage65.121.661.073.273.274.468.861.6
OpenVLA-OFT · per-suite results
OpenVLA-OFTSpatial88.340.080.598.397.396.393.984.0
OpenVLA-OFTObject38.925.499.073.797.672.371.866.5
OpenVLA-OFTGoal62.025.253.293.992.575.259.163.0
OpenVLA-OFTLong38.738.287.089.486.863.576.966.4
OpenVLA-OFTAverage56.431.979.588.793.375.874.269.6
VLANeXt · per-suite results
VLANeXtSpatial95.778.686.999.798.898.096.693.1
VLANeXtObject99.548.598.699.384.799.878.286.5
VLANeXtGoal96.663.651.597.570.896.863.976.2
VLANeXtLong69.772.090.386.975.881.784.379.7
VLANeXtAverage90.465.781.895.982.594.180.883.9

Best average Highlighting compares only the Average rows.

VLANeXt in action

From simulation
to real robots

Real-world experiments validate the core VLANeXt model on Franka single-arm and Aloha bimanual setups. These demonstrations and results are from the core-recipe study.

Table IV · Real-world evaluation

Single-arm and bimanual manipulation

Successful trials / 20 attempts per task. Each task uses 50 training demonstrations after DROID pretraining.

Real-world evaluation · successful trials out of 20 · Family Table IV
ModelSingle armBimanual arms
Clean tableDrawer manipulationClean tableLifting
OpenVLA-OFT7/207/205/209/20
π010/208/2010/2013/20
VLANeXt14/2011/2011/2015/20
Clean table · single arm
Clean table · single arm
Open drawer and place object · single arm
Open drawer and place object · single arm
Clean table · bimanual arms
Lifting · bimanual arms

LIBERO demonstrations

Completing different tasks

Pick up the black bowl between the plate and the ramekin and place it on the plate.
Pick up the tomato sauce and place it in the basket.
Open the middle drawer of the cabinet.
Put both the alphabet soup and the tomato sauce in the basket.

LIBERO-plus demonstrations

One task, seven types of perturbations

“Pick up the black bowl next to the plate and place it on the plate,” under seven distinct perturbation categories.

Citation

Build on VLANeXt

If you find our work useful, please cite both VLANeXt and VLANeXt Family.

VLANeXt & VLANeXt Family

@inproceedings{wu2026vlanext,
  title={VLANeXt: Recipes for Building Strong VLA Models},
  author={Xiao-Ming Wu and Bin Fan and Kang Liao and Jian-jian Jiang and Runze Yang and Yihang Luo and Zhonghua Wu and Wei-Shi Zheng and Chen Change Loy},
  booktitle={ICML},
  year={2026}
}

@article{wu2026vlanextfamily,
  title={VLANeXt Family: A Systematic Study of VLA Models from Core Recipes to Emerging Paradigms},
  author={Xiao-Ming Wu and Kang Liao and Yihang Luo and Bin Fan and Jian-jian Jiang and Runze Yang and Zhonghua Wu and Wei-Shi Zheng and Chen Change Loy},
  journal={arXiv preprint arXiv:2602.18532},
  year={2026},
  url={https://arxiv.org/abs/2602.18532}
}