ROBOTREPLICA / SITE 01 / SO-101

VLA-Replica

A Low-Cost, Reproducible Benchmark for Real-World Evaluation of Vision-Language-Action Models

SITE MAINTAINERS

Intelligent Robotics and Vision Lab at the University of Texas at Dallas

01 / ABSTRACT

Abstract

Vision-Language-Action (VLA) models have shown strong promise for general-purpose robotic manipulation, but their real-world evaluation remains limited by a lack of accessible, reproducible, and consistent benchmarks. Simulation benchmarks fail to capture real-world complexity, while existing real-world benchmarks often require expensive hardware, centralized evaluation, or are limited in task diversity.

We introduce VLA-Replica, a low-cost, easily reproducible real-world benchmark for evaluating VLA models. Built from off-the-shelf components, the system can be quickly assembled and replicated across laboratories, providing a consistent environment for policy evaluation anywhere in the world. It includes diverse manipulation tasks, a small-scale demonstration dataset for target-domain adaptation, and protocols for both in-distribution and out-of-distribution evaluation. Results across independently constructed setups demonstrate the reproducibility of the benchmark.

02 / OVERVIEW

Overview

One standardized platform, ten manipulation tasks, and a complete train-to-evaluation workflow.

VLA-Replica hardware, cameras, workspace, and task overview
The platform combines the SO-101 follower arm, a light box, top and wrist cameras, and a standardized manipulation workspace.
10manipulation tasks
500expert demonstrations
90reference scenes
ID + OODevaluation tracks

03 / BUILDING THE PLATFORM

Building the Platform

A user with no prior knowledge of the benchmark assembled the setup within one hour using the published guide.

Replicate the physical setup

  1. Source the SO-101, cameras, enclosure, lighting, and task objects.
  2. Assemble the workspace and fixed camera mounts.
  3. Calibrate the robot and cameras using the documented procedure.
  4. Verify the platform against the reference scenes.

04 / CAMERA CALIBRATION

Calibrate consistent camera views

We use an AprilTag to guide camera placement and overlay reference images on the live camera feeds to calibrate the top and wrist cameras.

This helps maintain consistent camera viewpoints across replicated setups.

AprilTag guidanceLive reference overlayTop + wrist cameras

05 / MANIPULATION TASKS

Manipulation Tasks

Ten tasks test physical skill, object interaction, and instruction following across training, in-distribution, and out-of-distribution conditions.

Open detailed task reference
01—04

Pick-and-Place

Place bread and bowls, stack blocks, and collect multiple objects.

05—07

Object Interaction

Fold a towel, open an oven, and erase a whiteboard.

08—10

Counting & Memory

Shake pepper, lift a bowl, and press a button a specified number of times.

Ten VLA-Replica manipulation tasks and their ID and OOD variations
Benchmark task suite. Each column shows a task with its training and evaluation variations.

06 / EXPERT DEMONSTRATIONS

Expert Demonstrations

We provide 50 demonstrations for each task for training or fine-tuning.

Download demonstration data
Examples of expert demonstrations

07 / TEST SCENES

Test Scene Reference Images

PLACEMENT GUIDE

Match each test scene precisely.

Use the corresponding reference image to reproduce the object positions and orientations before every evaluation rollout. This walkthrough demonstrates the placement process.

Ninety fixed reference scenes make physical evaluation repeatable. The first row of each image shows ID conditions and the second shows OOD conditions, except for tasks without an OOD variant.

Reference scenes for Put bread on plate
01Put bread on plate
Reference scenes for Put bowl on coaster
02Put bowl on coaster
Reference scenes for Stack blocks
03Stack blocks
Reference scenes for Fold towel
04Fold towel
Reference scenes for Open oven
05Open oven
Reference scenes for Clean whiteboard
06Clean whiteboard
Reference scenes for Pour pepper
07Pour pepper
Reference scenes for Lift bowl
08Lift bowl
Reference scenes for Press button
09Press button
Reference scenes for Collect blocks
10Collect blocks

08 / LEADERBOARD

Leaderboard

Verified results include compatible publicly released policies evaluated by RobotReplica, alongside submitted policies. Success rates use five rollouts per task; switch evaluation tracks, sort by any task, and use the video link to inspect every rollout for a method.

RankMethodVideo
1π₀.₅60.540.80.80.40.410.60.40.40.40.2↗
2MolmoAct2-SO100_10170.4610.80.40.40.60.40.60.20.20↗
3GR00T1.780.360.80.800.210.20.20.200.2↗
4π₀50.340.80.6000.80.20.40.20.20.2↗
5SmolVLA30.260.60.20.200.60.40.200.20.2↗
6ACT10.180.40000.40.40.20.20.20↗
7DiT-D20.160.4000.20.20.60.2000↗
8X-VLA40.140.40.2000.6000.200↗
9DiT-F20.120.40000.20.40.2000↗

Select Average or any task heading to rank methods. Use the arrow in the Video column to open all five rollout videos for every task.

TRAINING DETAILS

Training & Model Parameters

Implementation and training settings used for each evaluated policy.

ParameterACTDiT-DDiT-FSmolVLAX-VLAπ₀π₀.₅MolmoAct2-SO100_101GR00T1.7
Training typeFrom scratchFrom scratchFrom scratchFine-tuningFine-tuningFine-tuningFine-tuningFine-tuningFine-tuning
Dataset size500 demos500 demos500 demos500 demos500 demos500 demos500 demos500 demos500 demos
Batch size12812812812812816161632
Training steps40K40K40K40K40K40K40K40K40K
GPUs211441144
GPU typeA6000 AdaH200H200A6000 AdaA6000 AdaH200H200A6000 AdaA6000 Ada
Action chunk size323232323232323232
Number of action steps322424323232323232
Vision encoderDefaultDefaultDefaultDefaultTrainableFrozenDefaultDefaultDefault
ImplementationLeRobotLeRobotLeRobotLeRobotLeRobotLeRobotLeRobotLeRobotLeRobot
Learning rateDefaultDefaultDefaultDefaultDefaultDefaultDefaultDefaultDefault

References

  1. Zhao, Tony Z., Vikash Kumar, Sergey Levine, and Chelsea Finn. “Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.” arXiv preprint arXiv:2304.13705, 2023.
  2. Jones, Bryson. “Dissecting and Open-Sourcing Multitask Diffusion Transformer Policy.” Blog post, 2026.
  3. Shukor, Mustafa, Dana Aubakirova, Francesco Capuano, Pepijn Kooijmans, Steven Palma, Adil Zouitine, Michel Aractingi, et al. “SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.” arXiv preprint arXiv:2506.01844, 2025.
  4. Zheng, Jinliang, Jianxiong Li, Zhihao Wang, Dongxiu Liu, Xirui Kang, Yuchun Feng, Yinan Zheng, et al. “X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model.” arXiv preprint arXiv:2510.10274, 2025.
  5. Black, Kevin, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, et al. “π₀: A Vision-Language-Action Flow Model for General Robot Control.” arXiv preprint arXiv:2410.24164, 2024.
  6. Physical Intelligence, Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, et al. “π₀.₅: A Vision-Language-Action Model with Open-World Generalization.” arXiv preprint arXiv:2504.16054, 2025.
  7. Fang, Haoquan, Jiafei Duan, Donovan Clay, Sam Wang, Shuo Liu, Weikai Huang, Xiang Fan, et al. “MolmoAct2: Action Reasoning Models for Real-world Deployment.” arXiv preprint arXiv:2605.02881, 2026.
  8. NVIDIA. “GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.” arXiv preprint arXiv:2503.14734, 2025.

09 / CODE & RESOURCES

Run the benchmark

The public implementation includes training, calibration, scene references, and evaluation utilities.

View code on GitHub

10 / CITATION

BibTeX

Please cite VLA-Replica if it helps your research.

@misc{huang2026vlareplicalowcostreproduciblebenchmark,
  title={VLA-REPLICA: A Low-Cost, Reproducible Benchmark for
    Real-World Evaluation of Vision-Language-Action Models},
  author={Alex S. Huang and Jiahui Zhang and Shiqing Tang and Yu Xiang},
  year={2026},
  eprint={2605.20774},
  archivePrefix={arXiv},
  primaryClass={cs.RO},
  url={https://arxiv.org/abs/2605.20774}
}

11 / CONTACT

Evaluate with RobotReplica

Submit your policy information through the RobotReplica evaluation form. The team will coordinate the hosted evaluation and leaderboard submission with you.

Request an evaluation

ACKNOWLEDGEMENTS

This work was supported in part by the National Science Foundation under Grant Nos. 2346528 and 2520553, the NVIDIA Academic Grant Program Award, and gift funding from XPeng.