Researchers submit a policy to a site with matching robot hardware. The volunteer site team runs it on maintained real-world tasks.
A PHYSICAL BENCHMARK NETWORK FOR ROBOT MANIPULATION
Your policy.
Our robot.
One comparable result.
RobotReplica is a non-commercial, community-led research initiative connecting organizations that voluntarily host standardized robot manipulation benchmarks.
What RobotReplica does
RobotReplica proactively evaluates compatible public policies and checkpoints, verifies the results, and maintains robot-specific leaderboards.

PARTNER ORGANIZATIONS & PEOPLE
IRVL · UT DallasRichardson, TexasSITE 01: SO-101
GRILL · CornellIthaca, New YorkSITE 03: Tianji Marvin Arms
PRACTICE · UCLALos Angeles, CaliforniaSITE 05: AgileX PiPER
IRIS · ASUTempe, ArizonaSITE 06: Allegro HandCURRENT ROBOTREPLICA SITESSeven sites across the United States
Our current network includes seven physical benchmark sites. We welcome additional organizations, robot platforms, and evaluation sites to join RobotReplica.
Become a partnerOVERVIEW
A shared evaluation service
for real robots.
Robotics results are difficult to compare when every lab uses a different robot, scene, and protocol. RobotReplica keeps physical benchmark sites running so the community can evaluate on consistent hardware and tasks.
Physical benchmark sites
Each partner maintains a robot, workspace, task objects, cameras, and an evaluation protocol.
Robot-matched evaluation
Partner sites evaluate submitted policies and compatible publicly released policies on matching robot hardware.
Verified leaderboards
RobotReplica maintains a board for each benchmark site using verified results from both evaluation paths.
Researchers do not need to rebuild the full benchmark. The benchmark stays at the host site; policies travel to it.
CURRENT SITES
Start with the robot
you already use.
Seven organizations are building the first RobotReplica sites. Each site operates its hardware, tasks, and evaluation process; RobotReplica maintains the verified robot-specific leaderboards.
SITE 01Intelligent Robotics and Vision Lab @ UT Dallas
SO-101
RobotReplica SO-101 site
A hosted SO-101 evaluation site for reproducible VLA-Replica manipulation benchmarks.
Benchmark tasksVLA-Replica evaluates 10 tabletop manipulation tasks in both in-distribution and out-of-distribution scenes.
- SO-101 arm
- UT Dallas
- Accepting evaluations
Site maintainers
LIVE LEADERBOARD
SO-101 · VLA-Replica
Average policy success rates across the official benchmark. Ranked by the in-distribution result.
OpenArm
RobotReplica OpenArm
A hosted OpenArm evaluation site extending the network to bimanual, contact-rich manipulation.
Benchmark tasksInitial benchmark tasks are being designed for the standardized OpenArm Cell; the final protocol is still in development.
- Bimanual platform
- Open-source hardware
- Protocol in development
Site maintainers
Generalizable Robot Intelligence and Learning Lab @ Cornell University
Tianji Marvin Arms
RobotReplica Marvin site
A hosted Tianji Marvin dual-arm evaluation site equipped with two WUJI hands for reproducible, force-aware manipulation benchmarks.
Benchmark tasksThe team is developing a benchmark around the force-controlled, humanlike Marvin arm platform.
- Tianji Marvin arms
- Two WUJI hands
- In development
Site maintainers
Computational Design and Fabrication Group @ MIT + Perceptual Science Group @ MIT
Franka Panda
RobotReplica Franka Panda site
A hosted Franka Panda evaluation site extending the network with a new reproducible manipulation benchmark.
Benchmark tasksThe team is developing its robot workspace, benchmark tasks, scene specifications, and baseline evaluation protocol.
- Franka Panda arm
- MIT CSAIL
- In development
Site maintainers
AgileX PiPER
RobotReplica PiPER site
A UCLA-hosted PiPER evaluation site extending the network with a reproducible manipulation benchmark for the compact robot arm.
Benchmark tasksThe PRACTICE Lab team is developing the workspace, task suite, scene specifications, and evaluation protocol for its PiPER robot.
- AgileX PiPER
- UCLA
- In development
Site maintainers
Intelligent Robotics and Interactive Systems Lab @ Arizona State University
Allegro Hand
RobotReplica Allegro Hand site
An ASU-hosted Allegro Hand evaluation site for reproducible in-hand manipulation and object reorientation benchmarks.
Benchmark tasksThe IRIS Lab team is developing an in-hand manipulation benchmark that evaluates how reliably policies reorient objects with the Allegro Hand.
- Allegro Hand
- In-hand reorientation
- In development
Site maintainers
GRASP Laboratory @ University of Pennsylvania
Dual YAM Robots
RobotReplica YAM bimanual site
A Penn-hosted pair of YAM robots for reproducible bimanual manipulation benchmarks.
Benchmark tasksThe Penn GRASP team is developing a benchmark that evaluates coordinated manipulation with a pair of YAM robots.
- Two YAM robots
- Bimanual manipulation
- In development
Site maintainers
POLICY REGISTRY
Discover featured
robot policies.
Explore a curated directory of public robot policies and foundation models. Inclusion is for discovery and does not imply RobotReplica evaluation or hardware compatibility.
- Xiaomi-Robotics-1
- Qwen-RobotManip
- GR00T N1.7
- SmolVLA
SUBMITTED-POLICY WORKFLOW
From your model
to a verified score.
Find your robot
Choose a site that operates the same robot embodiment as your model or policy.
Contact the site
Share your policy, interface requirements, and the benchmark track you want to enter.
We run the evaluation
The host executes your policy on its maintained setup under a standardized protocol.
Compare the result
Your verified score is added to the leaderboard for that site and robot.
Evaluation details—policy interface, checkpoints, task coverage, and reporting—are coordinated directly with the selected host site.
PARTNER PLAYBOOK
Build a new
RobotReplica site.
A partner site turns a robot and workspace into a maintained community benchmark. These five stages provide the starting framework; RobotReplica coordinates details with each host.
STAGE 1 OF 5
Choose the robot
Select the robot embodiment your organization can maintain and support for repeated community evaluations.
STAGE 2 OF 5
Build the workspace
Create a stable physical workspace with consistent lighting, a light box where appropriate, and fixed external and wrist cameras.
STAGE 3 OF 5
Design the tasks
Define a diverse benchmark suite—typically around 10 tasks—that reflects the robot’s capabilities and useful manipulation skills.
STAGE 4 OF 5
Specify the scenes
Create repeatable scene setups for every task, including object placement and in-distribution and out-of-distribution variations.
See VLA-Replica scene referencesSTAGE 5 OF 5
Validate with a policy
Run a baseline policy across the full suite to verify task feasibility, evaluation criteria, and the end-to-end protocol.
Bring a new robot into the network.
Tell us about your organization, robot platform, workspace, and the benchmark tasks you want to develop.

























