A PHYSICAL BENCHMARK NETWORK FOR ROBOT MANIPULATION

Your policy.
Our robot.
One comparable result.

RobotReplica is a non-commercial, community-led research initiative connecting organizations that voluntarily host standardized robot manipulation benchmarks.

CORE ACTIVITIES

What RobotReplica does

01
FOR POLICY DEVELOPERSEvaluate submitted policies.

Researchers submit a policy to a site with matching robot hardware. The volunteer site team runs it on maintained real-world tasks.

02
FOR THE ROBOTICS COMMUNITYBenchmark public policies.

RobotReplica proactively evaluates compatible public policies and checkpoints, verifies the results, and maintains robot-specific leaderboards.

RobotReplica network connecting physical robot evaluation sites
A DISTRIBUTED PHYSICAL BENCHMARK NETWORKMULTIPLE ROBOTS

PARTNER ORGANIZATIONS & PEOPLE

OVERVIEW

A shared evaluation service
for real robots.

Robotics results are difficult to compare when every lab uses a different robot, scene, and protocol. RobotReplica keeps physical benchmark sites running so the community can evaluate on consistent hardware and tasks.

01

Physical benchmark sites

Each partner maintains a robot, workspace, task objects, cameras, and an evaluation protocol.

02

Robot-matched evaluation

Partner sites evaluate submitted policies and compatible publicly released policies on matching robot hardware.

03

Verified leaderboards

RobotReplica maintains a board for each benchmark site using verified results from both evaluation paths.

THE ROBOTREPLICA MODEL

Researchers do not need to rebuild the full benchmark. The benchmark stays at the host site; policies travel to it.

CURRENT SITES

Start with the robot
you already use.

Seven organizations are building the first RobotReplica sites. Each site operates its hardware, tasks, and evaluation process; RobotReplica maintains the verified robot-specific leaderboards.

VLA-Replica SO-101 evaluation setup showing the light box, top camera, camera mount, follower arm, and wrist webcamSITE 01
Accepting evaluationsRichardson, Texas

Intelligent Robotics and Vision Lab @ UT Dallas

SO-101

RobotReplica SO-101 site

A hosted SO-101 evaluation site for reproducible VLA-Replica manipulation benchmarks.

Benchmark tasksVLA-Replica evaluates 10 tabletop manipulation tasks in both in-distribution and out-of-distribution scenes.

Pick & placeBread on plate · Bowl on coaster · Lift bowlDexterousStack or collect blocks · Fold towel · Pour pepperInteractionOpen oven · Clean whiteboard · Press button
  • SO-101 arm
  • UT Dallas
  • Accepting evaluations

LIVE LEADERBOARD

SO-101 · VLA-Replica

Average policy success rates across the official benchmark. Ranked by the in-distribution result.

Rank / policyIDOOD
01π₀.₅ID54%OOD35%
02π₀ID34%OOD30%
03SmolVLAID26%OOD30%
Top 3 · success rate · 5 runs per taskView full leaderboard
In developmentSan Francisco, California

General Intelligence Labs

OpenArm

RobotReplica OpenArm

A hosted OpenArm evaluation site extending the network to bimanual, contact-rich manipulation.

Benchmark tasksInitial benchmark tasks are being designed for the standardized OpenArm Cell; the final protocol is still in development.

BimanualCoordinated grasping · Object handoffs · Two-arm placementContact-richInsertion and assembly · Tool use · Articulated objectsRobustnessObject variations · Scene distractors · Tasks at varied heights
  • Bimanual platform
  • Open-source hardware
  • Protocol in development
In developmentIthaca, New York

Generalizable Robot Intelligence and Learning Lab @ Cornell University

Tianji Marvin Arms

RobotReplica Marvin site

A hosted Tianji Marvin dual-arm evaluation site equipped with two WUJI hands for reproducible, force-aware manipulation benchmarks.

Benchmark tasksThe team is developing a benchmark around the force-controlled, humanlike Marvin arm platform.

RobotDual 7-DoF humanlike arms · Two WUJI handsControlFull-force control · Adaptive impedance · Compliant interactionBenchmarkTask suite and evaluation protocol in development
  • Tianji Marvin arms
  • Two WUJI hands
  • In development
In developmentCambridge, Massachusetts

Computational Design and Fabrication Group @ MIT + Perceptual Science Group @ MIT

Franka Panda

RobotReplica Franka Panda site

A hosted Franka Panda evaluation site extending the network with a new reproducible manipulation benchmark.

Benchmark tasksThe team is developing its robot workspace, benchmark tasks, scene specifications, and baseline evaluation protocol.

RobotFranka Panda 7-DoF research armWorkspaceControlled lighting · Fixed camera views · Repeatable scenesBenchmarkTask suite and evaluation protocol in development
  • Franka Panda arm
  • MIT CSAIL
  • In development
In developmentLos Angeles, California

PRACTICE Lab @ UCLA

AgileX PiPER

RobotReplica PiPER site

A UCLA-hosted PiPER evaluation site extending the network with a reproducible manipulation benchmark for the compact robot arm.

Benchmark tasksThe PRACTICE Lab team is developing the workspace, task suite, scene specifications, and evaluation protocol for its PiPER robot.

RobotAgileX PiPER robot armWorkspaceStandardized cameras · Repeatable scenes · Controlled setupBenchmarkTask suite and evaluation protocol in development
  • AgileX PiPER
  • UCLA
  • In development
In developmentTempe, Arizona

Intelligent Robotics and Interactive Systems Lab @ Arizona State University

Allegro Hand

RobotReplica Allegro Hand site

An ASU-hosted Allegro Hand evaluation site for reproducible in-hand manipulation and object reorientation benchmarks.

Benchmark tasksThe IRIS Lab team is developing an in-hand manipulation benchmark that evaluates how reliably policies reorient objects with the Allegro Hand.

RobotAllegro four-finger dexterous handCore skillIn-hand object reorientationBenchmarkObject set, target orientations, and evaluation protocol in development
  • Allegro Hand
  • In-hand reorientation
  • In development
In developmentPhiladelphia, Pennsylvania

GRASP Laboratory @ University of Pennsylvania

Dual YAM Robots

RobotReplica YAM bimanual site

A Penn-hosted pair of YAM robots for reproducible bimanual manipulation benchmarks.

Benchmark tasksThe Penn GRASP team is developing a benchmark that evaluates coordinated manipulation with a pair of YAM robots.

RobotTwo YAM robot armsCore skillCoordinated bimanual manipulationBenchmarkTask suite, scenes, and evaluation protocol in development
  • Two YAM robots
  • Bimanual manipulation
  • In development

POLICY REGISTRY

Discover featured
robot policies.

Explore a curated directory of public robot policies and foundation models. Inclusion is for discovery and does not imply RobotReplica evaluation or hardware compatibility.

  • Xiaomi-Robotics-1
  • Qwen-RobotManip
  • GR00T N1.7
  • SmolVLA
Explore featured policies

SUBMITTED-POLICY WORKFLOW

From your model
to a verified score.

1

Find your robot

Choose a site that operates the same robot embodiment as your model or policy.

2

Contact the site

Share your policy, interface requirements, and the benchmark track you want to enter.

3

We run the evaluation

The host executes your policy on its maintained setup under a standardized protocol.

4

Compare the result

Your verified score is added to the leaderboard for that site and robot.

Evaluation details—policy interface, checkpoints, task coverage, and reporting—are coordinated directly with the selected host site.

PARTNER PLAYBOOK

Build a new
RobotReplica site.

A partner site turns a robot and workspace into a maintained community benchmark. These five stages provide the starting framework; RobotReplica coordinates details with each host.

01

STAGE 1 OF 5

Choose the robot

Select the robot embodiment your organization can maintain and support for repeated community evaluations.

02

STAGE 2 OF 5

Build the workspace

Create a stable physical workspace with consistent lighting, a light box where appropriate, and fixed external and wrist cameras.

03

STAGE 3 OF 5

Design the tasks

Define a diverse benchmark suite—typically around 10 tasks—that reflects the robot’s capabilities and useful manipulation skills.

04

STAGE 4 OF 5

Specify the scenes

Create repeatable scene setups for every task, including object placement and in-distribution and out-of-distribution variations.

See VLA-Replica scene references
05

STAGE 5 OF 5

Validate with a policy

Run a baseline policy across the full suite to verify task feasibility, evaluation criteria, and the end-to-end protocol.

READY TO HOST?

Bring a new robot into the network.

Tell us about your organization, robot platform, workspace, and the benchmark tasks you want to develop.

Propose a partner site