ArmnetBench v0.1: Parallel Real-World Evaluation of Manipulation Policies on a Low-Cost Arm Farm
This benchmark addresses the bottleneck of real-world evaluation for generalist robot manipulation policies by providing a scalable, low-cost platform for parallel testing, though it is an incremental contribution as it primarily validates the infrastructure and provides initial comparisons.
ArmnetBench v0.1 introduces a low-cost robot arm farm for parallel real-world evaluation of manipulation policies, comparing 7 policies across 12 tasks with 2,518 policy rollouts and 600 demonstrations, all human-scored for success, suboptimal, or failure. The benchmark provides quality-labeled trajectories for downstream learning and releases 3,118 episodes in standard formats.
Real-world evaluation is a bottleneck in developing generalist robot manipulation policies. Each rollout requires physical hardware and an operator to set up, reset, and score it. We introduce ArmnetBench v0.1, a benchmark run on a fleet of low-cost SO-101 cells under light on-site supervision. v0.1 validates this arm farm end to end and compares 7 policies across 12 tasks with both single-arm and bimanual configurations. Each policy is trained or fine-tuned on 50 demonstrations per task; the benchmark contains 2,518 policy rollouts and 600 reference demonstrations. All 3,118 episodes carry a three-way label (successful, suboptimal, or failure). Policy rollouts are human-scored, while demonstrations are successful by construction. Beyond evaluation, its quality-labelled trajectories support downstream learning, from reward and predictive world models to policies trained on mixed-quality data. The leaderboard is an initial comparison under this shared budget. We release the 3,118 core episodes in LeRobot v3.0 and RoboMeter formats.