Frontier robot policies score ~97% on academic LIBERO — and 0/68 on RoboGate's industrial Pick&Place suite (NVIDIA Isaac Sim, Franka Panda). This is the public, truthful leaderboard of that gap.
→ Live leaderboard & charts: robogate.io/vla
15 real-inference policies — VLA, diffusion, and a world-model — all score 0/68. The collapse spans architectures: autoregressive VLAs and NVIDIA's Cosmos Policy 2B diffusion world-model fail identically. A scripted analytic-IK baseline scores 83.8% (57/68) on the same scenarios (100% on nominal), proving the task is solvable in-harness.
See LEADERBOARD.md for the full table, or leaderboard.json for machine-readable data.
| Policies evaluated | 15 (all 0/68) |
| Scripted baseline | 83.8% (57/68) |
| Suite | 68 industrial scenarios on NVIDIA Isaac Sim |
| Paper | robogate.io/paper |
Think your policy survives? Open a submission issue →
Give us the HuggingFace model ID and how to run inference. We run it on the real 68-scenario suite on our GPU host and publish the result here — whatever it is. See CONTRIBUTING.md.
Academic benchmarks overstate real-world capability (LIBERO 90%+ collapses under perturbation; NVIDIA's own research calls the missing standardized sim-to-real benchmark "a critical bottleneck"). RoboGate is a curated industrial-realism gauntlet with truthful, reproducible reporting — complementary to NVIDIA + Hugging Face's Isaac Lab-Arena / LeRobot EnvHub, not a competitor to their eval infrastructure.
This repo is the public leaderboard + submission surface. The full validation engine (CLI, runtime drift monitor, report generator) lives in the main RoboGate product. Want pre-deployment validation or runtime monitoring for your robot cell? → robogate.io
Maintained by AgentAI Co., Ltd. · NVIDIA Inception member · Data: CC BY 4.0