a few axes come into play when designing a model (an intelligent brain) with emergent capabilities (which performs tasks in environments not in the training dataset)
The 2 inputs
it’s simple:
- multi-sensory data (vision, touch, nerves, neurons, etc.)
- unique task-environment cells
these 2—scaled with orders of magnitude of data—can function as a very good pretraining-ready dataset.
I’m going to follow @GeneralistAI’s Gen-0 strategy of using just 270k hours.
What to measure
I won’t go into how the model should be trained; I have no expertise there.
But let me explain the metrics we can use once the model is trained to track intelligence—or what @GeneralistAI calls “physical commonsense”:
Behaviour Cloning
how autonomously and accurately the task is performed, compared to how a human would do it.
An additional metric to track is throughput against the standard operating procedure, comparing a human vs. the robot (speed, task accuracy, and reliability over the same long-horizon task are also very important).
Emergent Capability
this one is much more nuanced and is still being actively researched—
one possible way to track this would be: how many unique tasks can this omni or general brain do—tasks it has simply not seen in the dataset, in environments also not present in the dataset
The 270k-hour bet
I want to end by clearly stating: regardless of what these two metrics look like for current Gen-0 models and other physical-AI foundational model labs, I believe—and it goes without saying—that we’re going to see models trained on roughly the same amount of base pre-training data (270k hours) that absolutely crush today’s leading models on both metrics, due to simply one thing: Data Sample Efficiency.
Once that happens, physical commonsense should follow a more log-linear scaling law than the current power law and become much more data-sample-efficient from a capability perspective.
Where OpenRobot fits
OpenRobot plays a small part in helping new physical-AI teams by curating diverse, multimodal datasets—starting with stereo-RGBD (+ more senses incoming)—so we can effectively close the 0–1 stage for anyone training a new foundational model from scratch.
By the way, it takes only 270k hours to hit the two metrics I mentioned above.
As we’ve seen with the evolution of open models in LLMs, I think the physical-AI ecosystem will eventually converge on a similar imperative: institutions will collect and curate their own datasets so they can preserve physical processes in-house and remain data-sovereign.
This last thesis is also inspired by Jensen Huang’s framing at GTC: open foundations can be specialized with proprietary data, workflows and definitions of accuracy.
References
- @GeneralistAI — The Dark Matter of Robotics: Physical Commonsense https://generalistai.com/blog/physical-commonsense
- NVIDIA — The Future of AI Is Open and Proprietary (Jensen Huang at GTC) https://blogs.nvidia.com/blog/ai-future-open-and-proprietary/
If you’re training your first physical-AI foundational model, hmu for really good, sample-efficient stereo hardware!
Shipping to your nearest data-foundry :)