The Open Robot Company

Blogs From X

What’s needed to scale emergent capability in physical AI

The two inputs for emergent physical intelligence, what to measure after training, and the 270k-hour bet on sample efficiency.

a few axes come into play when designing a model (an intelligent brain) with emergent capabilities (which performs tasks in environments not in the training dataset)

The 2 inputs

it’s simple:

  1. multi-sensory data (vision, touch, nerves, neurons, etc.)
  2. unique task-environment cells

these 2—scaled with orders of magnitude of data—can function as a very good pretraining-ready dataset.

I’m going to follow @GeneralistAI’s Gen-0 strategy of using just 270k hours.

What to measure

I won’t go into how the model should be trained; I have no expertise there.

But let me explain the metrics we can use once the model is trained to track intelligence—or what @GeneralistAI calls “physical commonsense”:

Behaviour Cloning

how autonomously and accurately the task is performed, compared to how a human would do it.

An additional metric to track is throughput against the standard operating procedure, comparing a human vs. the robot (speed, task accuracy, and reliability over the same long-horizon task are also very important).

Emergent Capability

this one is much more nuanced and is still being actively researched—

one possible way to track this would be: how many unique tasks can this omni or general brain do—tasks it has simply not seen in the dataset, in environments also not present in the dataset

The 270k-hour bet

I want to end by clearly stating: regardless of what these two metrics look like for current Gen-0 models and other physical-AI foundational model labs, I believe—and it goes without saying—that we’re going to see models trained on roughly the same amount of base pre-training data (270k hours) that absolutely crush today’s leading models on both metrics, due to simply one thing: Data Sample Efficiency.

Once that happens, physical commonsense should follow a more log-linear scaling law than the current power law and become much more data-sample-efficient from a capability perspective.

Where OpenRobot fits

OpenRobot plays a small part in helping new physical-AI teams by curating diverse, multimodal datasets—starting with stereo-RGBD (+ more senses incoming)—so we can effectively close the 0–1 stage for anyone training a new foundational model from scratch.

By the way, it takes only 270k hours to hit the two metrics I mentioned above.

As we’ve seen with the evolution of open models in LLMs, I think the physical-AI ecosystem will eventually converge on a similar imperative: institutions will collect and curate their own datasets so they can preserve physical processes in-house and remain data-sovereign.

This last thesis is also inspired by Jensen Huang’s framing at GTC: open foundations can be specialized with proprietary data, workflows and definitions of accuracy.

References

  1. @GeneralistAI — The Dark Matter of Robotics: Physical Commonsense https://generalistai.com/blog/physical-commonsense
  2. NVIDIA — The Future of AI Is Open and Proprietary (Jensen Huang at GTC) https://blogs.nvidia.com/blog/ai-future-open-and-proprietary/

If you’re training your first physical-AI foundational model, hmu for really good, sample-efficient stereo hardware!

Shipping to your nearest data-foundry :)

Originally published on X

Open on X