Analysis
Google DeepMind's Gemini Robotics 2 extends the company's robotics AI model to full-body control of a humanoid robot -- walking, crouching, and manipulating objects -- where its predecessor handled only upper-body manipulation tasks. The model is video-native, meaning it learns to control a robot's movements from video input rather than requiring specialized robot-specific sensor data.
That's a meaningful technical jump. Most humanoid-robot AI to date has stitched together separate systems for locomotion and manipulation; a single model handling both end-to-end is closer to how the large language model wave collapsed dozens of narrow NLP tasks into one general-purpose architecture. It puts Google DeepMind in more direct competition with Tesla's Optimus program, Figure, and China's Unitree, all racing to prove a single foundation model can generalize across an entire robot's body.
For robotics investors, the stakes are about which layer captures value: if a small number of foundation models end up powering most humanoid robots the way a small number of LLMs now power most AI applications, the hardware companies building on top of those models may end up with less differentiated economics than the model layer itself. What to watch: whether Gemini Robotics 2 gets licensed to third-party humanoid-robot hardware makers, and how it compares head-to-head against Tesla's and Figure's own foundation models on real-world tasks.