Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
| Source: Google DeepMind Blog
Tags: Gemini Robotics, Google DeepMind, embodied AI, robotics, Gemini API, physical AI
Google DeepMind released Gemini Robotics ER 2, an embodied reasoning model now available via the Gemini API that acts as a high-level robot brain — enabling real-time video understanding, multi-step task planning, and multi-robot collaboration.
Details
Google DeepMind's Gemini Robotics ER 2 is the company's most capable embodied reasoning model for robotics, available today via the Gemini API and Google AI Studio, with Gemini Enterprise Agent Platform access in private preview. The model functions as a high-level planning layer: it handles human communication, physical world understanding, and multi-step task planning, then hands off motor execution to lower-level VLA models. Three capability upgrades over Gemini Robotics ER 1.6 stand out. First, continuous video feed processing lets robots track their own progress in real time, adapt when something goes wrong, and know precisely when to transition to the next step. Second, multi-robot collaboration is introduced — multiple robots can divide work in shared spaces and complete workflows that would exceed a single robot's capability. Third, the model can natively call external tools like Google Search or any user-defined function mid-task. The architecture overlaps reasoning and execution so the robot can 'think ahead' while still performing current actions — a meaningful latency improvement over sequential approaches. Google DeepMind is providing developer examples to support teams getting started with the API. This is an official release from Google DeepMind directly, not a third-party report.