Gemini Robotics ER 2: Video Understanding, Task Orchestration and Multi-Robot Collaboration Make a High-Level Brain for Robots
Google DeepMind has launched Gemini Robotics ER 2, which it calls its most capable embodied reasoning model. It is built to act as the high-level brain of a robot: it talks with people, understands the physical world, plans multi-step tasks, and hands motor execution to a lower-level vision-language-action (VLA) model. By watching continuous video, a robot can track its own progress, recover from mistakes and know when to move to the next step. The release also adds multi-robot collaboration. Developers can use the model now through the Gemini API and Google AI Studio, with a private preview on the Gemini Enterprise Agent Platform.
Google DeepMind has launched Gemini Robotics ER 2, which it describes as its most capable embodied reasoning model for robotics. The announcement post is credited to Steven Hansen and Peng Xu and carries a date of July 30, 2026. Its position is clear. It is not a model that drives motors directly. It is meant to be the high-level brain of a robot. What was announced According to Google, Gemini Robotics ER 2 lets a robot chat with people, understand the physical world and plan multi-step tasks. It then hands motor execution to any given lower-level vision-language-action (VLA) model. The model can also natively call tools such as Google Search to look up information, as well as any user-defined function. Google calls it a significant upgrade over Gemini Robotics ER 1.6. Three themes stand out: continuous video understanding, tool and task orchestration, and, for the first time, multi-robot collaboration.
On availability, the model is publicly available to developers through the Gemini API and Google AI Studio, and in private preview on the Gemini Enterprise Agent Platform. Google also shares examples of how to configure the model and prompt it to power more useful physical AI tasks. How it works: reasoning on top, control below Google argues that accurate spatial reasoning is not enough for robots that must help people in everyday settings. A robot must also think fast, timing its decisions to the real-time pace of the physical world. For that reason, ER 2 is designed so that a robot can think about what comes next while it is still performing its current action.
For developers, the recipe is an agentic setup. You declare low-level control interfaces, such as VLA models or navigation APIs, as tools. You then stream multimodal video, audio or text directly into the model. ER 2 decides which tool to call, with which arguments, and what to do next based on what it sees. This mirrors the function-calling pattern that software agents already use. The difference is that the tool is now a physical system that moves and grasps. Google also states that ER 2 integrates into the Gemini Live API. The source excerpt we reviewed is cut off at that sentence, so readers should consult the full documentation for details. The layering has a practical benefit. The reasoning model and the control policy can improve independently. A robot maker can gain better planning without retraining its motion model, and can swap the low-level controller without rewriting the high-level logic. The video feedback loop A classic plan-then-act pipeline often assumes that a command succeeded once it was sent. ER 2 instead watches continuous video feeds. With them, a robot can track its own progress, adapt if something goes wrong, and know exactly when to move on to the next step. This matters in the physical world, where an object can slip, a door can stay ajar, or a person can move the target. Only a high-level model that sees the outcome has a chance to notice such deviations and correct them.
How orchestration was evaluated Google says the model can be evaluated with robots in simulation, with real-world robot control, and even paired with a human controlling the robot remotely. Its stated result is that ER 2 consistently outperforms ER 1.6 for tool orchestration across three control modes: real VLA, simulated VLA and human tele-operation. To be accurate about what we know, the excerpt we reviewed does not include numeric benchmark scores, latency figures or pricing. We therefore cannot quantify the size of the gain. Multi-robot collaboration ER 2 introduces multi-robot collaboration. Robots can work together in shared spaces and complete complex workflows that a single robot could not do alone. In principle, one high-level model can split a job, assign roles and coordinate several robots working in sequence or in parallel. For warehouses, factories and laboratory automation, that is a step from single-robot demos toward coordinated work cells. The excerpt does not describe the coordination mechanism or any limit on fleet size, so developers will have to test those points themselves.
What it means for developers and enterprises For developers, the barrier drops. Through the Gemini API they can reach a planner that understands video and calls tools, without building a full embodied reasoning stack. For enterprises, the private preview offers a controlled way to assess safety, compliance and integration cost. For the wider ecosystem, decoupling high-level reasoning from low-level control encourages a clearer division of labor. Model vendors supply a general brain. Robot makers and VLA researchers focus on motion and manipulation. Challenges and outlook Several open questions remain. Latency comes first: a high-level model reached through a cloud API must prove it can serve tasks that need very fast reactions. Reliability and safety come next: if the video judgment is wrong, a robot may keep acting on a false belief, so hardware safety limits and human oversight stay essential. Cost and scale also matter, because multi-robot setups multiply inference calls, and pricing and quotas will shape commercial viability. Finally, the field needs open, reproducible benchmarks to compare embodied reasoning models.
Overall, Gemini Robotics ER 2 packages planning, supervision and collaboration for robots as a service that developers can call. Whether it delivers in messy real-world sites will depend on independent evaluations and on feedback from early customers.