Google DeepMind released Gemini Robotics ER 2, an embodied reasoning model for robots, on July 30[1]. It serves as a high-level brain that plans multi-step tasks and hands motor execution off to a separate model. The two headline changes are the ability to track its own progress by watching continuous video, and the ability for different types of robots to divide work between them. Developers can access it through the Gemini API and Google AI Studio[1].

Built to Avoid the Stop-and-Think Pause

For robots to be useful in the places people live, accurate spatial reasoning alone is not enough. Decisions have to keep pace with the real-time speed of the physical world[1].

Gemini Robotics ER 2 is designed around that timing problem. It handles conversation with people, understanding of the surrounding environment, and multi-step planning, then hands the actual motor control to a lower-level vision-language-action (VLA) model[1]. It can also natively call tools such as Google Search, or any function a developer defines[1].

To cut latency, ER 2 connects through the Gemini Live API using a bidirectional streaming endpoint[1]. Developers declare low-level control interfaces — VLA models or navigation APIs — as tools, and stream video, audio, or text straight into the model[1]. The result removes the jarring stop-and-think pauses that appear whenever a robot has to work out its next step[1]. Google DeepMind says tool orchestration accuracy beats the previous ER 1.6 across all three control modes tested: real VLA, simulated VLA, and human tele-operation[1].

As a demonstration, the company published a setup in which ER 2 orchestrates navigation and manipulator APIs on Boston Dynamics' Spot, producing a robot that fetches objects on a spoken request. The code is on GitHub[1].

Measuring "Is It Done Yet" From Video

One of the hardest problems in robotics is knowing when a task is actually finished[1]. ER 2 introduces two measures for this.

The first is progress classification. Each frame of the video feed is assigned to one of five bands — 0-20 percent, 20-40 percent, and so on — turning progress into a number[1]. When something fails midway, the robot can repeat just that step instead of restarting the whole workflow[1]. ER 2 reaches 57.4 percent accuracy on this task, ahead of the previous generation and competing frontier models[1].

The second is moment finding: pinpointing the exact frame where a critical event occurs, such as the instant coffee stops being poured into a cup[1]. ER 2 scores 91.3 percent accuracy with a mean absolute distance of 0.96 seconds[1]. Google DeepMind describes this as competing with much larger model categories at a fraction of the compute cost and 4x the execution speed[1]. Together, the two measures let a robot confirm that a job such as tightening a light bulb or tying a trash bag is complete to specification before moving on[1].

Different Robots Splitting the Work

No single robot suits every task: a wheeled rover excels indoors, while a humanoid handles uneven terrain[1]. ER 2 adds multi-robot collaboration, letting diverse machines communicate through a shared semantic understanding, hand off to each other, and finish work together[1]. Google DeepMind shows Apptronik's Apollo 2 humanoid working alongside Franka's F3 Duo[1].

Core spatial understanding was updated as well. Success and failure detection now runs on raw video rather than static snapshots, catching mid-execution problems like spills, slips, and misalignments[1]. Instrument reading extends beyond circular dials and sight glasses to digital displays, linear scales, rulers, and liquid thermometers, tested across 10 types of instruments[1].

Stopping When a Person Comes Close

On safety, Google DeepMind reports its best results yet on two benchmarks: Safety Instruction Following, which measures adherence to physical constraints, and Human Proximity, which measures awareness of nearby people[1]. The model halts a humanoid robot when someone is close by and autonomously resumes work only once the area is clear[1].

A new benchmark evaluates whether a foundation model can act as a safe VLA orchestrator, scoring it on enforcing safety constraints, monitoring the environment, assessing physical feasibility, and asking a human for clarification when uncertain[1]. In the Google DeepMind announcement, this framework is published as ASIMOV-Agentic, which also measures whether the agent can refuse unsafe tool calls coming from a VLA[2].

One of Three Models

ER 2 is not a standalone release. It is one part of Gemini Robotics 2, a family of three models[2]. The other two are Gemini Robotics 2, a VLA model that controls humanoid robots from feet to fingertips, and Gemini Robotics On-Device 2, a lightweight version that runs locally on the robot[2]. On-Device 2 can adapt to new bi-arm robots with different shapes, sensors, and degrees of freedom in a few hours, typically with fewer than 200 examples[2].

Availability differs across the three. ER 2 is out on the Gemini API and Google AI Studio, and in private preview on the Gemini Enterprise Agent Platform[1][2]. The VLA and On-Device models remain limited to early-access partners[2]. The closer a model gets to moving actual limbs, the more restricted it is — so for now, outside developers work with the planning and supervision layer that ER 2 provides.

Summary

Gemini Robotics ER 2 is the reasoning model responsible for planning and supervising a robot's work. It measures progress from video, pinpoints critical moments to within 0.96 seconds on average, and now coordinates handoffs between different types of robots. A low-latency connection through the Gemini Live API removes the pauses that used to appear whenever the robot stopped to think. It is available to developers through the Gemini API and Google AI Studio, while the VLA models that actually move the limbs remain at the early-access stage.

Source[1]: https://blog.google/innovation-and-ai/models-and-research/google-deepmind/gemini-robotics-er-2/

Source[2]: https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/