For AI to be useful and helpful to people in the physical realm, it has to demonstrate “embodied” reasoning—the humanlike ability to comprehend and react to the world— as well as safely take action to get things done. To that end, Google DeepMind introduced two new AI models, based on the Gemini 2.0 technology, which lay the foundation for a new generation of helpful robots.
Gemini Robotics, an advanced vision-language-action (VLA) model, was built on Gemini 2.0 with the addition of physical actions as a new output modality for the purpose of directly controlling robots. Gemini Robotics-ER, a Gemini model with advanced spatial understanding, enables roboticists to run their own programs using Gemini’s embodied reasoning (ER) abilities. Both models enable a variety of robots to perform a wide range of real-world tasks. As part of the effort, Google DeepMind is partnering with Apptronik to build the next generation of humanoid robots with Gemini 2.0 and working with a selected number of trusted testers, including Agile Robots, Agility Robots, Boston Dynamics, and Enchanted Tools, to guide the future of Gemini Robotics-ER.
Gemini Robotics Vision-Language-Action Model
To be useful and helpful to people, Google DeepMind believes that AI models for robotics need three principal qualities: they have to be general, meaning they’re able to adapt to different situations; they have to be interactive, meaning they can understand and respond quickly to instructions or changes in their environment; and they have to be dexterous, meaning they can do the kinds of things people generally can do with their hands and fingers, like carefully manipulate objects.
Generality
Gemini Robotics utilizes Gemini's deep world understanding to adapt to new situations and handle a wide range of tasks without prior training. It is also recognized for its ability to interact with unfamiliar objects, follow diverse instructions, and navigate new environments effectively.
Interactivity
To operate in the dynamic, physical world, robots must be able to seamlessly interact with people and their surrounding environment and adapt to changes on the fly. Gemini Robotics taps into Gemini’s advanced language understanding capabilities and can understand and respond to commands phrased in everyday, conversational language and in different languages.
It can interpret and respond to a significantly wider range of natural language instructions compared to the company's previous models, adjusting its behavior based on diverse inputs. Additionally, it continuously observes its surroundings, detects environmental or instruction changes, and modifies its actions accordingly.
Dexterity
Many everyday tasks that humans perform effortlessly require surprisingly fine motor skills and are still too difficult for robots. Gemini Robotics can tackle extremely complex, multi-step tasks that require precise manipulation. If an object slips from its grasp, or someone moves an item around, Gemini Robotics quickly replans and carries on—a crucial ability for robots in the real world, where surprises are the norm.
Multiple Embodiments
Gemini Robotics was designed to easily adapt to different robot types. The model was trained primarily on data from the bi-arm robotic platform, ALOHA 2, but the company has also demonstrated that it could control a bi-arm platform. Gemini Robotics can be specialized for more complex embodiments, such as the humanoid Apollo robot developed by Apptronik, with the goal of completing real world tasks.
Enhancing Gemini’s World Understanding
The company also introduced an advanced vision-language model called Gemini Robotics-ER (short for ‘“embodied reasoning”). This model is intended to enhance Gemini’s understanding of the world in ways necessary for robotics, focusing especially on spatial reasoning, and allows roboticists to connect it with their existing low-level controllers.
Gemini Robotics-ER improves Gemini 2.0’s existing abilities like pointing and 3D detection by a large margin. Combining spatial reasoning and Gemini’s coding abilities, Gemini Robotics-ER can instantiate entirely new capabilities on the fly. The company claims Gemini Robotics-ER can perform all the steps necessary to control a robot right out of the box, including perception, state estimation, spatial understanding, planning and code generation. In this end-to-end setup, the model achieves a success rate two to three times higher than Gemini 2.0. When code generation alone is insufficient, Gemini Robotics-ER leverages in-context learning, identifying patterns from a few human demonstrations to generate a solution.
Responsibly Advancing AI and Robotics
The physical safety of robots and the people around them is a longstanding, foundational concern in the science of robotics. That's why roboticists have classic safety measures such as avoiding collisions, limiting the magnitude of contact forces, and ensuring the dynamic stability of mobile robots. Gemini Robotics-ER can be interfaced with these low-level safety-critical controllers, specific to each embodiment. Expanding on Gemini’s core safety features, Gemini Robotics-ER models can assess whether a given action is safe within a specific context and generate appropriate responses accordingly.
To advance robotics safety research in both academia and industry, the company is releasing a new dataset aimed at evaluating and improving semantic safety in embodied AI and robotics. In previous work, they demonstrated how a Robot Constitution, inspired by Isaac Asimov’s Three Laws of Robotics, could help prompt a large language model to choose safer tasks for robots. Building on this, they have developed a framework that automatically generates data-driven constitutions—rules expressed directly in natural language—to guide a robot’s behavior. This framework enables users to create, modify, and apply constitutions, fostering the development of robots that are safer and better aligned with human values. Lastly, the new ASIMOV dataset is designed to help researchers systematically assess the safety implications of robotic actions in real-world scenarios.
Learn more about Industrial AI's Role in the Digital Transformation of Industries.