Embodied AI is the manifestation of Artificial Intelligence within a physical entity—the "body" to the AI's "mind." It is the specific subset of Physical Intelligence where the AI model is inextricably linked to a physical form factor (robot arm, humanoid, mobile base) and its specific sensor-actuator configuration. Embodied AI learns through interaction with the environment—"learning by doing"—often utilizing techniques like reinforcement learning to master fine motor skills and complex manipulation tasks that are difficult to hard-code.
While Physical Intelligence describes the capability (understanding the world), Embodied AI describes the implementation (the machine itself). Research distinguishes between the digital intent (the "brain" or Agentic AI) and the physical execution (the "muscles" or Embodied AI).
Embodied AI represents the transition of Foundation Models into Vision-Language-Action (VLA) models. An LLM outputs text; a VLA outputs motor torques and robotic trajectories. This shift is transformative because it moves automation from "structured" environments (cages, jigs, precise lighting) to "unstructured" environments (brownfield factories, warehouses with human traffic). Embodied AI allows robots to work alongside humans, adapting to the messiness of the real world rather than requiring the world to be organized around the robot.