Humanoid robots have spent years mastering the art of looking futuristic while performing suspiciously modest tasks. Pick up a cup. Move a block. Fold something badly. Try not to topple over.
Google DeepMind now wants to push them beyond that awkward phase.
The company has introduced Gemini Robotics 2, a new family of artificial intelligence models designed to help robots reason, move, manipulate objects, and cooperate across longer, more complicated assignments. Its headline capability is whole-body control. Instead of commanding only a robot’s arms or upper body, the technology can coordinate movements from feet to fingertips.
That means a humanoid robot can walk toward an object, bend down, maintain its balance, pick up the item, and place it somewhere else—all as parts of one continuous task.
It sounds simple. For a human, it is. For a tall collection of motors, cameras, sensors, and occasionally rebellious joints, it is a computational juggling act.
DeepMind’s broader ambition is even bigger: create general-purpose physical AI that can operate safely and usefully in spaces built for people.
Three Models, Three Different Jobs
Gemini Robotics 2 is not a single, monolithic robot brain. DeepMind has organized the system around three specialized models.
The flagship Gemini Robotics 2 model is a vision-language-action model, commonly shortened to VLA. It connects what a robot sees and hears with the physical actions it needs to perform.
Gemini Robotics ER 2 handles embodied reasoning. Think of it as the planner standing behind the curtain. It interprets instructions, studies the environment, breaks complicated jobs into manageable steps, monitors progress, and decides what should happen next.
Finally, Gemini Robotics On-Device 2 runs locally on compatible robotic hardware. It targets situations in which a machine needs fast responses, encounters an unreliable internet connection, or cannot depend on cloud processing.
Together, the models form a layered system: one plans, another converts plans into movement, and an efficient version can work directly aboard the machine.
As SiliconANGLE explains, this resembles the dual-system architecture already used in humanoid robotics, but DeepMind has expanded the formula with stronger reasoning, full-body control, and local deployment.
From Tabletop Tricks to Full-Body Movement
Earlier Gemini Robotics models concentrated heavily on tabletop manipulation and upper-body actions. Gemini Robotics 2 stretches that capability across the entire humanoid form.
DeepMind demonstrated the model using Apptronik’s Apollo 2 robot. In one example, Apollo received an instruction to put a watering can inside a green bin on a lower shelf.
Completing that command required much more than identifying the right object. The robot needed to move through the room, approach the watering can, bend toward the floor, grasp it, regain or maintain its balance, travel to the shelf, crouch, and place the can in the correct container.
The important part was continuity. Instead of treating walking, reaching, and object manipulation as entirely separate routines, the model coordinated them as one whole-body sequence.
According to The Verge, DeepMind acknowledges that the robots still need to improve their movement speed. The demonstrations look deliberate rather than human-fast.
Still, speed is only one piece of the puzzle. A robot that moves slowly but completes the correct job is useful. A robot that sprints confidently toward the wrong shelf is a workplace incident wearing expensive shoes.
Balance Is Part of Intelligence Too
Whole-body robotics creates a problem that chatbots never face: gravity has opinions.
When a humanoid bends forward, reaches sideways, lifts an object, or lowers itself toward a shelf, its center of gravity changes. Each movement affects every movement that follows. A badly timed reach can turn a sophisticated robot into an extremely costly floor decoration.
Gemini Robotics 2 coordinates locomotion and manipulation while accounting for balance. Its motor-control outputs cover a humanoid’s complete body rather than focusing solely on hands and arms.
That matters because real environments rarely present objects at the perfect height and distance. Tools sit on floors. Packages hide under tables. Supplies occupy crowded shelves. A genuinely useful robot must adjust its posture, navigate obstacles, and manipulate objects without demanding that humans redesign the room around it.
DeepMind’s approach attempts to unite those actions under one VLA model. The system sees the scene, interprets the instruction, and produces appropriate physical commands.
This does not mean every humanoid suddenly gains flawless gymnastics. The model’s performance still depends on the robot’s sensors, mechanical design, actuators, hands, training data, and safety systems.
Software may supply the intelligence, but hardware still has to survive the choreography.
Five-Fingered Hands Enter the Picture
Walking and crouching attract attention, but useful robotics often comes down to the hands.
Traditional industrial robots usually rely on specialized grippers. Those tools work wonderfully in predictable environments, especially when repeatedly handling objects of similar size and shape. Homes and general workplaces, however, contain a chaotic collection of bags, cables, switches, tools, containers, and other objects apparently designed to annoy robots.
Gemini Robotics 2 supports both grippers and more complex, five-fingered robotic hands.
DeepMind showed the model working with the 22-degree-of-freedom SharpaWave tactile hand. Demonstrated tasks included sealing a resealable plastic bag, tying flexible bag handles, manipulating cordage, and unscrewing a lightbulb.
These jobs require precise contact, coordinated finger movement, and continuous correction. A flexible bag does not behave like a rigid factory component. It bends, slips, wrinkles, and generally refuses to cooperate.
Humanoids Daily reports that DeepMind also demonstrated Gemini Robotics 2 across different hardware configurations, including Apollo 2 and a dual-arm Franka system.
That cross-platform capability supports an important goal: creating AI that can adapt to multiple robotic bodies instead of remaining trapped inside one machine.
ER 2 Becomes the High-Level Brain

Physical movement is only half the challenge. Robots must also understand what people want.
Gemini Robotics ER 2 serves as the system’s high-level reasoning layer. It accepts natural-language instructions, examines incoming visual information, plans multi-stage assignments, and communicates commands to lower-level control models or robotic tools.
DeepMind says the model can manage tasks lasting several minutes and involving hundreds of intermediate decisions.
That is a meaningful shift from the familiar pattern of giving a robot one narrowly defined command at a time. A person can request a broader outcome, while ER 2 determines the sequence needed to achieve it.
The model can also call external tools. If developers permit it, ER 2 can access services such as Google Search or use custom functions to retrieve information required for a job.
This capability builds on DeepMind’s earlier effort to connect embodied reasoning with digital tools. A robot may need to look up disposal rules before sorting waste, check instructions for unfamiliar equipment, or retrieve information that is not visible in its immediate surroundings.
Of course, tool access raises another question: what happens when retrieved information is incomplete, irrelevant, or unsafe?
That is why reasoning, verification, and refusal mechanisms matter just as much as the ability to search.
Robots That Know How Far They Have Come
Long tasks create a surprisingly tricky problem: the robot needs to know what it has already completed.
Gemini Robotics ER 2 analyzes continuous video streams to track progress. Instead of examining only isolated images, it follows how a scene changes while the robot works.
If an action fails, the model can identify the most recent successfully completed step and attempt a correction. The robot does not necessarily need to restart the entire assignment.
Imagine a machine organizing a workspace. It picks up several objects, places two correctly, drops a third, and then encounters an obstructed drawer. A weak system might lose its place or repeat completed actions. A stronger one recognizes the interruption, resolves it, and resumes from the appropriate point.
ER 2 also improves “moment finding”—recognizing when an important event occurs. A robot needs to know when a container is full, when an object has reached the correct position, or when a fastening task is genuinely finished.
These details sound painfully mundane. They are also essential.
A robot that cannot tell when it has finished pouring will turn your coffee request into a tabletop irrigation project.
Progress tracking gives the system a better chance of deciding when to continue, retry, stop, or ask a human for help.
Multi-Robot Teamwork Arrives
DeepMind is also introducing coordination between different robots.
Gemini Robotics ER 2 can divide a larger workflow into smaller assignments and direct multiple machines to work together. One robot might locate or transport an object, while another handles a manipulation task better suited to its design.
In DeepMind’s demonstrations, Apollo 2 collaborated with a dual-arm robot during a garage-cleaning scenario. The machines divided responsibilities and moved tools into a bin.
This matters because no single robotic form works perfectly everywhere. Humanoids can navigate spaces and use equipment designed for people, but wheeled robots may move more efficiently across flat floors. Fixed robotic arms offer stability and precision. Smaller mobile manipulators can squeeze into areas where a full-sized humanoid would struggle.
Multi-robot collaboration lets each machine contribute its strongest capability.
However, coordinated robotics also multiplies the possible failure points. Machines must share accurate information, avoid obstructing one another, manage task handoffs, and remain aware of nearby people.
The demonstration represents a technical step forward, not proof that autonomous robot crews are ready to roam every warehouse tomorrow morning.
Still, the direction is clear. DeepMind is building systems that treat robots less like isolated appliances and more like members of a coordinated physical workforce.
On-Device 2 Cuts the Cloud Cord
Not every robot can pause while waiting for the internet.
Gemini Robotics On-Device 2 is optimized to run directly on a robot’s local computer. That reduces reliance on cloud connectivity and can cut network-related delays.
Local operation matters in factories, remote facilities, mobile systems, and other settings where connections may become slow or unavailable. It may also help organizations keep certain sensor data and operational processes on site.
DeepMind says On-Device 2 can adapt to substantially different robotic bodies with only a few hours of additional data collection. According to the company’s demonstrations, developers can tune it using fewer than 200 examples in some cases.
The model inherits motion-transfer techniques from earlier Gemini Robotics work. Rather than training a completely new system from scratch whenever the hardware changes, developers can transfer existing capabilities and refine them for a new arrangement of sensors, joints, or manipulators.
DeepMind demonstrated this adaptability on several third-party platforms, including Dexmate, SO101, and Trossen systems.
Hardware independence remains an aspiration rather than a solved problem. Every robotic body has unique limits. Still, faster adaptation could lower one of the field’s biggest barriers: the time and data required to teach each machine separately.
Safety Moves to Center Stage
Giving AI control over a physical machine changes the stakes. A chatbot error may produce a bad paragraph. A large robot error may produce a dented wall—or a very uncomfortable safety meeting.
DeepMind says Gemini Robotics ER 2 improves its ability to detect nearby humans, trigger safety tools, and stop a robot when someone moves too close. The system can resume after the area becomes safe.
The company also introduced ASIMOV-Agentic, a benchmark designed to test safety orchestration under uncertainty. It evaluates whether an embodied reasoning system can reject unsafe tool calls, recognize when a requested task may be impossible, and seek human assistance when confidence falls.
These protections complement traditional robotic safety measures such as collision detection, restricted operating zones, emergency stops, speed limits, and mechanical safeguards.
No benchmark can guarantee flawless real-world behavior. Controlled evaluations cannot reproduce every cluttered room, damaged sensor, unexpected human action, or bizarre object a deployed machine might encounter.
The real test will arrive through repeated operation outside carefully managed demonstrations.
DeepMind’s safety work nevertheless highlights the right problem. A capable robot must know more than how to act. It must also recognize when not to act.
Who Can Use the Models?
Availability differs across the three releases.
Developers can access Gemini Robotics ER 2 through the Gemini API and Google AI Studio. DeepMind is also offering it through a private preview on the Gemini Enterprise Agent Platform.
The flagship Gemini Robotics 2 VLA model and Gemini Robotics On-Device 2 are initially going to early-access hardware partners rather than receiving an unrestricted public release.
That staged rollout makes sense. Deploying an embodied reasoning model through an API is not the same as handing general motor-control software to anyone with a humanoid robot and an adventurous attitude.
Partners must integrate the models with specific sensors, actuators, control systems, and physical safeguards. Developers also need to define which low-level actions and external tools ER 2 may call.
This means consumers should not expect Gemini-powered humanoids to appear in every household immediately. The technology remains focused on research, development, and carefully managed partner deployments.
Businesses may eventually apply such systems to manufacturing, logistics, maintenance, laboratory work, and other multi-stage physical processes. Domestic robots remain an obvious long-term possibility, but the gap between a polished laboratory demonstration and dependable household operation is enormous.
Homes, after all, contain children, pets, stairs, clutter, and cables placed precisely where no sensible engineer would put them.
A Milestone, Not the Finish Line

Gemini Robotics 2 addresses several of the field’s hardest problems at once: whole-body coordination, dexterous manipulation, extended task planning, progress tracking, rapid hardware adaptation, and collaboration between machines.
That combination makes the release notable.
Yet none of the source reports proves that DeepMind has solved general-purpose robotics. The demonstrations remain demonstrations. The robots move cautiously, access remains limited, and real-world reliability will depend on far more than model intelligence.
Deployment will require robust hardware, affordable machines, careful integration, strong cybersecurity, maintainable components, practical energy use, and safety performance that holds up across thousands of unpredictable situations.
Even so, the trajectory has changed.
Humanoid robotics is moving beyond machines that perform isolated party tricks in controlled spaces. Developers increasingly want systems that understand goals, manage longer workflows, recover from errors, and use their entire bodies in human environments.
Gemini Robotics 2 offers a clearer picture of what that future could look like: not one robot doing everything, but adaptable machines combining specialized models—and sometimes cooperating with one another—to complete complicated physical work.
The robot uprising remains postponed. The robot group project, however, has officially begun.
Sources
- The Verge — “Google DeepMind’s new AI model can control a robot’s entire body”
- SiliconANGLE — “Google DeepMind debuts Gemini Robotics 2 model series for humanoid robots”
- Humanoids Daily — “Google DeepMind Unveils Gemini Robotics 2”
Kingy Launch Brief
Put the week’s verified AI launches in your inbox.
One source-checked edition every Friday, with a clear try, watch or skip verdict. After subscribing, check your inbox and confirm your address.
Free · Fridays · Double opt-in · Unsubscribe anytime
