Google DeepMind has introduced Gemini Robotics 2, a VLA model designed for full-body control of robots, along with the ER 2 solution that facilitates complex planning and coordination among multiple devices.

https://www.youtube.com/watch?v=4lSQnrMC6nY

The Gemini Robotics 2 model transforms text and visual information into motor actions. It is engineered for a variety of robotic form factors, ranging from desktop manipulators to full-sized humanoids.

Unlike its predecessor, this AI now manages not just the upper body. As a result, robots have learned to walk, squat, bend, stretch, and lift objects even in confined spaces.

In demonstrations, the humanoid robot Apptronik Apollo 2, powered by Gemini Robotics 2, was seen picking up a watering can, removing a baseball glove from a shelf, and locating specific items.

Apollo 2's dexterity in action. Source: Google DeepMind.

DeepMind also announced enhancements in motor skills: the model now supports intricate hands with 22 degrees of freedom, enabling tasks such as tying knots, sealing packages, or twisting light bulbs.

The ER 2 model has also been updated, which DeepMind describes as the "high-level brain" for robots. This solution is capable of planning multi-step tasks, monitoring whether actions are completed or need repetition, and distributing tasks among several robots.

In one video, Apollo 2 instructs a dual-arm robot from Google to organize tools into a container while cleaning a garage.

The company also showcased the local model Gemini Robotics On-Device 2, which operates independently of internet connectivity and adapts to new robots with less than 200 examples in just a few hours of training.

Currently, Gemini Robotics 2 is available in private preview, while ER 2 can be accessed via the Gemini API and Google AI Studio. Trustworthy partners are testing On-Device 2.

It is worth noting that in July, Humanoid introduced an approach to training robots on actual production tasks.