Google has launched Gemini Robotics 2, an update to its artificial intelligence model specifically designed for robotics. This technology aims to enable machines to understand their surroundings, plan sequences of actions, and perform more intricate movements, reducing the need for strictly pre-programmed commands.
Integration of technologies for complex tasks
The system integrates computer vision, natural language processing, and motor control. This combination allows robots to handle a wider range of activities, such as full body locomotion, item handling, and interaction with other equipment.
Overcoming the limitation of fixed commands
Currently, many robots rely on meticulous programming or human supervision, which restricts their effectiveness in dynamic environments where there are constant changes, misaligned objects, or unforeseen events. Gemini Robotics 2 seeks to mitigate this problem by allowing the robot to analyze an instruction, assess the scenario, and determine the best series of movements to complete the task.
During a demonstration, the model commanded the Apollo 2 humanoid robot, manufactured by Apptronik, to locate a watering can, travel to a shelf, and place it in a designated container. Google itself acknowledged that obstacles persist, such as the need to accelerate movement speed.
The company stated: 'Now, our model can control complete humanoid robots for the first time, translating commands into intelligent whole-body control.'
AI Model Structure
The new generation of models consists of three distinct versions, each with a specific function: Gemini Robotics 2 converts visual data and verbal commands into physical actions; Gemini Robotics ER 2 focuses on planning, understanding the environment, and coordinating multi-phase tasks; and Gemini Robotics On-Device 2 operates directly on the hardware, minimizing dependence on internet connection.
According to Google, the model executed locally on the robot can quickly adapt to new platforms, requiring fewer than 200 training examples, even when there are variations in sensors, shape, or movement capability.
Challenges in precision and manipulation
To be truly useful in daily life, besides moving, robots must manipulate objects with great delicacy. In the tests conducted, the system controlled five-fingered robotic hands and mechanical grippers to perform tasks such as closing packages, fitting components, and organizing materials.
The results point to significant progress but also reveal limitations. Tasks requiring extremely fine finger movements remain more challenging compared to operations performed by conventional grippers. Google reported that 'We are still advancing in the level of precision and speed to achieve near-human dexterity.'
Additionally, the company introduced safety features, including mechanisms capable of detecting risks, suspending operations, and requesting human assistance when necessary.