→ Back to Home
Robotics

DeepMind's Gemini Robotics 2 Ushers in New Era of Whole-Body AI Control for Humanoid Robots

Google DeepMind has unveiled its next-generation robotics AI models: Gemini Robotics 2, Gemini Robotics ER (Embodied Reasoning) 2, and Gemini Robotics On-Device 2. The flagship Gemini Robotics 2 is a Vision-Language-Action (VLA) model that represents a significant leap by enabling end-to-end control of an entire robot body, including legs and feet, rather than being limited to upper-body manipulation. This model supports multi-robot collaboration and is designed for deployment across various robot form factors, including humanoids, dual-arm robots, and industrial manipulators. Simultaneously, Gemini Robotics ER 2 was introduced as a higher-level reasoning model, empowering robots to understand their surroundings, plan long-duration, complex tasks, and recover from execution errors. The third model, Gemini Robotics On-Device 2, is an optimized VLA model that runs locally on robotic devices, ensuring low-latency operations and adaptability to new robot types within hours, even without an internet connection. These models were showcased in a demonstration where a humanoid robot, equipped with a Sharpa dexterous hand, performed complex multi-finger tasks and navigated obstacles to complete a multi-step objective. This suite of new models fundamentally alters the landscape for robotics development and deployment. For engineers and integrators, it means a significant reduction in the complexity and time required to program advanced robotic behaviors. The ability of Gemini Robotics 2 to control whole-body movements and adapt to diverse platforms democratizes access to sophisticated robotic capabilities, making advanced automation more attainable for a wider range of industries. The embodied reasoning of ER 2 is crucial for robust real-world applications, as it allows robots to operate more autonomously and reliably in unpredictable environments, minimizing human intervention for error correction. Furthermore, the on-device capability of On-Device 2 addresses critical latency and connectivity challenges, opening doors for robotics in remote or sensitive environments where cloud dependency is not feasible. This directly impacts manufacturing, logistics, healthcare, and even domestic robotics, pushing the boundaries of what automated systems can achieve. This announcement from Google DeepMind is a direct continuation of the broader trend towards more generalized, AI-driven automation, deeply rooted in advancements in large language models (LLMs) and multimodal AI. Just as foundation models have revolutionized natural language processing and image generation, their application to robotics aims to create more versatile and intelligent physical agents. This mirrors the shift seen in cloud computing, where abstracting infrastructure with services like Kubernetes and serverless functions has enabled developers to focus on application logic rather than underlying hardware. Similarly, these robotics foundation models abstract away much of the low-level control and perception complexity, allowing developers to focus on high-level task definition. The emphasis on "on-device" deployment also aligns with the edge computing trend, bringing AI inference closer to the data source to reduce latency and improve reliability, a critical factor for real-time robotic operations. This convergence of AI, cloud principles, and edge computing is accelerating the development of truly autonomous and adaptable robotic systems. Practitioners should immediately begin exploring the Gemini Robotics SDK (if available to trusted testers) to understand its capabilities for their specific use cases. The ability to fine-tune models with as few as 50-100 demonstrations, as seen in previous Gemini Robotics iterations, suggests a lower barrier to entry for custom applications. Organizations should invest in upskilling their teams in AI model integration and data curation for fine-tuning, as the quality of demonstration data will be paramount. A key implication is the potential for significant cost savings in development and deployment, as a single generalized model can replace multiple task-specific programs. However, trade-offs include the computational resources required for training and inference, especially for the more complex ER 2 model, and the ongoing need for rigorous safety testing, particularly with whole-body control and autonomous error recovery. Practitioners should closely monitor DeepMind's release roadmap for wider access to these models and tools, and consider how these general-purpose AI capabilities can be integrated into existing robotic infrastructure to unlock new levels of automation and flexibility.
#robotics#ai#foundation models#deepmind#humanoid robots#embodied ai
Read original source