Gemini 3.5 Flash Gains Ability to See Screens and Operate Computers
Google has rolled out a groundbreaking update to its Gemini 3.5 Flash model, endowing it with the ability to "see" and interact with computer screens, effectively allowing it to operate a computer autonomously. This advancement marks a significant leap in AI capabilities, moving beyond conversational interfaces to direct interaction with digital environments. The updated model can now interpret visual information on a screen, navigate various applications, and execute actions independently to complete intricate tasks.
This new functionality is currently being made available to developers and enterprise clients through the Gemini API and the Gemini Enterprise Agent Platform. The integration of computer use directly into Gemini 3.5 Flash simplifies the process for developers who previously relied on a dedicated Gemini 2.5 computer use model for creating custom AI agents. Now, building agents that can reason, navigate, and take action across diverse digital settings becomes more streamlined and efficient.
To illustrate its potential, Google demonstrated Gemini 3.5 Flash's ability to perform tasks like finding the cheapest flights by browsing multiple booking websites, inputting dates, and sifting through available options. Another example involved the AI playing the game 2048, showcasing its capacity to make decisions and execute moves within a visual interface. These demonstrations highlight the model's versatility in handling both practical and recreational applications.
The introduction of such powerful capabilities naturally raises important questions regarding safety and control, particularly for enterprise users. Google has addressed these concerns by incorporating robust safeguards into the computer use feature of Gemini 3.5 Flash. These include targeted adversarial training to enhance the model's resilience and two key protective mechanisms. First, the model can be configured to require explicit user confirmation before proceeding with any sensitive or irreversible actions. Second, it is equipped to automatically halt tasks if it detects a prompt-injection attack, thereby mitigating potential misuse and ensuring a more secure operational environment. This dual approach aims to balance the revolutionary potential of autonomous AI with necessary ethical and security considerations.
Read original source