Google is taking a major leap in AI automation with its latest model, Google’s Gemini 2.5 Computer Use, designed to browse and interact with the web just like a human. The new model enables AI agents to perform real-world tasks—such as navigating websites, filling out forms, and submitting data—directly through a browser interface.
The tech giant says Google’s Gemini 2.5 Computer Use leverages advanced visual understanding and reasoning capabilities to interpret user requests and execute them seamlessly across websites. This allows the model to operate within human-designed web environments, even when there’s no API or direct integration available.
Beyond traditional automation, Gemini 2.5 can be applied to interface testing, task execution, and even web-based workflows—marking a step forward in Google’s vision of fully functional AI agents. Earlier versions of this technology powered Project Mariner, an experimental AI that could independently perform browser tasks like adding items to an online shopping cart based on user prompts.
This announcement follows closely on the heels of OpenAI’s Dev Day, where ChatGPT introduced new app integrations and enhanced agent features. Similarly, Anthropic rolled out a version of its Claude model with built-in computer-use capabilities last year, underscoring the fierce competition among AI leaders in the agentic automation space.
In Google’s demonstration videos—sped up for clarity—the AI is seen navigating websites, typing, and dragging elements. While Google’s Gemini 2.5 Computer Use isn’t yet optimized for full desktop control, it currently supports 13 actions, including opening web browsers, clicking buttons, and managing drag-and-drop operations.
The company claims its model “outperforms leading alternatives on multiple web and mobile benchmarks.” However, unlike ChatGPT Agents or Anthropic’s tools, Google’s model is limited to browser-level control and does not yet interact with operating systems directly.
Developers can now experiment with Google’s Gemini 2.5 Computer Use via Google AI Studio and Vertex AI. There’s also a live demo on Browserbase, where the model can perform tasks like “playing the 2048 game” or “browsing Hacker News for trending debates.”
This launch signals Google’s growing ambition to merge artificial intelligence with real-time web navigation—paving the way for a future where AI agents can execute complex online actions autonomously.