Google is transforming Gemini into a universal agent capable of working for several days

The new Gemini agent retains its memory over multiple days and delegates its tasks to specialized sub-agents while controlling computational costs.

Google Cloud is redefining Gemini around something broader than a conversational assistant. At Gemini at Work 2026, the company introduced the Gemini agent, a universal work agent designed to research information, handle knowledge work, create media, write and run code, operate enterprise software and delegate parts of a job to its own sub-agents.

The architectural shift is particularly important. Gemini becomes the persistent agent that holds context, tools and skills, while the underlying model becomes a separate choice. Google can route work across its own Gemini family and, in supported environments, Anthropic Claude models without moving the memory and workflows surrounding the agent.

The approach extends the broader Gemini Enterprise agent platform Google introduced earlier this year. One Gemini across Workspace, desktop and command line

Google wants the same agent to remain available across environments.

Gemini can operate on web, iOS, Android, Windows and macOS, as well as through the command line, Google Workspace, Microsoft 365, Slack and third-party applications. It can also run headlessly without a dedicated interface.

Execution is persistent in the cloud. Jobs that take hours or days can continue after a user closes their laptop while retaining the same context and memory.

Inside Workspace, the agent now works directly across Gmail, Drive, Docs, Slides, Sheets, Chat and Calendar while carrying the same memory, skills and enterprise controls between them. The model becomes interchangeable

Google is drawing a clearer distinction between Gemini the agent and the model executing a given task.

Gemini can dynamically choose which model fits a workload. Google's own model families remain central to the system, but Anthropic Claude models are already part of its multi-model strategy, with additional private and open models expected later.

Google's Gemini Enterprise release notes also confirm support for Claude Opus 5.5 and Claude Sonnet 5.5 in several developer environments.

The goal is partly economic. Simple workloads do not always justify running a premium reasoning model.

Google is therefore adding Smart Routing to send tasks toward an appropriate model, alongside real-time project spending caps that can pause agent execution once a predefined budget is reached. Gemini can create its own temporary workforce

Gemini is also becoming explicitly multi-agent.

For complex objectives, the primary agent can dynamically create specialized sub-agents and coordinate work across parallel and sequential steps lasting hours or days.

Google is introducing a more persistent category as well: coworker agents.

These agents can have defined organizational roles, persistent storage and even their own `@agents.company.com` email addresses.

Inside Workspace, a coworker agent can receive its own account, email, calendar, Drive and company-directory presence. Employees can add it to Chat spaces, mention it in conversations or assign it work through document comments.

The agent acts under its own identity rather than impersonating the employee who created it. Four types of memory

Google is also formalizing how Gemini remembers.

It describes session memory for an active job, semantic memory for structured knowledge accumulated from documents and interactions, procedural memory for how work gets done, including reusable skills, and episodic memory covering what the agent has done before.

Those memories sit alongside tools and skills.

Gemini can connect to systems including Confluence, Microsoft Office, Teams, Slack, Git, Jira, Salesforce, ServiceNow, BigQuery, Databricks, Postgres and Snowflake. It can also connect to internal and external MCP servers.

Skills are reusable instructions, knowledge and workflows that can be published at company, team or personal level.

The architecture increasingly resembles an agent harness: the foundation model is only one component inside a persistent layer of tools, memory, permissions and working methods. From natural-language questions to SQL and PySpark

Google is also giving Gemini domain-specific capabilities for data work.

Data and ML teams can ask it to generate PySpark, create notebooks, train models and troubleshoot pipelines. Business users can turn natural-language questions into reusable operational reports.

Gemini can identify relevant datasets through Knowledge Catalog, produce SQL, Spark or Python and execute the code in BigQuery or Managed Spark with Lightning Engine. It can then create charts and dashboards in the analytics tools teams already use.

Once an operational query is verified and saved, it can be rerun without repeatedly invoking a model and consuming additional tokens.

That is an important design choice: Google is turning successful generative analysis into deterministic reusable workflows wherever possible. Agents get their own identities

Agent autonomy also requires a different security model.

Google gives every agent a cryptographically attested identity governed through least-privilege permissions. The identity follows its actions into logs and even into virtual machines created to execute code.

Actions are written to an audit trail under the agent's identity, while access to outside systems can be propagated through standards such as OAuth.

Agents also execute inside an Agent Sandbox, with traffic routed through Agent Gateway, a policy enforcement layer Google describes as an AI network firewall.

This allows an organization to define a policy once and apply it across its agent fleet. Infrastructure designed for agent loops

Google is tying the shift to its latest infrastructure as well.

The TPU 8i is optimized specifically for inference and reinforcement learning, including latency-sensitive MoE and agent workloads. Google reports up to 80% better inference performance per dollar compared with the previous generation.

The design includes 384 MB of on-chip SRAM, 288 GB of HBM and 19.2 Tb/s interconnect bandwidth, alongside a dedicated Collective Acceleration Engine intended