Skip to main content
Back to News
analysis/AI Infrastructure

Google Puts Agent Message Queues Inside Spanner Transactions

Google made Spanner queues generally available so AI agents commit a state change and the async task it triggers in one transaction, not two systems.

Stefan Trbojevic

Stefan Trbojevic

3 October 20265 min read
LinkedIn
Abstract dark slate network topology with pathways of glowing nodes converging into one transactional core

The takeaway

Database-native queues remove the outbox pattern and its reconciliation tax for agents whose state already lives in the same database, but they do not make external side effects exactly-once. At-least-once delivery plus at-most-once acknowledgement still requires idempotent tool calls and a caught ASSERT_ROWS_MODIFIED error.

Why it matters for builders

If your agent keeps state in Spanner, transactional task dispatch is a genuine simplification: no relay, no outbox table, no separate broker to operate, and scheduled retries and human-approval timeouts become ordinary writes. Two caveats before you migrate. External actions still need idempotency keys, because the guarantee is at-least-once delivery, not exactly-once effects. And this is a platform bet: Spanner queues only pay off if the state they protect is already in Spanner, so teams on Postgres or DynamoDB will keep paying for a broker or an outbox until their own vendor ships the same primitive.

Google Puts Agent Message Queues Inside Spanner Transactions

Google Cloud today made Spanner queues generally available, folding asynchronous task messaging directly into its globally distributed database. The pitch is narrow and technical: an AI agent's state change and the action it triggers commit in a single transaction, or neither happens at all. For teams wiring agents into production systems, that removes one of the most common failure modes in agentic architecture.

Abstract diagram of a shared junction fusing two data lanes into one ordered flow

What Google shipped

Spanner queues are relational structures, not a sidecar. You define a queue with DDL, exactly like a table, with a payload column and a primary key. Enqueuing a message is an ordinary write inside a read-write transaction. Workers consume messages through a table-valued function, RECEIVE_QUEUE_NAME(), streamed to the client over the ExecuteStreamingSql API. Acknowledgement is a delete inside a transaction.

Four documented capabilities matter for agents. Atomic decide-and-act enqueue lets an agent update its memory or state tables and enqueue tasks to peer agents in one read-write transaction, backed by Spanner's strict serializability. Scheduled delivery uses a DeliverTime column, so delayed retries, agent check-ins, and SLA escalation timers can be enqueued transactionally instead of living in an external cron scheduler. Message leases, returned as a lease token and expiry with every leased task, support work that runs longer than the default window through RENEWLEASE_QUEUE_NAME(). And human-in-the-loop timeouts are handled natively: one transaction records the pending approval and schedules the escalation, and whichever fires first resolves the wait.

Because claim, acknowledge, and state all live in the same database, backlog inspection and execution history become ordinary GoogleSQL queries rather than a separate broker console.

The consistency gap it closes

The failure mode Google is attacking is architectural, and anyone running agents in production has met it. The agent keeps state in an operational database and dispatches asynchronous work through a separate queue. Those two systems have disjoint commit points, so two failure modes are always live.

If the database write succeeds and the dispatch fails, the agent has decided to act but nothing executes. If the dispatch succeeds and the state transaction rolls back, the agent executes against a state that never existed. Retries, speculation, and race conditions in multi-agent coordination amplify both.

The usual defence is the transactional outbox: write the intent to an outbox table in the same transaction, then run a relay that reads it and publishes downstream. It works, and it is infrastructure you own forever, plus idempotency keys on every consumer and a reconciliation job for whatever slipped through. Google calls this a reliability tax. The plainer description is that it is a second distributed system bolted onto the first to fix a boundary the database could not cross.

The interesting part for builders is the trade-off against durable execution engines. Platforms such as Restate and Temporal solve the same class of problem by owning orchestration in their own runtime, which is the better answer when a workflow spans many services over days. That trade-off was the subject of our analysis of Restate's $20M raise. Spanner queues are the database-native alternative: no new runtime, no relay, provided your state already lives in Spanner.

Abstract illustration of one pathway forking into parallel retry loops around a single core

Exactly-once is still your job

This is where vendor copy and engineering reality part company, and it is the most valuable thing in Google's own documentation.

Spanner queues guarantee at-least-once delivery and at-most-once acknowledgement. That combination does not make your external side effects exactly-once. Google says exactly-once processing is achievable, and it is, but only by design: the documentation states that applications must still be safe to retry, and recommends passing the task ID as the idempotency key to the external API being called.

Google also documents a race worth internalising. If a worker stalls, its lease can expire and another worker can process and delete the message. When the stalled worker returns and tries to acknowledge, ASSERT_ROWS_MODIFIED 1 fails the statement instead of silently overwriting newer state. That protection only works if you catch the error and abort the transaction.

Two operational limits matter at planning time: 100 queues for instances with at least one node, and 1,000 active receive queries per queue with identical arguments. Queues are also not change streams. Change streams do continuous capture for downstream replication, while queues are for leased transactional work with acknowledgement.

Abstract grid of separate platform slabs linked by thin luminous data pathways

The wider pattern

Google is not alone in deciding that the agent platform should be the platform you already run.

MongoDB opened the public preview of Atlas Agent Engine, a unified execution, memory, and governance layer for production agents, framed explicitly as putting agents to work without adopting a new stack. IBM shipped agent identity in watsonx Orchestrate, giving each agent a distinct identity registered with the identity provider a security team already operates, task-scoped short-lived tokens instead of inherited user permissions, and an audit trail that records which agent called which tool on whose behalf.

Strip the marketing and all three announcements say the same thing. What production agents are missing is not another runtime. It is queueing, memory, identity, and audit, and those belong in the database and directory you already trust.

What to watch

Watch whether other managed databases ship native queues, and whether durable execution vendors respond by making their engines cheaper to embed rather than adopt. Watch whether transactional task dispatch becomes a standard checkbox in agent framework evaluations, the way connection pooling once did. Watch whether agent runtime isolation and agent data primitives converge, as they did for sandboxes earlier this week. And watch the limits: 100 queues per instance is generous for one application and tight for a platform running thousands of agents.

Share𝕏

The Automation Brief

Read 5 AI stories instead of 50.

The essential moves in AI agents, models, automation and infrastructure — filtered for builders and operators, with the part that actually matters.

No noise. Unsubscribe anytime.

Editorial notes

Reported by

Stefan Trbojevic

Edited by

n8n Lab Editorial

Published

3 October 2026

Updated

3 October 2026

AI disclosure: AI assisted with research and drafting. Factual claims are reviewed by an editor.

n8n Lab is an independent service provider. We are not affiliated with, endorsed by, or sponsored by n8n GmbH. “n8n” is a trademark of n8n GmbH and is used here only to describe the platform-specific implementation and automation services we provide.