On this page
Short answer
Choose RabbitMQ when agents hand off jobs that run for seconds or minutes and must not be lost: per-message acknowledgements, redelivery, consumer timeouts, dead-lettering and routing by exchange are broker features, and Direct Reply-To covers request-reply. Choose Kafka when the event history is the product: agents read and replay a retained log, stream processing sits beside it, and Kafka 4.2 made share groups production-ready for queue-style work. Choose NATS when agents are spread across clouds, sites and devices and most traffic is request-reply or fan-out, with JetStream adding persistence where needed. Nothing stops you running two: a log for events and a queue for work.
RabbitMQ, Kafka and NATS on what agent systems need
| RabbitMQ | Apache Kafka | NATS with JetStream | |
|---|---|---|---|
| Task hand-off | Per-message acknowledgements; unacknowledged messages requeued when a channel closes | Offsets per partition; share groups add per-record accept, reject and, from 4.2, renew | JetStream consumers acknowledge each message; Core NATS has no acknowledgements |
| Long-running work | Consumer timeout, 30 minutes by default, set per queue or per consumer | Share group record lock, 30 seconds by default, extended with renew | AckWait, 30 seconds by default, reset by in-progress acknowledgements |
| Poison messages | Quorum queues dead-letter after 20 deliveries by default | Share groups stop after 5 delivery attempts by default | MaxDeliver and backoff per consumer; term ends retries |
| Request-reply | Built in: Direct Reply-To, no reply queue | Not built in; correlate request and reply topics in code | Built in: replies to an _INBOX subject, with a no-responders signal |
| Routing | Topic and headers exchanges; bindings change without touching producers | Topic and partition key; filtering happens in consumers | Subject hierarchy with wildcards; queue groups load-balance |
| Replay | Streams replay; queues delete on acknowledgement | Core design: retained log, seven days by default | Streams replay from start, sequence or time |
| Ordering | Queue holds publication order; redeliveries can arrive out of order | Within a partition | Stream sequence number per stored message |
| Backpressure | Consumer prefetch; publishers blocked by flow control and alarms | Consumers pull at their own pace | MaxAckPending, 1,000 by default |
What an agent system asks of a broker
Agent protocols do not replace the broker. The Model Context Protocol defines two standard transports, stdio and Streamable HTTP, and allows custom ones, so tool calls between an agent and an MCP server normally travel over stdio or HTTP rather than through a broker. The broker carries what sits around them: jobs queued for workers, results returned to the agent that asked, events fanned out to other agents, and a record of what happened. Each of those maps to a broker property: acknowledgement and redelivery, request-reply, routing, and retention with replay.
Task hand-off and long-running work
Agent tasks are slow, because a single step may wait on a model for minutes. RabbitMQ keeps a delivery until the consumer acknowledges it and returns it to the queue if the consumer disappears or exceeds the consumer timeout, which RabbitMQ 4.3 lets you set per queue or per consumer on quorum queues. Prefetch limits how many jobs one worker holds. Python teams usually reach this through Celery; our Celery and RabbitMQ guide for LLM task queues covers the settings.
Kafka consumer groups track one committed offset per partition, so a slow record usually delays the records behind it in that partition. Share groups, production-ready in Kafka 4.2, let several consumers work the same partitions with per-record acknowledgement and delivery counting; the record lock is 30 seconds by default, configurable up to an hour, and 4.2 added a renew acknowledgement for longer processing. NATS JetStream redelivers after AckWait, 30 seconds by default, and an in-progress acknowledgement resets the timer for long jobs. A WorkQueue stream deletes each message after one acknowledgement.
Request-reply for tool calls
When an agent calls a tool over a broker it needs the answer back. NATS has request-reply in its core protocol: the client creates a private _INBOX subject, the responder replies to it, and if nobody is subscribed the server signals no responders at once rather than letting the caller time out. Synadia's open-source Agent Protocol for NATS builds agent discovery, streamed replies and heartbeats on these primitives and documents its delivery as at-most-once. RabbitMQ Direct Reply-To delivers replies to the requester's channel without creating a reply queue, which keeps request-reply cheap at tens of thousands of clients. Kafka has no request-reply primitive; applications pair a request topic with a reply topic and correlate the messages themselves.
Routing by capability
Routing decides which agent or worker pool gets a task: by skill, model, tenant or GPU class. RabbitMQ does this on the broker. A topic exchange matches routing keys such as agent.summarize.gpu against bindings with * and # wildcards, so a new specialist agent is a new binding and producers do not change. NATS subjects give similar flexibility with * and > wildcards, and queue groups spread requests across instances of one agent. Kafka routes by topic and partition key; finer selection happens in consumers or in stream processing, which suits fan-out of events better than dispatch of individual tasks.
Replay, ordering and backpressure
Replay matters when a new agent must catch up on history or a run must be audited. It is Kafka's core design: records stay until retention removes them, seven days by default, and every consumer group reads at its own offset, in order within a partition. RabbitMQ streams and NATS JetStream streams both offer non-destructive reads and replay; RabbitMQ queues delete messages once acknowledged. For backpressure, RabbitMQ limits unacknowledged deliveries per consumer and blocks fast publishers through flow control, JetStream caps in-flight messages with MaxAckPending, and Kafka consumers simply pull at their own pace while the log absorbs bursts.
AI tooling around each broker
Vendors now ship agent tooling on top of these brokers. Described from each vendor's own pages:
- Confluent: Streaming Agents are event-driven agents that run on Flink in Confluent Cloud, with MCP tool calling and replay for debugging. Its open-source MCP server lets AI assistants work with Confluent Cloud, Confluent Platform and Apache Kafka; a fully managed version also exists.
- Synadia, the company behind NATS: the open-source Agent Protocol for NATS and agent SDKs for Python and TypeScript.
- Solace: Solace Agent Mesh, an agent development and runtime platform on Solace's event broker, with an open-source framework and support for A2A and MCP.
- StreamNative: Orca Agent Engine, in private preview, runs Google ADK and OpenAI Agents SDK agents on Pulsar and Kafka topics in StreamNative Cloud.
- Redis: Redis Agent Memory, in preview, stores session and long-term memory for agents. It is a memory service, not a broker.
Agents reach RabbitMQ through its standard clients and protocols, including AMQP 0-9-1, AMQP 1.0 and MQTT, and through task frameworks such as Celery.
Redis Streams, Pulsar and cloud queues
Redis Streams give consumer groups, acknowledgements and claiming of stuck entries, which is enough for light task hand-off where Redis already runs; durability depends on persistence settings. Apache Pulsar offers Shared and Key_Shared subscriptions with individual and negative acknowledgements, closer to queue semantics than classic Kafka consumers. Amazon SQS needs no broker at all, but a message is invisible for at most 12 hours while a consumer works on it and is kept for at most 14 days.
When each wins
- Agents hand off long, expensive jobs that must not be lost? RabbitMQ with quorum queues, or Kafka share groups if Kafka is already the platform.
- Tool calls and agent-to-agent requests across clouds, sites and devices? NATS.
- Agents react to business events and new agents must replay history? Kafka.
- Routing by skill, model or tenant without changing producers? RabbitMQ exchanges or NATS subjects.
- A small team that wants one lightweight system for messaging, persistence and a key-value store? NATS with JetStream.
See RabbitMQ for AI workloads and the guide to RabbitMQ for AI agents for the RabbitMQ side in depth. AceMQ provides 24/7 RabbitMQ support and Kafka support, so the advice does not depend on which broker you pick.
Frequently asked questions
What is the best message broker for AI agents?
It depends on the traffic. RabbitMQ fits task hand-off with acknowledgements and routing, Kafka fits event streams that agents replay, and NATS fits request-reply and fan-out across distributed sites. Some systems run a log and a queue side by side.
Can Kafka be used as a task queue for AI agents?
Yes, through share groups, production-ready since Kafka 4.2. They add per-record acknowledgement, delivery counting and lock renewal, so several consumers can process records from the same partitions.
Does NATS guarantee delivery of agent messages?
Core NATS delivers at most once to subscribers connected at the time. JetStream adds storage, acknowledgements, redelivery and replay.
Do AI agents need a message broker if they use MCP?
MCP defines how an agent talks to a tool server, over stdio or HTTP. A broker is still useful for queuing long jobs, fanning out events and keeping a replayable record.
Related
Where this gets done
The work behind this page, run by the same engineers who wrote it.
- 24/7 RabbitMQ support15-minute emergency SLA, versions back to 3.8.x
- Managed RabbitMQ servicesWe run the brokers, on your infrastructure or hosted
- RabbitMQ consultingArchitecture, migration and remediation from senior engineers
- RabbitMQ health checkEngineer-led assessment with a prioritised fix list
- Extended LTS support for RabbitMQ 3.xCVE backports for versions the community no longer patches
- RabbitMQ commercial licensingTanzu RabbitMQ licences from an authorized Broadcom partner
- RabbitMQ troubleshootingLive incidents and recurring faults
- RabbitMQ upgrades3.x to 4.x, planned and executed in your window
- RabbitMQ migrationsFrom IBM MQ, Kafka, cloud brokers or older RabbitMQ
- RabbitMQ implementation and architectureCluster design, DR and go-live
- RabbitMQ corporate trainingAdmin and developer courses taught by working engineers
Other RabbitMQ guides, comparisons and research
- GuideThe RabbitMQ Performance Tuning GuideRead the guide
- GuideThe RabbitMQ Streams GuideRead the guide
- GuideThe RabbitMQ Reliability Guide: Ten Failure Patterns and Their FixesRead the guide
- GuideThe RabbitMQ Disaster Recovery GuideRead the guide
- GuideThe RabbitMQ Clustering and Sizing GuideRead the guide
- GuideThe RabbitMQ on Kubernetes GuideRead the guide
- GuideThe RabbitMQ Migration GuideRead the guide
- GuideThe RabbitMQ Security and Hardening GuideRead the guide
- GuideThe RabbitMQ Monitoring and Alerting GuideRead the guide
- GuideAgent-to-Agent Messaging over AMQP and RabbitMQRead the guide
- ComparisonRabbitMQ vs Amazon SQS ComparedSee the comparison
- ComparisonManaged RabbitMQ Options ComparedSee the comparison
- ComparisonMessage Broker Support Options ComparedSee the comparison
- ResearchWhat Breaks in Production RabbitMQ: 145 Support Tickets, 2023 to 2026Read the research
- ResearchRabbitMQ in Production 2026: What 22 Assessed Estates Actually RunRead the research
- ResearchThe RabbitMQ CVE Register, 2026 EditionRead the research
Recent RabbitMQ articles
- RabbitMQ Cluster Operator vs Helm Chart on KubernetesOct 2026
- RabbitMQ Alternatives: 9 Options and When to StayOct 2026
- RabbitMQ consumer_timeout: Why Consumers VanishOct 2026
- RabbitMQ End of Life and End of Support Dates (3.6 to 4.3)Oct 2026
- VMware Licensing Cost in 2026Sep 2026
- The Tanzu Software in Your VCF You Are Not UsingSep 2026
Need this done on your cluster?
AceMQ's senior RabbitMQ engineers support 130+ enterprise clients in 26+ countries under a 15-minute emergency SLA, with direct escalation to the RabbitMQ core team.