Failover was a manual runbook nobody had rehearsed, and messages could not be replayed
The integration estate gained a rehearsed recovery procedure, a defensible answer to what happened during an outage, and a runbook its own engineers can execute without waiting on a vendor.
Overview
Boomi processes moved regulated transactions across systems, with RabbitMQ as the transport between them. Recovering from a data centre event meant a manual promotion nobody had practised, and once a message was consumed there was no way to replay it.
Challenge
Two gaps compounded each other. Failover depended on a human following an unrehearsed procedure, and because queues discard on consume there was no audit-friendly way to reconstruct what had flowed during an incident. Both had to be solved without breaking the controls the estate is audited against.
Environment
Boomi integration processes over RabbitMQ on a container platform, spanning a primary and a disaster recovery data centre.
Approach
AceMQ produced a disaster recovery design covering manual and automated failover, including what the commercial RabbitMQ distribution changes about the automated option. In parallel we assessed converting selected queues to streams so messages remain replayable after consumption, and specified a buffering layer so Boomi processes degrade rather than fail when the broker is unreachable.
Solution
- 1Disaster recovery design comparing manual promotion against automated failover
- 2Cluster sizing across the primary and DR data centres
- 3Queue-to-stream conversion assessment so messages stay replayable after consumption
- 4Review of replay against the estate's audit and retention controls
- 5Buffering layer so integration processes degrade instead of failing during an outage
- 6Operational runbook and training for the support and development teams
Outcome
The integration estate gained a rehearsed recovery procedure, a defensible answer to what happened during an outage, and a runbook its own engineers can execute without waiting on a vendor.
Technologies
Related Use Cases
Retry Automation and Downstream Back-Pressure Remediation
Reducing manual error-queue operations by improving retry handling, dead-lettering, and downstream flow management across RabbitMQ, BizTalk, and D365.
RabbitMQ Federation Remediation for Energy SCADA Systems
Resolving federation failures causing pipeline monitoring delays in critical SCADA infrastructure.
Ready for a Boomi Health Check?
AceMQ's senior Boomi engineers have handled this exact type of engagement before. Whether you need architectural guidance, hands-on remediation, or an ongoing managed partnership, we're ready to help.