A public sector agency ran a Cassandra cluster inherited from a vendor engagement that had ended years earlier. Documentation was thin and the operating team had never been given a baseline. AceMQ assessed the cluster against its current workload and projected growth.
Assessing an inherited cluster means separating deliberate design choices from accumulated accident. Replication factor and consistency levels needed checking against the durability the agency actually required, token distribution needed verification for hot spots, and JVM and garbage collection behavior needed measurement under real load rather than inference from settings.
On-premises Apache Cassandra cluster supporting case management and records retention workloads.
AceMQ collected evidence from the running cluster — token ownership, per-node load distribution, garbage collection pause distributions, compaction backlog, and pending task queues — and evaluated it against the agency's durability and availability requirements. Findings were ranked by how close each was to causing an incident.
The agency received a baseline it had never had, with the items nearest to causing an outage identified and sequenced. Several inherited configuration choices, including consistency levels that did not match the stated durability requirement, were corrected as a result.
Redesigning partition keys and clustering order to eliminate unbounded partitions and remove secondary index queries that were hitting every node.
Ongoing support for anti-entropy repair that never completed within gc_grace_seconds, leaving the cluster exposed to deleted data resurrecting.
Whether you need architecture advisory, 24/7 support, or full managed services, AceMQ has the expertise to help.