A retail analytics provider exposed Druid-backed dashboards to its own customers, so query latency was a product characteristic rather than an internal concern. Broker timeouts appeared unpredictably and defied the team's attempts to reproduce them. AceMQ took over support with named senior engineers and direct escalation.
Druid query latency depends on how many segments a query touches, how well those segments are sized, whether results come from cache, and whether historical processing threads are contended. The timeouts correlated with datasources whose segments were both too numerous and too small, which meant per-segment overhead dominated actual work. Cache configuration was also negating itself on the highest-volume query patterns.
Apache Druid on AWS serving customer-facing analytics dashboards with published latency expectations.
AceMQ instrumented per-query segment counts and scan times to distinguish queries that were genuinely expensive from queries paying overhead on fragmented segments. Support then covered both the operational response to timeouts and the structural work needed to stop generating them.
Broker timeouts became rare and p99 dashboard latency roughly halved as segment fragmentation was retired. Query laning keeps heavy ad-hoc analysis from affecting the customer-facing path even when it runs during business hours.
Resolving Kafka supervisor lag caused by task slot exhaustion, oversized ingestion tasks, and handoff failures to deep storage.
Assessment of historical tiering, retention rules, and replication factors to align infrastructure cost with how data is actually queried over time.
Whether you need architecture advisory, 24/7 support, or full managed services, AceMQ has the expertise to help.