An advertising exchange depended on Druid for near-real-time bid analytics. Kafka supervisor lag grew from seconds to hours during peak traffic, so the dashboards traders relied on showed stale data exactly when it mattered. AceMQ engaged under the emergency SLA.
The lag had several compounding sources. Middle Manager task slots were exhausted because ingestion tasks were running long and holding slots that pending tasks needed. Segments produced by those tasks were oversized, so handoff to deep storage and loading onto historicals was slow. Failed tasks were being retried into the same saturated slot pool.
Apache Druid on Kubernetes ingesting bid and impression streams from Kafka for real-time analytics.
AceMQ traced the ingestion pipeline end to end — supervisor, task assignment, segment publishing, handoff, and historical load — to find where time was actually accumulating rather than tuning the supervisor in isolation. Capacity and task shape were corrected together, since fixing one without the other simply relocates the queue.
Ingestion lag returned to the seconds range and held there through peak traffic. Trading dashboards now reflect current data during the periods when they are actually used.
Ongoing support for broker timeouts and unpredictable query latency driven by segment sizing, cache behavior, and processing thread contention.
Redesigning segment granularity, partitioning, and auto-compaction policy so segment counts stay bounded as historical data accumulates.
Whether you need architecture advisory, 24/7 support, or full managed services, AceMQ has the expertise to help.