A logistics provider had metrics and traces disappearing intermittently for short-lived workloads. The vendor's support confirmed nothing was wrong on the platform side, which left the customer's own agent and infrastructure layers as the place to look. AceMQ works those layers.
Short-lived batch pods were terminating before the agent flushed its buffer, so their final metrics and any traces from the last seconds of execution never arrived. Separately, a network policy change had started dropping a portion of egress traffic during peak windows, and the agent's retry behavior masked it as intermittent rather than systematic. The two problems looked like one flaky symptom, which is why it had gone unresolved for months.
Datadog agents deployed as a DaemonSet on Kubernetes with high pod churn, egress through a shared network path with policy enforcement.
AceMQ does not operate the vendor's platform — we work the layers the customer owns and the vendor's support will not investigate. We instrumented the agent's own telemetry, correlated gap windows against pod lifecycle events and egress metrics, and separated the two overlapping causes before fixing either.
Telemetry gaps for short-lived workloads stopped entirely, and peak-window loss went to effectively zero once the egress policy was corrected. The customer can now distinguish a real service outage from a collection failure, which was previously guesswork.
Ongoing support for a Datadog monitor estate producing more alerts than the on-call rotation could meaningfully act on.
Consulting engagement to govern custom metric cardinality, log ingest, and trace volume so observability spend tracks value instead of accident.
Whether you need architecture advisory, 24/7 support, or full managed services, AceMQ has the expertise to help.