Getting the on-call rotation back to alerts that mean something
Page volume dropped by roughly two thirds while the monitors covering genuine customer impact stayed in place. Acknowledgement times improved because engineers stopped treating pages as background noi…
Overview
A SaaS provider's on-call engineers were acknowledging alerts reflexively because most of them required no action. AceMQ provides ongoing support to keep the monitor estate meaningful, working alongside the customer's team on their own Datadog organization.
Challenge
Monitors had accumulated over years with no retirement process. Many used static thresholds set against traffic levels that no longer existed, so they fired on every normal peak. Composite monitors created alert storms where one underlying failure produced dozens of pages. Multi-alert monitors grouped by a tag that had since become high cardinality, generating one alert per instance rather than one per service. Nobody owned the estate, so nothing was ever removed.
Environment
Datadog monitoring a multi-region AWS and Kubernetes estate, alerts routed into an on-call rotation.
Approach
We work from alert history rather than opinion: which monitors fired, how often, and whether anyone did anything as a result. Monitors that never lead to action get retired or downgraded to a dashboard. The ones that matter get thresholds that reflect current behavior, appropriate grouping, and downtime rules that suppress the predictable noise.
Solution
- 124/7 support with a 15-minute emergency SLA and named senior engineers embedded in the customer's alert review cadence
- 2Alert history analyzed by monitor to separate the ones that drive action from the ones that are always acknowledged and closed
- 3Static thresholds replaced with anomaly or forecast-based conditions where the underlying signal has real seasonality
- 4Multi-alert grouping corrected so alerts arrive per service rather than per ephemeral instance
- 5Composite monitor logic restructured so one root failure produces one page instead of a storm
- 6Monitors defined as code with owners attached, so unowned monitors are visible and removable
Outcome
Page volume dropped by roughly two thirds while the monitors covering genuine customer impact stayed in place. Acknowledgement times improved because engineers stopped treating pages as background noise.
Technologies
Related Use Cases
Datadog Agent Telemetry Gap Remediation
Remediation of intermittent Datadog telemetry gaps traced to agent buffering, container lifecycle, and network egress behavior on the customer's own infrastructure.
Datadog APM Instrumentation Coverage Assessment
Assessment of APM instrumentation coverage and trace completeness across a service estate where distributed traces kept breaking at service boundaries.
Need Expert Datadog Support?
AceMQ's senior Datadog engineers have handled this exact type of engagement before. Whether you need architectural guidance, hands-on remediation, or an ongoing managed partnership, we're ready to help.