Separating the dashboards people rely on from the ones nobody has opened in a year
The estate collapsed to a small fraction of its original dashboard count, with the survivors owned and version controlled. Engineers stopped misreading broken panels as healthy services, which had bee…
Overview
A logistics provider had accumulated thousands of Grafana dashboards across dozens of teams with no ownership model. AceMQ assessed the estate to determine what was in real use, what was broken, and what should be standardized.
Challenge
Dashboards had been cloned repeatedly, so the same service had multiple slightly different views and no clear source of truth. A large share of panels referenced metrics that no longer existed after service rewrites, which meant engineers regularly looked at empty graphs and assumed the service was idle rather than the panel broken. There was no way to tell which dashboards mattered because usage was never tracked.
Environment
Grafana on AWS serving multiple business units, dashboards created ad hoc through the UI with no version control.
Approach
We pulled the full dashboard inventory through the API, cross-referenced every panel query against the metrics actually present in the datasource, and layered in view statistics to rank by real usage. That produced three lists — keep, fix, and retire — plus a template design so future dashboards converge instead of diverging.
Solution
- 1Full dashboard and panel inventory extracted via the Grafana API, including folder and permission structure
- 2Every panel query validated against live datasource metadata to flag panels referencing metrics that no longer exist
- 3View statistics analyzed to rank dashboards by actual usage rather than by who shouts loudest
- 4Duplication analysis identifying clone families that should collapse into one parameterized dashboard
- 5Reusable dashboard templates with variables designed to cover the common service patterns across teams
- 6Provisioning-as-code model recommended so dashboards live in version control with named owners
Outcome
The estate collapsed to a small fraction of its original dashboard count, with the survivors owned and version controlled. Engineers stopped misreading broken panels as healthy services, which had been a recurring cause of delayed incident detection.
Technologies
Related Use Cases
Grafana Observability Stack Consolidation
Consulting engagement to consolidate fragmented metrics, logs, and traces onto a single Grafana-based observability layer with consistent labeling.
Grafana Alerting and Datasource Support
Ongoing support for Grafana unified alerting, notification routing, and datasource reliability across an operational monitoring estate.
Ready for a Grafana Health Check?
AceMQ's senior Grafana engineers have handled this exact type of engagement before. Whether you need architectural guidance, hands-on remediation, or an ongoing managed partnership, we're ready to help.