A manufacturer running Prometheus at each plant plus a global federation layer found the global dashboards disagreed with local ones. AceMQ assessed the topology to determine whether federation was the right pattern at their scale.
Hierarchical federation was being used to pull raw series rather than aggregates, so the global instance was scraping an enormous payload from each site on every interval. Scrapes regularly exceeded the timeout and returned partial data, which is why global and local views disagreed. Sites with intermittent WAN links produced gaps that looked like plant outages. Nobody could tell which numbers to trust during an actual production incident.
Prometheus at multiple manufacturing sites with intermittent WAN connectivity to a central observability environment.
The assessment measured what federation was actually transferring against the scrape timeout budget, which made the partial-scrape problem plain. We then evaluated whether the customer's requirements were better served by remote write with a central store than by federation, and modeled the bandwidth and storage implications of each option against their real series counts.
The customer moved to site-level recording rules plus remote write, which eliminated the partial-scrape gaps entirely. Global and site dashboards now agree, and WAN outages backfill instead of leaving permanent holes in the record.
Consulting engagement to design multi-year Prometheus metric retention with downsampling and object storage, replacing oversized local disks.
Emergency remediation of Prometheus servers being OOM-killed repeatedly after a deployment introduced an unbounded label.
Whether you need architecture advisory, 24/7 support, or full managed services, AceMQ has the expertise to help.