Apache Pulsar Support

24/7 Apache Pulsar Support with a 15-Minute Emergency SLA

AceMQ supports Apache Pulsar's multi-layer architecture in production — BookKeeper bookies under disk and IO pressure causing cluster-wide write latency spikes, backlog quotas triggering unexpected producer throttling, and ZooKeeper metadata store issues cascading into broker instability. Every ticket reaches a named senior engineer who understands both the broker and bookie layers.

Senior Apache Pulsar engineers on call right now — 24/7/365
15 min emergency SLA24 /7 global coverage130 + enterprise customers26 + countries served

Trusted for mission-critical Apache Pulsar by teams in finance, healthcare, defense, and telecom

Escalation Path

Your first hour of a Apache Pulsar outage

Most vendors publish an SLA number. This is what actually happens, minute by minute, when you page a senior AceMQ engineer.

T+0

You page us

Phone, email, or Slack — any channel reaches the on-call senior engineer directly. No web form, no tier-1 queue.

T+15

Named engineer live

A senior engineer who already knows your environment joins a live bridge. Zero cold-start, no re-explaining your topology.

T+30

Root cause isolated

Direct broker access, log and metric review, and a working hypothesis with a rollback plan before we touch anything.

Post

Written RCA

Documented root cause, the fix applied, and the prevention steps — delivered after every P1, not just when asked.

Response Times

SLA tiers, contractually guaranteed

Every tier reaches a senior Apache Pulsar engineer. There is no tier-1 triage layer to get through.

P1 — Emergency
15 min

Production down, messages not flowing, cluster or broker failure

P2 — Critical
1 hour

Severe degradation, rising error rates, approaching capacity limits

P3 — High
4 hours

Performance issues, configuration problems, non-critical failures

P4 — Standard
Next day

Questions, guidance, best practices, non-urgent improvements

Incident Triage

Apache Pulsar problems we fix every week

These are real symptoms from real Apache Pulsar production environments — and the first thing our engineers check when one comes in.

Write latency spikes cluster-wide with no obvious broker cause
What we check firstBookKeeper bookie disk IO and journal/ledger volume separation. Pulsar's storage layer is a common blind spot for teams used to Kafka's single-layer model — a single overloaded bookie degrades writes for every topic with an ensemble that includes it, not just its own partitions.
Typical resolution1–3 hours
Producers suddenly get throttled or messages start getting rejected
What we check firstBacklog quota configuration and current backlog size against the configured retention policy. Whether the topic is set to producer_exception, producer_request_hold, or consumer_backlog_eviction determines whether this is throttling or actual message loss — and that policy choice is easy to get wrong.
Typical resolutionUnder 2 hours
Retention behaves unexpectedly — messages disappear early or never expire
What we check firstNamespace-level and topic-level policy precedence. Pulsar's policy hierarchy (namespace default vs. explicit topic override) is a frequent source of retention and deduplication behaving differently than the admin expected.
Typical resolution1–2 hours
Brief unavailability or latency blips during normal operation
What we check firstBundle unload and split events in broker logs. Pulsar rebalances topic ownership across brokers by splitting and unloading bundles automatically, and this is often mistaken for an incident when it's actually expected load-balancing behavior — though we verify it isn't masking a real imbalance.
Typical resolutionUnder 1 hour
Geo-replicated clusters drift out of sync
What we check firstReplication backlog and network latency between clusters, plus whether a replicator connection dropped and didn't recover cleanly. Geo-replication lag compounds quietly until a failover reveals how far behind the remote cluster actually is.
Typical resolution2–4 hours
Broker instability that seems to start with metadata operations
What we check firstZooKeeper (or configured metadata store) session timeouts and quorum health. Because brokers, bookies, and topic ownership all depend on the metadata store, a struggling ZooKeeper quorum cascades into broker-level symptoms that look unrelated at first glance.
Typical resolution2–4 hours
Old data offloaded to S3 can't be read back
What we check firstTiered storage offload configuration and the offload driver's credentials/permissions against the target bucket. A misconfigured offload threshold or a permissions change on the bucket leaves cold data present but inaccessible, which surfaces only when a consumer actually seeks into it.
Typical resolution2–4 hours

Resolution times reflect typical Apache Pulsar engagements under an active AceMQ support contract. Every P1 closes with a written root-cause analysis.

Not on the list? Tell us what's breaking
What's Included

Everything in your Apache Pulsar support contract

No add-on pricing for incidents. No per-ticket charges. One contract covers the whole surface.

Emergency Incident Response

A bookie under IO pressure, a broker instability cascade, or a metadata store quorum issue. A senior engineer joins a live bridge within 15 minutes with access across the broker and bookie layers — not a ticket acknowledgement.

Root Cause Analysis

Every P1 closes with a written RCA: what failed, why, the fix applied, and the specific config or infrastructure change that prevents recurrence.

Broker & BookKeeper Tuning

Journal/ledger disk separation, backlog quota and retention policy design, and bundle assignment tuning against your actual topic count and throughput profile.

Metadata Store Health Advisory

Proactive monitoring of ZooKeeper (or your configured metadata store) quorum health, since a single degraded metadata node can cascade into cluster-wide broker instability.

Capacity & Storage Reviews

Quarterly reviews of bookie disk headroom, backlog growth trends, and tiered storage offload behavior so you size ahead of demand rather than reacting to a throttled producer.

Kafka Migration & Multi-Layer Onboarding

Migration planning from Kafka, including mapping the broker/bookie architectural split for ops teams new to it, plus geo-replication design for multi-region deployments.

Anywhere You Run It

We support Apache Pulsar wherever it's deployed

Cloud, Kubernetes, bare metal, hybrid, and air-gapped — including environments where you can't give us outbound network access.

StreamNative CloudSelf-managed Pulsar clustersAWS (EC2, EKS)Google Cloud (GKE)Microsoft Azure (AKS)Kubernetes & OpenShiftBare metal & on-premiseTiered storage to S3/GCSGeo-replicated multi-regionAir-gapped / no outbound access
Why AceMQ

What you get that you don't get elsewhere

Named Engineers, Zero Cold Start

The same senior engineers stay on your account. They know your bundle layout, your bookie ensemble configuration, and your replication topology — so a P1 starts with diagnosis, not twenty minutes of explaining your environment.

No Tier-1 Triage Layer

You reach a senior Pulsar engineer directly by phone, email, or Slack. No help desk collecting information to pass along, no escalation approval process standing between you and someone who can fix it.

We Know the Storage Layer, Not Just the Broker

Pulsar's split broker/bookie architecture is exactly where teams coming from Kafka get surprised. We treat BookKeeper as a first-class diagnostic target, not an afterthought behind the broker logs.

Genuine Follow-the-Sun Coverage

Engineers across 26+ countries and every time zone. Your 3am latency spike is someone's mid-afternoon — no overnight skeleton crew, no waiting for a region to wake up.

Proactive, Not Just Reactive

Quarterly health checks on bookie disk and metadata store quorum health, plus shared intelligence across our support base. When a version-specific bug surfaces on one customer's cluster, every affected customer hears about it before it reaches their production.

Honest About the Operational Complexity

Pulsar's multi-layer architecture buys real flexibility but real operational surface. We tell customers plainly when that trade-off fits their use case and when a simpler system would serve them better.

FAQ

Apache Pulsar support questions

Your Pulsar Cluster's Storage Layer Shouldn't Be a Blind Spot

Whether you need emergency response tonight or a support contract that treats BookKeeper as seriously as the broker, AceMQ staffs every engagement with a named senior Pulsar engineer. Support quotes returned within 24 hours.

Get in Touch

Talk to a Support Expert

Send us a message and we'll follow up within one business day — or book a free 30-min consultation directly.

305-204-2607
info@acemq.com
66 W. Flagler St. 9th Floor
Miami, FL 33130

Prefer to talk now? Call us directly or use the consultation tab to find a time that works.

We respond within 1 business day.

Pick a time that works — no pressure, no pitch. Just 30 minutes with an expert.

We respond within 1 business day.