Tanzu GemFire Support

24/7 Tanzu GemFire Support with a 15-Minute Emergency SLA

AceMQ is Broadcom's VMware Expert Advantage Partner of the Year for the Americas. Our engineers run GemFire in trading, telecom, and payments environments where a two-second GC pause is an incident — heap pressure, members leaving on partitions, WAN sender backlog, and multi-hour disk-store recovery. Every ticket reaches a named senior engineer who already knows your region topology.

Senior Tanzu GemFire engineers on call right now — 24/7/365
15 min emergency SLA24 /7 global coverage130 + enterprise customers26 + countries served

Trusted for mission-critical Tanzu GemFire by teams in finance, healthcare, defense, and telecom

Escalation Path

Your first hour of a Tanzu GemFire outage

Most vendors publish an SLA number. This is what actually happens, minute by minute, when you page a senior AceMQ engineer.

T+0

You page us

Phone, email, or Slack — any channel reaches the on-call senior engineer directly. No web form, no tier-1 queue.

T+15

Named engineer live

A senior engineer who already knows your environment joins a live bridge. Zero cold-start, no re-explaining your topology.

T+30

Root cause isolated

Direct broker access, log and metric review, and a working hypothesis with a rollback plan before we touch anything.

Post

Written RCA

Documented root cause, the fix applied, and the prevention steps — delivered after every P1, not just when asked.

Response Times

SLA tiers, contractually guaranteed

Every tier reaches a senior Tanzu GemFire engineer. There is no tier-1 triage layer to get through.

P1 — Emergency
15 min

Production down, messages not flowing, cluster or broker failure

P2 — Critical
1 hour

Severe degradation, rising error rates, approaching capacity limits

P3 — High
4 hours

Performance issues, configuration problems, non-critical failures

P4 — Standard
Next day

Questions, guidance, best practices, non-urgent improvements

Incident Triage

Tanzu GemFire problems we fix every week

These are real symptoms from real Tanzu GemFire production environments — and the first thing our engineers check when one comes in.

Whole cluster stalls for seconds at a time; clients time out in bursts
What we check firstGC logs for full and mixed pause duration against member-timeout. A heap pause longer than member-timeout gets the paused member evicted from the distributed system, so the pause becomes a membership event — heap sizing and failure detection have to be tuned together, not separately.
Typical resolution1–3 hours
Member "unexpectedly left the distributed system" with no crash and no OOM
What we check firstLocator and peer logs for suspect processing, then member-timeout, ack-wait-threshold, and the network path between hosts. A brief partition or a paused JVM both look identical from the outside — the logs distinguish them, and the fix is different for each.
Typical resolution2–4 hours
WAN gateway sender queue growing; the remote site is falling further behind
What we check firstQueue size and batch statistics per gateway sender against receiver apply rate. Usually too few dispatcher threads, a batch-size tuned for a different message profile, or a receiver-side bottleneck — and with a persistent queue the backlog starts consuming disk while you diagnose.
Typical resolution2–4 hours
Entries the application expects are simply gone, with no error anywhere
What we check firstRegion eviction and expiration attributes. Heap LRU with a destroy action on a non-persistent region silently drops entries under memory pressure — which is correct behaviour and almost never what the team intended when they wanted overflow-to-disk.
Typical resolutionUnder 2 hours
Full cluster restart taking hours to return to service
What we check firstDisk-store oplog count and total size per member, and which member holds the most recent copy of each persistent region. Recovery replays oplogs and blocks until the newest copy is online — so a cluster with compaction disabled and years of oplogs recovers slowly by design.
Typical resolutionSame day
Serialization or ClassCastException errors after a rolling upgrade
What we check firstPDX type registry against the deployed class versions on each member. A field added or reordered on one side, a class that quietly moved off PDX to Java serialization, or read-serialized set inconsistently across members will all surface only once mixed versions are live.
Typical resolution1–3 hours
Function execution timing out on partitioned regions under load
What we check firstRegion colocation and the function's routing keys. A function that should execute on a single member but touches data across buckets turns one local call into a cross-member fan-out, and the timeout is the symptom rather than the problem.
Typical resolution1–3 hours

Resolution times reflect typical Tanzu GemFire engagements under an active AceMQ support contract. Every P1 closes with a written root-cause analysis.

Not on the list? Tell us what's breaking
What's Included

Everything in your Tanzu GemFire support contract

No add-on pricing for incidents. No per-ticket charges. One contract covers the whole surface.

Emergency Incident Response

Cluster stalled, members dropping, or a data grid that won't come back after a restart. A senior GemFire engineer joins a live bridge within 15 minutes with access to diagnose — not a ticket acknowledgement.

Root Cause Analysis

Every P1 closes with a written RCA: what failed, why, the fix applied, and the heap, timeout, or region configuration change that stops it recurring. Delivered as standard, not on request.

Latency & Heap Tuning

GC selection and heap sizing for sub-millisecond operation, off-heap configuration, eviction and expiration policy, OQL index design, and client pool sizing against your real P99 requirement rather than a default.

WAN & Multi-Site Operations

Gateway sender and receiver tuning, active-active conflict handling, persistent queue sizing, and site-failover procedures — including what to do when a site comes back after hours of divergence.

Persistence & Recovery Assurance

Disk-store layout, compaction policy, and rehearsed full-cluster restart procedures with measured recovery times, so your DR plan reflects how long recovery actually takes rather than an estimate.

Broadcom Licensing & Entitlement

As Broadcom's Expert Advantage Partner of the Year, we handle Tanzu GemFire subscription questions alongside the technical ones: entitlement reconciliation, true-up exposure, renewal terms, and quotes inside 24 hours.

Anywhere You Run It

We support Tanzu GemFire wherever it's deployed

Cloud, Kubernetes, bare metal, hybrid, and air-gapped — including environments where you can't give us outbound network access.

VMware vSphere & TanzuKubernetes & OpenShiftTanzu Platform / Cloud FoundryBare metal & on-premiseAWS (EC2, EKS)Microsoft Azure (AKS)Google Cloud (GKE)Hybrid & multi-site WANAir-gapped / no outbound accessGemFire 9.xGemFire 10.xApache Geode deployments
Why AceMQ

What you get that you don't get elsewhere

Broadcom Expert Advantage Partner of the Year

AceMQ holds Broadcom's top VMware partner award for the Americas in 2025. In practice that means an escalation path into Tanzu engineering when a problem is genuinely a product defect — and one team handling both your support and your licensing.

Named Engineers, Zero Cold Start

The same senior engineers stay on your account. They know your region types, your colocation model, and your heap profile — so a P1 starts with diagnosis instead of you narrating your topology while the cluster is down.

No Tier-1 Triage Layer

You reach a senior GemFire engineer directly by phone, email, or Slack. No help desk collecting information to pass along, no escalation approval between you and someone who can read a GC log against a membership timeline.

JVM Depth, Not Just Grid Configuration

Most GemFire incidents are JVM incidents wearing a distributed-systems costume. We diagnose across G1 and ZGC behaviour, off-heap allocation, safepoint pauses, and OS-level page cache — because that is where the pause that killed your member actually came from.

Genuine Follow-the-Sun Coverage

Engineers across 26+ countries and every time zone. Your 3am membership storm is someone's mid-afternoon — no overnight skeleton crew, no waiting for a region to wake up.

Proactive, Not Just Reactive

Quarterly health checks plus shared intelligence across our support base. When a version-specific bug surfaces on one customer's cluster, every affected customer hears about it before it reaches their production.

FAQ

Tanzu GemFire support questions

An In-Memory Grid Fails at Memory Speed. Support Should Match.

Whether you need emergency response tonight or a support contract that prevents the next membership storm, AceMQ staffs every engagement with a named senior GemFire engineer. Support quotes returned within 24 hours.

Get in Touch

Talk to a Support Expert

Send us a message and we'll follow up within one business day — or book a free 30-min consultation directly.

305-204-2607
info@acemq.com
66 W. Flagler St. 9th Floor
Miami, FL 33130

Prefer to talk now? Call us directly or use the consultation tab to find a time that works.

We respond within 1 business day.

Pick a time that works — no pressure, no pitch. Just 30 minutes with an expert.

We respond within 1 business day.