Clustered Pentaho execution moves work onto Carte slave servers, and when those slaves degrade the symptom is a transformation that simply never finishes. The master reports it as running indefinitely, and the operator's only lever is a restart.
Clustered transformations began hanging several times a week. Carte slave JVMs grew steadily through the day and were never reclaimed, orphaned transformation sessions accumulated in the slave's registry, and restarting the master did not clear them. The batch window began overrunning into the production shift.
Pentaho Data Integration with a Carte master and several slave servers on on-premises Linux hosts, feeding a manufacturing reporting warehouse.
AceMQ captured heap dumps and thread dumps from a hung slave rather than restarting it, which identified where rows were backing up and which step was blocking. We corrected the row-set sizing that was causing unbounded buffering, added lifecycle handling so completed transformations are cleaned from the slave registry, and containerized the slaves so a degraded instance is replaced rather than nursed.
The weekly hangs stopped and the batch window returned inside its allotted time. Slave restarts became a rare, automated event rather than a manual overnight task.
Ongoing support for a Pentaho Data Integration estate that must keep running reliably while a longer-term replacement is planned.
Resolving containers repeatedly OOMKilled because the JVM and Node runtimes inside them were sizing heap against host memory rather than the cgroup limit.
Whether you need architecture advisory, 24/7 support, or full managed services, AceMQ has the expertise to help.