Keeping a legacy Pentaho estate reliable while you plan its replacement
The recurring nightly failures stopped, and half-loaded target tables were eliminated by making partial failures abort the load. The estate is stable enough that the migration can proceed on plan…
Overview
Deciding to leave Pentaho does not make the nightly load optional. Most organizations need someone who can still fix Kettle problems at 3am for the eighteen to thirty-six months a migration realistically takes.
Challenge
The nightly load had become unreliable: transformations failing on memory pressure when input volumes grew, jobs that reported success while a partially failed step left the target table half-loaded, and JDBC driver behavior that had changed after a database upgrade. Failures surfaced as business complaints the next morning rather than as alerts overnight.
Environment
Pentaho Data Integration on-premises Linux servers, SQL Server and PostgreSQL sources and targets, nightly batch window feeding operational reporting.
Approach
AceMQ provides named senior engineers with a 15-minute emergency response SLA and 24/7 coverage — no tier-1 triage queue between the customer and someone who can actually read the transformation. We addressed the recurring failures at the source rather than restarting jobs, and rebuilt the error handling so partial loads fail cleanly.
Solution
- 1Corrected JVM heap and step-level row-set sizing for transformations failing under grown input volumes
- 2Replaced silent partial-load behavior with explicit error handling and transactional boundaries
- 3Fixed connection pool and JDBC driver settings destabilized by a database platform upgrade
- 4Added job-level and step-level alerting so overnight failures page rather than wait for morning
- 5Built a runbook for the failure patterns most likely to recur before the migration completes
Outcome
The recurring nightly failures stopped, and half-loaded target tables were eliminated by making partial failures abort the load. The estate is stable enough that the migration can proceed on plan rather than under emergency pressure.
Technologies
Related Use Cases
Pentaho Carte Cluster Stability Remediation
Resolving Carte slave server instability where long-running clustered transformations hung, leaked memory, and left orphaned carte sessions.
Pentaho to Modern ELT Stack Migration
Planning and executing a staged move off Pentaho Data Integration onto a modern ELT stack, without a big-bang cutover of hundreds of transformations.
Pentaho Lineage and Documentation Recovery Assessment
Reconstructing lineage and documentation for an undocumented estate of legacy Kettle transformations before anyone attempts to change or replace them.
Docker Disk Exhaustion and Build Cache Support
Stopping recurring build agent and node outages caused by unpruned Docker build cache, dangling images, and orphaned volumes filling the filesystem.
AWS Lambda Cost and Right-Sizing Assessment
Measuring memory, duration, and concurrency across a large Lambda estate to right-size functions that were provisioned by guesswork.
AWS Lambda SQS Redrive Loop Remediation
Breaking a redrive loop where SQS messages were reprocessed indefinitely because the queue visibility timeout was shorter than the Lambda function timeout.
Need Expert Pentaho Support?
AceMQ's senior Pentaho engineers have handled this exact type of engagement before. Whether you need architectural guidance, hands-on remediation, or an ongoing managed partnership, we're ready to help.