#12545·xgboost

[jvm-packages] Unbounded native memory growth per DMatrix used with Booster.predict() on Linux (2.1.4 – 3.4.0)

Author: coder644Created Sep 8, 2026Updated Sep 8, 2026

Summary

A long-running JVM service that repeatedly performs new DMatrix(...) → Booster.predict() → DMatrix.dispose() shows linear, unbounded native memory (RSS) growth on Linux/x86_64. The growth is invisible to GC and to heap analysis, and the process is eventually OOM-killed. Reproduced from 2.1.4 through 3.4.0 (the latest xgboost4j release on Maven Central), on JDK 8 and JDK 21. The leak is per-DMatrix (not per call), scales with matrix size, and does not reproduce on macOS.

Environment

Package ml.dmlc:xgboost4j_2.12 — 2.1.4 (our production version) and 3.4.0 (latest on Maven Central)
JDK Temurin 8, Temurin 21 (Linux, both reproduce); Corretto 17 (macOS, clean)
OS Linux x86_64 — Kubernetes pods in production; eclipse-temurin Docker images + libgomp1 for the reproducer
Also tested macOS 15 (aarch64) — no leak in any scenario, same jars

GitHub v3.4.1 (released 2026-08-15) is not yet published to Maven Central, so 3.4.0 is the newest version we could test from Maven.

Minimal reproducer

LeakRepro.java (self-contained: trains a tiny quantile booster, then runs 100k new DMatrix + predict + dispose cycles, printing iteration,RSS_KB every 10k):

java
import java.io.BufferedReader;
import java.io.InputStreamReader;
import java.lang.management.ManagementFactory;
import java.util.HashMap;
import java.util.Map;

import ml.dmlc.xgboost4j.java.Booster;
import ml.dmlc.xgboost4j.java.DMatrix;
import ml.dmlc.xgboost4j.java.XGBoost;

/**
 * Minimal reproducer for a native memory leak in xgboost4j on linux/x86_64:
 * RSS grows linearly (~0.5-0.9 KB per cycle) as long as the JVM keeps doing
 * "new DMatrix -> Booster.predict -> DMatrix.dispose" cycles, and never plateaus.
 * Construct+dispose without predict is flat; a single reused DMatrix is flat.
 */
public class LeakRepro {

    public static void main(String[] args) throws Exception {
        final int rows = 1440, cols = 14;
        float[] data = new float[rows * cols];
        float[] labels = new float[rows];
        java.util.Random rnd = new java.util.Random(42);
        for (int i = 0; i < data.length; i++) data[i] = rnd.nextFloat() * 100f;
        for (int i = 0; i < rows; i++) labels[i] = rnd.nextFloat() * 100f;

        // Train a small booster (quantile objective, as used by our production models).
        Map<String, Object> params = new HashMap<>();
        params.put("objective", "reg:quantileerror");
        params.put("quantile_alpha", 0.5);
        params.put("max_depth", 4);
        params.put("nthread", 2);
        DMatrix train = new DMatrix(data, rows, cols);
        train.setLabel(labels);
        Booster booster = XGBoost.train(train, params, 20, new HashMap<String, DMatrix>(), null, null);
        train.dispose();

        // Warm up: native lib load, JIT, OpenMP thread pool.
        for (int i = 0; i < 1000; i++) {
            DMatrix dm = new DMatrix(data, rows, cols);
            booster.predict(dm, false, 0);
            dm.dispose();
        }

        // Measured loop: new DMatrix + predict + dispose.
        for (int i = 1; i <= 100_000; i++) {
            DMatrix dm = new DMatrix(data, rows, cols);
            try {
                booster.predict(dm, false, 0);
            } finally {
                dm.dispose();
            }
            if (i % 10_000 == 0) {
                System.gc(); // keep heap noise down; native leak is GC-independent
                Thread.sleep(50);
                System.out.println(i + "," + rssKb());
            }
        }
        booster.dispose();
    }

    /** Process RSS in KB (ps works on linux and macOS). */
    static long rssKb() throws Exception {
        String pid = ManagementFactory.getRuntimeMXBean().getName().split("@")[0];
        Process p = Runtime.getRuntime().exec(new String[]{"/bin/sh", "-c", "ps -o rss= -p " + pid});
        BufferedReader r = new BufferedReader(new InputStreamReader(p.getInputStream()));
        String line = r.readLine();
        p.waitFor();
        return Long.parseLong(line.trim());
    }
}

One-command run (verified end-to-end):

bash
docker run --rm --platform linux/amd64 -v "$PWD":/repro -w /repro --memory 1g eclipse-temurin:8-jdk sh -c '
  apt-get update -qq && apt-get install -y -qq libgomp1
  M=https://repo1.maven.org/maven2
  curl -sLO $M/ml/dmlc/xgboost4j_2.12/3.4.0/xgboost4j_2.12-3.4.0.jar
  curl -sLO $M/com/esotericsoftware/kryo/5.6.2/kryo-5.6.2.jar
  curl -sLO $M/org/objenesis/objenesis/3.4/objenesis-3.4.jar
  curl -sLO $M/com/esotericsoftware/minlog/1.3.1/minlog-1.3.1.jar
  curl -sLO $M/commons-logging/commons-logging/1.2/commons-logging-1.2.jar
  javac -cp xgboost4j_2.12-3.4.0.jar:kryo-5.6.2.jar LeakRepro.java
  java -Xmx256m -cp .:xgboost4j_2.12-3.4.0.jar:kryo-5.6.2.jar:objenesis-3.4.jar:minlog-1.3.1.jar:commons-logging-1.2.jar LeakRepro
'

Observed behavior (linux/amd64, RSS in KB)

Run @ 10k cycles @ 100k cycles Slope
3.4.0, Temurin 8 92,512 120,996 ≈ 0.29 KB per DMatrix
3.4.0, Temurin 21 103,624 144,156 ≈ 0.41 KB per DMatrix

Every 10k-sample block is monotonically increasing; heap stays flat (-Xmx256m, System.gc() between samples). Long runs (100k cycles) across {2.1.4, 3.4.0} × {glibc, jemalloc} all stay linear with no plateau.

What the leak is (and is not)

Measured with an extended probe built on the same loop (same JDK/jar matrix; per-DMatrix slopes from 20k-cycle runs unless noted):

Experiment Result Implication
new DMatrix + dispose, no predict flat construction path is clean (post-#10307)
one reused DMatrix, predict × 3 per cycle flat not per-call; a stable matrix does not grow
new DMatrix + predict × 1 + dispose 0.52 KB/DMatrix ≈ same as × 3 leak is per DMatrix that predict() touched, not per call
matrix size 1440×14 → 60×14 0.75 → 0.14 KB/DMatrix retained memory scales with matrix size
LD_PRELOAD jemalloc still leaks (0.64–0.70 KB/DMatrix) true leak, not glibc arena retention
MALLOC_ARENA_MAX=1 still leaks not arena count
jemalloc exit stats (MALLOC_CONF=stats_print:true) exit-time Allocated of a leaking run ≈ clean-baseline run (~6.1–6.2 MB in all cases) the blocks are still owned and are destroyed at process exit — held by some long-lived container, not lost pointers
rebuild the Booster (dispose + reload model) every 1000 cycles slope roughly halves (0.75 → 0.30 KB/DMatrix on 3.4.0) but stays linear two layers: part attached to the Booster instance, part process-global
8 concurrent predict threads no amplification (0.40 KB/DMatrix) not thread-local accumulation
JVM-side audit (Arthas vmtool instance counts after forced GC) matches expected object counts exactly nothing retained on the Java side
all scenarios on macOS (aarch64, Corretto 17) flat, incl. 3.4.0 Linux-specific

Relation to earlier issues

  • #10300 / #10307: the DMatrix construction-path leak, fixed in 2.1.0 by converting the DMatrix-related JNI functions to std::unique_ptr. The note in #10307 says "Not all functions are protected yet. This PR converts DMatrix-related functions." Our evidence points at the predict path: a DMatrix that is never passed to predict() does not leak, while one that is does.
  • #12284 / #12286: error-path leak in XGBoosterPredictFromDense() with base margin. Different function (we call Booster.predict(DMatrix) on a DMatrix handle, no base margin) and different trigger (our leak is on the success path — every call succeeds).

Impact

In production (K8s, linux/amd64) the service creates one DMatrix per dashboard query and disposes it right after prediction; RSS grows by multiple GB per day, GC is ineffective, and the memory is only reclaimed when the pod restarts. The only effective mitigations we found are lowering the DMatrix creation rate or restarting the process.

Happy to run follow-ups if useful (valgrind/ASAN builds, a source build of master, larger matrices, flame graphs from nativemem profiling, …).