All articles
2026-08-11•37 min read•JVM · Performance · Kubernetes

Analyzing Java heap dumps in production — a practical guide to memory leaks and retained objects

Capture and analyze Java heap dumps from Docker and Kubernetes workloads with a repeatable Eclipse MAT workflow that finds the object retaining the memory.

pC
Prashant Chaturvedi
Engineer

A Spring Boot order service has been healthy for months. Then its resident memory starts climbing a few megabytes an hour — not enough to alert on any single sample, but enough that the pod crosses its memory limit every two or three days and Kubernetes restarts it. The restart clears the heap, the graph resets to baseline, and the cycle repeats. Someone suggests bumping -Xmx. That is the moment to stop guessing and take a heap dump.

A heap dump is a snapshot of every live object in the JVM at one instant: every instance, its field values, its class, and the references connecting it to other objects. It is the one diagnostic artifact that can answer the question a memory incident is really asking — what is holding all this memory reachable, and why? This guide is for experienced Java and Spring Boot engineers who need a repeatable way to answer that question in production: how to capture a dump without hurting the service, how to walk it in Eclipse Memory Analyzer (MAT) until you find the retaining owner, and how to tell a real leak from a cache doing its job. The commands were verified on JDK 21; everything here applies to Java 17 and later, running locally, in Docker, or in Kubernetes.

What a heap dump is — and when it is the wrong tool

A heap dump (usually an .hprof file) contains the object graph at capture time: instances, arrays, field values including references, class metadata, and thread-related objects. What it does not contain is history. It shows which objects are alive and what refers to them; it does not show when they were allocated, which line of code created them, or what happened between two moments. Those gaps matter, because they decide which tool you should reach for first.

ArtifactWhat it actually captures
Heap dumpEvery reachable object and the reference graph between them, at one instant
Thread dumpThe stack of every thread at one instant — where code is blocked or running
GC logsAllocation and collection events over time — pause times, occupancy after GC
JFR recordingSampled events over time — allocations, contention, old-object samples
JVM metricsAggregated time series — heap used/committed, GC count and time, pool sizes

The incident pattern decides whether a heap dump is worth its cost:

  • Memory climbs and stays high across full GCs — the signature of retained objects. A heap dump is the right tool.
  • Memory climbs but drops after each major GC — normal allocation churn. Look at allocation rates and sizing, not dumps.
  • OutOfMemoryError: Java heap space already fired — a heap dump is usually essential; the JVM can write one automatically.
  • Container OOMKilled (exit code 137) with flat JVM heap metrics — the heap probably is not the problem. Resident set size (RSS) includes metaspace, thread stacks, direct buffers, and JIT code; the JVM diagnostics article covers non-heap accounting. A heap dump will show a healthy heap and waste your incident window.
  • The process is hanging, not growing — take thread dumps, not heap dumps.

Yes

No

Yes

No

Yes

No

Yes

No

Memory symptom

Heap grows and stays high

after full GC?

Heap dump: retained objects likely

Container OOMKilled,

JVM heap flat?

Non-heap suspect: metaspace, threads,

direct buffers, RSS headroom

Heap drops after

each GC?

Allocation churn: check GC logs

and allocation rate, not a dump

OOM thrown by the JVM?

Correlate metrics, logs,

and deploys before dumping

Yes

No

Yes

No

Yes

No

Yes

No

Memory symptom

Heap grows and stays high

after full GC?

Heap dump: retained objects likely

Container OOMKilled,

JVM heap flat?

Non-heap suspect: metaspace, threads,

direct buffers, RSS headroom

Heap drops after

each GC?

Allocation churn: check GC logs

and allocation rate, not a dump

OOM thrown by the JVM?

Correlate metrics, logs,

and deploys before dumping

Two limitations to internalize before you start. First, a heap dump is a snapshot: one dump shows what is retained, two dumps show what is growing. For slow leaks, the second dump is often the one that convicts. Second, a dump of a 8 GB heap is a multi-gigabyte file that freezes the JVM while it is written. Both of those costs have to be planned for, which is what the next two sections are about.

The mental model: reachability, dominators, retained heap

Heap-dump analysis is applied graph theory. The vocabulary is small but load-bearing, and MAT will not make sense without it.

Generations and regions. At the level you need for dump analysis, the heap is where objects live; the collector (G1 by default on Java 17+) divides it into regions and tracks which objects are reachable. Young objects die cheap and often; objects that survive enough collections get promoted to the old generation. A memory leak is old-generation occupancy that grows because objects stay reachable past their useful life — GC is working correctly, it simply cannot discard what your code still points at. That is why the dump matters: it contains the pointers.

Reachability and GC roots. An object is reachable if a chain of references leads to it from a GC root — the entry points the collector trusts: live threads and their stacks, static fields of loaded classes, JNI references, and a few JVM-internal holders. An object with no path from any GC root is garbage, regardless of how many other garbage objects point at it.

Reference strength. Not all references hold their targets equally:

  • Strong — a normal field or variable. The collector never reclaims a strongly reachable object.
  • Soft (SoftReference) — reclaimed only when memory gets tight. Caches built on soft references shrink under pressure instead of causing OOME.
  • Weak (WeakReference) — reclaimed at the next GC if no strong path exists. WeakHashMap keys and ThreadLocalMap entries work this way.
  • Phantom (PhantomReference) — used for post-mortem cleanup; the object is already gone, the reference just lets you run logic afterward.

Leak hunting is almost entirely about strong references. In MAT, when you trace paths to GC roots you normally exclude weak and soft references — an object held only by weak references is scheduled to die and is not your leak.

Shallow vs. retained heap. Shallow heap is the memory the object itself occupies — its header and fields, including reference slots. Retained heap is the memory that would be freed if this object were garbage-collected: its shallow size plus everything it exclusively keeps alive. A HashMap object itself is tiny (a few dozen bytes shallow) but can retain gigabytes of entries. This is the distinction the whole workflow rests on: the object at the top of the shallow-heap table is often a byte[] or char[], while the object at the top of the retained-heap table is the thing your code owns and can fix.

Dominator tree. Object A dominates object B if every path from the GC roots to B passes through A. The dominator tree reorganizes the object graph by ownership: a dominator’s children are the objects it alone keeps alive. This is what makes it the primary leak-hunting view — it answers “who is responsible for this memory” rather than “who points at it”, which in a dense graph can be many objects.

Incoming and outgoing references. Incoming references are who points at an object; outgoing are what it points at. You read incoming references to walk toward the owner, outgoing to walk toward the payload.

A concrete graph to hold all of this together:

AuditTrail.java
import java.time.Instant;
import java.util.ArrayList;
import java.util.List;
/** Leaky version: a static list that grows with every request and is never trimmed. */
public final class AuditTrail {
public record RequestRecord(String endpoint, int status, Instant at) {}
// GC root: the class itself. Everything in this list stays reachable forever.
private static final List<RequestRecord> RECORDS = new ArrayList<>();
public static void record(String endpoint, int status) {
RECORDS.add(new RequestRecord(endpoint, status, Instant.now()));
}
}

GC roots

static field RECORDS

elementData[]

AuditTrail class

(loaded by app ClassLoader)

ArrayList

Object[]

RequestRecord #1

RequestRecord #2

RequestRecord #N

Instant

String endpoint

GC roots

static field RECORDS

elementData[]

AuditTrail class

(loaded by app ClassLoader)

ArrayList

Object[]

RequestRecord #1

RequestRecord #2

RequestRecord #N

Instant

String endpoint

RequestRecord instances are individually small — a couple of fields, tens of bytes. But AuditTrail is a loaded class, which makes its static RECORDS field a GC root, which makes the ArrayList and its backing array reachable, which keeps every RequestRecord, Instant, and String alive. The dominator tree collapses this into one line: AuditTrail (more precisely, the ArrayList it owns) retains the whole graph. The “large object” would be the Object[] backing array; the leak owner is the class that refuses to let it go.

Heap dumps are production data — treat them accordingly

Before the capture commands, the part that should shape your runbook: a heap dump contains whatever was in memory. For a typical service that includes database rows, request and response bodies, deserialized JSON trees, HTTP headers including Authorization tokens, decrypted secrets that were decrypted in-process, session data, and connection pool credentials held in fields. Assume a heap dump is as sensitive as your production database, because it frequently is.

Handling rules that hold up in a review:

  • Restrict access like any other production-data extract: named individuals, time-boxed, ideally under the same controls as production shell access.
  • Encrypt at rest and in transit. Store dumps in encrypted volumes or object storage with bucket policies; do not attach raw dumps to tickets, Slack, or email.
  • Copy, don’t move, then clean up. Keep an explicit inventory of where the dump exists (pod filesystem, /tmp on a node, your laptop, object storage) and delete each copy when the analysis closes. Note that heap dumps on Linux are created mode 600 — a small mercy, not a policy.
  • Retention policy. Decide in advance: e.g., dump deleted after the incident report ships, or kept N days under the security team’s standard. Undecided means forever.
  • Document what was taken. Incident records should note the dump’s timestamp, pod/node, service version, and where the copy lives.

Security note — sanitizing a heap dump after the fact is not realistic; there is no reliable scrubber for the HPROF format. The controls that work are prevention (don’t let secrets sit in memory longer than needed) and containment (treat the file as sensitive end to end). MAT can run OQL to prove whether a token made it into the dump — useful for scoping an exposure assessment, not for cleaning it.

Two operational costs:

Disk. A heap dump is roughly proportional to the live heap contents, and can approach or exceed -Xmx. In my verification run, a 64 MB heap produced a 114 MB .hprof. Plan for free disk at least equal to the configured max heap wherever the dump lands — a pod filesystem that fills mid-dump gives you a truncated file and a worse day. Gzip helps: jcmd’s -gz option compressed a 23 MB dump to 4.6 MB in testing.

Latency. Writing a dump stops the world for the duration — the JVM performs a full GC first (unless you pass -all to jcmd), then serializes the heap to disk. On multi-gigabyte heaps this is a multi-second to multi-minute pause. On a live traffic-serving pod, take the pod out of rotation first or accept the stall deliberately.

Production readiness checklist for the dump artifact itself — the engineering-side pre-flight list (flags, volumes, tooling) is in the Checklists section:

  • Dump access restricted to named on-call engineers, under the same controls as production shell access.
  • Dump storage encrypted at rest; the transfer path (object storage, shared volume) agreed with security in advance.
  • Retention and deletion policy written down — including who deletes each copy and when.
  • Free disk on the dump target at least equal to the configured max heap.
  • The team knows a dump freezes the JVM for the write, and has decided who pulls the pod from rotation first.

Capturing a heap dump

Automatic dumps on OutOfMemoryError

This should already be on. The cost is zero until it fires, and the day it fires is the day you needed it most:

JVM flags
-XX:+HeapDumpOnOutOfMemoryError
-XX:HeapDumpPath=/var/dumps

HeapDumpPath names either a directory — the JVM writes java_pid<pid>.hprof inside it — or an exact file path. The dump happens before the JVM exits, which creates the trap that catches most container deployments: this flag triggers on a JVM OutOfMemoryError, which is thrown when GC cannot reclaim enough heap. A Kubernetes pod killed by the cgroup OOM killer for exceeding its memory limit (RSS) never throws a Java OOME — the kernel just SIGKILLs it, no dump, no exception. Two consequences: set -Xmx comfortably below the container limit so Java heap exhaustion (if it comes) fires first, and remember that OOMKill diagnostics need non-heap accounting, not a heap dump. Related flags worth knowing: -XX:+CrashOnOutOfMemoryError writes a crash log and aborts (dump still happens first if enabled), and -XX:OnOutOfMemoryError='...' runs a command — sometimes used to trigger a thread dump or alert alongside the heap dump.

Manual capture with jcmd

jcmd is the right tool on any modern JDK — it talks to the JVM’s attach listener and is stable where jmap is not. Run it as the same OS user as the Java process:

terminal
# Find the pid (in a container the JVM is usually pid 1)
jcmd -l
# Live objects only (runs a full GC first) — what you want for leak hunting
jcmd <pid> GC.heap_dump /var/dumps/manual.hprof
# Include unreachable objects — skips the pre-dump GC; larger file, noisier dump
jcmd <pid> GC.heap_dump -all /var/dumps/manual-all.hprof
# Compress while writing (gzip level 1–9; 1 is fast and usually enough)
jcmd <pid> GC.heap_dump -gz=1 /var/dumps/manual.hprof.gz
# Overwrite an existing file (dumps refuse to clobber by default)
jcmd <pid> GC.heap_dump -overwrite /var/dumps/manual.hprof

The full GC before a normal dump is a feature: unreachable objects are noise for leak analysis, so a “live objects” dump is smaller and cleaner. Use -all when you specifically want to see garbage-in-flight — for example, to understand allocation pressure right before an OOME.

Two cheap alternatives before committing to a full dump: jcmd <pid> GC.class_histogram prints a live-object histogram to stdout (no file, much cheaper, enough to confirm which classes are accumulating), and jcmd <pid> VM.flags records the exact flags the process is running.

jmap: works, but prefer jcmd

terminal
jmap -dump:live,format=b,file=/var/dumps/jmap.hprof <pid>

jmap still works on JDK 17/21 and produces the same format. The caveats: it sits in the JDK’s unsupported-tools category, the live option forces a full GC pause the same way, it refuses to overwrite an existing file (verified on JDK 21: Unable to create ... File exists), it has to run as the same user with the attach mechanism available, and the -F force flag that old runbooks reference was removed in JDK 9 — if a playbook says jmap -F, the playbook predates your JDK. Reach for jmap when jcmd is missing; reach for jhsdb jmap --binaryheap when you have a core dump rather than a live process.

Docker

Inside the container, nothing changes — assuming the tools and filesystem allow it:

terminal — on the Docker host
docker exec <container> jcmd 1 GC.heap_dump /tmp/heap.hprof
docker cp <container>:/tmp/heap.hprof ./heap.hprof
docker exec <container> rm /tmp/heap.hprof

Three things bite in practice. First, jcmd must exist in the image: JRE-slim and distroless images ship java but not the JDK tools — jlink-built runtimes only include jcmd if the jdk.jcmd module was added. Check before the incident (docker exec <container> which jcmd), and if it’s absent either rebuild with the tools or copy in a matching jcmd binary — attach requires the tool and JVM to be version-compatible. Second, /tmp in a container is the container’s writable layer: it’s fine for a fast docker cp, but the dump dies with the container, and a large dump can exhaust the container filesystem quota. Third, jcmd must run under the same UID as the JVM process — a docker exec -u mismatch produces cryptic attach failures.

Kubernetes

terminal
# App container: pid is almost always 1
kubectl exec <pod> -c <container> -- jcmd 1 GC.heap_dump /tmp/heap.hprof
# Copy it out — requires tar in the image; note the leading slash handling
kubectl cp <pod>:/tmp/heap.hprof ./heap.hprof -c <container>
# Clean up the pod's filesystem
kubectl exec <pod> -c <container> -- rm /tmp/heap.hprof

The real problem in Kubernetes is where the dump can land. The container filesystem is ephemeral: if the pod restarts — which, in a memory incident, it is about to do — the dump is gone. The options, in order of reliability:

  1. A mounted volume. An emptyDir survives container restarts within the pod but not pod rescheduling; a PersistentVolumeClaim survives everything and can be read later from another pod. For services where heap dumps are routine, a small PVC mounted at /var/dumps is the boring, correct answer.
  2. Copy before restart. kubectl cp works as long as the container stays up and the image contains tar. If tar is absent, kubectl exec <pod> -- cat /tmp/heap.hprof > heap.hprof streams it out — slower, works everywhere.
  3. Ephemeral debug container. kubectl debug -it <pod> --image=<jdk image> --target=<container> adds a tools container that shares the target’s process namespace. Useful when the app image lacks jcmd; be aware the attach handshake relies on files under the target’s /tmp, so it works cleanly only if the containers share a /tmp mount — otherwise run jcmd through /proc/<pid>/root paths or accept kubectl exec into a tools-bearing image as the primary path.

A minimal deployment fragment that makes automatic dumps survivable — dump dir on a persistent volume, JDK_JAVA_OPTIONS to pass flags (JDK_JAVA_OPTIONS is read directly by the java launcher on JDK 9+; JAVA_TOOL_OPTIONS also works but leaks into every JDK tool invocation and prints a Picked up line to stderr):

deployment.yaml (excerpt)
spec:
containers:
- name: orders
image: registry.example.com/orders:1.4.2
env:
- name: JDK_JAVA_OPTIONS
value: >-
-Xmx1536m
-XX:+HeapDumpOnOutOfMemoryError
-XX:HeapDumpPath=/var/dumps
volumeMounts:
- name: dumps
mountPath: /var/dumps
resources:
limits:
memory: 2Gi # heap 1.5 Gi + headroom for non-heap
volumes:
- name: dumps
persistentVolumeClaim:
claimName: orders-dumps

Warning — kubectl exec ... jcmd GC.heap_dump and jmap -dump both freeze the JVM for the duration of the write. On a pod serving traffic, pull it from the Service endpoints first (kubectl label pod <pod> ... or let the readiness probe fail naturally) or schedule the capture during low traffic. Never run a dump on a pod that is one GC away from OOMKill without accepting it may die mid-dump.

What to collect alongside the dump

A heap dump without context is a tarball of objects. While the JVM is still alive, capture:

  • Pod/node state — kubectl describe pod for events, restart counts, and OOMKilled reasons; kubectl get pod -o yaml for limits and the running image digest.
  • JVM facts — jcmd <pid> VM.flags and VM.command_line for the real flags, GC.heap_info for pool sizes at capture time.
  • Thread dump — jcmd <pid> Thread.print (or Thread.dump_to_file for virtual-thread workloads), in case the retention is executor-bound.
  • GC logs and metrics — the heap-after-GC trend from Prometheus/Grafana is what proves “leak” rather than “spike”.
  • Application logs around the climb, plus the request-rate timeline — a leak’s slope often tracks traffic or a specific endpoint’s traffic.
  • Deploy and config changes — the single most productive correlation: what shipped when the slope changed?

A repeatable analysis workflow in Eclipse MAT

Eclipse Memory Analyzer (MAT) is the standard open-source tool — free, scriptable via ParseHeapDump.sh for headless runs, and built around exactly the dominator analysis this workflow needs. VisualVM opens dumps for quick histogram browsing, and the commercial tools (JProfiler, YourKit) add allocation tracking if you need to know where objects were created rather than just who holds them — but for “find the retaining owner”, MAT is the reference. JFR’s jdk.OldObjectSample event deserves a mention as the lightweight alternative: it samples objects that survive GC, which can flag a leak without ever taking a dump.

One setup step matters before you open anything: MAT is itself a Java application, and parsing a large dump needs heap roughly on the order of the live objects in it. Edit MemoryAnalyzer.ini in the MAT install directory and set -Xmx to something comfortably larger than the dump file (for a 4 GB dump, -Xmx12g is a reasonable start) — the parsing step fails or crawls otherwise. MAT writes index files next to the .hprof on first open, so the directory must be writable; keep them — they make reopens fast.

The workflow is ordered so that each stage answers one question and feeds the next.

Yes

No / unbounded

Confirm pattern in metrics

Open dump in MAT

Leak Suspects report

Histogram: who is numerous?

Dominator tree: who retains the most?

Path to GC roots: who anchors it?

Incoming references: identify owner object

Retention expected and bounded?

Size or tune it — not a leak

Identify code that adds but never removes

Validate against metrics and code

Second dump confirms growth

Yes

No / unbounded

Confirm pattern in metrics

Open dump in MAT

Leak Suspects report

Histogram: who is numerous?

Dominator tree: who retains the most?

Path to GC roots: who anchors it?

Incoming references: identify owner object

Retention expected and bounded?

Size or tune it — not a leak

Identify code that adds but never removes

Validate against metrics and code

Second dump confirms growth

Stage 1: Confirm the pattern before you open anything

Question: does this even look like retained memory?

Do: in Grafana/Prometheus, look at heap-after-GC (or old-generation occupancy) over the incident window. A leak shows a rising floor — each GC drops the line, but to a higher point every time. A workload spike shows a higher plateau that stabilizes.

Trap: dumping because “memory is high” without checking whether it stays high across GCs. Full heap at a traffic peak is a sizing question, not a leak.

Stage 2: Open the dump safely

Question: can MAT even parse this file?

Do: File → Open Heap Dump. Watch for the dump landing writable-adjacent (index files) and MAT’s own -Xmx. For dumps from jcmd -gz, MAT reads .hprof.gz directly. If parsing fails with heap errors, that is MAT’s heap, not the dump’s.

Trap: assuming a truncated dump (disk-full or killed mid-write) is valid. MAT usually fails fast on truncated files, but a partially-written tail can parse and produce quietly wrong object counts — check file size against expectations first.

Stage 3: Leak Suspects — read it, don’t believe it

Question: what does the automated analysis think dominates the heap?

Do: let the Leak Suspects report run on open (or Overview → Leak Suspects). It applies dominator analysis heuristics and names one or more “problem suspects” — typically a thread, classloader, or component retaining a large percentage.

Reading it: a suspect retaining 70% of the heap is worth chasing; the report’s “suspect” label is still a heuristic guess, not a verdict. The report names what accumulates, rarely why — a “big ConcurrentHashMap” suspect tells you where to start digging, not what to fix.

Trap: reporting “MAT found the leak” from this screen alone. A large intentional structure — a connection pool, a well-sized cache — will always top this report in a healthy application, so the suspect label is a lead, not a verdict.

Stage 4: The histogram — who is numerous

Question: which classes dominate by count and by total size?

Do: Toolbar → Histogram. Sort by Objects (count) and by Shallow Heap. Filter by regex on your package — in.o612.eng.orders.* — to separate your classes from the JDK’s.

Reading it: the top of any real application’s shallow-heap list is boringly consistent — byte[], char[], String, Object[], map nodes. What you’re hunting is an anomaly: a domain class (OrderContext, RequestRecord) with a count that resembles your request volume rather than your live-request count, or a collection class with an implausible instance count. In the verification demo for this article, GC.class_histogram showed AuditTrail$RequestRecord at rank 2 with exactly 200,000 instances — matching the 200,000 requests the process had handled. Counts that track lifetime totals, not current load, are the leak signature.

Trap: chasing the top shallow-heap row. byte[] at 40% of the heap is not the leak — it is the payload. The question is what retains those arrays.

Stage 5: Dominator tree — who retains it

Question: which single objects, if GC’d, would free the most memory?

Do: Toolbar → Dominator Tree. It lists objects ordered by retained heap — the biggest owners first.

Reading it: expand the top dominator. A healthy app’s top dominators are the intentional big structures: caches, connection pools, ForkJoinPool, the heap itself. A leak shows an owner that shouldn’t be big — a service singleton holding a ConcurrentHashMap that retains 1.2 GB of request contexts, say. Note whether retained size concentrates in one dominator (a leak or one oversized structure) or spreads thin (death by a thousand caches — a sizing problem).

Trap: stopping at the first big dominator when it’s the application ClassLoader or a thread — those dominate by structure, not by fault. Keep expanding until the dominator is something your code owns.

Stage 6: Path to GC roots — what anchors the object

Question: which strong-reference chain keeps this alive?

Do: right-click a suspect object (or a group in the dominator tree) → Path to GC Roots → exclude all phantom/weak/soft etc. references. For a class of leaked objects rather than one instance, use Merge Shortest Paths to GC Roots on a selection to see the common anchor.

Reading it: the path bottom-up reads as the retention story: OrderContext ← ConcurrentHashMap$Node ← table[] ← inFlight field ← OrderService ← Spring singleton ← Class ← GC root. Every arrow is a reference your code (or the framework) maintains. The leak is usually obvious in the middle of the chain — a field you forgot to clear, a map that never removes.

Trap: forgetting to exclude weak/soft references — a path through WeakReference or ThreadLocalMap$Entry proves non-retention, and a strong-reference claim built on it is wrong. Conversely, ThreadLocalMap$Entry rows with referent = null (the weak key already collected, the strong value stranded) are a real ThreadLocal leak signature.

Stage 7: Incoming references — name the owner

Question: which application object holds the retaining reference?

Do: right-click → List objects → with incoming references — and walk toward the root until the referrer is a class in your package, not a JDK collection. The dominator view often gets you here faster; incoming references confirm the field name (MAT shows fieldname on the reference edge when available).

Reading it: you want a sentence like “ConcurrentHashMap field inFlight on singleton OrderService retains 41,209 OrderContext instances.” Class-level suspicion becomes field-level evidence here — write that sentence down, it becomes the incident report’s core.

Trap: attributing the leak to the collection (ConcurrentHashMap, ArrayList) rather than the field and class that own it. The fix always lives with the owner.

Stage 8: Classify the retention

Question: is this retention accidental or intentional, bounded or unbounded?

  • Expected and bounded — a sized cache, a pool at its configured max, a bounded queue. Not a leak; tune or resize.
  • Expected but unbounded — a cache with no eviction, an in-memory queue with no cap. The design is the bug.
  • Accidental — listeners never deregistered, ThreadLocals never removed, error paths that skip cleanup, per-request state on a singleton. Classic leaks.
  • Transitional — objects mid-flight (active requests, in-progress batch). They show as big in any busy-app dump; correlate with request rate before accusing them.

Trap: calling every large dominator a leak. A bounded 2 GB cache doing exactly its job is a config discussion, not a bug.

Stage 9: Validate against the application

Question: does the owning code actually have a path that retains?

Do: read the code behind the identified field. There must exist an add without a matching remove on some path — and that path must be reachable under your incident’s traffic pattern. Then check the metrics: does the retained-object count growth correlate with a specific endpoint, consumer group, or schedule?

Trap: fixing the first plausible-looking path in code without confirming it’s the path production actually hits. Two dumps (next stage) or a targeted GC.class_histogram before/after the suspected endpoint is cheap confirmation.

Stage 10: Compare a second dump

Question: what is growing, not just big?

Do: capture a second dump N hours later under comparable traffic, open both in MAT, and compare histograms (Histogram view → the “compare” toolbar button → select the other dump) to get per-class instance and size deltas.

Reading it: the leak class shows a delta proportional to elapsed request count; stable caches and pools show flat deltas. This is the cleanest leak-vs-plateau discriminator short of fixing it and watching.

Trap: comparing dumps taken at different traffic levels without normalizing — a busy dump naturally holds more of everything. Compare deltas per unit of work (requests handled between dumps), not raw deltas.

Six leak patterns in Java, in code

Each of these compiles and runs (Java 17+; the Caffeine example needs com.github.ben-manes.caffeine:caffeine:3.2.x). The MAT observations describe what the verified demo code actually produces.

A. Static collection that only grows

The AuditTrail class from the mental-model section is the whole pattern: a static field is reachable for the life of its class loader — which, for application classes, is the life of the JVM — so a static collection is an implicit GC root for everything ever added.

LeakDemo.java — verified: OOME with heap dump on JDK 21
public final class LeakDemo {
public static void main(String[] args) {
while (true) {
AuditTrail.record("/api/orders", 200);
}
}
}
terminal — verified output
java -Xmx64m -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=./dumps LeakDemo
# java.lang.OutOfMemoryError: Java heap space
# Dumping heap to ./dumps/java_pid11930.hprof ...
# Heap dump file created [114242072 bytes]

In MAT, the dominator tree shows AuditTrail (as a class) dominating an ArrayList whose Object[] holds every RequestRecord; Path to GC Roots ends at system class → AuditTrail. The fix is not always “bound the list” — ask why per-request state is in memory at all; a debug ring buffer has a legitimate bounded form:

AuditTrailFixed.java
import java.time.Instant;
import java.util.ArrayDeque;
import java.util.Deque;
public final class AuditTrailFixed {
public record RequestRecord(String endpoint, int status, Instant at) {}
private static final int MAX_RECORDS = 10_000;
private static final Deque<RequestRecord> RECORDS = new ArrayDeque<>(MAX_RECORDS);
public static synchronized void record(String endpoint, int status) {
if (RECORDS.size() == MAX_RECORDS) {
RECORDS.pollFirst();
}
RECORDS.addLast(new RequestRecord(endpoint, status, Instant.now()));
}
}

B. Cache with no bound or eviction

PricingService.java — the leak shape
import java.math.BigDecimal;
import java.util.Map;
import java.util.concurrent.ConcurrentHashMap;
public final class PricingService {
public record PriceQuote(String sku, BigDecimal price) {}
private final Map<String, PriceQuote> cache = new ConcurrentHashMap<>();
public PriceQuote quote(String sku) {
return cache.computeIfAbsent(sku, this::loadFromPriceBook);
}
private PriceQuote loadFromPriceBook(String sku) {
return new PriceQuote(sku, BigDecimal.valueOf(sku.hashCode() & 0xFFFF));
}
}

“It is a cache” explains the retention; it does not justify it. A cache’s job is to trade memory for latency within a budget. With an unbounded map, the budget is -Xmx, and eviction is performed by the OOM killer. In MAT this looks identical to a deliberate cache — a service bean dominating a ConcurrentHashMap — which is why stage 8 (expected vs. accidental) has to be answered from the code, not the dump. The tell is the key set: SKU keys with cardinality equal to every SKU ever seen, or worse, keys built from request parameters (sku + ":" + region + ":" + customerTier) whose cardinality is effectively unbounded.

PricingServiceFixed.java
import com.github.benmanes.caffeine.cache.Cache;
import com.github.benmanes.caffeine.cache.Caffeine;
import java.math.BigDecimal;
import java.time.Duration;
public final class PricingServiceFixed {
public record PriceQuote(String sku, BigDecimal price) {}
private final Cache<String, PriceQuote> cache = Caffeine.newBuilder()
.maximumSize(50_000)
.expireAfterWrite(Duration.ofMinutes(30))
.build();
public PriceQuote quote(String sku) {
return cache.get(sku, this::loadFromPriceBook);
}
private PriceQuote loadFromPriceBook(String sku) {
return new PriceQuote(sku, BigDecimal.valueOf(sku.hashCode() & 0xFFFF));
}
}

Sizing trade-offs: maximumSize caps entry count — safe against cardinality, but you must estimate bytes-per-entry to know what 50,000 entries costs (and maximumWeight with a weigher exists for variable-size values). expireAfterWrite/expireAfterAccess cap staleness and age-out cold entries, but an expiry-only cache under a burst of distinct keys can still outgrow its budget between maintenance passes. Set both when memory pressure is the concern, and instrument cache.stats() — a cache with a 40% hit rate is memory spent on nothing.

C. ThreadLocal never removed on a pooled thread

RequestContext.java — the leak shape
import java.util.concurrent.ExecutorService;
public final class RequestContext {
private static final ThreadLocal<byte[]> REQUEST_PAYLOAD = new ThreadLocal<>();
public static void handle(ExecutorService pool, byte[] payload) {
pool.submit(() -> {
REQUEST_PAYLOAD.set(payload);
process(payload);
// No remove(): the payload stays reachable via the worker thread.
});
}
private static void process(byte[] payload) { /* business logic */ }
}

On a request-per-thread model a stray ThreadLocal dies with its thread. On a pool — servlet container threads, @Async executors, ForkJoinPool.commonPool — the worker thread outlives every task, so each worker’s ThreadLocalMap retains the last payload it processed. Leak rate is bounded by pool size (200 threads retain at most ~200 payloads), which makes this a slow, confusing leak rather than a cliff — and it makes heap per worker the metric to watch.

RequestContextFixed.java — the corrected handle()
public static void handle(ExecutorService pool, byte[] payload) {
pool.submit(() -> {
try {
REQUEST_PAYLOAD.set(payload);
process(payload);
} finally {
REQUEST_PAYLOAD.remove();
}
});
}

In the dump, payloads appear retained by java.lang.Thread via threadLocals → ThreadLocalMap → Entry[] → value. Entries whose referent is null mean the ThreadLocal key itself was GC’d (it is held weakly) while the value stays strongly reachable — those entries are uncollectable and unreachable through the API; only remove() on a live key or thread death cleans them. The same shape applies to SLF4J’s MDC, which is a ThreadLocal under the hood — an MDC.put without a matching MDC.clear() leaks strings on pooled workers.

D. Listener registered on a long-lived publisher

InventoryEvents.java — publisher lives as long as the app
import java.util.List;
import java.util.concurrent.CopyOnWriteArrayList;
import java.util.function.Consumer;
public final class InventoryEvents {
public record StockChanged(String sku, int newLevel) {}
private final List<Consumer<StockChanged>> listeners = new CopyOnWriteArrayList<>();
public void subscribe(Consumer<StockChanged> listener) { listeners.add(listener); }
public void publish(StockChanged event) { listeners.forEach(l -> l.accept(event)); }
}
LowStockAlerter.java — subscriber with real state
public final class LowStockAlerter {
private final byte[] dashboardState = new byte[256 * 1024];
public LowStockAlerter(InventoryEvents events) {
// The lambda captures `this`; the publisher retains every alerter created.
events.subscribe(event -> {
if (event.newLevel() < 5) {
notifyOperator(event.sku());
}
});
}
private void notifyOperator(String sku) { /* paging logic */ }
}

Every LowStockAlerter created — per request, per session, per reconnect — is captured by the lambda, which lives in the publisher’s listener list, which lives as long as the publisher. In the dump: one publisher bean dominating a CopyOnWriteArrayList full of LowStockAlerter$$Lambda entries, each retaining an alerter and its dashboardState. The fix is making deregistration part of the subscriber’s lifecycle:

InventoryEventsFixed.java — subscribe() returns a handle
public interface Subscription extends AutoCloseable {
@Override void close();
}
public Subscription subscribe(Consumer<StockChanged> listener) {
listeners.add(listener);
return () -> listeners.remove(listener);
}
LowStockAlerterFixed.java — closed when the owning session ends
public final class LowStockAlerterFixed implements AutoCloseable {
private final InventoryEventsFixed.Subscription subscription;
public LowStockAlerterFixed(InventoryEventsFixed events) {
this.subscription = events.subscribe(event -> { /* ... */ });
}
@Override
public void close() {
subscription.close();
}
}

In Spring, @PreDestroy on the subscriber (or wiring into the webSocket/connection lifecycle) is the usual place to call close(). WeakReference listeners are the lazy alternative — they clean themselves up, but silently drop events while a subscriber is dying-but-alive, and they obscure exactly the bug the explicit handle prevents.

E. ClassLoader leak on plugin reload

The subtlest of the six: every Class is strongly reachable from its defining ClassLoader, and a ClassLoader is reachable from every class it loaded — so retaining one instance of one plugin-loaded class pins the entire loader and every class in it.

Plugin.java — compiled to plugins/, off the app classpath
public final class Plugin implements Runnable {
private final byte[] moduleState = new byte[1024 * 1024];
@Override
public void run() {
// Registry lives in the app ClassLoader and outlives this loader.
PluginRegistry.register(this);
}
}
PluginRegistry.java — owned by the application ClassLoader
import java.util.List;
import java.util.concurrent.CopyOnWriteArrayList;
public final class PluginRegistry {
private static final List<Object> ACTIVE = new CopyOnWriteArrayList<>();
public static void register(Object plugin) {
ACTIVE.add(plugin);
}
public static int count() {
return ACTIVE.size();
}
}
PluginHost.java — a fresh loader per reload
import java.io.File;
import java.net.URL;
import java.net.URLClassLoader;
public final class PluginHost {
private final File pluginDir;
public PluginHost(File pluginDir) {
this.pluginDir = pluginDir;
}
public void reload() throws Exception {
URL[] urls = { pluginDir.toURI().toURL() };
try (URLClassLoader loader =
new URLClassLoader(urls, PluginHost.class.getClassLoader())) {
Class<?> pluginClass = Class.forName("Plugin", true, loader);
Runnable plugin = (Runnable) pluginClass.getDeclaredConstructor().newInstance();
plugin.run();
} // loader is closed — but not collectable: PluginRegistry holds the Plugin
}
}

Verified locally: five reload() calls left PluginRegistry.count() == 5 — five plugin instances, five Plugin classes, five URLClassLoaders, all pinned by a registry the application class loader owns. The same mechanism fires in real life through hot-reload contexts (devtools-style restart loaders), app-server redeploys, JDBC drivers that DriverManager registers globally, threads spawned by plugin code that never stop, and ThreadLocals set by old-loader code.

In MAT, the giveaway is duplicated classes: search the histogram for the plugin class name and you will see the same class name once per loader, each with a different ClassLoader instance. Path to GC Roots on a stale loader leads back through the registry, the JDBC driver table, or a live thread — whichever holds the pin. The fix is lifecycle discipline on unload: deregister from global registries, stop spawned threads, clear ThreadLocals, then verify in a dump that the old loader is collectable.

F. Request/response payloads kept “for debugging”

PayloadRecorder.java — the leak shape
import java.util.List;
import java.util.concurrent.CopyOnWriteArrayList;
public final class PayloadRecorder {
private static final List<byte[]> RECENT_REQUESTS = new CopyOnWriteArrayList<>();
public static void capture(byte[] requestBody) {
RECENT_REQUESTS.add(requestBody); // never trimmed
}
}

Payload retention has a legitimate variant — bounded capture is genuinely useful — so the discipline is double-bounding: cap entries and cap bytes per entry, since one 50 MB upload in a “recent requests” buffer is worse than a thousand small ones.

PayloadRecorderFixed.java
import java.util.ArrayDeque;
import java.util.Arrays;
import java.util.Deque;
public final class PayloadRecorderFixed {
private static final int MAX_ENTRIES = 128;
private static final int MAX_BODY_BYTES = 64 * 1024;
private static final Deque<byte[]> RECENT_REQUESTS = new ArrayDeque<>();
public static synchronized void capture(byte[] requestBody) {
if (RECENT_REQUESTS.size() == MAX_ENTRIES) RECENT_REQUESTS.pollFirst();
byte[] copy = Arrays.copyOf(requestBody, Math.min(requestBody.length, MAX_BODY_BYTES));
RECENT_REQUESTS.addLast(copy);
}
}

Distinguishing spike from leak for payloads: buffers held only during request processing disappear after GC and don’t show in a live-only dump; payloads in a dump were reachable at capture time, so a live dump full of byte[] means either retention or a genuine in-flight burst — the request-rate metric at capture time decides which. The same pattern family covers retained JsonNode trees, response bodies buffered for logging, and reactive DataBuffer chains — in reactive stacks look for buffers retained by Netty/reactor internals rather than application fields.

Retention patterns specific to Spring Boot services

The same handful of shapes produce most real-world Spring Boot leaks:

PatternRetention mechanismMAT signatureAction
Singleton bean holds per-request stateField on a singleton written per requestService bean dominates request-shaped objectsRemove the field; pass state through the call or a request-scoped bean
Unbounded cacheConcurrentHashMap/guava cache, no max size/weightMap dominates; key cardinality ≈ all keys ever seenCaffeine with maximumSize/maximumWeight + expiry; instrument hit rate
Scheduled job accumulates results@Scheduled task appends to a field for “later”Job bean dominates a growing collectionStream results out; keep only counters/aggregates
WebSocket/session map never cleanedMap<sessionId, …> with no @OnClose/disconnect removalEndpoint bean dominates sessions + per-session buffersRemove on close/transport error; add TTL sweeps for orphaned entries
Async executor + ThreadLocal/MDCPooled workers keep last task’s contextThread objects dominate payloads via ThreadLocalMaptry/finally remove; MDC clear decorators on the executor
Kafka listener retains recordsBatch/buffer fields on singleton consumer; dead-letter lists in memoryConsumer bean dominates ConsumerRecord/payload treesProcess and release; bound in-memory retry/dead-letter storage
HTTP client buffersResponse bodies cached/buffered for retry or loggingClient internals dominate byte[]/Netty buffersStream large bodies; cap logged/buffered body size
JPA persistence context in batchFirst-level cache accumulates every loaded entityPersistenceContext/Hibernate StatefulPersistenceContext dominates entitiesEntityManager.clear() per chunk; stateless sessions; smaller chunks
JSON tree deserializationWhole-document JsonNode for huge payloadsObjectNode/ArrayNode trees dominateStreaming parser (JsonParser) or targeted extraction
Metrics cardinalityLabel value per user/order/tenant → a meter per seriesMicrometer MeterRegistry dominates Meter instancesBound label values; drop high-cardinality tags at the source
In-memory queue/retry bufferBlockingQueue/lists as durable-ish buffersQueue dominates messages; grows when consumer stallsBound the queue + backpressure; move retries to durable storage

Every row reduces to the same rule: something with application lifetime holds something with request lifetime. The dump finds the holder; the fix restores the lifetime boundary.

The Kubernetes incident workflow, end to end

MAT workstationOn-call engineerKubernetesApp podPrometheus/GrafanaMAT workstationOn-call engineerKubernetesApp podPrometheus/GrafanaJVM pauses for the dump writeHeap-after-GC floor rises over hoursAlert: memory near limit / restart rate upkubectl exec: jcmd GC.class_histogramDomain classes accumulating → confirms suspicionkubectl exec: jcmd GC.heap_dump /var/dumps/heap.hprofkubectl cp pod:/var/dumps/heap.hprof ./Later — OOMKilled and restart, dump safe on PVC/local copyOpen dump → histogram → dominator tree → GC rootsRetaining owner identifiedCorrelate with deploys, config, traffic — validate in codeDeploy bounded/evicting fix behind flag or canaryPost-release: heap floor flat, no restartsSecond dump days later: delta confirms the fix
MAT workstationOn-call engineerKubernetesApp podPrometheus/GrafanaMAT workstationOn-call engineerKubernetesApp podPrometheus/GrafanaJVM pauses for the dump writeHeap-after-GC floor rises over hoursAlert: memory near limit / restart rate upkubectl exec: jcmd GC.class_histogramDomain classes accumulating → confirms suspicionkubectl exec: jcmd GC.heap_dump /var/dumps/heap.hprofkubectl cp pod:/var/dumps/heap.hprof ./Later — OOMKilled and restart, dump safe on PVC/local copyOpen dump → histogram → dominator tree → GC rootsRetaining owner identifiedCorrelate with deploys, config, traffic — validate in codeDeploy bounded/evicting fix behind flag or canaryPost-release: heap floor flat, no restartsSecond dump days later: delta confirms the fix

The order matters more than the tools: confirm the pattern in metrics before touching the pod, take the cheap histogram before the expensive dump, copy the dump out before the pod can restart, and close the loop with a second dump after the fix ships.

Do not jump to conclusions

  • High shallow heap ≠ the owner. byte[] and String will top every histogram you ever open. They are the cargo; find the ship.
  • Big retained heap still needs the path. A dominator tells you what would be freed, not why it is held. Without the GC-root path you cannot tell a leak from a cache.
  • Intentional structures still need bounds. Caches, pools, and queues retaining memory is their design — the question is never “why does it retain” but “what bounds it, and is the bound right”.
  • One dump is a snapshot. A single dump shows structure; growth needs a second dump, a GC.class_histogram time series, or JFR old-object samples.
  • Rising heap has innocent causes. Delayed GC (G1 tolerates a dirty old gen), a genuine traffic increase, larger payloads, warm-up after deploy, humongous-object fragmentation. Match the slope to a cause before writing a fix.
  • Container RSS ≠ Java heap. Metaspace, thread stacks, JIT code, direct buffers, and malloc’d native structures all count toward the cgroup limit and appear in no heap dump. OOMKilled with a flat heap graph is a native-memory incident — reach for NativeMemoryTracking, not MAT.
  • OOME has non-heap causes. Metaspace, unable to create native thread, Direct buffer memory, Requested array size exceeds VM limit — each is a different subsystem and a different fix; only Java heap space points at what this article covers.
  • MAT parses reachability, not intent. Leak Suspects names structures; it cannot know your cache TTL was deliberate or your queue was supposed to drain. Evidence identifies the owner; only context convicts.

Matching the evidence to the question

ArtifactBest question it answersStrengthsLimitationsCollect it when
Heap dumpWhat is reachable and what retains itComplete object graph; field-level attributionSnapshot only; pauses JVM; huge and sensitiveHeap stays high across GCs; JVM OOME fired
Thread dumpWhat is each thread doing nowCheap, instant, repeatableNo memory contents; per-instantHangs, thread-pool exhaustion, suspected executor retention
GC logsIs occupancy after each collection growing; how long pauses runEvent-level history of every GC; cheap to keep onNo object identity; verbose without toolingFirst — always; it decides whether a dump is warranted
JFR recordingWhich code allocates, which objects surviveProduction-safe continuous profiling; OldObjectSampleSampled; needs recording enabled before the incidentChronic low-grade leaks; allocation-rate questions
JVM metrics (Micrometer)Which memory area is growingAlways on; per-pool breakdownAggregates hide individual objectsContinuously; the tripwire for every other artifact

Checklists

Before production

  • -XX:+HeapDumpOnOutOfMemoryError + -XX:HeapDumpPath on every JVM workload, pointed at durable or promptly-copyable storage.
  • -Xmx bounded below the container limit with explicit non-heap headroom; MaxRAMPercentage or explicit flags chosen deliberately.
  • Dump volume sized ≥ max heap, monitored, and cleaned by policy.
  • jcmd available in-image or a documented kubectl debug path; verified with which jcmd, not assumed.
  • GC logging or heap metrics in place — you cannot confirm a leak pattern without a history to compare against.
  • Data-handling policy for dumps agreed with security: access, transfer, retention, deletion.
  • The capture-and-copy runbook rehearsed once — in the incident is too late.

During the incident

  • Confirm the pattern first: rising heap-after-GC floor in metrics, not just a high number.
  • Rule out OOMKill-vs-OOME: if the pod was OOMKilled with flat JVM heap, investigate native memory instead.
  • Cheap probe before heavy capture: jcmd <pid> GC.class_histogram.
  • Record context: VM.flags, Thread.print, pod events, deploy history, request rate.
  • Warn before the pause: dump the pod only after removing it from rotation or accepting the stall.
  • Copy the dump off the pod before it can restart; check that its size is plausible for the heap.
  • Log where every copy of the dump lives — it contains production data.

Heap dump analysis

  • MAT -Xmx raised above dump size; index files have a writable home.
  • Leak Suspects read but not trusted; histogram and dominator tree done by hand.
  • Suspects traced to GC roots with weak/soft/phantom references excluded.
  • Owner identified at field level: class + field + expected vs. actual cardinality.
  • Retention classified: expected-bounded / expected-unbounded / accidental / transitional.
  • Finding correlated with request rate, endpoint, schedule, or deploy — in that order.
  • Growth confirmed by second dump or histogram delta before the fix ships.

Before deploying the fix

  • The fix targets the retaining owner — the map field, the lifecycle, the bound — not the symptom (-Xmx bump).
  • Every path that adds has a matching remove — error paths and cancellation included.
  • New bounds have numbers behind them: expected cardinality × entry size ≤ budget.
  • Rollback plan exists and the flag/config toggle is tested.
  • Post-deploy evidence defined in advance: heap floor flat across N days, second dump delta gone, restart count zero.

Exercise: the order service

A Spring Boot order-processing service runs in Kubernetes with -Xmx2g inside a 2.5 Gi limit. Over roughly six hours of steady traffic, heap-after-GC climbs from ~700 MB to ~1.9 GB, then the pod restarts. A heap dump captured at 1.8 GB shows, in the dominator tree, that one OrderService singleton retains 68% of the heap via a ConcurrentHashMap named inFlight; its entries are OrderContext records each holding the request’s raw byte[] payload (~40 KB average, ~29,000 entries). Path to GC Roots runs: OrderContext ← Node ← table[] ← inFlight ← OrderService ← Spring context ← Class ← system class. There is no cache annotation anywhere near this field.

OrderService.java — the code behind the map
public void submit(String orderId, byte[] payload) {
inFlight.put(orderId, new OrderContext(orderId, payload, System.nanoTime()));
try {
validate(payload);
process(payload);
inFlight.remove(orderId); // happy path cleans up
} catch (InvalidOrderException e) {
reject(orderId); // returns without removing
}
}

Questions to reason through before reading further:

  1. What is retaining the OrderContext instances?
  2. Why can’t GC reclaim them?
  3. What is the appropriate code or design change?
  4. What evidence would prove the fix worked?

Reading the case

1. What retains them. The inFlight ConcurrentHashMap on the OrderService singleton — an application-lifetime bean holding request-lifetime entries.

2. Why GC can’t help. Every entry is strongly reachable through the singleton’s field; nothing about them is weak, expired, or unreachable. GC is doing its job — the objects are alive because a live map points at them.

3. The fix. The code removes entries on success but not on the rejection path — a leak proportional to the invalid-order rate, which also explains why the climb tracks traffic. The minimal fix is finally { inFlight.remove(orderId); }. The deeper question is whether a synchronous submit needs an inFlight map at all — if it exists for observability (counting/stuck-order detection), bound it or move tracking to a metric; if entries are meant to persist across processing stages, they need an explicit lifecycle and a TTL sweep for orphaned entries, because “remove on every exit path” is a bug that recurs every time someone adds a path.

4. Validation. Reproduce against staging traffic including invalid orders: inFlight size should return to ~0 when traffic stops, heap floor should stay flat across hours, and a second dump should show OrderContext counts tracking active requests rather than cumulative requests. A jcmd GC.class_histogram delta over time is the cheap version of the same proof.

Incident playbook

When the page says memory is climbing, in this order:

  1. Confirm the shape. Heap-after-GC floor rising over time → proceed. Flat heap with RSS growth → switch to native-memory diagnosis. Spike → wait and re-check.
  2. Preserve context. kubectl describe pod, jcmd <pid> VM.flags, Thread.print, deploy history, traffic graph.
  3. Cheap probe. jcmd <pid> GC.class_histogram — does a domain class track lifetime request count?
  4. Stabilize. If restart is imminent, remove the pod from rotation first; if the flag was set, check for an existing java_pid*.hprof from the last OOME.
  5. Capture. jcmd <pid> GC.heap_dump -gz=1 /var/dumps/heap.hprof.gz — warning issued, pause accepted, disk confirmed.
  6. Exfiltrate before restart. kubectl cp (or exec cat) the dump out; check the size is plausible before deleting the pod copy.
  7. Analyze. Histogram → dominator tree → path to GC roots (weak/soft excluded) → incoming references → owner at field level.
  8. Classify before fixing. Bounded/expected vs. unbounded/accidental; correlate the owner with an endpoint, schedule, or deploy.
  9. Fix the holder. Add the bound, the eviction, the remove() in finally, the lifecycle hook — not the bigger -Xmx.
  10. Verify the cure. Post-deploy: flat heap floor, zero restarts, and a second dump or histogram delta that no longer shows the class growing.
  11. Close out. Write down the retaining path in the incident report; delete every dump copy per the data policy; add the pattern to the team’s leak review checklist.

Where to go next

The companion piece on JVM internals and diagnostics covers the non-heap side of memory incidents — metaspace, threads, NMT — and the GC collectors whose logs tell you whether a dump is even warranted. For the follow-up skill after this one: get comfortable reading a dump you took on purpose, in staging, before the night you need it in production. The workflow is the same; the stakes are not.

JVMPerformanceKubernetes

Related Articles

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind