1 · End-to-end pipeline
+Two independent subsystems cooperate. LivenessTracker decides that + something is leaking and which class. ReferenceChainTracker then answers why it is + retained, by running a budgeted, resumable, tag-based breadth-first walk of the reachable object graph on + its own JVMTI-attached thread.
+ +into per-klass population rings"] + POP --> TREND["hasQualifyingGrowth():
least-squares slope + hysteresis"] + POP --> FLOOR["heapFloorRising() /
secondsToOOM() projection"] + TREND --> CAND["selectLeakCandidates()
top-5 by slope"] + FLOOR --> CAND + end + + subgraph RC["ReferenceChainTracker (retention explanation)"] + direction TB + SIG["hasLeakSignal()"] --> AFF["canAffordNewSearch()
pain budget + signal"] + AFF --> SCHED["shouldRunPass()"] + SCHED --> PASS["runPass() — one budgeted pass"] + PASS --> FRONT[("FrontierTable
tag -> parent_tag")] + FRONT --> PASS + PASS --> TERM["Termination check"] + TERM -->|"terminal"| REL["releaseSearchTags() -> restartSearch()"] + REL --> SCHED + end + + CAND --> SIG + CAND --> POLL["pollWatchedTargets():
tag candidate instances"] + POLL --> PASS + FRONT --> RECON["reconstructChain()
leaf -> root"] + RECON --> JFR["JFR: ReferenceChain /
ReferenceChainAbandoned"] +
One thread, one lock
+A single agent-owned pthread (threadLoop(), referenceChains.cpp:690)
+ wakes on an adaptive cadence, runs at most one pass, then polls targets. GC callbacks only bump an atomic epoch.
Tags, not handles
+Frontier identity is a JVMTI object tag — non-retaining, so the walk never keeps a dying object alive.
+ Dead objects vanish for free at the next GetObjectsWithTags.
Resumable by construction
+Every pass is bounded by an edge budget and a wall-clock deadline. Unfinished work stays in a FIFO
+ (_pending_expand) and is picked up by the next pass — including a partially-visited batch.
Cost is paid for, not assumed
+Two leaky buckets (PainBudget): one gates starting a search on accumulated safepoint cost,
+ one gates every pass on non-safepoint CPU cost.
-
+
- There is no root-seeded
FollowReferencesfirst pass any more.runPass()always + drivesrunPassManualWalk(); the only first-pass difference is that root enumeration is forced + (referenceChains.cpp:3665-3699). The header comment at + referenceChains.h:47-53 still describes the old shape.
+ SearchAbandonReason::TTLis not driven by_ttl_ms. That field is assigned in +start()and only reported in the abandoned event; the branch that storesTTLis the + no-progress detector (referenceChains.cpp:3807-3818).
+ CANARY_NO_PROGRESS_PASS_LIMIT = 3carries an in-source// TEMP: was 30, lowered for testing+ marker (referenceChains.h:2376), while the comment below it still reasons about a base of 30.
+
2 · Leak signal: how a class becomes a candidate
+A klass is only trusted as leaking after it clears a fill gate, a slope gate, a class-level hysteresis + gate and a per-thread hysteresis gate. The bar itself moves depending on whether the aggregate heap floor + corroborates the story.
+ +tracked objects per klass"] --> B["count_ring[30] push
(sample = distinct GC ages, not raw count)"] + B --> C{"ring_fill >= 10
KLASS_POPULATION_MIN_FILL_FOR_TREND"} + C -->|"no"| X1["no trend yet"] + C -->|"yes"| D["least-squares regression over the ring
cached_slope = recent_mean - earliest_mean"] + D --> E{"cached_slope >= max(0.15 * earliest_mean, 1)
LEAK_GROWTH_REL_MIN / ABS_MIN"} + E -->|"no"| F["consecutive_positive = 0 (hard reset)"] + E -->|"yes"| G["consecutive_positive++ (saturating)"] + + G --> H{"heapFloorRising()?"} + H -->|"yes, corroborated"| I["required_hysteresis = 3"] + H -->|"no"| J["required_hysteresis = 5"] + I --> K + J --> K{"consecutive_positive >= required_hysteresis
AND cached_slope > 0"} + K -->|"no"| X2["skip klass"] + K --> L{"any tid_trend with
consecutive_positive >= required_hysteresis"} + L -->|"no"| X3["skip klass — no single thread
owns the growth"] + L -->|"yes"| M["insert into top-5 by slope descending"] + M --> N["selectLeakCandidates() returns candidates"] +
Heap-floor corroboration
+heapFloorRising() (livenessTracker.cpp:1220-1259) demands
+ both a rising mean and a rising minimum, so a sawtooth workload whose troughs stay flat does not lower the bar.
| Gate | Constant | Value |
|---|---|---|
| Mean rise, relative | HEAP_FLOOR_GROWTH_REL_MIN | 0.02 |
| Mean rise, absolute | HEAP_FLOOR_GROWTH_ABS_MIN | 1 MiB |
| Floor rise, relative | HEAP_FLOOR_FLOOR_REL_MIN | 0.01 |
| Floor rise, absolute | HEAP_FLOOR_FLOOR_ABS_MIN | 512 KiB |
selectLeakCandidates() is slow and gated
+ (livenessTracker.cpp:1370). topKlassesByGenerationCount() is fast and
+ ungated — ranked on the single latest sample with no hysteresis at all
+ (livenessTracker.cpp:1480). The fast one is only ever consulted after the slow one
+ has already fired, so it never needs to wait out the same hysteresis twice
+ (referenceChains.cpp:4175-4201).
+ Qualifying growth bar: max(0.15 · earliest_mean, 1)
+ Below an earliest-window mean of ~6.7 tracked instances, the absolute floor
+ (LEAK_GROWTH_ABS_MIN = 1) sets the bar, not the 15% relative term
+ (LEAK_GROWTH_REL_MIN) — a klass with only a handful of instances still needs to grow by a whole
+ instance to qualify, it cannot coast in on a tiny relative slope.
3 · Urgency: the OOM projection and its latch
+secondsToOOM() projects when the live-heap floor will hit the tighter of the JVM max heap
+ and the container limit. The raw value swings by orders of magnitude between observations, so it is never compared
+ bare — it is read through a latch with separate arm and release bars.
30 post-GC samples"] --> B{"_gc_generations on
AND max_heap > 0"} + B -->|"no"| N1["return -1 (no projection)"] + B -->|"yes"| C["limit = min(JVM max heap, container limit)"] + C --> D{"ring_fill >= 10"} + D -->|"no"| N2["return -1 (INSUFFICIENT_FILL)"] + D -->|"yes"| E["regress bytes and time"] + E --> F{"bytes_delta > 0 AND time_delta > 0"} + F -->|"no"| N3["return -1 (NOT_RISING)"] + F -->|"yes"| G["re-regress the most recent half
(min fill 5)"] + G --> H{"recent half also rising?"} + H -->|"no"| N4["return -1 (RECENT_HALF_FLAT)
rejects a plateaued step change"] + H -->|"yes"| I["seconds = (limit - recent_mean) / rate"] +
also clears _urgent_search_spent, so a new episode earns a fresh entitlement" + note right of Latched: "a reading between 300 and 600 stays Latched
and resets _urgent_release_ticks to 0" + Latched --> Releasing: "reading at or over 600 (OOM_URGENT_RELEASE_S), or negative — ++_urgent_release_ticks" + Releasing --> Latched: "any reading back under the release bar" + Releasing --> Calm: "URGENT_RELEASE_CONSECUTIVE = 5 clear readings in a row" +
isUrgent() uses the latch above with a 300 s arm bar and is what suppresses the no-progress abandon and
+ authorises one out-of-band search per episode.
+ threadLoop()'s ramp uses a raw, unlatched, much wider window —
+ seconds_to_oom < OOM_RAMP_START_S = 1800 s — recomputed on every wake
+ (referenceChains.cpp:735-736). They arm at different distances and flap differently.
+ The OOM ramp (referenceChains.cpp:739-772)
+Inside the 30-minute window, with x = 1 - secondsToOOM()/1800 (0 at the far edge, 1 at OOM),
+ both the pause target and the cadence ramp exponentially toward their urgent ceilings:
| Quantity | Baseline | Urgent ceiling | Ramp |
|---|---|---|---|
| Per-pass pause target | _pause_target_ms | URGENT_PAUSE_TARGET_MS = 100 ms | base * (100/base)^x |
| Pass cadence | PASS_CADENCE_NS = 1 s | URGENT_CADENCE_NS = 10 ms | 1s * (10ms/1s)^x |
| Edge budget ceiling | _budget | min(_budget * 4, MAX_REFERENCE_CHAINS_BUDGET) | step, on target change |
The cadence ramps from the fixed baseline, not from the live _effective_cadence_ns —
+ anchoring on the moving value would compound the exponent across iterations. While urgent, the ramp owns
+ _effective_cadence_ns outright; updatePacing() silently resumes ownership the moment urgency clears.
Cadence ramp: 1000ms · (10ms/1000ms)^x
+ Fully determined by PASS_CADENCE_NS and URGENT_CADENCE_NS — no
+ assumed baseline. The dashed marker is where isUrgent()'s own, separately-latched threshold
+ (OOM_URGENT_THRESHOLD_S = 300 s) falls on this same x-axis, for scale.
4 · Pass scheduling: shouldRunPass()
+ Called once per BFS-thread wake. Three regimes — no search yet, search terminal, search running — + evaluated strictly in this order (referenceChains.cpp:870-1006).
+ += safepoint pain budget drained
AND hasLeakSignal()"} + A1 -->|"no"| F1["FALSE"] + A1 -->|"yes"| T1["_urgent_search_spent = _urgent_latched
TRUE — take the very first pass"] + + B1 -->|"yes"| B2{"_search_state == RUNNING?"} + + B2 -->|"no (terminal)"| C1{"_tags_released?"} + C1 -->|"no"| T2["TRUE — force runPass() to retry
releaseSearchTags(); restart is forbidden
until every tag is confirmed cleared"] + C1 -->|"yes"| C2{"canAffordNewSearch(now)"} + C2 -->|"no"| F2["FALSE — terminal, waiting"] + C2 -->|"yes"| T3["restartSearch()
TRUE"] + + B2 -->|"yes"| D0["canary_active = candidates outstanding
all_covered = leak tags all resolved
emergency = canary_active AND no candidate
progress for 3 passes"] + D0 --> D1["CPU pain refill multiplier:
all_covered -> 1x
emergency -> 100x
canary_active -> 15x
else -> 1x"] + D1 --> D2{"_cpu_pain_budget.canStartNow()"} + D2 -->|"no"| F3["FALSE — throttled"] + D2 -->|"yes"| D3{"gcFinishEpoch() changed
since last pass?"} + D3 -->|"yes"| T4["TRUE — GC trigger"] + D3 -->|"no"| D4{"canary_active?"} + D4 -->|"yes"| T5["TRUE — run back-to-back,
cadence bypassed"] + D4 -->|"no"| D5{"now - _last_pass_ns >=
_effective_cadence_ns"} + D5 -->|"yes"| T6["TRUE — cadence trigger"] + D5 -->|"no"| F4["FALSE — idle"] +
Order matters in the multiplier
+all_covered is tested before emergency, so a search whose leak tags are all resolved
+ drops back to 1× even if the canary is nominally stuck.
Two budgets, two scopes
+_safepoint_pain_budget is never consulted for a RUNNING search — only through
+ canAffordNewSearch() when starting or restarting one. _cpu_pain_budget gates each pass.
No early wake on GC
+onGCFinish() bumps an atomic epoch and nothing else. Waking the thread per GC would buy ≤1 s of
+ latency while collapsing the loop cadence to GC frequency (referenceChains.cpp:858-866).
Sleep only when idle
+threadLoop() skips its sleep entirely when a pass is about to run, so canary passes execute
+ back-to-back and the PID controller alone regulates cost (referenceChains.cpp:805-811).
hasLeakSignal()
+ (referenceChains.cpp:785-798): that signal answers "is there a leak candidate right now",
+ which is unrelated to whether an in-flight search still has pending frontier work. Gating every pass on it would
+ stall a search's own convergence whenever no candidate happens to be visible.
+ 5 · One search, pass by pass
+A search is a sequence of passes over a persistent frontier. Each pass performs up to four sub-phases, + each of which can truncate independently and hand its remainder to the next pass.
+ +return as a no-op"] + G1 -->|"yes"| R0["resolveLoadedClasses()
class_tag -> class-name dictionary id
(per-class scan skipped when the count is unchanged)"] + + R0 --> R1{"first pass?
(_search_started == false)"} + R1 -->|"yes"| R2["_search_started = true
_search_start_ns = now
force root enumeration"] + R1 -->|"no"| R3{"last root enum truncated
OR >= 2s since last
(ROOT_ENUM_MIN_INTERVAL_NS)"} + R3 -->|"yes"| R2 + R3 -->|"no"| R4["skip root enumeration this pass"] + + R2 --> W["runPassManualWalk()"] + R4 --> W + + subgraph WALK["runPassManualWalk — four sub-phases under one deadline"] + direction TB + W1["A · Root enumeration
IterateOverReachableObjects
budget = _first_pass_budget · no deadline"] + W2["B · Static-field sweep
admitStaticFieldRoots(), 512-class chunk
only when the loaded-class count changed"] + W3["C · Ordinary expansion
expandFrontier(_pending_expand / _priority_expand)"] + W4["D · Rotation
collect stale roots / leak accumulators / stale expanded
then a second expandFrontier() on the reserved budget"] + W1 --> W2 --> W3 --> W4 + end + + W --> WALK + WALK --> M0["_passes_run++ · _last_pass_gc_finish_epoch · _last_pass_ns"] + M0 --> M1{"root-enum pass?"} + M1 -->|"no"| M2["updatePacing(safepoint_ticks)"] + M1 -->|"yes"| M3["maybeRevokeBorrowForRootEnumPass()
(excluded from the PID signal)"] + M2 --> M4 + M3 --> M4["_search_pain_ms += safepoint ms
_cpu_pain_budget.spend(non-safepoint ms)"] + M4 --> M5["Termination decision chain — see tab 8"] +
| Sub-phase | Budget | Deadline | Why it exists |
|---|---|---|---|
| A · Root enumeration | +_first_pass_budget (auto = min(_budget*50, 200000)) |
+ none, by design | +Enumerates heap roots and stack refs. Re-run at most every 2 s; already-admitted roots short-circuit + cheaply, so a root missed this pass is picked up later, never lost. | +
| B · Static-field sweep | +chunk of 512 classes | +_pass_deadline_ns |
+ JVMTI has no jvmtiHeapRootKind for static fields, and expansion never descends from class
+ objects — without this, SomeClass.staticField → obj is structurally undiscoverable. |
+
| C · Ordinary expansion | +_effective_budget − rotation reserve |
+ _pass_deadline_ns |
+ The actual BFS: resolve a batch of frontier tags, walk exactly one hop from each. | +
| D · Rotation | +min(expand_budget/2, 288) reserved up front |
+ _pass_deadline_ns |
+ Re-observes entries whose recorded story may have gone stale: transient root kinds, leak-accumulating + signatures, long-expanded entries whose fields have since been mutated. | +
6 · expandFrontier(): the resumable batch loop
+ The heart of the resumability. Each iteration resolves a batch of frontier tags back to live objects,
+ runs one stop-the-world FollowReferences for the whole batch, and descends exactly one hop —
+ enforced by a membership gate on the batch's own tag set (referenceChains.cpp:2943-3303).
OS::nanotime() >= _pass_deadline_ns"} + L0 -->|"yes"| TR["truncated = true — break"] + L0 -->|"no"| L1["pick a lane:
_priority_expand vs _pending_expand,
alternating via _expand_lane_prefer_priority"] + L1 --> L2{"both lanes empty?"} + L2 -->|"yes"| DONE["no pending work — break"] + L2 -->|"no"| L3["batch_size = min(queue, budget, _gotw_batch_size)"] + L3 --> L4["GetObjectsWithTags(batch)
dead tags simply do not come back = free pruning"] + L4 --> L5["self-calibrate _gotw_batch_size:
per-call EMA scaled by the remaining
deadline window, clamped to [8, 512]"] + L5 --> L6["build jobjectArray holder
+ fill batch_tags membership set"] + L6 --> L7["FollowReferences(holder) — one STW HeapWalkOperation
heapReferenceCallback descends only into tags in batch_tags"] + L7 --> L8{"truncated?"} + + L8 -->|"no"| S1["for each tag in batch:
resolved -> markExpanded()
unresolved -> clear() = ABANDONED
pop_front() · progress = true"] + L8 -->|"yes, some batch entry visited"| S2["rolling resume:
settle only entries BEFORE
_last_visited_batch_tag;
leave it and the rest at the queue front"] + L8 -->|"yes, nothing visited"| S3["leave the whole batch queued"] + + S1 --> L0 + S2 --> TR + S3 --> TR +
One hop, guaranteed
+heapReferenceCallback returns JVMTI_VISIT_OBJECTS only when the visited object's tag is
+ in batch_tags. Everything else is admitted but not descended into
+ (referenceChains.cpp:2001-2017).
Rolling resume
+_last_visited_batch_tag is the cursor inside a truncated batch. Entries before it are settled;
+ the partially-visited one stays at the queue head and is re-walked next pass.
Deadline sampled cheaply
+Inside the callback the wall clock is only read every 4096th invocation
+ ((++counter & 0xFFF) == 0) — the loop top checks it per batch
+ (referenceChains.cpp:1647-1656).
Adaptive batch size
+_gotw_batch_size is tuned per call from measured GetObjectsWithTags cost against the
+ remaining deadline window, so a slow JVM naturally shrinks its batches rather than blowing the pause target.
Admission, in callback order
+past _pass_deadline_ns?"} + A1 -->|"yes"| Z1 + A1 -->|"no"| A2{"tag at or below MARKER_TAG_BASE
(canary marker)"} + A2 -->|"yes"| Z2["record chain link, set found bit,
do NOT descend"] + A2 -->|"no"| A3{"tag is negative (class object)"} + A3 -->|"yes"| Z3["descend only for the static-field
seed holder -> class edge"] + A3 -->|"no"| A4["compute parent_tag and depth
from the referrer's tag"] + A4 --> A5{"depth >= hop_cap"} + A5 -->|"yes"| Z4["drop — no admission, no descent"] + A5 -->|"no"| A6{"isLeakTag(tag)?"} + A6 -->|"yes"| Z5["allocate frontier tag, insert,
setLeakTag(), record instance, descend"] + A6 -->|"no"| A7{"tag == 0 (unseen)"} + A7 -->|"yes"| A8["admitObject()"] + A7 -->|"no"| A9["already admitted ->
improveChain() / reparentToDurableRoot()
/ maybeUpgradeRootAttachedRootKind()"] + A8 --> A10{"result"} + A10 -->|"BUDGET_EXHAUSTED"| Z6["truncated · abort"] + A10 -->|"FRONTIER_CAP_HIT"| Z7["frontier_cap_hit · abort
-> whole search abandoned"] + A10 -->|"ADMITTED"| A11["push onto priority or pending queue"] + A9 --> A12 + A11 --> A12{"tag in batch_tags?"} + A12 -->|"yes"| Z8["JVMTI_VISIT_OBJECTS — descend one hop"] + A12 -->|"no"| Z9["return 0 — do not descend"] +
ALREADY_ADMITTED is the idempotency that makes re-enumeration cheap.
+ admitObject() short-circuits on any non-zero tag
+ (referenceChains.cpp:2022-2057), which is precisely why root enumeration can be re-run on
+ every pass without re-paying for the graph it already discovered.
+ 7 · Root discovery and the chunked static-field sweep
+Two disjoint sources of root-attached entries. IterateOverReachableObjects reports stack
+ locals, JNI handles, monitors and thread roots — but never static fields, which need their own sweep.
jvmtiHeapRootKind -> jvmtiHeapReferenceKind"] + R2 --> R3["admitObject(parent_tag = 0, depth = 0, root_kind)"] + R3 --> R4{"ALREADY_ADMITTED?"} + R4 -->|"yes"| R5["maybeUpgradeRootAttachedRootKind()
durability tie-break"] + R4 -->|"no"| R6["new root-attached frontier entry"] + end + + subgraph SF["B · admitStaticFieldRoots — chunked, cursor-resumed"] + direction TB + S1["GetLoadedClasses()"] --> S2["partition app classes
(non-null loader) to the front"] + S2 --> S3["chunk = [cursor, cursor + 512)"] + S3 --> S4["holder array filled in REVERSE chunk order
so HotSpot's LIFO descent visits ascending"] + S4 --> S5["one FollowReferences with static_field_seed = true
empty batch_tags -> exactly one hop past each class"] + S5 --> S6{"truncated?"} + S6 -->|"yes"| S7["mark lap truncated;
cursor = chunk_start + classes_visited - 1
(redo the partial class)"] + S6 -->|"no"| S8["cursor = chunk_end"] + S7 --> S9 + S8 --> S9{"cursor reached class_count?"} + S9 -->|"yes"| S10["lap wrap: cursor = 0;
cycle_complete only if no chunk truncated"] + S9 -->|"no"| S11["resume here next pass"] + end +
FollowReferences over every class cannot finish inside a 5–50 ms deadline — the observed behaviour was
+ truncated = 1 on 275 of 275 passes with 0–1 edges admitted. Because a single-call sweep restarts at
+ class 0 every time, every class past the deadline point was permanently unreachable
+ (referenceChains.h:1936-1960).
+ Root-kind durability
+A root-attached entry records why it is reachable. When a more durable root is later observed + admitting the same object, the recorded kind is upgraded rather than keeping whichever root happened to be enumerated + first (referenceChains.h:249-278).
+| Tier | Kinds | Meaning |
|---|---|---|
| 3 — most durable | STATIC_FIELD, SYSTEM_CLASS | Genuine retention evidence. |
| 2 | JNI_GLOBAL | Durable, but owned outside the JVM heap. |
| 1 — transient | MONITOR, STACK_LOCAL, JNI_LOCAL, THREAD, OTHER | "First observed via", not "rooted by". THREAD and OTHER have no documented tier and are conservatively bucketed here. |
Two repair operations exist because a depth comparison alone cannot express both cases:
+ improveChain() replaces a shallow root-attached entry when a strictly deeper path reaches it, and
+ reparentToDurableRoot() handles the equal-depth case — a depth-1 entry parented to a transient root is
+ re-parented to a durable one. Without the latter, the real hotdog shape (a static singleton collection at depth 0, its
+ elements at depth 1) would keep a stack-local parent forever, because the depths tie.
8 · State machines
+Two independent state machines: one per frontier entry, one per search.
+ +Per-entry: FrontierEntryState
+ the re-walk marks it EXPANDED again" + FRONTIER --> EDGE: "markEdge() — reconstructChain() walked through this hop" + EXPANDED --> EDGE: "markEdge()" + EDGE --> EXPANDED: "rotation re-expansion overwrites the state" + EXPANDED --> ABANDONED: "releaseSearchTags() at search end" + EDGE --> ABANDONED: "releaseSearchTags() at search end" + ABANDONED --> [*]: "resetForRestart() zeroes _table_size — every slot becomes un-inserted" +
releaseSearchTags() leaves parent_tag, referrer_klass, depth and
+ root_kind intact, so reconstructChain() keeps working from memory after the search ends
+ (referenceChains.h:1966-1975).
+ Per-search: SearchState
+ stays RUNNING, rotation keeps re-observing" + RUNNING --> ABANDONED_TTL: "30 passes with no frontier growth AND not isUrgent()" + RUNNING --> COMPLETED_C: "all canary candidates found" + RUNNING --> ABANDONED_CS: "candidates outstanding AND 30 passes without frontier growth AND canary stall past canaryStuckPassLimit()" + + ABANDONED_FC --> RELEASE + ABANDONED_TTL --> RELEASE + ABANDONED_CS --> RELEASE + COMPLETED_G --> RELEASE + COMPLETED_C --> RELEASE + + note right of RELEASE: "releaseSearchTags() failed — shouldRunPass()
returns true to retry; restart stays blocked" + RELEASE --> RUNNING: "_tags_released AND canAffordNewSearch() -> restartSearch()" + + state "ABANDONED / FRONTIER_CAP" as ABANDONED_FC + state "ABANDONED / TTL (no-progress)" as ABANDONED_TTL + state "ABANDONED / CANARY_STUCK" as ABANDONED_CS + state "COMPLETED (graph exhausted)" as COMPLETED_G + state "COMPLETED (candidates found)" as COMPLETED_C + state "tag release + canary reset" as RELEASE +
| Reason | Value | Suppressed by isUrgent()? | Rationale |
|---|---|---|---|
NONE | 0 | — | Not abandoned. |
FRONTIER_CAP | 1 | No | The metadata table is full; nothing further can be admitted at all. |
TTL | 2 | Yes | A search still making real progress must not be killed just because the process is close to OOM. |
CANARY_STUCK | 3 | No | A candidate chase with zero discovery progress is provably not converging; continuing at urgency-boosted budget only burns pause budget the dying process needs. |
9 · Termination, tag release and restart
+Six checks, evaluated in a fixed priority order at the end of every pass + (referenceChains.cpp:3765-3856). Only the first match applies.
+ +enqueue abandoned event"] + C1 -->|"no"| C2{"no pending frontier
AND _watched_leak_klass_count == 0"} + C2 -->|"yes"| A2["COMPLETED — graph exhausted
within the caps"] + C2 -->|"no"| C3{"no pending frontier
but a watch is active"} + C3 -->|"yes"| A3["stay RUNNING — rotation keeps
re-observing mutated fields"] + C3 -->|"no"| C4{"_passes_since_last_progress >= 30
AND NOT isUrgent()"} + C4 -->|"yes"| A4["ABANDONED / TTL"] + C4 -->|"no"| C5{"all candidates found"} + C5 -->|"yes"| A5["COMPLETED + counter
REFERENCE_CHAIN_CANDIDATES_FOUND"] + C5 -->|"no"| C6{"candidates outstanding
AND 30 passes no frontier growth
AND canary stall >= canaryStuckPassLimit()"} + C6 -->|"yes"| A6["ABANDONED / CANARY_STUCK
_canary_stuck_restart_count++"] + C6 -->|"no"| A7["stay RUNNING"] + + A1 --> R + A2 --> R + A4 --> R + A5 --> R + A6 --> R + R["releaseSearchTags():
GetObjectsWithTags over every non-ABANDONED tag,
SetTag(obj, 0), then mark all ABANDONED"] --> R2{"batch call succeeded?"} + R2 -->|"no"| R3["_tags_released = false — mark NOTHING;
a failed batch says nothing about liveness,
and rewinding _next_tag while a live object
still carries a tag would break tag uniqueness"] + R2 -->|"yes"| R4["_tags_released = true
clear canary marker tags · zero candidate state
reset _canary_stuck_restart_count unless CANARY_STUCK"] +
Escalating patience
+canaryStuckPassLimit() = CANARY_NO_PROGRESS_PASS_LIMIT << min(_canary_stuck_restart_count, 8)
+ (referenceChains.h:2389-2393). Repeated CANARY_STUCK abandons double the stall
+ tolerance, up to 256×. _canary_stuck_restart_count is deliberately not reset by
+ restartSearch() — it must survive restarts or the escalation could never widen.
Escalation: canaryStuckPassLimit(n) = 3 << min(n, 8)
+ Each consecutive CANARY_STUCK abandon doubles the stall tolerance for the
+ next attempt at the same candidate chase, up to the MAX_CANARY_STUCK_BACKOFF_SHIFT = 8 cap
+ (restart count 8 and 9 tolerate the same 768 passes — the shift has saturated).
What a restart resets, and what it does not
+Reset (restartSearch())
+ _frontier->resetForRestart() (slot occupancy only, the allocation survives) · _next_tag = 1 ·
+ _search_started = false · state back to RUNNING · both expand queues and the membership index ·
+ leak-signature maps · _leak_tags_assigned/_resolved · _last_pass_* · _passes_run ·
+ the resolved/static-field class counts.
Persists across a restart
+The class-tag table and its allocator · _resolved_chains · _watched_leak_klass_ids ·
+ _canary_stuck_restart_count · the rotation cursors and the static-field sweep cursor ·
+ _pause_pid, _effective_budget, _effective_cadence_ns · the cached
+ java/lang/Object global ref · the FrontierTable allocation and capacity.
restartSearch() also spends the finished search's accumulated
+ _search_pain_ms into _safepoint_pain_budget before zeroing it
+ (referenceChains.cpp:1123) — a cheap search may restart again soon, an expensive one must wait
+ proportionally longer. It asserts on _tags_released.
10 · Pacing: PID controller, budget borrowing, pain budgets
+The configured constants are ceilings and baselines, not literal per-pass values. What each pass + actually spends is decided by a feedback loop over the previous pass's measured in-safepoint time.
+ +pass_wall_ticks total,
safepoint_ticks inside
IterateOverReachableObjects / FollowReferences"] --> SPLIT{"split"} + SPLIT -->|"safepoint ms"| PID["_pause_pid.compute(pass_ms)
target = _effective_pause_target_ms
positive signal = came in UNDER target"] + SPLIT -->|"safepoint ms"| PAIN1["_search_pain_ms +=
-> gates the NEXT search"] + SPLIT -->|"non-safepoint ms"| PAIN2["_cpu_pain_budget.spend()
-> gates the NEXT pass"] + + PID --> BORROW{"pass_ms at or under 50% of target
for 5 consecutive passes?"} + BORROW -->|"yes"| B1["_borrowed_budget += _budget * 0.25
capped at 3 * _budget"] + BORROW -->|"no"| B2["_borrowed_budget = 0 immediately"] + B1 --> CL + B2 --> CL["ceiling = _budget + _borrowed_budget
floor = min(2000, ceiling)
_effective_budget = clamp(prev + signal)"] + CL --> OV{"overflow = desired - clamped"} + OV -->|"negative — still over target at the floor"| W["widen _effective_cadence_ns
by 1ms per overflow edge, up to 4s"] + OV -->|"positive"| S["shorten _effective_cadence_ns,
down to 10ms"] + OV -->|"== 0"| K["leave the cadence alone"] + W --> NEXT + S --> NEXT + K --> NEXT["next pass: expand_budget = _effective_budget,
deadline = _effective_pause_target_ms,
sleep/gate = _effective_cadence_ns"] +
Clamp function: effective_budget = clamp(desired, floor, ceiling)
+ Illustrative baseline _budget = 800 edges/pass (below
+ MIN_EFFECTIVE_BUDGET = 2000). With no borrowed headroom, floor = min(2000, ceiling) = ceiling
+ — the clamp range collapses to a single point and _effective_budget is pinned regardless of the PID
+ signal. Once _borrowed_budget pushes the ceiling past 2000 (shown at 3× budget, the borrowing
+ ceiling), a real band opens up and the floor becomes MIN_EFFECTIVE_BUDGET itself.
Inverted sign convention
+Unlike ObjectSampler / MallocTracer / RateLimiter, which subtract the PID signal from an interval, this + controller adds it to a budget: a positive signal means the pass came in under target, so the next one may + do more work (referenceChains.h:2002-2012).
+One compute() is one pass
+ sampling_window = 1 and time_delta_coefficient = 1.0, because a pass is not a fixed
+ real-time window like the other three usages assume. Gains are 10/1/2 — deliberately not copied from the shared triple.
Root-enum passes are excluded
+They spend the deliberately oversized _first_pass_budget; feeding that into the PID would throttle
+ every cheap expansion pass that follows (referenceChains.cpp:3713-3728).
Borrowing is lost instantly
+A single pass that is not comfortably under target zeroes both the warm-up streak and the whole accumulated + borrow. Earning headroom takes 5 passes; losing it takes one.
+The two leaky buckets
+_safepoint_pain_budget | _cpu_pain_budget | |
|---|---|---|
| Gates | Starting or restarting a search | Running an individual pass |
| Charged | Once per finished search, in restartSearch(), with _search_pain_ms | Every pass, with the non-safepoint remainder |
| Refill rate | pain_budget_percent / 100 | Same base, times 1× / 15× (covering) / 100× (emergency) |
| Read at | canAffordNewSearch() | shouldRunPass(), RUNNING branch |
0.0 never drains
+ (painBudget.h:46-49), so once anything has been spent the bucket blocks permanently.
+ Leaky bucket: _safepoint_pain_budget over time
+ Illustrative simulation, not a captured log — refill rate 2% (an example
+ pain_budget_percent), two searches finishing at t=0 and t=1600ms and spending their accumulated
+ _search_pain_ms (60ms, then 45ms) into the bucket. canAffordNewSearch() only returns
+ true where the curve touches the zero baseline — everywhere the shaded area is above zero, a new search or
+ restart is blocked.
Auto-tuning at start()
+ autoTuneDefaults() (referenceChains.cpp:360-462) fills in only the
+ knobs the operator did not set explicitly, from max heap size and processor count:
| Knob | Formula |
|---|---|
| Edge budget | DEFAULT * sqrt(heap_mib / 512), clamped to [default, max] — keeps the pause proportional to sqrt(heap) |
| First-pass budget | budget * 10, capped |
| TTL | DEFAULT * (heap_mib / 512), clamped to [default, 30 min] |
| Frontier cap | scaled by the budget ratio in floating point (integer division here would undershoot by ~20%) |
| Pause target | DEFAULT * (1 + (nprocs-1)/3), capped at 50 ms |
| Pain budget % | DEFAULT * (1 + (nprocs-1)/4), capped at 5% |
11 · From frontier to emitted chain
+The frontier table doubles as a degenerate edge store: a chain is just the transitive closure of
+ parent_tag links from a target back to a root-attached entry.
markEdge(tag)
root_kind = entry.root_kind
tag = entry.parent_tag"] + P --> C{"tag == 0?"} + C -->|"no, and hops <= maxCapacity()"| L + C -->|"no, bound exceeded"| F2["corrupt or cyclic -> return false"] + C -->|"yes"| OUT["chain in leaf -> root order
out_root_kind = the root-attached entry's kind"] + OUT --> EV["buildChainEvent():
targetTag = leak_tag if set, else the frontier tag
depth · root_kind · chain"] + EV --> JFR["JFR ReferenceChain event"] + JFR --> JOIN["backend joins ReferenceChain.targetTag
to HeapLiveObject.leakTag"] +
FrontierEntry fields
+ | Field | Purpose |
|---|---|
parent_tag | Tag of the entry that discovered this one; 0 means root-attached. The only link the reconstruction walks. |
referrer_klass | StringDictionary id of this object's class name, resolved ahead of time — GetClassSignature is illegal inside a heap callback. |
depth | Hop count from the root. Drives the hop cap and improveChain()'s "deeper wins" comparison. |
state | FRONTIER / EXPANDED / EDGE / ABANDONED. |
leak_tag | LivenessTracker's per-instance tag, copied in at admission. Becomes ReferenceChain.targetTag, the backend's join key. 0 = ordinary BFS admission. |
root_kind | jvmtiHeapReferenceKind of the admitting edge, meaningful only when parent_tag == 0. |
class_tag | Raw negative JVMTI class tag — stable across StringDictionary regeneration, unlike referrer_klass. Needed by the retroactive leak-accumulation seed scan, which runs long after the live callback is gone. |
jobject is ever retained. Holding a live handle would defeat the entire point of using
+ non-retaining JVMTI tags for frontier identity — the walk would keep the leak alive.
+ The canary path
+Candidate instances are pre-tagged with distinct negative marker tags
+ (MARKER_TAG_BASE = -(1<<62)) applied to a specific representative object — matching by class alone
+ would record a chain for an unrelated, possibly short-lived instance of the same class. When the walk hits a marker it
+ records the chain link and the found bit, but does not descend.
+ buildCanaryChainEvent() therefore starts from the recorded parent tag (the negative marker itself is
+ rejected by lookup()), then prepends the candidate's own class and reverses to leaf→root
+ (referenceChains.h:2647-2712).
Beyond the representative, any object of a watched class discovered by the walk is auto-recorded, up to
+ MAX_DISCOVERED_INSTANCES_PER_CLASS = 8 per slot — each instance's chain is independently useful, since
+ different instances may be retained by different paths.
12 · Constants reference
+Values as they stand on this branch. Several are explicitly marked provisional / unbenchmarked in source.
+ +Leak detection — livenessTracker.h
+ | Constant | Value | Role |
|---|---|---|
MAX_KLASS_POPULATION_ENTRIES | 256 | Tracked klasses. |
KLASS_POPULATION_RING_SIZE | 30 | Trend window, in GC epochs. |
KLASS_POPULATION_MIN_FILL_FOR_TREND | 10 | Minimum fill before any slope is trusted. |
LEAK_GROWTH_REL_MIN / _ABS_MIN | 0.15 / 1 | Qualifying growth bar. |
LEAK_TREND_HYSTERESIS_BASE | 5 | Consecutive qualifying epochs required. |
LEAK_TREND_HYSTERESIS_CORROBORATED | 3 | Lowered bar when the heap floor is rising. |
TID_TREND_RING_SIZE / MIN_FILL | 16 / 6 | Per-thread trend window. |
MAX_TID_TRENDS | 8 | Threads tracked per klass. |
MAX_LEAK_CANDIDATES | 5 | Top-k candidate slots. |
HEAP_FLOOR_RECENT_HALF_MIN_FILL | 5 | Corroboration window for the OOM projection. |
Scheduling and urgency — referenceChains.h
+ | Constant | Value | Role |
|---|---|---|
PASS_CADENCE_NS | 1 s | Baseline cadence and ramp anchor. |
OOM_URGENT_THRESHOLD_S | 300 s | Latch arm bar for isUrgent(). |
OOM_URGENT_RELEASE_S | 600 s (2×) | Latch release bar. |
URGENT_RELEASE_CONSECUTIVE | 5 | Clear readings needed to unlatch. |
OOM_RAMP_START_S | 1800 s | Start of the unlatched exponential ramp in threadLoop(). |
URGENT_PAUSE_TARGET_MS | 100 ms | Ramp ceiling for the pause target. |
URGENT_CADENCE_NS | 10 ms | Ramp floor for the cadence. |
CANARY_PAIN_BUDGET_COVERING_MULTIPLIER | 15× | CPU refill while chasing candidates. |
CANARY_PAIN_BUDGET_REFILL_MULTIPLIER | 100× | CPU refill in the emergency case. |
Walk, budgets and termination — referenceChains.h
+ | Constant | Value | Role |
|---|---|---|
AUTO_FIRST_PASS_BUDGET_MULTIPLIER / _CAP | 50 / 200000 | Auto-scaled first-pass edge budget. |
ROOT_ENUM_MIN_INTERVAL_NS | 2 s | Minimum spacing between root enumerations. |
STATIC_FIELD_SWEEP_CHUNK_CLASSES | 512 | Classes per static-field chunk. |
STATIC_FIELD_SWEEP_NON_STATIC_CAP_PER_CLASS | 32 | Non-static edges admitted per class per lap. |
MIN_EFFECTIVE_BUDGET | 2000 | Budget floor under PID control. |
MIN/MAX_EFFECTIVE_CADENCE_NS | 10 ms / 4 s | Cadence clamp. |
CADENCE_NS_PER_EDGE_OVERFLOW | 1 ms / edge | How budget overflow converts into cadence widening. |
BORROW_WARMUP_PASSES / ceiling | 5 / 3× _budget | Budget borrowing. |
PRIORITY_EXPAND_CAP | 1024 | Fast-lane queue cap (2048-slot membership index). |
NO_PROGRESS_PASS_LIMIT | 30 | Passes without frontier growth before the TTL abandon. |
CANARY_NO_PROGRESS_PASS_LIMIT | 3 — marked TEMP: was 30 | Base of the canary stall limit, and the emergency threshold. |
MAX_CANARY_STUCK_BACKOFF_SHIFT | 8 | Escalation cap — up to 256× the base. |
MARKER_TAG_BASE | -(1<<62) | Canary marker tag namespace. |
LEAK_TAG_BASE / POOL_SIZE | 0x40000000 / 256 | LivenessTracker leak-tag pool. |
MAX_DISCOVERED_INSTANCES_PER_CLASS | 8 | Auto-recorded instances per watched class. |
INITIAL_TABLE_CAPACITY | 1024 | Frontier table start size; doubles to the configured cap. |