From 1b86543b93f19b9e2c4ea0391de3eb04a9115f52 Mon Sep 17 00:00:00 2001 From: Jaroslav Bachorik Date: Thu, 17 Sep 2026 23:33:13 +0200 Subject: [PATCH 1/2] Add reference-chains design and architecture docs --- .../LiveHeapReferenceChains-Implementation.md | 410 ++++++++++++ doc/architecture/LiveHeapReferenceChains.md | 610 ++++++++++++++++++ .../ReferenceChains-SignalsExplained.md | 366 +++++++++++ doc/reference-chains-collection-summary.md | 87 +++ doc/reference-chains-design.md | 259 ++++++++ 5 files changed, 1732 insertions(+) create mode 100644 doc/architecture/LiveHeapReferenceChains-Implementation.md create mode 100644 doc/architecture/LiveHeapReferenceChains.md create mode 100644 doc/architecture/ReferenceChains-SignalsExplained.md create mode 100644 doc/reference-chains-collection-summary.md create mode 100644 doc/reference-chains-design.md diff --git a/doc/architecture/LiveHeapReferenceChains-Implementation.md b/doc/architecture/LiveHeapReferenceChains-Implementation.md new file mode 100644 index 0000000000..b016cd10d3 --- /dev/null +++ b/doc/architecture/LiveHeapReferenceChains-Implementation.md @@ -0,0 +1,410 @@ +# Live Heap Reference Chains — As-Built Implementation Reference + +**Status:** matches `jb/reference-chains` as of 2026-09 +**Jira:** [PROF-15341](https://datadoghq.atlassian.net/browse/PROF-15341) + +This document is the detailed, as-built reference for the reference-chain walk +engine. It records the mechanisms exactly as implemented, with the rationale +that lives in the code comments distilled into one place. + +Companion documents, each with its own scope: + +| Document | Scope | +|---|---| +| `doc/reference-chains-design.md` | The original design decision: why bounded BFS-from-roots was chosen over the alternatives | +| `doc/reference-chains-collection-summary.md` | The four-component summary: detection signal, resumable walk, latency budgets, rotation | +| `doc/architecture/LiveHeapReferenceChains.md` | The full architecture: frontier/tag lifecycle, triggering, termination, data structures | +| `doc/architecture/ReferenceChains-SignalsExplained.md` | A guided tour of the *scheduling* side: cadence, leak-signal gate, pain budget, PID controller, canary backoff, OOM ramp, abort path | + +This document covers what none of the above covers in one place: the pass +structure, root discovery (including static-field roots), candidate-scoped +reach, the canary tagging mechanics, the rotation tiers, the as-built +pacing/budget arithmetic, and the `LivenessTracker` admission boost. Where a +topic belongs to the signals tour (scheduling behavior), this document +cross-references it instead of repeating it. + +--- + +## 1. Data flow + +```mermaid +flowchart TD + LT["LivenessTracker::selectLeakCandidates (population-slope ranking)"] -->|"ranked klass candidates"| PWT["ReferenceChainTracker::pollWatchedTargets"] + PWT -->|"pre-tag each candidate's representative with a distinct marker tag"| CHASE["canary candidate chase"] + BFS["BFS thread: threadLoop"] -->|"shouldRunPass gate (see SignalsExplained §4-§8)"| RP["runPass"] + RP -->|"every pass"| MW["runPassManualWalk"] + MW -->|"cadence-gated (>= 2s apart)"| RE["IterateOverReachableObjects seeds roots"] + MW -->|"when the loaded-class set changed since the last completed sweep"| SF["admitStaticFieldRoots: app-classes-first, chunked, resumable sweep"] + MW -->|"candidates open: before any breadth-first work"| PRONGS["candidate-scoped reach: walkCandidateThreadLocals + walkStaticFieldAnchors descend walks"] + MW --> EF["expandFrontier: batched array-holder FollowReferences"] + EF --> FT["FrontierTable: FRONTIER / EXPANDED / EDGE / ABANDONED"] + MW -->|"reserved rotation slice"| ROT["rotation: leak-accumulation (Tier 1/2) + stale-EXPANDED + root-kind + static-anchor FIFO"] + ROT --> EF + CHASE -->|"heapReferenceCallback prunes a marker-tagged candidate: chain link recorded"| PWT + PWT -->|"buildChainEvent (representative, pruned marker, and up to 8 auto-discovered instances/class)"| RC["cacheResolvedChain: one entry per klass id, cap 128"] + RC -->|"Profiler::dump, snapshot without clearing"| DR["drainPendingChainEvents"] + DR --> JFR["datadog.ReferenceChain / datadog.ReferenceChainAbandoned"] +``` + +Every pass is driven by the manual walk (`runPassManualWalk()`): pure JVMTI +heap calls, which run inside the `VM_HeapWalkOperation` safepoint. The walk +reads no raw oop — JVMTI's iterators apply the active collector's own +barriers, so the walk is correct on every collector, including ZGC, where +concurrent relocation would corrupt a raw-oop reader. The phases below run in this order, each with its own +deadline slice (the per-sub-operation deadline is reset, so an earlier phase +cannot eat a later phase's slice): + +1. root/stack-ref enumeration (cadence-gated), +2. candidate thread-local descend walks (when a chase is open), +3. static-field sweep (when the loaded-class set changed), +4. `expandFrontier()` (ordinary breadth-first progress), +5. static-anchor descend walks, +6. rotation re-expansion. + +A rotation slice is reserved up front, before any of the above can spend the +whole pass budget. The reservation is capped at half the expand budget: a +pacing-throttled pass degrades both sides proportionally instead of starving +ordinary expansion to zero (or rotation to zero). + +## 2. Root discovery + +### 2.1 Root/stack-ref enumeration + +`IterateOverReachableObjects` walks the roots and dispatches +`heapRootCallback()`/`stackRefCallback()`. Root enumeration pays its fixed +root-walk-and-dispatch cost in full on every run, regardless of budget, so it +is gated by `ROOT_ENUM_MIN_INTERVAL_NS` (2s): it does not re-fire at the +per-second pass cadence. A budget-exhausted truncation here retries on the +next pass; a frontier-cap hit abandons the search (below). Note that root +enumeration alone never discovers a root's transitive children — the +callbacks are given no oop, only a tag pointer — so `expandFrontier()` is +always needed for further progress, first pass or resumed. + +### 2.2 Static-field roots (`admitStaticFieldRoots()`) + +An object held only by `SomeClass.staticField` is not reachable through +`IterateOverReachableObjects`' root/stack-ref callbacks at all. This is +precisely the leak shape production pods showed: a growing `static final` +collection. The sweep: + +- Repartitions the per-call `GetLoadedClasses()` array app-classes-first (any + non-bootstrap classloader), in place, every call — `GetLoadedClasses()` + gives no ordering guarantee across calls, so the sweep reprioritizes each + time. A likely leak source is reached within the first chunks instead of + after every JDK platform class. +- Advances `STATIC_FIELD_SWEEP_CHUNK_CLASSES` (512) classes per pass through + a resumable cursor (`_static_field_sweep_cursor`). A single un-chunked + `FollowReferences` over every loaded class could not finish inside one + pass's safepoint deadline on a JVM with tens of thousands of classes. +- Retries a truncated chunk on the next pass; a lap that truncated even once + is not marked done (`_static_field_sweep_cycle_truncated`), per the + subsystem's "no silent truncation" contract. +- Only runs at all when the loaded-class set has *actually changed* since the + last completed lap (`_last_static_field_class_count` differs from the count + `resolveLoadedClasses()` refreshed this same pass). Otherwise it would + re-pay a stop-the-world walk over every class on every pass, forever. +- Admits `STATIC_FIELD` edges always; caps non-static class edges + (constant-pool, interface, superclass, classloader) at + `STATIC_FIELD_SWEEP_NON_STATIC_CAP_PER_CLASS` (32) per class, so one outlier + class cannot blow the chunk's deadline. + +### 2.3 Candidate-scoped reach (descend walks) + +Once a candidate chase is open, breadth-first progress alone can leave the +tagged instances queued behind the ordinary backlog for many passes. Each +open chase therefore gets bounded *descend walks* — `FollowReferences` from a +specific anchor, up to `DESCENT_HOPS` (16) hops below it (raised from 6 after +pod round 7: the static-`ExecutorService` → queue → task → accumulator → +list → chunk shape is ~7 deep; the walk is deadline-bounded per slice either +way, so a deeper cap just lets each bounded walk cover the whole holder +interior — already-admitted entries are skipped, so repeated passes march +deeper each time): + +- **Prong 1, thread-locals** (`walkCandidateThreadLocals()`): descend-walks + the qualifying threads' `Thread` objects — registered via + `Profiler::onThreadStart/onThreadEnd` + (`registerThreadObject()`/`unregisterThreadObject()`, a mutex-guarded + tid → global-ref map; a running thread's `Thread` object is reachable via + the VM anyway, so the strong ref does not distort reachability) — through + their `ThreadLocalMap` subgraphs. A `Thread → ThreadLocalMap → table[] → + Entry → value → holder → chunk` chain is 5-6 hops that ordinary BFS may + reach very late, if at all. Capped at `THREAD_WALK_MAX_ANCHORS` (4) per + pass, rotated fairly across the (klass, tid) pairs by a cursor. +- **Prong 2, static anchors** (`walkStaticFieldAnchors()`): descend-walks + root-attached static holders. Anchors are selected by + `collectStaticFieldAnchorsForRotation()` with *tiered* per-tier cursors + (fresh container-interface holders first); classes are classified once by + `reconcileAnchorClassShapes()` (capped at + `ANCHOR_SHAPE_RECONCILE_BUDGET` = 128 distinct classes per pass — + `resolveContainerInterfaceTags()`/`classImplementsContainerOrMap()` decide + whether a class implements a `Collection`/`Map`-like interface); at-risk + holders (frontier entries whose parent died) flow through a FIFO + (`pushAtRiskStaticAnchor()`/`drainStaticAnchorFifo()`, popped + `STATIC_ANCHOR_FIFO_DRAIN` = 16 per pass). Capped at + `STATIC_ANCHOR_ROTATION_BUDGET` (32) anchors per pass. Selection size and + walked size are decoupled: a truncation requeues un-walked anchors + (`requeueStaticAnchorFifoFront()`), so a larger selection never extends the + pause — it can only spend selection-scan time, which holds no safepoint. + +Both prongs resolve their anchor batch with a single `GetObjectsWithTags` +call (the call's O(tag_map) cost is the dominant term on a large tag map — +the same batching rationale as `expandFrontier()`). + +## 3. Frontier expansion and batching + +`expandFrontier()` resolves up to `budget` pending tags in *one* +`GetObjectsWithTags` call, packs the live objects into a single JNI array +(`holder`), and expands the whole batch via exactly one +`FollowReferences(initial_object = holder)` — one BFS level per +VM-safepoint operation rather than one safepoint per entry: + +```mermaid +sequenceDiagram + participant BFS as BFS thread + participant JVMTI as JVMTI + participant JNI as JNI holder array + participant CB as heapReferenceCallback + + BFS->>BFS: pull up to budget tags from front of _pending_expand + BFS->>JVMTI: GetObjectsWithTags, resolve which tags are still live + JVMTI-->>BFS: live jobject references + BFS->>JNI: EnsureLocalCapacity, then NewObjectArray to build holder + alt exception, EnsureLocalCapacity failure, or null holder + BFS->>BFS: ctx.truncated = true, retry this batch next pass + else holder built successfully + BFS->>JNI: SetObjectArrayElement per resolved object + BFS->>JVMTI: FollowReferences, initial_object = holder + JVMTI->>CB: heapReferenceCallback per outgoing edge + CB-->>JVMTI: descend only for batch_tags boundary objects + JVMTI-->>BFS: one BFS hop expanded for the whole batch + BFS->>BFS: markExpanded, admitObject appends children to _pending_expand + end +``` + +The JNI/JVMTI error handling is defensive by construction: a null `holder`, +a pending JNI exception after `NewObjectArray`/`SetObjectArrayElement`, or an +`EnsureLocalCapacity` failure all set `ctx.truncated = true` (retry next +pass) instead of marking the batch permanently `EXPANDED`. A failed +`java/lang/Object` class resolution with pending work also forces +`truncated = true`, so `runPass()` cannot mistake it for +`SearchState::COMPLETED`. + +## 4. The canary candidate chase + +`pollWatchedTargets()` pre-tags each candidate's *specific representative +object* (identity match — matching by class alone could record a chain for an +unrelated, possibly short-lived, instance of the same class) with a distinct +negative marker tag (`MARKER_TAG_BASE - i`; the negative range is disjoint +from the positive frontier-tag counter and the negative class-tag range). +When the walk's `heapReferenceCallback()` later encounters one, it *prunes* +it: records the candidate's referrer link (parent tag, referrer klass, +depth) for `buildCanaryChainEvent()`, and does not expand through it. + +The walk also auto-marks *any* instance of a watched class it discovers (up +to `MAX_DISCOVERED_INSTANCES_PER_CLASS` = 8 per slot): a leaking class's +many live instances each have independently useful chains — different +parents, different retention paths — and the pre-tagged representative is +just one sample. + +The chase's *scheduling* (back-to-back on progress, work-scaled exponential +backoff on no progress, pain-budget refill) is covered in +`ReferenceChains-SignalsExplained.md` §8, including the measured pod incident +(32 minutes at ~88 passes/min) that motivated the backoff. The *termination* +half is here: a chase that has made zero candidate-discovery progress for +`CANARY_NO_PROGRESS_PASS_LIMIT` (30) consecutive passes is abandoned with +`SearchAbandonReason::CANARY_STUCK`. Unlike the TTL check, this fires even +while `isUrgent()` holds: a zero-progress chase at urgency-boosted +budget/cadence only burns STW pause budget the process needs during the same +OOM approach the chase was launched to diagnose. The abandon limit *widens* +with repeated stuck restarts (`_canary_stuck_restart_count` survives +`restartSearch()`), so a permanently un-findable candidate cannot keep +cycling at the base limit. + +## 5. Rotation: rediscovering growth in an already-visited container + +The frontier walk visits each object once. That is insufficient for the leak +shape this feature targets: a `static final` collection field that is +*appended to*, not reassigned. The container is already `EXPANDED` long +before its element klass earns a leak signal. `runPass()`'s completion branch +therefore requires `_watched_leak_klass_count == 0` (no klass under active +leak watch) before moving to `SearchState::COMPLETED`; while a klass is +watched, the search stays `RUNNING` and rotation re-queues the growing +container: + +```mermaid +flowchart TD + LT2["LivenessTracker::topKlassesByGenerationCount"] -->|"refreshed once per tick,
only after hasLeakSignal() fires"| WK["_watched_leak_klass_ids
(max 5)"] + WK -->|"klass_id newly watched"| SEED["seedLeakAccumulationForNewlyWatchedKlass:
one-time scan of already-EXPANDED
frontier entries by class_tag"] + ADM["admitObject: ADMITTED"] -->|"class_tag of new object"| TLA["trackLeakAccumulation"] + SEED --> TLA + WK -->|"class_tag match?"| TLA + TLA --> T1["Tier 1: _leak_signature_totals
(leaf_klass_id, parent_class_id) -> count"] + TLA --> T2["Tier 2: _leak_parent_fanout
parent_tag -> count, within winning signature"] + T1 -->|"delta vs previous pass's snapshot"| RANK["collectLeakAccumulationCandidatesForRotation:
pick winning signature, then its top parent_tag(s)"] + T2 --> RANK + RANK -->|"re-queue for re-expansion,
budget 16/pass"| EF2["expandFrontier"] + EF2 -->|"new elements admitted"| ADM +``` + +Matching a newly-admitted object against a watched klass id uses a stable +JVMTI class tag (`classTagAllocator.h`, shared between +`ReferenceChainTracker` and `LivenessTracker`) rather than the classMap +`StringDictionary` id — that id is not guaranteed stable if the dictionary +is compacted/regenerated mid-search (observed live: the same class resolved +to two different classMap ids from two subsystems). + +The rotation tiers, with their per-pass budgets: + +| Tier | Budget/pass | What it selects | +|---|---|---| +| Leak-accumulation (Tier 1/2 above) | 16 | The containers of currently-flagged klasses — the targeted tier | +| Stale-EXPANDED (own cursor) | 256 | Any long-`EXPANDED` entry — low-priority eventual coverage of the whole table; deliberately *not* scaled with table size (two earlier versions tried; see git history) | +| Root-kind (own cursor) | 16 | Transient-root entries, so attribution converges to a durable root kind | +| Static-anchor FIFO | 16 drained | At-risk holders whose parent died, into the same `walkStaticFieldAnchors()` batch | + +The stale-EXPANDED tier's own cursor exists so a frontier table full of +long-lived infrastructure objects cannot permanently starve higher-tag +entries of ever being re-queued. + +## 6. Termination + +```mermaid +stateDiagram-v2 + direction LR + [*] --> RUNNING + RUNNING --> COMPLETED: frontier drained,
no truncation this pass,
no leak klass still watched + RUNNING --> ABANDONED_TTL: wall-clock TTL exceeded
with work still pending
(suppressed while urgent) + RUNNING --> ABANDONED_CAP: frontier-size cap hit + RUNNING --> ABANDONED_CANARY: zero candidate-discovery progress for
CANARY_NO_PROGRESS_PASS_LIMIT (30) passes + ABANDONED_TTL --> RUNNING: restartSearch + ABANDONED_CAP --> RUNNING: restartSearch + ABANDONED_CANARY --> RUNNING: restartSearch + COMPLETED --> RUNNING: restartSearch +``` + +The abandon reason is recorded (`SearchAbandonReason`) and surfaced as the +`datadog.ReferenceChainAbandoned` JFR event — no silent truncation. +`releaseSearchTags()` clears every live tag the search still owns once it +ends, without discarding the frontier table's own records: chain +reconstruction keeps working from memory after the search ends. + +Restarts are gated on an actual leak indication (a +`selectLeakCandidates()` candidate, or an urgent-latched seconds-to-OOM +projection — one search per latched episode via `_urgent_search_spent`) plus +the safepoint pain budget. This closes a structural gap: a one-shot walk +could finish before population-trend detection accumulated enough GC epochs +to flag a candidate, leaving anything allocated afterward permanently +undiscoverable. See `ReferenceChains-SignalsExplained.md` §5-§6 for the +gate's full behavior. + +## 7. Pacing and budgets (as-built arithmetic) + +- **PID controller on genuine in-safepoint time.** `updatePacing()` is fed + only each pass's in-safepoint ticks — measured across every + `IterateOverReachableObjects`/`FollowReferences` call the pass makes, + explicitly excluding `GetObjectsWithTags` (not a safepoint call) and every + bookkeeping line in between. Non-safepoint bookkeeping must not be + mistaken for pause-time-SLO pressure. The controller's overflow term also + widens/narrows the fallback cadence. +- **Separate CPU pain budget.** Non-safepoint pass cost is charged to + `_cpu_pain_budget`, a second `PainBudget` sharing the same + `painbudget=N` percent knob as the safepoint pain budget — one + operator-facing "acceptable background cost" percentage covers both leaky + buckets. +- **Budget borrowing.** A sustained run of comfortably-under-target passes + (after a warmup) earns extra headroom above `_budget` + (`_borrowed_budget`); any pass that is not comfortably under target + revokes it immediately (`maybeRevokeBorrowForRootEnumPass()` also revokes + on a root-enum pass), so `_budget` itself stays the ceiling the instant + the search stops proving it has room. +- **First-pass budget.** The search's one-shot root-seeded first pass draws + its own much larger edge budget (`firstpassbudget`, auto-scaled 10× from + `budget`, capped) exactly once — a steady-state budget sized for cheap + incremental expansion would truncate a cold full-graph walk long before + it reaches anything interesting — and its duration is excluded from the + pacing signal so it cannot throttle every cheap pass that follows. +- **Urgency ramp.** See `ReferenceChains-SignalsExplained.md` §9: pause + target and cadence ramp exponentially toward their ceilings as + `secondsToOOM()` falls inside `OOM_RAMP_START_S` (30 min); the budget + ceiling is held at 4× for the ramp's entire duration; the ramp owns + `_effective_cadence_ns` outright while active. +- **Auto-tuned defaults.** `autoTuneDefaults()` scales the defaults for + budget (√heap-proportional), first-pass budget, TTL, frontier cap, and + pause target (capped at 50ms) with the resolved max heap / container + limit, but only for sub-options the operator did not set explicitly. + +The urgency *latch* is covered in `ReferenceChains-SignalsExplained.md` §9. +On the `LivenessTracker` side, `secondsToOOM()` itself: accounts for the +container memory limit (not just `-Xmx`) when projecting the exhaustion +point; requires a confirmed rising heap-floor trend; and corroborates the +projection with a recent-half slope check, so a single stale sample cannot +spike or collapse it. + +## 8. `LivenessTracker` admission boost (chase phase) + +While a candidate chase is open, allocations by the candidates' *qualifying +threads* are admitted to liveness tracking at 100% (`admitForTracking()`, +published two-phase — slots first, then the count with RELEASE — by +`noteSelectedCandidates()`), and urgency admits everything. Without this, +the default 10% liveness subsampling thins small per-(klass, tid) +populations enough to make leak-tag correlation intermittently fail on real +pods (observed live as the intermittent zero-tag runs in +`LeakTagCorrelationReferenceChainTest`'s own lottery analysis). The boost +only ever adds admissions on top of the configured ratio — fail-open by +construction: a stale or missed boost cannot drop an allocation the ratio +would have admitted. The draw itself is the process-wide xorshift64 stream +from main's sampling refactor (#794): per-thread TLS state, an integer +threshold compare, no `` machinery. + +## 9. Output path and JFR persistence + +A resolved chain is cached per klass id (`_resolved_chains`, capped at +`MAX_RESOLVED_CHAINS` = 128 entries, drop-not-evict once full — surfaced as +`REFERENCE_CHAIN_EVENTS_DROPPED`) and re-stamped into *every* subsequent +dump the sample survives into (`drainPendingChainEvents()` snapshots without +clearing, mirroring how `LivenessTracker` re-emits live-object samples). A +long-lived leak's chain is therefore present in each JFR chunk, not only in +the chunk active when it was first reconstructed. The write happens on the +dump thread (`Profiler::dump()`), never on the BFS thread — the walk never +blocks on JFR I/O, and JFR writes never trigger a walk. + +## 10. Configuration + +`referencechains=true:hops=N:budget=N:ttl=N:framecap=N:pausetarget=N:painbudget=N:firstpassbudget=N` + +- Negative `hops`/`budget`/`framecap` values are floored (an unfloored + negative `hops` would wrap to ~4e9 as a `u32`, silently disabling the cap + it is meant to enforce); all three are also ceiling-clamped. +- `ttl <= 0` disables the wall-clock TTL cutoff. +- Unset values are auto-tuned (see §7). + +## 11. Why the tuning pass is not a JMH benchmark + +The defaults above are placeholders pending a measurement pass. That pass +cannot be a JMH benchmark, for three reasons: + +1. **The cost is not per-Java-operation.** JMH measures the throughput and + latency of a benchmark method. This subsystem's cost is (a) STW + safepoint pauses from `VM_HeapWalkOperation` — global stalls + attributable to no benchmark method — and (b) CPU on a dedicated native + background thread. Both reach a Java workload only as indirect + throughput degradation mixed with GC noise; JMH can neither observe + nor attribute them directly. +2. **Activation is leak-signal-gated.** No pass runs until the population + rings fill (~10 GC epochs), trend hysteresis clears (5 consecutive + qualifying epochs), and the search gate opens. A seconds-scale JMH + iteration measures the feature idle. Exercising the chase needs minutes + of continuous leaking allocation; run-to-run variance is then dominated + by GC cadence and by *when* hysteresis cleared — exactly the steady-state + assumption JMH's fork/iteration statistics make. +3. **The knobs control quantities JMH cannot see.** `pausetarget`, + `budget`, cadence, backoff, and the pain budgets regulate per-pass + safepoint duration, pass-cost EMA, and refill rates — all directly + observable as `jdk.ExecuteVMOperation[HeapWalkOperation]` JFR durations + and the tracker's own pass telemetry, which is what the shipped `utils/` + repro/sweep/report tooling consumes. + +A coarse JMH A/B (leaking workload, `referencechains` off vs on) remains +possible with the repo's existing `ddprof-stresstest` JMH setup, and would +serve as an end-to-end throughput regression guard. It cannot tune these +defaults. Tuning them requires the JFR-based measurement pass above. diff --git a/doc/architecture/LiveHeapReferenceChains.md b/doc/architecture/LiveHeapReferenceChains.md new file mode 100644 index 0000000000..5276da3f6b --- /dev/null +++ b/doc/architecture/LiveHeapReferenceChains.md @@ -0,0 +1,610 @@ +# Reference Chains for Surviving Live Heap Samples + +**Status:** Implemented (see "Implementation status" below) +**Date:** 2026-07-07 +**Jira:** [PROF-15341](https://datadoghq.atlassian.net/browse/PROF-15341) + +## Implementation status + +The "Chosen design" section below has been implemented following +`LiveHeapReferenceChains-ImplementationPlan.md` (kept locally, not committed) +(Phases 0-7). It is off by default; the shipping switch is the `referencechains` argument +parsed by `Arguments` (`arguments.cpp`'s `CASE("referencechains")`), e.g. +`referencechains=true:hops=64:budget=2000:ttl=60000:framecap=65536`. + +Read this status note alongside the actual code before relying on it, not instead of it: + +- **The BFS engine (frontier table, tag lifecycle, incremental resumption, termination, + JFR event shapes) is implemented and unit-tested** (`ddprof-lib/src/main/cpp/referenceChains.h`/ + `.cpp`, `ddprof-lib/src/test/cpp/referenceChains_ut.cpp`). +- **The lifecycle gap is closed: it now runs inside a live profiling session.** + `Profiler::start()` (`profiler.cpp`) calls `ReferenceChainTracker::instance()->start(args)` + (gated on `args._reference_chains`, independent of the CPU/wall/alloc engine mask, the + same way `malloc_tracer`/`NativeSocketSampler` are gated on their own flags) followed by + the new `ReferenceChainTracker::startThread()`, which spawns the BFS thread + (`threadLoop()`) - safe there because the JVM/JVMTI environment is already fully up by + that point in the lifecycle, unlike inside `start()` itself, which must stay callable + with no live JVM for `referenceChains_ut.cpp`'s tests. `Profiler::stop()` calls the + matching `stopThread()`/`stop()` pair. Because `start()` runs, `SetEventNotificationMode` + for the GC callbacks is now actually invoked, so `onGCStart()`/`onGCFinish()` fire and the + BFS thread's `shouldRunPass()` scheduling loop (GC-epoch signal or the fixed cadence) is + live. A `datadog.ReferenceChainAbandoned` event now reaches a real `Recording`: when a + dump is requested (`Profiler::dump()`, the same call site that already flushes + `LivenessTracker`) and the search's state is `SearchState::ABANDONED`, + `buildAbandonedEvent()`'s output is written via the new + `Profiler::writeReferenceChainAbandoned()` / `FlightRecorder::recordReferenceChainAbandoned()` + wrappers (mirroring `writeHeapUsage()`'s exact shape). +- **`buildChainEvent()` now has a call site: the target-selection feed is closed.** + `ReferenceChainTracker::pollWatchedTargets()` (referenceChains.cpp), called from + `threadLoop()` once per scheduling cycle after `runPass()`, is that feed. It polls + `LivenessTracker::selectLeakCandidates()` (the positive population-slope ranking, Open + Question 3 below) and, for each ranked klass's live representative instance that an + ordinary `runPass()` walk has *already* tagged (`getTag() > 0` - a read, never a `SetTag` + seed; see Open Question 3 for why the design's original seeding proposal was replaced), + calls `buildChainEvent(tag, ...)`, which is cached (`cacheResolvedChain()`) rather than + written immediately. The actual JFR write happens later and on a different thread: enqueued + via `enqueueChainEvent()` and drained by `Profiler::dump()` -> `drainPendingChainEvents()` -> + `Profiler::writeReferenceChain()` / `FlightRecorder::recordReferenceChain()` (`profiler.cpp` + lines ~1960-1974), decoupling the walk from JFR I/O. Deduplication is `_resolved_chains` + (an `unordered_map` keyed by klass_id, `referenceChains.h`), not a + per-search tag set, so a klass flagged across consecutive polls emits its chain only once. + This is gated on `_gc_generations` *and* + liveness tracking both being enabled - `referencechains=...` alone still gets the + whole-graph-only behavior (no target seeding), resolving Open Question 3's "still + undecided" fallback. See + `ddprof-test/src/test/java/com/datadoghq/profiler/referencechains/ReferenceChainTrackingTest.java`: + `shouldReportAbandonedSearchOnTinyFrontierCap` exercises the abandonment path. + `shouldReconstructReferrerChainToGcRoot` (Phase E's own exit criterion) is no longer + `@Disabled`: it allocates a growing, real population of a fixture class, drives GCs and + `Profiler::dump()` calls until `LivenessTracker::selectLeakCandidates()` trusts the resulting + trend, then asserts on the `datadog.ReferenceChain` event `pollWatchedTargets()` produces - + a real end-to-end exercise of this whole mechanism against a live JVM, not a synthetic + frontier fixture. +- **Phase 5's tuning defaults are provisional, not empirically finalized.** The hop cap, + per-pass budget, TTL, and frontier-size cap (`arguments.h`'s `DEFAULT_REFERENCE_CHAINS_*` + constants) are explicitly-labeled placeholders; no benchmark against this codebase has + run yet (see Open Question 2 below and the implementation plan's Phase 5). + +## Goal + +For a subset of live-heap samples that survive past their allocation window, produce +a **reference chain** — a sequence of referrer *types* (not full field-level paths, not +necessarily to *all* GC roots) connecting the sampled instance back to *a* GC root. This +is diagnostic information ("what kind of object chain is keeping this alive"), not a +heap-dump-grade exact retainer analysis. + +## Constraints + +- Must run cheaply, with as short a safepoint / STW contribution as possible. +- Must work on stock vendor JDKs the agent attaches to — no forked/patched JVM builds. +- Exhaustive (all-roots, full-path) chains are explicitly **not** required; referrer-type-only, + bounded-depth, best-effort chains are acceptable. + +## Approaches considered + +Three approaches were evaluated; two are ruled out as launch requirements for concrete, +evidence-backed reasons. One sub-idea (Approach C's `ParallelObjectIterator` variant) is +explicitly kept open as a conditional future option; see its discussion below. + +| # | Approach | Completeness | Complexity | Feasibility | Status | +|---|---|---|---|---|---| +| A | Full JVMTI `FollowReferences` reverse-graph walk, piggybacked on an already-scheduled major GC | 4/5 | 4/5 | 2/5 | Rejected | +| B | Bounded BFS-from-roots with frontier pruning (JFR "leak profiler" technique, adapted) | 3/5 | 3/5\* | 4/5 | **Chosen** | +| C | Hook GC mark/copy closures (G1, ZGC) to record parent pointers inline during marking | 2/5 | 5/5 | 1/5 | Rejected | + +Scale (1-5 for each column): Completeness — higher is more complete (5 = closest to exhaustive all-roots/full-path); Complexity — higher is more complex to implement/maintain (5 = most complex, lower is better); Feasibility — higher is more feasible to ship on stock vendor JDKs (5 = most feasible). No single column dominates the decision; see the per-approach rationale below for why B was chosen despite not scoring highest on every column. + +\* This 3/5 reflects only the single-pass BFS sketched at selection time. The "Chosen +design" section below replaces that sketch with an incremental, resumable BFS +(JVMTI-tag-based frontier persistence across GC cycles, an agent-thread-driven pass that +gets its safepoint transparently from the JVMTI heap-walk call it makes, GC-callback +signaling, and explicit termination/tag-cleanup bookkeeping), which is materially more +complex than this score suggests — closer to 4/5 in implementation and maintenance +effort. The score is left unchanged above (it documents the state of the comparison at +decision time) rather than retroactively edited. + +### A — Full reverse-reachability walk (rejected) + +Safepoint length scales with live-set size regardless of how the walk is triggered. +Modern regionalized collectors (G1, Shenandoah) rarely perform a true full-heap walk +during ordinary major GCs, so "ride an already-paid pause" is not a reliable amortization +strategy. Cost is fundamentally at odds with the "short safepoint" constraint. + +### B — Bounded BFS-from-roots (chosen) + +Mirrors OpenJDK's own `jdk.OldObjectSample` leak-profiler implementation +(`src/hotspot/share/jfr/leakprofiler/chains/{edgeStore,bfsClosure,dfsClosure}.cpp`): +a `VM_Operation`-driven BFS from GC roots, retaining only edges on the frontier toward a +small, fixed sample set, with a hard hop cap (HotSpot itself caps chains at ~200 hops, +split 100/100 from leaf and from root). We can go cheaper than JFR because only the +**referrer class**, not object identity or field name, is needed — the `EdgeStore` +degenerates to `(referrer_klass, parent_tag, depth)` records, where `parent_tag` links +each record back to the record that discovered it, enabling chain reconstruction. + +Adopting this pattern is a re-scoping of proven, shipping HotSpot code, not a novel +algorithm design. + +### C — GC mark/copy closure piggyback (rejected) + +Investigated specifically for G1 and ZGC on the premise that per-edge referrer +information is already available inside the collector's own marking/evacuation closures +(`G1ParCopyClosure::do_oop_work`, ZGC's `ZMarkConcurrentRootsIteratorClosure` / +load-barrier closures), so recording it would cost nothing beyond what the GC already +pays. + +Rejected because there is no stable, externally reachable hook into these closures: + +- They are internal, template-instantiated C++ classes compiled into `libjvm.so` at + HotSpot build time — not a registrable/pluggable extension point. +- This differs categorically from `VMStructs`-style introspection already used in this + codebase (`ddprof-lib/src/main/cpp/hotspot/vmStructs.cpp`), which reads VM state + passively via an officially exported offset table. Intercepting a GC closure's + *behavior* would require either shipping a patched OpenJDK build (a fork/maintenance + commitment far beyond anything in this codebase) or binary-patching unversioned, + per-build-mangled function addresses — not shippable across JDK point releases. + +A related idea — using HotSpot's internal `ParallelObjectIterator` +(landed via [JDK-8322043](https://www.mail-archive.com/serviceability-dev@openjdk.org/msg12977.html), +used by `VM_HeapDumper` to partition heap regions across GC worker threads for parallel +heap dumping) to shrink Approach B's safepoint by parallelizing the walk — was also +investigated. Same verdict: it is an internal C++ class, not exposed via JVMTI, with no +stable ABI for an attached agent to call. Symbol-sniffing internal HotSpot functions *is* +an established pattern in this codebase (`VMStructs::findHeapUsageFunc`, +`vmStructs.cpp:489-509`), but that precedent covers a single leaf virtual method with a +value/POD-ish return; `ParallelObjectIterator` is a multi-class subsystem that coordinates +the VM's own GC worker threads under safepoint control — an order of magnitude larger +fragility surface, with a much higher blast radius if a layout assumption is wrong (GC +worker-thread coordination corruption vs. a bad JMX stat). Not pursued as a launch +requirement; revisit only if Approach B's single-threaded pause proves to be a measured +bottleneck, and treat it as an isolated, heavily version/flag-gated fast path with +automatic fallback — never a dependency. + +## Chosen design: incremental, resumable bounded BFS + +A single-pass bounded BFS still means one pause sized to whatever budget is configured. +The refinement below spreads that budget across multiple short passes instead of one +contiguous one, trading a possibly-higher *aggregate* STW total for a much better +*latency distribution* — no single long tail pause. + +### Why the frontier can survive across passes: JVMTI object tags + +The obstacle to pausing and resuming a BFS is that the frontier (the worklist of +not-yet-expanded objects) is normally a set of raw addresses, and a moving/compacting GC +between passes can relocate or collect any of them. + +JVMTI object tags solve this: + +- Tags are identity-based and GC-move-transparent — a tagged object can be re-resolved + after a GC regardless of where it moved. +- Tags are **non-retaining** — tagging does not keep an object alive. This is a *new* + subsystem dependency, not a reuse of one: the existing live-object sampler + (`LivenessTracker`, `livenessTracker.cpp`) does not use JVMTI tags at all — it + correlates sampled objects via JNI weak global references (`NewWeakGlobalRef`) held in + its own index table with its own locking and GC-triggered cleanup + (`livenessTracker.cpp:327-357`, `:53-70`). Non-retention is a property both mechanisms + happen to share, not evidence that this reuses proven infrastructure. +- **Investigated replacing tags outright with `LivenessTracker`'s weak-ref + index-table + pattern — resolved as "adopt the table pattern, keep the tags."** The pattern cannot + fully substitute for tags: `FollowReferences`/`IterateThroughHeap` (the calls that + actually discover a frontier object's referrers) can filter/report against a *tagged* + object set natively; a JNI weak-ref table has no hook into that machinery, so the + frontier-discovery step would still need tagged objects regardless of what stores the + metadata. What the investigation *does* carry over: `LivenessTracker`'s proven + `TrackingEntry`-style slot table — CAS-based index allocation, a signal-safe `SpinLock` + (`spinLock.h`), doubling-resize, and GC-epoch-triggered cleanup + (`livenessTracker.cpp:152-176`, `:213-278`, `:369-409`) — is a better-precedented design + for the frontier's *metadata* storage than inventing one from scratch, since a JVMTI tag + is a single `jlong` with no room for `(parent_tag, referrer_klass, depth)` on its own. + Recommendation: use the tag as an index into a `TrackingEntry`-style table (fields: + `parent_tag`, `referrer_klass`, `depth`) rather than encoding all three into the tag + value or a from-scratch hashmap. See "Frontier metadata storage" below and Open + Question 4. +- Non-retention gives incremental resumption a useful side effect for free: if a frontier + object dies between passes, it simply fails to re-resolve on the next pass. That branch + of the search is pruned automatically, with no extra liveness bookkeeping required. + +### Data structures + +- **Frontier**: a set of `(tag, parent_tag, referrer_klass, depth)` records. `tag` is the + JVMTI tag assigned to a not-yet-expanded object; `parent_tag` links back for chain + reconstruction; `depth` supports the hop cap. +- **EdgeStore**: accumulates `(referrer_klass, parent_tag, depth)` per discovered edge for + objects that are on a path toward a target sample. Keyed by tag, not address — + degenerate relative to JFR's `EdgeStore` since object identity/field names are not + required, but it retains the same `parent_tag` linkage field as the Frontier so a chain + can be walked back from a target sample to a root by following `parent_tag` across + EdgeStore records. + +### Frontier metadata storage: reusing `LivenessTracker`'s table pattern + +A JVMTI tag is one `jlong` — it can identify a frontier object and make it visible to +`FollowReferences`/`IterateThroughHeap`, but it cannot itself hold the three fields +(`parent_tag`, `referrer_klass`, `depth`) each Frontier/EdgeStore record needs. Two ways +to close that gap were considered: + +1. Build a bespoke hashmap keyed by tag value, from scratch. +2. Reuse `LivenessTracker`'s existing slot-table design (`livenessTracker.h:21-30` + `TrackingEntry`, `livenessTracker.cpp:152-176` sizing, `:213-278`/`:369-409` CAS slot + allocation and doubling resize, `spinLock.h`'s signal-safe `SpinLock`), using the tag + value as the slot index instead of a `jweak` as the identity handle. + +(2) is the better-precedented choice — it's shipping code, already exercises the exact +"per-slot payload, GC-cycle-driven cleanup, contention-safe locking" shape this needs — +provided the sizing formula is **not** copied as-is. `LivenessTracker` sizes its table +from `max_heap / sampling_interval` (a flat allocation-sample rate, +`livenessTracker.cpp:152-176`, capped at `MAX_TRACKING_TABLE_SIZE = 262144`, +`livenessTracker.h:39`); a BFS frontier's width is driven by per-hop fan-out in the object +graph, not by an allocation rate, and multiple concurrent searches (one per live-heap +sample being chased, see Open Question 3) each need their own capacity — the existing +formula does not transfer and a new one is an open question (folded into Open Question 2). + +### Algorithm + +1. Seed the frontier from GC roots (first pass) or from the persisted frontier + (resumed pass). +2. Resolve currently-live tagged frontier objects. Objects that fail to resolve are + dropped (dead — free pruning). +3. Expand the frontier up to a fixed per-pass budget (edge count or time slice). +4. Newly discovered objects are tagged and added to the frontier for the next pass. +5. Persist the frontier (native memory owned by the agent, not thread-local scratch) and + return control to the VM. +6. Repeat until: a target sample is reached, the hop cap is hit, or a per-search + abandonment limit (see Termination) is exceeded. + +### Triggering passes: resolved — the profiler never schedules its own safepoint + +Investigated whether pass-continuation work could ride the JVMTI +`GarbageCollectionStart`/`GarbageCollectionFinish` callbacks — the same callback this +codebase already uses to call `_heap_usage_func` (`vmStructs.cpp`) — instead of each pass +paying for its own safepoint. + +**Correction to an earlier framing in this doc**: a pass does not run "inside a dedicated +`VM_Operation::doit()`" that the profiler constructs — HotSpot's `VM_Operation`/ +`VMThread::execute()` machinery is internal, unexported C++ with no agent-facing entry +point; nothing outside HotSpot can submit one. What actually happens, confirmed against +`src/hotspot/share/prims/jvmtiTagMap.cpp` and `jvmtiEnv.cpp`: +`SetTag`/`GetTag` need no safepoint at all — they take only a `MutexLocker` over a +JVM-internal "hot lock" on the tag map. `FollowReferences`/`IterateThroughHeap` **do** +bring the VM to a safepoint, but the JVM does this internally and transparently +(`VM_HeapWalkOperation`/`VM_HeapIterateOperation`, dispatched via +`VMThread::execute()` *inside* HotSpot's own implementation of those calls) the moment an +ordinary attached agent thread calls them — the calling thread simply blocks until the +walk finishes. A pass is therefore: an agent-owned, already-attached thread (the same kind +`LivenessTracker` already runs on, see `livenessTracker.cpp:303-409`) calling +`FollowReferences`/`IterateThroughHeap` directly; the safepoint is a side effect of that +call, not something the profiler builds or schedules. + +**Confirmed the VM is genuinely at a safepoint (all mutators stopped) for the full +duration of both callbacks**, on every collector: + +- JVMTI spec: *"This event is sent while the VM is still stopped... the event handler + must not use JNI functions and must not use JVM TI functions except those which + specifically allow such use (see the raw monitor, memory management, and environment + local storage functions)."* +- openjdk/jdk source: delivery is synchronous on the VMThread + (`src/hotspot/share/prims/jvmtiExport.cpp:2752-2790`, comment *"this event is posted + from VM-Thread"*); every call site is inside a safepoint-executing `VM_Operation::doit()`, + backed by explicit asserts — e.g. Parallel GC's + `assert(SafepointSynchronize::is_at_safepoint())` (`gc/parallel/psScavenge.cpp:305-306`), + G1's `assert_at_safepoint_on_vm_thread()` (`gc/g1/g1VMOperations.cpp:141-157`), + Shenandoah and ZGC wrapping the same `SvcGCMarker` only inside their respective + `VM_Operation`/`VM_ZOperation::doit()` paths. Stable JDK 11 → mainline, across + Serial/Parallel/G1/Shenandoah/ZGC. + +**But this does not make the GC-triggered callback itself usable as the execution vehicle +for a pass.** The "functions which specifically allow such use" are exactly two: +`Allocate` and `Deallocate` (the entire **Memory Management** category). `SetTag`, +`GetTag`, `GetObjectsWithTags`, `FollowReferences`, and `IterateThroughHeap` are all in +the **Heap** category, which is *not* on that allowlist — calling any of them from inside +`GarbageCollectionStart`/`Finish` is exactly what the restriction forbids. The spec's own +prescribed escape hatch — notify a raw monitor from the callback, do the real work on a +separate agent thread — doesn't preserve "the pass rides the GC's own pause" property +either: the woken agent thread runs after the GC's collection pause has already ended +(mutators resumed), so calling `FollowReferences`/`IterateThroughHeap` there triggers a +**new**, separate safepoint of its own (per the corrected mechanism above) rather than +reusing the GC's. + +The only way to fold the tag/walk work into the GC's own STW window would be to bypass +the official JVMTI entry points and reach into HotSpot's internal `JvmtiTagMap` directly +via symbol-sniffing — reintroducing exactly the fragility class already rejected for +Approach C (unversioned internal C++ state, no stable ABI). Doing that here would undo +the reason C was rejected. + +**Conclusion: "no new marginal safepoints" is not achievable while staying within +official JVMTI usage.** Each pass still triggers its own safepoint — transparently, via +whichever agent thread calls `FollowReferences`/`IterateThroughHeap` for that pass, not +via anything the profiler schedules itself. The GC callbacks remain useful only as a +low-cost *signal* ("a GC just happened, a pass may be worth running soon") — not as the +execution vehicle for the pass itself. This does not change the core incremental design +(frontier persistence via JVMTI tags, self-pruning of dead branches, per-pass budget) — +it only removes the "zero marginal safepoints" claim from the cost/benefit case. The +design's actual value remains what it was framed as: trading one long pause for several +short, independently-triggered ones — a latency-distribution improvement, not a +total-STW reduction. + +### Termination and abandonment + +Because passes are spread across a mutating heap, a search that never reaches a root or +the hop cap could otherwise persist indefinitely, accumulating abandoned frontier state +across GC cycles. Required cutoffs: + +- Hop cap (as in Approach B's single-pass form). +- A hard cap on passes-per-search or wall-clock TTL from first observation (value TBD — + see Open Question 2). +- Explicit reporting of abandoned searches (no silent truncation) so this shows up as a + measurable "chain not found within budget" outcome rather than being indistinguishable + from "no chain exists." +- Tag release: on abandonment or completion, every JVMTI tag this search assigned to + frontier/`EdgeStore` objects (`SetTag(obj, 0)`) must be cleared before the search's + state is discarded. Without this, an abandoned search leaves its tags in place + indefinitely, directly aggravating the tag-table sizing/contention risk raised in + Open Question 4. +- A hard cap on frontier size (record count or native-memory footprint). The hop cap and + pass/TTL cap bound how *long* a search runs, but not how *wide* the frontier can grow + within that time — a wide fan-out graph could accumulate an unbounded number of + `(tag, parent_tag, referrer_klass, depth)` records before either cutoff is hit. When + the cap is reached, stop admitting new frontier entries for that search and report it + as an abandoned/truncated search (value TBD — see Open Question 2). + +### Correctness note: chains are historical, not a single consistent snapshot + +A chain built across multiple passes stitches together `"A referenced B"` facts observed +at different points in time, not one frozen graph. For the stated purpose — explaining, +by referrer type, what typically retains this class of surviving object — this is +sufficient, and is not meaningfully weaker than a single-pass walk: GC roots (e.g. thread +stack frames) are themselves a live-changing set across a single pause's boundary, so +"one true snapshot" is already an approximation in the single-pass case. Any +documentation or output surface built on this must describe results as an **observed** +retaining path, not a claim about the object's current exact retention state. + +### Cost/benefit summary + +- **Does not reduce total STW time.** Each safepoint/callback entry pays fixed + synchronization overhead; K short increments likely sum to equal or *more* aggregate + pause time than one contiguous walk covering the same work. +- **Improves latency distribution.** No single long tail pause — the thing most likely to + actually affect deployed application health (p99 latency, heartbeat timeouts), even + when total accumulated pause-ms is flat or slightly worse. + +## Non-goals + +- Exhaustive paths to all GC roots. +- Field-level or object-identity-level chains (referrer *type* only). +- Any GC-internal-closure hook (Approach C) or internal parallel-iteration API use as a + launch dependency. + +## Open questions before implementation + +1. ~~Confirm `GarbageCollectionStart`/`GarbageCollectionFinish` callback timing relative to + safepoint release.~~ **Resolved** (see Triggering section): the callback is genuinely + at a safepoint, but the JVMTI Heap-category functions needed to do frontier work + (`SetTag`/`GetTag`/`FollowReferences`/`IterateThroughHeap`) are not in the callback's + allowed function set. Each pass instead triggers its own safepoint transparently, via + whichever agent thread calls `FollowReferences`/`IterateThroughHeap` for that pass — + the profiler never constructs a `VM_Operation` itself (see correction in Triggering + section). The "no new marginal safepoints" framing is dropped; the design's value is + latency distribution, not total-STW reduction. +2. Choose per-pass budget defaults (edge count vs. time slice), hop cap, the + passes-per-search/wall-clock TTL abandonment cutoff, and the frontier-size/memory cap + (see Termination and "Frontier metadata storage") — needs measurement against + representative heap shapes and per-hop fan-out, not a guess; `LivenessTracker`'s + flat-sample-rate sizing formula does not transfer to a graph-search frontier. + **Not resolved — provisional defaults only, no measurement has occurred.** The + implementation currently ships explicitly-labeled "provisional default pending Phase 5 + empirical tuning" constants (`arguments.h`: `DEFAULT_REFERENCE_CHAINS_HOP_CAP = 200`, + citing this doc's own JFR ~200-hop/100-100 precedent; `DEFAULT_REFERENCE_CHAINS_BUDGET + = 1000`; `DEFAULT_REFERENCE_CHAINS_TTL_MS = 60000`; `DEFAULT_REFERENCE_CHAINS_FRONTIER_CAP + = 65536`, sized as a fraction of `LivenessTracker::MAX_TRACKING_TABLE_SIZE` rather than + derived from any BFS-specific measurement; plus `referenceChains.h`'s + `FrontierTable::INITIAL_TABLE_CAPACITY = 1024` and + `ReferenceChainTracker::PASS_CADENCE_NS` = 1 s). These let the subsystem run and be + tested end-to-end, but none are backed by a benchmark against this codebase — do not + describe them as measured. The real resolution path is + `LiveHeapReferenceChains-BenchmarkPlan.md` (kept locally, not committed), + which specifies the JMH/async-profiler matrix and decision rule Phase 5 still needs to + execute; this question stays open until that plan is actually run. + + **Pause-time-SLO feedback loop — SHIPPED, reusing the existing `PidController`.** + Implemented in + `LiveHeapReferenceChains-RemainingWorkPlan.md` (kept locally, not committed)'s + Phase D (`ReferenceChainTracker::updatePacing()`, `referenceChains.cpp`). This does not + replace the hop/TTL/frontier-cap constants raised in the first half of this question — only + the per-pass edge-count budget and the pass cadence, per the shipped mechanism below. + - New config sub-option `referencechains=...:pausetarget=` (`arguments.cpp`'s + `CASE("referencechains")` parser, field `Arguments::_reference_chains_pause_target_ms`), + defaulting to `arguments.h`'s `DEFAULT_REFERENCE_CHAINS_PAUSE_TARGET_MS = 5` — explicitly + labeled provisional/un-benchmarked, the same way the other + `DEFAULT_REFERENCE_CHAINS_*` constants are. Choosing its real value is still a Phase-5-style + empirical question, not resolved by this mechanism landing. + - `ReferenceChainTracker::start()` (re)constructs its own `PidController` instance + (`_pause_pid`) targeting `_pause_target_ms`, with its own gain triple — + **not** `ObjectSampler`/`MallocTracer`'s shared, uncited P=31/I=511/D=3/cutoff=15s + (`NativeSocketSampler` uses its own `RateLimiter`, not this `PidController` triple, and + was never part of the shared instance). The caveat this question raised about that triple + being copy-pasted, not independently derived, still stands for those two; it was not + "resolved," just not repeated here. The new instance uses P=10/I=1/D=2/window=1/cutoff=5s + — smaller proportional gain than the shared triple because a pass-duration-ms error is + single/low-double-digit in magnitude, unlike the shared triple's event-count scale + (`referenceChains.cpp`'s `start()`, inline comment on each gain). Gain *convergence* is + verified by gtest (three `ReferenceChainsTest` cases: steady-state at the ceiling, over- + ceiling, under-ceiling — see Phase D's exit criteria below), not by a live benchmark + against representative heap shapes; that remains a + `LiveHeapReferenceChains-BenchmarkPlan.md` (kept locally, not committed) item, + not fully closed by this mechanism landing. + - Measurement point: `runPass()` (`referenceChains.cpp`) times its own root + `IterateOverReachableObjects` call (first pass) or `expandFrontier()`'s + `GetObjectsWithTags`+`FollowReferences` pair (resumed pass) — already the thread blocked + inside the safepoint those calls trigger (Triggering section) — and converts to whole + milliseconds before feeding `_pause_pid.compute()` (matching every other `PidController` + caller in this codebase, which all feed integer counts). + - `updatePacing(u64 pass_wall_ns)` folds the budget-scaling and cadence-widening/relaxing + decisions into that one `compute()` call per pass: the signal is added to + `_effective_budget` (the *value* `runPass()` now passes to `expandFrontier()` instead of + the fixed `_budget`) and clamped to `[MIN_EFFECTIVE_BUDGET = 50, _budget + _borrowed_budget]` + — a later addition lets the ceiling temporarily borrow above `_budget` (up to + `BORROW_CEILING_MULTIPLIER`x) rather than capping hard at the config value; see the + rotation-budget-starvation fix in git history for why. Whatever the clamp could not + absorb (`overflow`) drives `_effective_cadence_ns` (see Open + Question 5 below for why cadence is folded into this same output rather than a second + controller): `CADENCE_NS_PER_EDGE_OVERFLOW = 1ms/edge` scales the unabsorbed overflow into + a nanosecond adjustment, widening `_effective_cadence_ns` toward + `MAX_EFFECTIVE_CADENCE_NS = 4s` when still over-ceiling even at the budget floor, relaxing + it toward `MIN_EFFECTIVE_CADENCE_NS = 10ms` when comfortably under-ceiling even at the + budget ceiling. + - The hop cap and the frontier-size hard cap (Termination section) are untouched by this + mechanism — they stay fixed correctness/memory-safety bounds, not controller-tuned, exactly + as this question originally specified. + - One known, deliberate scope limit carried over from the plan: `buildAbandonedEvent()`'s + `datadog.ReferenceChainAbandoned` event still reports the static config ceiling `_budget`, + not the adaptive `_effective_budget` — changing that event's semantics was out of Phase D's + stated scope. +3. Decide the sample-batching policy: one incremental search per live-heap sample, or + batched multi-target BFS sharing a single frontier walk (batching amortizes better but + couples unrelated samples' termination conditions together). + **Shipped, but not in either form this question anticipated.** The implemented + `ReferenceChainTracker::runPass()` (referenceChains.cpp) does not target any sample at + all: it runs a single, singleton-owned search that walks the whole root-reachable graph + (bounded by the hop/budget/frontier caps) with no per-sample seeding. Reconstructing a + chain for a specific tag is a separate, read-only step (`buildChainEvent(target_tag, ...)`) + applied after (or during) that one shared search - closer in spirit to "batched" (one + frontier walk can answer for many targets) than "one search per sample", but arrived at + by omission (the target-sample feed did not exist at that time, see the implementation + plan's Phase 7 report) rather than a deliberate batching design. + Whether this generalizes to true multi-target batching (explicit seeding from multiple + samples, coordinated termination) is still open and deferred, consistent with this + question's original framing. + + **Target-selection policy — SHIPPED (positive population-slope ranking).** Implemented in + `LiveHeapReferenceChains-RemainingWorkPlan.md` (kept locally, not committed)'s + Phases A-C. The missing piece above was *which* tag(s) `buildChainEvent()` should + reconstruct for. As shipped: per klass, `LivenessTracker` tracks a rolling window of its + live tracked-instance population count, sampled once per `LivenessTracker::cleanup_table()` + epoch advance (the same GC-epoch cadence that already recomputes survivor status, + `livenessTracker.cpp` — a *different*, slower cadence than `ReferenceChainTracker`'s own BFS pass cadence, Open + Question 5; "past N passes" below means GC epochs observed by `LivenessTracker`, not BFS + passes). A klass whose population trend over that window is positive — new instances + arriving faster than old ones are dying — is a leak candidate; a bounded cache/pool also + holds old objects but its population stabilizes or shrinks. Of all klasses with a + positive trend, seed only the top 3–5 by trend magnitude for the next BFS pass — this + doubles as the per-pass seeding cap this question's last bullet asks for, so no separate + budget constant is needed. + + Mechanics (as shipped): + - `LivenessTracker` resolves the class lazily at JFR-flush time (`flush_table()`) to keep + the allocation-sampling path free of a `GetObjectClass` call. Per-klass population is + computed by resolving the klass once per *surviving* entry inside the existing + `cleanup_table()` epoch-advance pass instead (`resolveKlassId()`, the same + `GetObjectClass` + `Class.getName()` + `Profiler::lookupClass()` sequence + `flush_table()` uses) — cost scales with live-table size × GC frequency, not with + allocation rate, so it does not touch the hot sampling path. This whole step is gated on + `_gc_generations` (`cleanup_table()`), so a plain liveness session pays none of it. + - A fixed-capacity table (`KlassPopulationEntry _klass_population[]`, + `MAX_KLASS_POPULATION_ENTRIES = 256`) keyed by klass `StringDictionary` id holds a + `KLASS_POPULATION_RING_SIZE = 30`-slot ring of per-epoch population counts plus one + `jweak` of a currently-live representative instance (minted fresh in + `foldKlassCountsLocked()`, deliberately *not* aliasing the source `TrackingEntry::ref`, + which `cleanup_table()` can reap out from under it). When full, the + least-recently-updated entry is evicted (`recordKlassPopulationSampleLocked()`), the + same fixed-capacity/LRU shape every other table in this design uses. Counts are + accumulated into a reused scratch array (`accumulateKlassCount()`) during the survivor + loop, then folded into the ring at the end of the pass. + - Trend/slope: computed by the shared `ringThirdsStats()` helper (`livenessTracker.cpp`) — + average (and minimum) of the earliest third of the filled window vs. the most recent third + (cheap, allocation-free, avoids full least-squares regression). Trend is trusted only once + the ring reaches a minimum fill (`KLASS_POPULATION_MIN_FILL_FOR_TREND = 10` samples) to + avoid noise right after a klass starts being tracked. A klass only qualifies as a leak + candidate once its growth (`hasQualifyingGrowth()`, which also checks a floor-rise bar to + reject oscillations whose peak passes a magnitude test but whose baseline never rises) has + held for `consecutive_positive` consecutive epochs at or above a hysteresis threshold — + `LEAK_TREND_HYSTERESIS_BASE = 5`, lowered to `LEAK_TREND_HYSTERESIS_CORROBORATED = 3` when + the aggregate post-GC heap floor is itself rising (`heapFloorRising()`, fed by a + lock-free, single-writer `_heap_floor_ring`). This closes the false-positive gap a + single-epoch positive-slope test had: a see-saw/oscillating population could trigger a + search without any real longer-term growth. The exact window size (up to 30) and the + "top 3–5" cutoff are starting points, not measured values — a separate tuning pass, not + folded into Open Question 2's pause-time work (this is a leak-detection sensitivity + tradeoff, not a safepoint-cost tradeoff). See + [reference-chains-collection-summary.md](../reference-chains-collection-summary.md) for + the full mechanism. + - `LivenessTracker::selectLeakCandidates(KlassCandidate *out, int max)` returns, on + demand under a shared lock, the positive-slope klasses ranked by magnitude descending, + capped at `min(max, MAX_LEAK_CANDIDATES = 5)` — this top-N cutoff *is* the per-pass + seeding cap, so no separate budget constant is needed. Each `KlassCandidate` carries the + klass id and its representative `jweak`. + - Known limitation, stated rather than solved: if population trends positive across + *many* klasses simultaneously, that's more likely heap-wide growth (warm-up, load + increase) than several independent leaks. The top-3–5 ranking limits how many candidates + get chased, but does not distinguish this case from true multi-leak; a future refinement. + - The leak *judgment* is retrospective (needs a full window of GC-epoch history to see the + trend) but the *reconstruction target* does not need to be the exact instance that built + up the trend — any currently-live tracked instance of the flagged klass is evidence of + the same leak. **The bridging step is a READ, not a `SetTag` write** (a correction to + this doc's original proposal, found while grounding + `LiveHeapReferenceChains-RemainingWorkPlan.md` (kept locally, not committed); + see its "Correction to the design doc's Open Question 3 mechanism"). Pre-`SetTag`ing a + candidate before the forward walk reached it would make `heapReferenceCallback()`'s + `*tag_ptr == 0` branch — the *only* branch that records `parent_tag`/`depth` — skip it, + yielding an empty/root chain. Instead `ReferenceChainTracker::pollWatchedTargets()` + resolves each candidate's `jweak`, and if `runPass()`'s whole-graph walk has already + tagged it (`getTag() > 0`, a pure read), calls `buildChainEvent(tag, ...)` and emits the + chain. A candidate still at tag 0 is retried on a later poll, since the whole-graph walk + eventually visits every root-reachable object (barring the hop/budget/frontier caps). No + new backward-walk primitive is needed; this reuses what the frontier table already + records. + - This couples two independently-flagged, independently-scheduled subsystems + (`referencechains=...` vs. `_record_liveness`/`_gc_generations`) that had no existing + relationship — `LivenessTracker` identifies objects via `jweak` and never calls JVMTI + `SetTag`/`GetTag`, while `ReferenceChainTracker` identifies objects purely via JVMTI tags + it assigns during its own traversal. `pollWatchedTargets()` bridging them is the one new + piece of machinery this adds; everything else reuses existing structures. + - **Resolved:** `referencechains=...` gets this target-seeding behavior only when liveness + tracking *and* `_gc_generations` are both enabled (`LivenessTracker::gcGenerationsEnabled()`, + checked in `pollWatchedTargets()`); otherwise it falls back to the whole-graph-only + behavior (no target seeding), which is the doc's originally-stated fallback. +4. ~~Decide whether the frontier should use JVMTI object tags at all, or adopt + `LivenessTracker`'s weak-ref + index-table pattern instead.~~ **Resolved** (see + "Frontier metadata storage"): tags stay, because `FollowReferences`/ + `IterateThroughHeap` need tagged objects to filter/report frontier membership and a + weak-ref table has no hook into that machinery — but the per-tag *metadata* + (`parent_tag`, `referrer_klass`, `depth`) should be stored in a + `LivenessTracker`-style slot table (tag value as index) rather than a bespoke + structure, reusing its proven `SpinLock`/CAS-allocation/resize code. Remaining open + item: the table-sizing formula, folded into Open Question 2. +5. Decide the actual pass-scheduling policy now that GC callbacks can only be a signal, + not a vehicle: e.g. a background agent thread woken by the GC-callback signal that + then calls `FollowReferences`/`IterateThroughHeap` for the next pass (paying its own + transparent safepoint), vs. a fixed-cadence timer independent of GC activity. Needs a + cost model for how many such safepoints per second are acceptable before this stops + being "more palatable" than one larger pause. + **A decision shipped, but not the cost-modeled one this question asks for.** + `ReferenceChainTracker::shouldRunPass()` (referenceChains.cpp) combines both candidates + rather than choosing between them: it triggers a pass when the GC-finish epoch has + advanced since the last pass, *or* a fixed `PASS_CADENCE_NS` (1 second, explicitly + labeled provisional in `referenceChains.h`) has elapsed, whichever comes first. No + safepoints-per-second/per-pause-duration measurement backs the 1-second cadence value - + it was chosen only so an idle search still makes progress without polling tightly. The + cost model this question actually asks for is still open, deferred to Phase 5's + benchmark plan (`LiveHeapReferenceChains-BenchmarkPlan.md` (kept locally, not committed)), + which has not been run. + + **SHIPPED — folded into Open Question 2's pause-time-SLO feedback loop, not solved + separately.** Implemented in the same `ReferenceChainTracker::updatePacing()` + (`referenceChains.cpp`, Phase D) described under Open Question 2: one `PidController` + `compute()` call per pass drives both that question's budget adjustment and this question's + cadence adjustment from the single measured per-pass safepoint duration, rather than two + independently-tuned mechanisms. `shouldRunPass()` and `threadLoop()` now compare against + `_effective_cadence_ns` in place of the fixed `PASS_CADENCE_NS` constant (which survives + only as `_effective_cadence_ns`'s starting value in `start()` and as the unit + `MAX_EFFECTIVE_CADENCE_NS` scales from); `threadLoop()`'s own sleep between iterations uses + `_effective_cadence_ns` too; so a controller-driven relaxed cadence actually shortens how + long an idle, no-GC-event search waits between passes, not just what the comparison in + `shouldRunPass()` reads. The GC-finish-epoch trigger in `shouldRunPass()` remains unconditional + on cadence, exactly as before — cadence only governs the fixed-interval fallback for an idle + search, per this question's original framing. See Open Question 2 above for the concrete + clamp/overflow mechanics (`MIN_EFFECTIVE_CADENCE_NS = 10ms`, `MAX_EFFECTIVE_CADENCE_NS = 4s`, + `CADENCE_NS_PER_EDGE_OVERFLOW`) and the gain-tuning caveat, which applies identically here + since it is the same controller instance. The cost-modeled "how many safepoints/sec is + acceptable" question this Open Question originally asked for is answered structurally (the + controller widens cadence exactly when passes are running long relative to the configured + ceiling) rather than by a specific measured number — that number is still a + `LiveHeapReferenceChains-BenchmarkPlan.md` (kept locally, not committed) item. diff --git a/doc/architecture/ReferenceChains-SignalsExplained.md b/doc/architecture/ReferenceChains-SignalsExplained.md new file mode 100644 index 0000000000..d45657e469 --- /dev/null +++ b/doc/architecture/ReferenceChains-SignalsExplained.md @@ -0,0 +1,366 @@ +# How reference-chain hunting decides when to run and when to back off + +*A guided tour of the signals in `ReferenceChainTracker` (`referenceChains.cpp`/`.h`), for +readers with no prior context on this subsystem.* + +## 1. What problem this is solving + +Java heaps leak. When they do, the useful question isn't "how big is the heap" — it's +*"what is holding onto these objects and refusing to let go?"* Answering that means +walking live references backwards from a suspect object to a GC root, i.e. reconstructing +a **reference chain**. + +The obstacle: walking the heap graph (JVMTI's `FollowReferences`/`IterateThroughHeap`) +requires the JVM to stop every thread at a safepoint — a Stop-The-World (STW) pause, the +same kind a GC pause is. A profiler that stops the world to investigate a leak is +trading one problem for another. So the whole design of this subsystem is really an +answer to one question: + +> **How do we get a useful heap walk without stopping the world for longer, or more +> often, than the leak justifies?** + +Everything below is the machinery that answers that question — split into two halves: +*when should the next slice of walking happen* (triggering), and *how big/frequent +should that slice be allowed to get* (pacing/pausing). + +## 2. The core trick: one long pause becomes many short ones + +A single "walk the whole reachable heap and find the chain" pass could take seconds on a +large heap — an unacceptable pause. Instead, this subsystem does **bounded, resumable +BFS**: each *pass* only visits a limited number of objects (a *budget*), remembers where +it left off using JVMTI object tags, and picks the walk back up on the next pass. So a +"search" for a leak's chain is really a sequence of many short passes, each its own small +safepoint, spread out over time. + +This reframes the whole design problem from "avoid the pause" (impossible — see §3) to +"decide, pass by pass, whether *now* is a good time to spend one of these small pauses, +and how big it should be." + +## 3. Why GC callbacks can only ever be a *signal*, never the *work* + +The natural instinct: "GC just ran, the heap just changed — hook the GC callback and do +the walk right there, for free, since the VM is already stopped." + +This doesn't work, and the reason is worth understanding because it shapes the rest of +the design. JVMTI's spec is explicit: inside `GarbageCollectionStart`/ +`GarbageCollectionFinish`, an agent may call only the **Memory Management** category +(`Allocate`/`Deallocate`). Everything a heap walk needs — `SetTag`, `GetTag`, +`GetObjectsWithTags`, `FollowReferences`, `IterateThroughHeap` — is in the **Heap** +category, which is explicitly *not* on that allow-list. Calling any of them from inside +the GC callback is exactly what the restriction forbids (confirmed against +`GCCallbackGuard` in `referenceChains.cpp`, which asserts this in debug builds). + +So the GC callback can do exactly one cheap, legal thing: bump an atomic counter. + +```cpp +void ReferenceChainTracker::onGCFinish() { + GCCallbackGuard guard; // "we are inside the forbidden window" + atomicIncRelaxed(_gc_finish_epoch, (u64)1); +} +``` + +That's it. No walk, no tag calls, nothing heap-related — just "a GC finished, epoch N". +This is the first, and most important, idea to internalize: + +> **A GC callback is a doorbell, not a worker.** It tells a separate thread "something +> happened, go check if it's worth acting on" — it never does the acting itself. + +The actual walk happens later, on a dedicated background thread, deliberately outside +the safepoint the GC callback fired inside. That thread calling `FollowReferences` +triggers its *own*, independent safepoint — the GC's pause and the walk's pause are two +separate STW events, not one shared one. (An earlier version of this design hoped to +"ride" the GC's own pause for free; that turned out to be architecturally impossible +without reaching into unversioned HotSpot internals — see +`doc/architecture/LiveHeapReferenceChains.md`'s Triggering section for the full +investigation.) + +## 4. The scheduling loop: cadence + epoch, not "run on every GC" + +The background thread (`threadLoop()`) wakes roughly once a second and asks +`shouldRunPass()`: "should I spend a pass right now?" Two independent signals feed that +decision: + +1. **The GC-finish epoch changed** since the last pass. A GC just happened — the heap + graph likely moved, so a fresh pass is probably worth its cost. +2. **A fixed cadence has elapsed** since the last pass, even with no new GC. Two distinct + roles, easy to conflate: + - For a **workload with no/rare GCs** it is the fallback the naive reader expects — + without it, the search would stall forever waiting for a signal that never comes. + - For an **in-progress search** it is the *crawl's pacing knob*, not a signal re-check: + passes between GCs are the only thing that drains the search's own frontier backlog + (its "found but not yet expanded" objects). GC time does not advance the crawl — + pass time does — so a RUNNING search legitimately runs cadence-driven passes with + zero new signal, and the pause-time controller (§7) widens/narrows exactly this + cadence. The only genuinely stale re-check is the terminal state's cheap + restart-gate re-evaluation (§5), two atomic loads, deliberately kept unconditional. + +```cpp +u64 gc_finish_epoch = gcFinishEpoch(); +if (gc_finish_epoch != _last_pass_gc_finish_epoch) { + return true; // "a GC just happened, a pass may be worth running soon" +} +... +return cadence_elapsed; +``` + +Notice what's deliberately *not* here: the loop does **not** wake up early on every GC. +`onGCFinish()` only bumps a counter — it never calls `pthread_kill()` to interrupt the +sleeping thread (the only early-wake signal is shutdown's abort, §10). Why swallow the +up-to-~1s latency instead of reacting instantly? Because under a GC-heavy workload, +waking on *every* GC would collapse the loop's cadence down to GC frequency — each wake +is a full iteration of scheduling logic, not free. The design accepts "at most ~1s of +extra latency" in exchange for not turning a GC storm into a scheduling storm. This is +a recurring theme in this subsystem: **every signal is deliberately made cheap to check +and expensive to over-react to.** + +One caveat the epoch trigger carries: it is an *unconditional* bypass — while it is +set, no cadence check applies. Minor young GCs bump the epoch as readily as majors, so +once the adaptive cadence (§7) has shrunk to its floor, a workload that GCs more often +than the wake interval effectively drives one pass per wake, at up to GC frequency. +That is bounded and acceptable for the crawl lanes (each pass is pause-budgeted by the +PID controller); the one lane that needs a rate bound of its own — the canary chase — +gets one explicitly (§8). + +## 5. Not every wake actually walks: the leak-signal gate + +Waking up and checking cheap counters is fine to do often. Actually walking the heap +costs real STW time, so before a **brand-new search** (or a **restart** of one that +finished) is allowed to begin, a second, independent gate applies: +`LivenessTracker`'s population-trend signal — "is there currently a class whose live +object count looks like it's growing without bound?" (`hasLeakSignal()`). + +```cpp +bool ReferenceChainTracker::hasLeakSignal() { + if (!LivenessTracker::instance()->gcGenerationsEnabled()) { + return true; // no trend signal available at all -> don't gate on it + } + ... + int n = LivenessTracker::instance()->selectLeakCandidates(probe, 1); + return n > 0; +} +``` + +This matters because a full-heap BFS is expensive relative to a targeted one, and there's +no point paying that cost speculatively, with nothing to justify it. If the leak-tracking +feature isn't even enabled, this reduces to "always true" — the walk runs unconditionally, +exactly as a simpler standalone version of this feature would. + +Once a *search is already running*, though, this leak-signal gate is deliberately **not** +re-applied per pass — an in-progress search's own frontier (its list of "found but not yet +expanded" objects) is allowed to keep converging pass after pass regardless of whether a +fresh leak candidate happens to be visible right now. Gating an already-running search on +"is there a candidate this instant" would stall its progress for no good reason; the gate's +job is to decide whether *starting* new expensive work is worth it, not to second-guess +work already committed to. + +The same "a potential leak is currently detected" signal also raises *liveness tracking +fidelity* itself: once candidate selection has fired, the qualifying tids' allocations are +admitted to the tracking table at 100% instead of the configured live-samples ratio +(default 10% — a 90% probabilistic drop that thins small per-(klass, tid) populations; +`LivenessTracker::admitForTracking()`). The raise is bounded by the candidate threads' +own allocation rate, not the process's whole allocation rate, so the tracking table's +cost scales with the leak's own threads; and it is refreshed — and cleared — poll by poll +with the candidate selection, so it never outlives the chase. The OOM urgency ramp +(§9) raises admission to 100% for *all* allocations for the same reason it drops every +pacing rule: the process is expected to die soon, and maximizing what the last chapter +captures outweighs the tracking table's transient volume. + +## 6. Paying for it: the "pain budget" leaky bucket + +Even with a real leak signal, restarting a full search back-to-back forever would be +its own kind of runaway cost. The safety valve here is a **leaky bucket over cost**, not +over time — `PainBudget` (`painBudget.h`): + +```cpp +// spend(): "that last search cost N milliseconds of wall-clock work" +// canStartNow(): "has that debt drained back to ~0 yet, at refill_rate?" +``` + +The intuition: if a search finished *cheaply*, it can restart again almost immediately. +If it was *expensive*, the next restart has to wait proportionally longer. This is a much +better model than a fixed cooldown timer, because "how expensive was the last search" +is exactly the thing worth reacting to — a fixed cooldown would either be too +conservative after a cheap search or too permissive after an expensive one. + +`canAffordNewSearch()` combines both gates from §5–6: the leak signal has to say "worth +it" *and* the pain budget has to say "affordable" before a new/restarted search is +allowed to begin. + +## 7. Self-tuning the pause itself: a PID controller on pause time + +So far: *when* to run a pass. Now: *how expensive should that pass be allowed to get?* + +Each pass has a target STW duration (`_pause_target_ms` — an operator-configured ceiling, +e.g. "no single pass should take more than N ms"). After every pass, the tracker measures +how long it actually took and feeds that into a small PID controller +(`_pause_pid.compute()`), which nudges the *per-pass budget* — how many objects the next +pass is allowed to visit — up or down: + +- Pass ran comfortably under target → budget can grow a bit (there's headroom). +- Pass ran over target → budget shrinks, so the *next* pass is smaller and faster. + +This is the same self-correcting idea used elsewhere in this profiler for sampling rates +(`ObjectSampler`, `MallocTracer`) — measure the actual cost, compare to a target, adjust +the knob that controls the next iteration's cost, repeat. It means the operator doesn't +have to hand-tune a "safe" fixed budget for every heap size and object graph shape; the +controller finds it empirically, pass by pass. + +There's a second knob the same controller feeds: if the budget is already at its floor +and *still* over target, that's a sign the real problem isn't "how much work per pass" — +it's "passes are happening too close together." In that case the controller widens the +*cadence* (the sleep interval between passes) instead. Conversely, if a pass finishes +comfortably under target even at the configured budget ceiling, the idle cadence is +shortened — there's slack to use it to converge faster. Either way, the same measured +signal (pass duration vs. target) drives both "how much work per pass" and "how often to +even try." + +There's also a small "savings account" on top of this (budget-borrowing): a *sustained* +run of comfortably-under-target passes slowly raises the ceiling itself, not just the +budget inside it — but a single pass that isn't comfortably under target revokes that +extra headroom immediately. The asymmetry is deliberate: earning slack should take +sustained good behavior; losing it should be instant, so a run of easy passes can never +turn into an excuse for one expensive one. + +## 8. When "wait for the next cadence tick" isn't good enough: canary mode + +The scheduling described in §4–7 assumes a slow, whole-heap background search. But +sometimes there's a much sharper signal available: `LivenessTracker` has already flagged +*specific* suspect objects (a small, pre-tagged "canary" set) worth confirming quickly. In +that mode: + +```cpp +bool canary_active = _candidate_count > 0 && + __builtin_popcountll(_candidate_found_bits) < (u64)_candidate_count; +if (canary_active) { + return true; // run the next pass immediately, no cadence wait +} +``` + +While a canary search is active, cadence is bypassed — passes run back-to-back *while the +chase is fresh or making candidate progress*, because the PID controller (§7) is already +keeping each individual pass's pause small, and a chase that is genuinely close to its +target should resolve in a handful of passes. + +Back-to-back-forever, however, is exactly the wrong promise for the failure mode this +feature actually meets in production: a candidate the crawl cannot reach soon (a leak +holder buried behind a deep frontier backlog — the coverage lottery inherent to an +external JVMTI agent with no reverse-edge primitive). Measured live on a production-like +analyzer pod: one such un-findable candidate held the chase open for 32 minutes at ~88 +passes/min — a full core of engine work — because the bypass applied unconditionally and +the loop skips its sleep whenever a pass will run. "Well under 20ms per 60s recording" +is only ever true for *findable* candidates. + +So the canary lane carries its own rate bound — a **progress-driven, work-scaled +exponential backoff**. The inter-pass spacing is a *multiple of the measured cost of a +pass itself* (an EMA of each pass's wall duration), not a fixed wall-clock constant: +- Every pass that makes candidate progress (a candidate found, or a new candidate + admitted) resets the spacing multiplier to 1 — back-to-back. At multiplier 1 the + next pass starts as soon as the last one ended, which is harmless by construction + for a cheap pass, and exactly the fast-resolution burst a chase that is genuinely + close should get. +- Every pass with *no* candidate progress doubles the multiplier, capped at 16. A + permanently stuck chase therefore settles at one pass per 16 × (its own cost) — + it keeps ticking indefinitely (abandonment is a separate, frontier-aware detector's + job, below), but its steady burn is structurally bounded to ~1/16 of a core on pass + work, *whatever that work is*. +- The work-scaling is the point: a fixed cap only binds when it exceeds the pass's own + duration — the production pod's passes ran 0.7–4s, so a 1s cap would have changed + nothing at all (the loop is work-bound, never sleep-bound, when the pass exceeds the + cap), while the same 1s cap starved a genuinely reachable deep chase whose passes + cost milliseconds. Scaling to the measured cost gives the expensive-pod chase a real + bound and the cheap-but-deep chase its density from the same law. +- The GC-epoch trigger deliberately does **not** bypass the backoff: a GC-heavy workload + bumps the epoch on virtually every wake, so letting GCs override the spacing would make + the backoff unreachable on exactly the deployments that burn the most. +- The OOM urgency ramp (§9) overrides it entirely — imminent OOM remains the one regime + where the chase burns budget back-to-back. + +The pain-budget refill rate is raised 100x while a chase is open, but note what that +means now: **not** a rate control (the backoff is the rate bound) — just a double-throttle +guard, so the conservative base refill rate tuned for the ordinary ~1 pass/s crawl doesn't +starve a chase the backoff has already paced. The earlier covering-vs-emergency refill +distinction existed to feed the unbounded back-to-back mode and is gone with it. + +## 9. The panic button: ramping up as OOM approaches + +All of the above optimizes for "acceptable background cost most of the time." But if +`LivenessTracker` projects the heap is genuinely on a collision course with +`OutOfMemoryError` within the next ~30 minutes (`secondsToOOM()`), "acceptable background +cost" is the wrong objective — the process might not survive long enough for a leisurely +search to finish. So there's a third mode: an **urgency ramp**. + +As projected time-to-OOM shrinks from 30 minutes toward zero, the pause-time target and +the pass cadence both ramp *exponentially* toward much more aggressive ceilings: + +```cpp +double x = 1.0 - seconds_to_oom / OOM_RAMP_START_S; // 0 at 30min out, 1 at OOM +target_ms = pause_target * pow(URGENT_PAUSE_TARGET_MS / pause_target, x); +cadence_ns = pow(URGENT_CADENCE_NS / PASS_CADENCE_NS, x) * PASS_CADENCE_NS; +``` + +The reasoning behind exponential (rather than linear) ramping: at 30 minutes out, the +situation still might resolve itself (a GC frees the suspect objects, the trend reverses) +— stay cheap. In the last seconds before OOM, the process is likely to die anyway, so it's +worth spending far more of the pause-time budget to collect a usable chain *before that +happens* than to protect a latency budget for a process that may not be there to benefit +from it. Held flat-out cheap the whole time, this urgency signal would arrive too late to +matter; held aggressive the whole time, it would waste budget on every one of the many +false alarms a rising trend that later reverses produces. + +Two more details make this practical rather than flappy: + +- **Hysteresis, not a bare threshold.** `isUrgent()` *latches* on when + time-to-OOM drops below a threshold, and only *releases* after several consecutive + observations comfortably clear of a separate (higher) release bar. A single noisy + reading crossing back and forth across one threshold would otherwise thrash the ramp + on and off every second. +- **One search per urgency episode.** Once an urgent episode has spent its one + authorized search, further ticks within the same episode don't keep tearing down and + restarting it from scratch — the per-candidate probe (§5) remains the only trigger + until the episode actually clears. + +## 10. Stopping mid-pass: the abort path + +Everything above is about *starting* passes thoughtfully. There's also a clean way to +*stop* one that's already in flight — needed when the profiler itself is shutting down +(or a test needs to reset state) while a `FollowReferences` call is still blocked inside +the JVM. + +```cpp +_abort_pass_requested.store(true, std::memory_order_relaxed); +pthread_kill(_thread, WAKEUP_SIGNAL); // interrupts a sleeping thread promptly +``` + +The callback JVMTI invokes for each visited object checks this flag and returns an abort +code the moment it sees it set — since nothing outside the JVM can interrupt a call +already inside `FollowReferences`, the flag has to be checked *from inside* the callback +JVMTI itself is driving. `pthread_kill` with a no-op-handler signal only helps the *other* +common case — a thread parked in `OS::sleep()` between passes — wake up promptly instead +of waiting out the rest of its interval. + +## 11. Putting it together + +| Question | Signal | Where | +|---|---|---| +| Did the heap graph just change? | GC-finish epoch bump | `onGCFinish()` → `shouldRunPass()` | +| No GC signal — is it time anyway? | Fixed/adaptive cadence elapsed | `shouldRunPass()` | +| Is a *new* search worth starting at all? | LivenessTracker population trend | `hasLeakSignal()` | +| Can we afford to spend that cost right now? | Pain-budget leaky bucket | `canAffordNewSearch()` | +| How big/frequent should passes be, steady-state? | PID controller on measured pause time | `updatePacing()` | +| Are we chasing specific known suspects? | Canary candidate set + progress-driven backoff | `canary_active` bypass + work-scaled `_canary_backoff_mult` | +| Is OOM close enough to abandon caution? | secondsToOOM() latch/release | urgency ramp in `threadLoop()` | +| Need to stop a pass already in flight? | Abort flag + wakeup signal | `_abort_pass_requested` | + +The unifying idea across all eight mechanisms: **every trigger is a cheap check, and +every response is proportional to real, measured cost** — never a fixed guess. GC +callbacks stay legal by doing nothing but incrementing a counter. Whether to search at +all is gated on an independent leak-trend signal, not "because a GC happened." Whether a +search can *restart* is gated on how expensive it actually was last time, not a flat +cooldown. How big a pass gets is tuned from its own measured pause time, not a static +config value. And the one scenario where none of that caution applies — imminent OOM — is +its own explicitly separate, hysteretic escalation path, not a tweak to the steady-state +knobs. + +That's what makes several short, adaptively-sized pauses a genuinely better trade than +one long one: the *decision* of when to pay each of those small costs is never blind — +it's always backed by a signal that says this particular pause is likely to be worth it. diff --git a/doc/reference-chains-collection-summary.md b/doc/reference-chains-collection-summary.md new file mode 100644 index 0000000000..0354a8cbe7 --- /dev/null +++ b/doc/reference-chains-collection-summary.md @@ -0,0 +1,87 @@ +# Reference Chain Collection: Design Summary + +## Problem + +Given a JVM heap with objects suspected of leaking (e.g., klasses whose live population grows monotonically across GC generations), reconstruct a **referrer chain** from a GC root down to a representative instance of the suspect klass — without pausing the JVM for longer than a small, bounded budget, and without assuming the entire heap graph can be walked in one pass. + +Three constraints drive the design: + +1. **Detecting *which* klasses are worth walking** must be near-free and based on survivorship trend, not raw allocation volume. +2. **The walk itself** (JVMTI `FollowReferences`) can be arbitrarily expensive on a large heap, so it must be interruptible and resumable. +3. **Total STW/JVMTI-callback time per pass** must stay under a small budget so the profiler doesn't visibly perturb the target application. + +--- + +## Component 1: Surviving-Generation Signal (`LivenessTracker`) + +Rather than triggering a heap walk on every allocation or every GC, the tracker maintains a **per-klass population history** and only nominates a klass as a "leak candidate" once it shows a **sustained positive trend across GC generations** — i.e., its live (surviving) instance count keeps growing generation over generation, not just spiking transiently. + +**Mechanics:** + +- Population sampling is driven off the existing allocation-sampling hot path (`track()`), but the actual **per-klass counts are only folded into history at `cleanup_table()`'s GC-epoch-advance pass** — i.e., once per GC, not once per allocation. This keeps the hot path allocation-free and cheap. +- Each klass gets a small **ring buffer of recent per-epoch surviving counts** (`KLASS_POPULATION_RING_SIZE = 30` samples). A ring, not an unbounded history, because we only care about recent trend, not lifetime totals. +- A klass's trend is only trusted once its ring has a **minimum fill (`KLASS_POPULATION_MIN_FILL_FOR_TREND = 10` samples)** — avoids false-positive trend detection on a klass that's simply new to being tracked (too few points to fit a slope to). +- `selectLeakCandidates()` computes a slope over each ring and returns the **top-N klasses by slope magnitude** (`MAX_LEAK_CANDIDATES = 5`), each paired with a live representative instance (a `jweak`) discovered during sampling — this weak reference is what seeds the walk in Component 2. +- A klass only qualifies once its growth (`hasQualifyingGrowth()`) has held for `consecutive_positive` epochs at or above a **hysteresis threshold** — `LEAK_TREND_HYSTERESIS_BASE = 5` by default, lowered to `LEAK_TREND_HYSTERESIS_CORROBORATED = 3` when the aggregate post-GC heap floor is itself rising (`heapFloorRising()`, fed by a lock-free, single-writer `_heap_floor_ring` populated from `onGC()`). Because the aggregate heap-floor signal can't attribute growth to any one klass, it only ever raises or lowers the bar uniformly for the whole scan — it never reorders or singles out individual candidates. +- The whole table (`_klass_population`, up to `MAX_KLASS_POPULATION_ENTRIES = 256` entries) is a flat array scanned linearly — deliberately no index structure, since 256 entries is cheap to scan and this stays off the allocation hot path. +- Everything under this table (population array, size counter) is guarded by a single `SpinLock` (`_table_lock`) — the *same* lock `cleanup_table()` already holds for its epoch-advance pass, rather than adding a second lock. **Any code path that mutates this table (including test-only reset seams) must take that lock — mutating `_klass_population_size` or the array unguarded is a data race against the epoch-advance pass**, discovered in practice while hardening test seams. + +**Why this design:** it decouples "is this klass suspicious" (cheap, GC-cadence, statistical) from "reconstruct why it's suspicious" (expensive, JVMTI, on-demand) — the expensive walk only ever runs against klasses that have already earned a positive trend signal, not against every allocation site. + +--- + +## Component 2: Resumable Frontier Walk (`ReferenceChainTracker`) + +Once a klass is nominated, a **persistent background BFS thread** reconstructs a path from a GC root to a tagged instance of that klass, using JVMTI's `FollowReferences`/heap-tag mechanism — but broken into many small, budgeted passes rather than one unbounded walk. + +**Mechanics:** + +- The tracker is a **process-wide singleton** with its own thread (`threadLoop()`), woken on a fixed cadence (`effectiveCadenceNs`) rather than synchronously from allocation or GC callbacks — decouples walk progress from the rate of GC/allocation events. +- `runPass()` dispatches on `_search_started`: + - **First pass for a search**: enumerates heap roots via `IterateOverReachableObjects()` (`heapRootCallback()`/`stackRefCallback()`), tagging root-referenced objects as it goes. + - **Every subsequent pass**: calls `expandFrontier()`, which resumes from a **persisted frontier** (the previous pass's boundary tags) instead of re-walking from roots. This is the resumability mechanism: each pass advances the frontier outward by one bounded increment and stops. +- Each pass is capped by an **edge-admission budget** (`effectiveBudget`, e.g. `edges_admitted` capped at a configured value like 4000/200000/500 depending on test config) — `expandFrontier()`'s nested loops (`while (!ctx.truncated && progress)` outer, `for (jlong tag : candidate_tags)` inner) both check a truncation flag and bail out the moment the budget is exhausted, so a single pass's JVMTI-callback time is bounded regardless of heap size. +- **Cooperative abort**: an `std::atomic _abort_pass_requested` flag, checked inside `heapReferenceCallback()` (the JVMTI callback invoked per edge), lets `stopThread()` interrupt an **in-flight** walk promptly — set before `pthread_kill(WAKEUP_SIGNAL)`/`pthread_join()`, cleared by `startThread()`. Without this, a `FollowReferences` call already in progress at JVM shutdown or profiler restart can't be interrupted, and `pthread_join()` blocks indefinitely (a real, previously-diagnosed shutdown hang). +- Search state is a small state machine: `RUNNING → {ABANDONED | COMPLETED}`. `RUNNING` can **restart itself** (fresh root walk) once a candidate's chain is found and its tags released, gated by `canAffordNewSearch()`'s **pacing budget** — self-throttling, not unconditional: a search won't restart back-to-back if it would blow the perturbation budget. A terminal state (`ABANDONED`/`COMPLETED`) is not final: `shouldRunPass()` restarts the search from it via `restartSearch()` once the pain budget has drained and a leak indication is (still) present. +- `runPass()` only moves to `COMPLETED` once the frontier is fully drained **and** `_watched_leak_klass_count == 0` (no klass currently under active leak watch, Component 4). A fully-drained frontier while a klass is still watched leaves `_search_state` at `RUNNING` instead: the walk has visited every reachable object once, but a leak-shaped klass keeps growing by **mutating an already-visited container** (e.g. appending to a `static final` collection field long after the walk first admitted it), which a one-time visit can never observe again. Rotation (Component 4) is what re-observes those already-`EXPANDED` entries on later passes. +- A search is marked `ABANDONED` (with a reason code) if it runs out of frontier budget without completing — e.g. hitting a frontier-cap under a tiny configured budget. This is a deliberate, observable outcome, not a silent failure — surfaced so operators can distinguish "the walk gave up" from "the walk is still in progress." + +**Why this design:** treating the walk as a resumable state machine (persisted frontier + tags) rather than one atomic call means a heap graph of unbounded size never forces an unbounded pause — cost is amortized across many cheap passes, each individually bounded and individually abortable. + +--- + +## Component 3: Latency Budget Enforcement + +The system enforces its "don't perturb the app" guarantee at **three independent layers**, not just one: + +1. **Per-pass edge budget** (`effectiveBudget`) — caps JVMTI callback invocations per pass (Component 2). +2. **Pain budget** (`_pain_budget`/`_search_pain_ms`, spent via `_pain_budget.spend(...)`) — tracks cumulative walk cost against a wall-clock ceiling; used by `canAffordNewSearch()` to decide whether a new search/restart is affordable right now, not just whether the current pass fit its edge budget. This is what prevents "many cheap passes" from silently adding up to an expensive aggregate cost. +3. **Pass cadence** (`effectiveCadenceNs`) — the background thread only wakes and attempts a pass on a fixed cadence (plus GC-epoch-triggered wakeups), rather than continuously spinning, bounding CPU overhead between passes. + +Together these mean: a single pass is bounded (edge budget), a sequence of passes is bounded (pain budget), and idle overhead between passes is bounded (cadence) — the three layers target three different ways an unbounded-cost walk could otherwise leak into the target application's latency. + +--- + +## Component 4: Rediscovering Growth in an Already-Visited Container + +Once a klass is leak-flagged, `pollWatchedTargets()` refreshes `_watched_leak_klass_ids` (up to `MAX_WATCHED_LEAK_KLASSES = 5`, matching `LivenessTracker::MAX_LEAK_CANDIDATES`) from `LivenessTracker::topKlassesByGenerationCount()` — a faster, un-hysteresis-gated ranking than the `selectLeakCandidates()` canary set, but only consulted once `hasLeakSignal()` has already fired via that slower path. + +**Why a matching mechanism is needed at all:** the real leak shape this targets is a `static final` collection field that gets *appended to*, not reassigned — the container itself was already admitted and `EXPANDED` in an early pass, long before `selectLeakCandidates()`'s hysteresis authorized watching its element klass. New elements can only be rediscovered by re-expanding that already-visited container, not by discovering a brand-new root. + +**Mechanics:** + +- Matching a newly-admitted object against `_watched_leak_klass_ids` must use a class identity that stays valid for the object's whole lifetime, not `referrer_klass` — a classMap `StringDictionary` id that can differ for the same class at different times if that dictionary is compacted/regenerated. Both `ReferenceChainTracker` and `LivenessTracker` mint from a single shared, process-wide `ClassTagAllocator` (`classTagAllocator.h`) and store the resulting stable `class_tag` (`FrontierEntry::class_tag`, `KlassPopulationEntry::stable_class_tag`) instead. +- `trackLeakAccumulation()` runs on every successful admission (`admitObject()`'s `ADMITTED` result) and aggregates, per `(leaf_klass_id, parent_class_id)` signature, how many admitted children of a watched leaf klass were observed under a parent of that class (`_leak_signature_totals`, ranked by delta against the previous pass's snapshot — Tier 1), and per parent *object* tag, how many such children that specific parent holds (`_leak_parent_fanout` — Tier 2, ranked within the winning Tier-1 signature). +- `seedLeakAccumulationForNewlyWatchedKlass()` runs once, the moment a klass_id first enters `_watched_leak_klass_ids`: it scans the whole frontier table for already-`EXPANDED` entries whose `class_tag` matches, since a container that was fully admitted before its element klass started being watched would otherwise never get its first Tier-1/Tier-2 data point. +- `collectLeakAccumulationCandidatesForRotation()` re-queues the Tier-2 winner(s) for re-expansion (`LEAK_ACCUMULATION_ROTATION_BUDGET = 16` per pass) — this is what actually re-visits the growing container and picks up elements appended since its first expansion. + +**Two additional robustness fixes surfaced only under a real growing-collection repro, not by the unit suite alone:** + +- **Urgent-signal latch.** `isUrgent()` used to be a bare `secondsToOOM() < OOM_URGENT_THRESHOLD_S` comparison; that projection is derived from a short ring of heap deltas and can swing by orders of magnitude between consecutive observations of the same steadily-growing heap. A bare comparison flapped, and each flap back to "urgent" bypassed the per-klass hysteresis gate in `hasLeakSignal()` and restarted the search — which discards the frontier table and the Tier-1/Tier-2 accumulators above, so they never got the several passes they need to converge. `isUrgent()` now latches on first crossing and only releases after `URGENT_RELEASE_CONSECUTIVE` (5) consecutive observations at or above `OOM_URGENT_RELEASE_S` (2× the threshold); `_urgent_search_spent` limits each latched episode to authorizing one restart. +- **Classmap-generation sync at startup.** `LivenessTracker::initialize()` now seeds `_last_class_map_generation` from the real classMap generation instead of leaving it at the default `0`. Previously, the first `cleanup_table()` call after any profiler start saw `current_generation != 0`, treated it as a classMap reset, and wiped `_klass_population` — discarding any population history folded in between `initialize()` and that first `cleanup_table()` call. + +--- + +## Output Path + +Once a candidate's chain is fully reconstructed, `pollWatchedTargets()` builds a chain event (`buildChainEvent()`), which is enqueued (`enqueueChainEvent()`) and later drained (`drainPendingChainEvents()`, called from `Profiler::dump()`, not from the BFS scheduling thread) into `Profiler::writeReferenceChain()` — ultimately surfaced as a `datadog.ReferenceChain` JFR event on the next `Profiler::dump()`. This keeps the expensive walk and the (comparatively cheap, already-existing) JFR-write path decoupled — the walk never blocks on JFR I/O, and JFR writes never trigger a walk. diff --git a/doc/reference-chains-design.md b/doc/reference-chains-design.md new file mode 100644 index 0000000000..059a181189 --- /dev/null +++ b/doc/reference-chains-design.md @@ -0,0 +1,259 @@ +# Reference Chains for Surviving Live Heap Samples + +**Status:** Implemented (see `doc/reference-chains-collection-summary.md` for the as-built design) +**Date:** 2026-07-07 +**Jira:** TBD + +## Goal + +For a subset of live-heap samples that survive past their allocation window, produce +a **reference chain** — a sequence of referrer *types* (not full field-level paths, not +necessarily to *all* GC roots) connecting the sampled instance back to *a* GC root. This +is diagnostic information ("what kind of object chain is keeping this alive"), not a +heap-dump-grade exact retainer analysis. + +## Constraints + +- Must run cheaply, with as short a safepoint / STW contribution as possible. +- Must work on stock vendor JDKs the agent attaches to — no forked/patched JVM builds. +- Exhaustive (all-roots, full-path) chains are explicitly **not** required; referrer-type-only, + bounded-depth, best-effort chains are acceptable. + +## Approaches considered + +Three approaches were evaluated; two are ruled out as launch requirements for concrete, +evidence-backed reasons. One sub-idea (Approach C's `ParallelObjectIterator` variant) is +explicitly kept open as a conditional future option; see its discussion below. + +| # | Approach | Completeness | Complexity | Feasibility | Status | +|---|---|---|---|---|---| +| A | Full JVMTI `FollowReferences` reverse-graph walk, piggybacked on an already-scheduled major GC | 4/5 | 4/5 | 2/5 | Rejected | +| B | Bounded BFS-from-roots with frontier pruning (JFR "leak profiler" technique, adapted) | 3/5 | 3/5\* | 4/5 | **Chosen** | +| C | Hook GC mark/copy closures (G1, ZGC) to record parent pointers inline during marking | 2/5 | 5/5 | 1/5 | Rejected | + +\* This 3/5 reflects only the single-pass BFS sketched at selection time. The "Chosen +design" section below replaces that sketch with an incremental, resumable BFS +(JVMTI-tag-based frontier persistence across GC cycles, a dedicated `VM_Operation` per +pass, GC-callback signaling, and explicit termination/tag-cleanup bookkeeping), which is +materially more complex than this score suggests — closer to 4/5 in implementation and +maintenance effort. The score is left unchanged above (it documents the state of the +comparison at decision time) rather than retroactively edited. + +### A — Full reverse-reachability walk (rejected) + +Safepoint length scales with live-set size regardless of how the walk is triggered. +Modern regionalized collectors (G1, Shenandoah) rarely perform a true full-heap walk +during ordinary major GCs, so "ride an already-paid pause" is not a reliable amortization +strategy. Cost is fundamentally at odds with the "short safepoint" constraint. + +### B — Bounded BFS-from-roots (chosen) + +Mirrors OpenJDK's own `jdk.OldObjectSample` leak-profiler implementation +(`src/hotspot/share/jfr/leakprofiler/chains/{edgeStore,bfsClosure,dfsClosure}.cpp`): +a `VM_Operation`-driven BFS from GC roots, retaining only edges on the frontier toward a +small, fixed sample set, with a hard hop cap (HotSpot itself caps chains at ~200 hops, +split 100/100 from leaf and from root). We can go cheaper than JFR because only the +**referrer class**, not object identity or field name, is needed — the `EdgeStore` +degenerates to `(referrer_klass, parent_ref, depth)` records, where `parent_ref` links +each record back to the record that discovered it, enabling chain reconstruction. + +Adopting this pattern is a re-scoping of proven, shipping HotSpot code, not a novel +algorithm design. + +### C — GC mark/copy closure piggyback (rejected) + +Investigated specifically for G1 and ZGC on the premise that per-edge referrer +information is already available inside the collector's own marking/evacuation closures +(`G1ParCopyClosure::do_oop_work`, ZGC's `ZMarkConcurrentRootsIteratorClosure` / +load-barrier closures), so recording it would cost nothing beyond what the GC already +pays. + +Rejected because there is no stable, externally reachable hook into these closures: + +- They are internal, template-instantiated C++ classes compiled into `libjvm.so` at + HotSpot build time — not a registrable/pluggable extension point. +- This differs categorically from `VMStructs`-style introspection already used in this + codebase (`ddprof-lib/src/main/cpp/hotspot/vmStructs.cpp`), which reads VM state + passively via an officially exported offset table. Intercepting a GC closure's + *behavior* would require either shipping a patched OpenJDK build (a fork/maintenance + commitment far beyond anything in this codebase) or binary-patching unversioned, + per-build-mangled function addresses — not shippable across JDK point releases. + +A related idea — using HotSpot's internal `ParallelObjectIterator` +(landed via [JDK-8322043](https://www.mail-archive.com/serviceability-dev@openjdk.org/msg12977.html), +used by `VM_HeapDumper` to partition heap regions across GC worker threads for parallel +heap dumping) to shrink Approach B's safepoint by parallelizing the walk — was also +investigated. Same verdict: it is an internal C++ class, not exposed via JVMTI, with no +stable ABI for an attached agent to call. Symbol-sniffing internal HotSpot functions *is* +an established pattern in this codebase (`VMStructs::findHeapUsageFunc`, +`vmStructs.cpp:489-509`), but that precedent covers a single leaf virtual method with a +value/POD-ish return; `ParallelObjectIterator` is a multi-class subsystem that coordinates +the VM's own GC worker threads under safepoint control — an order of magnitude larger +fragility surface, with a much higher blast radius if a layout assumption is wrong (GC +worker-thread coordination corruption vs. a bad JMX stat). Not pursued as a launch +requirement; revisit only if Approach B's single-threaded pause proves to be a measured +bottleneck, and treat it as an isolated, heavily version/flag-gated fast path with +automatic fallback — never a dependency. + +## Chosen design: incremental, resumable bounded BFS + +A single-pass bounded BFS still means one pause sized to whatever budget is configured. +The refinement below spreads that budget across multiple short passes instead of one +contiguous one, trading a possibly-higher *aggregate* STW total for a much better +*latency distribution* — no single long tail pause. + +### Why the frontier can survive across passes: JVMTI object tags + +The obstacle to pausing and resuming a BFS is that the frontier (the worklist of +not-yet-expanded objects) is normally a set of raw addresses, and a moving/compacting GC +between passes can relocate or collect any of them. + +JVMTI object tags solve this: + +- Tags are identity-based and GC-move-transparent — a tagged object can be re-resolved + after a GC regardless of where it moved. +- Tags are **non-retaining** — tagging does not keep an object alive. This is the same + property the existing live-object sampler in this codebase already relies on, so this + is a new *use* of an existing mechanism, not new risk surface. +- Non-retention gives incremental resumption a useful side effect for free: if a frontier + object dies between passes, it simply fails to re-resolve on the next pass. That branch + of the search is pruned automatically, with no extra liveness bookkeeping required. + +### Data structures + +- **Frontier**: a set of `(tag, parent_tag, referrer_klass, depth)` records. `tag` is the + JVMTI tag assigned to a not-yet-expanded object; `parent_tag` links back for chain + reconstruction; `depth` supports the hop cap. +- **EdgeStore**: accumulates `(referrer_klass, parent_tag, depth)` per discovered edge for + objects that are on a path toward a target sample. Keyed by tag, not address — + degenerate relative to JFR's `EdgeStore` since object identity/field names are not + required, but it retains the same `parent_tag` linkage field as the Frontier so a chain + can be walked back from a target sample to a root by following `parent_tag` across + EdgeStore records. + +### Algorithm + +1. Seed the frontier from GC roots (first pass) or from the persisted frontier + (resumed pass). +2. Resolve currently-live tagged frontier objects. Objects that fail to resolve are + dropped (dead — free pruning). +3. Expand the frontier up to a fixed per-pass budget (edge count or time slice). +4. Newly discovered objects are tagged and added to the frontier for the next pass. +5. Persist the frontier (native memory owned by the agent, not thread-local scratch) and + return control to the VM. +6. Repeat until: a target sample is reached, the hop cap is hit, or a per-search + abandonment limit (see Termination) is exceeded. + +### Triggering passes: resolved — cannot avoid dedicated safepoints + +Investigated whether pass-continuation work could ride the JVMTI +`GarbageCollectionStart`/`GarbageCollectionFinish` callbacks — the same callback this +codebase already uses to call `_heap_usage_func` (`vmStructs.cpp`) — instead of +scheduling a dedicated `VM_Operation` per pass. + +**Confirmed the VM is genuinely at a safepoint (all mutators stopped) for the full +duration of both callbacks**, on every collector: + +- JVMTI spec: *"This event is sent while the VM is still stopped... the event handler + must not use JNI functions and must not use JVM TI functions except those which + specifically allow such use (see the raw monitor, memory management, and environment + local storage functions)."* +- openjdk/jdk source: delivery is synchronous on the VMThread + (`src/hotspot/share/prims/jvmtiExport.cpp:2752-2790`, comment *"this event is posted + from VM-Thread"*); every call site is inside a safepoint-executing `VM_Operation::doit()`, + backed by explicit asserts — e.g. Parallel GC's + `assert(SafepointSynchronize::is_at_safepoint())` (`gc/parallel/psScavenge.cpp:305-306`), + G1's `assert_at_safepoint_on_vm_thread()` (`gc/g1/g1VMOperations.cpp:141-157`), + Shenandoah and ZGC wrapping the same `SvcGCMarker` only inside their respective + `VM_Operation`/`VM_ZOperation::doit()` paths. Stable JDK 11 → mainline, across + Serial/Parallel/G1/Shenandoah/ZGC. + +**But this does not make the callback usable as the execution vehicle for a pass.** The +"functions which specifically allow such use" are exactly two: `Allocate` and +`Deallocate` (the entire **Memory Management** category). `SetTag`, `GetTag`, +`GetObjectsWithTags`, `FollowReferences`, and `IterateThroughHeap` are all in the +**Heap** category, which is *not* on that allowlist — calling any of them from inside +`GarbageCollectionStart`/`Finish` is exactly what the restriction forbids. The spec's own +prescribed escape hatch — notify a raw monitor from the callback, do the real work on a +separate agent thread — doesn't preserve the "rides the pause" property either: by the +time the woken agent thread runs, `VM_Operation::doit()` has already returned and the +safepoint has been released, so the tagging/walk work ends up running concurrently with +resumed mutators, not during the STW window. + +The only way to do the tag/walk work *while actually inside* the callback's STW window +would be to bypass the official JVMTI entry points and reach into HotSpot's internal +`JvmtiTagMap` directly via symbol-sniffing — reintroducing exactly the fragility class +already rejected for Approach C (unversioned internal C++ state, no stable ABI). Doing +that here would undo the reason C was rejected. + +**Conclusion: "no new marginal safepoints" is not achievable while staying within +official JVMTI usage.** Each pass needs its own dedicated, budget-capped `VM_Operation`. +The GC callbacks remain useful only as a low-cost *signal* ("a GC just happened, a pass +may be worth scheduling soon") — not as the execution vehicle for the pass itself. This +does not change the core incremental design (frontier persistence via JVMTI tags, +self-pruning of dead branches, per-pass budget) — it only removes the "zero marginal +safepoints" claim from the cost/benefit case. The design's actual value remains what it +was framed as: trading one long pause for several short, independently-scheduled ones — +a latency-distribution improvement, not a total-STW reduction. + +### Termination and abandonment + +Because passes are spread across a mutating heap, a search that never reaches a root or +the hop cap could otherwise persist indefinitely, accumulating abandoned frontier state +across GC cycles. Required cutoffs: + +- Hop cap (as in Approach B's single-pass form). +- A hard cap on passes-per-search or wall-clock TTL from first observation. +- Explicit reporting of abandoned searches (no silent truncation) so this shows up as a + measurable "chain not found within budget" outcome rather than being indistinguishable + from "no chain exists." + +### Correctness note: chains are historical, not a single consistent snapshot + +A chain built across multiple passes stitches together `"A referenced B"` facts observed +at different points in time, not one frozen graph. For the stated purpose — explaining, +by referrer type, what typically retains this class of surviving object — this is +sufficient, and is not meaningfully weaker than a single-pass walk: GC roots (e.g. thread +stack frames) are themselves a live-changing set across a single pause's boundary, so +"one true snapshot" is already an approximation in the single-pass case. Any +documentation or output surface built on this must describe results as an **observed** +retaining path, not a claim about the object's current exact retention state. + +### Cost/benefit summary + +- **Does not reduce total STW time.** Each safepoint/callback entry pays fixed + synchronization overhead; K short increments likely sum to equal or *more* aggregate + pause time than one contiguous walk covering the same work. +- **Improves latency distribution.** No single long tail pause — the thing most likely to + actually affect deployed application health (p99 latency, heartbeat timeouts), even + when total accumulated pause-ms is flat or slightly worse. + +## Non-goals + +- Exhaustive paths to all GC roots. +- Field-level or object-identity-level chains (referrer *type* only). +- Any GC-internal-closure hook (Approach C) or internal parallel-iteration API use as a + launch dependency. + +## Open questions before implementation + +1. ~~Confirm `GarbageCollectionStart`/`GarbageCollectionFinish` callback timing relative to + safepoint release.~~ **Resolved** (see Triggering section): the callback is genuinely + at a safepoint, but the JVMTI Heap-category functions needed to do frontier work + (`SetTag`/`GetTag`/`FollowReferences`/`IterateThroughHeap`) are not in the callback's + allowed function set, so each pass still needs its own dedicated `VM_Operation`. The + "no new marginal safepoints" framing is dropped; the design's value is latency + distribution, not total-STW reduction. +2. Choose per-pass budget defaults (edge count vs. time slice) and hop cap — needs + measurement against representative heap shapes, not a guess. +3. Decide the sample-batching policy: one incremental search per live-heap sample, or + batched multi-target BFS sharing a single frontier walk (batching amortizes better but + couples unrelated samples' termination conditions together). +4. Decide behavior when JVMTI tagging is already saturated by the existing live-object + sampler (tag-table sizing/contention) — this reuses infrastructure that has other + consumers in this codebase. +5. Decide the actual pass-scheduling policy now that GC callbacks can only be a signal, + not a vehicle: e.g. a background thread woken by the GC-callback signal that then + requests its own bounded `VM_Operation`, vs. a fixed-cadence timer independent of GC + activity. Needs a cost model for how many dedicated small safepoints per second are + acceptable before this stops being "more palatable" than one larger pause. From 19b9d3f3b1c2368cb6e86543f7b31a2abdf8f4ed Mon Sep 17 00:00:00 2001 From: Jaroslav Bachorik Date: Fri, 18 Sep 2026 14:56:27 +0200 Subject: [PATCH 2/2] Drop references to uncommitted plan docs and the superseded branch --- .../LiveHeapReferenceChains-Algorithm.html | 1450 +++++++++++++++++ .../LiveHeapReferenceChains-Implementation.md | 2 +- doc/architecture/LiveHeapReferenceChains.md | 58 +- 3 files changed, 1475 insertions(+), 35 deletions(-) create mode 100644 doc/architecture/LiveHeapReferenceChains-Algorithm.html diff --git a/doc/architecture/LiveHeapReferenceChains-Algorithm.html b/doc/architecture/LiveHeapReferenceChains-Algorithm.html new file mode 100644 index 0000000000..e6412ffea4 --- /dev/null +++ b/doc/architecture/LiveHeapReferenceChains-Algorithm.html @@ -0,0 +1,1450 @@ + + + + + +Reference Chains — Leak Triggers & Resumable Heap Walk + + + + +
+

Reference Chains: leak-detection triggers & the resumable heap walk

+

Extracted from ddprof-lib/src/main/cpp/referenceChains.{h,cpp}, + livenessTracker.{h,cpp}, painBudget.h — branch jb/reference-chains-pi. + Every claim carries a file:line anchor.

+
+ + + +
+ + +
+

1 · End-to-end pipeline

+

Two independent subsystems cooperate. LivenessTracker decides that + something is leaking and which class. ReferenceChainTracker then answers why it is + retained, by running a budgeted, resumable, tag-based breadth-first walk of the reachable object graph on + its own JVMTI-attached thread.

+ +
+flowchart LR + subgraph LT["LivenessTracker (leak detection)"] + direction TB + GC["GC epoch advance"] --> POP["cleanup_table(): fold survivors
into per-klass population rings"] + POP --> TREND["hasQualifyingGrowth():
least-squares slope + hysteresis"] + POP --> FLOOR["heapFloorRising() /
secondsToOOM() projection"] + TREND --> CAND["selectLeakCandidates()
top-5 by slope"] + FLOOR --> CAND + end + + subgraph RC["ReferenceChainTracker (retention explanation)"] + direction TB + SIG["hasLeakSignal()"] --> AFF["canAffordNewSearch()
pain budget + signal"] + AFF --> SCHED["shouldRunPass()"] + SCHED --> PASS["runPass() — one budgeted pass"] + PASS --> FRONT[("FrontierTable
tag -> parent_tag")] + FRONT --> PASS + PASS --> TERM["Termination check"] + TERM -->|"terminal"| REL["releaseSearchTags() -> restartSearch()"] + REL --> SCHED + end + + CAND --> SIG + CAND --> POLL["pollWatchedTargets():
tag candidate instances"] + POLL --> PASS + FRONT --> RECON["reconstructChain()
leaf -> root"] + RECON --> JFR["JFR: ReferenceChain /
ReferenceChainAbandoned"] +
+ +
+
+

One thread, one lock

+

A single agent-owned pthread (threadLoop(), referenceChains.cpp:690) + wakes on an adaptive cadence, runs at most one pass, then polls targets. GC callbacks only bump an atomic epoch.

+
+
+

Tags, not handles

+

Frontier identity is a JVMTI object tag — non-retaining, so the walk never keeps a dying object alive. + Dead objects vanish for free at the next GetObjectsWithTags.

+
+
+

Resumable by construction

+

Every pass is bounded by an edge budget and a wall-clock deadline. Unfinished work stays in a FIFO + (_pending_expand) and is picked up by the next pass — including a partially-visited batch.

+
+
+

Cost is paid for, not assumed

+

Two leaky buckets (PainBudget): one gates starting a search on accumulated safepoint cost, + one gates every pass on non-safepoint CPU cost.

+
+
+ +
+ Three places where the in-source narrative has drifted from the code — verified, worth knowing before reading the headers: +
    +
  1. There is no root-seeded FollowReferences first pass any more. runPass() always + drives runPassManualWalk(); the only first-pass difference is that root enumeration is forced + (referenceChains.cpp:3665-3699). The header comment at + referenceChains.h:47-53 still describes the old shape.
  2. +
  3. SearchAbandonReason::TTL is not driven by _ttl_ms. That field is assigned in + start() and only reported in the abandoned event; the branch that stores TTL is the + no-progress detector (referenceChains.cpp:3807-3818).
  4. +
  5. CANARY_NO_PROGRESS_PASS_LIMIT = 3 carries an in-source // TEMP: was 30, lowered for testing + marker (referenceChains.h:2376), while the comment below it still reasons about a base of 30.
  6. +
+
+
+ + +
+

2 · Leak signal: how a class becomes a candidate

+

A klass is only trusted as leaking after it clears a fill gate, a slope gate, a class-level hysteresis + gate and a per-thread hysteresis gate. The bar itself moves depending on whether the aggregate heap floor + corroborates the story.

+ +
+flowchart TB + A["GC epoch: cleanup_table() folds surviving
tracked objects per klass"] --> B["count_ring[30] push
(sample = distinct GC ages, not raw count)"] + B --> C{"ring_fill >= 10
KLASS_POPULATION_MIN_FILL_FOR_TREND"} + C -->|"no"| X1["no trend yet"] + C -->|"yes"| D["least-squares regression over the ring
cached_slope = recent_mean - earliest_mean"] + D --> E{"cached_slope >= max(0.15 * earliest_mean, 1)
LEAK_GROWTH_REL_MIN / ABS_MIN"} + E -->|"no"| F["consecutive_positive = 0 (hard reset)"] + E -->|"yes"| G["consecutive_positive++ (saturating)"] + + G --> H{"heapFloorRising()?"} + H -->|"yes, corroborated"| I["required_hysteresis = 3"] + H -->|"no"| J["required_hysteresis = 5"] + I --> K + J --> K{"consecutive_positive >= required_hysteresis
AND cached_slope > 0"} + K -->|"no"| X2["skip klass"] + K --> L{"any tid_trend with
consecutive_positive >= required_hysteresis"} + L -->|"no"| X3["skip klass — no single thread
owns the growth"] + L -->|"yes"| M["insert into top-5 by slope descending"] + M --> N["selectLeakCandidates() returns candidates"] +
+ +
+ Latency to first candidate ≈ 10 ring pushes to fill the trend window, plus 3–5 further qualifying + epochs of hysteresis — roughly 13–15 GC epochs. This is exactly why the search must be restartable: a search that + completes the whole reachable graph before the trend detector has spoken would otherwise never look again. +
+ +

Heap-floor corroboration

+

heapFloorRising() (livenessTracker.cpp:1220-1259) demands + both a rising mean and a rising minimum, so a sawtooth workload whose troughs stay flat does not lower the bar.

+ + + + + + + + +
GateConstantValue
Mean rise, relativeHEAP_FLOOR_GROWTH_REL_MIN0.02
Mean rise, absoluteHEAP_FLOOR_GROWTH_ABS_MIN1 MiB
Floor rise, relativeHEAP_FLOOR_FLOOR_REL_MIN0.01
Floor rise, absoluteHEAP_FLOOR_FLOOR_ABS_MIN512 KiB
+ +
+ Two rankings, deliberately different. selectLeakCandidates() is slow and gated + (livenessTracker.cpp:1370). topKlassesByGenerationCount() is fast and + ungated — ranked on the single latest sample with no hysteresis at all + (livenessTracker.cpp:1480). The fast one is only ever consulted after the slow one + has already fired, so it never needs to wait out the same hysteresis twice + (referenceChains.cpp:4175-4201). +
+ +
+

Qualifying growth bar: max(0.15 · earliest_mean, 1)

+

Below an earliest-window mean of ~6.7 tracked instances, the absolute floor + (LEAK_GROWTH_ABS_MIN = 1) sets the bar, not the 15% relative term + (LEAK_GROWTH_REL_MIN) — a klass with only a handful of instances still needs to grow by a whole + instance to qualify, it cannot coast in on a tiny relative slope.

+
+
+
+ + +
+

3 · Urgency: the OOM projection and its latch

+

secondsToOOM() projects when the live-heap floor will hit the tighter of the JVM max heap + and the container limit. The raw value swings by orders of magnitude between observations, so it is never compared + bare — it is read through a latch with separate arm and release bars.

+ +
+flowchart TB + A["_heap_floor_ring / _time_ring
30 post-GC samples"] --> B{"_gc_generations on
AND max_heap > 0"} + B -->|"no"| N1["return -1 (no projection)"] + B -->|"yes"| C["limit = min(JVM max heap, container limit)"] + C --> D{"ring_fill >= 10"} + D -->|"no"| N2["return -1 (INSUFFICIENT_FILL)"] + D -->|"yes"| E["regress bytes and time"] + E --> F{"bytes_delta > 0 AND time_delta > 0"} + F -->|"no"| N3["return -1 (NOT_RISING)"] + F -->|"yes"| G["re-regress the most recent half
(min fill 5)"] + G --> H{"recent half also rising?"} + H -->|"no"| N4["return -1 (RECENT_HALF_FLAT)
rejects a plateaued step change"] + H -->|"yes"| I["seconds = (limit - recent_mean) / rate"] +
+ +
+stateDiagram-v2 + direction LR + [*] --> Calm + Calm --> Latched: "secondsToOOM() in [0, 300) — OOM_URGENT_THRESHOLD_S
also clears _urgent_search_spent, so a new episode earns a fresh entitlement" + note right of Latched: "a reading between 300 and 600 stays Latched
and resets _urgent_release_ticks to 0" + Latched --> Releasing: "reading at or over 600 (OOM_URGENT_RELEASE_S), or negative — ++_urgent_release_ticks" + Releasing --> Latched: "any reading back under the release bar" + Releasing --> Calm: "URGENT_RELEASE_CONSECUTIVE = 5 clear readings in a row" +
+ +
+ Two distinct notions of "urgent" coexist. + isUrgent() uses the latch above with a 300 s arm bar and is what suppresses the no-progress abandon and + authorises one out-of-band search per episode. + threadLoop()'s ramp uses a raw, unlatched, much wider window — + seconds_to_oom < OOM_RAMP_START_S = 1800 s — recomputed on every wake + (referenceChains.cpp:735-736). They arm at different distances and flap differently. +
+ +

The OOM ramp (referenceChains.cpp:739-772)

+

Inside the 30-minute window, with x = 1 - secondsToOOM()/1800 (0 at the far edge, 1 at OOM), + both the pause target and the cadence ramp exponentially toward their urgent ceilings:

+ + + + + + + +
QuantityBaselineUrgent ceilingRamp
Per-pass pause target_pause_target_msURGENT_PAUSE_TARGET_MS = 100 msbase * (100/base)^x
Pass cadencePASS_CADENCE_NS = 1 sURGENT_CADENCE_NS = 10 ms1s * (10ms/1s)^x
Edge budget ceiling_budgetmin(_budget * 4, MAX_REFERENCE_CHAINS_BUDGET)step, on target change
+

The cadence ramps from the fixed baseline, not from the live _effective_cadence_ns — + anchoring on the moving value would compound the exponent across iterations. While urgent, the ramp owns + _effective_cadence_ns outright; updatePacing() silently resumes ownership the moment urgency clears.

+ +
+

Cadence ramp: 1000ms · (10ms/1000ms)^x

+

Fully determined by PASS_CADENCE_NS and URGENT_CADENCE_NS — no + assumed baseline. The dashed marker is where isUrgent()'s own, separately-latched threshold + (OOM_URGENT_THRESHOLD_S = 300 s) falls on this same x-axis, for scale.

+
+
+
+ + +
+

4 · Pass scheduling: shouldRunPass()

+

Called once per BFS-thread wake. Three regimes — no search yet, search terminal, search running — + evaluated strictly in this order (referenceChains.cpp:870-1006).

+ +
+flowchart TB + S(["shouldRunPass(now_ns)"]) --> B1{"_search_started?"} + + B1 -->|"no"| A1{"canAffordNewSearch(now)
= safepoint pain budget drained
AND hasLeakSignal()"} + A1 -->|"no"| F1["FALSE"] + A1 -->|"yes"| T1["_urgent_search_spent = _urgent_latched
TRUE — take the very first pass"] + + B1 -->|"yes"| B2{"_search_state == RUNNING?"} + + B2 -->|"no (terminal)"| C1{"_tags_released?"} + C1 -->|"no"| T2["TRUE — force runPass() to retry
releaseSearchTags(); restart is forbidden
until every tag is confirmed cleared"] + C1 -->|"yes"| C2{"canAffordNewSearch(now)"} + C2 -->|"no"| F2["FALSE — terminal, waiting"] + C2 -->|"yes"| T3["restartSearch()
TRUE"] + + B2 -->|"yes"| D0["canary_active = candidates outstanding
all_covered = leak tags all resolved
emergency = canary_active AND no candidate
progress for 3 passes"] + D0 --> D1["CPU pain refill multiplier:
all_covered -> 1x
emergency -> 100x
canary_active -> 15x
else -> 1x"] + D1 --> D2{"_cpu_pain_budget.canStartNow()"} + D2 -->|"no"| F3["FALSE — throttled"] + D2 -->|"yes"| D3{"gcFinishEpoch() changed
since last pass?"} + D3 -->|"yes"| T4["TRUE — GC trigger"] + D3 -->|"no"| D4{"canary_active?"} + D4 -->|"yes"| T5["TRUE — run back-to-back,
cadence bypassed"] + D4 -->|"no"| D5{"now - _last_pass_ns >=
_effective_cadence_ns"} + D5 -->|"yes"| T6["TRUE — cadence trigger"] + D5 -->|"no"| F4["FALSE — idle"] +
+ +
+
+

Order matters in the multiplier

+

all_covered is tested before emergency, so a search whose leak tags are all resolved + drops back to 1× even if the canary is nominally stuck.

+
+
+

Two budgets, two scopes

+

_safepoint_pain_budget is never consulted for a RUNNING search — only through + canAffordNewSearch() when starting or restarting one. _cpu_pain_budget gates each pass.

+
+
+

No early wake on GC

+

onGCFinish() bumps an atomic epoch and nothing else. Waking the thread per GC would buy ≤1 s of + latency while collapsing the loop cadence to GC frequency (referenceChains.cpp:858-866).

+
+
+

Sleep only when idle

+

threadLoop() skips its sleep entirely when a pass is about to run, so canary passes execute + back-to-back and the PID controller alone regulates cost (referenceChains.cpp:805-811).

+
+
+ +
+ Why a RUNNING search is not gated on hasLeakSignal() + (referenceChains.cpp:785-798): that signal answers "is there a leak candidate right now", + which is unrelated to whether an in-flight search still has pending frontier work. Gating every pass on it would + stall a search's own convergence whenever no candidate happens to be visible. +
+
+ + +
+

5 · One search, pass by pass

+

A search is a sequence of passes over a persistent frontier. Each pass performs up to four sub-phases, + each of which can truncate independently and hand its remainder to the next pass.

+ +
+flowchart TB + P(["runPass(jvmti, jni)"]) --> G0{"enabled AND frontier exists?"} + G0 -->|"no"| Z0["bail"] + G0 -->|"yes"| G1{"_search_state == RUNNING?"} + G1 -->|"no"| Z1["retry releaseSearchTags() if needed;
return as a no-op"] + G1 -->|"yes"| R0["resolveLoadedClasses()
class_tag -> class-name dictionary id
(per-class scan skipped when the count is unchanged)"] + + R0 --> R1{"first pass?
(_search_started == false)"} + R1 -->|"yes"| R2["_search_started = true
_search_start_ns = now
force root enumeration"] + R1 -->|"no"| R3{"last root enum truncated
OR >= 2s since last
(ROOT_ENUM_MIN_INTERVAL_NS)"} + R3 -->|"yes"| R2 + R3 -->|"no"| R4["skip root enumeration this pass"] + + R2 --> W["runPassManualWalk()"] + R4 --> W + + subgraph WALK["runPassManualWalk — four sub-phases under one deadline"] + direction TB + W1["A · Root enumeration
IterateOverReachableObjects
budget = _first_pass_budget · no deadline"] + W2["B · Static-field sweep
admitStaticFieldRoots(), 512-class chunk
only when the loaded-class count changed"] + W3["C · Ordinary expansion
expandFrontier(_pending_expand / _priority_expand)"] + W4["D · Rotation
collect stale roots / leak accumulators / stale expanded
then a second expandFrontier() on the reserved budget"] + W1 --> W2 --> W3 --> W4 + end + + W --> WALK + WALK --> M0["_passes_run++ · _last_pass_gc_finish_epoch · _last_pass_ns"] + M0 --> M1{"root-enum pass?"} + M1 -->|"no"| M2["updatePacing(safepoint_ticks)"] + M1 -->|"yes"| M3["maybeRevokeBorrowForRootEnumPass()
(excluded from the PID signal)"] + M2 --> M4 + M3 --> M4["_search_pain_ms += safepoint ms
_cpu_pain_budget.spend(non-safepoint ms)"] + M4 --> M5["Termination decision chain — see tab 8"] +
+ + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
Sub-phaseBudgetDeadlineWhy it exists
A · Root enumeration_first_pass_budget (auto = min(_budget*50, 200000))none, by designEnumerates heap roots and stack refs. Re-run at most every 2 s; already-admitted roots short-circuit + cheaply, so a root missed this pass is picked up later, never lost.
B · Static-field sweepchunk of 512 classes_pass_deadline_nsJVMTI has no jvmtiHeapRootKind for static fields, and expansion never descends from class + objects — without this, SomeClass.staticField → obj is structurally undiscoverable.
C · Ordinary expansion_effective_budget − rotation reserve_pass_deadline_nsThe actual BFS: resolve a batch of frontier tags, walk exactly one hop from each.
D · Rotationmin(expand_budget/2, 288) reserved up front_pass_deadline_nsRe-observes entries whose recorded story may have gone stale: transient root kinds, leak-accumulating + signatures, long-expanded entries whose fields have since been mutated.
+
+ + +
+

6 · expandFrontier(): the resumable batch loop

+

The heart of the resumability. Each iteration resolves a batch of frontier tags back to live objects, + runs one stop-the-world FollowReferences for the whole batch, and descends exactly one hop — + enforced by a membership gate on the batch's own tag set (referenceChains.cpp:2943-3303).

+ +
+flowchart TB + E(["expandFrontier(hop_cap, budget)"]) --> L0{"deadline reached?
OS::nanotime() >= _pass_deadline_ns"} + L0 -->|"yes"| TR["truncated = true — break"] + L0 -->|"no"| L1["pick a lane:
_priority_expand vs _pending_expand,
alternating via _expand_lane_prefer_priority"] + L1 --> L2{"both lanes empty?"} + L2 -->|"yes"| DONE["no pending work — break"] + L2 -->|"no"| L3["batch_size = min(queue, budget, _gotw_batch_size)"] + L3 --> L4["GetObjectsWithTags(batch)
dead tags simply do not come back = free pruning"] + L4 --> L5["self-calibrate _gotw_batch_size:
per-call EMA scaled by the remaining
deadline window, clamped to [8, 512]"] + L5 --> L6["build jobjectArray holder
+ fill batch_tags membership set"] + L6 --> L7["FollowReferences(holder) — one STW HeapWalkOperation
heapReferenceCallback descends only into tags in batch_tags"] + L7 --> L8{"truncated?"} + + L8 -->|"no"| S1["for each tag in batch:
resolved -> markExpanded()
unresolved -> clear() = ABANDONED
pop_front() · progress = true"] + L8 -->|"yes, some batch entry visited"| S2["rolling resume:
settle only entries BEFORE
_last_visited_batch_tag;
leave it and the rest at the queue front"] + L8 -->|"yes, nothing visited"| S3["leave the whole batch queued"] + + S1 --> L0 + S2 --> TR + S3 --> TR +
+ +
+
+

One hop, guaranteed

+

heapReferenceCallback returns JVMTI_VISIT_OBJECTS only when the visited object's tag is + in batch_tags. Everything else is admitted but not descended into + (referenceChains.cpp:2001-2017).

+
+
+

Rolling resume

+

_last_visited_batch_tag is the cursor inside a truncated batch. Entries before it are settled; + the partially-visited one stays at the queue head and is re-walked next pass.

+
+
+

Deadline sampled cheaply

+

Inside the callback the wall clock is only read every 4096th invocation + ((++counter & 0xFFF) == 0) — the loop top checks it per batch + (referenceChains.cpp:1647-1656).

+
+
+

Adaptive batch size

+

_gotw_batch_size is tuned per call from measured GetObjectsWithTags cost against the + remaining deadline window, so a slow JVM naturally shrinks its batches rather than blowing the pause target.

+
+
+ +

Admission, in callback order

+
+flowchart TB + C(["heapReferenceCallback(kind, tag_ptr, referrer_tag_ptr)"]) --> A0{"_abort_pass_requested"} + A0 -->|"yes"| Z1["truncated · VISIT_ABORT"] + A0 -->|"no"| A1{"every 4096th call:
past _pass_deadline_ns?"} + A1 -->|"yes"| Z1 + A1 -->|"no"| A2{"tag at or below MARKER_TAG_BASE
(canary marker)"} + A2 -->|"yes"| Z2["record chain link, set found bit,
do NOT descend"] + A2 -->|"no"| A3{"tag is negative (class object)"} + A3 -->|"yes"| Z3["descend only for the static-field
seed holder -> class edge"] + A3 -->|"no"| A4["compute parent_tag and depth
from the referrer's tag"] + A4 --> A5{"depth >= hop_cap"} + A5 -->|"yes"| Z4["drop — no admission, no descent"] + A5 -->|"no"| A6{"isLeakTag(tag)?"} + A6 -->|"yes"| Z5["allocate frontier tag, insert,
setLeakTag(), record instance, descend"] + A6 -->|"no"| A7{"tag == 0 (unseen)"} + A7 -->|"yes"| A8["admitObject()"] + A7 -->|"no"| A9["already admitted ->
improveChain() / reparentToDurableRoot()
/ maybeUpgradeRootAttachedRootKind()"] + A8 --> A10{"result"} + A10 -->|"BUDGET_EXHAUSTED"| Z6["truncated · abort"] + A10 -->|"FRONTIER_CAP_HIT"| Z7["frontier_cap_hit · abort
-> whole search abandoned"] + A10 -->|"ADMITTED"| A11["push onto priority or pending queue"] + A9 --> A12 + A11 --> A12{"tag in batch_tags?"} + A12 -->|"yes"| Z8["JVMTI_VISIT_OBJECTS — descend one hop"] + A12 -->|"no"| Z9["return 0 — do not descend"] +
+ +
+ ALREADY_ADMITTED is the idempotency that makes re-enumeration cheap. + admitObject() short-circuits on any non-zero tag + (referenceChains.cpp:2022-2057), which is precisely why root enumeration can be re-run on + every pass without re-paying for the graph it already discovered. +
+
+ + +
+

7 · Root discovery and the chunked static-field sweep

+

Two disjoint sources of root-attached entries. IterateOverReachableObjects reports stack + locals, JNI handles, monitors and thread roots — but never static fields, which need their own sweep.

+ +
+flowchart TB + subgraph RE["A · IterateOverReachableObjects"] + direction TB + R1["heapRootCallback / stackRefCallback"] --> R2["translateHeapRootKind():
jvmtiHeapRootKind -> jvmtiHeapReferenceKind"] + R2 --> R3["admitObject(parent_tag = 0, depth = 0, root_kind)"] + R3 --> R4{"ALREADY_ADMITTED?"} + R4 -->|"yes"| R5["maybeUpgradeRootAttachedRootKind()
durability tie-break"] + R4 -->|"no"| R6["new root-attached frontier entry"] + end + + subgraph SF["B · admitStaticFieldRoots — chunked, cursor-resumed"] + direction TB + S1["GetLoadedClasses()"] --> S2["partition app classes
(non-null loader) to the front"] + S2 --> S3["chunk = [cursor, cursor + 512)"] + S3 --> S4["holder array filled in REVERSE chunk order
so HotSpot's LIFO descent visits ascending"] + S4 --> S5["one FollowReferences with static_field_seed = true
empty batch_tags -> exactly one hop past each class"] + S5 --> S6{"truncated?"} + S6 -->|"yes"| S7["mark lap truncated;
cursor = chunk_start + classes_visited - 1
(redo the partial class)"] + S6 -->|"no"| S8["cursor = chunk_end"] + S7 --> S9 + S8 --> S9{"cursor reached class_count?"} + S9 -->|"yes"| S10["lap wrap: cursor = 0;
cycle_complete only if no chunk truncated"] + S9 -->|"no"| S11["resume here next pass"] + end +
+ +
+ Why chunking is not an optimisation but a correctness fix. On a JVM with ~34k loaded classes a single + FollowReferences over every class cannot finish inside a 5–50 ms deadline — the observed behaviour was + truncated = 1 on 275 of 275 passes with 0–1 edges admitted. Because a single-call sweep restarts at + class 0 every time, every class past the deadline point was permanently unreachable + (referenceChains.h:1936-1960). +
+ +

Root-kind durability

+

A root-attached entry records why it is reachable. When a more durable root is later observed + admitting the same object, the recorded kind is upgraded rather than keeping whichever root happened to be enumerated + first (referenceChains.h:249-278).

+ + + + + + + +
TierKindsMeaning
3 — most durableSTATIC_FIELD, SYSTEM_CLASSGenuine retention evidence.
2JNI_GLOBALDurable, but owned outside the JVM heap.
1 — transientMONITOR, STACK_LOCAL, JNI_LOCAL, THREAD, OTHER"First observed via", not "rooted by". THREAD and OTHER have no documented tier and are conservatively bucketed here.
+

Two repair operations exist because a depth comparison alone cannot express both cases: + improveChain() replaces a shallow root-attached entry when a strictly deeper path reaches it, and + reparentToDurableRoot() handles the equal-depth case — a depth-1 entry parented to a transient root is + re-parented to a durable one. Without the latter, the real hotdog shape (a static singleton collection at depth 0, its + elements at depth 1) would keep a stack-local parent forever, because the depths tie.

+
+ + +
+

8 · State machines

+

Two independent state machines: one per frontier entry, one per search.

+ +

Per-entry: FrontierEntryState

+
+stateDiagram-v2 + direction LR + [*] --> FRONTIER: "insert() from admitObject(), leak-tag interception or canary pruning" + FRONTIER --> EXPANDED: "markExpanded() — its batch's FollowReferences completed" + FRONTIER --> ABANDONED: "clear() — GetObjectsWithTags did not resolve the tag (object died)" + note right of EXPANDED: "rotation re-queues the tag, and
the re-walk marks it EXPANDED again" + FRONTIER --> EDGE: "markEdge() — reconstructChain() walked through this hop" + EXPANDED --> EDGE: "markEdge()" + EDGE --> EXPANDED: "rotation re-expansion overwrites the state" + EXPANDED --> ABANDONED: "releaseSearchTags() at search end" + EDGE --> ABANDONED: "releaseSearchTags() at search end" + ABANDONED --> [*]: "resetForRestart() zeroes _table_size — every slot becomes un-inserted" +
+
+ ABANDONED means "the JVMTI tag was released", not "the record was discarded". + releaseSearchTags() leaves parent_tag, referrer_klass, depth and + root_kind intact, so reconstructChain() keeps working from memory after the search ends + (referenceChains.h:1966-1975). +
+ +

Per-search: SearchState

+
+stateDiagram-v2 + direction TB + [*] --> RUNNING: "first shouldRunPass() that clears canAffordNewSearch()" + + RUNNING --> ABANDONED_FC: "frontier_cap_hit — checked first, wins over everything" + RUNNING --> COMPLETED_G: "not truncated AND no watched leak klass" + note right of RUNNING: "not truncated BUT a watch is active —
stays RUNNING, rotation keeps re-observing" + RUNNING --> ABANDONED_TTL: "30 passes with no frontier growth AND not isUrgent()" + RUNNING --> COMPLETED_C: "all canary candidates found" + RUNNING --> ABANDONED_CS: "candidates outstanding AND 30 passes without frontier growth AND canary stall past canaryStuckPassLimit()" + + ABANDONED_FC --> RELEASE + ABANDONED_TTL --> RELEASE + ABANDONED_CS --> RELEASE + COMPLETED_G --> RELEASE + COMPLETED_C --> RELEASE + + note right of RELEASE: "releaseSearchTags() failed — shouldRunPass()
returns true to retry; restart stays blocked" + RELEASE --> RUNNING: "_tags_released AND canAffordNewSearch() -> restartSearch()" + + state "ABANDONED / FRONTIER_CAP" as ABANDONED_FC + state "ABANDONED / TTL (no-progress)" as ABANDONED_TTL + state "ABANDONED / CANARY_STUCK" as ABANDONED_CS + state "COMPLETED (graph exhausted)" as COMPLETED_G + state "COMPLETED (candidates found)" as COMPLETED_C + state "tag release + canary reset" as RELEASE +
+ + + + + + + + + +
ReasonValueSuppressed by isUrgent()?Rationale
NONE0—Not abandoned.
FRONTIER_CAP1NoThe metadata table is full; nothing further can be admitted at all.
TTL2YesA search still making real progress must not be killed just because the process is close to OOM.
CANARY_STUCK3NoA candidate chase with zero discovery progress is provably not converging; continuing at urgency-boosted budget only burns pause budget the dying process needs.
+
+ + +
+

9 · Termination, tag release and restart

+

Six checks, evaluated in a fixed priority order at the end of every pass + (referenceChains.cpp:3765-3856). Only the first match applies.

+ +
+flowchart TB + T(["end of runPass()"]) --> C1{"frontier_cap_hit"} + C1 -->|"yes"| A1["ABANDONED / FRONTIER_CAP
enqueue abandoned event"] + C1 -->|"no"| C2{"no pending frontier
AND _watched_leak_klass_count == 0"} + C2 -->|"yes"| A2["COMPLETED — graph exhausted
within the caps"] + C2 -->|"no"| C3{"no pending frontier
but a watch is active"} + C3 -->|"yes"| A3["stay RUNNING — rotation keeps
re-observing mutated fields"] + C3 -->|"no"| C4{"_passes_since_last_progress >= 30
AND NOT isUrgent()"} + C4 -->|"yes"| A4["ABANDONED / TTL"] + C4 -->|"no"| C5{"all candidates found"} + C5 -->|"yes"| A5["COMPLETED + counter
REFERENCE_CHAIN_CANDIDATES_FOUND"] + C5 -->|"no"| C6{"candidates outstanding
AND 30 passes no frontier growth
AND canary stall >= canaryStuckPassLimit()"} + C6 -->|"yes"| A6["ABANDONED / CANARY_STUCK
_canary_stuck_restart_count++"] + C6 -->|"no"| A7["stay RUNNING"] + + A1 --> R + A2 --> R + A4 --> R + A5 --> R + A6 --> R + R["releaseSearchTags():
GetObjectsWithTags over every non-ABANDONED tag,
SetTag(obj, 0), then mark all ABANDONED"] --> R2{"batch call succeeded?"} + R2 -->|"no"| R3["_tags_released = false — mark NOTHING;
a failed batch says nothing about liveness,
and rewinding _next_tag while a live object
still carries a tag would break tag uniqueness"] + R2 -->|"yes"| R4["_tags_released = true
clear canary marker tags · zero candidate state
reset _canary_stuck_restart_count unless CANARY_STUCK"] +
+ +

Escalating patience

+

canaryStuckPassLimit() = CANARY_NO_PROGRESS_PASS_LIMIT << min(_canary_stuck_restart_count, 8) + (referenceChains.h:2389-2393). Repeated CANARY_STUCK abandons double the stall + tolerance, up to 256×. _canary_stuck_restart_count is deliberately not reset by + restartSearch() — it must survive restarts or the escalation could never widen.

+ +
+

Escalation: canaryStuckPassLimit(n) = 3 << min(n, 8)

+

Each consecutive CANARY_STUCK abandon doubles the stall tolerance for the + next attempt at the same candidate chase, up to the MAX_CANARY_STUCK_BACKOFF_SHIFT = 8 cap + (restart count 8 and 9 tolerate the same 768 passes — the shift has saturated).

+
+
+ +

What a restart resets, and what it does not

+
+
+

Reset (restartSearch())

+

_frontier->resetForRestart() (slot occupancy only, the allocation survives) · _next_tag = 1 · + _search_started = false · state back to RUNNING · both expand queues and the membership index · + leak-signature maps · _leak_tags_assigned/_resolved · _last_pass_* · _passes_run · + the resolved/static-field class counts.

+
+
+

Persists across a restart

+

The class-tag table and its allocator · _resolved_chains · _watched_leak_klass_ids · + _canary_stuck_restart_count · the rotation cursors and the static-field sweep cursor · + _pause_pid, _effective_budget, _effective_cadence_ns · the cached + java/lang/Object global ref · the FrontierTable allocation and capacity.

+
+
+

restartSearch() also spends the finished search's accumulated + _search_pain_ms into _safepoint_pain_budget before zeroing it + (referenceChains.cpp:1123) — a cheap search may restart again soon, an expensive one must wait + proportionally longer. It asserts on _tags_released.

+
+ + +
+

10 · Pacing: PID controller, budget borrowing, pain budgets

+

The configured constants are ceilings and baselines, not literal per-pass values. What each pass + actually spends is decided by a feedback loop over the previous pass's measured in-safepoint time.

+ +
+flowchart LR + M["measure the pass:
pass_wall_ticks total,
safepoint_ticks inside
IterateOverReachableObjects / FollowReferences"] --> SPLIT{"split"} + SPLIT -->|"safepoint ms"| PID["_pause_pid.compute(pass_ms)
target = _effective_pause_target_ms
positive signal = came in UNDER target"] + SPLIT -->|"safepoint ms"| PAIN1["_search_pain_ms +=
-> gates the NEXT search"] + SPLIT -->|"non-safepoint ms"| PAIN2["_cpu_pain_budget.spend()
-> gates the NEXT pass"] + + PID --> BORROW{"pass_ms at or under 50% of target
for 5 consecutive passes?"} + BORROW -->|"yes"| B1["_borrowed_budget += _budget * 0.25
capped at 3 * _budget"] + BORROW -->|"no"| B2["_borrowed_budget = 0 immediately"] + B1 --> CL + B2 --> CL["ceiling = _budget + _borrowed_budget
floor = min(2000, ceiling)
_effective_budget = clamp(prev + signal)"] + CL --> OV{"overflow = desired - clamped"} + OV -->|"negative — still over target at the floor"| W["widen _effective_cadence_ns
by 1ms per overflow edge, up to 4s"] + OV -->|"positive"| S["shorten _effective_cadence_ns,
down to 10ms"] + OV -->|"== 0"| K["leave the cadence alone"] + W --> NEXT + S --> NEXT + K --> NEXT["next pass: expand_budget = _effective_budget,
deadline = _effective_pause_target_ms,
sleep/gate = _effective_cadence_ns"] +
+ +
+

Clamp function: effective_budget = clamp(desired, floor, ceiling)

+

Illustrative baseline _budget = 800 edges/pass (below + MIN_EFFECTIVE_BUDGET = 2000). With no borrowed headroom, floor = min(2000, ceiling) = ceiling + — the clamp range collapses to a single point and _effective_budget is pinned regardless of the PID + signal. Once _borrowed_budget pushes the ceiling past 2000 (shown at 3× budget, the borrowing + ceiling), a real band opens up and the floor becomes MIN_EFFECTIVE_BUDGET itself.

+
+
+ +
+
+

Inverted sign convention

+

Unlike ObjectSampler / MallocTracer / RateLimiter, which subtract the PID signal from an interval, this + controller adds it to a budget: a positive signal means the pass came in under target, so the next one may + do more work (referenceChains.h:2002-2012).

+
+
+

One compute() is one pass

+

sampling_window = 1 and time_delta_coefficient = 1.0, because a pass is not a fixed + real-time window like the other three usages assume. Gains are 10/1/2 — deliberately not copied from the shared triple.

+
+
+

Root-enum passes are excluded

+

They spend the deliberately oversized _first_pass_budget; feeding that into the PID would throttle + every cheap expansion pass that follows (referenceChains.cpp:3713-3728).

+
+
+

Borrowing is lost instantly

+

A single pass that is not comfortably under target zeroes both the warm-up streak and the whole accumulated + borrow. Earning headroom takes 5 passes; losing it takes one.

+
+
+ +

The two leaky buckets

+ + + + + + + + +
_safepoint_pain_budget_cpu_pain_budget
GatesStarting or restarting a searchRunning an individual pass
ChargedOnce per finished search, in restartSearch(), with _search_pain_msEvery pass, with the non-safepoint remainder
Refill ratepain_budget_percent / 100Same base, times 1× / 15× (covering) / 100× (emergency)
Read atcanAffordNewSearch()shouldRunPass(), RUNNING branch
+
+ Degenerate configuration: a refill rate of exactly 0.0 never drains + (painBudget.h:46-49), so once anything has been spent the bucket blocks permanently. +
+ +
+

Leaky bucket: _safepoint_pain_budget over time

+

Illustrative simulation, not a captured log — refill rate 2% (an example + pain_budget_percent), two searches finishing at t=0 and t=1600ms and spending their accumulated + _search_pain_ms (60ms, then 45ms) into the bucket. canAffordNewSearch() only returns + true where the curve touches the zero baseline — everywhere the shaded area is above zero, a new search or + restart is blocked.

+
+
+ +

Auto-tuning at start()

+

autoTuneDefaults() (referenceChains.cpp:360-462) fills in only the + knobs the operator did not set explicitly, from max heap size and processor count:

+ + + + + + + + + + +
KnobFormula
Edge budgetDEFAULT * sqrt(heap_mib / 512), clamped to [default, max] — keeps the pause proportional to sqrt(heap)
First-pass budgetbudget * 10, capped
TTLDEFAULT * (heap_mib / 512), clamped to [default, 30 min]
Frontier capscaled by the budget ratio in floating point (integer division here would undershoot by ~20%)
Pause targetDEFAULT * (1 + (nprocs-1)/3), capped at 50 ms
Pain budget %DEFAULT * (1 + (nprocs-1)/4), capped at 5%
+
+ + +
+

11 · From frontier to emitted chain

+

The frontier table doubles as a degenerate edge store: a chain is just the transitive closure of + parent_tag links from a target back to a root-attached entry.

+ +
+flowchart TB + T["target tag"] --> L{"lookup(tag) succeeds?"} + L -->|"no"| F1["return false — never fabricate a partial chain"] + L -->|"yes"| P["push entry.referrer_klass onto the chain
markEdge(tag)
root_kind = entry.root_kind
tag = entry.parent_tag"] + P --> C{"tag == 0?"} + C -->|"no, and hops <= maxCapacity()"| L + C -->|"no, bound exceeded"| F2["corrupt or cyclic -> return false"] + C -->|"yes"| OUT["chain in leaf -> root order
out_root_kind = the root-attached entry's kind"] + OUT --> EV["buildChainEvent():
targetTag = leak_tag if set, else the frontier tag
depth · root_kind · chain"] + EV --> JFR["JFR ReferenceChain event"] + JFR --> JOIN["backend joins ReferenceChain.targetTag
to HeapLiveObject.leakTag"] +
+ +

FrontierEntry fields

+ + + + + + + + + + + +
FieldPurpose
parent_tagTag of the entry that discovered this one; 0 means root-attached. The only link the reconstruction walks.
referrer_klassStringDictionary id of this object's class name, resolved ahead of time — GetClassSignature is illegal inside a heap callback.
depthHop count from the root. Drives the hop cap and improveChain()'s "deeper wins" comparison.
stateFRONTIER / EXPANDED / EDGE / ABANDONED.
leak_tagLivenessTracker's per-instance tag, copied in at admission. Becomes ReferenceChain.targetTag, the backend's join key. 0 = ordinary BFS admission.
root_kindjvmtiHeapReferenceKind of the admitting edge, meaningful only when parent_tag == 0.
class_tagRaw negative JVMTI class tag — stable across StringDictionary regeneration, unlike referrer_klass. Needed by the retroactive leak-accumulation seed scan, which runs long after the live callback is gone.
+
+ No jobject is ever retained. Holding a live handle would defeat the entire point of using + non-retaining JVMTI tags for frontier identity — the walk would keep the leak alive. +
+ +

The canary path

+

Candidate instances are pre-tagged with distinct negative marker tags + (MARKER_TAG_BASE = -(1<<62)) applied to a specific representative object — matching by class alone + would record a chain for an unrelated, possibly short-lived instance of the same class. When the walk hits a marker it + records the chain link and the found bit, but does not descend. + buildCanaryChainEvent() therefore starts from the recorded parent tag (the negative marker itself is + rejected by lookup()), then prepends the candidate's own class and reverses to leaf→root + (referenceChains.h:2647-2712).

+

Beyond the representative, any object of a watched class discovered by the walk is auto-recorded, up to + MAX_DISCOVERED_INSTANCES_PER_CLASS = 8 per slot — each instance's chain is independently useful, since + different instances may be retained by different paths.

+
+ + +
+

12 · Constants reference

+

Values as they stand on this branch. Several are explicitly marked provisional / unbenchmarked in source.

+ +

Leak detection — livenessTracker.h

+ + + + + + + + + + + + + + +
ConstantValueRole
MAX_KLASS_POPULATION_ENTRIES256Tracked klasses.
KLASS_POPULATION_RING_SIZE30Trend window, in GC epochs.
KLASS_POPULATION_MIN_FILL_FOR_TREND10Minimum fill before any slope is trusted.
LEAK_GROWTH_REL_MIN / _ABS_MIN0.15 / 1Qualifying growth bar.
LEAK_TREND_HYSTERESIS_BASE5Consecutive qualifying epochs required.
LEAK_TREND_HYSTERESIS_CORROBORATED3Lowered bar when the heap floor is rising.
TID_TREND_RING_SIZE / MIN_FILL16 / 6Per-thread trend window.
MAX_TID_TRENDS8Threads tracked per klass.
MAX_LEAK_CANDIDATES5Top-k candidate slots.
HEAP_FLOOR_RECENT_HALF_MIN_FILL5Corroboration window for the OOM projection.
+ +

Scheduling and urgency — referenceChains.h

+ + + + + + + + + + + + + +
ConstantValueRole
PASS_CADENCE_NS1 sBaseline cadence and ramp anchor.
OOM_URGENT_THRESHOLD_S300 sLatch arm bar for isUrgent().
OOM_URGENT_RELEASE_S600 s (2×)Latch release bar.
URGENT_RELEASE_CONSECUTIVE5Clear readings needed to unlatch.
OOM_RAMP_START_S1800 sStart of the unlatched exponential ramp in threadLoop().
URGENT_PAUSE_TARGET_MS100 msRamp ceiling for the pause target.
URGENT_CADENCE_NS10 msRamp floor for the cadence.
CANARY_PAIN_BUDGET_COVERING_MULTIPLIER15×CPU refill while chasing candidates.
CANARY_PAIN_BUDGET_REFILL_MULTIPLIER100×CPU refill in the emergency case.
+ +

Walk, budgets and termination — referenceChains.h

+ + + + + + + + + + + + + + + + + + + + +
ConstantValueRole
AUTO_FIRST_PASS_BUDGET_MULTIPLIER / _CAP50 / 200000Auto-scaled first-pass edge budget.
ROOT_ENUM_MIN_INTERVAL_NS2 sMinimum spacing between root enumerations.
STATIC_FIELD_SWEEP_CHUNK_CLASSES512Classes per static-field chunk.
STATIC_FIELD_SWEEP_NON_STATIC_CAP_PER_CLASS32Non-static edges admitted per class per lap.
MIN_EFFECTIVE_BUDGET2000Budget floor under PID control.
MIN/MAX_EFFECTIVE_CADENCE_NS10 ms / 4 sCadence clamp.
CADENCE_NS_PER_EDGE_OVERFLOW1 ms / edgeHow budget overflow converts into cadence widening.
BORROW_WARMUP_PASSES / ceiling5 / 3× _budgetBudget borrowing.
PRIORITY_EXPAND_CAP1024Fast-lane queue cap (2048-slot membership index).
NO_PROGRESS_PASS_LIMIT30Passes without frontier growth before the TTL abandon.
CANARY_NO_PROGRESS_PASS_LIMIT3 — marked TEMP: was 30Base of the canary stall limit, and the emergency threshold.
MAX_CANARY_STUCK_BACKOFF_SHIFT8Escalation cap — up to 256× the base.
MARKER_TAG_BASE-(1<<62)Canary marker tag namespace.
LEAK_TAG_BASE / POOL_SIZE0x40000000 / 256LivenessTracker leak-tag pool.
MAX_DISCOVERED_INSTANCES_PER_CLASS8Auto-recorded instances per watched class.
INITIAL_TABLE_CAPACITY1024Frontier table start size; doubles to the configured cap.
+
+ +
+ +
+ Generated from source on branch jb/reference-chains-pi. Line references were accurate at extraction time; + re-verify against the working tree before quoting them. +
+ + + + + diff --git a/doc/architecture/LiveHeapReferenceChains-Implementation.md b/doc/architecture/LiveHeapReferenceChains-Implementation.md index b016cd10d3..73658f8ae3 100644 --- a/doc/architecture/LiveHeapReferenceChains-Implementation.md +++ b/doc/architecture/LiveHeapReferenceChains-Implementation.md @@ -1,6 +1,6 @@ # Live Heap Reference Chains — As-Built Implementation Reference -**Status:** matches `jb/reference-chains` as of 2026-09 +**Status:** Implemented as of 2026-09 **Jira:** [PROF-15341](https://datadoghq.atlassian.net/browse/PROF-15341) This document is the detailed, as-built reference for the reference-chain walk diff --git a/doc/architecture/LiveHeapReferenceChains.md b/doc/architecture/LiveHeapReferenceChains.md index 5276da3f6b..dea8f10bb8 100644 --- a/doc/architecture/LiveHeapReferenceChains.md +++ b/doc/architecture/LiveHeapReferenceChains.md @@ -6,9 +6,8 @@ ## Implementation status -The "Chosen design" section below has been implemented following -`LiveHeapReferenceChains-ImplementationPlan.md` (kept locally, not committed) -(Phases 0-7). It is off by default; the shipping switch is the `referencechains` argument +The "Chosen design" section below has been implemented. It is off by default; the +shipping switch is the `referencechains` argument parsed by `Arguments` (`arguments.cpp`'s `CASE("referencechains")`), e.g. `referencechains=true:hops=64:budget=2000:ttl=60000:framecap=65536`. @@ -60,10 +59,10 @@ Read this status note alongside the actual code before relying on it, not instea trend, then asserts on the `datadog.ReferenceChain` event `pollWatchedTargets()` produces - a real end-to-end exercise of this whole mechanism against a live JVM, not a synthetic frontier fixture. -- **Phase 5's tuning defaults are provisional, not empirically finalized.** The hop cap, +- **The tuning defaults are provisional, not empirically finalized.** The hop cap, per-pass budget, TTL, and frontier-size cap (`arguments.h`'s `DEFAULT_REFERENCE_CHAINS_*` constants) are explicitly-labeled placeholders; no benchmark against this codebase has - run yet (see Open Question 2 below and the implementation plan's Phase 5). + run yet (see Open Question 2 below). ## Goal @@ -390,8 +389,8 @@ retaining path, not a claim about the object's current exact retention state. representative heap shapes and per-hop fan-out, not a guess; `LivenessTracker`'s flat-sample-rate sizing formula does not transfer to a graph-search frontier. **Not resolved — provisional defaults only, no measurement has occurred.** The - implementation currently ships explicitly-labeled "provisional default pending Phase 5 - empirical tuning" constants (`arguments.h`: `DEFAULT_REFERENCE_CHAINS_HOP_CAP = 200`, + implementation currently ships explicitly-labeled "provisional default pending empirical + tuning" constants (`arguments.h`: `DEFAULT_REFERENCE_CHAINS_HOP_CAP = 200`, citing this doc's own JFR ~200-hop/100-100 precedent; `DEFAULT_REFERENCE_CHAINS_BUDGET = 1000`; `DEFAULT_REFERENCE_CHAINS_TTL_MS = 60000`; `DEFAULT_REFERENCE_CHAINS_FRONTIER_CAP = 65536`, sized as a fraction of `LivenessTracker::MAX_TRACKING_TABLE_SIZE` rather than @@ -399,15 +398,12 @@ retaining path, not a claim about the object's current exact retention state. `FrontierTable::INITIAL_TABLE_CAPACITY = 1024` and `ReferenceChainTracker::PASS_CADENCE_NS` = 1 s). These let the subsystem run and be tested end-to-end, but none are backed by a benchmark against this codebase — do not - describe them as measured. The real resolution path is - `LiveHeapReferenceChains-BenchmarkPlan.md` (kept locally, not committed), - which specifies the JMH/async-profiler matrix and decision rule Phase 5 still needs to - execute; this question stays open until that plan is actually run. + describe them as measured. The real resolution path is a JMH/async-profiler benchmark + matrix with an explicit decision rule, still to be executed; this question stays open + until that benchmark work is actually run. **Pause-time-SLO feedback loop — SHIPPED, reusing the existing `PidController`.** - Implemented in - `LiveHeapReferenceChains-RemainingWorkPlan.md` (kept locally, not committed)'s - Phase D (`ReferenceChainTracker::updatePacing()`, `referenceChains.cpp`). This does not + Implemented as `ReferenceChainTracker::updatePacing()` (`referenceChains.cpp`). This does not replace the hop/TTL/frontier-cap constants raised in the first half of this question — only the per-pass edge-count budget and the pass cadence, per the shipped mechanism below. - New config sub-option `referencechains=...:pausetarget=` (`arguments.cpp`'s @@ -427,9 +423,8 @@ retaining path, not a claim about the object's current exact retention state. single/low-double-digit in magnitude, unlike the shared triple's event-count scale (`referenceChains.cpp`'s `start()`, inline comment on each gain). Gain *convergence* is verified by gtest (three `ReferenceChainsTest` cases: steady-state at the ceiling, over- - ceiling, under-ceiling — see Phase D's exit criteria below), not by a live benchmark - against representative heap shapes; that remains a - `LiveHeapReferenceChains-BenchmarkPlan.md` (kept locally, not committed) item, + ceiling, under-ceiling), not by a live benchmark + against representative heap shapes; that benchmark work remains open, not fully closed by this mechanism landing. - Measurement point: `runPass()` (`referenceChains.cpp`) times its own root `IterateOverReachableObjects` call (first pass) or `expandFrontier()`'s @@ -454,10 +449,10 @@ retaining path, not a claim about the object's current exact retention state. - The hop cap and the frontier-size hard cap (Termination section) are untouched by this mechanism — they stay fixed correctness/memory-safety bounds, not controller-tuned, exactly as this question originally specified. - - One known, deliberate scope limit carried over from the plan: `buildAbandonedEvent()`'s + - One known, deliberate scope limit: `buildAbandonedEvent()`'s `datadog.ReferenceChainAbandoned` event still reports the static config ceiling `_budget`, - not the adaptive `_effective_budget` — changing that event's semantics was out of Phase D's - stated scope. + not the adaptive `_effective_budget` — making that event track the adaptive value was + deliberately left out of scope. 3. Decide the sample-batching policy: one incremental search per live-heap sample, or batched multi-target BFS sharing a single frontier walk (batching amortizes better but couples unrelated samples' termination conditions together). @@ -468,15 +463,13 @@ retaining path, not a claim about the object's current exact retention state. chain for a specific tag is a separate, read-only step (`buildChainEvent(target_tag, ...)`) applied after (or during) that one shared search - closer in spirit to "batched" (one frontier walk can answer for many targets) than "one search per sample", but arrived at - by omission (the target-sample feed did not exist at that time, see the implementation - plan's Phase 7 report) rather than a deliberate batching design. + by omission (no target-sample feed existed) rather than a deliberate batching design. Whether this generalizes to true multi-target batching (explicit seeding from multiple samples, coordinated termination) is still open and deferred, consistent with this question's original framing. - **Target-selection policy — SHIPPED (positive population-slope ranking).** Implemented in - `LiveHeapReferenceChains-RemainingWorkPlan.md` (kept locally, not committed)'s - Phases A-C. The missing piece above was *which* tag(s) `buildChainEvent()` should + **Target-selection policy — SHIPPED (positive population-slope ranking).** + The missing piece above was *which* tag(s) `buildChainEvent()` should reconstruct for. As shipped: per klass, `LivenessTracker` tracks a rolling window of its live tracked-instance population count, sampled once per `LivenessTracker::cleanup_table()` epoch advance (the same GC-epoch cadence that already recomputes survivor status, @@ -539,9 +532,8 @@ retaining path, not a claim about the object's current exact retention state. trend) but the *reconstruction target* does not need to be the exact instance that built up the trend — any currently-live tracked instance of the flagged klass is evidence of the same leak. **The bridging step is a READ, not a `SetTag` write** (a correction to - this doc's original proposal, found while grounding - `LiveHeapReferenceChains-RemainingWorkPlan.md` (kept locally, not committed); - see its "Correction to the design doc's Open Question 3 mechanism"). Pre-`SetTag`ing a + this doc's original proposal, found while grounding the implementation): + Pre-`SetTag`ing a candidate before the forward walk reached it would make `heapReferenceCallback()`'s `*tag_ptr == 0` branch — the *only* branch that records `parent_tag`/`depth` — skip it, yielding an empty/root chain. Instead `ReferenceChainTracker::pollWatchedTargets()` @@ -583,13 +575,12 @@ retaining path, not a claim about the object's current exact retention state. labeled provisional in `referenceChains.h`) has elapsed, whichever comes first. No safepoints-per-second/per-pause-duration measurement backs the 1-second cadence value - it was chosen only so an idle search still makes progress without polling tightly. The - cost model this question actually asks for is still open, deferred to Phase 5's - benchmark plan (`LiveHeapReferenceChains-BenchmarkPlan.md` (kept locally, not committed)), - which has not been run. + cost model this question actually asks for is still open, deferred to a benchmark + run that has not happened yet. **SHIPPED — folded into Open Question 2's pause-time-SLO feedback loop, not solved separately.** Implemented in the same `ReferenceChainTracker::updatePacing()` - (`referenceChains.cpp`, Phase D) described under Open Question 2: one `PidController` + (`referenceChains.cpp`) described under Open Question 2: one `PidController` `compute()` call per pass drives both that question's budget adjustment and this question's cadence adjustment from the single measured per-pass safepoint duration, rather than two independently-tuned mechanisms. `shouldRunPass()` and `threadLoop()` now compare against @@ -606,5 +597,4 @@ retaining path, not a claim about the object's current exact retention state. since it is the same controller instance. The cost-modeled "how many safepoints/sec is acceptable" question this Open Question originally asked for is answered structurally (the controller widens cadence exactly when passes are running long relative to the configured - ceiling) rather than by a specific measured number — that number is still a - `LiveHeapReferenceChains-BenchmarkPlan.md` (kept locally, not committed) item. + ceiling) rather than by a specific measured number — that benchmark work is still open.