Skip to content

feat(analysis): bind Rubin loading uncertainty to an analysis-run profile - #374

Draft
seonghobae wants to merge 9 commits into
mainfrom
feat/rubin-loading-uncertainty-analysis-run-gap-006
Draft

seonghobae wants to merge 9 commits into
mainfrom
feat/rubin-loading-uncertainty-analysis-run-gap-006

Conversation

@seonghobae

@seonghobae seonghobae commented Aug 31, 2026 •

Copy link
Copy Markdown
Contributor

Consolidation status

rubin_loading_uncertainty_v1 / tepp.rubin_loading_uncertainty.v1 is a Draft Analysis Run composition over protected-main psychometric_core::{recover_loading_point_estimate_mean, combine_draw_level_ols_loadings}. It is not a second estimator, Mislevy person-level plausible-value implementation, ESEM/DSEM sampler, CWC profile, or independent protected-main landing authority.

Current #374 head is 7fb76d1d5b338ecfb652a11a136c984c092ada15 on protected base main@a243f18da4a4ca8a8d068c39922537f1f8ed6ad0. The surviving Analysis Run vehicle is #416@03f8de2ed0a0fb842d2022d411814e440df7cfb4. Preserve every valid source/test/ADR/doctoring/scientific-evidence delta through ordinary non-force conflict resolution. Simple Close is not consolidation.

Scientific / temporal repair

The branch repairs real predecessor defects: per-row immutable snapshot_id + AvailableTime; instant-based KnowledgeCutoff binding; cross-snapshot fail-closed and same-snapshot future censoring; 256-draw and 1,000,000 admitted-cell application resource envelopes; reachable artifact count validation; exact binary64 Rubin-total integrity; symmetric 256 KiB artifact wire admission; protected-main robust point-estimate invocation; fallible digest propagation; and terminal validation_status = "validated" separated from artifact inference.

Repair lineage remains 152cb391... RED → dd46f8eb... causal repair → 8f0e7433... contract migration → ADR 0034 Proposed at a4a0ab01... → doctoring b3094612... → dependency cleanup 7fb76d1d.... Historical replay requires same-snapshot rows with AvailableTime > KnowledgeCutoff to leave the earlier scientific matrix/result unchanged; cross-snapshot rows are provenance violations, not censored history. The 256-draw and one-million-cell ceilings are resource limits, not psychometric recommendations.

Scientific acceptance and activation gaps

Issue #503 owns recovery/coverage acceptance. Draft child #504 is now exact 3d153b382e2382efb532f49bb1e2a8b93be0b73f. It provides eight 512-replication known-truth scenarios and a separate 512-replication leakage-safe rolling-origin study, with explicit attempted/recovered/failed accounting, robust point-estimate bias/RMSE and Monte Carlo uncertainty, Rubin Ubar/B/T, and a scoped large-sample normal coverage diagnostic centered on Qbar.

#504 also repaired its own provenance semantics. RED 16e3c151d6cadd79c1ec8e32d74a8401333be42c demonstrates that different generated Monte Carlo populations were initially aliased to one immutable snapshot/run identity. Repair ed2ac416ac42cb15909cefde0581af408de652e7 gives every generated population a distinct snapshot and every executor call a distinct run/idempotency identity. Rolling-origin views share one replication-specific snapshot because they are historical views of one population, but use distinct run receipts. CodeRabbit's valid scenario-description finding was repaired at ba90324ccb072ae24dcc1f7f54cbc579d219fe4e: the 8 scenarios are n × m × paired(lambda,sigma), not a 16-cell four-factor cross.

Issue #505 owns a separate Validation Evidence / Projection-policy gap. The product can currently combine arbitrary finite complete-data draw matrices after temporal/resource admission but does not bind a versioned draw-generator/imputer contract or validation-evidence identity for the generator/analysis pairing. Correct Qbar/Ubar/B/T arithmetic therefore must not be promoted into generally validated inferential semantics for arbitrary supplied draws. #504's ADR 0034 at 3d153b... records this fail-closed activation boundary under ADR 0014 and scopes the checked Monte Carlo evidence to its declared generator/design.

Do not describe #374/#504 as scientifically accepted or release-ready while #503 lacks exact-head completion, and do not treat #503/#504 as closing #505. LLM judgment and synthetic unit fixtures cannot substitute for claim-specific validation evidence.

Documentation / overlap authority

ADR 0034 remains Proposed; the newer scientific-evidence child carries current #503/#505 and replicate-identity wording. Shared docs/adr/README.md, docs/TRACEABILITY.md, docs/product-technical-gap-baseline.md, and CHANGELOG current-state reconciliation remain #435 / surviving-lane single-writer work. This branch also contains unrelated psychometric_core/src/error.rs MANIFESTVARstd Display-coverage overlap independently owned by #353; exclude that overlap from this scientific fold rather than copying it.

Review / fold gate

Live ruleset 18156473 requires one qualifying current-head approval, stale-review dismissal after pushes, all review threads resolved, central required workflows, and non-fast-forward protection. Checks/reviews from #374 do not transfer to #504 or the final fold head.

Before successor-based Close, verify on #416/successor exact head that the temporal provenance, resource bounds, Rubin-total integrity, robust point-estimate owner call, provider/domain status separation, bounded wire/digest path, #503 scientific evidence including per-replication snapshot/run identity, #505 activation boundary, ADR/doctoring, TRACEABILITY, and every other valid delta survived ordinary conflict resolution. Then reacquire exact-head format/lint/test/rustdoc, authored line+branch coverage, security/SAST/CodeQL/documentation gates, review resolution, and qualifying independent approval.

No self-approval, force push, destructive rebase, skip/xfail, coverage-denominator manipulation, scanner suppression, no-op rerun, mutable dependency consumption, provider-routing copy, or predecessor evidence transfer is authorized.

…file

Operators still cannot request the already-merged psychometric_core
draw-mean OLS loadings and Rubin T combination as a digest-bound
analysis-run output. Bind them jointly as rubin_loading_uncertainty_v1
/ tepp.rubin_loading_uncertainty.v1 (ADR 0034). Cutoff-filter
observations, refuse raw proportions and single-draw inputs, and keep
the claim boundary as complete-data OLS combination rather than
Mislevy person-level plausible values.

Not a new ESEM/DSEM estimator, not CWC, not a Driver p.16 std restore,
and not persistence.
@coderabbitai

coderabbitai Bot commented Aug 31, 2026 •

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@devin-ai-integration devin-ai-integration Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note

This report is out of date. Scroll down for Devin Review's latest report on this PR.

Devin Review found 3 potential issues.

Devin Review

Comment thread crates/analysis_engine/src/rubin_loading_artifact.rs
Comment thread crates/analysis_engine/tests/rubin_loading_execution_contract.rs
Comment thread crates/analysis_engine/src/rubin_loading_artifact.rs Outdated

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 5 new potential issues.

Devin Review

Comment thread crates/analysis_engine/src/rubin_loading_artifact.rs Outdated
Comment thread crates/analysis_engine/src/rubin_loading_artifact.rs Outdated
Comment thread crates/analysis_engine/src/rubin_loading_artifact.rs Outdated
Comment thread crates/analysis_engine/src/rubin_loading_artifact.rs
Comment thread crates/analysis_engine/src/rubin_loading_artifact.rs
…p-006

Resolve the CHANGELOG.md append conflict by keeping both entries.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@seonghobae

Copy link
Copy Markdown
Contributor Author

Restack on protected main (a243f18)

Non-force merge of origin/main (merge commit ec2bf9c4); the only conflict was the CHANGELOG.md append, both entries kept. ADR 0034 already carries an admitted maturity value (active-PR).

Local evidence on the pushed head (toolchain 1.98.0): cargo test -p analysis_engine -p psychometric_core 403 passed / 0 failed, clippy -D warnings clean on both crates, cargo fmt --all --check clean, documentation/workspace/docstring contracts PASS, git diff --check clean.

🤖 Generated with Claude Code

Copy link
Copy Markdown
Contributor Author

Scientific acceptance child #504 now carries issue #503 candidate evidence on exact head ac324d74911f15c840a5c1d1c419ef192ea7d0fa, stacked directly on this PR's 7fb76d1d5b338ecfb652a11a136c984c092ada15 head.

The child adds 8 × 512 repeated-sampling known-truth scenarios plus a 512-replication rolling-origin replay contract. It consumes validation_core for bias/RMSE/coverage/Monte-Carlo summaries, keeps robust point_estimate_mean as the recovery target, and centers the diagnostic interval on Rubin Qbar with T. #503 remains open until exact-head checks and review are GREEN.

Do not close #374 as superseded unless the surviving #416/fold head is verified to inherit both this PR's temporal/resource/artifact repairs and #504's scientific evidence/test/docs. Predecessor check receipts do not transfer.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant