Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
40 commits
Select commit Hold shift + click to select a range
d25df8e
fix(ci): fail closed to explicitly free orchestrator routes
seonghobae Sep 2, 2026
feb3399
test(ci): lock free-only discovered-agent admission
seonghobae Sep 2, 2026
a63c454
docs(ci): document free-only hourly route admission
seonghobae Sep 2, 2026
68b9c6c
docs(adr): bind hourly gateway to explicit free admission
seonghobae Sep 2, 2026
aef7181
test(docs): preserve scientific queue lineage markers
seonghobae Sep 2, 2026
26af18f
fix(ci): reject discovery rows without price evidence
seonghobae Sep 2, 2026
8594325
test(ci): reject models without price fields
seonghobae Sep 2, 2026
e3a5f56
test(ci): make legacy discovery fixtures explicitly free
seonghobae Sep 2, 2026
f8749d3
docs(ci): record explicit free hourly admission in Unreleased
seonghobae Sep 2, 2026
737c01f
docs(governance): route semantic LLM work through released orchestrator
seonghobae Sep 2, 2026
687d780
test(governance): reproduce stale direct-provider authority
seonghobae Sep 2, 2026
602a0ca
docs(llm): make released orchestrator the routing authority
seonghobae Sep 2, 2026
2567ecd
docs(prd): add released orchestrator routing amendment
seonghobae Sep 2, 2026
224664e
docs(prd): supersede direct provider execution authority
seonghobae Sep 2, 2026
5ed99cb
docs(trd): require released orchestrator contract for LLM work
seonghobae Sep 2, 2026
8fbfe6e
refactor(architecture): separate DDD contexts and released owner boun…
seonghobae Sep 2, 2026
d4b5f16
test(ci): reproduce unreleased provider bootstrap contract
seonghobae Sep 2, 2026
947df0b
fix(ci): delegate hourly routing to released orchestrator
seonghobae Sep 2, 2026
ee30948
test(ci): retire consumer-side free-route authority
seonghobae Sep 2, 2026
37a5000
refactor(ci): retire TEPP-owned provider routing bootstrap
seonghobae Sep 2, 2026
f03c050
test(ci): pin secure released-gateway and base-head contract
seonghobae Sep 2, 2026
152a8a3
docs(ci): align hourly runbook to released gateway
seonghobae Sep 2, 2026
b991fa4
docs(adr): move hourly routing to released owner contract
seonghobae Sep 2, 2026
639faec
docs(prd): fix released-routing evidence and whitespace
seonghobae Sep 2, 2026
484c509
docs(doctoring): trace released-orchestrator ownership repair
seonghobae Sep 2, 2026
17d4b8c
docs: remove LLM contract trailing whitespace
seonghobae Sep 2, 2026
bb178de
docs: preserve LLM status metadata layout
seonghobae Sep 2, 2026
33a68ec
docs: repair LLM and cutoff architecture authority
seonghobae Sep 2, 2026
9ff5261
docs: keep architecture free of volatile PR state
seonghobae Sep 2, 2026
6d756d0
test(ci): require HTTPS-only gateway redirects
seonghobae Sep 2, 2026
f1da3f2
fix(ci): restrict gateway redirects to HTTPS
seonghobae Sep 2, 2026
1f0d2dd
test(ci): require HTTPS-only OpenCode redirects
seonghobae Sep 2, 2026
4475542
fix(ci): keep OpenCode redirects on HTTPS
seonghobae Sep 2, 2026
4248b33
test(llm): pin contributing guide to released orchestrator
seonghobae Sep 2, 2026
01f45a9
fix(llm): remove direct provider credential guidance
seonghobae Sep 2, 2026
5af6795
fix(actions): centralize hourly development admission
seonghobae Sep 4, 2026
03876fb
Revert "fix(actions): centralize hourly development admission"
seonghobae Sep 4, 2026
5b2637f
fix(actions): retain centralized hourly admission
seonghobae Sep 4, 2026
e0dc955
Merge origin/main into fix/hourly-orchestrator-free-admission-20260902
seonghobae Sep 13, 2026
0dcd198
test(quality): expect central admission marker after main restack
seonghobae Sep 13, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
272 changes: 133 additions & 139 deletions .github/workflows/hourly-nim-product-development.yml

Large diffs are not rendered by default.

6 changes: 3 additions & 3 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,14 +15,14 @@ TEPP is the Temporal Event Psychometrics Platform: a multilingual, temporal, rel
7. Multilingual measurement uses one shared latent semantic space. Language-specific morphology and lexical emissions may vary, but equivalent meanings must be aligned and tested for measurement invariance.
8. Production line and branch coverage are 100%. All public modules, traits, structs, enums, functions, methods, error variants, configuration fields, and safety contracts have complete docstrings.
9. Scientific acceptance requires realistic synthetic truth: parameter recovery, RMSE, bias, interval coverage, temporal ordering, graph recovery, invariance, and CPU/GPU parity. Skipped or ignored GPU tests are not evidence.
10. LLM live tests use `NVIDIA_NIM_API_KEY`. `COPILOT_GITHUB_TOKEN` is prohibited. Existing independent review-agent credentials must not be repurposed.
11. LLM orchestration allocates test-time computation between direct routing and deeper multi-agent workflows. Workflow depth, decomposition, access lists, recursion, role-specific reasoning effort, verification/adjudication, and comparable-budget ablations are recorded. LLM output never replaces deterministic/statistical scientific authority.
10. Every semantic LLM operation and every model-backed GitHub Actions workflow goes through a released, versioned `contextual-orchestrator` contract. Actions use the `orchestrator/free` route through the gateway credential only; TEPP must not select a provider/model/group, declare a paid fallback, call providers directly, or consume provider API keys such as NVIDIA NIM, OpenRouter, OpenAI, or Bytez credentials. If the released orchestrator contract cannot supply the required capability, fail closed and repair the canonical owner before consumer adoption. `COPILOT_GITHUB_TOKEN` is prohibited. Independent review-agent credentials must not be repurposed as execution credentials.
Comment thread
devin-ai-integration[bot] marked this conversation as resolved.
Comment thread
coderabbitai[bot] marked this conversation as resolved.
Comment thread
seonghobae marked this conversation as resolved.
11. LLM orchestration allocates test-time computation between direct routing and deeper multi-agent workflows. Workflow depth, decomposition, access lists, recursion, role-specific reasoning effort, verification/adjudication, and comparable-budget ablations are recorded. LLM output never replaces deterministic/statistical scientific authority. Model timeout defaults must not terminate reasoning/stream/tool-call work merely because elapsed time is long; user cancellation, provider termination, and explicit administrative limits remain distinct outcomes.
Comment thread
devin-ai-integration[bot] marked this conversation as resolved.
12. Database object names contain at least two words and use `snake_case` by default. CamelCase or PascalCase is permitted only where language conventions require it.
13. Every scientific or standards claim is traced to an authoritative primary source and cited in APA 7th style in `docs/research/`.
14. Changes that alter latent-variable meaning, temporal semantics, event ontology, multilingual invariance, estimator targets, privacy authority, or service authority require an ADR and a PRD version change when the approved product/measurement target changes.
15. Do not blanket-mask PII when doing so destroys valid authorship, temporal, longitudinal, event, entity-role, or multiple-membership measurement. Use purpose-bound authorization, opaque analytical identifiers, separately protected identity mapping, encryption, selective disclosure, retention/deletion, and auditable privileged access.
16. Design toward CSAP and SOC 2 evidence readiness and align AI governance with current published ISO/NIST guidance where applicable, but never claim certification, attestation, conformance, or legal sufficiency without external evidence.
17. Preserve standalone operation and modular MSA composition. `naruon`, `contextual-orchestrator`, and other CWL services integrate through versioned APIs/artifacts; no direct cross-service application-table access is permitted.
17. Preserve standalone operation and modular MSA composition. `naruon`, `contextual-orchestrator`, and other CWL services integrate only through released/versioned APIs or immutable artifacts plus explicit ACLs; no direct cross-service application-table access or mutable sibling-head dependency is permitted.
18. Documents, external metadata, serialized payloads, model checkpoints, and LLM outputs are untrusted until their owning boundary validates identity, provenance, size/depth, authorization, and scientific semantics.
19. Figma/Product Design becomes authoritative only for a stable product interaction contract; UI design never overrides the PRD, data model, numerical/scientific contract, or protected-main implementation truth.

Expand Down
322 changes: 134 additions & 188 deletions ARCHITECTURE.md

Large diffs are not rendered by default.

1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,7 @@ All notable changes to TEPP are documented here. The format follows Keep a Chang

## [Unreleased]

- Hourly contextual-orchestrator bootstrap now admits only discovery rows whose provider-reported prompt and completion token prices are both present and exactly `0.0` before cheapest ranking. Paid, partial, missing, and fully unpriced rows stay out of the general-chat pool; an empty explicitly-free pool fails closed with no paid fallback. Secret-free discovery evidence records admitted free candidates and excluded non-free candidates. ADR 0017 and the hourly runbook document the same boundary. The checksum-pinned contextual-orchestrator revision is unchanged; unknown price is not treated as free.
- Removed the repository-local hourly PR-maintenance caller now covered by the central required scheduler, retired stale workflow registrations, narrowed documentation triggers, keyed PR concurrency by fixed workflow name, repository, and pull-request number without cancelling non-PR runs, and combined line/branch coverage on one sequential runner while preserving both 100% gates and diagnostics.

- `event_core` adds bounded Allen interval-consistency classification, atomic path-consistency closure, contradiction/resource refusals, and an explicit dependency-error fallback without claiming unrestricted global satisfiability.
Expand Down
9 changes: 5 additions & 4 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,10 +45,11 @@ Use primary papers, international standards, official specifications, and offici

- Treat all model output as untrusted structured input.
- Preserve exact source spans and evidence identifiers.
- Use `NVIDIA_NIM_API_KEY` for approved live tests.
- Never use or introduce `COPILOT_GITHUB_TOKEN`.
- Record provider, model, prompt hash, reasoning effort, workflow depth, tools/access list, seed where supported, latency, token usage, and cost.
- Include direct-routing versus orchestrated and reasoning-effort ablations.
- Route every semantic LLM operation and model-backed GitHub Actions workflow through a released, versioned `contextual-orchestrator` contract. GitHub Actions use only `orchestrator/free` through the contextual-orchestrator gateway credential.
- Do not select or hard-code a provider, model, provider group, or paid fallback in TEPP, and do not expose provider API keys to TEPP workflows. If a released orchestrator contract cannot provide the required capability, fail closed and repair the canonical owner before adopting the change here.
- Never use or introduce `COPILOT_GITHUB_TOKEN`, and never repurpose independent review-agent credentials as execution credentials.
- Record the contextual-orchestrator release/contract identity, route, prompt hash, reasoning effort, workflow depth, tools/access list, seed where supported, latency, token usage, and cost. Record provider/model identity only when the released orchestrator returns it as execution provenance; it is evidence, not TEPP routing authority.
- Include direct-routing versus orchestrated and reasoning-effort ablations where scientifically relevant. LLM output never replaces numerical estimation or scientific acceptance.

## Database naming

Expand Down
52 changes: 32 additions & 20 deletions docs/LLM_ORCHESTRATION.md
Original file line number Diff line number Diff line change
@@ -1,13 +1,14 @@
# TEPP LLM Orchestration and Test-Time Compute Contract

**Status:** Partial — `tepp_api::route_orchestration` is the governed selector; live provider execution is not yet shipped.
**Last reviewed:** 2026-08-13
**Status:** Partial — TEPP owns semantic-task policy/evidence contracts; provider routing requires a released `contextual-orchestrator` contract and is not yet deployable from the current repository state.
Comment thread
seonghobae marked this conversation as resolved.

**Last reviewed:** 2026-09-02

## 1. Purpose

TEPP uses LLMs only for bounded semantic unitization, candidate-model review, evidence-grounded interpretation, and independent claim verification. Statistical estimation, temporal eligibility, event-relation validity, measurement invariance, numerical acceptance, and release authority remain deterministic/Rust or governed human authority.

The product must allocate test-time compute adaptively rather than assume that either one frontier model or a large fixed multi-agent graph is always best.
The product may allocate test-time compute adaptively rather than assume that either one frontier model or a large fixed multi-agent graph is always best. TEPP decides the semantic task, evidence/access policy, scientific risk, and admissible orchestration mode; provider/model/group routing is a `contextual-orchestrator` responsibility.

## 2. Research basis

Expand All @@ -21,21 +22,22 @@ These results motivate experiments; they do not prove that deeper orchestration

## 3. Orchestration modes

| Mode | Typical TEPP use | Default compute |
| Mode | Typical TEPP use | TEPP policy request |
|---|---|---|
| direct | simple span classification, deterministic-schema fill, low-ambiguity label | one model call |
| direct | simple span classification, deterministic-schema fill, low-ambiguity label | one governed semantic call |
| verify | interpretation or classification with material unsupported-claim risk | producer + independent verifier |
| committee | K/model interpretation with scientific ambiguity | blinded parallel raters + adjudication |
| conductor | complex evidence synthesis or multi-stage semantic reasoning | adaptive roles/topology under explicit budget |
| abstain | provider/evidence/validation insufficient | no forced answer |
| abstain | evidence/contract/capability insufficient | no forced answer |

`tepp_api::route_orchestration` is the governed selector. It chooses the cheapest mode expected to satisfy the quality/risk profile, but latency is not the primary objective. Quality, evidence support, calibration, disagreement, controllability, and reproducibility dominate. The returned plan is a proposal: `scientific_authority_code` remains `deterministic_statistical_gates`.
`tepp_api::route_orchestration` governs only the TEPP-side task/mode proposal. It does not select a provider, provider group, concrete model, or paid fallback. Provider-neutral routing and execution are delegated through a released `contextual-orchestrator` API/client/schema. Quality, evidence support, calibration, disagreement, controllability, and reproducibility dominate; `scientific_authority_code` remains `deterministic_statistical_gates`.

## 4. Explicit experimental variables

Every orchestration benchmark records and can ablate:

- model/provider pool;
- the released orchestrator contract/version and routing receipt;
- orchestrator-selected provider/model identity as observed provenance, not TEPP routing configuration;
- direct versus multi-agent topology;
- workflow stage count;
- worker count;
Expand All @@ -47,9 +49,9 @@ Every orchestration benchmark records and can ablate:
- total token/call/compute budget;
- verification/adjudication policy;
- stopping rule;
- provider failure/fallback behavior.
- failure/abstention outcome.

Comparisons must use approximately comparable budgets or report the budget difference explicitly.
Comparisons must use approximately comparable budgets or report the budget difference explicitly. Fugu/Conductor/TRINITY experiments are ablations over a released orchestration boundary; they do not authorize branch-local provider routing.

## 5. Role-specific effort

Expand All @@ -66,11 +68,13 @@ The policy is empirically calibrated and versioned rather than hard-coded as a p

## 6. Evidence and trust boundary

LLM calls receive only the minimum evidence bundle needed for the assigned role. Documents are untrusted observations and cannot alter orchestration policy, tools, credentials, model pool, access lists, or scientific gates.
LLM calls receive only the minimum evidence bundle needed for the assigned role. Documents are untrusted observations and cannot alter orchestration policy, tools, credentials, routing, access lists, or scientific gates.

Each call records:

```text
orchestrator_contract_version
orchestrator_receipt_id
provider_id
model_id
model_revision_or_endpoint
Expand All @@ -90,7 +94,7 @@ duration_record
verdict_status
```

Raw credentials are never model-visible. Model outputs are proposals, not source facts.
Provider/model fields are returned provenance. They are not TEPP provider-selection inputs. Raw provider credentials are never TEPP model-visible or required by TEPP semantic clients. Model outputs are proposals, not source facts.

## 7. Quality metrics

Expand All @@ -105,33 +109,41 @@ At minimum evaluate:
- prompt-injection success rate;
- abstention quality;
- token/call/compute cost;
- provider/model failure resilience;
- orchestrator failure resilience;
- repeated-run variance.

For model-selection review, statistical predictive/recovery/stability/invariance gates run before LLM judgment. LLM preference cannot rescue a statistically rejected candidate.

## 8. contextual-orchestrator boundary

`contextual-orchestrator` is the preferred CWL provider-neutral integration when available. TEPP owns the evidence bundle, statistical/model-selection policy, scientific acceptance, artifact provenance, and allowed role/access configuration. The orchestrator owns provider routing/orchestration execution within the supplied policy. Neither service reads the other's application database directly.
`contextual-orchestrator` is the canonical CWL owner of provider discovery, key auto-discovery, provider/model/group routing, paid/free policy, request-family adaptation, fallback, streaming/tool-call lifecycle, and provider execution. TEPP owns evidence bundles, semantic-task policy, statistical/model-selection policy, scientific acceptance, artifact provenance, and allowed role/access configuration. Neither service reads the other's application database directly.

Production TEPP integration consumes only a **released, versioned** contextual-orchestrator API/client/schema through an ACL with immutable artifact identity and provenance. A mutable protected-main commit, open PR head, or checksum-pinned source snapshot without an immutable release is candidate evidence, not production dependency authority. At the 2026-09-02 review, contextual-orchestrator has no GitHub release, so production semantic execution through this boundary remains fail-closed until a compatible release exists and is verified/adopted.

`orchestrator_live` exposes a loopback-only `POST /v1/interpretation-runs` listener that records the selected mode and budget. Listener output is hypothetical and cannot become scientific authority. Production TLS and provider execution remain outside this crate.
`orchestrator_live` may expose a local adapter/listener for contract tests and hypothetical planning, but it is not a second provider router and cannot become scientific authority. Production provider execution stays behind the released contextual-orchestrator boundary.

## 9. Development and live-test credentials

Live model tests use GitHub Secret `NVIDIA_NIM_API_KEY` with the minimum runtime mapping required by the selected adapter. `COPILOT_GITHUB_TOKEN` is prohibited. Existing independent review-agent credentials are separate and must not be renamed, copied, or repurposed for TEPP model execution.
TEPP semantic clients use only the contextual-orchestrator gateway credential appropriate to the released contract. Model-backed GitHub Actions request `orchestrator/free`; they do not receive or select provider credentials, providers, models, provider groups, or paid fallbacks. `COPILOT_GITHUB_TOKEN` is prohibited. Independent review-agent credentials remain separate and cannot be renamed, copied, or repurposed for model execution.

If the released orchestrator cannot expose a required capability, TEPP fails closed and the missing contract/capability is repaired at the contextual-orchestrator owner before consumer adoption. TEPP does not compensate by importing a provider SDK or secret.

## 10. Acceptance, timeout, and failure semantics

A workflow fails closed or abstains when required evidence is missing, the released orchestration contract is unavailable, results violate schema, model disagreement exceeds policy, injection tests trigger, or verifier support is insufficient. Provider outage/routing recovery is owned by contextual-orchestrator; TEPP receives the governed result or a typed unresolved/failure outcome and never silently chooses a replacement provider or paid route.

## 10. Acceptance and fallback
Reasoning, streaming, and tool-call work is not terminated merely because an arbitrary elapsed-time default expires. User cancellation, provider-declared termination, and explicit administrative timeout are separate typed outcomes. Long-running OpenCode/Strix/Noema-style work must remain possible when the released contract supports it.

A workflow fails closed or abstains when required evidence is missing, provider results violate schema, model disagreement exceeds policy, injection tests trigger, or verifier support is insufficient. Provider outage may route to an allowed alternative, return deferred/unresolved state, or use deterministic fallback where scientifically valid. It never silently changes the estimand or fabricates semantic evidence.
No LLM path may silently change an estimand, perform numerical scientific acceptance, authoritatively activate a candidate, or fabricate semantic evidence.

## 11. Required ablation before production claim

Before claiming an orchestration mode materially improves TEPP, compare at least:

1. strongest approved single-model direct baseline;
1. strongest approved direct baseline exposed by the released orchestrator;
2. direct + verifier;
3. fixed role-based multi-agent workflow;
4. adaptive/learned-conductor-style workflow where available;
5. at least two reasoning-effort/budget settings.

Report uncertainty and failure modes, not only the best benchmark score.
Report uncertainty, routing receipts, contract version, budget, and failure modes, not only the best benchmark score.
Loading
Loading