Skip to content

Commit 5eeade5

Browse files
Merge pull request #146 from CodeWithJuber/fix/hermetic-test-suite
feat: TypeSafe System One (Jev) typed proposers + hermetic test suite
2 parents 2e98d86 + 418b9c7 commit 5eeade5

17 files changed

Lines changed: 755 additions & 20 deletions

‎.github/workflows/ci.yml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -48,7 +48,7 @@ jobs:
4848
shell: bash
4949
run: |
5050
set +e
51-
node --test test/*.test.js 2>&1 | tee /tmp/win-test.log
51+
node --test --import ./test/_setup.js test/*.test.js 2>&1 | tee /tmp/win-test.log
5252
ec=${PIPESTATUS[0]}
5353
echo "===== FAILING TESTS (name + file) ====="
5454
grep -nE '^not ok ' /tmp/win-test.log | head -80

‎ARCHITECTURE.md‎

Lines changed: 19 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -280,6 +280,22 @@ fails safe to the stock ID on no gateway / unreachable `/v1/models` / no family
280280
resolved `tier→model` mapping for verification. The `MODELS` export shape is unchanged: this is a
281281
resolution-time layer, not a table edit.
282282

283+
**Typed proposers via TypeSafe System One (`src/jev.js`).** Two of the substrate's proposer
284+
judgments are not text-generation tasks at all: `route`'s complexity band is a classification
285+
(cheap/mid/premium), and preflight's assumption gate is four independent yes/no readings (one
286+
per rubric dimension). When `TYPESAFE_API_KEY` is set (same `FORGE_LLM=1` opt-in), those two
287+
faculties ask Jev instead of a text model — one batched `POST /v1/systemone` returning typed
288+
`choice`/`noul` answers with probability distributions and confidence in ~150ms, versus seconds
289+
of text plus JSON parsing. The module reuses the adjudicate contract verbatim: opt-in, fail-safe
290+
(null → text-LLM fallback → deterministic rubric; a null never moves a verdict), zero-dependency
291+
(the `llm.js` spawned-child pattern, key in child env as `_FORGE_JEV_KEY`), and secret-refusing
292+
on the outgoing state. Jev answers are validated against the questions asked — a choice naming
293+
an option we never offered is garble and fails safe. The reconciles are untouched: `BAND_FLOOR`
294+
still floors the routing band, the assumption gate still bounds completeness to ±band, and
295+
clarifying free-text questions stay with the deterministic rubric, because a System One model
296+
judges but does not author prose. Provenance records which proposer answered
297+
(`llm.provider: "jev"` in `forge route --json`, `assumption.provenance.provider` in preflight).
298+
283299
**Intent cards (`src/intent.js`).** Prompt → intent by the same exemplar k-NN math as
284300
model routing — a labeled bank (English + Hinglish rows) under overlap similarity with a
285301
confidence gate, NOT a keyword DFA. Note `intentGrams` ≠ `contentGrams`: route.js stops
@@ -581,18 +597,20 @@ from the tree it describes.
581597
```mermaid
582598
%%{init: {'theme':'base','themeVariables':{'primaryColor':'#201a15','primaryTextColor':'#f2ede7','primaryBorderColor':'#372c22','lineColor':'#f26430','secondaryColor':'#272019','tertiaryColor':'#171310','edgeLabelBackground':'#201a15','clusterBkg':'#171310','clusterBorder':'#4a3b2e','fontFamily':'ui-sans-serif, system-ui, sans-serif','fontSize':'14px'},'flowchart':{'curve':'basis','padding':10,'nodeSpacing':36,'rankSpacing':44}}}%%
583599
flowchart LR
600+
test["test<br/>105 files"]
584601
test["test<br/>106 files"]
585602
src["src<br/>97 files"]
586603
landing["landing<br/>61 files"]
587604
research["research<br/>35 files"]
588605
global["global<br/>3 files"]
589606
bench["bench<br/>2 files"]
590607
scripts["scripts<br/>2 files"]
608+
_remember[".remember<br/>1 file"]
591609
docs["docs<br/>1 file"]
610+
test -- 201 --> src
592611
examples["examples<br/>1 file"]
593612
test -- 206 --> src
594613
bench -- 7 --> src
595-
examples -- 4 --> src
596614
test -- 2 --> scripts
597615
scripts --> src
598616
src --> global

‎CHANGELOG.md‎

Lines changed: 37 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -8,6 +8,43 @@ to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
88

99
### Added
1010

11+
- **TypeSafe System One (Jev) as the fast proposer.** Where forge's LLM layer asked a text
12+
model for a judgment that is really a classification or a yes/no — `route`'s complexity band
13+
and preflight's assumption gate — it can now ask Jev instead: typed `choice`/`noul` answers
14+
with real probability distributions and confidence in ~150ms, rather than seconds of text
15+
generation followed by JSON parsing. The new `src/jev.js` client follows the existing
16+
proposer contract exactly: opt-in (`FORGE_LLM=1` plus `TYPESAFE_API_KEY`, overridable via
17+
`TYPESAFE_BASE_URL`), fail-safe (any error → null → text-LLM fallback → deterministic
18+
rubric, and a null never changes a verdict), zero-dependency (one raw HTTPS POST through
19+
the child-process-fetch pattern, the key travelling via child env — never argv, never
20+
logged), and secret-refusing on the way out. Routing keeps its `BAND_FLOOR` reconcile and
21+
gains `llm.provider: "jev"` plus confidence in `forge route --json`; the assumption gate
22+
scores all four rubric dimensions in one batched call (free-text clarifying questions stay
23+
with the deterministic rubric — a System One model judges, it does not author prose).
24+
`test/_setup.js` now scrubs `TYPESAFE_*` so the suite stays hermetic with the key exported.
25+
26+
### Fixed
27+
28+
- **The test suite is hermetic.** It inherited the developer's environment, so it was green
29+
in CI and red on any machine where forge was actually installed and enabled — the two
30+
things a maintainer does. An exported `FORGE_LLM=1` both flipped the "llm off by default"
31+
assertion in `test/substrate.test.js` and made the faculties fire real model calls, and a
32+
real `~/.forge` reached `doctor()`'s machine-scoped install check through
33+
`test/doctor.test.js`. Wall time was 593s with two failures. A new `test/_setup.js`,
34+
preloaded via `--import` into every test process, scrubs `FORGE_*`/provider env by prefix,
35+
sandboxes `$HOME` to a throwaway tmpdir, and forces the keyless HTTP runner instead of
36+
shelling out to a real `claude` binary: **0 failures in ~40s**. `test/hermetic.test.js`
37+
pins the scrub list against `envVarsRead()` so the two cannot drift, and fails loudly if
38+
anyone drops the `--import` wiring. Two assertions were wrong rather than merely leaky and
39+
were corrected: `doctor` asserted a global `failed === 0` to prove a local property about
40+
`na` rows, and a comment in `substrate` claimed no runner reaches the real CLI — the
41+
opposite of the truth, and the reason that file spent 85s on live calls.
42+
43+
### Documentation
44+
45+
- `CLAUDE.md`: Biome 2.5.2 → 2.5.5 (matching the pin), "600+ tests" → "1000+", and the lint
46+
command `npx biome check` → `npm run check` — the documented command fails outright, since
47+
the npx package is `@biomejs/biome`, not `biome`.
1148
- **OpenClaw is a first-class emit target — the compiler's tenth tool.** Instructions need
1249
no new file: OpenClaw appends the execution folder's `AGENTS.md` after its configured
1350
agent-workspace files as project context, so the canonical source reaches it the same way
@@ -198,7 +235,6 @@ to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
198235
every PreToolUse hook and recompiled the same trigger-glob RegExp each time; compiled
199236
globs are now cached in a module-level map bounded by the distinct globs in the
200237
lesson set.
201-
202238
## [0.27.4] - 2026-08-04
203239

204240
### Fixed

‎CLAUDE.md‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -3,14 +3,14 @@
33
## Stack
44

55
- Node.js >=20, pure ESM (`"type": "module"`), zero runtime dependencies.
6-
- Linter/formatter: Biome 2.5.2 (dev dependency).
6+
- Linter/formatter: Biome 2.5.5 (dev dependency).
77
- Types: TypeScript via JSDoc annotations — no `.ts` files, checked by `tsc`.
88

99
## Commands
1010

1111
- Install: `npm ci`
12-
- Test: `npm test` (node:test, 600+ tests)
13-
- Lint + format: `npx biome check` (or `npm run check`)
12+
- Test: `npm test` (node:test, 1000+ tests)
13+
- Lint + format: `npm run check` (the npx package is `@biomejs/biome`, not `biome`)
1414
- Typecheck: `npm run typecheck`
1515
- Build pages: `npm run pages:build`
1616

‎docs/GUIDE.md‎

Lines changed: 13 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1452,6 +1452,17 @@ exposes `llm.provenance` per faculty (`llm-cleared` / `llm-tightened` / `llm-rai
14521452
conservative tighten-/raise-only mode. Each faculty pairs a pure `*LLM` proposer with a
14531453
`reconcile` step — extend by adding both, never by trusting the model's answer directly.
14541454

1455+
**TypeSafe System One (Jev) is the preferred proposer when configured.** Where the judgment
1456+
is already a classification or a yes/no — `route`'s complexity band (a `choice` over
1457+
cheap/mid/premium) and preflight's assumption gate (one batched `noul` per rubric dimension) —
1458+
`src/jev.js` asks Jev instead of a text model: typed answers with real probability
1459+
distributions and confidence in ~150ms, rather than seconds of generation followed by JSON
1460+
parsing. Set `TYPESAFE_API_KEY` (plus the same `FORGE_LLM=1` opt-in) and the two proposers
1461+
prefer it automatically; the text-LLM runner remains the fallback on any failure, and the
1462+
deterministic rubrics still judge. Free-text clarifying questions stay with the rubric — a
1463+
System One model judges, it does not author prose. `forge route --json` shows which proposer
1464+
answered under `llm.provider` (`jev` / `text`) with Jev's confidence.
1465+
14551466
### Support a new tool
14561467

14571468
Add an emitter module in `src/emit/<tool>.js` (mirror an existing one like
@@ -1478,6 +1489,8 @@ code reads but this table misses fails CI on the forge repo):
14781489
| `OPENROUTER_API_KEY` | OpenRouter provider |
14791490
| `OPENAI_API_KEY` | OpenAI provider (OpenAI-compatible chat/completions); low-configuration auto-detect fallback after Anthropic |
14801491
| `GEMINI_API_KEY` / `GOOGLE_API_KEY` | Google Gemini provider via its OpenAI-compatible endpoint; low-configuration auto-detect fallback after Anthropic |
1492+
| `TYPESAFE_API_KEY` | TypeSafe System One (Jev) — with `FORGE_LLM=1`, the route/assumption proposers prefer typed ~150ms judgments over a text round-trip; unset = text-LLM proposer only |
1493+
| `TYPESAFE_BASE_URL` | override the Jev endpoint (default `https://api.typesafe.ai`) — staging/self-hosted |
14811494
| `FORGE_LLM` | `1` enables the LLM proposer layer (off = fully deterministic) |
14821495
| `FORGE_LLM_AMBIENT` | `1` lets the ambient hook use the proposer too |
14831496
| `FORGE_LLM_HTTP` | `1` forces direct HTTP (Anthropic Messages or OpenAI-compatible, per the resolved provider) instead of the `claude` CLI; automatic when the CLI is absent |

‎mintlify/concepts/model-routing.mdx‎

Lines changed: 16 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -24,6 +24,22 @@ under an overlap-similarity metric with a confidence gate — not a keyword look
2424
genuinely needs it.
2525
</Note>
2626

27+
## Optional typed proposer — TypeSafe Jev
28+
29+
The rubric is the judge; an optional **proposer** layer can refine it. With `FORGE_LLM=1`
30+
plus `TYPESAFE_API_KEY` set, `route` asks TypeSafe's System One model (Jev) for the
31+
complexity band as a typed `choice` with probabilities and confidence in ~150ms — instead
32+
of a multi-second text-LLM round-trip. The same applies to preflight's assumption gate,
33+
scored as one batched `noul` per rubric dimension.
34+
35+
<Note>
36+
Fail-safe by construction: any Jev error falls back to the text-LLM proposer, then to
37+
the deterministic rubric — and a miss never changes a verdict. Without the key, behavior
38+
is byte-identical. `forge route --json` shows which proposer answered under
39+
`llm.provider` (`jev` / `text`), with Jev's confidence. `TYPESAFE_BASE_URL` overrides
40+
the endpoint for staging or self-hosted deployments.
41+
</Note>
42+
2743
## Intent, then tier
2844

2945
Routing shares its math with intent detection (`src/intent.js`): a prompt maps to an

‎mintlify/concepts/pre-action-gate.mdx‎

Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -48,6 +48,15 @@ flowchart TD
4848
</Step>
4949
</Steps>
5050

51+
## Optional model proposers
52+
53+
Every phase above is a deterministic rubric by default. `FORGE_LLM=1` adds a thin
54+
**proposer** layer that can refine — never decide — the gate. With `TYPESAFE_API_KEY`
55+
also set, the route and assumption proposers prefer TypeSafe's System One (Jev): typed
56+
`choice` / `noul` answers with probabilities in ~150ms instead of a text round-trip.
57+
Any failure falls back to the deterministic path, so the flags are safe to leave off or
58+
on.
59+
5160
## Blast radius
5261

5362
**Blast radius** — the set of files an edit is predicted to impact, read from the code

‎package.json‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -52,14 +52,14 @@
5252
"scripts"
5353
],
5454
"scripts": {
55-
"test": "node --test test/*.test.js",
55+
"test": "node --test --import ./test/_setup.js test/*.test.js",
5656
"bench": "node bench/bench.mjs",
5757
"lint": "biome lint .",
5858
"format": "biome format --write .",
5959
"check": "biome check .",
6060
"check:fix": "biome check --write .",
6161
"typecheck": "tsc -p tsconfig.json",
62-
"coverage": "node --test --experimental-test-coverage test/*.test.js",
62+
"coverage": "node --test --experimental-test-coverage --import ./test/_setup.js test/*.test.js",
6363
"bump": "node scripts/bump.mjs",
6464
"forge": "node src/cli.js",
6565
"pages:build": "node scripts/build-pages.mjs",

‎src/docs_check.js‎

Lines changed: 3 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -20,6 +20,7 @@ const DOC_FILES = ["README.md", "docs/GUIDE.md", "ARCHITECTURE.md", "ROADMAP.md"
2020
// values injected by host tools rather than set by users.
2121
const INTERNAL_ENV = new Set([
2222
"_FORGE_LLM_KEY",
23+
"_FORGE_JEV_KEY",
2324
"FORGE_EMBED_KEY",
2425
// Test-only override of the settings.json path `forge init` targets — plumbing for
2526
// exercising merge/remove/exit-code behavior without touching the real ~/.claude.
@@ -31,7 +32,8 @@ const INTERNAL_ENV = new Set([
3132

3233
// Prefixes that mark an env var as OURS to document. A doc may freely mention other
3334
// tools' vars (GITHUB_TOKEN, PATH) — those aren't claims about forge's own surface.
34-
const ENV_PREFIX_RE = /\b((?:FORGE|ANTHROPIC|LITELLM|OPENROUTER|ENABLE_CORTEX)_[A-Z0-9_]+)\b/g;
35+
const ENV_PREFIX_RE =
36+
/\b((?:FORGE|ANTHROPIC|LITELLM|OPENROUTER|ENABLE_CORTEX|TYPESAFE)_[A-Z0-9_]+)\b/g;
3537

3638
function readDoc(root, rel) {
3739
const p = join(root, rel);

0 commit comments

Comments
 (0)