The repository directory is named hamlet; the Python distribution and the only live source tree
are townlet. Same project.
Townlet is a deep reinforcement-learning substrate expressed as configuration. An environment —
variables, observation layout, substrate topology, affordances, effects, items and reward
function — is written in YAML, compiled into a single frozen, hash-carrying CompiledUniverse
artifact, and executed GPU-natively against torch tensors.
The point is authoring. The project's endorsed vision (docs/product/vision.md) puts the pivot in
one line: from game as experience to writing a game as experience. The change it exists to make
is that someone with an idea for a mechanic or a game system can turn it into a running, trainable,
reproducible RL environment by writing config — no environment subclass, no observation-tensor
plumbing, no reward-function code. Each subsystem below exists to move a category of "you must
write Python for this" into "you can declare this."
The survival world shipped in configs/default_curriculum — eight meters, fourteen affordances,
one 8×8 grid — is intended as the first-class demonstration of that idea, not as the product
itself.
- Version 0.1.0, classified
Development Status :: 3 - Alpha. There are no release tags; the repository's only tags, locally and onorigin, are the two oracle tags below —oracle-2026-08-13(0e875d7a) andoracle-2026-08-17(4222a917). maincarries the recovery. The 168 commits of theproject-recoveryrewrite were merged through PR #32 (merge commit07b26ed5) on 2026-08-15, after both of the merge gates it was held behind were satisfied — CI restoration, and a claim-by-claim re-verification of this file. Before that merge,main's tip was dated 2025-11-28 and described a system that no longer existed. A second merge landed through PR #35 at4222a917on 2026-08-16, carrying theslow-marker deletion and the integration-test repairs, and a third through PR #36 at04062872on 2026-08-19 (41 commits: the blind re-run instrument, protocol Appendices A.6.1/B, and the previous re-verification of this file). The nightly onmainhas been green since the second merge — thirteen consecutive runs to 2026-08-28, the last nine against04062872(see Continuous integration). Repair continues on aproject-recovery*branch and reachesmainonly through the same gates, so between mergesmaintrails the branch by design — it is the last state that passed both gates, not the newest state. The unit-3 token cut described below has landed on the branch; whether it is onmainis answered bygit log main -1 -- src/townlet/universe/dto/token_spec.py, not by this file.- The project is mid strangler rewrite behind a pinned oracle. Tag
oracle-2026-08-17(commit4222a917, the merge commit of PR #35) freezes the previous system as the specification for preserved behaviour; it supersededoracle-2026-08-13(0e875d7a) on 2026-08-17, which stays as history. Fromdocs/oracle/ORACLE.md: "The oracle never mutates" — it moves forward to a new tag, never edits the old one — and "a diff against the oracle is a defect in the rebuild unless the register says otherwise." Accepted differences are recorded indocs/oracle/known-divergences.md(eleven entries; what a green harness run means has changed since the 2026-08-26 token cut — see Run something). - CI runs, and all three per-push gates are green at
1065dbf0— the last commit CI had reported on in full when this file was stamped, read at 2026-08-29T08:51Z. A commit cannot report on itself: the commit carrying this file is pushed after it is written, so its own runs land afterwards. Lint, Config Validation and Tests fire on every push, and that has been true since 2026-08-15 and not before; across the recovery branches, 283 of 356 completed runs passed. Sincea725bf66the default suite deselects nothing: theslowmarker that had kept 31 red integration tests out of every per-push gate is deleted, 29 of them are repaired and 2 deleted as dead, and the Tests job runs the whole suite. Read this as young rather than settled: before 2026-08-15 nothing had run on the recovery at all, nothing had passed anywhere since 2025-11-28, and the Lint gate has twice gone red for a stretch that nobody watching noticed — seven pushes on 2026-08-19, then 47 consecutive pushes, 2026-08-22 to 2026-08-29, while the product workspace recorded the branch as green (docs/product/current-state.md, 2026-08-26). Both were caught by a re-verification, not by the gate. See Continuous integration. - Where this file calls something shipped, it means present and wired, not mature. The project came out of a long stretch of intermittent attention and is unfinished in places; the specific gaps are under Known rough edges.
- No backwards compatibility: no fallbacks, no deprecation cycle, no migration paths. Breaking changes land directly.
Every command, file path, count and quotation below was executed or read against the working tree
at commit 1065dbf0 (2026-08-29, with Lint, Config Validation and Tests all green on it);
the repository-state facts under Continuous integration
were read from the GitHub API at 2026-08-29T08:27Z. Dates here are UTC, which is why the second merge is
dated 2026-08-16 although its commit carries 2026-08-17 in local time. This is the fifth full
claim-by-claim re-verification of this file. The first (1b25c99d) found ten claims stale in a
single day, every one because the recovery had fixed the thing being described; the second
(33bfff51, at the first merge) found five more in four commits; the third (905acd96, at the
second merge) found twenty-one in 27; the fourth (4a225d84, at the third merge) found eighteen
and four omissions in 43, and its adversarial pass found ten factual defects in the sweep's own
corrections before they were applied — which is the argument for the method over a re-read, and
the reason a sweep that finds nothing is treated here as a sweep that was not run. This one
covers 143 commits, the longest gap yet, and found 33 stale, wrong or misleading claims and
sixteen material omissions; its adversarial pass found six defects in the draft's own
corrections, the same shape as last time.
There are deliberately no test counts, coverage percentages, observation widths or
training-performance figures here — see Numbers. This file decays fast because the
project moves fast: it is a status report stamped at a commit, not a standing description, and
it is re-verified by sweep — not by re-reading — whenever it is published.
configs/default_curriculum/stratum.yaml declares the world's physics. This is the entire file:
stratum:
version: "1.0"
substrate:
type: grid
grid:
topology: square
width: 8
height: 8
boundary: clamp
distance_metric: manhattan
observation_encoding: relative
diagonals: true
vision_support: both
temporal_support: enabled
observation_mode:
mode: full_autoobservation_encoding and observation_mode configured the old raster observation spec and
nothing reads them now — measured, scaled and relative compile to a byte-identical
TokenSpec. They still parse, which makes them a No-Defaults violation rather than a choice;
filed as hamlet-6a4a6596bd, not removed here because it would touch every pack's
stratum.yaml.
Nothing else defines the substrate. The grid is 8×8 for every level in the pack: stratum.yaml
exists only at pack root, and the level loader reads no substrate file from a level directory.
Replacing that substrate block re-substrates the same experiment. The measured case is recorded
as Trial 001 in docs/product/metrics.md: the survival pack was moved to a 6-dimensional
gridnd world by editing about six lines of stratum.yaml, with zero lines changed under
src/townlet/ — it compiled, reset and stepped, the movement vocabulary auto-expanded to
DIM0_NEG … DIM5_POS, and the entire domain carried over. That trial also found the real caveat:
gridnd has no partial-vision support, so the pack's POMDP levels have to be switched to
active_vision: global or the whole-pack compile fails.
Rewards are declarative in the same way. This is
configs/default_curriculum/levels/L1_full_observability/drive.yaml, in full:
drive:
version: '1.0'
modifiers:
energy_crisis:
bar: energy
ranges:
- name: range_0
min: 0.0
max: 0.2
multiplier: 0.0
- name: range_1
min: 0.2
max: 1.0
multiplier: 1.0
extrinsic:
type: constant_base_with_shaped_bonus
base_reward: 0.01
bar_bonuses:
- bar: energy
center: 0.0
scale: 0.5
- bar: health
center: 0.0
scale: 0.5
variable_bonuses: []
apply_modifiers: []
intrinsic:
strategy: adaptive_rnd
base_weight: 0.1
apply_modifiers:
- energy_crisis
adaptive_config:
enabled: true
threshold: 100.0
decay_rate: 0.995
min_weight: 0.01
shaping:
- type: approach_reward
weight: 0.01
target_affordance: EAT
max_distance: 5.0
- type: completion_bonus
weight: 0.1
affordance: SLEEP
composition:
normalize: false
clip: null
log_components: true
log_modifiers: trueThere are no reward classes to subclass — src/townlet/environment/reward_strategy.py does not
exist and no RewardStrategy remains anywhere under src/. The compiler turns that YAML into a
GPU-native computation graph and hashes it into the compiled artifact, so a checkpoint knows which
reward function produced it. Two fields in that file are inert: composition.normalize and
composition.clip validate but have no reader (see Known rough edges).
configs/default_curriculum/
experiment.yaml # which levels the pack declares
stratum.yaml # substrate and topology (its observation keys are inert — see above)
environment.yaml # meter observation types (range_type), VFS variables, cascade graph
actions.yaml # substrate and custom actions, action labels
brain.yaml # architecture, optimizer, loss, Q-learning, replay
items.yaml effects.yaml vfs_profiles.yaml
transition_rules.yaml variables_reference.yaml action_labels.yaml # optional, compiled
presentation.yaml # optional; rendering hints for the live viewer — never compiled
levels/<level>/
curriculum.yaml # vision and temporal switches
bars.yaml # meters, bounds, cascades
affordances.yaml # interactions (interaction_type required), costs, hours, placement
drive.yaml # reward specification
training.yaml # hyperparameters, enabled actions
brain.yaml # optional; a COMPLETE replacement brain, never a partial patch
A level directory carries those five required files plus two optional ones: an items.yaml
declaring level-scoped item spawns, and — since d60104f0 (2026-08-22, PDR-0027) — a
complete brain.yaml that forks the whole brain for that level (one pack does:
configs/test/set_encoder_smoke); the loader (src/townlet/universe/raw_configs_v21.py) reads
nothing else from it. The shared catalogs vfs_profiles.yaml and effects.yaml are rejected
outright at level scope, and a level items.yaml must declare the v1.0 ItemsAppearance schema
(src/townlet/universe/loaders/preflight.py). Without a level brain.yaml the architecture
is pack-level: a level's training.yaml overrides exactly five scalars of the pack brain.yaml
— gamma, target-update frequency, the double-DQN flag, learning rate and replay capacity — and
none of them live in the architecture block. Of the optional pack-root files,
transition_rules.yaml (typed social-residue rules, 7e989e8c) is carried by no pack in
configs/; variables_reference.yaml, where a pack declares the extents that size its
zone, group and message scopes, by twenty.
Three declarations are worth knowing about because older docs predate them. Each meter's
observation type is declared per meter as range_type in environment.yaml — a closed,
parameterized vocabulary of nine normalization kinds (minmax, log_scaled, zscore, …); the
shipped pack uses minmax with clip: false for all eight. range_type no longer
reaches the observation (unit-3 token cut, 2026-08-26): a meter token's value lane is the
meter's clamped position within its bars.yaml [min, max], derived from the bars declaration
alone, so declaring cyclical_sin_cos on a clock meter or one_hot on a categorical one now
changes nothing an agent sees. The declaration still parses and still feeds the compiled
normalization spec. That is a designer-facing capability that silently stopped working, filed
as hamlet-1e335e0363 and not accepted as final — either the meter token carries the declared
normalization, or the surface is retired with a superseding record. Each environment.yaml variable
declares a semantic_type from a closed vocabulary, and each affordance declares its
interaction_type (instant, multi_tick or dual) — required, no default. And a pack may carry
an optional presentation.yaml at its root: the live-inference server reads it to render meters and
affordances (labels, colours, icons, plain/percent/currency formats), the compiler never does, so it enters no hash and
cannot change behaviour (src/townlet/demo/presentation.py,
docs/config-schemas/presentation.md — archived 2026-08-24, back at the live path since
931e26d8 on 2026-08-26, the only file in that directory with no staleness banner).
No shipped pack carries one; without it the viewer renders every meter honestly from its declared
bounds — a bar as a fraction of the declared range, the plain value, no % or $ inferred from a
name.
About the shipped curriculum, since older docs oversell it. default_curriculum declares five
levels, but bars.yaml, affordances.yaml and drive.yaml are byte-identical across all five,
and the substrate is pack-level. Only two levels change the world the agent sees:
L2_partial_observability (active_vision: partial) and L3_temporal_mechanics
(active_temporal: true, day_length: 24). Corrected 2026-08-26 at the unit-3 token cut:
this paragraph used to say observation_schema_hash took three distinct values across the five
levels — one shared by L0_0/L0_5/L1, one for L2, one for L3 — and read that as "three distinct
observation surfaces". It now takes exactly one: compiling all five gives an identical
observation_schema_hash, layout_hash, token_type_schema_hash and total_dims. The
compiled observation surface is the same at every level. Partial observability is a runtime
visibility filter over an unchanged TokenSpec (out-of-range spatial tokens have presence and
payload zeroed), not a compiled difference; and the day/night phase is now an authored global
in the pack-root vfs_profiles.yaml, so every level carries it, active_temporal or not. So
the "five documented levels are three distinct universes" reading — still true of the
mechanics — is no longer visible in the observation artifact at all, and a level's
distinguishing feature has to be looked for in curriculum.yaml, not in a hash. Within the
first group, L0_5_dual_resource and L1_full_observability differ, outside comments, in one
line of one file (output_subdir), and L0_0_minimal differs from them in training hyperparameters — the
double-DQN flag, target-update frequency, batch size, intrinsic-annealing floor, episode budget
and checkpoint interval — and in enabling one fewer custom action (REST only, against REST and
MEDITATE) — which does move its action_schema_hash and vfs_hash.
Python 3.13 or newer (.python-version pins 3.13) and uv. A GPU is
optional; the runtime falls back to CPU.
git clone https://github.com/foundryside-dev/hamlet
cd hamlet
uv sync --all-extrasNo branch checkout is needed: main carries the recovery. Ongoing repair lands on a
project-recovery* branch first and reaches main only through the merge gates, so main trails
active work by design — it is the last state that passed both gates, not the newest state.
Use --all-extras: it is what all four CI workflows specify, and a bare uv sync installs runtime
dependencies only — pytest, black, ruff and mypy live in the dev extra. No PYTHONPATH export is
needed; the project installs editable. There are no console-script entry points, so everything runs
as uv run scripts/<name>.py or uv run python -m townlet.<module>.
Check that a pack compiles (no cache written):
uv run python -m townlet.universe validate configs/default_curriculum \
--primary-level L1_full_observabilityCompile it, and inspect the artifact:
uv run python -m townlet.universe compile configs/default_curriculum \
--primary-level L1_full_observability
uv run python -m townlet.universe inspect configs/default_curriculum \
--primary-level L1_full_observability --format json--primary-level is required by compile and validate, and by inspect whenever you point it
at a pack rather than at a .msgpack artifact. Compiling a pack compiles every level in it, but
the cache holds one artifact per primary level, at
<pack>/.compiled/universe-<level>.msgpack (gitignored).
Train. --config takes the pack root, not a level directory, and --level and --inference-port
are both required:
uv run scripts/run_demo.py --config configs/default_curriculum \
--level L1_full_observability --episodes 10000 --inference-port 8766This runs training and a WebSocket inference server in one process and writes
runs/<output_subdir>/<timestamp>/, where output_subdir is read from the level's
training.yaml (run_metadata.output_subdir, required — there is no fallback; each
default_curriculum level sets it to its own directory name). The run directory holds
checkpoints/ with .sha256 sidecars, metrics.db, tensorboard/, training.log, and a
config_snapshot/ of the configuration that produced them.
Serve a trained checkpoint on its own — six positional arguments, the last two being the pack and the level:
uv run python -m townlet.demo.live_inference <checkpoint_dir> 8766 0.2 10000 \
configs/default_curriculum L1_full_observabilityRun the differential harness that adjudicates the rewrite. It creates or reuses a detached git
worktree at the oracle tag under .oracle/ (checking on reuse that it sits at the tag's commit and
is clean), runs the same logical pack, level, seed, agent count and step count on both trees as
subprocesses — each side is a (code root, pack root) pair: the oracle side reads its frozen copies
of the packs under oracle_fixtures/, the rebuild side reads the live configs/ — and compares
four env-step trace streams (obs, actions, dones, rewards; trace format v4 added
actions as an adjudicated stream, 9e7197e6) plus the compiled provenance hashes. Each side
draws its own seeded actions by default; --scripted makes the oracle side record its actions
and the rebuild side replay them verbatim, so the world's dynamics are compared under identical
inputs. Verdicts and a report.json land under runs/differential/<run-id>/; a --cell names
both device rows of a pack/level, so this prints two verdicts:
uv run python -m townlet.oracle.harness --cell default_curriculum:L0_0_minimalIts declared matrix is twenty cells: five levels of default_curriculum × {cpu, cuda}, three
single-axis packs under configs/differential/ × {cpu, cuda}, and two packs whose
vfs_profiles.yaml declares variables (configs/test/items_smoke, configs/test/effects_smoke)
× {cpu, cuda} — originally the only runnable packs that exposed VFS profile variables, added
under PDR-0074 so the cut that split the old obs_vfs block was visible to the harness at
all. The CUDA cells
are always declared and reported SKIPPED rather than silently dropped when --cuda is absent. It
exits 0 only when every cell is AGREE, SKIPPED, or DIVERGED_AS_REGISTERED — the last meaning the
cell's declared binding to a divergence-register entry matched narrowly, in one of three shapes.
Old-side crash: the oracle side crashed without producing a trace, the registered signature
appearing in the final exception text of its stderr, and the rebuild side ran and produced a
valid trace. Hash-only: both sides ran, exactly the enumerated provenance hashes differ — no more
and no fewer — and every trace stream matches byte-for-byte. Stream-scoped (afa09b81,
2026-08-22, built for the token cut): both sides ran, exactly the enumerated trace streams
diverge — shape changes included — and every other stream matches byte-for-byte. A cell may bind
a hash entry and a stream entry together; compare_traces labels that hash+stream. Behaviour
is never suppressed: an undeclared stream difference is DIVERGE, an undeclared hash moving is
HASH_MISMATCH, and a declared divergence that fails to manifest is REGISTERED_DIVERGENCE_ABSENT
— all red. An unmatched red of any kind
still fails, and an empty or all-SKIPPED run exits 1 so that doing nothing cannot look green.
What exit 0 means has changed since oracle-2026-08-17. At the re-tag, the sixteen
default_curriculum and differential cells declared nothing and their fixtures were byte copies
of the live packs, so a green run meant old and new agree. That is no longer true — the
matrix.py docstring says so in those words. Since the token cut was bound on 2026-08-26
(7b432de3), every one of the twenty cells binds three register entries — DIV-009
(six pre-cut compiler landings that moved provenance, not behaviour), DIV-010 (the engine
tick variable), and DIV-008 (the token cut, bound twice: as a hash entry naming five
fields — the three entries' union is the eight movers observed — and as the stream entry
naming obs alone). A green run now means everything diverged exactly as registered: obs
is permitted and required to differ on every cell, and actions, dones and rewards —
undeclared — are held byte-exact, which is what makes the token design's acceptance criterion
machine-checked rather than argued. The acceptance runs, runs/differential/20260826-172349
(--scripted) and 20260826-172441 (plain), both exit 0: ten CPU cells
DIVERGED_AS_REGISTERED with shape hash+stream, a streams key present and containing
exactly obs; ten CUDA cells SKIPPED, so the finding is CPU-only. Fixtures are no longer
byte copies either: vfs_profiles.yaml differs on the four standing and differential packs
(the authored time_of_day_phase global), effects.yaml and vfs_profiles.yaml on
effects_smoke, effects.yaml plus a fixture-only level brain.yaml on items_smoke
(DIV-007); a declared input delta and a declared output delta remain two decisions, and
neither blesses the other. Two cells no longer measure the axis they were added for
(div003_scaled, items_smoke) and are demoted as evidence in writing (PDR-0124; tickets
under Known rough edges).
The register holds eleven entries, in its own lifecycle vocabulary: DIV-001 and DIV-002
are checkpoint-boundary, tag-stamped, and cannot appear in an env-step trace; DIV-003,
DIV-004 and DIV-005 are retired at the tag; DIV-006 and DIV-011 are retired into
DIV-008; DIV-007, DIV-008, DIV-009 and DIV-010 are built. Any DIVERGE or
HASH_MISMATCH with no matched entry is a rebuild defect or a missing register entry — both
findings.
uv run ruff check .
uv run black --check src tests
uv run mypy src/townlet --show-error-codes
uv run python scripts/no_defaults_lint.py src/townlet/ --whitelist .defaults-whitelist.txt
uv run python scripts/validate_compiler_cli.py
uv run pytest
cd frontend && npm test # vitest; local only — no workflow runs ituv run pytest runs the whole suite: since a725bf66 there is no slow marker and no -m in
addopts, so the default command deselects nothing. The frontend gate exists as of a5cca764 —
frontend/package.json had never been in the repository before that commit.
Four GitHub Actions workflows exist, all specifying uv sync --all-extras on Python 3.13: Lint,
Config Validation, Tests, and Full Test Suite.
Three of the four run on every push, and 283 of the 356 completed runs on the
recovery branches have passed (read from the GitHub API at 2026-08-29T08:27Z). That became true on
2026-08-15 and had never been true before: between 2025-11-28 and that date no workflow had run
against the recovery at all — the 168 commits it merged in PR #32 landed across only seven pushed
shas, all on 2026-08-15, and nothing before them was checked. 116 shas have been checked in
all across project-recovery and project-recovery-2 — treat the gates as restored, not as
seasoned. And one gap is open on main itself: the third merge, 04062872, triggered no
per-push Lint, Tests or Config Validation run there, where both earlier merges had fired all
three within seconds; the nightlies are the only reads that tip has (hamlet-83c8e3b50e, P1;
the deciding test is the next merge).
- Lint, Config Validation and Tests trigger on
pushtomainand toproject-recovery*(and onpull_request). The glob is deliberate: the original defect was that the recovery branch simply was not named in the trigger list, so naming the next branch individually would have rebuilt the same trap on the next rename. The first runs in the recovery's history were green — Lint 1m11s, Config Validation 1m14s, Tests 24m21s. Of the 356 completed runs since, 73 have failed: 60 Lint, 13 Tests, and no Config Validation. The Lint reds come in two streaks. The first, thirteen through 2026-08-19, was mostlyruffline-length violations in trial probe scripts — one stood red for seven consecutive pushes before a merge gate caught it — plus the no-defaults linter once at8c5fa2c8, on product source. The second is the one to read: Lint was red for 47 consecutive pushes, from7dc6f66c(2026-08-22) to237b0c38(2026-08-29), last green before it0b659130(2026-08-21). Three steps took turns being the red one — the no-defaults linter, Black (64 files) andruff— and because the workflow runsruff→ Black → mypy → no-defaults and stops at the first failure, each red hid whatever was broken behind it. Fixed in two commits:237b0c38blacked the 64 files;b915139emade three real defaults required (TokenTypeSchema.slot_bindings,owner_capacity,source_map) and whitelisted the fifteen structural hits (eleven whitelist entries) with their reasons in.defaults-whitelist.txt. The Tests reds: the wall-clock ratio flake atbf0f2fe4, a hosted-runner loss ate65f59e1, two 2026-08-22 runs wheretest_masked_loss_during_trainingcompleted no episode (11c80a94,c9c1d50e; passing again fromba2766e6, cause not established), eight consecutive 2026-08-24 runs where a test openeddocs/architecture/vfs.mdafter its promotion toVFS.md(fixed478ab7ee), and one at9563dc45where the pack-freeze guard refused fixture drift no cell declared — closed by bindingDIV-008at7b432de3. No red has been shown to be a product regression, but the 2026-08-22 pair is unexplained rather than cleared, and the honest reading of the second Lint streak is that a gate nobody watches is not a gate. - The Tests job now runs the whole suite. Until
a725bf66the defaultpytestinvocation carried-m "not slow", and theslowmarker covered four files — three of them holding 31 tests that had been failing unseen (test_temporal_mechanics.py,test_training_loop.py,test_recurrent_networks.py): some broken by the recovery's own constructor and layout changes, some already stale onmainbefore it began (the temporal file assertedBed/Jobagainst a pack that namedSLEEP/WORKatf0a9ae8a). No per-push gate had ever run them; the first post-merge nightly was what surfaced the red. Of the 31, 29 were repaired and 2 deleted as dead (2ba1f530,e62a5e4a), the marker was deleted, and the per-push Tests job now executes the rest — every Tests run sincea725bf66has deselected nothing. - Full Test Suite — the same suite on a nightly trigger — had its nightly 06:00 UTC cron deleted during the
recovery and restored at the 2026-08-15 merge. The reason is worth knowing, because it is a
property of GitHub rather than of this repo: the scheduler reads the workflow file from the
default branch, so while the recovery lived on a branch, an enabled cron would have kept
testing a
mainfrozen atf0a9ae8a, ~160 commits behind the branch. That workflow has never passed — every scheduled run since 2025-11-03 was red, the last 64 of them (2025-11-28 to 2026-01-30) against an untouchedmain— until GitHub's dormancy rule disabled it; re-enabling it from the branch would only have resumed that stream against the wrong tree. The workflow isactiveagain, and it now passes. It fired twice againstmainat07b26ed5and was red both times, with the same 31 failures — the tests above, which thatmainstill deselected from every other gate and had not yet repaired. Since the second merge carried the marker deletion and the repairs tomainat4222a917, every run has been green — thirteen in a row: aworkflow_dispatchon 2026-08-17 (run 31981122221), three scheduled runs against4222a917(2026-08-17 to -08-19), and nine scheduled runs against04062872, the third merge, from 2026-08-20 to 2026-08-28 (runs 32340503178 through 33198018719; the last two fired at 17:17Z and 18:09Z rather than the usual ~06:35Z). Each merge puts commits onmainthat the nightly has never run against, so the first nightly after any merge is always the next reading to check. Sincea725bf66the nightly and the per-push Tests job are the same bareuv run pytestand differ only in trigger. - Three of the four — every one except Lint — run
scripts/validate_compiler_cli.pybefore their other steps, and no step setscontinue-on-error, so that script gates the rest. It exits 0, sweeping every pack it does not explicitly exclude. Read the exclusions, because one of them matters:EXCLUDED_DIRSnamestemplates,aspatial_testandreference_config, and onlyaspatial_testexists — the other two are dead names. Soconfigs/aspatial_test, one of the packs this file names as a working non-Town universe, is never validated by CI. It does validate by hand (exit 0), which is how the claim below is supported; it is simply not gated.
What CI does not cover, stated so the green is not read as wider than it is: the harness that
adjudicates the rewrite (townlet.oracle.harness) is run locally by the operator, not in CI; the
frontend's npm test and npm run build run locally only — no workflow installs Node; and two
members of the default suite — wall-clock ratio assertions (a 5% VFS-overhead ratio and a 1.5×
scripted-kernel ratio) taken under always-on coverage instrumentation — are flaky by
construction; one of them is the bf0f2fe4 red above. Tracked as one defect
(hamlet-f9090ec3e8).
src/townlet/ is the only source tree — there is no src/hamlet/ — and holds 16 packages. The
load-bearing ones:
universe/— the universe compiler (UAC). Parses and cross-validates a pack, resolves its references and shared schemas, compiles every level, and emits oneCompiledUniverse: a frozen dataclass carrying 19 declared*_hashfields plus per-level metadata (17 before the token cut;token_type_schema_hashandlayout_hashare the two it added). Not all of them are enforced — see checkpoint identity below. The cache is keyed on a config hash and a provenance id (compiler version, git sha, python, torch and pydantic versions), and is discarded when any config file's mtime is newer than the artifact's. The 2026-08-24 cleanup (312d0fe0,PDR-0121) gave it one authoritative stage enum (universe/stages.py), an error-code registry (universe/error_codes.py) and aSourceMap, so a compile error now carriesfile:line; the never-called cues seam (CuesCompiler,config/cues.py) was deleted.universe/dto/token_spec.py— the observation ABI, since the unit-3 token cut (4dde71a2, 2026-08-26). An observation is a set of typed tokens, not a raster: seven engine token types in a fixed order (self,meter,affordance,agent,item,effect,variable_element), each with a payload width fixed across all universes and a per-universe compiled capacity, serialized flat with a presence feature leading every row.environment/token_publishers.pyfills the flat view; partial observability zeroes out-of-range spatial tokens and never reshapes the tensor.token_type_schema_hashis the transfer contract (measured identical across a 2-D grid, a 3-D cubic grid and an aspatial universe);layout_hashis the per-universe flat-net contract. The old fixed-width superset and its per-level activity mask —ObservationSpec,ObservationActivity,vfs/observation_builder.py— are deleted, not wrapped.vfs/— variables and compiled transition programs (VTC). Access control is enforced at runtime, not merely declared:VariableRegistryraisesPermissionErrorwhen a reader or writer is not on the variable's list. The compiled transition schedule is built intoVectorizedHamletEnvand drives the ordered phases of the step loop; since7cbfbff8affordance occupancy is one of its phases, so contention is authorable fromactions.yaml.environment/dac_engine.py— declarative rewards (DAC), compiled from a level'sdrive.yaml.agent/— brain-as-code, layer 2.brain.yamlselects architecture, optimizer and loss throughnetwork_factory.py,optimizer_factory.pyandloss_factory.py. Since9a0007detoken_set(TokenSetQNetwork, with a declaredmeanorattentionaggregator,PDR-0112) is a shipped architecture andset_encoderno longer builds. Census, not intent: of the 39brain.yamlfiles inconfigs/, 29 declarefeedforward(everydefault_curriculumlevel included), 5dueling, 5token_set(thetoken_transfer_*andset_encoder_smokefixtures), nonerecurrent.environment/,population/,substrate/— the vectorized torch runtime. Device is an explicit parameter:VectorizedHamletEnvrequires one and raises rather than picking a default.training/checkpoint_utils.py— checkpoint identity. One shared gate,assert_checkpoint_identity, called by both the training-resume path (demo/runner.py) and the serving path (demo/live_inference.py). Eleven of the*_hashfields are stamped into a checkpoint (CHECKPOINT_FORMAT_VERSION = 4since the token cut; a version-3 checkpoint refuses loudly), and eight of those are hard-compared on load —vfs_hash,drive_hash, the effectivebrain_hash, the four per-level content hashes, and one of the two token hashes chosen by architecture: atoken_setnetwork comparestoken_type_schema_hash, every other reader compareslayout_hash, because a flat reader's dims are positional — alongside action count andprimary_level, so a checkpoint refuses to load into a universe it does not match, including a different level of the same pack.pack_brain_hashis stamped and required present but compared only to state a brain-lineage fork (PDR-0027); the old observation-dim and observation-field-UUID legs are gone with their producer. What is not enforced is recorded rather than hidden:observation_schema_hashis stamped and never compared, and the five pack-level hashes (experiment,stratum,environment,actions,items) are computed and serialized and compared by no checkpoint consumer — only the differential harness reads them, as a provenance diff between two compiles —DIV-001indocs/oracle/known-divergences.md.oracle/— the differential harness described above.demo/— training runner and live-inference server. The server is the only reader of the optionalpresentation.yaml(demo/presentation.py, DTO inconfig/presentation_config.py): it validates the file against the compiled universe's meter and affordance names at startup and raisesPresentationErroron an unknown one, forwards each meter's declared bounds and lethal edges to the viewer on connect, and no compiler stage or hash ever sees it.
Delivered and wired at this commit: a YAML pack compiles to a frozen, hash-carrying artifact
(configs/default_curriculum and configs/L5_multi_agent both validate clean); that artifact
drives the vectorized torch environment; the observation is a compiled token set whose
replacement of the raster ABI the oracle harness adjudicated on all ten CPU cells with world
dynamics byte-exact (PDR-0124); reward functions are specified in config, with no Python
reward classes left to subclass; VFS access control is enforced at runtime; all nine variable
scopes construct at runtime; temporality is authored (the day/night phase is a pack-level
variable over the engine's tick, not an engine block); and the training entry
point runs end to end, writing a run directory whose training.log records Training loop completed
normally beside a config_snapshot/ of the pack that produced it.
One measurement the cut left open was ruled, not hidden: the token design (PDR-0114) carried
a reversal trigger at 8× the pre-cut observation width, and L1 compiles to total_dims 1132
against a pre-cut 120 — 9.43×, 82% of it the affordance block. The trigger fired, was
escalated with three levers drilled and refuted (PDR-0124), and on 2026-08-29 the owner
ruled option 4: the 8× cap and the engine constants stay unchanged, and the 9.43× reading
is carried as recorded debt into migration unit 5, to be re-measured after the pack migration
moves the census.
Intent, not yet built — stated plainly because older docs blur the line:
- Brain-as-code layers 1 and 3. The behaviour contract (panic thresholds, forbidden actions,
personality dials, allowed goals) and the declarative think-loop graph are specified in
docs/architecture/archive/hld/02-brain-as-code.md(archived 2026-08-24; the current honest treatment isdocs/architecture/BAC.md) and have no implementation: their identifiers appear in zero files undersrc/andconfigs/. Layer 2, the network/optimizer/loss surface, is real, and since the token cut includes a token-native architecture (token_set, above); a token-native recurrent brain is still intent (migration unit 4) —RecurrentSpatialQNetworksurvives as a block reader over theself,meterandaffordancetoken rows. - One standard compiler for both halves of an experiment. The universe compiles to an artifact;
the brain rides inside it as a validated
BrainConfigplus abrain_hash, rather than compiling to an artifact of its own.CompiledBrainexists only indocs/. - A second demonstrator that varies the domain rather than the substrate. The substrate axis
has a measured witness (Trial 001, above).
docs/product/vision.mdnames four existing packs as domain-varying candidates of unverified depth (aspatial_test,L5_multi_agent,simple,reference— the pack isreference/model_pack). All four validate clean at this commit; whether any varies the domain enough to count as a witness is unassessed, and that — not the compile status — is the open question. - The "Low Energy Delirium" reward-hacking lesson described in older docs. It is not implemented: no level of the shipped curriculum declares the multiplicative reward the lesson depends on.
- The frontend builds and has a test gate, but the gate is local only.
frontend/package.jsonand its lockfile were restored ata5cca764— before that commit neither had ever been in the repository, sonpm run devcould not run althoughscripts/run_demo.py --helptold you to. Nownpm run buildsucceeds andnpm testruns the vitest suite (three files underfrontend/src/); no CI workflow installs Node or runs either. One component is dead code:AffordanceGraph.vueis mounted behind anaffordance_graphmessage that no server emits (hamlet-102db4c2e0). - A compiled pack can fail to cache without failing the command, and it is a class of failure
rather than one pack. (Fixed 2026-08-21, commit
03764c6b— agent profiles now serialize, the field is typedCompiledGlobalProfile | None, and a failed cache write fails the compile. Re-verified 2026-08-24: a pack with a non-nullagent_profilecompiles and writes its cache artifact. The record below is kept as stamped at 2026-08-20.)configs/reference/model_packcompiles, printsCompilation succeeded, and exits 0 — while its cache artifact is not written: serialization raisescan not serialize 'CompiledGlobalProfile' object, the failure is downgraded to a log warning the CLI never displays, and nothing propagates it to the exit code.inspectthen fails withArtifact not found. The trigger is a non-emptyagent_profile.variablesin a pack'svfs_profiles.yaml: packs declaring zero agent-profile variables cache normally, packs declaring one or more do not, and adding a single agent-profile variable to a pack that caches is enough to reproduce it. The error names the global class becauseuniverse/compiled.py(then line 123) typed the field asagent_profile: Any | None = None # TODO: Add CompiledAgentProfile type— the untyped field is the root cause, and the message points at the wrong half of the config. Compiling every pack inconfigs/from a cleared cache, exactly two fail this way today. CI cannot see any of it, because the gate runsvalidate, which writes no cache. This is the project's recurring shape — a failure that is not loud — and it is tracked as a defect rather than left as folklore. (Pack census re-taken 2026-08-29:configs/holds 38 directories carrying anexperiment.yaml; 19 are fixtures underconfigs/test/, three declared expected-to-fail (set_encoder_smokeand the threetoken_transfer_*fixtures are new since 2026-08-20); the other 19 aredefault_curriculum,L5_multi_agent,aspatial_test,simple,reference/model_pack, threedifferential/div003_*harness packs, and eleven authoring-trial packs — twotrial002_*and ninetrial_*, the ninth the blind re-run packtrial_b_blind_organism— for the trials indocs/product/trials/. All 35 non-negative packs validate clean at this commit.reference/model_packvalidates and compiles, but itsitems.yamldeclares aspawn_effectshape the runtime rejects, so env construction raises where the compiler passed —hamlet-5a87550adb, the same shape again.) - The declarable surface exceeds the exercised surface. Measured on 2026-08-20 by compiling
the 30 packs in
configs/that then compiled — every pack except the three negative fixtures — and counting rules in-process (the sample is now 35; the count is not re-taken):- Two of the nine compiled transition-program families —
action_writeandsocial_residue— carried zero rules in every one of them. Both now have an authoring surface no shipped pack uses: custom actions carry awrites:list (universe/compilers/actions.pyno longer hardcodeswrites=(), which was whyaction_writeused to be unproducible), and social-residue rules live in a typed pack-roottransition_rules.yaml(7e989e8c); nowrites:key and notransition_rules.yamlexists underconfigs/, so by census both families are still empty everywhere. (A third,interaction_progress, was also empty everywhere until 2026-08-15; repairingconfigs/reference/model_packbrought the only pack that exercises it back into the measured set, where it carries two progress rules and two completion-bonus rules — still the only pack inconfigs/that produces any, after ten trial packs were added to the sample. The surface did not change — the sample did.) drive.yaml'sintrinsic.strategyacceptsicmandcount_based, which have no implementation anywhere — those tokens occur only insideconfig/drive_as_code.py, in theLiteral, its docstring and an unreadicm_configfield.composition.normalizeandcomposition.clipvalidate but have no reader;dac_engine.pytakes onlylog_componentsandlog_modifiersfrom that block.type: grid3dwas deleted from the substrate schema (it never had aSubstrateFactory.buildbranch, so it could only compile toward a guaranteed crash); 3-D grids aretype: gridwithtopology: cubic.- Three of the nine declared variable scopes —
zone,groupandmessage— used to validate and compile clean and then hard-crash at environment construction, because nothing passed the registry itsnum_zones/num_groups/num_message_slotsextents and no YAML could set them (found by Trial K,docs/product/trials/). Fixed at6b752b3c(2026-08-21,hamlet-9e1ae3b7a2closed): a pack declares anextents:block invariables_reference.yaml, the compiler carries it into the level metadata,environment/vectorized_env.pypasses it to the registry, and a pack that declares one of those scopes without extents is refused loudly. Two packs declare them (L5_multi_agent,trial_o_bidding_blind). - The four VFS variables
configs/default_curriculum/environment.yamlused to declare —deficit_energy,deficit_satiation,time_since_last_eat,time_since_last_sleep— were observed but written by nothing, so agents saw frozen zeros in slots the ABI claimed were live. Deleted 2026-08-22 (hamlet-dc8f887cd5); the shipped pack now declares no custom variables. Trial L (docs/product/trials/0001/L-20260818.md) demonstrated the counter mechanic is authorable without them: a bar with a negative passive rate advances per tick, anon_startmodifyresets it on use.
- Two of the nine compiled transition-program families —
- Nine defects landed open with the token cut (
PDR-0124), all still in triage. Two are inline above — inertobservation_encoding(hamlet-6a4a6596bd, P1) andrange_typereaching nothing (hamlet-1e335e0363, P1). The rest: item tokens carry no declared identity (hamlet-559cc74246, P1); the indistinguishability refusal has no declared-parameter escape hatch (hamlet-2aca57c0f0, P1); effects surviveenv.reset()(hamlet-d76684f549, P1); affordance tokens skip the indistinguishability check and drop costs, hours and non-on_startinteractions (hamlet-81bf807963, P2); effect slots migrate columns on expiry and over-capacity items drop silently (hamlet-4538ba909f, P2); everyvariable_elementslot in the fleet is an inert declaration (hamlet-aba6171ff7, P2); and thereference/model_packenv-construction crash above (hamlet-5a87550adb, P2). A tenth — L3 unobservable after the cut — blocked and was fixed first, by declaration (9563dc45). - Documentation outside
docs/product/anddocs/oracle/is being reconciled, and the rewrite is blocked.scripts/README.mdstill documentsscripts/validate_configs.pyandscripts/validate_substrates.py, neither of which is present;CLAUDE.md, regenerated and corrected repeatedly, still describesCuesCompileras "instantiated atcompiler.py:69" afterbb43e024deleted it. The 2026-08-24 archive and the 2026-08-26 recovery are under Documentation; the corpus rewrite itself is gated onhamlet-ad2773718a(generate from the consuming code paths, not from the Pydantic models). An independent VFS gap analysis (docs/product/assessments/vfs-gap-analysis-20260821.md) scored 129 spec cells as 63 WORKS / 9 INERT / 26 BLOCKED / 31 ABSENT. The count of confirmed-false claims in canonical docs is tracked as a product guardrail indocs/product/metrics.md. - The recording subsystem is slated for removal once its intent is captured
(
pyproject.toml,recordingextra). Its MP4 export also shells out to anffmpegbinary that is not a Python dependency.
This README states no test count, no coverage percentage, no observation-vector width and no
learning-curve figures. Numbers like those start decaying the moment they are written, which is
how the last set went wrong: the README that sat on main until the 2026-08-15 merge (f0a9ae8a)
badged a test count and a coverage percentage and stated an observation width; docs/product/metrics.md
records that coverage figure as measured-false, and the width it gave is not what the compiler
reports for any default_curriculum level at this commit. Read them off the tree instead: compile a level and read
Observation Dim from the summary (the field is metadata.observation_dim, set from
token_spec.total_dims — plural — and it is a property of the pack, not a constant of the
project; the attribute was observation_spec.total_dims until the 2026-08-26 token cut deleted
that artifact, and token_spec.census says where the width goes, type by type); run
uv run pytest for the suite; and read docs/product/metrics.md for measurements
stamped with the commit and date they were taken at. Two trigger denominators live there, not
here: the frozen L2 pre-raster baseline (docs/product/baselines/2026-08-l2-preraster/,
PDR-0122, five seeds) and the 9.43× width reading above.
Current and maintained as part of the recovery:
docs/product/vision.md— purpose, audiences, anti-goals. Owner-endorsed; it separates what is shipped from what is intended, and tags each claim with how it was established. Re-stamped 2026-08-24 (PDR-0119): the loop ends with the trained model leaving for the designer's own game — train here, deploy there.docs/product/current-state.md— where the rewrite stands.docs/product/roadmap.md— the current bet list, stated as intent rather than dates.docs/product/metrics.md— dated measurements and the documentation-truth guardrail.docs/product/decisions/— every product decision as a numbered record with its reversal trigger;docs/product/prds/anddocs/product/trials/hold the authoring-trial instrument and the per-trial records this file cites for measured authorability claims;docs/product/assessments/holds the independent audits (the VFS gap analysis among them) anddocs/product/baselines/the frozen measurements that arm reversal triggers.docs/oracle/ORACLE.mdanddocs/oracle/known-divergences.md— the rewrite's rules and its accepted divergences.docs/README.md— the map of the rest ofdocs/, with a trust level per directory.
Subsystem detail lives in docs/architecture/ and docs/config-schemas/. On 2026-08-24
(PDR-0118) the old architecture corpus — including
docs/architecture/archive/UNIVERSE_AS_CODE.md (corrected 2026-08-16) and
docs/architecture/archive/vfs-current-implementation.md (corrected then and again on
2026-08-17, when the compiled observation field gained a typed feature) — was archived
wholesale to docs/architecture/archive/ and replaced by a six-document HLD set reviewed
against source that day: HLD.md, STRATA.md, UAC.md, BAC.md, COMPILER.md, and
VFS.md (the former vfs.md, promoted). A same-day recut (c4e8bd58, "zzz. archive") then
swept most of the rest of docs/ — docs/config-schemas/ included — into docs/zzz. archive/
(the literal directory name; about 400 markdown files sit there as history). On 2026-08-26
(PDR-0125, owner-authorised) 53 files were
recovered to their live paths with 51 dated staleness banners, all thirteen
docs/config-schemas/ files among them: twelve carry a 2026-08-26 banner — nine naming how
they are wrong (variables.md is wholesale 2025-11 stale; affordances.md documents a schema
wired to nothing; expressions.md calls nine shipped functions "planned"), three carrying a
✅ (transition_rules.md verified accurate; bars.md and training.md accurate but for one
known error each) — and only presentation.md carries none. Treat both archives as history,
never as a record of what shipped.
MIT. See LICENSE — Copyright (c) 2025 John.