Release compilation runs a bounded size optimizer after emission. Debug and hot reload bypass all optimizer scans. The objective is a smaller complete Wasm file; execution-speed optimization remains the engine's job. Binaryen is not a compiler dependency.
Direct emission also uses struct.new_default for payload-free first enum
variants in both profiles. This skips explicit zero/null operands without
adding an analysis pass. Release string emission chooses passive GC initializer
data by encoded cost instead of using a fixed minimum string length.
Before each of the two bounded cleanup sweeps, a module-wide analysis removes unobservable globals and substitutes short, proven constant initial values. The cleanup sweeps simplify integer instructions, fold adjacent integer constants and propagate uniformly constant integer locals, remove unreachable code and unused locals, simplify structured control flow, and share identical return sequences. The second sweep handles patterns exposed by the first. Compilation does not iterate to a fixed point, and it stops early when a sweep does not shrink the module.
Function sharing then merges bodies that differ only in integer constants. Function sharing initially retains each original function as a small wrapper passing its constants to a shared helper. The subsequent inliner can absorb single-reference helpers and remaps surviving function indices and reports. Exports and function-reference semantics remain intact; inlined origins are retained separately in the report.
Every cleanup body and complete module must shrink. Function-sharing costs include wrappers, helper bodies, new signatures/declarations, and section encoding overhead. The implementation preserves integer overflow and traps, does not fold floating-point operations, and retains effects of discarded values. Control-flow changes remap branch targets. Shared return sequences exclude bodies with non-defaultable locals; function sharing excludes tail calls and exception/continuation constructs. The tests cover those boundaries, GC/reference values, multi-value returns, large indices, and observable effects.
The promotion was measured against master at 1b65a40 and the full pipeline
at 3ecf397 on experiment/wasm-size-optimizations. The experiment and its
Binaryen comparisons remain on that branch. Commit 3f3c9be records a
reproducible comparison with individual passes disabled.
| Real script | Previous master | Promoted Release | Saved | Reduction |
|---|---|---|---|---|
| Minish Cap | 35,111 | 32,648 | 2,463 | 7.0% |
| Lunistice | 31,413 | 29,163 | 2,250 | 7.2% |
| Neon White | 5,073 | 4,768 | 305 | 6.0% |
| A Hat in Time | 50,205 | 48,536 | 1,669 | 3.3% |
The selected passes retain 98.2% / 98.4% of the full experimental savings on Minish Cap / Lunistice. Inlining and temporary sinking stay experimental: together, they save only another 46 / 36 bytes compared with this pipeline. Temporary sinking alone contributes 36 / 17 bytes in the two-sweep configuration. The remaining size benefit does not justify promoting that additional analysis and maintenance surface yet. The original straight-line propagation experiment stays off master; a separately measured constant-local analysis is described below. Redundant null rewriting is omitted because direct emitter fixes already provide its savings in both profiles.
Integer constant folding stays because it saves 113 bytes on Minish Cap and fits into the existing instruction scan. A second cleanup sweep saves another 282 / 41 bytes on the primary scripts when measured with temporary sinking enabled; it exposes useful simplifications without introducing another analysis. Pass savings interact and must not be added together as independent totals.
A fresh Binaryen 132 pass comparison on the promoted Release output identified instruction simplification as the largest remaining individual opportunity on the primary scripts. After subtracting Binaryen's no-pass re-encoding, it saved 371 / 315 bytes on Minish Cap / Lunistice; constant propagation saved 224 / 14, local simplification/coalescing 179 / 157, and code folding 92 / 125.
The existing Release cleanup now also inverts integer comparisons followed by
eqz, removes redundant Boolean normalization where the value or its consumer
permits it, and simplifies lossless extend/wrap pairs. Floating-point comparison
inversion is excluded because NaNs invalidate the usual inequality identities.
Identical if arms share one body while retaining the condition's evaluation
and the original typed label as a block. Negated conditions with an else can
swap complete arms. Branch depths, block parameters and result values remain
intact. Arm comparisons examine at most 256 instructions, and overlapping inner
matches wait for the existing second cleanup sweep. A more elaborate nested
rewrite added no real-script savings and was discarded.
Redundant null assertions are removed from statically non-null constructor/cast results and before GC reads/writes when intervening operands are only individual nontrapping constants or reads. Checks stay before calls, writes and potentially trapping operands, preserving the first trap and preceding observable effects. The pipeline still has two bounded cleanup sweeps and the same strict body and module size gates. Debug bypasses these changes.
| Real script | Before | After | Additional saving |
|---|---|---|---|
| Minish Cap | 32,648 | 32,488 | 160 |
| Lunistice | 29,163 | 29,025 | 138 |
| Neon White | 4,768 | 4,692 | 76 |
| A Hat in Time | 48,536 | 48,449 | 87 |
Seven new runtime tests cover signed/unsigned boundaries, non-Boolean conditions, truncation, unordered NaNs, typed branch parameters/results, nested selected calls, preserved condition traps, GC constructor results, packed array reads, and null checks before calls or division traps. All 501 library, 696 compiler, 20 binary and five example tests pass, as do Clippy, the browser-target check, and 36 baseline-versus-optimized corpus runtime invocations. All 176 maintained modules validate; all 211 runtime scenarios and the Debug/Release profile check pass. The unchanged Unity size gate also passes, including individual function/type budgets and Lunistice base/DLC behavior; no baseline refresh was needed.
Final optimized-host seven-sample medians, with no concurrent build or runtime suite, are:
| Real script | Passes disabled | Complete Release backend | Added backend time |
|---|---|---|---|
| Minish Cap | 3.101 ms | 6.717 ms | 3.616 ms |
| Lunistice | 16.138 ms | 19.389 ms | 3.251 ms |
| Neon White | 0.501 ms | 1.033 ms | 0.532 ms |
| A Hat in Time | 4.387 ms | 7.360 ms | 2.973 ms |
These include lowering/emission but exclude parsing/type checking. The complete
pipeline remains close to the original promotion's 6.448 / 19.161 ms totals on
the primary scripts. Cross-run differences include measurement variation and
do not isolate the new rules' cost. Samples remain in target/size-check/timing.json.
Binaryen's remaining instruction-only savings fall to 194 / 195 bytes on the
primary scripts. Full -Oz on the new output reaches 29,638 / 26,344 bytes;
these are diagnostic comparisons, not an assertion that the remaining passes
compose additively. The next measured candidates are constant propagation on
Minish Cap and further instruction/local simplification on both primary scripts.
Minish Cap retains packed constants in locals and repeatedly extracts their
halves with shifts and truncation. Binaryen's --precompute-propagate exposed
this opportunity. The new Release cleanup collects integer locals whose
explicit assignments all store the same literal, then replaces reads dominated
by a store. A conflicting or nonconstant write disqualifies the entire local.
Facts become available only after a store, including for parameters and
default-initialized locals. Scope exits and else discard facts established
inside that scope, so a branch that skips initialization cannot inherit them.
Outer facts survive calls and loops because all writes agree. Unsupported
exception/continuation control is excluded. This is a bounded linear scan,
without iterative control-flow analysis or heap/global assumptions.
Adjacent integer wrap/extend and zero tests now fold as well. Existing arithmetic, local removal and branch cleanup consume the exposed constants. Large literals may temporarily expand individual reads; the existing strict final body and module size gates reject a result that does not shrink. Debug still bypasses all optimization scans.
| Real script | Previous master | New Release | Additional saving |
|---|---|---|---|
| Minish Cap | 32,488 | 32,086 | 402 |
| Lunistice | 29,025 | 28,933 | 92 |
| Neon White | 4,692 | 4,692 | 0 |
| A Hat in Time | 48,449 | 48,325 | 124 |
Constant-local propagation without the new unary folding saved 285 / 64 / 0 / 84 bytes respectively. Together they save 8.6% / 7.9% on the primary scripts relative to pre-optimizer master. No fixture in the nine-script corpus grows; automatic Unity and the two collection fixtures save 298 / 307 / 307 additional bytes, and both small async fixtures are unchanged.
Six new runtime tests cover packed constants, both conditional arms, skipped initialization through branches/tables, loop backedges, conflicting writes, parameter/default values, repeated equal writes, signed/unsigned conversions, division traps and rejection of larger expanded literals. The existing Never-emission test now recognizes its marker as either an i32 or i64 constant, because folding an extension legitimately changes the instruction width.
Validation passes: 507 library tests, all 696 compiler tests (including the updated marker check), 20 binary tests and five example tests; Clippy; and the browser compiler's wasm32 target check. All 176 maintained modules validate, all 211 runtime scenarios and the Debug/Release profile check pass, and all 36 baseline-versus-optimized corpus runtime invocations pass. The unchanged Unity size gate passes, including per-function/type budgets and Lunistice base/DLC behavior; no baseline refresh was needed.
Optimized-host seven-sample medians, measured without a concurrent build or runtime suite, are 2.929 -> 6.300 ms for Minish Cap and 17.086 -> 20.343 ms for Lunistice (passes disabled -> complete Release backend). Neon White measures 0.523 -> 1.112 ms and A Hat in Time 4.360 -> 7.682 ms. These include lowering/emission but exclude parsing/type checking; cross-run variation means they do not isolate the new analysis's cost. The complete primary-script backend remains in the same range as the preceding pipeline. The large synthetic timings were noisy, so they are not used to infer pass overhead.
After this change Binaryen's standalone constant-propagation pass offers no net
saving relative to its no-pass roundtrip on the primary scripts. Instruction
simplification still saves 194 / 193 bytes, local simplification/coalescing
161 / 155, and code folding 92 / 125. Full -Oz reaches 29,653 / 26,350 bytes,
leaving a 2,433 / 2,583-byte gap. These pass results are diagnostic and do not
compose additively; Binaryen's own final output can change with input shape.
A local-copy propagation trial was rejected: running it early slightly grew
all four original real-script outputs, and moving it after control cleanup
saved only 0 / 4 / 0 / 12 bytes. Collection-fixture gains did not justify that
additional analysis. Binaryen's --vacuum instead exposed inexpensive missing
rules in the existing cleanup: empty else arms, discarded global reads, and
general constant-condition branches.
The promoted rules remove an empty else, replace a completely empty untyped
if with a drop of its condition, and discard unused global reads. They retain
condition calls and traps, global writes and typed block parameters. A literal
condition selects one complete arm; the replacement initially retains the
original typed block label, so branch payloads and depths remain valid. Existing
label cleanup then removes unused labels. This extends the existing bounded
scans, without another pass or any Debug work.
| Real script | Previous master | New Release | Additional saving |
|---|---|---|---|
| Minish Cap | 32,086 | 31,965 | 121 |
| Lunistice | 28,933 | 28,684 | 249 |
| Celeste external port | 32,364 | 32,275 | 89 |
| Neon White | 4,692 | 4,675 | 17 |
| A Hat in Time | 48,325 | 48,173 | 152 |
The Celeste source is live_split_celeste_port.split from the sibling porting
workspace, SHA-256
B8C09AD7A525B2200F58FE1902D365D11B07A6A2FE4A0467891BD59DA0D9378A.
Its unoptimized Release is 35,462 bytes: the complete pipeline saves 3,187 bytes
(9.0%). The other primary scripts now save 9.0% / 8.7% relative to pre-optimizer
master. None of the ten measured scripts grows. Celeste's baseline and optimized
modules validate in both wasmparser and Node, and its Debug equivalence check
passes. There is no maintained Celeste behavioral harness in this repository.
Five new runtime tests cover typed parameters and multi-value branch payloads, truthy non-Boolean constants, outer branch targets, selected and unselected traps, condition effects, implicit typed else values, non-defaultable local initialization, and retained global writes. The earlier assignment-factoring test now uses an opaque local condition so it continues to test factoring, rather than having the new constant-arm rule erase the conditional first. All 512 library, 696 compiler, 20 binary and five example tests pass, along with Clippy, formatting and the browser-target check. All 176 maintained modules validate, all 211 runtime scenarios and the Debug/Release profile check pass, and all 36 baseline-versus-optimized corpus runtime invocations pass. The runner additionally validates both Celeste modules and reports its missing behavioral harness explicitly. The unchanged Unity size gate passes, including per-function/type budgets and Lunistice base/DLC behavior; no baseline refresh was needed.
Optimized-host seven-sample medians, with no concurrent build or runtime suite:
| Primary script | Passes disabled | Complete Release backend |
|---|---|---|
| Minish Cap | 3.050 ms | 6.681 ms |
| Lunistice | 16.682 ms | 19.842 ms |
| Celeste | 2.577 ms | 5.293 ms |
These include lowering/emission and exclude parsing/type checking. The totals remain in the previous pipeline's range; cross-run variation does not isolate the incremental cost of individual rules. Debug bypasses all these scans.
On the new output, Binaryen's instruction pass saves another 194 / 186 / 412
bytes on Minish Cap / Lunistice / Celeste, after subtracting its no-pass
roundtrip. Local simplification/coalescing saves 160 / 116 / 43, code folding
92 / 114 / 92, and global simplification 60 / 3 / 252. Full -Oz reaches
29,653 / 26,354 / 28,882 bytes, leaving gaps of 2,312 / 2,330 / 3,393 bytes.
Instruction simplification, with Celeste included, is the next largest common
opportunity among these measured individual passes; results do not add linearly.
Celeste's instruction diff includes repeated all-zero struct constructors
replaced by struct.new_default (eleven six-field and six three-field examples),
108 removed null assertions, and eleven load/widen pairs combined into
i64.load32_u. Default struct construction is a concrete candidate for a direct
emitter improvement that could benefit both profiles without adding a pass.
For a payload-free first enum variant, the tag is zero and every payload slot
already receives its Wasm default. Emitting struct.new_default replaces those
explicit operands and struct.new, retaining the same type and fresh allocation.
This is a constant-time choice at the constructor site, skips the old field
emission loop, and applies to Debug as well as Release. Constructors with payload
expressions retain their existing emission, including effects and negative zero.
Contextual optional/result conversions still run outside this emission routine.
| Real script | Previous Release | New Release | Release saving | Debug code-section saving |
|---|---|---|---|---|
| Minish Cap | 31,965 | 31,933 | 32 | 32 |
| Lunistice | 28,684 | 28,660 | 24 | 25 |
| Celeste external port | 32,275 | 32,117 | 158 | 223 |
| Neon White | 4,675 | 4,675 | 0 | 0 |
| A Hat in Time | 48,173 | 47,785 | 388 | 388 |
Debug savings above use executable code rather than varying DWARF metadata. No measured module grows. Automatic Unity and the two collection fixtures also save 1,610 / 1,746 / 1,746 Release bytes; the real-script results justify the change independently. Existing optimization passes can amplify or absorb direct emission savings, so Debug and Release deltas need not match.
Relative to the original pre-optimizer outputs, cumulative savings are now 3,178 bytes (9.1%) for Minish Cap, 2,753 (8.8%) for Lunistice, and 3,345 (9.4%) for Celeste. These include the emitter improvement; the current pass-disabled outputs themselves are smaller than the original baseline.
The added source-level regression runs with both profiles and with optimization enabled/disabled. It checks mixed and packed payload fields, nonzero tags, effectful zero payloads, the sign of negative zero, and optional/result wrapping. All 513 library, 696 compiler, 20 binary and five example tests pass, alongside Clippy and the browser-target check. The corpus runner passes all 36 maintained behavioral invocations and validates Celeste, which still lacks a maintained gameplay harness here. Its external source remains unchanged.
All 176 maintained modules validate, all 211 runtime scenarios and the Debug/Release profile check pass. The Unity gate required a reviewed baseline refresh: four automatic-profile fixtures gain four shared IL2CPP helpers and lose one shared Mono helper, adding three function/type entries (20 type-section bytes and six function-section bytes). The two Mono Linux build bodies change from 28-byte wrappers plus a 98-byte shared helper to 95/105-byte direct bodies, a local 46-byte cost. The changed constructor shapes enable different sharing groups; no new runtime dependency, scratch memory or source fixture is retained. Against the immediately previous compiler, the three automatic-profile/metadata fixtures shrink by 1,592 bytes each and automatic Lunistice by 1,628 bytes. The baseline is refreshed to the measured output, tightening the complete-module budgets as well. Its recheck and Lunistice base/DLC behavioral checks pass.
Binaryen -Oz still produces 29,653 / 26,354 / 28,882 bytes for Minish Cap /
Lunistice / Celeste, leaving gaps of 2,280 / 2,306 / 3,235 bytes. Instruction
simplification now saves 162 / 162 / 254 bytes relative to Binaryen's no-pass
roundtrip. Widened loads remain a possible direct-emission follow-up. No new
compiler timing claim is made for this shortcut; it introduces no optimizer scan.
Looking only at isolated instruction passes underestimated the remaining
opportunities. Binaryen 132's src/passes/pass.cpp, SimplifyGlobals.cpp, and
Inlining.cpp show how its optimizing global/inlining passes rerun function
cleanup after exposing new opportunities. A fresh experiment replayed every
prefix of the open-world -Oz pipeline in one process, tested each standalone
pass with the same optimization/shrink settings, and disabled major pass families
within the full pipeline. Every resulting module validated in Node; each complete
replayed pipeline matched -Oz in size.
On master 40dd52a, disabling these families increased -Oz output by:
| Family disabled | Minish Cap | Lunistice | Celeste | A Hat in Time |
|---|---|---|---|---|
| Local simplification, reuse and common expressions | 2,646 | 3,356 | 1,061 | 1,856 |
| Inlining with cleanup | 504 | 898 | 571 | 533 |
| Global simplification and ordering | 401 | 38 | 1,462 | 3,711 |
| Branch/control simplification and folding | 533 | 848 | 826 | 933 |
| GC/reference optimizations | 183 | -49 | 259 | 21 |
These are interacting pipeline dependencies, not additive savings estimates for our compiler. In particular, disabling local cleanup also affects cleanup after inlining. Standalone optimizing inlining saved 1,396 / 1,820 / 1,466 bytes on Minish Cap / Lunistice / Celeste; much of that includes general function cleanup. The earlier narrow inlining prototype's small gain is not an upper bound.
Global cleanup was selected for its large measured Celeste/A Hat in Time benefit
and comparatively small implementation. The new Release analysis counts all
reads and checks every write against the global's literal initializer. Private,
unread globals can be removed; globals that only ever hold their initial value
can also be removed when replacing reads does not increase instruction size.
Deleted stores become drop, preserving evaluation, calls and traps. Existing
instruction/control cleanup then removes redundant constants and dead branches.
Another bounded sweep can discover globals made unread by that cleanup.
A concrete source is settings storage: the emitter reserves both current and previous values even when no code reads the previous value. Global cleanup removes unused storage and maintenance writes without changing host settings calls. Global counts fall from 197 to 105 in Celeste, 461 to 245 in A Hat in Time, 84 to 52 in Minish Cap, and 20 to 17 in Lunistice.
Imports, exports, shared globals and references in module initializers/offsets remain pinned. Unknown global-reference instructions are pinned conservatively. Nonliteral or allocating initializers stay intact. Floating constants are compared by bits; there is no floating-point arithmetic folding. Surviving indices are remapped through the Wasm reencoder, including module-level references. Function indices and report metadata remain unchanged. Each rewrite must shrink the whole module; Debug does not execute any of this analysis.
| Real script | Previous master | New Release | Additional saving | Binaryen -Oz |
Remaining gap |
|---|---|---|---|---|---|
| Minish Cap | 31,933 | 31,532 | 401 | 29,653 | 1,879 |
| Lunistice | 28,660 | 28,632 | 28 | 26,354 | 2,278 |
| Celeste external port | 32,117 | 30,648 | 1,469 | 28,882 | 1,766 |
| Neon White | 4,675 | 4,664 | 11 | — | — |
| A Hat in Time | 47,785 | 44,120 | 3,665 | 42,014 | 2,106 |
The ten-fixture corpus has no growth, and Debug executable/stable metadata remains identical with optimization enabled or disabled. The external Celeste source is unchanged; it receives compilation and validation, not a gameplay test. All 36 maintained baseline/optimized runtime invocations pass. Focused tests cover calls and traps from removed stores, imported/exported globals, mutations between calls, negative zero/NaN bits, large literals, and global initializer/data-offset remapping. The Debug-profile compiler regression now allows additional internal globals to be eliminated; it still explicitly checks that the debug-only binding was erased from Release lowering.
All 517 library, 696 compiler, 20 binary and five example tests pass, alongside Clippy, formatting and the browser-target check. All 176 maintained modules validate and all 211 runtime scenarios plus the Debug/Release profile check pass. The unchanged Unity size gate and Lunistice base/DLC behavior pass; no baseline refresh is needed.
Optimized-host seven-sample medians, with no concurrent build or runtime suite:
| Primary script | Passes disabled | Complete Release backend |
|---|---|---|
| Minish Cap | 2.971 ms | 7.616 ms |
| Lunistice | 16.579 ms | 20.688 ms |
| Celeste | 2.614 ms | 6.269 ms |
A Hat in Time measures 4.450 -> 9.336 ms. These include lowering/emission and exclude parsing/type checking. The primary-script totals are approximately 0.85–0.98 ms above the earlier control-cleanup measurements; cross-run differences do not isolate this pass's cost. Large collection timing was noisy and is not used to infer overhead. The byte gains justify the added Release work.
Rerunning Binaryen attribution on this output leaves only 11 / 11 / 0 / 53 bytes of benefit from its global family on Minish Cap / Lunistice / Celeste / A Hat in Time. This addresses the measured opportunity. The next substantial direction is local/control simplification that also enables profitable inlining, especially for Lunistice; widened loads are a much smaller priority.
Reproduce the attribution with a generated size corpus (optionally including the external Celeste manifest):
node scripts/binaryen-pass-attribution.mjs C:/Projekte/binaryen/bin/wasm-opt.exeThis writes per-prefix, standalone and disabled-family modules plus report.json
under target/binaryen-attribution. Binaryen is an offline reference tool only.
Further Binaryen source/output inspection separated the effects of inlining
from the cleanup it triggers. On 1091435, standalone plain inlining grew
Minish Cap / Lunistice / Celeste to 33,034 / 31,225 / 31,290 bytes, whereas
optimizing inlining reached 30,101 / 26,718 / 29,139. Restricting the latter to
single-caller candidates (plus Binaryen's trivial-wrapper rule) retained almost
all its savings. The missing ingredient is still cleanup around expanded code.
Two local experiments were not promoted:
- Retrying the earlier inliner against all defined functions, including generated helpers, with the current local/control cleanup saved only 167 / 95 / 115 bytes on the primary scripts. Trying all call sites improved Celeste by only another six bytes. That gain does not yet justify importing the inlining machinery.
- Inferring integer arguments identical at every direct call saved only 7 / 8 / 7 bytes on the primary scripts. Larger automatic-Unity and collection savings did not justify another whole-module analysis.
The selected change extends structured control cleanup. Short, closed integer
expressions and reference reads can replace an if/else with a select when
both arms are nontrapping and effect-free. Arm and condition scans are bounded;
condition writes to an arm's locals prevent reordering. Calls, stores, loads,
division, casts that can trap, and allocations are not speculated. Reference
results use typed select. A separate three-byte gain from recognizing calls
inside conditions was discarded rather than adding call metadata for it.
Fallthrough cleanup removes a branch/return only when the immediately following block/function ends reach the same destination. It uses wasmparser's instruction arities and tracks structured operand-stack heights. Every crossed frame must have the same stack base and exact result types; equal arity alone is insufficient for GC reference subtypes. Conditional branches become a drop of the evaluated condition. Loop backedges, branches that discard extra operands, and unsupported exception/continuation control remain intact. This adds a Release-only body analysis inside the existing bounded cleanup, with the existing body/module size gates. It does not change function indices, signatures or Debug emission.
| Real script | Previous master | New Release | Additional saving | Binaryen -Oz |
Remaining gap |
|---|---|---|---|---|---|
| Minish Cap | 31,532 | 31,477 | 55 | 29,651 | 1,826 |
| Lunistice | 28,632 | 28,530 | 102 | 26,351 | 2,179 |
| Celeste external port | 30,648 | 30,502 | 146 | 28,882 | 1,620 |
| Neon White | 4,664 | 4,646 | 18 | 3,936 | 710 |
| A Hat in Time | 44,120 | 43,943 | 177 | 42,014 | 1,929 |
No measured fixture grows. Conditional-expression selection alone accounts for
26 / 50 / 108 / 116 bytes on Minish Cap / Lunistice / Celeste / A Hat in Time;
fallthrough cleanup provides the remaining gains. This is a smaller incremental
step than globals, and the measurements do not claim it closes the inlining gap.
The wider inliner and constant-argument trial are preserved only as local
experiments under target, not enabled compiler passes.
Six new runtime tests cover conditional arithmetic, mutations in the condition, unselected calls and traps, nullable reference selections, branch conditions and effects, discarded stack operands, loop backedges, typed block parameters, multiple results, and incompatible intermediate reference-result types. All 523 library, 696 compiler, 20 binary and five example tests pass. The corpus passes its Debug equivalence checks and all 36 maintained runtime invocations; the external Celeste port validates and remains unchanged, without a maintained gameplay harness here.
All 176 maintained modules validate, all 211 runtime scenarios and the Debug/Release profile check pass, as do Clippy, the browser compiler check and the Unity size gate including Lunistice base/DLC behavior. No Unity baseline refresh was needed. The optimized host reproduces identical sizes for all ten corpus fixtures.
Seven warmed, alternating optimized-host samples measured these backend medians:
| Real script | Passes disabled | Complete Release backend |
|---|---|---|
| Minish Cap | 2.986 ms | 8.335 ms |
| Lunistice | 16.096 ms | 21.148 ms |
| Celeste | 4.851 ms | 11.682 ms |
| Neon White | 0.517 ms | 1.423 ms |
| A Hat in Time | 4.629 ms | 10.575 ms |
These include lowering/emission and exclude parsing/type checking. They measure the entire pipeline, not the isolated cost of this change. Cross-run timing variation, particularly in Celeste's passes-disabled baseline, prevents treating differences from the previous measurement as this pass's cost. Debug still bypasses all optimization passes.
A separate diagnostic Binaryen run with --closed-world -Oz --converge reached
28,834 / 24,196 / 28,297 / 40,104 bytes for Minish Cap / Lunistice / Celeste /
A Hat in Time. These reference modules validate in Node but were not run through
the behavioral harness. They are not the ordinary -Oz comparison in the table
above. The larger reduction, especially in Lunistice, motivates further study
of interacting type/call simplification and repeated cleanup; it does not imply
that a single missing pass will achieve it.
Function-level Binaryen inspection exposed another direct-emission opportunity:
GC strings shorter than 32 bytes still used one i32.const per UTF-8 byte,
followed by array.new_fixed. Many metadata and display strings are shorter
than that cutoff but substantially cheaper as passive data. For example, each
ASCII byte at or above 64 needs a two-byte signed LEB operand in addition to the
constant opcode, whereas passive data stores the byte once.
Release emission now compares the two encodings as each literal is emitted. The comparison includes constant operands, segment indices, newly stored bytes, segment headers, the DataCount section, and a conservative allowance for growth of the data-section size prefix. Existing pooled bytes are reused. Only a strictly smaller estimate is accepted; empty and tiny literals stay inline. This extends the existing literal pool without adding a module pass or a Binaryen dependency. Each use still allocates a fresh GC array, and passive initializer data remains available across calls and suspension. It does not occupy linear memory. Debug retains its existing emission and hot-reload behavior.
Measured against master 3b849bb, with the same real-script sources:
| Real script | Previous Release | New Release | Additional saving | New Binaryen -Oz |
Remaining gap |
|---|---|---|---|---|---|
| Minish Cap | 31,477 | 30,516 | 961 | 28,687 | 1,829 |
| Lunistice | 28,530 | 27,949 | 581 | 25,781 | 2,168 |
| Celeste external port | 30,502 | 30,352 | 150 | 28,730 | 1,622 |
| Neon White | 4,646 | 4,488 | 158 | 3,775 | 713 |
| A Hat in Time | 43,943 | 43,943 | 0 | 42,014 | 1,929 |
No measured fixture grows. Automatic Unity and the collection fixtures also shrink, but the real-script savings justify the change independently. These savings also improve the input to Binaryen: its new outputs are smaller too, so the remaining optimization gap is largely unchanged. A refreshed pass-family comparison still identifies local/control cleanup and its interaction with inlining as the larger remaining opportunities.
Validation covers complete encoded module sizes around signed/unsigned LEB boundaries, section overhead, duplicate literals, ASCII and multibyte UTF-8. A source-level runtime test checks short literals, empty strings, byte lengths, repeated allocation and suspension in both profiles. The existing static-data test now explicitly checks active segments: passive GC data is allowed, while GC-only literals must still stay out of linear memory.
All 525 library, 697 compiler, 20 binary and five example tests pass, as do Clippy and the browser compiler check. All 176 maintained modules validate, all 211 runtime scenarios and the Debug/Release profile check pass, and the size corpus passes its 36 behavioral invocations. Celeste remains unchanged and validates, without a maintained gameplay harness here.
The Unity baseline requires a reviewed refresh because bytes move from code
into passive data. All 38 complete modules shrink relative to the stored
baseline; function/type counts, helper sets, source fingerprints, linear static
data, scratch/read capacities and memory-page counts are unchanged. Data-section
growth is expected, and local Map/Set fixtures gain a three-byte DataCount
section. Several discovery/scanning helpers grow by one byte because an existing
pooled string's offset crosses the signed-LEB 64-byte boundary. For example,
Lunistice's UnityDiscoverIl2Cpp64::poll differs from 3b849bb only by changing
that offset from 56 to 87. Both Lunistice editions pass before refreshing the
baseline. The refresh also records the earlier control/global savings already
on master; the incremental real-script table above isolates this change.
Binaryen's remaining control-flow differences include a value-producing if
whose first arm immediately branches out, and chains of pure selections. The
existing Release cleanup now converts the first pattern into br_if followed
by the other arm. It retains a typed block around that arm until ordinary label
cleanup proves the label unused. This preserves internal branch targets, values
below the condition, and loop backedges. The rule excludes parameterized ifs,
whose bare branch can carry a payload, and branches to the if's own label.
Expression analysis now recognizes both ordinary and typed select operands.
It rewrites completed inner selections before inspecting their parents, so a
chain can simplify in one traversal. This replaces the old deferred list of
nonoverlapping edits; no new pass or cleanup sweep is added. The existing
32-instruction arm and 256-instruction condition limits remain, as do the
nontrapping/effect-free arm requirement and rejection of conflicting condition
writes. Debug emission is unchanged.
Measured against master 77bea0a:
| Real script | Previous Release | New Release | Additional saving | Binaryen -Oz |
Remaining gap |
|---|---|---|---|---|---|
| Minish Cap | 30,516 | 30,492 | 24 | 28,662 | 1,830 |
| Lunistice | 27,949 | 27,920 | 29 | 25,781 | 2,139 |
| Celeste external port | 30,352 | 30,210 | 142 | 28,721 | 1,489 |
| Neon White | 4,488 | 4,482 | 6 | 3,775 | 707 |
| A Hat in Time | 43,943 | 43,834 | 109 | 41,999 | 1,835 |
No measured fixture grows. Early-exit rewriting alone contributes 24 / 4 / 49 / 16 bytes on Minish Cap / Lunistice / Celeste / A Hat in Time; nested selections provide the remainder. The change is small in implementation scope and has its clearest real-script benefit in Celeste and A Hat in Time. Full Binaryen pass-family comparisons still show larger interacting local/inlining savings; this is incremental control cleanup, not a replacement for that work.
Three new runtime tests cover early-exit effects, discarded stack operands, internal typed labels, loop backedges, parameterized-if rejection, a 17-way selection chain, and writes inside selection conditions. The GC reference test also exercises nested typed selections in both the condition and an arm. All 528 library, 697 compiler, 20 binary and five example tests pass, as do Clippy, the browser compiler check, Debug-equivalence checks, and the size corpus's 36 behavioral invocations. Celeste's external source remains unchanged and validates; it still has no maintained gameplay harness here. All 176 maintained modules validate, all 211 runtime scenarios and the Debug/Release profile check pass, and the optimized Unity gate passes both Lunistice editions without a baseline refresh.
A local-lifetime experiment reused slots for nonoverlapping lexical intervals,
conservatively widening them across loops and protecting reads of implicit
defaults. Alone it saved 68 / 28 / 28 bytes on Lunistice / Celeste / A Hat in
Time but grew Minish Cap by 83 bytes by interfering with function sharing.
Giving the earlier inliner full module type information, two cleanup sweeps,
local reuse, and temporary sinking improved the all-call-site trial to
292 / 300 / 351 / 237 bytes saved on Minish Cap / Lunistice / Celeste /
A Hat in Time. Those experimental modules validate, but the larger machinery
and extra analyses are still not promoted. Sources and logs remain under
target/lifetimes-trial and target/lifetimes-inline-sinking.log locally.
Binaryen's instruction diff instead exposed a smaller implementation opportunity:
some ordinary struct constructors still push every zero/null field explicitly.
The existing Release instruction cleanup now uses struct.new_default when
every operand is a literal Wasm default. It retains the exact struct type and
a fresh allocation. Packed integer fields, positive floating-point zero and
nullable references are supported; negative zero, NaNs, nonzero fields and
effectful computations are not replaced. This complements the earlier direct
default-enum emitter shortcut; the general operand check remains Release-only.
The same cleanup now removes left-hand integer identities such as 0 + index,
1 * value and -1 & value when the other operand is a nontrapping, zero-input
push. In particular, local.tee is not such a push. These rules extend the
existing instruction traversal and add no module pass. Debug is unchanged.
Measured against master 8bc799d:
| Real script | Previous Release | New Release | Additional saving | Binaryen -Oz |
Remaining gap |
|---|---|---|---|---|---|
| Minish Cap | 30,492 | 30,421 | 71 | 28,662 | 1,759 |
| Lunistice | 27,920 | 27,836 | 84 | 25,781 | 2,055 |
| Celeste external port | 30,210 | 30,102 | 108 | 28,721 | 1,381 |
| Neon White | 4,482 | 4,451 | 31 | 3,775 | 676 |
| A Hat in Time | 43,834 | 43,667 | 167 | 41,999 | 1,668 |
No measured fixture grows. Struct-default rewriting alone contributes 68 / 66 / 90 / 164 bytes on Minish Cap / Lunistice / Celeste / A Hat in Time. Binaryen's final outputs on those four scripts are unchanged, so these savings reduce the measured gap rather than also moving the reference result.
Three new runtime tests check packed/float/reference fields, fresh allocation identity, retained effects, negative-zero and NaN payload bits, both integer widths and boundaries, and the distinction between local reads and tees.
Validation passed: 531 library, 697 compiler, 20 binary and five example tests; Clippy with warnings denied; the browser compiler wasm32 check; all 176 maintained modules and 211 runtime scenarios; the Debug/Release profile check; and 36 corpus runtime invocations. The Unity size gate and Lunistice base/DLC behavior passed without changing the baseline. All ten corpus fixtures validate and retain Debug equivalence. Celeste received compilation, validation and Debug-equivalence checks; no maintained gameplay harness is available here.
The follow-up developed on experiment/wasm-multivalue-inlining extends
f2ed1f2 with single-reference inlining and the cleanup that makes it profitable.
The selected pipeline is promoted together; all additional analysis is Release-only.
Binaryen's standalone optimizing inliner saved 1,395 / 1,660 / 1,181 / 1,111 bytes beyond its no-pass rewrite on Minish Cap / Lunistice / Celeste / A Hat in Time. Omitting individual nested passes confirmed that the benefit spans local cleanup, branch cleanup, instruction rewriting and GC allocation cleanup. No single additional small rule explains the whole gap.
The selected pipeline combines:
- Single-reference direct-call inlining, including multiple results, with exact
argument order and fresh locals on repeated calls. Exports,
ref.func, element references, recursion and unsupported control constructs protect callees. - Backward liveness over the structured control-flow graph, retaining loop-carried values and initial defaults. Dead stores retain their producers and their effects/traps. Locals share a slot only with the same type and no conflicting live value; parameter slots retain their indices.
- Bounded straight-line copy forwarding, expression movement through pure pushes, and removal of temporary set/get pairs that can stay on the stack.
- Removal of unused whole recursive type groups. A retained group keeps all its members, including unused members, preserving recursive type identity.
Inlining may temporarily add up to 64 bytes to a caller/callee pair, because later signature removal and cleanup can recover more. The complete candidate is accepted only if it is strictly smaller than the existing pipeline's output. Candidate rewriting has a 5 MB input-body budget, local analysis has body/local limits and a worklist budget, and cleanup remains two bounded sweeps. Module signatures are parsed once for the per-candidate cleanup context.
| Real script | Previous f2ed1f2 |
New Release | Saved | Binaryen -Oz on new output |
Remaining gap |
|---|---|---|---|---|---|
| Minish Cap | 30,421 | 29,898 | 523 | 28,658 | 1,240 |
| Lunistice | 27,836 | 26,971 | 865 | 25,867 | 1,104 |
| Celeste external port | 30,102 | 29,569 | 533 | 28,713 | 856 |
| A Hat in Time | 43,667 | 43,193 | 474 | 42,022 | 1,171 |
| Neon White | 4,451 | 4,227 | 224 | 3,813 | 414 |
The new Binaryen reference differs slightly from its output on master: inlining and sharing choices interact. The remaining gap is measured from each new artifact, not obtained by subtracting savings from an old reference.
Removing stack/copy cleanup produces 29,931 / 27,090 / 29,585 / 43,197 bytes on the four larger scripts. Removing slot coalescing produces 30,032 / 27,362 / 29,766 / 43,302; removing liveness entirely produces 30,238 / 27,803 / 29,819 / 43,411. Trying all call sites instead of only single-reference callees did not consistently improve the real scripts, so that larger trial is excluded.
The sidecar report now distinguishes actual indexed function bodies from
inlined_functions. function_names() includes both for code-demand checks;
it does not pretend an inlined helper still has a function index. Report
requests continue to produce identical Wasm. The Unity gate still checks both
code provenance and actual per-body sizes, with a reviewed baseline required
when helpers move into their callers.
Validation passed: 544 library tests (the 543-test suite plus the focused recursive-group test), 697 compiler tests, 20 binary tests and five example tests; Clippy with warnings denied; browser compiler wasm32 checking; 176 validated runtime fixtures, 211 runtime scenarios and the profile check; and 36 corpus runtime invocations. Coverage includes argument order/traps, early returns and branch tables, repeated calls with fresh locals, function-reference nulls, multiple results, loop-carried/default values, dead-store effects, source-local writes inside moved expressions, and stale call edges removed by constant cleanup. All ten measured fixtures validate, none grows, and stable Debug output is unchanged.
The Unity baseline was refreshed after reviewing all changes. Every complete module shrinks or stays equal; source fingerprints, scratch/memory requirements, type/function counts, retained helper provenance and section budgets have no growth. The old per-body gate correctly reports larger callers absorbing former helper bodies. Both Lunistice editions pass, and the refreshed gate passes without relaxing its checks.
The extra work increases Release compilation cost. In the optimized-host Unity run, Lunistice takes 90.2 ms versus 29.7 ms in the prior recorded baseline; automatic Lunistice takes 383.2 versus 127.1 ms, and the nested metadata fixture takes 545.1 versus 111.2 ms. These are whole-compile single-run measurements, not an isolated same-run pass benchmark. They nonetheless show that repeatedly constructing instruction-level liveness graphs for expanded callers adds measurable work. These subsecond Release times are acceptable for the size benefit; they are not a promotion blocker. Debug has no added analysis. Future work should prioritize the remaining size gap, while keeping analysis bounded and checking compile times for unusually large scripts.
Local comparison artifacts and ablation logs are under target/big-savings.
The external Celeste file remains unchanged and has no maintained gameplay
harness here. Debug and hot reload bypass the entire candidate pipeline.
A fresh Binaryen 132 attribution run after b725c29 identified locals and
control flow as the largest remaining families: disabling local passes in
-Oz costs 505 / 425 / 222 / 344 bytes, and disabling control passes costs
360 / 487 / 422 / 584 bytes on Minish Cap / Lunistice / Celeste / A Hat in
Time. These are interacting pipeline differences, not additive estimates.
Source inspection of Binaryen's ReorderLocals.cpp and the emitted branch
sequences guided the following changes to our existing Release cleanup.
- Local layout chooses the smallest encoding among the original order, frequency order, type grouping, and type grouping within equal LEB-index widths. The cost includes declaration groups and every local reference. Parameters stay fixed, types stay exact, and no lifetimes are merged here.
- An inert literal branch payload can move before a closed condition, turning
if; literal; br; endintoliteral; condition; br_if; drop. This preserves condition calls and traps, zero-arity loop targets and values below the expression. It excludes branches to the removed if label, parameterized ifs, unknown stack effects and conditions exceeding the scan bound. - A typed block starting with an early conditional exit can become a typed
if/else. The old continuation retains its label depth. Both the condition and early value must have exact, self-contained stack effects; discarded extra operands and branches within these expressions disqualify the rewrite. - Cleanup permits up to six shrinking sweeps instead of two, stopping as soon as another sweep offers no reduction. This exposes additional wins on automatic-profile Unity code; it does not change the four primary results.
| Real script | Before | After | Saved | Binaryen -Oz on new output |
Remaining gap |
|---|---|---|---|---|---|
| Minish Cap | 29,898 | 29,748 | 150 | 28,656 | 1,092 |
| Lunistice | 26,971 | 26,797 | 174 | 25,867 | 930 |
| Celeste | 29,569 | 29,448 | 121 | 28,715 | 733 |
| A Hat in Time | 43,193 | 43,057 | 136 | 42,018 | 1,039 |
| Neon White | 4,227 | 4,180 | 47 | — | — |
Automatic Lunistice saves 845 bytes. Collections / nested collections save 798 / 919; async fixtures each save eight. All ten complete modules shrink, validate, and retain identical stable Debug output. Extra analysis remains Release-only.
Ablations on Minish Cap / Lunistice / Celeste / A Hat in Time: type grouping alone saved 12 / 61 / 10 / 18 bytes; adding literal conditional branches reached 130 / 167 / 112 / 84, and adding leading-exit restructuring reached the final 150 / 174 / 121 / 136. Frequency-aware layout selection retained those corpus results while protecting large local-index boundaries. Increasing sweeps from two to six adds another 125 bytes on automatic Lunistice and 211 / 317 on the collection fixtures. Each transformation remains bounded and subject to the existing body/module size gates.
The full suite passes 548 library, 697 compiler, 20 binary and five example
tests. Clippy with warnings denied and the browser wasm32 check pass, as do
176 runtime-fixture validations, 211 runtime scenarios, the profile check and
36 corpus runtime invocations. Both Lunistice editions and the refreshed Unity
gate pass. New runtime regressions cover mixed local types
and defaults across index 127, fixed parameters, condition effects/traps,
function-reference null payloads, loop backedges and continuation branch depths.
Celeste has compilation/validation coverage but no maintained gameplay harness.
The Unity baseline was reviewed and refreshed. Every complete module shrinks
or stays equal. Source fingerprints are unchanged, with no increases in types,
functions, helper provenance, memory or section budgets. Two helper bodies have small
optimization-interaction regressions: ModuleElfIdentitySegment grows 12 bytes,
while its module shrinks nine; UnityCollectionTypesArrayElementClass grows
five in four map/set fixtures whose modules each shrink by over 200 bytes.
Those explicit per-body changes are accepted in the new baseline; the gate's
checks remain unchanged.
The optimized-host run measured 95.4 ms for Lunistice (previous baseline 90.2), 406.1 ms for automatic Lunistice (383.2), and 653.9 ms for nested metadata (545.1). These are cross-run whole-compile measurements, not isolated pass costs. They remain well below the compile-time range of concern. Debug adds no work.
Local measurements are under target/local-layout; the fresh attribution run
is under target/after-inline-attribution. A concrete remaining opportunity
in Binaryen's argument-elimination output is specializing formatting helpers
when every caller supplies the same numeric radix, alongside removing unused
helper parameters. Its end-to-end size benefit still needs a separate trial.
Binaryen's argument-elimination output specializes integer formatting when all calls pass the same radix, removes unused helper parameters, and follows those changes with local/constant cleanup. The new Release candidate applies those call-site facts before inlining, where they provide more benefit.
Only private functions referenced exclusively by ordinary direct calls qualify.
Exports, element/global ref.func references, tail-call targets, and unsupported
function subtypes keep their signatures. All calls must agree on an integer
constant before its parameter is specialized. An unused parameter can also be
removed, even when callers supply different values. In either case, the caller's
argument must be an inert one-instruction push: executed calls, loads, allocations
and trapping expressions are never erased. A bounded backwards stack walk can
locate earlier arguments across complete expressions; control boundaries,
unknown stack effects and shared multi-result producers stop that walk.
Unwritten constant parameters become literals. Written ones use initialized private locals, retaining fresh values for each activation; unused parameter stores retain their producers as drops or stack values. Function indices remain stable during this pass, and the existing type pruning removes unused signatures. Up to three shrinking rounds expose constants passed through helper chains. The existing integer folder now covers all signed/unsigned i32/i64 comparisons, so newly constant range checks can actually disappear.
Early transformations can affect later inlining and function sharing. Therefore the compiler evaluates the existing pipeline and the specialized pipeline and retains only the smaller complete module, including its matching report. No interim signature or function-body saving is sufficient on its own. Debug bypasses this entire process.
| Real script | Before | After | Saved | Binaryen -Oz on new output |
Remaining gap |
|---|---|---|---|---|---|
| Minish Cap | 29,748 | 29,660 | 88 | 28,657 | 1,003 |
| Lunistice | 26,797 | 26,701 | 96 | 25,870 | 831 |
| Celeste | 29,448 | 29,398 | 50 | 28,718 | 680 |
| A Hat in Time | 43,057 | 43,031 | 26 | 42,018 | 1,013 |
| Neon White | 4,180 | 4,149 | 31 | — | — |
| Automatic Lunistice | 116,062 | 115,571 | 491 | — | — |
Collections / nested collections save 562 / 623 bytes and async / async-loop save 29 / 60. All ten outputs validate, none grows, and stable Debug output remains identical. The external Celeste file is unchanged and still lacks a maintained gameplay harness.
Experiments show why placement and argument analysis matter: late specialization with the added comparison folds saved only 59 / 35 / 36 / 0 bytes on Minish Cap / Lunistice / Celeste / A Hat in Time. Moving it before inlining reached 88 / 47 / 36 / 26; allowing up to three shrinking rounds raised Lunistice to 66. Looking past complex later arguments reached the final results above. These are cumulative pipeline measurements, not additive pass contributions.
Six focused argument tests exercise constant agreement, differing call sites,
unused and overwritten parameters, local remapping, recursive calls, host
side effects, traps, exported/referenced functions and pinned tail calls. The
comparison regression checks every integer comparison against Wasmtime at
signed/unsigned boundary values. Measurement and experiment logs are under
target/arguments.
Validation passes 555 library tests, 697 compiler tests, 20 binary tests and five example tests, plus Clippy with warnings denied, the browser compiler's wasm32 check, 176 runtime-fixture validations, 211 runtime scenarios, the profile check and 36 corpus runtime invocations. Both Lunistice editions and the refreshed Unity gate pass.
The Unity baseline refresh accepts one per-body regression: StringFind grows
134 to 138 bytes in nested metadata, while its complete module shrinks by 489.
Every complete module shrinks or stays equal. Source fingerprints are unchanged;
there is no growth in retained helper provenance, type/function counts, memory
requirements or section budgets. No regression-gate checks are relaxed.
The optimized-host comparison measured 165.9 ms for Lunistice versus 93.5 ms in the prior recorded baseline, 986.2 ms for automatic Lunistice versus 409.1, and 1,001.6 ms for nested metadata versus 547.8. These cross-run whole-compile measurements include the extra candidate-pipeline comparison and are not isolated pass timings. This is an accepted Release-only cost for enforcing the stronger size guard; Debug adds no analysis.
A fresh standalone Binaryen comparison on the final output leaves only
14 / 38 / 15 / 16 bytes in dae-optimizing beyond Binaryen's no-pass rewrite
on Minish Cap / Lunistice / Celeste / A Hat in Time. Optimizing inlining still
saves 743 / 0 / 395 / 461 in isolation, while local CSE saves 172 on Lunistice.
These are candidates for the next investigation; standalone pass savings do
not compose and are not guaranteed attainable by a single local change.
The next investigation followed Binaryen's optimizing inliner into its nested cleanup passes. On the previous Minish Cap output, it removes just one small shared addition helper; most of its 743-byte improvement comes from cleanup inside the large caller. Disabling individual nested passes identified local coalescing, instruction simplification and scalar replacement as useful targets. Expanding more shared callees ourselves made Lunistice and A Hat in Time larger, so that experiment was rejected.
Release cleanup now replaces private struct allocations with field locals when
every reference use is a field read or write. The reference local must have one
allocation site, no aliases or escaping uses, and the allocation must dominate
all accesses within structured scopes. Objects allocated inside loops initialize
their fields on every iteration. Packed, shared and non-defaultable field
layouts remain conservative exclusions. Immediate struct.new; struct.get
projections also avoid constructing an object, while retaining evaluation of
every field, including unused fields with effects or traps.
Two related cleanup rules share identical closed tails of if arms and keep
local values on the operand stack across balanced straight-line statements.
Branch targets and non-defaultable local initialization prevent unsafe tail
movement. Existing control-flow liveness decides whether the first local read
can disappear or must become a tee for subsequent reads. These transformations
retain the existing body/module size gates and run only in Release.
| Real script | Before | After | Saved | Binaryen -Oz on new output |
Remaining gap |
|---|---|---|---|---|---|
| Minish Cap | 29,660 | 29,433 | 227 | 28,657 | 776 |
| Lunistice | 26,701 | 26,583 | 118 | 25,856 | 727 |
| Celeste | 29,398 | 29,201 | 197 | 28,714 | 487 |
| A Hat in Time | 43,031 | 42,984 | 47 | 42,018 | 966 |
| Neon White | 4,149 | 4,109 | 40 | — | — |
| Automatic Lunistice | 115,571 | 114,792 | 779 | — | — |
Collections and nested collections save 865 and 943 bytes; async fixtures stay unchanged. All ten modules validate, none grows, and stable Debug output remains identical. The external Celeste source remains untouched and has no maintained gameplay harness.
The staged experiment separated the improvements: stack forwarding saved
28 / 14 / 6 bytes on Minish Cap / Lunistice / Celeste; shared branch tails added
104 on Lunistice. Private struct locals then saved another 165 on Minish Cap and
180 on Celeste. Immediate projections added 34 / 11 on those scripts and 725 on
automatic Lunistice. These are sequential pipeline measurements, not independent
pass savings. Experiments and comparison artifacts are under target/inline-next.
Nine focused execution tests cover field mutation and constructor effects, loop initialization, discarded-field traps, null reads, identity uses, packed field truncation, branch effects and exits, non-defaultable locals, and repeated local lifetimes with later reads. They compare original and optimized execution including host call traces and trap kinds.
Validation passes 564 library tests, 697 compiler tests, 20 binary tests and five example tests, Clippy with warnings denied, and the browser compiler's wasm32 check. All 176 runtime fixtures validate; all 211 scenarios and the profile check pass, as do 36 corpus runtime invocations.
The Unity baseline review accepts four collection callback bodies growing by two bytes each (IL2CPP map/set and Mono map/set). Their complete modules shrink by 147 / 147 / 376 / 376 bytes. These are interactions between successive cleanup and inlining choices; per-invocation body gates do not guarantee every final body beats its previous compiler version. Every complete module and section stays equal or shrinks, source fingerprints stay identical, and the gate reports no other regressions. No regression checks are relaxed.
Optimized-host whole-compilation measurements are 169.0 ms for Lunistice, 1,106.0 ms for automatic Lunistice and 1,079.5 ms for nested metadata, versus 172.1 / 716.1 / 1,075.6 in the previous recorded baseline. These are cross-run measurements, not isolated pass costs. Debug adds no analysis.
On the new output, standalone Binaryen heap-to-local replacement saves nothing on Minish Cap, while local CSE still saves 170 / 206 bytes on Lunistice / A Hat in Time relative to Binaryen's no-pass rewrite. This suggests local expression reuse as a next candidate; standalone pass savings do not compose.
Following Binaryen's local-CSE results, Release now reuses already-computed
integer calculations, casts and eligible field/memory reads. A selected first
occurrence stores its result with local.tee; later occurrences read that
temporary. The cost model includes instruction encoding, local-index LEB sizes
and declaration costs, and reuses temporary slots with non-overlapping lifetimes.
The existing exact function-body and complete-module gates make the final
decision. This is a final pass, so it cannot disturb earlier specialization,
inlining or function-sharing choices. Debug does not run it.
Expression keys include local assignment versions and a state generation for mutable reads. Calls and stores invalidate state-dependent expressions; immutable field reads and pure arithmetic can survive unrelated effects. Only values computed before a conditional/block are available to both arms and its continuation. Loops start with fresh facts. Shared memory/globals and unsupported exception/control constructs conservatively exclude reuse.
Overlapping expressions retain their original dominating computation. Selecting a larger expression may allow a nested original to remain, but may never move a smaller expression's cache initialization into just one conditional arm. The scan is bounded by body/local/expression limits, available-fact limits and a budget for copying expression keys at control boundaries.
| Real script | Before | After | Saved | Binaryen -Oz on new output |
Remaining gap |
|---|---|---|---|---|---|
| Minish Cap | 29,433 | 29,217 | 216 | 28,493 | 724 |
| Lunistice | 26,583 | 26,217 | 366 | 25,618 | 599 |
| Celeste | 29,201 | 29,050 | 151 | 28,573 | 477 |
| A Hat in Time | 42,984 | 42,677 | 307 | 41,711 | 966 |
| Neon White | 4,109 | 4,088 | 21 | — | — |
| Automatic Lunistice | 114,792 | 113,195 | 1,597 | — | — |
Collections / nested collections save 1,893 / 2,163 bytes; async / async-loop
save 1 / 19. The first strictly straight-line prototype saved only 32 / 113 /
11 / 12 on Minish Cap / Lunistice / Celeste / A Hat in Time. Tracking dominating
expressions through structured conditionals and accounting for existing local
declaration groups made the larger savings possible. Experiment artifacts are
under target/cse-next.
Eight focused execution tests cover local writes and operand-stack snapshots, calls, memory/global/struct mutations, loops, conditional joins, overlapping expressions, first-trap ordering and non-defaultable reference initialization. They compare returned values, host-call traces and trap kinds. A shared-memory fixture verifies that expression reuse is disabled independently of ordinary unused-local compaction.
Validation passes 572 library tests, 697 compiler tests, 20 binary tests and five example tests, Clippy with warnings denied, the browser compiler's wasm32 check, and 36 corpus runtime invocations. All ten corpus modules validate and shrink; stable Debug output remains identical. The external Celeste source fingerprint is unchanged; its validation does not replace a gameplay harness. All 176 maintained runtime fixtures validate, and all 211 execution scenarios plus the Debug/Release profile check pass.
The Unity gate and both Lunistice editions pass with no function-body, section or whole-module growth. Its refreshed baseline records only equal or smaller outputs; source fingerprints, function/type counts and retention budgets stay unchanged. Optimized-host whole-compilation measurements are 173.7 ms for Lunistice, 1,050.6 ms for automatic Lunistice and 1,026.4 ms for nested metadata, versus 172.3 / 1,058.9 / 1,073.9 in the preceding baseline. These cross-run measurements show no material slowdown and are not isolated pass timings.
On the new output, Binaryen's standalone local CSE saves only four bytes on Minish Cap, grows Lunistice/Celeste, and still saves 149 on A Hat in Time relative to its no-pass rewrite. Optimizing inlining plus its nested cleanup still saves 476 / 0 / 243 / 400 on Minish Cap / Lunistice / Celeste / A Hat in Time. Those figures identify remaining investigation targets, not additive pass benefits.
Expression reuse can expose new dead stores, copies, discarded calculations and control-flow simplifications. Release now revisits the existing cleanup passes for at most six rounds, stopping at the first round that does not shrink the complete module. Each rewritten function must also shrink. Function indices, signatures and report mappings do not change in this final stage.
The added unused-value cleanup traces operands of discarded, nontrapping calculations. It removes arithmetic and comparisons but retains operand calls, assignments, allocations, potentially trapping operations and their original execution order. It can remove discarded floating-point calculations, including NaN and infinity results; it does not constant-fold floating-point values. Integer division/remainder, nonsaturating float-to-integer conversions, loads, GC reads and casts remain observable. Unknown instructions, control boundaries and multiple-result producers conservatively stop operand tracking.
Late local liveness also preserves structural initialization of non-defaultable reference locals. A store can be dead at runtime yet required by Wasm validation when both conditional arms overwrite the local before a read at their join.
Against 75512bd, the complete-module results are:
| Fixture | Before | After | Saved |
|---|---|---|---|
| minish_cap | 29,217 | 29,164 | 53 |
| unity_explicit | 26,217 | 26,183 | 34 |
| native | 4,088 | 4,046 | 42 |
| large_native | 42,677 | 42,659 | 18 |
| unity_automatic | 113,195 | 113,141 | 54 |
| collections | 142,810 | 142,712 | 98 |
| nested_collections | 152,158 | 152,058 | 100 |
| async | 1,920 | 1,910 | 10 |
| async_loop | 2,786 | 2,776 | 10 |
| celeste | 29,050 | 29,014 | 36 |
Binaryen 132 -Oz produces 28,493 / 25,618 / 28,573 / 41,711 bytes on
the new Minish Cap / Lunistice / Celeste / A Hat in Time output. Its remaining
advantage is 671 / 565 / 441 / 948 bytes. These are modest incremental gains;
the experiments did not uncover another large general-purpose saving.
Seven focused execution tests cover discarded arithmetic, float exceptional values, ordered calls, eager select operands, assignments, integer/conversion and GC null traps, control/multiple-result boundaries, and structural reference initializers. Validation passes 579 library, 697 compiler, 20 binary and five example tests, Clippy with warnings denied and the browser compiler wasm32 check. All ten corpus modules validate with identical stable Debug output, and all 36 maintained corpus scenarios pass. Celeste has no maintained gameplay harness; its external source fingerprint is unchanged. All 176 maintained runtime fixtures validate, and all 211 execution scenarios plus the Debug/Release profile check pass.
The Unity gate and Lunistice base/DLC behavior pass with no function-body, section or whole-module growth. Optimized-host full compilations measured 185.1 ms for Lunistice, 1,214.9 ms for automatic Lunistice and 1,047.3 ms for nested metadata, versus 173.3 / 1,077.2 / 1,004.6 ms in the previous checked-in baseline. These are cross-run measurements, not isolated pass timings.
The larger experiments did not justify promotion. Conditional facts added only
6 / 14 / 6 / 6 bytes of savings on Minish Cap / Lunistice / A Hat in Time / Celeste
beyond cleanup, despite larger synthetic benefits. Prioritizing copy hints and
interference degree during local allocation made Minish Cap 22 bytes larger for
a four-byte A Hat in Time saving. Tiny multi-use forwarding wrappers saved only
15 more bytes on A Hat in Time, with no primary Minish Cap, Lunistice or Celeste
benefit. Running discarded-value cleanup earlier did not improve the primary
scripts. These experiments are excluded; artifacts are under target/post-cse.
The next investigation compared all four real scripts with Binaryen 132,
replayed every -Oz prefix, disabled related groups of passes, and inspected
individual function changes. Source inspection used Binaryen checkout
79dfe6b412a3c22bfdb190ed6a4d79adf734db5d, particularly SimplifyLocals.cpp,
RemoveUnusedBrs.cpp, SSAify.cpp and the default pass schedule in pass.cpp.
Artifacts are under target/next-biggest; the full pipeline replay exactly
matches -Oz on each script.
On 0a0686d, disabling control-flow passes cost 236 / 275 / 391 / 209 bytes on
Minish Cap / Lunistice / A Hat in Time / Celeste. Disabling local-variable passes
cost 283 / 273 / 278 / 153 bytes. Disabling optimizing inlining cost only
41 / 0 / 62 / 37. These are interacting pipeline ablations, not additive savings.
The apparently large standalone optimizing-inlining wins mostly include its
nested cleanup; they do not establish that more inlining is the dominant gap.
A second diagnostic fed Binaryen's individual pass results through our optimizer.
Against matching no-op writer settings, simplify-locals-nostructure unlocked
155 / 200 / 229 / 62 additional bytes. Merely splitting locals with ssa-nomerge
unlocked 37 / 22 / 18 / 53. This points to expression movement and control-flow
reasoning before investing in a larger SSA conversion or another allocation
heuristic. The diagnostic used the intermediate structured-condition prototype;
these numbers are evidence for the chosen direction, not further savings on the
final implementation.
The attribution tool now records encodingBaseline, a no-pass rewrite with
optimization level 2 and shrink level 2. Binaryen's writer also responds to
these settings: its ordinary roundtrip differs even without an optimization
pass. Standalone and prefix deltas must use the matching encoding baseline.
Whole-module comparisons always use the actual compiler output.
The promoted implementation runs after the established Release pipeline and its final cleanup, for at most six additional shrinking rounds:
- Conditional selection follows closed nested
if/blockexpressions and floating-point conditions, while checking local writes and rejecting calls, global writes and escaping labels. Nontrapping arms can be evaluated eagerly; allocations, casts and loads are not speculated. A bounded 256-instruction arm limit handles longer chains that stopped at the earlier 32-instruction limit. - Temporary expressions can move to their first use across independent calculations. One side must be nontrapping and independent of external state, or both sides must avoid external writes and one must be nontrapping. Reads and writes of locals are checked in both directions. A potentially trapping expression cannot cross another potentially trapping expression or an observable effect. A retained tee preserves later uses and initialization.
The transformations share the existing nontrapping-operation classification and preserve floating-point calculations. Each changed body and the complete module must shrink against the completed previous pipeline. Movement stays within bounded straight-line regions. Debug and hot reload do not run this analysis.
Early application initially saved 100 / 67 / 177 / 54 bytes on the four real scripts, but grew the Mach-O UUID and PE debug-ID fixtures by 2 / 8 bytes. Locally smaller code interfered with later optimization choices. Running the new transformations only after the established pipeline removes this ordering regression, even though it gives up some of the early prototype's gains. The final gate reports no body, section or module growth; the early version is not promoted.
Against 0a0686d, complete-module results are:
| Fixture | Before | After | Saved |
|---|---|---|---|
| minish_cap | 29,164 | 29,092 | 72 |
| unity_explicit | 26,183 | 26,151 | 32 |
| native | 4,046 | 4,046 | 0 |
| large_native | 42,659 | 42,496 | 163 |
| unity_automatic | 113,141 | 113,057 | 84 |
| collections | 142,712 | 142,549 | 163 |
| nested_collections | 152,058 | 151,891 | 167 |
| async | 1,910 | 1,908 | 2 |
| async_loop | 2,776 | 2,774 | 2 |
| celeste | 29,014 | 28,970 | 44 |
Binaryen -Oz on the new output gives 28493 / 25618 / 41711 / 28573 bytes
for Minish Cap / Lunistice / A Hat in Time / Celeste, leaving gaps of
599 / 533 / 785 / 397 bytes. Local-variable and control-flow groups remain the
largest broad opportunities: their final-pipeline ablations cost
250 / 236 / 264 / 109 and 200 / 267 / 250 / 209 bytes, respectively. Inlining
still contributes only 41 / 0 / 62 / 37 bytes.
Eight focused execution tests cover both directions of local dependencies, global mutations and mutating calls, nested conditions and escaping returns, floating-point NaNs/infinities/signed zero, and null and memory traps. The final pipeline passes 587 library, 697 compiler, 20 binary and five example tests, Clippy with warnings denied, and the browser compiler wasm32 check. All ten corpus modules validate with unchanged stable Debug output; all 36 maintained corpus runtime scenarios pass. All 176 maintained runtime fixtures validate; all 211 execution scenarios and the Debug/Release profile check pass. Celeste's source fingerprint is unchanged; it still has no maintained gameplay harness.
The optimized-host Unity gate and Lunistice base/DLC behavior pass without any function-body, section or module growth. Whole-compilation times on the final gate run were 229.6 / 1,347.8 / 1,101.9 ms for Lunistice / automatic Lunistice / nested metadata. These are cross-run measurements rather than isolated pass timings. The refreshed baseline preserves source fingerprints, function/type counts and retention information.
The next deeper target is movement through control-flow boundaries and stronger proofs about GC reads and dependencies. The current implementation deliberately does not reorder potentially trapping reads against each other, or move a producer into a branch where its execution or definite initialization could change. A broader effect analysis should be measured against these specific remaining differences rather than adding more independent peephole rules.
Optimized-host seven-sample medians for the promoted pipeline:
| Real script | Passes disabled | Promoted Release backend | Added backend time |
|---|---|---|---|
| Minish Cap | 2.978 ms | 6.448 ms | 3.470 ms |
| Lunistice | 16.505 ms | 19.161 ms | 2.656 ms |
| Neon White | 0.505 ms | 1.005 ms | 0.500 ms |
| A Hat in Time | 4.556 ms | 7.469 ms | 2.913 ms |
These measurements include lowering and emission but exclude parsing/type
checking. Previous full-experiment totals were 7.591 / 20.856 ms on the primary
scripts. The selected pipeline is faster in these measurements while retaining
almost all size savings; cross-run differences are not isolated per-pass costs.
All nine fixtures and individual samples are in target/size-check/timing.json.
The optimized-host corpus check also validates identical sizes to the native
development build. Debug does not pay these optimizer costs.
The promoted pipeline has focused instruction/control/function-sharing tests and a source-level behavioral test covering evaluation order, GC values, fallible results and suspension. The real-script regression checks module validation, size reduction, unchanged original function indices and report metadata, and identical output with or without requesting a report.
The opt-in nine-fixture corpus records complete module bytes and section payload
sizes, validates baseline and optimized modules, and checks Debug equivalence.
DWARF .debug_info variable entries already have nondeterministic ordering in
unoptimized builds, so that comparison excludes only this section; executable
code, names, source maps and line tables must match. A smaller deterministic
fixture also checks the entire Debug binary byte for byte. The original pass
promotion made no Debug emitter changes; the later default-enum shortcut above
intentionally benefits both profiles.
Validation passes: 494 library, 696 compiler, 20 binary and four example tests; Clippy; formatting/diff checks; and the browser compiler's wasm32 target check. All 176 maintained modules validate, all 211 runtime scenarios and the Debug/Release profile check pass. The size-corpus runner additionally passes 36 behavioral invocations comparing unoptimized and promoted output.
$env:CARGO_PROFILE_DEV_DEBUG = '0'
$env:CARGO_PROFILE_TEST_DEBUG = '0'
$env:CARGO_INCREMENTAL = '0'
cargo test --lib --test compiler --bin splitc --bin splitls --examples --offline
cargo clippy --all-targets --offline -- -D warnings
cargo check -p splitscript-vscode-wasm --target wasm32-unknown-unknown --offline
cargo test --lib write_size_corpus --offline -- --ignored --nocapture
node scripts/wasm-size-runtime.mjs
cargo test --profile max-opt --lib measure_optimization_overhead --offline -- --ignored --nocaptureArtifacts stay under target/size-check. Timing uses seven warmed, alternating
samples per mode, excludes parsing/type checking, and includes lowering and
emission. The maintained runtime runner compares baseline and optimized output
through the existing behavioral scenarios rather than checking bytes alone.
External real scripts can join these opt-in measurements without becoming
repository fixtures. Set SPLITSCRIPT_SIZE_EXTRA_CORPUS to a JSON manifest of
[artifact_name, source_path] pairs before either measurement command. Names
must be unique ASCII letters/digits/underscores; paths may be absolute or
relative to the repository. For the sibling Celeste porting workspace:
'[["celeste", "../vibe-asl-porting/ports/live_split_celeste_port.split"]]' |
Set-Content target/extra-size-corpus.json
$env:SPLITSCRIPT_SIZE_EXTRA_CORPUS = 'target/extra-size-corpus.json'
cargo test --lib write_size_corpus --offline -- --ignored --nocapture
node scripts/wasm-size-runtime.mjs
cargo test --profile max-opt --lib measure_optimization_overhead --offline -- --ignored --nocaptureThe external modules receive the same size, Wasm validation and Debug-equivalence checks. The runtime runner validates external modules with Node and explicitly reports when there is no maintained behavioral harness; this is not an in-game test. The source files remain untouched. Normal tests and default measurements do not require the external workspace.