Skip to content

feat: compiler and register VM - #199

Open
ascandone wants to merge 35 commits into
mainfrom
feat/register-vm-ir
Open

ascandone wants to merge 35 commits into
mainfrom
feat/register-vm-ir

Conversation

@ascandone

@ascandone ascandone commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

The VM half of the series in one PR (after #189 specs and #190 funds engine): the register VM, its IR, the numscript compiler, the public compile/exec API, the CLI wiring, and the difftest VM leg, so the differential harness validates all of it inside this same PR. Cut from feat/vm's final tree state by path checkout; originally split as #199/#200, folded together since neither half is useful alone.

What this adds

  • internal/vm: the register VM: opcode set, executor, program and vars encode/decode, and an opt-in static verifier for bytecode that didn't just come out of this repo's compiler. Executes on internal/funds, the same engine the tree-walking interpreter uses since feat: replace the funds execution engine with internal/funds #190.
    • Wire format: a NUMB (program) / NVAR (vars) header with a major.minor bytecode version (currently 1.0; a reader accepts its own major and any minor up to its own), followed by tagged sections. Unknown sections are skipped unless they carry the must-understand bit.
    • Verifier (Verify / VerifyWithVars): instruction stream decodes, opcodes known, pool and register indices in range, jumps land on instruction boundaries, flag operands are 0/1, no read before write on any path, the current asset is set before use, and mark regions (oneof) are balanced on every path with no send, save or asset change inside one. Exec does not call it, and keeps its own runtime guards.
  • internal/ir: the IR the compiler lowers to: instruction set, assembly to bytecode with register allocation, a typechecker, and a textual format (grammar in IR.g4).
  • internal/compiler: lowers parsed numscript to that IR, with its own bidirectional checker (internal/typecheck) and the shared builtin names (internal/builtins; analysis' constants now alias them). Fixture tests run the whole interpreter spec corpus through compiler+VM, except asset scaling, which the compiler doesn't lower yet (blacklisted in scripts_test.go), plus a corpus-wide verifier pass.
  • Public API (numscript.go): Compile, program/vars encode-decode, NewVm/ExecVm returning the full ExecutionResult contract (postings with scopes, tx metadata, account metadata rows), opt-in VerifyCompiledProgram(WithVars), and Vm* aliases for the VM's execution error types so callers can classify failures with errors.As.
  • Docs: compiler-architecture.md, instruction-encoding.md (opcode values and operand layouts, kept in sync with instruction.go and the verifier's decoder), ir-textual-format.md, bytecode-verifier.md.
  • The difftest VM leg: every generated script runs against oracle, interpreter and VM. runVM compiles, encodes vars, verifies the bytecode (a verifier failure is an unconditional mismatch) and executes through the public ExecVm, so every generated script also exercises the public contract. Case carries the three pairwise verdicts; the sweep, the fuzzer and every deterministic pin check all three legs. The VM sits on the a-side of each comparison, so the VM rejecting a script another engine ran is a mismatch.

Differential results

3000-seed sweep on the current head: zero mismatches on all three legs.

Leg Tolerated
oracle vs new, oracle vs vm the same sets on both: 250 oracle compile rejections, 17 oracle resolve rejections, 85 negative amount vs missing funds, 535 numscript-only scripts (oracle skipped)
new vs vm 100 vm register capacity
  • VM capacity tolerances: vm register capacity (one-byte register operands) and vm program size (uint16 jump targets). Both are hard bounds of the encoding; the assembler fails closed with a typed error (ir.ErrRegisterBankOverflow / ir.ErrProgramTooLarge), tolerated by name and counted. Any other VM-side rejection is a mismatch.
  • No oracle tolerances on new vs vm: that leg uses CompareEngines, so an interpreter-only compile or resolve rejection, or any of the legacy-machine tolerances, is a mismatch. Only the two vm capacity bounds are tolerated.
  • Deterministic pins: the VM sides with the interpreter on every known numscript-vs-ledger divergence (negative dest/src max, negative overdraft cap, allotment sums, and the uncovered-shape agreements: percent portions, origin vars in cap positions, chained origins, ...).
  • Gate check: flipping one comparison in the compiler's minInt fails the sweep on 1057 script-legs, all attributed to the two VM legs, so the harness detects VM-only corruption and attributes it correctly.

Semantics

The compiler+VM matches the tree-walking interpreter on the spec corpus (minus asset scaling) and the generated corpus, including the recent main-side decisions: a negative destination max clamps to zero, allotments past 100% next to remaining reject at run time (Op_AssertLeftover), as the interpreter does since #198, and a negative allotment clause portion is rejected (Op_AssertNonNegativePortion).

Interpolating an account variable whose origin is another account or scoped(...) (e.g. account $a = @src then @dest:$a) posts to dest:src on both engines. A non-empty scope is rejected at run time with CannotCastScopedAccountToString, as in the interpreter, through a new ASSERT_UNSCOPED instruction (0x0A). The scenarios of internal/compiler/e2e_test.go that specs can express now live as spec fixtures (#209) and run on both engines; the Go tests left assert errors by struct equality.

92 files, ~21.9k lines added, of which ~3.8k are generated ANTLR. Root, internal/oracle and internal/difftest suites green; lint clean.

@NumaryBot NumaryBot added risk: medium NumaryBot classified this pull request as medium risk. bot-reviewed NumaryBot completed its review workflow for the current head. review-inconclusive and removed risk: medium NumaryBot classified this pull request as medium risk. bot-reviewed NumaryBot completed its review workflow for the current head. review-inconclusive labels Sep 25, 2026
@ascandone ascandone changed the title feat: register VM and IR layer feat: compiler and register VM Sep 25, 2026
@NumaryBot NumaryBot added risk: medium NumaryBot classified this pull request as medium risk. bot-reviewed NumaryBot completed its review workflow for the current head. review-inconclusive and removed risk: medium NumaryBot classified this pull request as medium risk. bot-reviewed NumaryBot completed its review workflow for the current head. review-inconclusive labels Sep 25, 2026
* Give the bytecode a major and a minor version number

The bytecode header used to carry one version number, and a reader only
refused bytecode with a higher number than its own. Older bytecode was
accepted even when the meaning of its instructions had changed, which is
exactly what happened between version 1 and version 2.

The header now carries two numbers, major and minor. The rule for a reader:

- same major, and a minor that is not newer than its own: accepted;
- a newer minor: refused. We do not assume the reader knows the new
  instructions, even if it might happen to;
- any other major: refused. Existing instructions may mean something else.

Decisions:

- Each number is 16 bits, so the header grows from 8 to 10 bytes. One byte
  each would probably have been enough, but the header size is the one
  thing that cannot be changed later without also changing the magic word,
  and the minor number moves on every addition. Two extra bytes per program
  is nothing.
- The two numbers are written as two plain fields, major first, rather than
  packed into one. Easier to read in a hex dump, no bit shifting.
- Bytecode written before this change has the old header layout and is
  refused as "different major". Nothing stored depends on it yet, so there
  is no conversion.
- Refusals return one error type carrying both the bytecode's version and
  the version this build supports, so a caller can tell the two apart.
- The "peek" helpers read the version off the header without checking it,
  so a caller can report which version an unreadable file was written with.
  Checking stays the decoder's job.
- The current version is 2.0. Version 1 was the first format; version 2
  changed what several instructions mean, which is what a major bump is for.
- This number describes the bytecode format only. The compiler and the
  library keep their own version numbers, and a new compiler release does
  not by itself mean a new bytecode version.

* Export the bytecode version from the public package

A program that stores compiled bytecode and runs it later, possibly with
another build, needs to check that the two match. Until now the version
lived in an internal package, so a caller could not read it without a
workaround such as encoding an empty value and reading the version back.

The root package now exposes:

- BytecodeVersion, the major/minor pair, with the compatibility rule;
- CurrentBytecodeVersion, what this build writes and the newest it reads;
- PeekCompiledProgramVersion and PeekVarsVersion, to read the version off
  raw bytes without decoding the rest;
- UnsupportedBytecodeVersionError, the error the decoders return.

Decisions:

- The names say "bytecode" rather than "format", to match how the rest of
  the public API talks about compiled programs.
- The peek helpers report the version but do not check it. A caller that
  wants to explain why a file is unreadable needs the number; the decoder
  already does the checking.
- One test drives the whole path through the public API, from Compile to a
  refused program, so the exported surface cannot drift from the internal
  one unnoticed.
A caller that decodes+verifies a compiled program once and reuses it many
times (e.g. an LRU keyed on the compiled bytes) currently has to guess
whether a later Vars is still compatible by reaching into Vars.StringsPool/
IntsPool lengths itself, duplicating this package's notion of "shape" and
silently going stale if that representation ever changes.

VerifyWithVars (and the public VerifyCompiledProgramWithVars alias) now also
returns a VerifiedVarsInfo on success, with unexported fields and a
CheckVars(vars) method: an O(1) check for "would VerifyWithVars(program,
vars) still succeed" without re-running the whole-program static pass or
looking at Vars' internals. A false result isn't a rejection — it just means
the caller should call VerifyWithVars again for an authoritative answer.

All existing call sites are updated for the new two-value return; none of
them use CheckVars yet.
regPool.endInstr ranged over the pendingFree map to push freed slots onto
freeList, which is popped as a stack: whenever several registers died on the
same instruction (any binary op consuming its last operands), the next
allocation reused a slot chosen by map iteration order, so the same script
compiled to different — equivalent — bytecode from run to run. Freed regs are
now pushed in ascending Reg order.

RunState.AccountBalances ranged over the balances map, so both the returned
slice and the order of Store reads (hence which error surfaces when several
reads fail) varied between runs. Entries are now visited in (asset, color)
order.

Both come with regression tests that fail on the previous code. An audit of
the remaining map iterations on the compile and execution paths found no
other order-dependent output: the rest either sort afterwards (VM/interpreter
account metadata, scaling pairs) or feed order-independent sets and maps.
Comment thread internal/vm/ir_test.go Outdated
Comment thread internal/vm/vm.go
Comment thread compiler-architecture.md Outdated

@NumaryBot NumaryBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The required automated review completed with no remaining findings.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bot-reviewed NumaryBot completed its review workflow for the current head. review-approved The NumaryBot review gate is satisfied for the current head. risk: medium NumaryBot classified this pull request as medium risk.

Development

Successfully merging this pull request may close these issues.

3 participants