Skip to content

ssa: describe audited runtime function contracts - #2510

Draft
zhouguangyuan0718 wants to merge 2 commits into
xgo-dev:mainfrom
zhouguangyuan0718:codex/runtime-function-attributes-20260906
Draft

zhouguangyuan0718 wants to merge 2 commits into
xgo-dev:mainfrom
zhouguangyuan0718:codex/runtime-function-attributes-20260906

Conversation

@zhouguangyuan0718

@zhouguangyuan0718 zhouguangyuan0718 commented Sep 5, 2026 •

Copy link
Copy Markdown
Collaborator

Runtime declarations currently omit contracts that their implementations already guarantee, so callers in separate modules cannot eliminate redundant comparisons, reuse read-only results, or exploit nonnegative lengths. Attach a first batch of audited LLVM 22 return, parameter, and function attributes when creating the resolved runtime symbol, covering both definitions and imported/linkname declarations.

  • AssertNilDerefPtr: returned input and nonnull result. The input remains nullable and the panic-producing call remains observable.
  • CStrCopy: returned destination, write-only access, return-only capture, and read/write memory effects that include the aggregate string source.
  • memequal, string comparisons, typed move/clear, MapLen, ChanCap, and hash helpers: precise pointer access/capture and memory effects, plus nofree, nosync, nounwind, and willreturn. Length/capacity results receive a target-width nonnegative range. String aggregates and global hash seeds require effects beyond argument memory.
  • Unconditional panic entries: cold and noreturn, without purity or nounwind promises. Rethrow, conditional checks, throw, and locking ChanLen are excluded.

Preserve overlap and zero-length semantics by leaving copy/comparison pointer arguments nullable and without noalias. Match reflect.typedmemmove after linkname resolution. Allocation contracts and constructor changes are outside this PR.

Use Context.CreateConstantRangeAttribute from xgo-dev/llvm#52. LLGo contains no private cgo bridge for this API. Until the binding change is released, go.mod temporarily replaces github.com/xgo-dev/llvm with the personal branch pinned at 4de18f5844ba28c8f608312e7a1f0606ade54d41 (v0.0.0-20260905230337-4de18f5844ba); both module checksums are recorded in go.sum. Replace this temporary pin with the upstream binding release when available.

Validation on macOS arm64 with Go 1.27.0 and LLVM/Clang 22.1.8:

  • go test ./ssa ./cl -count=1 passed against the pinned binding commit.
  • Declaration/definition attribute checks for amd64, arm64, 386, and wasm64, including nullable inputs and excluded functions.
  • LLVM verifier, O2/LTO optimization, and assembly emission for amd64, 386, and wasm64: preserve the nil-check call while folding its result, fold returned-copy identity and nonnegative-range checks, eliminate redundant reads, and retain reads across argument/global writes.
  • llgo test -v ./internal/test and llgo test -tags=nogc -v ./internal/test passed from runtime/, covering recoverable nil/panic behavior, overlapping moves through both runtime and reflect entries, zero-byte equality, typed clearing, C strings, and map/channel queries.

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: runtime function attributes

This PR attaches LLVM function attributes (memory effects, capture info, range, noreturn/cold, nonnull/returned) to a fixed set of runtime symbols so the optimizer can reason about them across go:linkname boundaries. The design is careful and well-tested (multi-target IR checks, O2/LTO fold verification, linkname-aliasing coverage).

Verified sound:

  • Parameter-index → LLVM-index mapping lines up with Go param positions (no sret/byval expansion for these signatures).
  • Linkname resolution happens in rtFunc before NewFunc, so the contract map is correctly keyed by resolved symbol names; package-qualified keys prevent cross-package collisions.
  • The bit constants (runtimeMemoryRead=0x555, runtimeArgRead/ReadWrite, runtimeCaptureReturn=0xf) correctly encode LLVM's MemoryEffects/CaptureInfo layouts, and all accompanying comments are accurate.
  • The range cgo bridge is ABI-sound; the bits!=32&&bits!=64 guard prevents shift UB.
  • No security or performance concerns (the per-function cost is a single map miss with an early return before any closure allocation).

The findings below are hardening/maintainability suggestions, not active defects in the current tested configuration.

Comment thread ssa/runtime_attrs.go
runtimeCaptureReturn = 0xf
)

func (p Program) addRuntimeAttributes(fn llvm.Value, name string) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hardening: no signature/arity validation before applying per-parameter pointer attributes.

addRuntimeAttributes keys purely on the resolved name and then blindly applies pointer(2, "writeonly"), pointer(3, "readonly"), etc. The reflect.typedmemmove entry is a bare pkg.name reachable by the real reflect package; today its signature matches the move contract, but nothing enforces that. If a contract-named symbol ever has a different shape (signature drift, or a future name collision), pointer attributes would be attached to a non-pointer parameter or an out-of-range index, producing a silent miscompile that is very hard to trace.

Consider asserting parameter count and that the targeted indices are pointer-typed before adding pointer attributes (mirroring the width guard already in addNonNegativeReturnRange). Cheap, and it turns a latent silent-miscompile into a loud panic.

Comment thread ssa/runtime_range.go Outdated

func addNonNegativeReturnRange(ctx llvm.Context, fn llvm.Value, bits int) {
if bits != 32 && bits != 64 {
panic("ssa: unsupported runtime integer width")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Minor: the panic message "ssa: unsupported runtime integer width" doesn't identify the offending function. If a mismatched symbol reaches the length contract (see the related note in runtime_attrs.go), the failure surfaces here without pointing at which function caused it. Including the function name would make the failure actionable.

Comment thread ssa/runtime_attrs.go
if !ok {
return
}
add := func(index int, name string, value uint64) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Minor readability: the closure parameter name string shadows the enclosing function's name argument (the symbol name). Harmless since the outer name isn't used after the map lookup, but shadowing an attribute-kind string over a symbol-name string is easy to misread. Renaming the closure param to attr/kind would remove the ambiguity.

@codecov

codecov Bot commented Sep 5, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.61905% with 1 line in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
ssa/runtime_attrs.go 97.56% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@github-actions

github-actions Bot commented Sep 5, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

4980c24d45ce | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 19848 B 0 B / +0.0% 387 B 0 B / +0.0% 439.196 ms -17.49 ms / -3.8% (better) 1.257 ms -29.24 us / -2.3% (better)
Linux cprintf-lto 19600 B 0 B / +0.0% 368 B 0 B / +0.0% 450.760 ms +2.284 ms / +0.5% (worse) 1.236 ms -15.39 us / -1.2% (better)
Linux fmtprintf 1612872 B -4136 B / -0.3% (better) 488657 B -4541 B / -0.9% (better) 3.624 s +38.5 ms / +1.1% (worse) 3.183 ms +58.13 us / +1.9% (worse)
Linux fmtprintf-lto 1476976 B -1768 B / -0.1% (better) 440698 B -2281 B / -0.5% (better) 10.331 s +10.57 ms / +0.1% (worse) 2.964 ms +93.46 us / +3.3% (worse)
Linux println 62728 B +128 B / +0.2% (worse) 14557 B -415 B / -2.8% (better) 444.733 ms +3.795 ms / +0.9% (worse) 1.553 ms +17.38 us / +1.1% (worse)
Linux println-lto 54568 B -88 B / -0.2% (better) 12202 B -335 B / -2.7% (better) 652.753 ms -6.821 ms / -1.0% (better) 1.542 ms -21.08 us / -1.3% (better)
macOS cprintf 84480 B 0 B / +0.0% 17117 B 0 B / +0.0% 680.713 ms +20.08 ms / +3.0% (worse) 3.792 ms +231.1 us / +6.5% (worse)
macOS cprintf-lto 84288 B 0 B / +0.0% 12881 B 0 B / +0.0% 790.106 ms -64.43 ms / -7.5% (better) 8.305 ms +3.118 ms / +60.1% (worse)
macOS fmtprintf 1457328 B -16080 B / -1.1% (better) 860184 B -6976 B / -0.8% (better) 3.747 s +61.98 ms / +1.7% (worse) 5.395 ms +214 us / +4.1% (worse)
macOS fmtprintf-lto 1159568 B 0 B / +0.0% 844868 B -3612 B / -0.4% (better) 9.802 s +44.37 ms / +0.5% (worse) 4.832 ms +343.5 us / +7.7% (worse)
macOS println 115120 B +320 B / +0.3% (worse) 34373 B -556 B / -1.6% (better) 852.344 ms +227 ms / +36.3% (worse) 7.141 ms +3.275 ms / +84.7% (worse)
macOS println-lto 118736 B 0 B / +0.0% 32577 B -304 B / -0.9% (better) 923.342 ms -7.964 ms / -0.9% (better) 3.882 ms -95.38 us / -2.4% (better)
Windows MinGW cprintf 19456 B 0 B / +0.0% 4550 B 0 B / +0.0% 1.191 s +36.12 ms / +3.1% (worse) 4.008 ms +613.9 us / +18.1% (worse)
Windows MinGW cprintf-lto 17920 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.350 s +158.5 ms / +13.3% (worse) 4.533 ms +1.126 ms / +33.1% (worse)
Windows MinGW fmtprintf 1882112 B -4096 B / -0.2% (better) 589030 B -4736 B / -0.8% (better) 4.185 s +36.15 ms / +0.9% (worse) 8.990 ms +697 us / +8.4% (worse)
Windows MinGW fmtprintf-lto 1936384 B -3072 B / -0.2% (better) 552566 B -4752 B / -0.9% (better) 10.196 s +125.8 ms / +1.2% (worse) 8.918 ms +706.9 us / +8.6% (worse)
Windows MinGW println 72704 B +1024 B / +1.4% (worse) 23782 B -256 B / -1.1% (better) 1.315 s +141.2 ms / +12.0% (worse) 6.692 ms -191 us / -2.8% (better)
Windows MinGW println-lto 66048 B 0 B / +0.0% 20806 B -208 B / -1.0% (better) 1.508 s +152.8 ms / +11.3% (worse) 6.901 ms +219.4 us / +3.3% (worse)
Windows MinGW 386 cprintf 37888 B 0 B / +0.0% 5326 B 0 B / +0.0% 1.059 s -31.33 ms / -2.9% (better) 5.213 ms +40.3 us / +0.8% (worse)
Windows MinGW 386 cprintf-lto 20992 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.105 s +1.586 ms / +0.1% (worse) 5.114 ms -594.8 us / -10.4% (better)
Windows MinGW 386 fmtprintf 1839104 B -4096 B / -0.2% (better) 463834 B -4560 B / -1.0% (better) 4.045 s +121.7 ms / +3.1% (worse) 11.111 ms +727 us / +7.0% (worse)
Windows MinGW 386 fmtprintf-lto 2167296 B +1024 B / +0.04727% (worse) 459998 B -3416 B / -0.7% (better) 9.317 s -171.8 ms / -1.8% (better) 10.244 ms -1.916 ms / -15.8% (better)
Windows MinGW 386 println 88064 B +1536 B / +1.8% (worse) 20070 B -384 B / -1.9% (better) 1.070 s +16.57 ms / +1.6% (worse) 8.375 ms +59.6 us / +0.7% (worse)
Windows MinGW 386 println-lto 72192 B +1024 B / +1.4% (worse) 18038 B -356 B / -1.9% (better) 1.284 s +14.94 ms / +1.2% (worse) 8.451 ms -61.2 us / -0.7% (better)
Windows MinGW ARM64 cprintf 18944 B 0 B / +0.0% 4436 B 0 B / +0.0% 1.463 s +11.11 ms / +0.8% (worse) 6.516 ms -20 us / -0.3% (better)
Windows MinGW ARM64 cprintf-lto 17920 B 0 B / +0.0% 4368 B 0 B / +0.0% 1.503 s +21.46 ms / +1.4% (worse) 6.978 ms +136.3 us / +2.0% (worse)
Windows MinGW ARM64 fmtprintf 1768960 B -5120 B / -0.3% (better) 500840 B -5672 B / -1.1% (better) 4.330 s +109.4 ms / +2.6% (worse) 13.835 ms +34.6 us / +0.3% (worse)
Windows MinGW ARM64 fmtprintf-lto 1864704 B -2048 B / -0.1% (better) 483252 B -3856 B / -0.8% (better) 9.802 s -139 ms / -1.4% (better) 13.498 ms +222.6 us / +1.7% (worse)
Windows MinGW ARM64 println 69120 B 0 B / +0.0% 22368 B -320 B / -1.4% (better) 1.456 s +7.069 ms / +0.5% (worse) 11.822 ms +512.7 us / +4.5% (worse)
Windows MinGW ARM64 println-lto 65536 B +512 B / +0.8% (worse) 20088 B -172 B / -0.8% (better) 1.666 s +9.31 ms / +0.6% (worse) 11.930 ms +628.5 us / +5.6% (worse)
Windows MSVC cprintf 120320 B 0 B / +0.0% 65782 B 0 B / +0.0% 853.200 ms -61.91 ms / -6.8% (better) 3.303 ms -61.5 us / -1.8% (better)
Windows MSVC cprintf-lto 119808 B 0 B / +0.0% 65718 B 0 B / +0.0% 880.950 ms -53.12 ms / -5.7% (better) 3.450 ms +128.4 us / +3.9% (worse)
Windows MSVC fmtprintf 1611264 B -5632 B / -0.3% (better) 684534 B -4736 B / -0.7% (better) 3.567 s +32.57 ms / +0.9% (worse) 8.649 ms -57.2 us / -0.7% (better)
Windows MSVC fmtprintf-lto 1624064 B -3072 B / -0.2% (better) 655014 B -4688 B / -0.7% (better) 8.676 s +93.42 ms / +1.1% (worse) 8.904 ms +239.4 us / +2.8% (worse)
Windows MSVC println 193536 B +512 B / +0.3% (worse) 119190 B -256 B / -0.2% (better) 867.746 ms -31.24 ms / -3.5% (better) 7.169 ms +102.9 us / +1.5% (worse)
Windows MSVC println-lto 189952 B -512 B / -0.3% (better) 116614 B -208 B / -0.2% (better) 1.051 s -6.259 ms / -0.6% (better) 7.060 ms -44.6 us / -0.6% (better)
Windows MSVC 386 cprintf 9728 B 0 B / +0.0% 3931 B 0 B / +0.0% 1.124 s +176.9 ms / +18.7% (worse) 7.003 ms +1.489 ms / +27.0% (worse)
Windows MSVC 386 cprintf-lto 9216 B 0 B / +0.0% 3853 B 0 B / +0.0% 1.132 s +10.15 ms / +0.9% (worse) 6.572 ms +553.8 us / +9.2% (worse)
Windows MSVC 386 fmtprintf 1182208 B -4096 B / -0.3% (better) 447200 B -4561 B / -1.0% (better) 3.770 s +79.65 ms / +2.2% (worse) 12.313 ms +399.3 us / +3.4% (worse)
Windows MSVC 386 fmtprintf-lto 1239040 B -2560 B / -0.2% (better) 440209 B -3316 B / -0.7% (better) 8.541 s -37.67 ms / -0.4% (better) 12.146 ms +1.087 ms / +9.8% (worse)
Windows MSVC 386 println 34816 B 0 B / +0.0% 18916 B -397 B / -2.1% (better) 1.091 s +150.2 ms / +16.0% (worse) 11.562 ms +1.859 ms / +19.2% (worse)
Windows MSVC 386 println-lto 33280 B -512 B / -1.5% (better) 17116 B -347 B / -2.0% (better) 1.295 s +189.2 ms / +17.1% (worse) 10.700 ms +702.7 us / +7.0% (worse)
Windows MSVC ARM64 cprintf 11264 B 0 B / +0.0% 3976 B 0 B / +0.0% 2.231 s +150.8 ms / +7.3% (worse) 8.376 ms +1.405 ms / +20.1% (worse)
Windows MSVC ARM64 cprintf-lto 10752 B 0 B / +0.0% 3868 B 0 B / +0.0% 2.245 s +45.99 ms / +2.1% (worse) 8.119 ms +929.5 us / +12.9% (worse)
Windows MSVC ARM64 fmtprintf 1358336 B -4608 B / -0.3% (better) 501340 B -4992 B / -1.0% (better) 6.657 s +46.26 ms / +0.7% (worse) 16.260 ms -1.112 ms / -6.4% (better)
Windows MSVC ARM64 fmtprintf-lto 1396224 B -2560 B / -0.2% (better) 483960 B -3908 B / -0.8% (better) 16.043 s -241.3 ms / -1.5% (better) 16.238 ms -625.4 us / -3.7% (better)
Windows MSVC ARM64 println 42496 B 0 B / +0.0% 22024 B -240 B / -1.1% (better) 2.216 s +154 ms / +7.5% (worse) 14.279 ms +394.4 us / +2.8% (worse)
Windows MSVC ARM64 println-lto 40960 B +512 B / +1.3% (worse) 19984 B -172 B / -0.9% (better) 2.584 s +201.3 ms / +8.4% (worse) 14.729 ms +1.917 ms / +15.0% (worse)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.510 ns/op -0.09 ns/op / -0.6% (better)
Linux BenchmarkMergeCompilerFlags 213.600 ns/op +15.7 ns/op / +7.9% (worse)
Linux BenchmarkMergeLinkerFlags 136.600 ns/op +9.4 ns/op / +7.4% (worse)
Linux BenchmarkChannelBuffered 69.790 ns/op -0.29 ns/op / -0.4% (better)
Linux BenchmarkChannelHandoff 14618 ns/op -2800 ns/op / -16.1% (better)
Linux BenchmarkDefer 46.620 ns/op -0.43 ns/op / -0.9% (better)
Linux BenchmarkDirectCall 1.164 ns/op -0.004 ns/op / -0.3% (better)
Linux BenchmarkGlobalRead 1.552 ns/op +0.386 ns/op / +33.1% (worse)
Linux BenchmarkGlobalWrite 7.758 ns/op -0.015 ns/op / -0.2% (better)
Linux BenchmarkGoroutine 28024 ns/op +6897 ns/op / +32.6% (worse)
Linux BenchmarkInterfaceCall 7.767 ns/op +1.16 ns/op / +17.6% (worse)
Linux BenchmarkRuntimeGetG 2.378 ns/op -0.066 ns/op / -2.7% (better)
macOS BenchmarkLookupPCRandom 16.080 ns/op +0.29 ns/op / +1.8% (worse)
macOS BenchmarkMergeCompilerFlags 143.700 ns/op +26 ns/op / +22.1% (worse)
macOS BenchmarkMergeLinkerFlags 92.690 ns/op +16.18 ns/op / +21.1% (worse)
macOS BenchmarkChannelBuffered 32 ns/op -1.08 ns/op / -3.3% (better)
macOS BenchmarkChannelHandoff 10256 ns/op -2091 ns/op / -16.9% (better)
macOS BenchmarkDefer 37.830 ns/op -4.28 ns/op / -10.2% (better)
macOS BenchmarkDirectCall 1.120 ns/op -0.125 ns/op / -10.0% (better)
macOS BenchmarkGlobalRead 1.199 ns/op -0.42 ns/op / -25.9% (better)
macOS BenchmarkGlobalWrite 1.269 ns/op -0.296 ns/op / -18.9% (better)
macOS BenchmarkGoroutine 51118 ns/op +15427 ns/op / +43.2% (worse)
macOS BenchmarkInterfaceCall 4.793 ns/op -0.945 ns/op / -16.5% (better)
macOS BenchmarkRuntimeGetG 2.546 ns/op -0.008 ns/op / -0.3% (better)
Windows MinGW BenchmarkLookupPCRandom 12.920 ns/op -0.07 ns/op / -0.5% (better)
Windows MinGW BenchmarkMergeCompilerFlags 623.200 ns/op +4.1 ns/op / +0.7% (worse)
Windows MinGW BenchmarkMergeLinkerFlags 540.700 ns/op +5.7 ns/op / +1.1% (worse)
Windows MinGW BenchmarkChannelBuffered 38.460 ns/op +1.38 ns/op / +3.7% (worse)
Windows MinGW BenchmarkChannelHandoff 916.500 ns/op -24.3 ns/op / -2.6% (better)
Windows MinGW BenchmarkDefer 58.610 ns/op +0.97 ns/op / +1.7% (worse)
Windows MinGW BenchmarkDirectCall 1.546 ns/op -0.313 ns/op / -16.8% (better)
Windows MinGW BenchmarkGlobalRead 1.857 ns/op -0.002 ns/op / -0.1% (better)
Windows MinGW BenchmarkGlobalWrite 2.458 ns/op +0.008 ns/op / +0.3% (worse)
Windows MinGW BenchmarkGoroutine 79556 ns/op -2595 ns/op / -3.2% (better)
Windows MinGW BenchmarkInterfaceCall 9.611 ns/op +0.305 ns/op / +3.3% (worse)
Windows MinGW BenchmarkRuntimeGetG 2.172 ns/op +0.311 ns/op / +16.7% (worse)
Windows MinGW 386 BenchmarkLookupPCRandom 26.460 ns/op -2.84 ns/op / -9.7% (better)
Windows MinGW 386 BenchmarkMergeCompilerFlags 717.200 ns/op -159.8 ns/op / -18.2% (better)
Windows MinGW 386 BenchmarkMergeLinkerFlags 635.400 ns/op -155.4 ns/op / -19.7% (better)
Windows MinGW 386 BenchmarkChannelBuffered 43.010 ns/op -0.04 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkChannelHandoff 980.200 ns/op -77.8 ns/op / -7.4% (better)
Windows MinGW 386 BenchmarkDefer 44.990 ns/op +3.84 ns/op / +9.3% (worse)
Windows MinGW 386 BenchmarkDirectCall 1.548 ns/op 0 ns/op / +0.0%
Windows MinGW 386 BenchmarkGlobalRead 1.547 ns/op -0.006 ns/op / -0.4% (better)
Windows MinGW 386 BenchmarkGlobalWrite 7.770 ns/op -0.011 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGoroutine 86497 ns/op +2587 ns/op / +3.1% (worse)
Windows MinGW 386 BenchmarkInterfaceCall 9.501 ns/op -0.103 ns/op / -1.1% (better)
Windows MinGW 386 BenchmarkRuntimeGetG 2.170 ns/op +0.311 ns/op / +16.7% (worse)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.080 ns/op +0.01 ns/op / +0.1% (worse)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 574.700 ns/op +11.2 ns/op / +2.0% (worse)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 536 ns/op -5.8 ns/op / -1.1% (better)
Windows MinGW ARM64 BenchmarkChannelBuffered 43.980 ns/op +0.1 ns/op / +0.2% (worse)
Windows MinGW ARM64 BenchmarkChannelHandoff 2435 ns/op +216 ns/op / +9.7% (worse)
Windows MinGW ARM64 BenchmarkDefer 56.220 ns/op -0.24 ns/op / -0.4% (better)
Windows MinGW ARM64 BenchmarkDirectCall 0.589 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkGlobalRead 0.884 ns/op +0.219 ns/op / +32.9% (worse)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.663 ns/op +0.0737 ns/op / +12.5% (worse)
Windows MinGW ARM64 BenchmarkGoroutine 57116 ns/op -4031 ns/op / -6.6% (better)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.718 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.769 ns/op +0.001 ns/op / +0.1% (worse)
Windows MSVC BenchmarkLookupPCRandom 12.590 ns/op +0.2 ns/op / +1.6% (worse)
Windows MSVC BenchmarkMergeCompilerFlags 536.600 ns/op -4.4 ns/op / -0.8% (better)
Windows MSVC BenchmarkMergeLinkerFlags 460.300 ns/op -4 ns/op / -0.9% (better)
Windows MSVC BenchmarkChannelBuffered 39.580 ns/op -0.08 ns/op / -0.2% (better)
Windows MSVC BenchmarkChannelHandoff 1305 ns/op +127 ns/op / +10.8% (worse)
Windows MSVC BenchmarkDefer 54.410 ns/op +0.29 ns/op / +0.5% (worse)
Windows MSVC BenchmarkDirectCall 1.745 ns/op -0.002 ns/op / -0.1% (better)
Windows MSVC BenchmarkGlobalRead 1.748 ns/op +0.001 ns/op / +0.1% (worse)
Windows MSVC BenchmarkGlobalWrite 2.788 ns/op 0 ns/op / +0.0%
Windows MSVC BenchmarkGoroutine 63238 ns/op -1450 ns/op / -2.2% (better)
Windows MSVC BenchmarkInterfaceCall 10.140 ns/op +0.343 ns/op / +3.5% (worse)
Windows MSVC BenchmarkRuntimeGetG 1.813 ns/op 0 ns/op / +0.0%
Windows MSVC 386 BenchmarkLookupPCRandom 26.560 ns/op +0.06 ns/op / +0.2% (worse)
Windows MSVC 386 BenchmarkMergeCompilerFlags 743.700 ns/op +27 ns/op / +3.8% (worse)
Windows MSVC 386 BenchmarkMergeLinkerFlags 689.200 ns/op +18.8 ns/op / +2.8% (worse)
Windows MSVC 386 BenchmarkChannelBuffered 43.530 ns/op -5.24 ns/op / -10.7% (better)
Windows MSVC 386 BenchmarkChannelHandoff 949.100 ns/op -26.5 ns/op / -2.7% (better)
Windows MSVC 386 BenchmarkDefer 44.890 ns/op -7.72 ns/op / -14.7% (better)
Windows MSVC 386 BenchmarkDirectCall 1.548 ns/op -0.073 ns/op / -4.5% (better)
Windows MSVC 386 BenchmarkGlobalRead 2.172 ns/op +0.049 ns/op / +2.3% (worse)
Windows MSVC 386 BenchmarkGlobalWrite 7.785 ns/op -0.193 ns/op / -2.4% (better)
Windows MSVC 386 BenchmarkGoroutine 86864 ns/op -2469 ns/op / -2.8% (better)
Windows MSVC 386 BenchmarkInterfaceCall 9.596 ns/op -0.944 ns/op / -9.0% (better)
Windows MSVC 386 BenchmarkRuntimeGetG 2.169 ns/op -0.455 ns/op / -17.3% (better)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.100 ns/op +0.04 ns/op / +0.3% (worse)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 577.500 ns/op +2.3 ns/op / +0.4% (worse)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 547.400 ns/op +5.1 ns/op / +0.9% (worse)
Windows MSVC ARM64 BenchmarkChannelBuffered 45.920 ns/op -0.26 ns/op / -0.6% (better)
Windows MSVC ARM64 BenchmarkChannelHandoff 2659 ns/op -48 ns/op / -1.8% (better)
Windows MSVC ARM64 BenchmarkDefer 61.820 ns/op -3.37 ns/op / -5.2% (better)
Windows MSVC ARM64 BenchmarkDirectCall 0.589 ns/op -0.0005 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkGlobalRead 0.592 ns/op -0.0717 ns/op / -10.8% (better)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.795 ns/op -0.002 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkGoroutine 57471 ns/op -870 ns/op / -1.5% (better)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.727 ns/op -0.005 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.799 ns/op +0.029 ns/op / +1.6% (worse)

Timer runtime benchmarks

Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 901.100 ns/op -2.9 ns/op / -0.3% (better)
Linux AfterFuncZeroDelivery/LLGo 39744 ns/op +4540 ns/op / +12.9% (worse)
Linux CreateStop/Go 289.200 ns/op -0.5 ns/op / -0.2% (better)
Linux CreateStop/LLGo 1938 ns/op +330 ns/op / +20.5% (worse)
Linux RearmStopped/Go 114.900 ns/op -0.8 ns/op / -0.7% (better)
Linux RearmStopped/LLGo 1394 ns/op -61 ns/op / -4.2% (better)
Linux ResetActive/Go 67.390 ns/op -1.2 ns/op / -1.7% (better)
Linux ResetActive/LLGo 712.700 ns/op -1.3 ns/op / -0.2% (better)
Linux ResetHeap1024/Go 67.140 ns/op -0.09 ns/op / -0.1% (better)
Linux ResetHeap1024/LLGo 192.100 ns/op +2.6 ns/op / +1.4% (worse)
macOS AfterFuncZeroDelivery/Go 561 ns/op +68.9 ns/op / +14.0% (worse)
macOS AfterFuncZeroDelivery/LLGo 98451 ns/op +25019 ns/op / +34.1% (worse)
macOS CreateStop/Go 189.200 ns/op +34 ns/op / +21.9% (worse)
macOS CreateStop/LLGo 895.100 ns/op +414 ns/op / +86.1% (worse)
macOS RearmStopped/Go 72.660 ns/op +8.2 ns/op / +12.7% (worse)
macOS RearmStopped/LLGo 630.500 ns/op +273.6 ns/op / +76.7% (worse)
macOS ResetActive/Go 53.340 ns/op +5.76 ns/op / +12.1% (worse)
macOS ResetActive/LLGo 230.900 ns/op +73.7 ns/op / +46.9% (worse)
macOS ResetHeap1024/Go 52.680 ns/op +5.29 ns/op / +11.2% (worse)
macOS ResetHeap1024/LLGo 101.100 ns/op +1.6 ns/op / +1.6% (worse)
Windows MinGW AfterFuncZeroDelivery/Go 559.800 ns/op -1 ns/op / -0.2% (better)
Windows MinGW AfterFuncZeroDelivery/LLGo 162256 ns/op -1049 ns/op / -0.6% (better)
Windows MinGW CreateStop/Go 113.400 ns/op -3.2 ns/op / -2.7% (better)
Windows MinGW CreateStop/LLGo 452.700 ns/op +16.2 ns/op / +3.7% (worse)
Windows MinGW RearmStopped/Go 31.460 ns/op +0.21 ns/op / +0.7% (worse)
Windows MinGW RearmStopped/LLGo 286.500 ns/op -26 ns/op / -8.3% (better)
Windows MinGW ResetActive/Go 20.070 ns/op +0.03 ns/op / +0.1% (worse)
Windows MinGW ResetActive/LLGo 165.200 ns/op -8 ns/op / -4.6% (better)
Windows MinGW ResetHeap1024/Go 20.540 ns/op -0.11 ns/op / -0.5% (better)
Windows MinGW ResetHeap1024/LLGo 134.300 ns/op -7.5 ns/op / -5.3% (better)
Windows MinGW 386 AfterFuncZeroDelivery/Go 950.500 ns/op +10.2 ns/op / +1.1% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 179089 ns/op +1993 ns/op / +1.1% (worse)
Windows MinGW 386 CreateStop/Go 192 ns/op -17.3 ns/op / -8.3% (better)
Windows MinGW 386 CreateStop/LLGo 2071 ns/op -113 ns/op / -5.2% (better)
Windows MinGW 386 RearmStopped/Go 63.180 ns/op -0.49 ns/op / -0.8% (better)
Windows MinGW 386 RearmStopped/LLGo 363.300 ns/op -9.1 ns/op / -2.4% (better)
Windows MinGW 386 ResetActive/Go 38.870 ns/op -0.26 ns/op / -0.7% (better)
Windows MinGW 386 ResetActive/LLGo 978.800 ns/op +8.4 ns/op / +0.9% (worse)
Windows MinGW 386 ResetHeap1024/Go 39.320 ns/op -0.04 ns/op / -0.1% (better)
Windows MinGW 386 ResetHeap1024/LLGo 196.700 ns/op +3.2 ns/op / +1.7% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 662.500 ns/op -2.3 ns/op / -0.3% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 134576 ns/op -1641 ns/op / -1.2% (better)
Windows MinGW ARM64 CreateStop/Go 200.400 ns/op +1.6 ns/op / +0.8% (worse)
Windows MinGW ARM64 CreateStop/LLGo 385.100 ns/op -27.8 ns/op / -6.7% (better)
Windows MinGW ARM64 RearmStopped/Go 70.570 ns/op +0.04 ns/op / +0.1% (worse)
Windows MinGW ARM64 RearmStopped/LLGo 278 ns/op -1.5 ns/op / -0.5% (better)
Windows MinGW ARM64 ResetActive/Go 31.160 ns/op +0.02 ns/op / +0.1% (worse)
Windows MinGW ARM64 ResetActive/LLGo 124.700 ns/op +1 ns/op / +0.8% (worse)
Windows MinGW ARM64 ResetHeap1024/Go 31.150 ns/op +0.01 ns/op / +0.03211% (worse)
Windows MinGW ARM64 ResetHeap1024/LLGo 136.200 ns/op -2.2 ns/op / -1.6% (better)
Windows MSVC AfterFuncZeroDelivery/Go 514.400 ns/op +24.9 ns/op / +5.1% (worse)
Windows MSVC AfterFuncZeroDelivery/LLGo 121379 ns/op -953 ns/op / -0.8% (better)
Windows MSVC CreateStop/Go 117.300 ns/op +1.8 ns/op / +1.6% (worse)
Windows MSVC CreateStop/LLGo 448.200 ns/op -40.5 ns/op / -8.3% (better)
Windows MSVC RearmStopped/Go 31.550 ns/op -0.12 ns/op / -0.4% (better)
Windows MSVC RearmStopped/LLGo 298.400 ns/op +2 ns/op / +0.7% (worse)
Windows MSVC ResetActive/Go 19.140 ns/op +0.02 ns/op / +0.1% (worse)
Windows MSVC ResetActive/LLGo 147.100 ns/op -0.5 ns/op / -0.3% (better)
Windows MSVC ResetHeap1024/Go 19.260 ns/op +0.1 ns/op / +0.5% (worse)
Windows MSVC ResetHeap1024/LLGo 145.600 ns/op -3.3 ns/op / -2.2% (better)
Windows MSVC 386 AfterFuncZeroDelivery/Go 959 ns/op +7.8 ns/op / +0.8% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 194256 ns/op +4834 ns/op / +2.6% (worse)
Windows MSVC 386 CreateStop/Go 191 ns/op -0.6 ns/op / -0.3% (better)
Windows MSVC 386 CreateStop/LLGo 1556 ns/op -60 ns/op / -3.7% (better)
Windows MSVC 386 RearmStopped/Go 62.970 ns/op -0.23 ns/op / -0.4% (better)
Windows MSVC 386 RearmStopped/LLGo 329.900 ns/op -10.7 ns/op / -3.1% (better)
Windows MSVC 386 ResetActive/Go 39.140 ns/op +0.09 ns/op / +0.2% (worse)
Windows MSVC 386 ResetActive/LLGo 969.800 ns/op +682.5 ns/op / +237.6% (worse)
Windows MSVC 386 ResetHeap1024/Go 39.380 ns/op +0.07 ns/op / +0.2% (worse)
Windows MSVC 386 ResetHeap1024/LLGo 176.900 ns/op +0.5 ns/op / +0.3% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 670.500 ns/op +6.3 ns/op / +0.9% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 143679 ns/op +10685 ns/op / +8.0% (worse)
Windows MSVC ARM64 CreateStop/Go 197.400 ns/op +0.3 ns/op / +0.2% (worse)
Windows MSVC ARM64 CreateStop/LLGo 406.200 ns/op -28 ns/op / -6.4% (better)
Windows MSVC ARM64 RearmStopped/Go 70.550 ns/op +0.03 ns/op / +0.04254% (worse)
Windows MSVC ARM64 RearmStopped/LLGo 300.200 ns/op -4.3 ns/op / -1.4% (better)
Windows MSVC ARM64 ResetActive/Go 31.060 ns/op -0.08 ns/op / -0.3% (better)
Windows MSVC ARM64 ResetActive/LLGo 131.100 ns/op -9.1 ns/op / -6.5% (better)
Windows MSVC ARM64 ResetHeap1024/Go 31.110 ns/op -0.08 ns/op / -0.3% (better)
Windows MSVC ARM64 ResetHeap1024/LLGo 141.200 ns/op -2.3 ns/op / -1.6% (better)

Compared with bf3071fdd386 measured in the same runner job.

@github-actions

github-actions Bot commented Sep 5, 2026

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

b094e1e9de03 | workflow run | long-term charts

WebAssembly output sizes

Profile and compiler Wasm module vs base Generated JS glue vs base
ec32/LLGo 70937 B -81 B / -0.1% (better) 70588 B 0 B / +0.0%
ec64/LLGo 75319 B +45 B / +0.1% (worse) 73884 B 0 B / +0.0%
js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
js/LLGo 65030 B -482 B / -0.7% (better) 68499 B 0 B / +0.0%
wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
wasip1/LLGo 71287 B -497 B / -0.7% (better) 0 B 0 B / 0.0%
wc32/LLGo 74990 B -98 B / -0.1% (better) 0 B 0 B / 0.0%

LLGo WebAssembly build measurements

Profile Build vs base
ec32 4.195 s -1.023 ms / -0.02437% (better)
ec64 3.955 s +9.304 ms / +0.2% (worse)
js 4.219 s -44.28 ms / -1.0% (better)
wasip1 2.911 s -53.56 ms / -1.8% (better)
wc32 2.848 s +41.98 ms / +1.5% (worse)

Compared with bf3071fdd386 measured in the same runner job.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant