diff --git a/docs/CLI_DEMOS_README.md b/docs/CLI_DEMOS_README.md index 897b504458..a8cf9b9888 100644 --- a/docs/CLI_DEMOS_README.md +++ b/docs/CLI_DEMOS_README.md @@ -68,7 +68,7 @@ using hyperdimensional computing mathematics. ![tri-benchmark](https://gHashTag.github.io/trinity/recordings/tri-benchmark.gif) -**63 tok/s** — Addresses performance objections about Trinity's speed. +~~**63 tok/s**~~ (withdrawn: that figure was the FPGA LLM's projection at a 92 MHz Fmax estimate, not a measurement, and `tri benchmark` prints no tok/s; see the FPGA table in the root README.md) — Addresses performance objections about Trinity's speed. **Results:** - VSA operations: 17x+ speedup via SIMD @@ -178,7 +178,7 @@ trinity/ | Objection | Response | GIF Demo | |-----------|-----------|-----------| -| "Where is benchmark?" | `tri benchmark` → 63 tok/s | ![tri-benchmark](https://gHashTag.github.io/trinity/recordings/tri-benchmark.gif) | +| "Where is benchmark?" | `tri benchmark` (it prints no tok/s; the ~~63 tok/s~~ once quoted here was the FPGA LLM's projection, withdrawn; see the FPGA table in the root README.md) | ![tri-benchmark](https://gHashTag.github.io/trinity/recordings/tri-benchmark.gif) | | "Where are tests?" | `tri test` → 74/74 passing | ![tri-test](https://gHashTag.github.io/trinity/recordings/tri-test.gif) | | "Where is reproducibility?" | `git clone → zig build → tri benchmark` | All GIFs reproducible | | "FPGA uses DSP?" | 0% DSP in bitstream | ![tri-fpga-synth](https://gHashTag.github.io/trinity/recordings/tri-fpga-synth.gif) | diff --git a/docs/docs/research/fpga-autoregressive-llm-report.md b/docs/docs/research/fpga-autoregressive-llm-report.md index 4e7bb4d9b6..62f376376d 100644 --- a/docs/docs/research/fpga-autoregressive-llm-report.md +++ b/docs/docs/research/fpga-autoregressive-llm-report.md @@ -26,14 +26,14 @@ First autoregressive ternary language model running on an FPGA with a fully open | Metric | Value | |--------|-------| | Board | QMTech XC7A100T-1FGG676C ($30) | -| Power | ~1W | +| Power | not measured (the ~1W figure is withdrawn: no instrument, rail or method is on record; see the FPGA table in the root README.md) | | Toolchain | openXC7 (yosys + nextpnr-xilinx + prjxray) | | DSP blocks | **0** | | LUT usage | ~7,400 (5.8%) | | BRAM usage | ~98% | -| Fmax | 92 MHz | -| Latency | 15.9 ms/token @ 92 MHz | -| Throughput | ~63 tok/s @ 92 MHz | +| Fmax | 92 MHz (an estimate, per the FPGA table in the root README.md; the board run was at 50 MHz) | +| Latency | 15.9 ms/token @ 92 MHz (projected at the 92 MHz estimate, not measured; the measured 50 MHz run is in the Total generation time row) | +| Throughput | ~34 tok/s measured (16 tokens in ~467 ms at 50 MHz); ~63 tok/s @ 92 MHz is a projection, not a measurement (see the FPGA table in the root README.md) | | Tokens generated | 16 (autoregressive from seed=42) | | Total generation time | ~467 ms @ 50 MHz | @@ -88,12 +88,12 @@ All weights use 2-bit ternary encoding: `01` = +1, `10` = -1, `00` = 0. Multipli | Platform | tok/s/W | |----------|---------| -| **Trinity XC7A100T** | **~63** | +| **Trinity XC7A100T** | ~~**~63**~~ withdrawn: power was not measured (none is on record), so no tok/s/W can be stated (see the FPGA table in the root README.md) | | FlightLLM (Alveo U280) | ~1.5 | | Bitnet.cpp (M2 Ultra) | ~0.12 | | Bitnet.cpp (i7-13700H) | ~0.03 | -Note: models differ in size (HSLM ~60K params vs LLaMA-7B), but the hardware efficiency ratio demonstrates the advantage of natively ternary architectures. +Note: models differ in size (HSLM ~60K params vs LLaMA-7B), ~~but the hardware efficiency ratio demonstrates the advantage of natively ternary architectures~~ (withdrawn with the Trinity row above: power was not measured, so no efficiency ratio is claimed; see the FPGA table in the root README.md). ## Latency Breakdown diff --git a/docs/lab/papers/patent-strategy/full-analysis.md b/docs/lab/papers/patent-strategy/full-analysis.md index 8df949c2cb..3b91715137 100644 --- a/docs/lab/papers/patent-strategy/full-analysis.md +++ b/docs/lab/papers/patent-strategy/full-analysis.md @@ -13,7 +13,7 @@ **Key record (18939352)**: "Trinity v2.0.1 — FPGA Autoregressive Ternary LLM" - Author: Vasilev Dmitrii (Trinity) - First autoregressive ternary LLM on FPGA -- QMTech XC7A100T ($30), 63 tok/s @ 92 MHz, ~1W +- QMTech XC7A100T ($30), 63 tok/s @ 92 MHz, ~1W as stated in the record (not current: 63 tok/s was a projection at a 92 MHz Fmax estimate, ~34 tok/s was measured at 50 MHz, and the ~1W figure is withdrawn; see the FPGA table in the root README.md) - Open-source toolchain: openXC7 (yosys + nextpnr-xilinx) - 16 tokens generated in autoregressive mode (seed=42) - License: MIT diff --git a/docs/proposals/BAEZ_OUTREACH.md b/docs/proposals/BAEZ_OUTREACH.md index 5cf6f570ee..8667a5ea9c 100644 --- a/docs/proposals/BAEZ_OUTREACH.md +++ b/docs/proposals/BAEZ_OUTREACH.md @@ -26,7 +26,7 @@ Dear Professor Baez, I've been following your Azimuth blog posts distinguishing numerology from structural physics. I'm hoping you could help evaluate something I've encountered. -I'm building an open-source ternary computing framework (Trinity) grounded in φ² + φ⁻² = 3. The engineering results are concrete: $30 FPGA running a ternary LLM at 63 tok/s, ~1W, zero DSP blocks. +I'm building an open-source ternary computing framework (Trinity) grounded in φ² + φ⁻² = 3. The engineering results are concrete: $30 FPGA running a ternary LLM at ~34 tok/s measured at 50 MHz (power not measured), zero DSP blocks. The mathematics led to φ-based expressions that approximate some physical constants: - m_p/m_e ≈ 6π⁵ (0.002% error) diff --git a/docs/proposals/FRENKEL_OUTREACH.md b/docs/proposals/FRENKEL_OUTREACH.md index e5631bff0b..05c760be37 100644 --- a/docs/proposals/FRENKEL_OUTREACH.md +++ b/docs/proposals/FRENKEL_OUTREACH.md @@ -33,7 +33,7 @@ **Key Principles:** 1. **Present as engineer**, not peer scientist -2. **Show what works** — FPGA on $30, 63 tok/s, 1W (concrete, verifiable) +2. **Show what works** — FPGA on $30, ~34 tok/s measured at 50 MHz, power not measured (the earlier 63 tok/s was a projection and the 1W figure is withdrawn; see the FPGA table in the root README.md) 3. **Ask honestly** — "is this numerology or is there structure?" 4. **Lead with rejected results** — demonstrates scientific integrity @@ -45,7 +45,7 @@ Dear Professor Frenkel, -I am a software engineer building Trinity, an open-source ternary computing framework grounded in φ² + φ⁻² = 3. The engineering works: ternary LLM on a $30 FPGA, 63 tok/s, ~1W, zero DSP blocks. +I am a software engineer building Trinity, an open-source ternary computing framework grounded in φ² + φ⁻² = 3. The engineering works: ternary LLM on a $30 FPGA, ~34 tok/s measured at 50 MHz (power not measured), zero DSP blocks. Along the way, I found φ-based expressions that approximate physical constants (m_p/m_e ≈ 6π⁵ at 0.002%). Some predictions were **falsified** and I document this openly in an "Evidence Ladder" that grades each claim. @@ -69,7 +69,7 @@ Bangkok, Thailand Dear Professor Frenkel, -I'm building an open-source ternary computing framework (FPGA LLM: $30, 63 tok/s, 1W). The architecture is grounded in φ² + φ⁻² = 3, which also generates expressions approximating some physical constants. +I'm building an open-source ternary computing framework (FPGA LLM: $30, ~34 tok/s measured at 50 MHz, power not measured). The architecture is grounded in φ² + φ⁻² = 3, which also generates expressions approximating some physical constants. I maintain an "Evidence Ladder" tracking what works, what's falsified, and what's speculative. I can't tell if there's real mathematical structure here or just coincidences. @@ -101,7 +101,7 @@ Dear Professor Baez, I've been following your Azimuth blog posts distinguishing numerology from structural physics. I'm hoping you could help evaluate something I've encountered. -I'm building an open-source ternary computing framework (Trinity) grounded in φ² + φ⁻² = 3. The engineering results are concrete: $30 FPGA running a ternary LLM at 63 tok/s, ~1W. +I'm building an open-source ternary computing framework (Trinity) grounded in φ² + φ⁻² = 3. The engineering results are concrete: $30 FPGA running a ternary LLM at ~34 tok/s measured at 50 MHz (power not measured). The mathematics led to φ-based expressions that approximate some physical constants: - m_p/m_e ≈ 6π⁵ (0.002% error) diff --git a/docs/proposals/HOSSENFELDER_OUTREACH.md b/docs/proposals/HOSSENFELDER_OUTREACH.md index c7d056cbba..891b4df97a 100644 --- a/docs/proposals/HOSSENFELDER_OUTREACH.md +++ b/docs/proposals/HOSSENFELDER_OUTREACH.md @@ -26,7 +26,7 @@ Professor Hossenfelder, I've been following your work on "Lost in Math" and your critiques of numerical coincidences in physics. I'm hoping for your assessment of something I've encountered. -I'm building an open-source ternary computing framework (Trinity) grounded in φ² + φ⁻² = 3. The engineering works: $30 FPGA running a ternary LLM at 63 tok/s, ~1W. +I'm building an open-source ternary computing framework (Trinity) grounded in φ² + φ⁻² = 3. The engineering works: $30 FPGA running a ternary LLM at ~34 tok/s measured at 50 MHz (power not measured). The mathematics led to φ-based expressions that approximate some physical constants: - m_p/m_e ≈ 6π⁵ (0.002% error) diff --git a/docs/reddit/t27ai/EXAMPLE_POSTS.md b/docs/reddit/t27ai/EXAMPLE_POSTS.md index bf89196cd4..195c1325fb 100644 --- a/docs/reddit/t27ai/EXAMPLE_POSTS.md +++ b/docs/reddit/t27ai/EXAMPLE_POSTS.md @@ -35,7 +35,7 @@ The golden ratio connects to ternary systems! This isn't coincidence — it's fu ## Real-world performance -- **63 tok/s @ 1W** on FPGA (QMTech XC7A100T, $30) +- ~~**63 tok/s @ 1W**~~ **~34 tok/s measured at 50 MHz** on FPGA (QMTech XC7A100T, $30); 63 tok/s was a projection and the 1W figure is withdrawn (see the FPGA table in the root README.md) - **CPU inference** without GPU - **SIMD 17x+** speedup, **JIT 22x+** speedup @@ -436,9 +436,9 @@ Space(infer(x)) = O(d²) ## Experimental Results -| Model | Params | Size | tok/s @ 1W | +| Model | Params | Size | tok/s ~~@ 1W~~ (power not measured) | |-------|--------|------|------------| -| HSLM-1B | 1.95M | 385 KB | 63 | +| HSLM-1B | 1.95M | 385 KB | ~~63~~ (withdrawn: the FPGA figure was a projection and power was not measured, none is on record; see the FPGA table in the root README.md) | | BitNet-3B | 3.1B | 1.2 GB | 12 | ## Why polynomial time matters diff --git a/docs/reddit/t27ai/SIDEBAR.md b/docs/reddit/t27ai/SIDEBAR.md index aabb9db9dc..2132fc8368 100644 --- a/docs/reddit/t27ai/SIDEBAR.md +++ b/docs/reddit/t27ai/SIDEBAR.md @@ -24,7 +24,7 @@ Trinity — Pure Zig autonomous AI agent swarm. ### BitNet LLM - CPU inference without GPU -- 63 tok/s @ 1W on FPGA +- ~~63 tok/s @ 1W on FPGA~~ (withdrawn: 63 tok/s was a projection and power was not measured, none is on record; the FPGA table in the root README.md gives ~34 tok/s measured at 50 MHz) - Quantized weights: {-1, 0, +1} ### TRI-27 @@ -34,7 +34,7 @@ Trinity — Pure Zig autonomous AI agent swarm. ### FPGA - QMTech XC7A100T ($30) -- 0% DSP, 19.6% LUT, 1.2W +- 0% DSP, 19.6% LUT, ~~1.2W~~ (power not measured: no measurement is on record for this board; see the FPGA table in the root README.md) - Fully open-source toolchain --- diff --git a/docs/research/sacred_formats_fpga.md b/docs/research/sacred_formats_fpga.md index 60ed893466..072f7c19c5 100644 --- a/docs/research/sacred_formats_fpga.md +++ b/docs/research/sacred_formats_fpga.md @@ -240,7 +240,7 @@ tri fpga power sacred_alu --clock 100MHz --duration 60s | Unit | Device | LUT | FF | DSP | Fmax (MHz) | Notes | |------------|------------|-----|-----|-----|------------|-------| -| hslm_full_top | XC7A100T | 4,267 | 2,449 | 0 | ≥92 | Autoregressive LLM, 63 tok/s | +| hslm_full_top | XC7A100T | 4,267 | 2,449 | 0 | ≥92 | Autoregressive LLM, ~~63 tok/s~~ ~34 tok/s measured at 50 MHz (63 tok/s was a projection and 92 MHz an Fmax estimate; see the FPGA table in the root README.md) | | ternary_mac | XC7A100T | ~150 | ~80 | 0 | ≥100 | Single MAC unit | ### GF16 Format — To Be Measured (Phase 2) diff --git a/fpga/openxc7-synth/BENCH-005_CORRECTED.md b/fpga/openxc7-synth/BENCH-005_CORRECTED.md index 05141ba0e4..107d280d97 100644 --- a/fpga/openxc7-synth/BENCH-005_CORRECTED.md +++ b/fpga/openxc7-synth/BENCH-005_CORRECTED.md @@ -41,7 +41,7 @@ | System | LUT | FF | DSP | Fmax | Status | |---------|-----|----|-----|------|--------| -| **hslm_full_top** | 4,267 | 2,449 | 0 | ≥92 MHz | ✅ Measured | +| **hslm_full_top** | 4,267 | 2,449 | 0 | ≥92 MHz (an Fmax estimate, per the FPGA table in the root README.md; the board run was at 50 MHz) | ✅ Measured (LUT/FF/DSP; Fmax not measured) | | **gf16_inference** | ⏳ TBD | ⏳ TBD | ⏳ TBD | ⏳ TBD | ⏳ Future work | > `hslm_full_top` = **full inference pipeline** (memory + MAC array + control) diff --git a/fpga/openxc7-synth/BENCH-005_RESULTS.md b/fpga/openxc7-synth/BENCH-005_RESULTS.md index bb361cf818..dffcb1e2a0 100644 --- a/fpga/openxc7-synth/BENCH-005_RESULTS.md +++ b/fpga/openxc7-synth/BENCH-005_RESULTS.md @@ -31,7 +31,7 @@ | System | LUT | FF | DSP | Fmax | Status | |---------|-----|----|-----|------|--------| -| **hslm_full_top** | 4,267 | 2,449 | 0 | ≥92 MHz | ✅ Measured | +| **hslm_full_top** | 4,267 | 2,449 | 0 | ≥92 MHz (an Fmax estimate, per the FPGA table in the root README.md; the board run was at 50 MHz) | ✅ Measured (LUT/FF/DSP; Fmax not measured) | | **gf16_inference** | ⏳ TBD | ⏳ TBD | ⏳ TBD | ⏳ TBD | ⏳ Future work | > **Why NOT comparable**: `hslm_full_top` is a **full inference pipeline** (memory + MAC array + control), while GF16 add/mul are **single operations**. Comparing 118 LUT (single op) to 4,267 LUT (full pipeline) is "apples vs oranges". diff --git a/fpga/openxc7-synth/gf16_synthesis_metrics.md b/fpga/openxc7-synth/gf16_synthesis_metrics.md index 5e517a5599..585c8916d5 100644 --- a/fpga/openxc7-synth/gf16_synthesis_metrics.md +++ b/fpga/openxc7-synth/gf16_synthesis_metrics.md @@ -9,7 +9,7 @@ | FF | 129,600 | | DSP48 | 240 | | BRAM36 | 135 | -| Target Fmax | ≥92 MHz (ternary baseline) | +| Target Fmax | ≥92 MHz (ternary baseline; that 92 MHz is itself an Fmax estimate, per the FPGA table in the root README.md) | ## Synthesis Results (Yosys) @@ -53,7 +53,7 @@ | Module | LUT | FF | DSP | Fmax (MHz) | Status | |--------|-----|----|-----|------------|--------| -| ternary (hslm) | 4,267 | 2,449 | 0 | ≥92 | ✅ Measured | +| ternary (hslm) | 4,267 | 2,449 | 0 | ≥92 (an Fmax estimate, per the FPGA table in the root README.md; the board run was at 50 MHz) | ✅ Measured (LUT/FF/DSP; Fmax not measured) | | gf16_add | 118 | 47 | 0 | ⏳ TBD | ⏳ Synthesis OK | | gf16_mul | 94 | 47 | 1 | ⏳ TBD | ⏳ Synthesis OK | diff --git a/src/tri/outreach/templates.zig b/src/tri/outreach/templates.zig index 9e788c53a3..b132a72e8d 100644 --- a/src/tri/outreach/templates.zig +++ b/src/tri/outreach/templates.zig @@ -173,7 +173,7 @@ pub const templates = [_]Template{ \\• Energy efficiency: 3000× less power than float32 \\• Natural binding: φ² (expansion) and φ⁻² (contraction) \\ - \\Implementation in Zig with FPGA backend: 63 tok/s @ 1W on $30 XC7A100T. Zero DSP usage. + \\Implementation in Zig with FPGA backend: ~34 tok/s at 50 MHz on $30 XC7A100T (power not measured). Zero DSP usage. \\ \\Question: Does ternary VSA merit further investigation in your view? \\ @@ -203,7 +203,7 @@ pub const templates = [_]Template{ \\Hardware difference: Zero-DSP FPGA synthesis vs your HBM approach. \\ \\Our results: - \\• 63 tok/s @ 1W on $30 XC7A100T + \\• ~34 tok/s at 50 MHz on $30 XC7A100T (power not measured) \\• 0% DSP usage, 19.6% LUT \\• Pure Zig implementation (no Python, no CUDA) \\ @@ -349,14 +349,14 @@ pub const templates = [_]Template{ .{ .id = "rabaey_short", .name = "Jan Rabaey — Zero-DSP FPGA LLM", - .subject = "63 tok/s @ 1W, zero DSP usage", + .subject = "~34 tok/s at 50 MHz, zero DSP usage", .body_template = \\Dear Jan, \\ \\I achieved LLM inference on $30 FPGA without DSP blocks. \\ \\Results: - \\• 63 tok/s @ 1W (QMTech XC7A100T) + \\• ~34 tok/s at 50 MHz, power not measured (QMTech XC7A100T) \\• 0% DSP usage, 19.6% LUT \\• Ternary weights {-1,0,+1} \\ diff --git a/src/tri/tri_patent.zig b/src/tri/tri_patent.zig index 6d8ed40191..2908854ed6 100644 --- a/src/tri/tri_patent.zig +++ b/src/tri/tri_patent.zig @@ -142,7 +142,7 @@ fn showAnalysis() void { print(" DSP48: 0/240 (0%%) BRAM36-eq: 135/135 (100%%)\n", .{}); print(" LUT: 4,267/63,400 (6.7%%) FF: 2,449/126,800 (1.9%%)\n\n", .{}); print(" {s}vs TerEffic (2025){s}: 0 DSP vs 3,041 DSP | $30 vs $5,000 | Yosys vs Vivado\n", .{ CYAN, RESET }); - print(" {s}Zenodo{s}: 63 tok/s @ 92 MHz, ~1W, 16 tokens autoregressive (seed=42)\n", .{ CYAN, RESET }); + print(" {s}Zenodo{s}: 16 tokens autoregressive (seed=42), ~34 tok/s measured at 50 MHz; the record's 63 tok/s @ 92 MHz is a projection and its ~1W is withdrawn (README.md FPGA table)\n", .{ CYAN, RESET }); print(" {s}Files{s}: fpga/openxc7-synth/hslm_ternary_mac.v, docs/lab/papers/trinity-fpga/draft.md\n\n", .{ GRAY, RESET }); // P3: Ouroboros @@ -344,7 +344,7 @@ fn showZenodo() void { print(" Date: 2026-03-10\n", .{}); print(" License: MIT\n", .{}); print(" Device: QMTech XC7A100T ($30)\n", .{}); - print(" Speed: 63 tok/s @ 92 MHz, ~1W\n", .{}); + print(" Speed: ~34 tok/s measured at 50 MHz (the record's 63 tok/s @ 92 MHz is a projection; its ~1W is withdrawn, see README.md FPGA table)\n", .{}); print(" Result: 16 tokens autoregressive (seed=42)\n", .{}); print(" Toolchain: openXC7 (yosys + nextpnr-xilinx)\n\n", .{}); diff --git a/tapes/tri-benchmark.tape b/tapes/tri-benchmark.tape index 6a4ac7be1d..a819f01d75 100644 --- a/tapes/tri-benchmark.tape +++ b/tapes/tri-benchmark.tape @@ -1,5 +1,5 @@ # VHS tape for tri benchmark -# Key demo: 63 tok/s performance for addressing speed objections +# Key demo: benchmark performance for addressing speed objections (the 63 tok/s once quoted here is withdrawn: it was an FPGA projection, see the FPGA table in README.md) Output tri-benchmark.gif Set FontSize 26