Skip to content

Repository files navigation

WISP

Stream what shouldn't run.

Tests Version CUDA Vulkan CPU--only DGX Spark ROCm License Stars

WISP — Stream what shouldn't run

WISP runs the largest open-source AI models ever built on the hardware sitting on your desk.

A 744B parameter frontier model. A 671B reasoning engine. A 2.8 trillion parameter behemoth, weights already public.

Not a slow demo. Not quantized to uselessness. The full model, at full frontier intelligence, streaming expert weights across GPU VRAM, system RAM, and your NVMe SSD in real time.


v3.0 — What Shipped

Ten ships across August–September 2026. Where something is built but not compiled, planned but not run on real weights, or a preview, this section says so.

CPU-only mode. AVX2 + FMA SIMD kernels, selected at runtime, and a two-tier RAM → NVMe hierarchy: any x86_64 CPU with AVX2 (Intel since 2013, AMD since 2015) runs WISP without a GPU. Measured on the build machine: 57.6 GB/s all-core RAM bandwidth. From that bandwidth wisp doctor estimates Mixtral-8x7B at ~6 tok/s once its experts are in RAM — an estimate, not a benchmark; a CPU-only run on real weights has not been timed yet.

Vulkan backend (compile with the Vulkan SDK). Compute shaders designed for AMD R9700 (RDNA4), the Ryzen AI Halo iGPU, Intel Arc and any Vulkan 1.2 GPU, with a zero-copy path on unified-memory devices that maps device memory directly instead of staging copies. Build with cmake -DWISP_VULKAN=ON (needs the Vulkan SDK's headers and glslc). Detection and auto-config work on every build; the backend itself has not been compiled or run on hardware yet.

New model adapters. DeepSeek-V4-Flash (284B / 13B active, 258 lookups/token) and GLM-5.3-Flash (320B / 18B active) are planning-only: wisp info sizes hardware for them and wisp convert refuses them before downloading. GLM-5.3-Flash's fp8 block scales are already decoded; its hyper-connections are not. Qwen3-Coder-480B-A35B (480B / 35B active, Apache 2.0) converts through the standard qwen3_moe path.

fp8-e4m3 decode. The official DeepSeek-V3 and R1 checkpoints ship fp8-e4m3 weights with a 128×128 block scale for each. WISP multiplies every block by its scale at convert time and refuses any fp8 weight whose scale is missing. Tested on synthetic checkpoints; a full 671B conversion has not been run yet.

Honest sizes. Expert and dense sizes are derived from each model's published config rather than borrowed estimates, and download sizes are read from the Hugging Face file listings. GLM-5.2's dense weights (~34.6 GB at fp16) and DeepSeek's (~31.8 GB) exceed consumer GPU VRAM. Dense quantization is planned for v3.1.

Convert preflight. Before a single byte downloads: a disk-space check covering download plus converted output, refusal of planning-only models, a refusal to overwrite a converted model, and a [y/N] summary of what is about to happen. --yes skips the prompt for scripts; --force re-converts or repairs.

Kimi K3 tensor names verified. Checked against the real moonshotai/Kimi-K3 checkpoint index and its modeling code: layers live under language_model.model.layers.N, routed experts under block_sparse_moe.experts.E.w1/w2/w3. Layer 0 is a dense MLP, so 1,472 lookups/token (92 MoE layers × top-16). The experts are MXFP4-packed, so wisp convert refuses K3 until MXFP4 decode lands — along with the latent-MoE and released-KDA pieces listed under KDA.

Web dashboard. /dashboard: one self-contained HTML page, zero external dependencies, so it works air-gapped. The expert heatmap draws every expert as a cell, colored by tier, with a white flash on each activation. Verified live on Mixtral-8x7B (2026-09-12): 64 experts flash per token, hit rate 35% → 65.5% over 125 tokens, TTFT 23.5 s cold.

Multi-node cluster (preview). Coordinator/worker architecture: a shard planner divides expert layers across nodes, workers serve their expert files, and the coordinator checks them at startup (--cluster-role, --cluster-nodes, --cluster-coordinator), exposes /v1/cluster and shows a node panel on the dashboard. In v3.0 inference still reads experts from local disk; remote worker expert fetch is v3.1.

Speculative decoding is opt-in. --speculative enables it, --no-speculative skips every drafter check. Nothing downloads on a first request: WISP asks, with the size, before fetching a drafter, and is_drafter_complete() rejects partial downloads.


The Numbers

Model Params Active/token Lookups/token Disk (int4) Status
GLM-5.2 744B 40B 600 ~414 GB ✅ Ready
DeepSeek-V3 671B 37B 464 ~374 GB ✅ Ready — fp8†
DeepSeek-R1 671B 37B 464 ~374 GB ✅ Ready — fp8†
Mixtral-8x7B 47B 13B 64 26.6 GB* ✅ Verified end-to-end
Mixtral-8x22B 141B 39B 112 ~90 GB ✅ Ready
Kimi K3 2.8T 104B 1,472 ~1.4 TB est. ⚠️ Names verified — MXFP4 pending
Qwen3-235B 235B 22B 752 ~130 GB ✅ Ready
Qwen3-2.4T 2.4T ~22B est. TBD ~1.2 TB est. ✅ Adapter ready — weights pending
Qwen3-Coder-480B 480B 35B 496 ~268 GB ✅ Ready
DeepSeek-V4-Flash 284B 13B 258 ~158 GB 📐 Planning only
GLM-5.3-Flash 320B 18B 336 ~178 GB 📐 Planning — fp8 ✓, hyper-connections pending
GLM-5.3 744B ~40B 600 ~414 GB ⚡ Weights released — adapter in progress§

*Measured from a real conversion — every expert file verified. "Verified end-to-end" means: downloaded, converted, and generated real tokens through the full C + CUDA engine on consumer hardware.

†fp8-e4m3 dequantized at convert time: e4m3 values plus a float32 weight_scale_inv holding one multiplier per 128×128 block. Before 2026-09-11 wisp convert dropped those scales, and its expert pattern also matched them as expert weights, so the converted model verified clean and generated garbage. Tested on synthetic checkpoints; full conversion not yet verified on real 671B weights. DeepSeek's sizes are derived from the official config — a 24.8 MB expert (3 × 2048 × 7168 weights in WISP's int4) — not yet measured from a conversion.

‡Dense weights for GLM-5.2 (~34.6 GB) and DeepSeek (~31.8 GB) at fp16 exceed consumer GPU VRAM. Dense quantization: v3.1. Current workaround: run with --cpu-only on a machine with 64 GB+ RAM, which keeps the dense weights in system memory; WISP's planner fits them at 64 GB, not at 48 GB.

§GLM-5.3 is planned against GLM-5.2's shape: Z.ai describes it as the same base model and the two published configs match. The adapter is still a stub, so wisp convert refuses it.

Lookups count routed experts only: shared experts run for every token, so they stay with the dense weights and are never streamed. Disk is WISP's int4 output, computed from the published shapes. It is not the download size. Qwen3-Coder-480B, DeepSeek-V4-Flash and GLM-5.3-Flash were checked against each model's published config.json and model card on 2026-09-11; Kimi K3's tensor names and layer layout against moonshotai/Kimi-K3 on 2026-09-14.

  • Qwen3-Coder-480B is the same qwen3_moe architecture as Qwen3-235B, so it takes the same converter path at a bigger shape (160 experts, 62 layers).
  • DeepSeek-V4-Flash and GLM-5.3-Flash are planning only. V4-Flash ships fp4 experts under DeepSeek's own tensor names and uses CSA/HCA compressed attention. GLM-5.3-Flash uses a KDA/DSA hybrid and a vision tower. Both use mHC hyper-connections, which the engine does not implement yet. GLM-5.3-Flash's fp8 weights use the same weight_scale_inv block scales as DeepSeek-V3, which the converter decodes. V4-Flash names its scales .scale, and that layout is not wired in.

July 2026 — The Biggest Week in Open Source AI

Kimi K3 (2.8T parameters) dropped July 17 — the largest open source model ever released. Qwen3.8 (2.4T parameters) announced July 19 — open weights coming soon. WISP was built for exactly this moment. Both are MoE models. Both stream with WISP.


How It Works

Every MoE (Mixture-of-Experts) model activates only a small fraction of its parameters for each token. GLM-5.2 activates ~5.4%. Kimi K3 activates 3.7% (104B of 2.8T). WISP exploits this with a self-organizing 3-tier cache:

Token arrives
    ↓
Model's own router selects which experts to activate
    ↓
WISP checks VRAM first  → instant if hit
    ↓
Then RAM                → one PCIe copy if hit
    ↓
Then NVMe SSD           → streams if cold
    ↓
LRU cache promotes hot experts upward automatically
After 10-15 min: high cache hit rate, self-organized
(measured: 68.8% hit rate after only 80 tokens)
No configuration. No preset modes. Fully automatic.

Absorbed MLA Attention

GLM-5.2 and DeepSeek use Multi-head Latent Attention (MLA). WISP implements true absorbed MLA — the KV cache stores the compressed c_kv latent instead of expanded K,V tensors. Result: ~70KB per token KV cache instead of ~5MB. This is what makes 1M-token context feasible in RAM.

KDA — Kimi Delta Attention

Kimi K3 runs linear attention in 69 of its 93 layers. Instead of a KV cache that grows with the conversation, each head carries a fixed [d_k, d_v] state matrix updated by a delta rule:

qn, kn = q/||q||, k/||k||     L2-normalized per head
b      = sigmoid(W_beta x)    per-channel write gate, in (0,1)
u      = S^T kn               what memory currently holds at this key
S      = S + b * kn (v - u)^T write the PREDICTION ERROR
o      = S^T qn               read, using the new state
y      = silu(g) * o          output gate

Three details carry the whole thing:

It writes v - u, not v. That is what makes it a delta rule — the state is corrected by exactly the amount it was wrong by, so re-writing a key that is already stored is a no-op instead of doubling it. Plain linear attention (S += v k^T) accumulates and saturates.

q and k are L2-normalized, which bounds the update. Without it the state diverges: measured at inf within 200 tokens at moderate input scale. A constant-size state is worthless if its contents blow up.

The read happens after the write, so a token can see its own value. Reading first makes the first token of every conversation emit exactly zero.

Cost is O(n) in sequence length and memory is constant — that is the property that makes 1M-token context tractable at all.

WISP implements KDA as a CUDA kernel with a matching pure-PyTorch fallback for CPU-only mode. The kernel normalizes internally rather than trusting its callers, so the C engine, the Python bindings and the fallback cannot drift apart. Tests assert the two paths agree numerically at head_dim 32, 100 and 200, that prefill lands on the same state as sequential decode, that repeated writes converge on the stored value, that the state stays bounded over 2,000 tokens, and that batch entries do not contaminate each other. The other 24 K3 layers are Gated MLA and already run through WISP's absorbed-MLA path.

Status: ✅ Kernel complete. Tensor names verified against moonshotai/Kimi-K3 (2026-09-14). Blocked on MXFP4 expert decode — v3.1. MXFP4 is the first blocker, not the only one: the released checkpoint also runs its experts in a 3,584-dim latent (routed_expert_down_proj / routed_expert_up_proj), and its KDA layers add short convolutions and a learned decay (A_log, dt_bias, f_a_proj, f_b_proj) beyond the recurrence above. wisp convert refuses K3 until those exist, before downloading anything.

Double-Buffer Async Pipeline

While the GPU computes token N (2-8ms), the C engine loads token N+1's predicted experts from SSD into pinned RAM, and the transfer stream moves them to VRAM (0.1-0.3ms) — fully hidden inside the compute window. GPU never waits on predicted experts. Works on all consumer hardware. No GPUDirect Storage required.

Speculative Decoding

Same-family small models draft 3 tokens simultaneously. The main model verifies all 3 in one parallel forward pass. At 39-55% acceptance rate: 2.2-2.8x effective throughput. Zero quality loss — the rejection-sampling scheme provably preserves the main model's output distribution (Leviathan 2023).

Opt-in: pass --speculative to enable. Default is off — no automatic downloads.

Display Auto-Detection

WISP detects whether your monitor is on the GPU or the motherboard (via the driver itself) and reserves VRAM accordingly. Move the monitor to the motherboard port → WISP auto-detects → full VRAM dedicated to inference. Override anytime with --display-mode gpu|igpu|auto.

Learning Cache — it gets faster the more you use it

WISP records which experts your sessions actually activate, in {model_dir}/.wisp_usage. On the next startup the hottest ones are queued for pre-warming before the first token is generated.

session 1   cold start; the LRU discovers your domain
session 2   last session's top experts pre-warmed at startup
session 7   near-instant warm start

Expert selection happens inside the C router, so the engine keeps a ring-buffer log of every (layer, expert) access that the runtime drains every 10 tokens — that stream feeds both the next-token prefetch predictor and the cross-session cache. Measured on a real Mixtral run: 12 tokens produced 768 expert observations and 238 tracked experts, of which 107 were pre-warmed on the following startup.

wisp cache --model ./models/glm-5.2/ --show    # what it has learned
wisp cache --model ./models/glm-5.2/ --reset   # start over

Ranking blends frequency with a recency decay, so a cache trained on three weeks of Rust adapts when you switch to prose instead of staying stuck on the old domain. The file is plain JSON — inspect it, diff it, or delete it.

OpenAI-Compatible API Server

wisp serve --model ./models/glm-5.2/ --port 8080

Anything that speaks the OpenAI API now speaks to WISP — Cursor, Continue.dev, Open WebUI, LM Studio frontends, the openai package:

from openai import OpenAI
client = OpenAI(base_url="http://localhost:8080/v1", api_key="wisp")
client.chat.completions.create(
    model="glm-5.2",
    messages=[{"role": "user", "content": "Write a quicksort"}],
    stream=True,
)

/v1/chat/completions (streaming + non-streaming), /v1/models, /health, /v1/stats for WISP's own numbers (tok/s, tier hit breakdown, learning-cache state, TTFT), /v1/expert-state for the live expert map, and the web dashboard at /dashboard. Streaming is genuinely incremental — the engine's generator runs on a worker thread feeding the event loop, so the first token reaches the client as soon as it exists rather than after the whole completion. Prompts are rendered with each family's own chat template. Requests are serialized: one engine, one KV cache, so concurrent decoding would interleave two conversations.

Install the extra: pip install -e '.[server]'

Stability Guarantees

  • System RAM is never allowed to fill: max(6GB, 25%) is always reserved for the OS, and a runtime watermark evicts RAM-tier experts past 80% usage (SSD stays authoritative, so eviction is always safe).
  • VRAM planning is hard-capped at 75% of the card, and the allocator degrades gracefully (evict-and-retry) if the driver's real ceiling is lower than the plan.

Web Dashboard

wisp serve --model ./models/glm-5.2/ --port 8080

Open http://localhost:8080/dashboard in any browser. Works from any device on the same network when the server is started with --public — open http://<this machine's IP>:8080/dashboard there. --public has no authentication: anyone who can reach the port can use your GPU, so only do this on a network you trust. Zero external dependencies — works on air-gapped machines.

Expert heatmap, one row per layer and one column per expert:

Blue  = in VRAM   (fast)
Amber = in RAM    (medium)
Grey  = on NVMe   (cold)
White = activated this token (flashes, fades over 800ms)

Brighter means an expert has fired more often; hover a cell for its layer, expert, tier and heat. Tiers come from the C caches themselves and heat from the learning cache. Until the model has loaded, the heatmap is empty rather than simulated.

Live stats: tok/s, cache hit rate, TTFT, token count. Tier monitor: VRAM / RAM / NVMe usage bars. Streaming chat: test inference directly from the browser. Everything it draws is plain JSON at /v1/expert-state and /v1/stats.

Verified on Mixtral-8x7B (2026-09-12): 64 experts flash per token. Hit rate 35% → 65.5% over 125 tokens. TTFT 23.5s cold start.


Desktop GUI

wisp-gui is a native desktop app that connects to the wisp serve API server running locally.

pip install -e '.[gui]'
wisp-gui

Point it at your converted model directory. The GUI starts wisp serve internally on a private loopback port.

Features:

  • Token stream with live tok/s counter
  • Tier monitor: VRAM / RAM / NVMe usage live from /v1/stats
  • Cache hit rate and expert observation count
  • Temperature, max_tokens, system prompt controls
  • Dark theme

Three columns: model + generation controls on the left, chat in the middle, live tier and engine monitors on the right. The tier bars show VRAM and RAM occupancy against this machine's real capacities, refreshed every two seconds.

The GUI is not a wrapper around the CLI — it never shells out or parses terminal output. It hosts the very same WispServer in a background thread and talks to it over the OpenAI API, so the path it exercises is byte-for-byte the one Cursor and Open WebUI use. Closing the window shuts the server down cooperatively, which is what lets the learning cache persist.

Screenshots welcome — open a GitHub Issue to submit yours.


Hardware Platforms

NVIDIA CUDA — Consumer GPUs

Standard 3-tier: VRAM → RAM → NVMe. RTX 3080+ recommended. RTX 5070 (12GB) verified end-to-end on Mixtral-8x7B. CUDA 12.0+ | 13.x. CUDA 12.8+ required for RTX 50 series.

This is the only path that has been run end-to-end on real weights.

Vulkan — AMD, Intel, Any GPU (compile with SDK)

Compute shaders designed for:

  • AMD R9700 (RDNA4) via Mesa/RADV or the AMD driver
  • AMD Ryzen AI Halo iGPU (Radeon 8060S, 40 CUs)
  • Intel Arc via the Intel Vulkan driver
  • Any GPU with a Vulkan 1.2 driver

Build: cmake -DWISP_VULKAN=ON (requires the Vulkan headers, glslc, and the Vulkan loader — all in the Vulkan SDK). Detection and auto-config work on all builds. Pass --vulkan to force the Vulkan path. The backend has not been compiled or run on hardware yet; compare its output against the CUDA or CPU path before trusting it.

AMD Ryzen AI Halo — Unified Memory

Up to 128GB LPDDR5x unified pool, Radeon 8060S iGPU. WISP detects it automatically (the 8060S by name, an integrated device, and a pool of at least 64GB) and plans 2-tier mode: unified pool → NVMe SSD. Zero-copy GPU access via Vulkan (requires a Vulkan build); without one, WISP runs CPU-only. Two Halo units over the TCP cluster would plan a 256GB pool — see the cluster preview for what that does today.

NVIDIA DGX Spark — Unified Memory

128GB coherent unified pool, 273 GB/s, GB10. ARM64 (aarch64). $4,699. WISP detects it automatically — all three of aarch64, a GB10-class device name, and a ~128GB reported pool must agree — and switches to 2-tier unified mode: CPU and GPU share the same physical memory, so the RAM→VRAM copy the double buffer exists to hide simply does not happen; prefetch from NVMe still matters. Sparks carry ConnectX-7 200 GbE, the link the 2–4 node cluster is planned around: 2 nodes = 256GB, 4 nodes = 512GB of pooled memory once remote expert fetch lands (v3.1).

wisp doctor   # shows "DGX Spark detected — unified memory mode"

Detection and tier planning are implemented; no Spark has been benchmarked. The figures above are hardware specifications, not WISP measurements.

CPU-Only — No GPU Required

Any x86_64 with AVX2 (Intel since 2013, AMD since 2015). 2-tier: RAM → NVMe. No CUDA or Vulkan needed. Use: wisp serve --cpu-only

Estimates from wisp doctor, which measures RAM bandwidth (57.6 GB/s on a 32GB DDR5-6000 build machine) — not timed runs:

  • Fastest CPU model: Mixtral-8x7B, ~6 tok/s once its experts are in RAM (25% of its experts fit in RAM on that 32GB machine)
  • GLM-5.2 CPU-only: ~1.15 tok/s with experts in RAM, ~0.36 tok/s cold

AMD Radeon AI PRO R9700 — ROCm

32GB GDDR6. 640 GB/s. 1531 TOPS INT4. PCIe 5.0. $1,299. Dual-slot blower — built for dense multi-GPU stacking.

WISP v2.0 detects R9700 via rocm-smi and sets auto-config. Full ROCm compute kernels (HIP port of CUDA kernels): v3.1, alongside the Vulkan path above. Stacking projections are in the R9700 table. ROCm contributions welcome — see CONTRIBUTING.md.

Setup: scripts/install_rocm.sh (Ubuntu) or install_rocm.ps1 (Windows). Both are detection-only; they do not make inference run on the card.


Multi-Node Cluster (Preview)

Coordinator/worker architecture for 2–4 nodes. The coordinator serves the normal OpenAI API. Workers hold their shard of the expert pool. Clients see a standard /v1/chat/completions endpoint.

v3.0: coordinator and worker architecture, shard planner, /v1/cluster endpoint, dashboard node panel. Expert data is still served from local disk: the coordinator runs inference from its own full copy of the model, and workers are planned and health-checked only.

v3.1: remote worker expert fetch via HTTP — the full distributed pool becomes active.

What works today:

  • Shard planning: layers are split into contiguous ranges, as evenly as the count allows (75 layers over 4 workers: 19, 19, 19, 18).
  • Workers serve their expert files over HTTP (POST /shard/experts) and report the layers they actually hold, their memory and the model (GET /shard/health).
  • The coordinator checks every worker at startup, prints the plan and the pooled memory, exposes GET /v1/cluster, and the dashboard shows each node as online or unreachable.

What does not work yet:

  • Expert routing. The C engine reads experts from its own model directory, and nothing feeds it bytes fetched from a worker. Routing needs a C-side expert-source hook.
  • RDMA. The transport is HTTP/JSON over TCP; --cluster-fabric rdma is refused rather than silently falling back.
  • Authentication. A worker serves expert weights to anyone who can reach its port. Run clusters on a network you trust.

DGX Spark (2–4 nodes, 200 GbE)

Node 0 (coordinator):

wisp serve --model ./models/glm-5.2/ \
  --cluster-role coordinator \
  --cluster-nodes http://spark1:8080,http://spark2:8080

Nodes 1+ (workers — --public so the coordinator can reach them):

wisp serve --model ./models/glm-5.2/ --public \
  --cluster-role worker --cluster-node-id 1 \
  --cluster-coordinator http://spark0:8080

Memory pools (projected, for when remote fetch lands):

  • 2 × DGX Spark: 256 GB — about two-thirds of GLM-5.2's experts (19,200 × 21.2 MB ≈ 380 GB)
  • 4 × DGX Spark: 512 GB — all of GLM-5.2's experts
  • Kimi K3's experts (~1.4 TB) exceed even four nodes, and K3 does not convert yet

For scale, not a measurement: a fully cold GLM-5.2 token needs 600 experts × 21.2 MB = 12.7 GB. Over a 200 Gb/s link (25 GB/s line rate) that is at least half a second, before HTTP overhead and base64's extra third.

Ryzen AI Halo (2–4 nodes, 10 GbE TCP)

Same commands. Inter-node bandwidth: ~1.25 GB/s (10 GbE), a tenth of a Spark link, so a fully cold GLM-5.2 token over the wire would take roughly ten seconds.


Performance

First verified on: R7 9800X3D | RTX 5070 12GB | 32GB DDR5-6000 | PCIe 4.0 NVMe

Model Cold Warm Hot +MTP Effective
Mixtral-8x7B 0.75 tok/s ✅ measured est. 2-5 tok/s* est. 5-10 tok/s* est. ~2x*
GLM-5.2 est. 0.7 tok/s* est. 5.5 tok/s* est. 8.5 tok/s* est. ~14 tok/s*
DeepSeek-V3/R1 est. 0.8 tok/s* est. 5.8 tok/s* est. 9.0 tok/s* est. ~14.5 tok/s*
Kimi K3 est. 0.2 tok/s* est. 3 tok/s* est. 6 tok/s* est. ~10 tok/s*

Cold = empty cache, first run. Warm = 15 min same domain. Hot = repeated patterns, cache fully warmed. MTP = with speculative decoding active. * = estimated from physics (expert size × lookups ÷ transfer bandwidth); real benchmarks land in v2.1. Measured Mixtral data (2026-07-19): 81 tokens at 0.75 tok/s from a completely cold engine, 68.8% cache hit rate after 80 tokens; a 300-token cold run averaged 0.56 tok/s on an earlier engine build. We do not publish a bold number we didn't measure.

Cold → warm progression, every figure measured (Mixtral-8x7B, R7 9800X3D + RTX 5070):

2026-07-19  cold engine   81 tokens averaging 0.75 tok/s
                          68.8% cache hit rate after 80 tokens
2026-09-12  cold start    TTFT 23.5 s (prefill streams experts off NVMe)
            live run      hit rate 35% → 65.5% over 125 tokens,
                          0.56 tok/s engine average, plain decoding

Warm and hot steady-state rates on long runs have not been measured yet.

Why Mixtral is Slower Than GLM-5.2 (yes, slower)

Mixtral 8x7B experts = 99MB each (measured). GLM-5.2 experts = 20.2MB each — 4.9× smaller.

Cold decode is transfer-bound: every uncached expert crosses PCIe or comes off the NVMe. Smaller experts = less data per token = faster streaming and faster cache warm-up. Mixtral is WISP's proving ground; GLM-5.2's fine-grained experts are where the architecture truly sings.


Quick Start

git clone https://github.com/zeroextub-collab/wisp
cd wisp
pip install -e .            # builds the C + CUDA engine (see Installation)
wisp doctor                 # confirm the toolchain found everything

wisp convert --model glm-5.2 --output ./models/
wisp chat --model ./models/glm-5.2/
# Desktop GUI
pip install -e '.[gui]'
wisp-gui
# CPU-only (no GPU needed)
wisp chat --model ./models/mixtral-8x7b/ --cpu-only

# Web dashboard
wisp serve --model ./models/mixtral-8x7b/ --port 8080
# then open http://localhost:8080/dashboard

# Speculative decoding (opt-in; asks before downloading the ~14 GB drafter)
wisp chat --model ./models/mixtral-8x7b/ --speculative

Not on PyPI yet. WISP builds a CUDA extension against your local toolchain, so it installs from source today — pip install wisp-engine will not work. (Note also that the name wisp on PyPI belongs to an unrelated project.) Prebuilt wheels are tracked in the roadmap.


Requirements

Component Minimum Recommended
RAM 16 GB 32 GB+
NVMe SSD PCIe 3.0, 300 GB free PCIe 5.0, 2 TB dedicated
GPU None (CPU-only works) RTX 3080+ / 8 GB+ VRAM
CUDA 12.0+ 12.8+ (required for RTX 50 series)
Python 3.10+ 3.11
OS Windows 10+ / Ubuntu 20.04+ Windows 11 / Ubuntu 22.04
Vulkan SDK 1.2+ optional — only for a -DWISP_VULKAN=ON build
OpenMP ships with the compiler (MSVC's 2.0 is enough) required for -DWISP_CPU_ONLY=ON builds

Installation

Windows (x64 Native Tools Command Prompt)

git clone https://github.com/zeroextub-collab/wisp
cd wisp
powershell -File scripts\install.ps1

Linux

git clone https://github.com/zeroextub-collab/wisp
cd wisp
bash scripts/install.sh

Manual

pip install torch --index-url https://download.pytorch.org/whl/cu128
pip install -e .
wisp doctor   # verify everything works

Usage

# Convert a model (downloads + converts to WISP format,
# resumable, SHA256-verified, integrity-checked at the end)
wisp convert --model glm-5.2 --output ./models/

# One-shot inference
wisp run --model ./models/glm-5.2/ \
         --prompt "Write a Python web scraper" \
         --stream

# Interactive chat (/clear /stats /quit)
wisp chat --model ./models/glm-5.2/

# Benchmark your hardware (cold -> warm -> hot)
wisp benchmark --model ./models/glm-5.2/ --runs 3

# Check system compatibility
wisp doctor

# Show tier allocation for your hardware (works pre-download)
wisp info --model glm-5.2

# Print the tensor names a checkpoint actually uses, and which of them
# WISP matched (use before converting an unreleased/unverified model)
wisp inspect --source ./shards --model kimi-k3

# Verify model integrity, expert file by expert file
wisp verify --model ./models/glm-5.2/

# Share a converted model so others skip re-conversion
wisp upload --model ./models/glm-5.2/ --repo you/glm-5.2-wisp

GUI

pip install -e '.[gui]'
wisp-gui

Opens the desktop app; point it at a converted model directory.

Python API

from wisp import WispEngine

engine = WispEngine("./models/glm-5.2/")

# Streaming
for token in engine.stream("Explain quantum entanglement"):
    print(token, end="", flush=True)

# One-shot
result = engine.generate("Write a sorting algorithm",
                         max_new_tokens=500,
                         temperature=0.7)
print(result)

Multi-GPU Support

WISP automatically detects and configures multiple GPUs. Zero manual configuration needed.

Setup Strategy Best For
Single GPU 8-12GB Dense + expert LRU cache Getting started
Single GPU 24GB+ Dense + large expert cache GLM-5.2 smooth
Dual GPU same size GPU0 dense, GPU1 pure cache 2× cache hits
Dual GPU diff sizes Bigger=dense, smaller=overflow Flexible
3+ GPUs Pipeline parallelism Maximum throughput

NVIDIA Stacked Setups (projected)

Setup Combined VRAM GLM-5.2 Kimi K3
2× RTX 4090 48GB Smooth warm Feasible
4× RTX 4090 96GB Near full cache Good
4× RTX 6000 Ada 192GB Full expert cache Strong

AMD Radeon AI PRO R9700 (ROCm detection + auto-config in v2.0. Full HIP compute kernels + Vulkan: v3.1)

The R9700 is purpose-built for exactly what WISP does: 32GB GDDR6 per card, dual-slot blower for dense stacking, PCIe 5.0, 1531 TOPS INT4.

R9700 Stack Combined VRAM GLM-5.2 Coverage Projected tok/s
1× R9700 32GB Dense + 1,257 experts 8-12
2× R9700 64GB Dense + 3,085 experts 18-25
4× R9700 128GB Dense + 6,741 experts 38-52

Note: WISP now detects the R9700 through ROCm — wisp doctor reports the card, its VRAM, and the gfx1201 target, and the tier planner sizes for it. It does not run on it: the compute kernels are CUDA and the HIP port is separate work. All R9700 throughput figures above remain projections from the tier math, not measurements. scripts/install_rocm.sh sets up detection. Contributions welcome.


Architecture

User prompt
    ↓
WISP CLI / Python API
    ↓
Auto-Config Engine
  (profiles hardware once, calculates optimal tier split,
   detects display on GPU/iGPU, sets safe VRAM/RAM budgets)
    ↓
Universal Runtime (C + CUDA)
  ├── Double-Buffer Async Prefetch
  │     hides SSD transfer inside GPU compute time
  ├── 3-Tier LRU Cache
  │     VRAM (fastest) → RAM (fast) → SSD (cold)
  │     scratch-ring serving + hit-gated promotion:
  │     cool experts never churn the VRAM cache
  ├── Speculative Decoding
  │     same-family drafter → 2.2-2.8x throughput
  ├── Multi-GPU Router
  │     auto-selects: single / dual / pipeline strategy
  └── RAM Watermark Monitor
        evicts experts past 80% RAM — the desktop never starves
    ↓
The model's own router drives expert selection.
WISP just delivers them as fast as possible.

Three layers, one job each: Python orchestrates (download, convert, configure — things that run once), C owns the hot path (64-1,472 expert fetches per token, cache coordination, prefetch threads), CUDA owns the math (absorbed MLA / GQA / KDA attention, int4 dequant, fused SwiGLU FFN, router top-K). The engine is model-agnostic: adapters map every family onto one canonical weight layout at conversion time, so adding a model is one adapter file — not a new engine.

On a coherent-memory machine the tier planner emits a two-tier plan instead (unified pool → NVMe); nothing else in the stack changes.


Roadmap

v1.0 — Shipped (August 1, 2026)

  • Universal 3-tier MoE streaming engine (C + CUDA)
  • Mixtral-8x7B and 8x22B full support — end-to-end verified
  • GLM-5.2 (744B), DeepSeek-V3/R1 (671B) adapters
  • Absorbed MLA attention (~70KB/token KV cache)
  • KDA linear attention kernel (CUDA + PyTorch fallback)
  • Double-buffer async prefetch pipeline
  • Speculative decoding (39-55% acceptance, 2.2-2.8x throughput)
  • OpenAI-compatible API server (wisp serve)
  • Learning cache (cross-session expert pre-warming)
  • Display auto-detection (VRAM conflict prevention)
  • Multi-GPU support (single / dual / pipeline)
  • 121 tests

v2.0 — Shipped (August 2026)

  • Kimi K3 (2.8T): KDA converter weight mapping + C forward pass branch. Kernel numerically verified; checkpoint tensor names verified (2026-09-14), with wisp inspect and WISP_KDA_NAMES as the escape hatch
  • Qwen3-235B and Qwen3-2.4T adapters
  • Desktop GUI (wisp-gui) — OpenAI client + live dashboard
  • DGX Spark unified memory mode (auto-detected, not benchmarked)
  • AMD R9700 ROCm detection + auto-config (detection only)
  • 197 tests

v3.0 — Current (September 14, 2026)

  • CPU-only mode: AVX2+FMA kernels, RAM→NVMe 2-tier
  • Vulkan backend: shaders + detection (compile with SDK; not yet run on hardware)
  • fp8-e4m3 decode with 128×128 block scales
  • DeepSeek V4 Flash adapter (284B/13B, 258 lookups; planning only)
  • Qwen3-Coder-480B adapter (Apache 2.0)
  • GLM-5.3-Flash planning (fp8 unblocked)
  • Convert preflight: disk check, [y/N], --yes, --force
  • Kimi K3 tensor names verified (moonshotai/Kimi-K3)
  • Web dashboard: /dashboard, expert heatmap, zero CDN
  • Multi-node cluster: coordinator/worker, preview
  • Speculative decoding opt-in (--speculative)
  • Honest sizes: derived from published configs, download sizes from the Hub
  • 436 tests

v3.1 — Next

  • Dense layer quantization (GLM-5.2 + DeepSeek on consumer hardware — the last major blocker)
  • GLM-5.3 full adapter (weights released 2026-08-25)
  • MXFP4 decode (the first of Kimi K3's blockers; latent MoE and the released KDA layer remain)
  • Remote worker expert fetch (completes cluster)
  • Full ROCm HIP + Vulkan compiled kernels (AMD R9700)
  • GLM-5.3-Flash full support (hyper-connections)
  • macOS / Metal backend

v4.0 — Community Driven

  • Generic MoE adapter (any HuggingFace MoE model)
  • GPUDirect Storage (Linux, datacenter GPUs)
  • Plugin system for community model profiles
  • Expert atlas (crowd-sourced topic affinity map)

Still unmeasured — help wanted

  • Real GLM-5.2 benchmark numbers
  • Real DeepSeek-V3 benchmark numbers
  • Warm/hot steady-state Mixtral numbers (long runs)
  • A timed CPU-only run on real weights, to replace the wisp doctor estimates
  • A Vulkan backend build and run on any GPU
  • A DGX Spark run, to replace the unified-mode projections

Credits

WISP would not exist without Colibrì.

JustVugg built Colibrì in July 2026 — a 2,400-line pure-C engine that proved a 744B parameter model could run on 25GB of consumer RAM by streaming expert weights from disk. Before Colibrì, everyone said this was impossible. After Colibrì, we built WISP.

jlnsrk converted the GLM-5.2 weights to a format the community could actually use.

matey-0 (Mateo Grgić) fixed the MTP head from int4 (0-4% acceptance) to int8 (39-59% acceptance) — turning speculative decoding from broken to genuinely useful.

→ github.com/JustVugg/colibri

WISP shares zero code with Colibrì. Complete independent reimplementation. JustVugg showed us what was possible.

Research

  • Leviathan et al. 2023 — Speculative Decoding
  • GLM team — GLM-5.2 and IndexShare MoE architecture
  • DeepSeek team — DeepSeek-V3/R1 MoE + Multi-head Latent Attention
  • Moonshot AI — Kimi K3: Open Frontier Intelligence (arXiv:2607.24653) — KDA hybrid linear attention, Stable LatentMoE routing
  • The llama.cpp community — proof that consumer hardware deserves frontier models

Full tribute: CREDITS.md


Contributing

See CONTRIBUTING.md for how to add new model adapters, CUDA kernels, and platform support.

Community priorities:

  • ROCm HIP kernel port (AMD R9700 full compute)
  • GUI improvements and screenshots
  • DGX Spark benchmarks
  • Real benchmark submissions from different hardware
  • Kimi K3 tensor names from a real checkpoint (wisp inspect)
  • Model integrity verification improvements

License

MIT License — see LICENSE.

Built by Saksham (Zero_planck) github.com/zeroextub-collab/wisp

Releases

Packages

Contributors

Languages