Stream what shouldn't run.
WISP runs the largest open-source AI models ever built on the hardware sitting on your desk.
A 744B parameter frontier model. A 671B reasoning engine. A 2.8 trillion parameter behemoth, weights already public.
Not a slow demo. Not quantized to uselessness. The full model, at full frontier intelligence, streaming expert weights across GPU VRAM, system RAM, and your NVMe SSD in real time.
Ten ships across August–September 2026. Where something is built but not compiled, planned but not run on real weights, or a preview, this section says so.
CPU-only mode.
AVX2 + FMA SIMD kernels, selected at runtime, and a two-tier RAM → NVMe
hierarchy: any x86_64 CPU with AVX2 (Intel since 2013, AMD since 2015) runs
WISP without a GPU. Measured on the build machine: 57.6 GB/s all-core RAM
bandwidth. From that bandwidth wisp doctor estimates Mixtral-8x7B at
~6 tok/s once its experts are in RAM — an estimate, not a benchmark; a
CPU-only run on real weights has not been timed yet.
Vulkan backend (compile with the Vulkan SDK).
Compute shaders designed for AMD R9700 (RDNA4), the Ryzen AI Halo iGPU,
Intel Arc and any Vulkan 1.2 GPU, with a zero-copy path on unified-memory
devices that maps device memory directly instead of staging copies.
Build with cmake -DWISP_VULKAN=ON (needs the Vulkan SDK's headers and
glslc). Detection and auto-config work on every build; the backend itself
has not been compiled or run on hardware yet.
New model adapters.
DeepSeek-V4-Flash (284B / 13B active, 258 lookups/token) and GLM-5.3-Flash
(320B / 18B active) are planning-only: wisp info sizes hardware for them
and wisp convert refuses them before downloading. GLM-5.3-Flash's fp8
block scales are already decoded; its hyper-connections are not.
Qwen3-Coder-480B-A35B (480B / 35B active, Apache 2.0) converts through the
standard qwen3_moe path.
fp8-e4m3 decode. The official DeepSeek-V3 and R1 checkpoints ship fp8-e4m3 weights with a 128×128 block scale for each. WISP multiplies every block by its scale at convert time and refuses any fp8 weight whose scale is missing. Tested on synthetic checkpoints; a full 671B conversion has not been run yet.
Honest sizes. Expert and dense sizes are derived from each model's published config rather than borrowed estimates, and download sizes are read from the Hugging Face file listings. GLM-5.2's dense weights (~34.6 GB at fp16) and DeepSeek's (~31.8 GB) exceed consumer GPU VRAM. Dense quantization is planned for v3.1.
Convert preflight.
Before a single byte downloads: a disk-space check covering download plus
converted output, refusal of planning-only models, a refusal to overwrite a
converted model, and a [y/N] summary of what is about to happen. --yes
skips the prompt for scripts; --force re-converts or repairs.
Kimi K3 tensor names verified.
Checked against the real moonshotai/Kimi-K3 checkpoint index and its
modeling code: layers live under
language_model.model.layers.N, routed experts under
block_sparse_moe.experts.E.w1/w2/w3. Layer 0 is a dense MLP, so
1,472 lookups/token (92 MoE layers × top-16). The experts are
MXFP4-packed, so wisp convert refuses K3 until MXFP4 decode lands —
along with the latent-MoE and released-KDA pieces listed under
KDA.
Web dashboard.
/dashboard: one self-contained HTML page, zero external dependencies, so
it works air-gapped. The expert heatmap draws every expert as a cell,
colored by tier, with a white flash on each activation. Verified live on
Mixtral-8x7B (2026-09-12): 64 experts flash per token, hit rate 35% →
65.5% over 125 tokens, TTFT 23.5 s cold.
Multi-node cluster (preview).
Coordinator/worker architecture: a shard planner divides expert layers
across nodes, workers serve their expert files, and the coordinator checks
them at startup (--cluster-role, --cluster-nodes,
--cluster-coordinator), exposes /v1/cluster and shows a node panel on
the dashboard. In v3.0 inference still reads experts from local disk;
remote worker expert fetch is v3.1.
Speculative decoding is opt-in.
--speculative enables it, --no-speculative skips every drafter check.
Nothing downloads on a first request: WISP asks, with the size, before
fetching a drafter, and is_drafter_complete() rejects partial downloads.
| Model | Params | Active/token | Lookups/token | Disk (int4) | Status |
|---|---|---|---|---|---|
| GLM-5.2 | 744B | 40B | 600 | ~414 GB | ✅ Ready |
| DeepSeek-V3 | 671B | 37B | 464 | ~374 GB | ✅ Ready — fp8† |
| DeepSeek-R1 | 671B | 37B | 464 | ~374 GB | ✅ Ready — fp8† |
| Mixtral-8x7B | 47B | 13B | 64 | 26.6 GB* | ✅ Verified end-to-end |
| Mixtral-8x22B | 141B | 39B | 112 | ~90 GB | ✅ Ready |
| Kimi K3 | 2.8T | 104B | 1,472 | ~1.4 TB est. | |
| Qwen3-235B | 235B | 22B | 752 | ~130 GB | ✅ Ready |
| Qwen3-2.4T | 2.4T | ~22B est. | TBD | ~1.2 TB est. | ✅ Adapter ready — weights pending |
| Qwen3-Coder-480B | 480B | 35B | 496 | ~268 GB | ✅ Ready |
| DeepSeek-V4-Flash | 284B | 13B | 258 | ~158 GB | 📐 Planning only |
| GLM-5.3-Flash | 320B | 18B | 336 | ~178 GB | 📐 Planning — fp8 ✓, hyper-connections pending |
| GLM-5.3 | 744B | ~40B | 600 | ~414 GB | ⚡ Weights released — adapter in progress§ |
*Measured from a real conversion — every expert file verified. "Verified end-to-end" means: downloaded, converted, and generated real tokens through the full C + CUDA engine on consumer hardware.
†fp8-e4m3 dequantized at convert time: e4m3 values plus a float32
weight_scale_inv holding one multiplier per 128×128 block. Before
2026-09-11 wisp convert dropped those scales, and its expert pattern also
matched them as expert weights, so the converted model verified clean and
generated garbage. Tested on synthetic checkpoints; full conversion not yet
verified on real 671B weights. DeepSeek's sizes are derived from the
official config — a 24.8 MB expert (3 × 2048 × 7168 weights in WISP's int4)
— not yet measured from a conversion.
‡Dense weights for GLM-5.2 (~34.6 GB) and DeepSeek (~31.8 GB) at fp16
exceed consumer GPU VRAM. Dense quantization: v3.1. Current workaround: run
with --cpu-only on a machine with 64 GB+ RAM, which keeps the dense
weights in system memory; WISP's planner fits them at 64 GB, not at 48 GB.
§GLM-5.3 is planned against GLM-5.2's shape: Z.ai describes it as the same
base model and the two published configs match. The adapter is still a
stub, so wisp convert refuses it.
Lookups count routed experts only: shared experts run for every token, so
they stay with the dense weights and are never streamed. Disk is WISP's
int4 output, computed from the published shapes. It is not the download
size. Qwen3-Coder-480B, DeepSeek-V4-Flash and GLM-5.3-Flash were checked
against each model's published config.json and model card on 2026-09-11;
Kimi K3's tensor names and layer layout against moonshotai/Kimi-K3 on
2026-09-14.
- Qwen3-Coder-480B is the same
qwen3_moearchitecture as Qwen3-235B, so it takes the same converter path at a bigger shape (160 experts, 62 layers). - DeepSeek-V4-Flash and GLM-5.3-Flash are planning only.
V4-Flash ships fp4 experts under DeepSeek's own tensor names and uses
CSA/HCA compressed attention. GLM-5.3-Flash uses a KDA/DSA hybrid and a
vision tower. Both use mHC hyper-connections, which the engine does not
implement yet. GLM-5.3-Flash's fp8 weights use the same
weight_scale_invblock scales as DeepSeek-V3, which the converter decodes. V4-Flash names its scales.scale, and that layout is not wired in.
July 2026 — The Biggest Week in Open Source AI
Kimi K3 (2.8T parameters) dropped July 17 — the largest open source model ever released. Qwen3.8 (2.4T parameters) announced July 19 — open weights coming soon. WISP was built for exactly this moment. Both are MoE models. Both stream with WISP.
Every MoE (Mixture-of-Experts) model activates only a small fraction of its parameters for each token. GLM-5.2 activates ~5.4%. Kimi K3 activates 3.7% (104B of 2.8T). WISP exploits this with a self-organizing 3-tier cache:
Token arrives
↓
Model's own router selects which experts to activate
↓
WISP checks VRAM first → instant if hit
↓
Then RAM → one PCIe copy if hit
↓
Then NVMe SSD → streams if cold
↓
LRU cache promotes hot experts upward automatically
After 10-15 min: high cache hit rate, self-organized
(measured: 68.8% hit rate after only 80 tokens)
No configuration. No preset modes. Fully automatic.
GLM-5.2 and DeepSeek use Multi-head Latent Attention (MLA).
WISP implements true absorbed MLA — the KV cache stores the
compressed c_kv latent instead of expanded K,V tensors.
Result: ~70KB per token KV cache instead of ~5MB.
This is what makes 1M-token context feasible in RAM.
Kimi K3 runs linear attention in 69 of its 93 layers. Instead of a KV
cache that grows with the conversation, each head carries a fixed
[d_k, d_v] state matrix updated by a delta rule:
qn, kn = q/||q||, k/||k|| L2-normalized per head
b = sigmoid(W_beta x) per-channel write gate, in (0,1)
u = S^T kn what memory currently holds at this key
S = S + b * kn (v - u)^T write the PREDICTION ERROR
o = S^T qn read, using the new state
y = silu(g) * o output gate
Three details carry the whole thing:
It writes v - u, not v. That is what makes it a delta rule —
the state is corrected by exactly the amount it was wrong by, so
re-writing a key that is already stored is a no-op instead of doubling
it. Plain linear attention (S += v k^T) accumulates and saturates.
q and k are L2-normalized, which bounds the update. Without it the
state diverges: measured at inf within 200 tokens at moderate input
scale. A constant-size state is worthless if its contents blow up.
The read happens after the write, so a token can see its own value. Reading first makes the first token of every conversation emit exactly zero.
Cost is O(n) in sequence length and memory is constant — that is the property that makes 1M-token context tractable at all.
WISP implements KDA as a CUDA kernel with a matching pure-PyTorch fallback for CPU-only mode. The kernel normalizes internally rather than trusting its callers, so the C engine, the Python bindings and the fallback cannot drift apart. Tests assert the two paths agree numerically at head_dim 32, 100 and 200, that prefill lands on the same state as sequential decode, that repeated writes converge on the stored value, that the state stays bounded over 2,000 tokens, and that batch entries do not contaminate each other. The other 24 K3 layers are Gated MLA and already run through WISP's absorbed-MLA path.
Status: ✅ Kernel complete. Tensor names verified against
moonshotai/Kimi-K3 (2026-09-14). Blocked on MXFP4 expert decode — v3.1.
MXFP4 is the first blocker, not the only one: the released checkpoint also
runs its experts in a 3,584-dim latent (routed_expert_down_proj /
routed_expert_up_proj), and its KDA layers add short convolutions and a
learned decay (A_log, dt_bias, f_a_proj, f_b_proj) beyond the
recurrence above. wisp convert refuses K3 until those exist, before
downloading anything.
While the GPU computes token N (2-8ms), the C engine loads token N+1's predicted experts from SSD into pinned RAM, and the transfer stream moves them to VRAM (0.1-0.3ms) — fully hidden inside the compute window. GPU never waits on predicted experts. Works on all consumer hardware. No GPUDirect Storage required.
Same-family small models draft 3 tokens simultaneously. The main model verifies all 3 in one parallel forward pass. At 39-55% acceptance rate: 2.2-2.8x effective throughput. Zero quality loss — the rejection-sampling scheme provably preserves the main model's output distribution (Leviathan 2023).
Opt-in: pass --speculative to enable.
Default is off — no automatic downloads.
WISP detects whether your monitor is on the GPU or the
motherboard (via the driver itself) and reserves VRAM
accordingly. Move the monitor to the motherboard port →
WISP auto-detects → full VRAM dedicated to inference.
Override anytime with --display-mode gpu|igpu|auto.
WISP records which experts your sessions actually activate, in
{model_dir}/.wisp_usage. On the next startup the hottest ones are
queued for pre-warming before the first token is generated.
session 1 cold start; the LRU discovers your domain
session 2 last session's top experts pre-warmed at startup
session 7 near-instant warm start
Expert selection happens inside the C router, so the engine keeps a ring-buffer log of every (layer, expert) access that the runtime drains every 10 tokens — that stream feeds both the next-token prefetch predictor and the cross-session cache. Measured on a real Mixtral run: 12 tokens produced 768 expert observations and 238 tracked experts, of which 107 were pre-warmed on the following startup.
wisp cache --model ./models/glm-5.2/ --show # what it has learned
wisp cache --model ./models/glm-5.2/ --reset # start overRanking blends frequency with a recency decay, so a cache trained on three weeks of Rust adapts when you switch to prose instead of staying stuck on the old domain. The file is plain JSON — inspect it, diff it, or delete it.
wisp serve --model ./models/glm-5.2/ --port 8080Anything that speaks the OpenAI API now speaks to WISP — Cursor,
Continue.dev, Open WebUI, LM Studio frontends, the openai package:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8080/v1", api_key="wisp")
client.chat.completions.create(
model="glm-5.2",
messages=[{"role": "user", "content": "Write a quicksort"}],
stream=True,
)/v1/chat/completions (streaming + non-streaming), /v1/models,
/health, /v1/stats for WISP's own numbers (tok/s, tier hit
breakdown, learning-cache state, TTFT), /v1/expert-state for the live
expert map, and the web dashboard at /dashboard.
Streaming is genuinely incremental —
the engine's generator runs on a worker thread feeding the event loop,
so the first token reaches the client as soon as it exists rather than
after the whole completion. Prompts are rendered with each family's own
chat template. Requests are serialized: one engine, one KV cache, so
concurrent decoding would interleave two conversations.
Install the extra: pip install -e '.[server]'
- System RAM is never allowed to fill:
max(6GB, 25%)is always reserved for the OS, and a runtime watermark evicts RAM-tier experts past 80% usage (SSD stays authoritative, so eviction is always safe). - VRAM planning is hard-capped at 75% of the card, and the allocator degrades gracefully (evict-and-retry) if the driver's real ceiling is lower than the plan.
wisp serve --model ./models/glm-5.2/ --port 8080Open http://localhost:8080/dashboard in any browser.
Works from any device on the same network when the server is started with
--public — open http://<this machine's IP>:8080/dashboard there.
--public has no authentication: anyone who can reach the port can use
your GPU, so only do this on a network you trust.
Zero external dependencies — works on air-gapped machines.
Expert heatmap, one row per layer and one column per expert:
Blue = in VRAM (fast)
Amber = in RAM (medium)
Grey = on NVMe (cold)
White = activated this token (flashes, fades over 800ms)
Brighter means an expert has fired more often; hover a cell for its layer, expert, tier and heat. Tiers come from the C caches themselves and heat from the learning cache. Until the model has loaded, the heatmap is empty rather than simulated.
Live stats: tok/s, cache hit rate, TTFT, token count.
Tier monitor: VRAM / RAM / NVMe usage bars.
Streaming chat: test inference directly from the browser.
Everything it draws is plain JSON at /v1/expert-state and /v1/stats.
Verified on Mixtral-8x7B (2026-09-12): 64 experts flash per token. Hit rate 35% → 65.5% over 125 tokens. TTFT 23.5s cold start.
wisp-gui is a native desktop app that connects to the wisp serve
API server running locally.
pip install -e '.[gui]'
wisp-guiPoint it at your converted model directory. The GUI starts wisp serve
internally on a private loopback port.
Features:
- Token stream with live tok/s counter
- Tier monitor: VRAM / RAM / NVMe usage live from
/v1/stats - Cache hit rate and expert observation count
- Temperature, max_tokens, system prompt controls
- Dark theme
Three columns: model + generation controls on the left, chat in the middle, live tier and engine monitors on the right. The tier bars show VRAM and RAM occupancy against this machine's real capacities, refreshed every two seconds.
The GUI is not a wrapper around the CLI — it never shells out or parses
terminal output. It hosts the very same WispServer in a background
thread and talks to it over the OpenAI API, so the path it exercises is
byte-for-byte the one Cursor and Open WebUI use. Closing the window
shuts the server down cooperatively, which is what lets the learning
cache persist.
Screenshots welcome — open a GitHub Issue to submit yours.
Standard 3-tier: VRAM → RAM → NVMe. RTX 3080+ recommended. RTX 5070 (12GB) verified end-to-end on Mixtral-8x7B. CUDA 12.0+ | 13.x. CUDA 12.8+ required for RTX 50 series.
This is the only path that has been run end-to-end on real weights.
Compute shaders designed for:
- AMD R9700 (RDNA4) via Mesa/RADV or the AMD driver
- AMD Ryzen AI Halo iGPU (Radeon 8060S, 40 CUs)
- Intel Arc via the Intel Vulkan driver
- Any GPU with a Vulkan 1.2 driver
Build: cmake -DWISP_VULKAN=ON (requires the Vulkan headers, glslc, and
the Vulkan loader — all in the Vulkan SDK).
Detection and auto-config work on all builds.
Pass --vulkan to force the Vulkan path.
The backend has not been compiled or run on hardware yet; compare its
output against the CUDA or CPU path before trusting it.
Up to 128GB LPDDR5x unified pool, Radeon 8060S iGPU. WISP detects it automatically (the 8060S by name, an integrated device, and a pool of at least 64GB) and plans 2-tier mode: unified pool → NVMe SSD. Zero-copy GPU access via Vulkan (requires a Vulkan build); without one, WISP runs CPU-only. Two Halo units over the TCP cluster would plan a 256GB pool — see the cluster preview for what that does today.
128GB coherent unified pool, 273 GB/s, GB10. ARM64 (aarch64). $4,699. WISP detects it automatically — all three of aarch64, a GB10-class device name, and a ~128GB reported pool must agree — and switches to 2-tier unified mode: CPU and GPU share the same physical memory, so the RAM→VRAM copy the double buffer exists to hide simply does not happen; prefetch from NVMe still matters. Sparks carry ConnectX-7 200 GbE, the link the 2–4 node cluster is planned around: 2 nodes = 256GB, 4 nodes = 512GB of pooled memory once remote expert fetch lands (v3.1).
wisp doctor # shows "DGX Spark detected — unified memory mode"Detection and tier planning are implemented; no Spark has been benchmarked. The figures above are hardware specifications, not WISP measurements.
Any x86_64 with AVX2 (Intel since 2013, AMD since 2015).
2-tier: RAM → NVMe. No CUDA or Vulkan needed.
Use: wisp serve --cpu-only
Estimates from wisp doctor, which measures RAM bandwidth (57.6 GB/s on a
32GB DDR5-6000 build machine) — not timed runs:
- Fastest CPU model: Mixtral-8x7B, ~6 tok/s once its experts are in RAM (25% of its experts fit in RAM on that 32GB machine)
- GLM-5.2 CPU-only: ~1.15 tok/s with experts in RAM, ~0.36 tok/s cold
32GB GDDR6. 640 GB/s. 1531 TOPS INT4. PCIe 5.0. $1,299. Dual-slot blower — built for dense multi-GPU stacking.
WISP v2.0 detects R9700 via rocm-smi and sets auto-config.
Full ROCm compute kernels (HIP port of CUDA kernels): v3.1, alongside the
Vulkan path above. Stacking projections are in the
R9700 table.
ROCm contributions welcome — see CONTRIBUTING.md.
Setup: scripts/install_rocm.sh (Ubuntu) or install_rocm.ps1
(Windows). Both are detection-only; they do not make inference run on
the card.
Coordinator/worker architecture for 2–4 nodes.
The coordinator serves the normal OpenAI API.
Workers hold their shard of the expert pool.
Clients see a standard /v1/chat/completions endpoint.
v3.0: coordinator and worker architecture, shard planner,
/v1/cluster endpoint, dashboard node panel.
Expert data is still served from local disk: the coordinator runs inference
from its own full copy of the model, and workers are planned and
health-checked only.
v3.1: remote worker expert fetch via HTTP — the full distributed pool becomes active.
What works today:
- Shard planning: layers are split into contiguous ranges, as evenly as the count allows (75 layers over 4 workers: 19, 19, 19, 18).
- Workers serve their expert files over HTTP (
POST /shard/experts) and report the layers they actually hold, their memory and the model (GET /shard/health). - The coordinator checks every worker at startup, prints the plan and
the pooled memory, exposes
GET /v1/cluster, and the dashboard shows each node as online or unreachable.
What does not work yet:
- Expert routing. The C engine reads experts from its own model directory, and nothing feeds it bytes fetched from a worker. Routing needs a C-side expert-source hook.
- RDMA. The transport is HTTP/JSON over TCP;
--cluster-fabric rdmais refused rather than silently falling back. - Authentication. A worker serves expert weights to anyone who can reach its port. Run clusters on a network you trust.
Node 0 (coordinator):
wisp serve --model ./models/glm-5.2/ \
--cluster-role coordinator \
--cluster-nodes http://spark1:8080,http://spark2:8080Nodes 1+ (workers — --public so the coordinator can reach them):
wisp serve --model ./models/glm-5.2/ --public \
--cluster-role worker --cluster-node-id 1 \
--cluster-coordinator http://spark0:8080Memory pools (projected, for when remote fetch lands):
- 2 × DGX Spark: 256 GB — about two-thirds of GLM-5.2's experts (19,200 × 21.2 MB ≈ 380 GB)
- 4 × DGX Spark: 512 GB — all of GLM-5.2's experts
- Kimi K3's experts (~1.4 TB) exceed even four nodes, and K3 does not convert yet
For scale, not a measurement: a fully cold GLM-5.2 token needs 600 experts × 21.2 MB = 12.7 GB. Over a 200 Gb/s link (25 GB/s line rate) that is at least half a second, before HTTP overhead and base64's extra third.
Same commands. Inter-node bandwidth: ~1.25 GB/s (10 GbE), a tenth of a Spark link, so a fully cold GLM-5.2 token over the wire would take roughly ten seconds.
First verified on: R7 9800X3D | RTX 5070 12GB | 32GB DDR5-6000 | PCIe 4.0 NVMe
| Model | Cold | Warm | Hot | +MTP Effective |
|---|---|---|---|---|
| Mixtral-8x7B | 0.75 tok/s ✅ measured | est. 2-5 tok/s* | est. 5-10 tok/s* | est. ~2x* |
| GLM-5.2 | est. 0.7 tok/s* | est. 5.5 tok/s* | est. 8.5 tok/s* | est. ~14 tok/s* |
| DeepSeek-V3/R1 | est. 0.8 tok/s* | est. 5.8 tok/s* | est. 9.0 tok/s* | est. ~14.5 tok/s* |
| Kimi K3 | est. 0.2 tok/s* | est. 3 tok/s* | est. 6 tok/s* | est. ~10 tok/s* |
Cold = empty cache, first run. Warm = 15 min same domain. Hot = repeated patterns, cache fully warmed. MTP = with speculative decoding active. * = estimated from physics (expert size × lookups ÷ transfer bandwidth); real benchmarks land in v2.1. Measured Mixtral data (2026-07-19): 81 tokens at 0.75 tok/s from a completely cold engine, 68.8% cache hit rate after 80 tokens; a 300-token cold run averaged 0.56 tok/s on an earlier engine build. We do not publish a bold number we didn't measure.
Cold → warm progression, every figure measured (Mixtral-8x7B, R7 9800X3D + RTX 5070):
2026-07-19 cold engine 81 tokens averaging 0.75 tok/s
68.8% cache hit rate after 80 tokens
2026-09-12 cold start TTFT 23.5 s (prefill streams experts off NVMe)
live run hit rate 35% → 65.5% over 125 tokens,
0.56 tok/s engine average, plain decoding
Warm and hot steady-state rates on long runs have not been measured yet.
Mixtral 8x7B experts = 99MB each (measured). GLM-5.2 experts = 20.2MB each — 4.9× smaller.
Cold decode is transfer-bound: every uncached expert crosses PCIe or comes off the NVMe. Smaller experts = less data per token = faster streaming and faster cache warm-up. Mixtral is WISP's proving ground; GLM-5.2's fine-grained experts are where the architecture truly sings.
git clone https://github.com/zeroextub-collab/wisp
cd wisp
pip install -e . # builds the C + CUDA engine (see Installation)
wisp doctor # confirm the toolchain found everything
wisp convert --model glm-5.2 --output ./models/
wisp chat --model ./models/glm-5.2/# Desktop GUI
pip install -e '.[gui]'
wisp-gui# CPU-only (no GPU needed)
wisp chat --model ./models/mixtral-8x7b/ --cpu-only
# Web dashboard
wisp serve --model ./models/mixtral-8x7b/ --port 8080
# then open http://localhost:8080/dashboard
# Speculative decoding (opt-in; asks before downloading the ~14 GB drafter)
wisp chat --model ./models/mixtral-8x7b/ --speculativeNot on PyPI yet. WISP builds a CUDA extension against your local toolchain, so it installs from source today —
pip install wisp-enginewill not work. (Note also that the namewispon PyPI belongs to an unrelated project.) Prebuilt wheels are tracked in the roadmap.
| Component | Minimum | Recommended |
|---|---|---|
| RAM | 16 GB | 32 GB+ |
| NVMe SSD | PCIe 3.0, 300 GB free | PCIe 5.0, 2 TB dedicated |
| GPU | None (CPU-only works) | RTX 3080+ / 8 GB+ VRAM |
| CUDA | 12.0+ | 12.8+ (required for RTX 50 series) |
| Python | 3.10+ | 3.11 |
| OS | Windows 10+ / Ubuntu 20.04+ | Windows 11 / Ubuntu 22.04 |
| Vulkan SDK | 1.2+ | optional — only for a -DWISP_VULKAN=ON build |
| OpenMP | ships with the compiler (MSVC's 2.0 is enough) | required for -DWISP_CPU_ONLY=ON builds |
git clone https://github.com/zeroextub-collab/wisp
cd wisp
powershell -File scripts\install.ps1git clone https://github.com/zeroextub-collab/wisp
cd wisp
bash scripts/install.shpip install torch --index-url https://download.pytorch.org/whl/cu128
pip install -e .
wisp doctor # verify everything works# Convert a model (downloads + converts to WISP format,
# resumable, SHA256-verified, integrity-checked at the end)
wisp convert --model glm-5.2 --output ./models/
# One-shot inference
wisp run --model ./models/glm-5.2/ \
--prompt "Write a Python web scraper" \
--stream
# Interactive chat (/clear /stats /quit)
wisp chat --model ./models/glm-5.2/
# Benchmark your hardware (cold -> warm -> hot)
wisp benchmark --model ./models/glm-5.2/ --runs 3
# Check system compatibility
wisp doctor
# Show tier allocation for your hardware (works pre-download)
wisp info --model glm-5.2
# Print the tensor names a checkpoint actually uses, and which of them
# WISP matched (use before converting an unreleased/unverified model)
wisp inspect --source ./shards --model kimi-k3
# Verify model integrity, expert file by expert file
wisp verify --model ./models/glm-5.2/
# Share a converted model so others skip re-conversion
wisp upload --model ./models/glm-5.2/ --repo you/glm-5.2-wisppip install -e '.[gui]'
wisp-guiOpens the desktop app; point it at a converted model directory.
from wisp import WispEngine
engine = WispEngine("./models/glm-5.2/")
# Streaming
for token in engine.stream("Explain quantum entanglement"):
print(token, end="", flush=True)
# One-shot
result = engine.generate("Write a sorting algorithm",
max_new_tokens=500,
temperature=0.7)
print(result)WISP automatically detects and configures multiple GPUs. Zero manual configuration needed.
| Setup | Strategy | Best For |
|---|---|---|
| Single GPU 8-12GB | Dense + expert LRU cache | Getting started |
| Single GPU 24GB+ | Dense + large expert cache | GLM-5.2 smooth |
| Dual GPU same size | GPU0 dense, GPU1 pure cache | 2× cache hits |
| Dual GPU diff sizes | Bigger=dense, smaller=overflow | Flexible |
| 3+ GPUs | Pipeline parallelism | Maximum throughput |
| Setup | Combined VRAM | GLM-5.2 | Kimi K3 |
|---|---|---|---|
| 2× RTX 4090 | 48GB | Smooth warm | Feasible |
| 4× RTX 4090 | 96GB | Near full cache | Good |
| 4× RTX 6000 Ada | 192GB | Full expert cache | Strong |
AMD Radeon AI PRO R9700 (ROCm detection + auto-config in v2.0. Full HIP compute kernels + Vulkan: v3.1)
The R9700 is purpose-built for exactly what WISP does: 32GB GDDR6 per card, dual-slot blower for dense stacking, PCIe 5.0, 1531 TOPS INT4.
| R9700 Stack | Combined VRAM | GLM-5.2 Coverage | Projected tok/s |
|---|---|---|---|
| 1× R9700 | 32GB | Dense + 1,257 experts | 8-12 |
| 2× R9700 | 64GB | Dense + 3,085 experts | 18-25 |
| 4× R9700 | 128GB | Dense + 6,741 experts | 38-52 |
Note: WISP now detects the R9700 through ROCm —
wisp doctorreports the card, its VRAM, and the gfx1201 target, and the tier planner sizes for it. It does not run on it: the compute kernels are CUDA and the HIP port is separate work. All R9700 throughput figures above remain projections from the tier math, not measurements.scripts/install_rocm.shsets up detection. Contributions welcome.
User prompt
↓
WISP CLI / Python API
↓
Auto-Config Engine
(profiles hardware once, calculates optimal tier split,
detects display on GPU/iGPU, sets safe VRAM/RAM budgets)
↓
Universal Runtime (C + CUDA)
├── Double-Buffer Async Prefetch
│ hides SSD transfer inside GPU compute time
├── 3-Tier LRU Cache
│ VRAM (fastest) → RAM (fast) → SSD (cold)
│ scratch-ring serving + hit-gated promotion:
│ cool experts never churn the VRAM cache
├── Speculative Decoding
│ same-family drafter → 2.2-2.8x throughput
├── Multi-GPU Router
│ auto-selects: single / dual / pipeline strategy
└── RAM Watermark Monitor
evicts experts past 80% RAM — the desktop never starves
↓
The model's own router drives expert selection.
WISP just delivers them as fast as possible.
Three layers, one job each: Python orchestrates (download, convert, configure — things that run once), C owns the hot path (64-1,472 expert fetches per token, cache coordination, prefetch threads), CUDA owns the math (absorbed MLA / GQA / KDA attention, int4 dequant, fused SwiGLU FFN, router top-K). The engine is model-agnostic: adapters map every family onto one canonical weight layout at conversion time, so adding a model is one adapter file — not a new engine.
On a coherent-memory machine the tier planner emits a two-tier plan instead (unified pool → NVMe); nothing else in the stack changes.
- Universal 3-tier MoE streaming engine (C + CUDA)
- Mixtral-8x7B and 8x22B full support — end-to-end verified
- GLM-5.2 (744B), DeepSeek-V3/R1 (671B) adapters
- Absorbed MLA attention (~70KB/token KV cache)
- KDA linear attention kernel (CUDA + PyTorch fallback)
- Double-buffer async prefetch pipeline
- Speculative decoding (39-55% acceptance, 2.2-2.8x throughput)
- OpenAI-compatible API server (
wisp serve) - Learning cache (cross-session expert pre-warming)
- Display auto-detection (VRAM conflict prevention)
- Multi-GPU support (single / dual / pipeline)
- 121 tests
- Kimi K3 (2.8T): KDA converter weight mapping + C forward pass
branch. Kernel numerically verified; checkpoint tensor names
verified (2026-09-14), with
wisp inspectandWISP_KDA_NAMESas the escape hatch - Qwen3-235B and Qwen3-2.4T adapters
- Desktop GUI (
wisp-gui) — OpenAI client + live dashboard - DGX Spark unified memory mode (auto-detected, not benchmarked)
- AMD R9700 ROCm detection + auto-config (detection only)
- 197 tests
- CPU-only mode: AVX2+FMA kernels, RAM→NVMe 2-tier
- Vulkan backend: shaders + detection (compile with SDK; not yet run on hardware)
- fp8-e4m3 decode with 128×128 block scales
- DeepSeek V4 Flash adapter (284B/13B, 258 lookups; planning only)
- Qwen3-Coder-480B adapter (Apache 2.0)
- GLM-5.3-Flash planning (fp8 unblocked)
- Convert preflight: disk check, [y/N], --yes, --force
- Kimi K3 tensor names verified (moonshotai/Kimi-K3)
- Web dashboard: /dashboard, expert heatmap, zero CDN
- Multi-node cluster: coordinator/worker, preview
- Speculative decoding opt-in (--speculative)
- Honest sizes: derived from published configs, download sizes from the Hub
- 436 tests
- Dense layer quantization (GLM-5.2 + DeepSeek on consumer hardware — the last major blocker)
- GLM-5.3 full adapter (weights released 2026-08-25)
- MXFP4 decode (the first of Kimi K3's blockers; latent MoE and the released KDA layer remain)
- Remote worker expert fetch (completes cluster)
- Full ROCm HIP + Vulkan compiled kernels (AMD R9700)
- GLM-5.3-Flash full support (hyper-connections)
- macOS / Metal backend
- Generic MoE adapter (any HuggingFace MoE model)
- GPUDirect Storage (Linux, datacenter GPUs)
- Plugin system for community model profiles
- Expert atlas (crowd-sourced topic affinity map)
- Real GLM-5.2 benchmark numbers
- Real DeepSeek-V3 benchmark numbers
- Warm/hot steady-state Mixtral numbers (long runs)
- A timed CPU-only run on real weights, to replace the
wisp doctorestimates - A Vulkan backend build and run on any GPU
- A DGX Spark run, to replace the unified-mode projections
WISP would not exist without Colibrì.
JustVugg built Colibrì in July 2026 — a 2,400-line pure-C engine that proved a 744B parameter model could run on 25GB of consumer RAM by streaming expert weights from disk. Before Colibrì, everyone said this was impossible. After Colibrì, we built WISP.
jlnsrk converted the GLM-5.2 weights to a format the community could actually use.
matey-0 (Mateo Grgić) fixed the MTP head from int4 (0-4% acceptance) to int8 (39-59% acceptance) — turning speculative decoding from broken to genuinely useful.
→ github.com/JustVugg/colibri
WISP shares zero code with Colibrì. Complete independent reimplementation. JustVugg showed us what was possible.
- Leviathan et al. 2023 — Speculative Decoding
- GLM team — GLM-5.2 and IndexShare MoE architecture
- DeepSeek team — DeepSeek-V3/R1 MoE + Multi-head Latent Attention
- Moonshot AI — Kimi K3: Open Frontier Intelligence (arXiv:2607.24653) — KDA hybrid linear attention, Stable LatentMoE routing
- The llama.cpp community — proof that consumer hardware deserves frontier models
Full tribute: CREDITS.md
See CONTRIBUTING.md for how to add new model adapters, CUDA kernels, and platform support.
Community priorities:
- ROCm HIP kernel port (AMD R9700 full compute)
- GUI improvements and screenshots
- DGX Spark benchmarks
- Real benchmark submissions from different hardware
- Kimi K3 tensor names from a real checkpoint (
wisp inspect) - Model integrity verification improvements
MIT License — see LICENSE.
Built by Saksham (Zero_planck) github.com/zeroextub-collab/wisp
