diff --git a/CHANGELOG.md b/CHANGELOG.md index a3758637..73c742e2 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,6 +9,10 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ## [Unreleased] +## [0.56.1] — 2026-09-21 + +A transformers-only release against **SKaiNET engine 0.56.0** (unchanged). + ### Added — Moonshine v2 streaming: every checkpoint of the family, exported from the published artifact - **`MoonshineV2ExportCli`** (`:llm-inference:moonshine:exportMoonshineV2`, JVM): a Hugging Face diff --git a/README.md b/README.md index a2bb80bd..de2f7a96 100644 --- a/README.md +++ b/README.md @@ -109,8 +109,19 @@ Honest status — see the project-status note at the top of this README. ## Current release -The current release is **0.56.0**, back in lock-step with **SKaiNET 0.56.0** (the engine skipped -0.55.0 to realign the two version lines). +The current release is **0.56.1** (against **SKaiNET 0.56.0** — a transformers-only release, same +pattern as 0.54.1: no new engine version needed). + +**Moonshine v2 streaming, every checkpoint of the family, exported from the published module.** +`MoonshineV2ExportCli` turns a Hugging Face `moonshine_streaming` snapshot into the five StableHLO +graphs of the streaming contract (frontend, encoder, adapter, masked prefill, dynamic-cache step) +plus the host-side embedding and vocabulary tables — no Python, nothing to configure per language +or size, and it fails if a checkpoint tensor goes unused. The model gains per-layer attention bands +(the German checkpoints need them) and split encoder/decoder widths (the small ones). See +*Export Moonshine v2 Streaming to StableHLO* in the docs. + +It builds on **0.56.0**, which put the two version lines back in lock-step with **SKaiNET 0.56.0** +(the engine skipped 0.55.0 to realign them). **Grouped-query attention without head expansion, and the compiled leg of SKEEP-005.** The engine's `scaledDotProductAttention` is grouped-query native, so `MultiHeadAttention` and @@ -243,7 +254,7 @@ The recommended way to consume is via the BOM. It pins every published `skainet- ```kotlin dependencies { - implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.0")) + implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.1")) // Versions resolved from the BOM: implementation("sk.ainet.transformers:skainet-transformers-core") diff --git a/docs/modules/ROOT/nav.adoc b/docs/modules/ROOT/nav.adoc index ff388395..1b71bb16 100644 --- a/docs/modules/ROOT/nav.adoc +++ b/docs/modules/ROOT/nav.adoc @@ -18,6 +18,7 @@ * xref:how-to/run-unified-cli.adoc[Use the Unified CLI] * xref:how-to/benchmarking.adoc[Run Benchmarks] * xref:how-to/compile-model-for-android.adoc[Compile a Model for Android (IREE)] +* xref:how-to/export-moonshine-v2-streaming.adoc[Export Moonshine v2 Streaming to StableHLO] .Reference * xref:reference/architecture.adoc[Architecture Overview] diff --git a/docs/modules/ROOT/pages/how-to/export-moonshine-v2-streaming.adoc b/docs/modules/ROOT/pages/how-to/export-moonshine-v2-streaming.adoc new file mode 100644 index 00000000..b8d90b58 --- /dev/null +++ b/docs/modules/ROOT/pages/how-to/export-moonshine-v2-streaming.adoc @@ -0,0 +1,100 @@ += Export Moonshine v2 Streaming to StableHLO +:description: Turn a Hugging Face moonshine_streaming checkpoint into the five StableHLO graphs of the streaming speech-to-text contract — from the published module, without Python. + +`skainet-transformers-inference-moonshine` carries the Moonshine **v2 streaming** model in the SKaiNET NN DSL +and, since **0.56.1**, a JVM command that exports it: a Hugging Face snapshot in, five StableHLO graphs and two +host-side tables out. Any StableHLO compiler can take it from there (see +xref:how-to/compile-model-for-android.adoc[Compile a Model for Android (IREE)]). + +== What you need + +* JDK 21+. +* A `moonshine_streaming` snapshot directory with `config.json`, `model.safetensors` and `tokenizer.json` — + for example https://huggingface.co/moonshine-ai/moonshine-streaming-tiny[`moonshine-ai/moonshine-streaming-tiny`] + (English) or https://huggingface.co/moonshine-ai/moonshine-streaming-tiny-de[`moonshine-ai/moonshine-streaming-tiny-de`] + (German), both published under the MIT licence at the time of writing. Check the model card of the exact + checkpoint and revision you download. + +== Run it + +From a checkout of this repository: + +[source,bash] +---- +MOONSHINE_MODEL_DIR=/path/to/snapshot \ +MOONSHINE_OUT_DIR=build/mlir/moonshine-v2 \ + ./gradlew :llm-inference:moonshine:exportMoonshineV2 +---- + +From your own build, resolve the published module and run its main class: + +[source,kotlin] +---- +val exportTool by configurations.creating +dependencies { + exportTool(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.1")) + exportTool("sk.ainet.transformers:skainet-transformers-inference-moonshine") +} +tasks.register("exportMoonshine") { + classpath = exportTool + mainClass.set("sk.ainet.models.moonshine.MoonshineV2ExportCliKt") + maxHeapSize = "12g" // the graphs carry the weights as constants + environment("MOONSHINE_MODEL_DIR", "/path/to/snapshot") + environment("MOONSHINE_OUT_DIR", layout.buildDirectory.dir("moonshine").get().asFile.path) +} +---- + +A tiny checkpoint exports in well under a minute. + +== What you get + +[cols="2,3,5"] +|=== +| File | Entry point | Inputs → outputs + +| `frontend.mlir` | `@moonshine_v2_frontend` | audio `[1, samples]` (16 kHz mono f32) → features `[1, frames, dim]` +| `encoder.mlir` | `@moonshine_v2_encoder` | features `[1, frames, dim]` → memory (sliding-window attention with bounded lookahead) +| `adapter.mlir` | `@moonshine_v2_adapter` | memory, positions `[1, frames]` → position-aware memory at decoder width +| `prefill.mlir` | `@moonshine_v2_decoder_prefill` | token embedding `[1,1,D]`, padded memory `[1,N,D]`, additive mask `[1,1,1,N]` → logits + per-layer self and cross K/V +| `with_past.mlir` | `@moonshine_v2_decoder_with_past` | token embedding, RoPE cos/sin `[1, headDim]`, self K/V (dynamic length), cross K/V, mask → logits + new self K/V +| `dec_embed.bin` | — | token embedding table `[vocab, D]`, little-endian f32 — the decode loop looks rows up on the host +| `vocab.bin` | — | `u32` piece count, then per token id a `u16` byte length and the UTF-8 piece +| `manifest.json` | — | geometry read from the checkpoint, the shapes used, the files written +|=== + +Weights are folded into the graphs as `stablehlo.constant`. A compiler can move them back out into a shared +archive — with IREE, compile the two decoder graphs with `--iree-opt-export-parameters=model=params.irpa`; the +archives are identical, keep one. + +== Options + +[cols="2,1,5"] +|=== +| Variable | Default | Meaning + +| `MOONSHINE_MODEL_DIR` | — | snapshot directory (required) +| `MOONSHINE_OUT_DIR` | `build/mlir/moonshine-v2` | output directory +| `MOONSHINE_GRAPH` | `all` | or a comma-separated subset of `frontend`, `encoder`, `adapter`, `prefill`, `with_past`, `embeddings`, `vocab` +| `MOONSHINE_FE_SAMPLES` | `21760` | frontend input length (multiple of 80; 320 samples = one feature frame) +| `MOONSHINE_ENC_FRAMES` | `64` | encoder and adapter window in frames (1.28 s) +| `MOONSHINE_MAX_MEM` | `256` | encoder-memory length the decoder graphs are padded to (5.12 s); shorter utterances are masked +| `MOONSHINE_STATIC_PAST` | dynamic | fix the step graph's self-attention cache length, for targets without dynamic shapes +|=== + +There is nothing to set per language or model size: widths, layer counts, per-layer attention bands +(`encoder_config.sliding_windows`), vocabulary size and the length of the positional table come from the +snapshot. With `all`, the export stops with an error if a tensor of the checkpoint was not used — a new +checkpoint layout fails loudly instead of silently dropping weights. + +== From code + +[source,kotlin] +---- +MoonshineV2Checkpoint(File("/path/to/snapshot")).use { checkpoint -> + val harness = MoonshineV2ExportHarness(checkpoint, MoonshineV2ExportOptions(maxMemoryFrames = 512)) + harness.export(outDir, setOf(MoonshineV2ExportHarness.Graph.PREFILL, MoonshineV2ExportHarness.Graph.WITH_PAST)) +} +---- + +`MoonshineV2HfWeightMap` (common code) is the DSL-parameter → checkpoint-tensor map the export uses; a runtime +that loads the weights eagerly can use the same map. diff --git a/docs/modules/ROOT/pages/reference/moonshine-encoder.adoc b/docs/modules/ROOT/pages/reference/moonshine-encoder.adoc index 85352d52..d277d75a 100644 --- a/docs/modules/ROOT/pages/reference/moonshine-encoder.adoc +++ b/docs/modules/ROOT/pages/reference/moonshine-encoder.adoc @@ -24,7 +24,7 @@ a DSL decoder is future work. [source,kotlin] ---- dependencies { - implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.0")) + implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.1")) implementation("sk.ainet.transformers:skainet-transformers-inference-moonshine") } ---- diff --git a/docs/modules/ROOT/pages/tutorials/android-getting-started.adoc b/docs/modules/ROOT/pages/tutorials/android-getting-started.adoc index 8d563154..0dabafff 100644 --- a/docs/modules/ROOT/pages/tutorials/android-getting-started.adoc +++ b/docs/modules/ROOT/pages/tutorials/android-getting-started.adoc @@ -55,7 +55,7 @@ dependencies { // self-registers it on ART at process start — nothing to call. runtimeOnly("sk.ainet.core:skainet-backend-jni-cpu") - implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.0")) + implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.1")) implementation("sk.ainet.transformers:skainet-transformers-core") implementation("sk.ainet.transformers:skainet-transformers-runtime-kllama") implementation("sk.ainet.transformers:skainet-transformers-inference-llama") diff --git a/docs/modules/ROOT/pages/tutorials/getting-started-java.adoc b/docs/modules/ROOT/pages/tutorials/getting-started-java.adoc index 7ece806b..f0698292 100644 --- a/docs/modules/ROOT/pages/tutorials/getting-started-java.adoc +++ b/docs/modules/ROOT/pages/tutorials/getting-started-java.adoc @@ -25,7 +25,7 @@ In your `build.gradle.kts`: [source,kotlin] ---- dependencies { - implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.0")) + implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.1")) implementation("sk.ainet.transformers:skainet-transformers-runtime-kllama") implementation("sk.ainet.transformers:skainet-transformers-agent") @@ -41,7 +41,7 @@ Or in Maven (Maven needs the `-jvm` classifier suffix on platform artifacts): sk.ainet.transformers skainet-transformers-bom - 0.56.0 + 0.56.1 pom import diff --git a/docs/modules/ROOT/pages/tutorials/getting-started-leaf.adoc b/docs/modules/ROOT/pages/tutorials/getting-started-leaf.adoc index f69ed9fc..3c93e112 100644 --- a/docs/modules/ROOT/pages/tutorials/getting-started-leaf.adoc +++ b/docs/modules/ROOT/pages/tutorials/getting-started-leaf.adoc @@ -34,7 +34,7 @@ encoder output to the advertised dimensionality. The runtime applies it automati [source,kotlin] ---- dependencies { - implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.0")) + implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.1")) implementation("sk.ainet.transformers:skainet-transformers-providers") } ---- diff --git a/docs/modules/ROOT/pages/tutorials/llama3-tool-calling.adoc b/docs/modules/ROOT/pages/tutorials/llama3-tool-calling.adoc index 39dee205..7f049afb 100644 --- a/docs/modules/ROOT/pages/tutorials/llama3-tool-calling.adoc +++ b/docs/modules/ROOT/pages/tutorials/llama3-tool-calling.adoc @@ -52,7 +52,7 @@ The pieces you need live in three modules: [source,kotlin] ---- dependencies { - implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.0")) + implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.1")) implementation("sk.ainet.transformers:skainet-transformers-runtime-kllama") implementation("sk.ainet.transformers:skainet-transformers-agent") diff --git a/gradle.properties b/gradle.properties index 00437628..9a0b17b6 100644 --- a/gradle.properties +++ b/gradle.properties @@ -1,5 +1,5 @@ GROUP=sk.ainet.transformers -VERSION_NAME=0.56.0 +VERSION_NAME=0.56.1 POM_DESCRIPTION=SKaiNET-transformers diff --git a/llm-inference/gemma/README.md b/llm-inference/gemma/README.md index 1b2be385..64e2b43d 100644 --- a/llm-inference/gemma/README.md +++ b/llm-inference/gemma/README.md @@ -4,7 +4,7 @@ Reusable **Gemma** model (incl. the FunctionGemma tool-calling fine-tune) author a portable graph producer with **no runtime/board/Torq code**. Pair it with the runtime module below to decode on-device. -- **Coordinate:** `sk.ainet.transformers:skainet-transformers-inference-gemma:0.56.0` +- **Coordinate:** `sk.ainet.transformers:skainet-transformers-inference-gemma:0.56.1` - **Targets:** `android`, `iosArm64`, `iosSimulatorArm64`, `macosArm64`, `linuxX64`, `linuxArm64` (broadly portable — mobile through server). - **Entry point:** `gemmaNetwork()` / `GemmaNetworkLoader` (loads a GGUF, builds the DSL graph, incl. the @@ -24,8 +24,8 @@ FunctionGemma has a one-liner facade in `…:skainet-transformers-runtime-kgemma ```kotlin dependencies { - implementation("sk.ainet.transformers:skainet-transformers-inference-gemma:0.56.0") - implementation("sk.ainet.transformers:skainet-transformers-runtime-gemma-iree:0.56.0") // on-device decode + implementation("sk.ainet.transformers:skainet-transformers-inference-gemma:0.56.1") + implementation("sk.ainet.transformers:skainet-transformers-runtime-gemma-iree:0.56.1") // on-device decode } ```