Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,10 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

## [0.56.1] — 2026-09-21

A transformers-only release against **SKaiNET engine 0.56.0** (unchanged).

### Added — Moonshine v2 streaming: every checkpoint of the family, exported from the published artifact

- **`MoonshineV2ExportCli`** (`:llm-inference:moonshine:exportMoonshineV2`, JVM): a Hugging Face
Expand Down
17 changes: 14 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -109,8 +109,19 @@ Honest status — see the project-status note at the top of this README.

## Current release

The current release is **0.56.0**, back in lock-step with **SKaiNET 0.56.0** (the engine skipped
0.55.0 to realign the two version lines).
The current release is **0.56.1** (against **SKaiNET 0.56.0** — a transformers-only release, same
pattern as 0.54.1: no new engine version needed).

**Moonshine v2 streaming, every checkpoint of the family, exported from the published module.**
`MoonshineV2ExportCli` turns a Hugging Face `moonshine_streaming` snapshot into the five StableHLO
graphs of the streaming contract (frontend, encoder, adapter, masked prefill, dynamic-cache step)
plus the host-side embedding and vocabulary tables — no Python, nothing to configure per language
or size, and it fails if a checkpoint tensor goes unused. The model gains per-layer attention bands
(the German checkpoints need them) and split encoder/decoder widths (the small ones). See
*Export Moonshine v2 Streaming to StableHLO* in the docs.

It builds on **0.56.0**, which put the two version lines back in lock-step with **SKaiNET 0.56.0**
(the engine skipped 0.55.0 to realign them).

**Grouped-query attention without head expansion, and the compiled leg of SKEEP-005.** The engine's
`scaledDotProductAttention` is grouped-query native, so `MultiHeadAttention` and
Expand Down Expand Up @@ -243,7 +254,7 @@ The recommended way to consume is via the BOM. It pins every published `skainet-

```kotlin
dependencies {
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.0"))
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.1"))

// Versions resolved from the BOM:
implementation("sk.ainet.transformers:skainet-transformers-core")
Expand Down
1 change: 1 addition & 0 deletions docs/modules/ROOT/nav.adoc
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,7 @@
* xref:how-to/run-unified-cli.adoc[Use the Unified CLI]
* xref:how-to/benchmarking.adoc[Run Benchmarks]
* xref:how-to/compile-model-for-android.adoc[Compile a Model for Android (IREE)]
* xref:how-to/export-moonshine-v2-streaming.adoc[Export Moonshine v2 Streaming to StableHLO]

.Reference
* xref:reference/architecture.adoc[Architecture Overview]
Expand Down
100 changes: 100 additions & 0 deletions docs/modules/ROOT/pages/how-to/export-moonshine-v2-streaming.adoc
Original file line number Diff line number Diff line change
@@ -0,0 +1,100 @@
= Export Moonshine v2 Streaming to StableHLO
:description: Turn a Hugging Face moonshine_streaming checkpoint into the five StableHLO graphs of the streaming speech-to-text contract — from the published module, without Python.

`skainet-transformers-inference-moonshine` carries the Moonshine **v2 streaming** model in the SKaiNET NN DSL
and, since **0.56.1**, a JVM command that exports it: a Hugging Face snapshot in, five StableHLO graphs and two
host-side tables out. Any StableHLO compiler can take it from there (see
xref:how-to/compile-model-for-android.adoc[Compile a Model for Android (IREE)]).

== What you need

* JDK 21+.
* A `moonshine_streaming` snapshot directory with `config.json`, `model.safetensors` and `tokenizer.json` —
for example https://huggingface.co/moonshine-ai/moonshine-streaming-tiny[`moonshine-ai/moonshine-streaming-tiny`]
(English) or https://huggingface.co/moonshine-ai/moonshine-streaming-tiny-de[`moonshine-ai/moonshine-streaming-tiny-de`]
(German), both published under the MIT licence at the time of writing. Check the model card of the exact
checkpoint and revision you download.

== Run it

From a checkout of this repository:

[source,bash]
----
MOONSHINE_MODEL_DIR=/path/to/snapshot \
MOONSHINE_OUT_DIR=build/mlir/moonshine-v2 \
./gradlew :llm-inference:moonshine:exportMoonshineV2
----

From your own build, resolve the published module and run its main class:

[source,kotlin]
----
val exportTool by configurations.creating
dependencies {
exportTool(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.1"))
exportTool("sk.ainet.transformers:skainet-transformers-inference-moonshine")
}
tasks.register<JavaExec>("exportMoonshine") {
classpath = exportTool
mainClass.set("sk.ainet.models.moonshine.MoonshineV2ExportCliKt")
maxHeapSize = "12g" // the graphs carry the weights as constants
environment("MOONSHINE_MODEL_DIR", "/path/to/snapshot")
environment("MOONSHINE_OUT_DIR", layout.buildDirectory.dir("moonshine").get().asFile.path)
}
----

A tiny checkpoint exports in well under a minute.

== What you get

[cols="2,3,5"]
|===
| File | Entry point | Inputs → outputs

| `frontend.mlir` | `@moonshine_v2_frontend` | audio `[1, samples]` (16 kHz mono f32) → features `[1, frames, dim]`
| `encoder.mlir` | `@moonshine_v2_encoder` | features `[1, frames, dim]` → memory (sliding-window attention with bounded lookahead)
| `adapter.mlir` | `@moonshine_v2_adapter` | memory, positions `[1, frames]` → position-aware memory at decoder width
| `prefill.mlir` | `@moonshine_v2_decoder_prefill` | token embedding `[1,1,D]`, padded memory `[1,N,D]`, additive mask `[1,1,1,N]` → logits + per-layer self and cross K/V
| `with_past.mlir` | `@moonshine_v2_decoder_with_past` | token embedding, RoPE cos/sin `[1, headDim]`, self K/V (dynamic length), cross K/V, mask → logits + new self K/V
| `dec_embed.bin` | — | token embedding table `[vocab, D]`, little-endian f32 — the decode loop looks rows up on the host
| `vocab.bin` | — | `u32` piece count, then per token id a `u16` byte length and the UTF-8 piece
| `manifest.json` | — | geometry read from the checkpoint, the shapes used, the files written
|===

Weights are folded into the graphs as `stablehlo.constant`. A compiler can move them back out into a shared
archive — with IREE, compile the two decoder graphs with `--iree-opt-export-parameters=model=params.irpa`; the
archives are identical, keep one.

== Options

[cols="2,1,5"]
|===
| Variable | Default | Meaning

| `MOONSHINE_MODEL_DIR` | — | snapshot directory (required)
| `MOONSHINE_OUT_DIR` | `build/mlir/moonshine-v2` | output directory
| `MOONSHINE_GRAPH` | `all` | or a comma-separated subset of `frontend`, `encoder`, `adapter`, `prefill`, `with_past`, `embeddings`, `vocab`
| `MOONSHINE_FE_SAMPLES` | `21760` | frontend input length (multiple of 80; 320 samples = one feature frame)
| `MOONSHINE_ENC_FRAMES` | `64` | encoder and adapter window in frames (1.28 s)
| `MOONSHINE_MAX_MEM` | `256` | encoder-memory length the decoder graphs are padded to (5.12 s); shorter utterances are masked
| `MOONSHINE_STATIC_PAST` | dynamic | fix the step graph's self-attention cache length, for targets without dynamic shapes
|===

There is nothing to set per language or model size: widths, layer counts, per-layer attention bands
(`encoder_config.sliding_windows`), vocabulary size and the length of the positional table come from the
snapshot. With `all`, the export stops with an error if a tensor of the checkpoint was not used — a new
checkpoint layout fails loudly instead of silently dropping weights.

== From code

[source,kotlin]
----
MoonshineV2Checkpoint(File("/path/to/snapshot")).use { checkpoint ->
val harness = MoonshineV2ExportHarness(checkpoint, MoonshineV2ExportOptions(maxMemoryFrames = 512))
harness.export(outDir, setOf(MoonshineV2ExportHarness.Graph.PREFILL, MoonshineV2ExportHarness.Graph.WITH_PAST))
}
----

`MoonshineV2HfWeightMap` (common code) is the DSL-parameter → checkpoint-tensor map the export uses; a runtime
that loads the weights eagerly can use the same map.
2 changes: 1 addition & 1 deletion docs/modules/ROOT/pages/reference/moonshine-encoder.adoc
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,7 @@ a DSL decoder is future work.
[source,kotlin]
----
dependencies {
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.0"))
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.1"))
implementation("sk.ainet.transformers:skainet-transformers-inference-moonshine")
}
----
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -55,7 +55,7 @@ dependencies {
// self-registers it on ART at process start — nothing to call.
runtimeOnly("sk.ainet.core:skainet-backend-jni-cpu")

implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.0"))
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.1"))
implementation("sk.ainet.transformers:skainet-transformers-core")
implementation("sk.ainet.transformers:skainet-transformers-runtime-kllama")
implementation("sk.ainet.transformers:skainet-transformers-inference-llama")
Expand Down
4 changes: 2 additions & 2 deletions docs/modules/ROOT/pages/tutorials/getting-started-java.adoc
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ In your `build.gradle.kts`:
[source,kotlin]
----
dependencies {
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.0"))
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.1"))

implementation("sk.ainet.transformers:skainet-transformers-runtime-kllama")
implementation("sk.ainet.transformers:skainet-transformers-agent")
Expand All @@ -41,7 +41,7 @@ Or in Maven (Maven needs the `-jvm` classifier suffix on platform artifacts):
<dependency>
<groupId>sk.ainet.transformers</groupId>
<artifactId>skainet-transformers-bom</artifactId>
<version>0.56.0</version>
<version>0.56.1</version>
<type>pom</type>
<scope>import</scope>
</dependency>
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -34,7 +34,7 @@ encoder output to the advertised dimensionality. The runtime applies it automati
[source,kotlin]
----
dependencies {
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.0"))
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.1"))
implementation("sk.ainet.transformers:skainet-transformers-providers")
}
----
Expand Down
2 changes: 1 addition & 1 deletion docs/modules/ROOT/pages/tutorials/llama3-tool-calling.adoc
Original file line number Diff line number Diff line change
Expand Up @@ -52,7 +52,7 @@ The pieces you need live in three modules:
[source,kotlin]
----
dependencies {
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.0"))
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.1"))

implementation("sk.ainet.transformers:skainet-transformers-runtime-kllama")
implementation("sk.ainet.transformers:skainet-transformers-agent")
Expand Down
2 changes: 1 addition & 1 deletion gradle.properties
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
GROUP=sk.ainet.transformers
VERSION_NAME=0.56.0
VERSION_NAME=0.56.1

POM_DESCRIPTION=SKaiNET-transformers

Expand Down
6 changes: 3 additions & 3 deletions llm-inference/gemma/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ Reusable **Gemma** model (incl. the FunctionGemma tool-calling fine-tune) author
a portable graph producer with **no runtime/board/Torq code**. Pair it with the runtime module below to
decode on-device.

- **Coordinate:** `sk.ainet.transformers:skainet-transformers-inference-gemma:0.56.0`
- **Coordinate:** `sk.ainet.transformers:skainet-transformers-inference-gemma:0.56.1`
- **Targets:** `android`, `iosArm64`, `iosSimulatorArm64`, `macosArm64`, `linuxX64`, `linuxArm64` (broadly
portable — mobile through server).
- **Entry point:** `gemmaNetwork()` / `GemmaNetworkLoader` (loads a GGUF, builds the DSL graph, incl. the
Expand All @@ -24,8 +24,8 @@ FunctionGemma has a one-liner facade in `…:skainet-transformers-runtime-kgemma

```kotlin
dependencies {
implementation("sk.ainet.transformers:skainet-transformers-inference-gemma:0.56.0")
implementation("sk.ainet.transformers:skainet-transformers-runtime-gemma-iree:0.56.0") // on-device decode
implementation("sk.ainet.transformers:skainet-transformers-inference-gemma:0.56.1")
implementation("sk.ainet.transformers:skainet-transformers-runtime-gemma-iree:0.56.1") // on-device decode
}
```

Expand Down
Loading