Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 13 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,19 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

## [0.56.2] — 2026-09-22

A transformers-only release against **SKaiNET engine 0.56.0** (unchanged).

### Added

- **`IreeMoonshineStream`** (`llm-runtime:iree-android`, `libskainet_moonshine_stream.so`, arm64-v8a +
armeabi-v7a, Vulkan + local-task): the streaming Moonshine v2 speech-to-text runtime over the five
graphs of `MoonshineV2ExportCli` — PCM in, cumulative partial transcripts out, exact final on
`finish()`. The counterpart of `IreeKvSession` for ASR: the last piece a Moonshine cartridge needed
from a released artifact. `native/build-moonshine-stream.sh` builds it with the same image and
cache as the other two libraries.

### Added

- **`IreeMoonshineStream`** (`llm-runtime:iree-android`, `libskainet_moonshine_stream.so`, arm64-v8a +
Expand Down
24 changes: 13 additions & 11 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -109,19 +109,21 @@ Honest status — see the project-status note at the top of this README.

## Current release

The current release is **0.56.1** (against **SKaiNET 0.56.0** — a transformers-only release, same
The current release is **0.56.2** (against **SKaiNET 0.56.0** — a transformers-only release, same
pattern as 0.54.1: no new engine version needed).

**Moonshine v2 streaming, every checkpoint of the family, exported from the published module.**
`MoonshineV2ExportCli` turns a Hugging Face `moonshine_streaming` snapshot into the five StableHLO
graphs of the streaming contract (frontend, encoder, adapter, masked prefill, dynamic-cache step)
plus the host-side embedding and vocabulary tables — no Python, nothing to configure per language
or size, and it fails if a checkpoint tensor goes unused. The model gains per-layer attention bands
(the German checkpoints need them) and split encoder/decoder widths (the small ones). See
*Export Moonshine v2 Streaming to StableHLO* in the docs.
**`IreeMoonshineStream`: the streaming Moonshine v2 speech-to-text runtime.** `libskainet_moonshine_stream.so`
(`llm-runtime:iree-android`, arm64-v8a + armeabi-v7a, Vulkan or CPU) drives the five graphs of
`MoonshineV2ExportCli` — frontend, encoder, adapter, masked prefill, dynamic with-past step — plus the
shared decoder parameter archive as a real streaming loop on the device: PCM in, cumulative partial
transcripts out, one exact full re-decode on `finish()`. The Kotlin binding is shaped like `IreeKvSession`:
files by absolute path, no `Context`. With the exporter (0.56.1) this is the last piece a Moonshine
cartridge needed from a released artifact rather than from C source of its own.

It builds on **0.56.0**, which put the two version lines back in lock-step with **SKaiNET 0.56.0**
(the engine skipped 0.55.0 to realign them).
It builds on **0.56.1**, which released `MoonshineV2ExportCli` (a Hugging Face snapshot in, the five
StableHLO graphs and two host-side tables out, no Python, nothing to configure per language or size), and
**0.56.0**, which put the two version lines back in lock-step with **SKaiNET 0.56.0** (the engine skipped
0.55.0 to realign them).

**Grouped-query attention without head expansion, and the compiled leg of SKEEP-005.** The engine's
`scaledDotProductAttention` is grouped-query native, so `MultiHeadAttention` and
Expand Down Expand Up @@ -254,7 +256,7 @@ The recommended way to consume is via the BOM. It pins every published `skainet-

```kotlin
dependencies {
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.1"))
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.2"))

// Versions resolved from the BOM:
implementation("sk.ainet.transformers:skainet-transformers-core")
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -32,7 +32,7 @@ From your own build, resolve the published module and run its main class:
----
val exportTool by configurations.creating
dependencies {
exportTool(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.1"))
exportTool(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.2"))
exportTool("sk.ainet.transformers:skainet-transformers-inference-moonshine")
}
tasks.register<JavaExec>("exportMoonshine") {
Expand Down
2 changes: 1 addition & 1 deletion docs/modules/ROOT/pages/reference/moonshine-encoder.adoc
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,7 @@ a DSL decoder is future work.
[source,kotlin]
----
dependencies {
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.1"))
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.2"))
implementation("sk.ainet.transformers:skainet-transformers-inference-moonshine")
}
----
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -55,7 +55,7 @@ dependencies {
// self-registers it on ART at process start — nothing to call.
runtimeOnly("sk.ainet.core:skainet-backend-jni-cpu")

implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.1"))
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.2"))
implementation("sk.ainet.transformers:skainet-transformers-core")
implementation("sk.ainet.transformers:skainet-transformers-runtime-kllama")
implementation("sk.ainet.transformers:skainet-transformers-inference-llama")
Expand Down
4 changes: 2 additions & 2 deletions docs/modules/ROOT/pages/tutorials/getting-started-java.adoc
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ In your `build.gradle.kts`:
[source,kotlin]
----
dependencies {
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.1"))
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.2"))

implementation("sk.ainet.transformers:skainet-transformers-runtime-kllama")
implementation("sk.ainet.transformers:skainet-transformers-agent")
Expand All @@ -41,7 +41,7 @@ Or in Maven (Maven needs the `-jvm` classifier suffix on platform artifacts):
<dependency>
<groupId>sk.ainet.transformers</groupId>
<artifactId>skainet-transformers-bom</artifactId>
<version>0.56.1</version>
<version>0.56.2</version>
<type>pom</type>
<scope>import</scope>
</dependency>
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -34,7 +34,7 @@ encoder output to the advertised dimensionality. The runtime applies it automati
[source,kotlin]
----
dependencies {
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.1"))
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.2"))
implementation("sk.ainet.transformers:skainet-transformers-providers")
}
----
Expand Down
2 changes: 1 addition & 1 deletion docs/modules/ROOT/pages/tutorials/llama3-tool-calling.adoc
Original file line number Diff line number Diff line change
Expand Up @@ -52,7 +52,7 @@ The pieces you need live in three modules:
[source,kotlin]
----
dependencies {
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.1"))
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.2"))

implementation("sk.ainet.transformers:skainet-transformers-runtime-kllama")
implementation("sk.ainet.transformers:skainet-transformers-agent")
Expand Down
2 changes: 1 addition & 1 deletion gradle.properties
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
GROUP=sk.ainet.transformers
VERSION_NAME=0.56.1
VERSION_NAME=0.56.2

POM_DESCRIPTION=SKaiNET-transformers

Expand Down
6 changes: 3 additions & 3 deletions llm-inference/gemma/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ Reusable **Gemma** model (incl. the FunctionGemma tool-calling fine-tune) author
a portable graph producer with **no runtime/board/Torq code**. Pair it with the runtime module below to
decode on-device.

- **Coordinate:** `sk.ainet.transformers:skainet-transformers-inference-gemma:0.56.1`
- **Coordinate:** `sk.ainet.transformers:skainet-transformers-inference-gemma:0.56.2`
- **Targets:** `android`, `iosArm64`, `iosSimulatorArm64`, `macosArm64`, `linuxX64`, `linuxArm64` (broadly
portable — mobile through server).
- **Entry point:** `gemmaNetwork()` / `GemmaNetworkLoader` (loads a GGUF, builds the DSL graph, incl. the
Expand All @@ -24,8 +24,8 @@ FunctionGemma has a one-liner facade in `…:skainet-transformers-runtime-kgemma

```kotlin
dependencies {
implementation("sk.ainet.transformers:skainet-transformers-inference-gemma:0.56.1")
implementation("sk.ainet.transformers:skainet-transformers-runtime-gemma-iree:0.56.1") // on-device decode
implementation("sk.ainet.transformers:skainet-transformers-inference-gemma:0.56.2")
implementation("sk.ainet.transformers:skainet-transformers-runtime-gemma-iree:0.56.2") // on-device decode
}
```

Expand Down
Loading