Motivation
Dictation currently always runs the medium streaming Moonshine model (ModelSpec.stt("en", JNI.MOONSHINE_MODEL_ARCH_MEDIUM_STREAMING, false) in MoonshineModel.kt). We want to measure whether smaller models are fast enough / faster on target devices, and what accuracy we trade away. A model selector would let us switch models and compare.
Current behavior
MoonshineModel.spec() hardcodes MOONSHINE_MODEL_ARCH_MEDIUM_STREAMING.
- The SDK also exposes
MOONSHINE_MODEL_ARCH_TINY, MOONSHINE_MODEL_ARCH_TINY_STREAMING, MOONSHINE_MODEL_ARCH_BASE, MOONSHINE_MODEL_ARCH_BASE_STREAMING, MOONSHINE_MODEL_ARCH_SMALL_STREAMING.
- Because
spec() determines the cache directory, only the medium model is ever downloaded or used.
Desired behavior
- A way to pick which Moonshine model the IME uses (at minimum: tiny/base/small/medium streaming variants), so we can A/B latency and accuracy of dictation on real devices.
- Download flow in setup must cover the selected model (per-model cache directories already fall out of
ModelCache.directoryFor(context, spec(), null)).
- Keep the Moonshine contract intact: spec → configure →
load() → start(), blocking calls serialized on the IME's worker executor; if the model can change while the IME is running, stop/reload must respect the existing STOPPING barrier.
Open questions
- UI surface: this MVP deliberately excludes settings screens. Should the selector be (a) a debug-only affordance (e.g.
BuildConfig.DEBUG-gated), (b) a simple spinner on the setup screen, or (c) driven by a build config field for per-build testing?
- Should selection persist across IME restarts, and how (no transcript/settings persistence is stored today)?
- What metrics do we actually want to capture while testing — just subjective latency, or also measured time-to-first-partial / time-to-final?
Acceptance criteria
Motivation
Dictation currently always runs the medium streaming Moonshine model (
ModelSpec.stt("en", JNI.MOONSHINE_MODEL_ARCH_MEDIUM_STREAMING, false)inMoonshineModel.kt). We want to measure whether smaller models are fast enough / faster on target devices, and what accuracy we trade away. A model selector would let us switch models and compare.Current behavior
MoonshineModel.spec()hardcodesMOONSHINE_MODEL_ARCH_MEDIUM_STREAMING.MOONSHINE_MODEL_ARCH_TINY,MOONSHINE_MODEL_ARCH_TINY_STREAMING,MOONSHINE_MODEL_ARCH_BASE,MOONSHINE_MODEL_ARCH_BASE_STREAMING,MOONSHINE_MODEL_ARCH_SMALL_STREAMING.spec()determines the cache directory, only the medium model is ever downloaded or used.Desired behavior
ModelCache.directoryFor(context, spec(), null)).load()→start(), blocking calls serialized on the IME's worker executor; if the model can change while the IME is running, stop/reload must respect the existing STOPPING barrier.Open questions
BuildConfig.DEBUG-gated), (b) a simple spinner on the setup screen, or (c) driven by a build config field for per-build testing?Acceptance criteria