Skip to content

Add model selector to compare speed of smaller Moonshine models #1

Description

@monneyboi

Motivation

Dictation currently always runs the medium streaming Moonshine model (ModelSpec.stt("en", JNI.MOONSHINE_MODEL_ARCH_MEDIUM_STREAMING, false) in MoonshineModel.kt). We want to measure whether smaller models are fast enough / faster on target devices, and what accuracy we trade away. A model selector would let us switch models and compare.

Current behavior

  • MoonshineModel.spec() hardcodes MOONSHINE_MODEL_ARCH_MEDIUM_STREAMING.
  • The SDK also exposes MOONSHINE_MODEL_ARCH_TINY, MOONSHINE_MODEL_ARCH_TINY_STREAMING, MOONSHINE_MODEL_ARCH_BASE, MOONSHINE_MODEL_ARCH_BASE_STREAMING, MOONSHINE_MODEL_ARCH_SMALL_STREAMING.
  • Because spec() determines the cache directory, only the medium model is ever downloaded or used.

Desired behavior

  • A way to pick which Moonshine model the IME uses (at minimum: tiny/base/small/medium streaming variants), so we can A/B latency and accuracy of dictation on real devices.
  • Download flow in setup must cover the selected model (per-model cache directories already fall out of ModelCache.directoryFor(context, spec(), null)).
  • Keep the Moonshine contract intact: spec → configure → load() → start(), blocking calls serialized on the IME's worker executor; if the model can change while the IME is running, stop/reload must respect the existing STOPPING barrier.

Open questions

  • UI surface: this MVP deliberately excludes settings screens. Should the selector be (a) a debug-only affordance (e.g. BuildConfig.DEBUG-gated), (b) a simple spinner on the setup screen, or (c) driven by a build config field for per-build testing?
  • Should selection persist across IME restarts, and how (no transcript/settings persistence is stored today)?
  • What metrics do we actually want to capture while testing — just subjective latency, or also measured time-to-first-partial / time-to-final?

Acceptance criteria

  • Model choice is configurable in-app or per-build (per decision above)
  • Setup downloads the selected model; IME uses the same one (no medium/other mismatch, since a missing model means a blocking download/failure at dictation time)
  • Switching models mid-session behaves sanely (reload or restart, no leaked transcriber)
  • Default remains the current medium streaming model unless changed

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions