From 7c53cfe225b99f2b0741f6ca97954ede5601ba2a Mon Sep 17 00:00:00 2001 From: Shinji Watanabe Date: Tue, 22 Sep 2026 05:28:19 -0400 Subject: [PATCH] Let the ASR and ST demos run on either OWSM v4 One line at the top, as the POWSM notebook has: MODEL, with the encoder-decoder named beside it. Both decode through decode_long, which takes either checkpoint, so nothing else had to change - what was missing was the reader knowing they could. asr_demo also says what best_path means on each: the whole model on the encoder-only checkpoint, and the CTC branch on the encoder-decoder, which answers what that branch was trained on rather than what the task asks. Co-Authored-By: Claude Opus 5 --- Demos/asr_demo.ipynb | 21 ++++++++++++++++----- Demos/st_demo.ipynb | 11 ++++++++--- 2 files changed, 24 insertions(+), 8 deletions(-) diff --git a/Demos/asr_demo.ipynb b/Demos/asr_demo.ipynb index 102eca59..f8f9169b 100644 --- a/Demos/asr_demo.ipynb +++ b/Demos/asr_demo.ipynb @@ -59,9 +59,14 @@ "source": [ "## Transcribe it\n", "\n", + "There are two OWSM v4 checkpoints, and this notebook runs on either: change\n", + "`MODEL` and nothing else.\n", "[`owsm_ctc_v4_1B`](https://huggingface.co/espnet/owsm_ctc_v4_1B) is\n", - "encoder-only: `decode_long` reads a recording of any length, in overlapping\n", - "windows, and decodes each on the CTC head with no beam search.\n" + "encoder-only — `decode_long` reads a recording of any length, in overlapping\n", + "windows, and decodes each on the CTC head with no beam search — and\n", + "[`owsm_v4_medium_1B`](https://huggingface.co/espnet/owsm_v4_medium_1B) is the\n", + "encoder-decoder, which searches, is slower, and is the one that can be given\n", + "a text prompt." ] }, { @@ -72,10 +77,11 @@ "source": [ "from espnet2.bin.s2t_inference import Speech2Text\n", "\n", - "s2t = Speech2Text.from_pretrained(\"espnet/owsm_ctc_v4_1B\", device=\"cpu\")\n", + "MODEL = \"espnet/owsm_ctc_v4_1B\" # or \"espnet/owsm_v4_medium_1B\"\n", + "s2t = Speech2Text.from_pretrained(MODEL, device=\"cpu\")\n", "\n", "segments = s2t.decode_long(\"sample.wav\", lang_sym=\"\", task_sym=\"\")\n", - "print(\" \".join(text for _, _, text in segments))\n" + "print(\" \".join(text for _, _, text in segments))" ] }, { @@ -85,7 +91,12 @@ "## Let it work out the language\n", "\n", "`` asks the model to identify the language instead of being told it.\n", - "The answer comes back as the first symbol of the decoded text.\n" + "The answer comes back as the first symbol of the decoded text.\n", + "\n", + "`best_path` is CTC decoding with no search. On the encoder-only checkpoint\n", + "that is the whole model; on the encoder-decoder it reads the CTC *branch*,\n", + "which answers what that branch was trained on rather than what the task\n", + "symbol asks for — call `s2t(...)` there instead, and pay for the search." ] }, { diff --git a/Demos/st_demo.ipynb b/Demos/st_demo.ipynb index 0c8748b7..8cc3e829 100644 --- a/Demos/st_demo.ipynb +++ b/Demos/st_demo.ipynb @@ -59,7 +59,11 @@ "\n", "The task symbol carries the target: `` transcribes, ``\n", "translates to German, and so on. The model's token list is where the\n", - "available pairs are written, so ask it rather than a table.\n" + "available pairs are written, so ask it rather than a table.\n", + "\n", + "Either OWSM v4 checkpoint does this, and `MODEL` is the only line to change:\n", + "the encoder-only one below, or `espnet/owsm_v4_medium_1B`, which searches and\n", + "is slower." ] }, { @@ -70,10 +74,11 @@ "source": [ "from espnet2.bin.s2t_inference import Speech2Text\n", "\n", - "s2t = Speech2Text.from_pretrained(\"espnet/owsm_ctc_v4_1B\", device=\"cpu\")\n", + "MODEL = \"espnet/owsm_ctc_v4_1B\" # or \"espnet/owsm_v4_medium_1B\"\n", + "s2t = Speech2Text.from_pretrained(MODEL, device=\"cpu\")\n", "\n", "targets = [t for t in s2t.s2t_model.token_list if t.startswith(\"