From 334562cf5048b26e7576c0c2cc7f988dff3b9c0e Mon Sep 17 00:00:00 2001 From: Shinji Watanabe Date: Tue, 22 Sep 2026 04:47:04 -0400 Subject: [PATCH 1/3] A demo for POWSM's four tasks Phone recognition has a command, an MCP tool and, with espnet#6792, a Space; this is its notebook, and it is the only one of the four places where all four of POWSM's tasks can be read side by side. hears the phones, transcribes and is visibly worse at it - which is what a phonetic model is - and then the two that take something written with the audio: is given the words and answers with the phones, is given the phones and answers with the words. Same six seconds of speech throughout, so the four answers can be compared. Named s2t_pr_demo.ipynb for the rule in Demos/README.md: _, and phone recognition is a variant of s2t rather than a task of its own. Pins espnet==202610.post2, which carries the page and the decode_window that passes a prompt on; until that release is on PyPI the job says so and skips, which is what allow_unreleased_pin is for. Co-Authored-By: Claude Opus 5 --- .github/workflows/s2t_pr_demo.yml | 31 +++++ Demos/README.md | 1 + Demos/s2t_pr_demo.ipynb | 221 ++++++++++++++++++++++++++++++ README.md | 1 + 4 files changed, 254 insertions(+) create mode 100644 .github/workflows/s2t_pr_demo.yml create mode 100644 Demos/s2t_pr_demo.ipynb diff --git a/.github/workflows/s2t_pr_demo.yml b/.github/workflows/s2t_pr_demo.yml new file mode 100644 index 00000000..43d85bd7 --- /dev/null +++ b/.github/workflows/s2t_pr_demo.yml @@ -0,0 +1,31 @@ +# Demos/s2t_pr_demo.ipynb, every Sunday. The badge on this workflow is what +# the README shows beside that notebook, which is why it has a file of its own: +# GitHub's badge is per workflow, and cannot show one job of a matrix. + +name: s2t_pr_demo + +on: + schedule: + # Sundays, 05:00 UTC + - cron: "0 5 * * 0" + workflow_dispatch: + pull_request: + paths: + - "Demos/s2t_pr_demo.ipynb" + - "tools/**" + - ".github/workflows/s2t_pr_demo.yml" + - ".github/workflows/_run_notebook.yml" + +permissions: + contents: read + +jobs: + run: + uses: ./.github/workflows/_run_notebook.yml + with: + notebook: Demos/s2t_pr_demo.ipynb + # POWSM's and reach the page through espnet2.bin.demo, and + # this notebook pins the release that carries them (espnet#6791 cuts + # it). Until it is on PyPI the job says so and skips. Delete this line + # when 202610.post2 is published. + allow_unreleased_pin: true diff --git a/Demos/README.md b/Demos/README.md index c52ee9c8..8cbe9946 100644 --- a/Demos/README.md +++ b/Demos/README.md @@ -40,6 +40,7 @@ the answer now. | [`asr_streaming_demo.ipynb`](asr_streaming_demo.ipynb) | [![asr_streaming_demo](https://github.com/espnet/notebook/actions/workflows/asr_streaming_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/asr_streaming_demo.yml) | Watch the words appear while the audio is still arriving | | [`st_demo.ipynb`](st_demo.ipynb) | [![st_demo](https://github.com/espnet/notebook/actions/workflows/st_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/st_demo.yml) | Translate English speech into German, French and Chinese — the same model | | [`s2t_align_demo.ipynb`](s2t_align_demo.ipynb) | [![s2t_align_demo](https://github.com/espnet/notebook/actions/workflows/s2t_align_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/s2t_align_demo.yml) | Line text up with the audio it was said in, and score how well they agree | +| [`s2t_pr_demo.ipynb`](s2t_pr_demo.ipynb) | [![s2t_pr_demo](https://github.com/espnet/notebook/actions/workflows/s2t_pr_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/s2t_pr_demo.yml) | Hear the phones, and cross between phones and words with POWSM | | [`tts_demo.ipynb`](tts_demo.ipynb) | [![tts_demo](https://github.com/espnet/notebook/actions/workflows/tts_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/tts_demo.yml) | Type a sentence, hear it spoken — one English voice, then 128 of them | | [`enh_demo.ipynb`](enh_demo.ipynb) | [![enh_demo](https://github.com/espnet/notebook/actions/workflows/enh_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/enh_demo.yml) | Pull speech out of noise, and measure how much it helped | | [`spk_demo.ipynb`](spk_demo.ipynb) | [![spk_demo](https://github.com/espnet/notebook/actions/workflows/spk_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/spk_demo.yml) | Turn a voice into a vector, and score two recordings against each other | diff --git a/Demos/s2t_pr_demo.ipynb b/Demos/s2t_pr_demo.ipynb new file mode 100644 index 00000000..33d4fe82 --- /dev/null +++ b/Demos/s2t_pr_demo.ipynb @@ -0,0 +1,221 @@ +{ + "cells": [ + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "# Phones, and the two ways across to words\n", + "\n", + "[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/espnet/notebook/blob/master/Demos/s2t_pr_demo.ipynb) [![s2t_pr_demo](https://github.com/espnet/notebook/actions/workflows/s2t_pr_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/s2t_pr_demo.yml)\n", + "\n", + "[POWSM](https://arxiv.org/abs/2510.24992) is a phonetic foundation model:\n", + "it hears speech and answers in phones. POWSM-CTC is the encoder-only\n", + "variant, and it does four things with one recording —\n", + "\n", + "| task | you give | it answers |\n", + "| :-- | :-- | :-- |\n", + "| `` | the audio | the words |\n", + "| `` | the audio | the phones, in IPA |\n", + "| `` | the audio **and the words** | the phones |\n", + "| `` | the audio **and the phones** | the words |\n", + "\n", + "— which is what this notebook runs, one after the other, on the same six\n", + "seconds of speech. CPU is enough." + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## Install" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "%pip install -q \"espnet==202610.post2\" espnet_model_zoo librosa" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## A recording" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "import librosa\n", + "from IPython.display import Audio, display\n", + "\n", + "!wget -q -O sample.wav https://github.com/espnet/espnet/raw/master/test_utils/ctc_align_test.wav\n", + "speech, rate = librosa.load(\"sample.wav\", sr=16000)\n", + "display(Audio(speech, rate=rate))" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## The model\n", + "\n", + "POWSM reads a 20-second window, so a shorter clip is padded to it. The\n", + "checkpoint also carries its own spelling of \"work the language out\n", + "yourself\" — `` here, where OWSM writes `` — so ask the model\n", + "rather than typing a symbol." + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "import librosa\n", + "from espnet2.bin.s2t_inference import Speech2Text\n", + "\n", + "s2t = Speech2Text.from_pretrained(\"espnet/powsm_ctc\", device=\"cpu\")\n", + "\n", + "window = s2t.preprocessor_conf[\"speech_length\"]\n", + "nolang = s2t.no_language()\n", + "padded = librosa.util.fix_length(speech, size=rate * window)\n", + "print(f\"{window} s window, language symbol {nolang}, CTC-only: {s2t.ctc_only}\")" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## ``: the phones\n", + "\n", + "The task it is built for. POWSM writes each phone between slashes, so that\n", + "a phone spelled like a BPE token is still one token; spaced out is easier\n", + "to read and is what anything counting or aligning them wants." + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "import re\n", + "\n", + "\n", + "def phones(decoded):\n", + " \"\"\"The phones out of POWSM's /p//h//o/ spelling.\"\"\"\n", + " return \" \".join(re.findall(r\"/([^/]+)/\", decoded))\n", + "\n", + "\n", + "def run(task, text_prev=\"\"):\n", + " decoded = s2t.best_path(\n", + " padded, text_prev=text_prev, lang_sym=nolang, task_sym=task\n", + " )[0][0]\n", + " return decoded.split(\">\")[-1].strip() if \"<\" in decoded else decoded\n", + "\n", + "\n", + "said_in_phones = phones(run(\"\"))\n", + "print(said_in_phones)" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## ``: the words, for contrast\n", + "\n", + "A phonetic model will transcribe, and it is worse at it than a model built\n", + "for text. That is not a bug to report; it is what the model is for." + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "print(run(\"\"))" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## ``: phones for words you give\n", + "\n", + "Now the audio is not the only input. Hand it the words that were said, and\n", + "it answers with the phones for *those* words rather than for whatever it\n", + "thought it heard." + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "said = \"the sale of the hotels is part of holiday's strategy\"\n", + "\n", + "print(phones(run(\"\", text_prev=said)))" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## ``: words for phones you give\n", + "\n", + "The other direction. The prompt is phones, in the spelling the model was\n", + "trained on — each between slashes." + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "as_prompt = \"/\" + \"//\".join(said_in_phones.split()) + \"/\"\n", + "print(as_prompt[:70], \"...\")\n", + "\n", + "print(run(\"\", text_prev=as_prompt))" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## Where next\n", + "\n", + "- **From the terminal**: `espnet phonemize sample.wav` is ``, in one line\n", + "- **In the browser**: the [powsm-ctc Space](https://huggingface.co/spaces/espnet/powsm-ctc)\n", + " offers all four tasks, and `espnet demo --model espnet/powsm_ctc` is the same page\n", + "- **From an assistant**: the MCP server's `phonemize` tool\n", + "- **Alignment**, with the same kind of model: [`s2t_align_demo.ipynb`](s2t_align_demo.ipynb)\n", + "- **The encoder-decoder POWSM**, [`espnet/powsm`](https://huggingface.co/espnet/powsm),\n", + " answers the same four tasks with a search rather than a CTC pass" + ] + } + ], + "metadata": { + "colab": { + "provenance": [] + }, + "kernelspec": { + "display_name": "Python 3", + "name": "python3" + }, + "language_info": { + "name": "python" + } + }, + "nbformat": 4, + "nbformat_minor": 0 +} diff --git a/README.md b/README.md index 69d7cc3a..16831c85 100644 --- a/README.md +++ b/README.md @@ -18,6 +18,7 @@ One per task, flat in [`Demos/`](Demos), each short enough to read in a sitting. | [`asr_streaming_demo.ipynb`](Demos/asr_streaming_demo.ipynb) | [![asr_streaming_demo](https://github.com/espnet/notebook/actions/workflows/asr_streaming_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/asr_streaming_demo.yml) | Watch the words appear while the audio is still arriving | | [`st_demo.ipynb`](Demos/st_demo.ipynb) | [![st_demo](https://github.com/espnet/notebook/actions/workflows/st_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/st_demo.yml) | Translate English speech into German, French and Chinese — the same model | | [`s2t_align_demo.ipynb`](Demos/s2t_align_demo.ipynb) | [![s2t_align_demo](https://github.com/espnet/notebook/actions/workflows/s2t_align_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/s2t_align_demo.yml) | Line text up with the audio it was said in, and score how well they agree | +| [`s2t_pr_demo.ipynb`](Demos/s2t_pr_demo.ipynb) | [![s2t_pr_demo](https://github.com/espnet/notebook/actions/workflows/s2t_pr_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/s2t_pr_demo.yml) | Hear the phones, and cross between phones and words with POWSM | | [`tts_demo.ipynb`](Demos/tts_demo.ipynb) | [![tts_demo](https://github.com/espnet/notebook/actions/workflows/tts_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/tts_demo.yml) | Type a sentence, hear it spoken — one English voice, then 128 of them | | [`enh_demo.ipynb`](Demos/enh_demo.ipynb) | [![enh_demo](https://github.com/espnet/notebook/actions/workflows/enh_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/enh_demo.yml) | Pull speech out of noise, and measure how much it helped | | [`spk_demo.ipynb`](Demos/spk_demo.ipynb) | [![spk_demo](https://github.com/espnet/notebook/actions/workflows/spk_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/spk_demo.yml) | Turn a voice into a vector, and score two recordings against each other | From 7a6745bae956f22b0267cfd0cf9015e7cb7a4032 Mon Sep 17 00:00:00 2001 From: Shinji Watanabe Date: Tue, 22 Sep 2026 05:17:17 -0400 Subject: [PATCH 2/3] Let the reader run this on either POWSM MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit `MODEL` at the top, and nothing else to change: `espnet/powsm_ctc` or `espnet/powsm`. What makes the swap work is `decode_window`, which reads a CTC-only checkpoint off its CTC head and searches one that has a decoder, so the notebook does not have to know which it is holding - `best_path`, which this used before, would have answered the CTC branch's own task on the encoder-decoder and quietly returned phones for every task. The note says what the trade is, measured rather than remembered: the encoder-decoder's phones are finer - `pʰ`, `tʰ`, `oʊ` where the CTC writes `p`, `t`, `o` - and it is about four times slower on a CPU. Co-Authored-By: Claude Opus 5 --- Demos/s2t_pr_demo.ipynb | 24 +++++++++++++++++------- 1 file changed, 17 insertions(+), 7 deletions(-) diff --git a/Demos/s2t_pr_demo.ipynb b/Demos/s2t_pr_demo.ipynb index 33d4fe82..c39deba9 100644 --- a/Demos/s2t_pr_demo.ipynb +++ b/Demos/s2t_pr_demo.ipynb @@ -66,10 +66,20 @@ "source": [ "## The model\n", "\n", + "There are two POWSMs, and this notebook runs on either: change `MODEL` below\n", + "and nothing else. `powsm_ctc` reads its window in one pass and is what\n", + "`espnet phonemize` loads; `powsm` is the encoder-decoder, whose phones are\n", + "finer - aspiration and proper diphthongs, `pʰ` and `oʊ` where the CTC writes\n", + "`p` and `o` - and which is about four times slower on a CPU.\n", + "\n", + "`decode_window` is what makes the swap work: a CTC-only checkpoint is read\n", + "off its CTC head, one with a decoder is searched, and the caller does not\n", + "have to know which it is holding.\n", + "\n", "POWSM reads a 20-second window, so a shorter clip is padded to it. The\n", - "checkpoint also carries its own spelling of \"work the language out\n", - "yourself\" — `` here, where OWSM writes `` — so ask the model\n", - "rather than typing a symbol." + "checkpoint also carries its own spelling of \"work the language out yourself\"\n", + "- `` here, where OWSM writes `` - so ask the model rather than\n", + "typing a symbol." ] }, { @@ -81,7 +91,8 @@ "import librosa\n", "from espnet2.bin.s2t_inference import Speech2Text\n", "\n", - "s2t = Speech2Text.from_pretrained(\"espnet/powsm_ctc\", device=\"cpu\")\n", + "MODEL = \"espnet/powsm_ctc\" # or \"espnet/powsm\", the encoder-decoder\n", + "s2t = Speech2Text.from_pretrained(MODEL, device=\"cpu\")\n", "\n", "window = s2t.preprocessor_conf[\"speech_length\"]\n", "nolang = s2t.no_language()\n", @@ -115,9 +126,8 @@ "\n", "\n", "def run(task, text_prev=\"\"):\n", - " decoded = s2t.best_path(\n", - " padded, text_prev=text_prev, lang_sym=nolang, task_sym=task\n", - " )[0][0]\n", + " \"\"\"One window, decoded the way this checkpoint has to be.\"\"\"\n", + " decoded = s2t.decode_window(padded, nolang, task, text_prev)\n", " return decoded.split(\">\")[-1].strip() if \"<\" in decoded else decoded\n", "\n", "\n", From 197d7ffc884224743e9da1a5d22c46246f16ca7d Mon Sep 17 00:00:00 2001 From: Shinji Watanabe Date: Tue, 22 Sep 2026 14:25:20 -0400 Subject: [PATCH 3/3] Three corrections from POWSM's author @chinjouli, reviewing: - the two models share a phone set, and the decoder splits diphthongs by design, so "finer" was me reading an inventory difference out of an output difference. What differs is which phones each chose. - "worse at transcription" needed its condition: against English, where a model trained for text has seen far more. On a language general corpora barely cover, POWSM is competitive. - "whatever it thought it heard" made the audio sound discarded. The text grounds the phones; the model is still listening, and that is where the pronunciation comes from - which is the interesting part of the task and what my sentence had talked the reader out of. RESULT ok, 16 cells, after the change. Co-Authored-By: Claude Opus 5 --- Demos/s2t_pr_demo.ipynb | 21 +++++++++++++-------- 1 file changed, 13 insertions(+), 8 deletions(-) diff --git a/Demos/s2t_pr_demo.ipynb b/Demos/s2t_pr_demo.ipynb index c39deba9..cb9be4cc 100644 --- a/Demos/s2t_pr_demo.ipynb +++ b/Demos/s2t_pr_demo.ipynb @@ -68,9 +68,10 @@ "\n", "There are two POWSMs, and this notebook runs on either: change `MODEL` below\n", "and nothing else. `powsm_ctc` reads its window in one pass and is what\n", - "`espnet phonemize` loads; `powsm` is the encoder-decoder, whose phones are\n", - "finer - aspiration and proper diphthongs, `pʰ` and `oʊ` where the CTC writes\n", - "`p` and `o` - and which is about four times slower on a CPU.\n", + "`espnet phonemize` loads; `powsm` is the encoder-decoder, which searches and\n", + "is about four times slower on a CPU. Both write the same phone set — a\n", + "diphthong is two symbols in either, by design — so what differs between them\n", + "is which phones they choose, not what they can say.\n", "\n", "`decode_window` is what makes the swap work: a CTC-only checkpoint is read\n", "off its CTC head, one with a decoder is searched, and the caller does not\n", @@ -141,8 +142,12 @@ "source": [ "## ``: the words, for contrast\n", "\n", - "A phonetic model will transcribe, and it is worse at it than a model built\n", - "for text. That is not a bug to report; it is what the model is for." + "A phonetic model will transcribe. On English — where a model trained for\n", + "text has seen far more — it is the weaker choice, and what comes back below\n", + "shows it. On a language that general speech corpora barely cover, the\n", + "comparison is a different one: POWSM is competitive there, because it was\n", + "trained on a phone-annotated collection rather than on whatever happens to\n", + "be plentiful." ] }, { @@ -160,9 +165,9 @@ "source": [ "## ``: phones for words you give\n", "\n", - "Now the audio is not the only input. Hand it the words that were said, and\n", - "it answers with the phones for *those* words rather than for whatever it\n", - "thought it heard." + "Now the audio is not the only input. The words you hand it ground the\n", + "answer — the phones come out for *those* words — and it is still listening\n", + "to the audio, which is where the pronunciation comes from." ] }, {