diff --git a/.github/workflows/s2t_pr_demo.yml b/.github/workflows/s2t_pr_demo.yml new file mode 100644 index 00000000..43d85bd7 --- /dev/null +++ b/.github/workflows/s2t_pr_demo.yml @@ -0,0 +1,31 @@ +# Demos/s2t_pr_demo.ipynb, every Sunday. The badge on this workflow is what +# the README shows beside that notebook, which is why it has a file of its own: +# GitHub's badge is per workflow, and cannot show one job of a matrix. + +name: s2t_pr_demo + +on: + schedule: + # Sundays, 05:00 UTC + - cron: "0 5 * * 0" + workflow_dispatch: + pull_request: + paths: + - "Demos/s2t_pr_demo.ipynb" + - "tools/**" + - ".github/workflows/s2t_pr_demo.yml" + - ".github/workflows/_run_notebook.yml" + +permissions: + contents: read + +jobs: + run: + uses: ./.github/workflows/_run_notebook.yml + with: + notebook: Demos/s2t_pr_demo.ipynb + # POWSM's and reach the page through espnet2.bin.demo, and + # this notebook pins the release that carries them (espnet#6791 cuts + # it). Until it is on PyPI the job says so and skips. Delete this line + # when 202610.post2 is published. + allow_unreleased_pin: true diff --git a/Demos/README.md b/Demos/README.md index c52ee9c8..8cbe9946 100644 --- a/Demos/README.md +++ b/Demos/README.md @@ -40,6 +40,7 @@ the answer now. | [`asr_streaming_demo.ipynb`](asr_streaming_demo.ipynb) | [![asr_streaming_demo](https://github.com/espnet/notebook/actions/workflows/asr_streaming_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/asr_streaming_demo.yml) | Watch the words appear while the audio is still arriving | | [`st_demo.ipynb`](st_demo.ipynb) | [![st_demo](https://github.com/espnet/notebook/actions/workflows/st_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/st_demo.yml) | Translate English speech into German, French and Chinese — the same model | | [`s2t_align_demo.ipynb`](s2t_align_demo.ipynb) | [![s2t_align_demo](https://github.com/espnet/notebook/actions/workflows/s2t_align_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/s2t_align_demo.yml) | Line text up with the audio it was said in, and score how well they agree | +| [`s2t_pr_demo.ipynb`](s2t_pr_demo.ipynb) | [![s2t_pr_demo](https://github.com/espnet/notebook/actions/workflows/s2t_pr_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/s2t_pr_demo.yml) | Hear the phones, and cross between phones and words with POWSM | | [`tts_demo.ipynb`](tts_demo.ipynb) | [![tts_demo](https://github.com/espnet/notebook/actions/workflows/tts_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/tts_demo.yml) | Type a sentence, hear it spoken — one English voice, then 128 of them | | [`enh_demo.ipynb`](enh_demo.ipynb) | [![enh_demo](https://github.com/espnet/notebook/actions/workflows/enh_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/enh_demo.yml) | Pull speech out of noise, and measure how much it helped | | [`spk_demo.ipynb`](spk_demo.ipynb) | [![spk_demo](https://github.com/espnet/notebook/actions/workflows/spk_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/spk_demo.yml) | Turn a voice into a vector, and score two recordings against each other | diff --git a/Demos/s2t_pr_demo.ipynb b/Demos/s2t_pr_demo.ipynb new file mode 100644 index 00000000..cb9be4cc --- /dev/null +++ b/Demos/s2t_pr_demo.ipynb @@ -0,0 +1,236 @@ +{ + "cells": [ + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "# Phones, and the two ways across to words\n", + "\n", + "[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/espnet/notebook/blob/master/Demos/s2t_pr_demo.ipynb) [![s2t_pr_demo](https://github.com/espnet/notebook/actions/workflows/s2t_pr_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/s2t_pr_demo.yml)\n", + "\n", + "[POWSM](https://arxiv.org/abs/2510.24992) is a phonetic foundation model:\n", + "it hears speech and answers in phones. POWSM-CTC is the encoder-only\n", + "variant, and it does four things with one recording —\n", + "\n", + "| task | you give | it answers |\n", + "| :-- | :-- | :-- |\n", + "| `` | the audio | the words |\n", + "| `` | the audio | the phones, in IPA |\n", + "| `` | the audio **and the words** | the phones |\n", + "| `` | the audio **and the phones** | the words |\n", + "\n", + "— which is what this notebook runs, one after the other, on the same six\n", + "seconds of speech. CPU is enough." + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## Install" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "%pip install -q \"espnet==202610.post2\" espnet_model_zoo librosa" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## A recording" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "import librosa\n", + "from IPython.display import Audio, display\n", + "\n", + "!wget -q -O sample.wav https://github.com/espnet/espnet/raw/master/test_utils/ctc_align_test.wav\n", + "speech, rate = librosa.load(\"sample.wav\", sr=16000)\n", + "display(Audio(speech, rate=rate))" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## The model\n", + "\n", + "There are two POWSMs, and this notebook runs on either: change `MODEL` below\n", + "and nothing else. `powsm_ctc` reads its window in one pass and is what\n", + "`espnet phonemize` loads; `powsm` is the encoder-decoder, which searches and\n", + "is about four times slower on a CPU. Both write the same phone set — a\n", + "diphthong is two symbols in either, by design — so what differs between them\n", + "is which phones they choose, not what they can say.\n", + "\n", + "`decode_window` is what makes the swap work: a CTC-only checkpoint is read\n", + "off its CTC head, one with a decoder is searched, and the caller does not\n", + "have to know which it is holding.\n", + "\n", + "POWSM reads a 20-second window, so a shorter clip is padded to it. The\n", + "checkpoint also carries its own spelling of \"work the language out yourself\"\n", + "- `` here, where OWSM writes `` - so ask the model rather than\n", + "typing a symbol." + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "import librosa\n", + "from espnet2.bin.s2t_inference import Speech2Text\n", + "\n", + "MODEL = \"espnet/powsm_ctc\" # or \"espnet/powsm\", the encoder-decoder\n", + "s2t = Speech2Text.from_pretrained(MODEL, device=\"cpu\")\n", + "\n", + "window = s2t.preprocessor_conf[\"speech_length\"]\n", + "nolang = s2t.no_language()\n", + "padded = librosa.util.fix_length(speech, size=rate * window)\n", + "print(f\"{window} s window, language symbol {nolang}, CTC-only: {s2t.ctc_only}\")" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## ``: the phones\n", + "\n", + "The task it is built for. POWSM writes each phone between slashes, so that\n", + "a phone spelled like a BPE token is still one token; spaced out is easier\n", + "to read and is what anything counting or aligning them wants." + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "import re\n", + "\n", + "\n", + "def phones(decoded):\n", + " \"\"\"The phones out of POWSM's /p//h//o/ spelling.\"\"\"\n", + " return \" \".join(re.findall(r\"/([^/]+)/\", decoded))\n", + "\n", + "\n", + "def run(task, text_prev=\"\"):\n", + " \"\"\"One window, decoded the way this checkpoint has to be.\"\"\"\n", + " decoded = s2t.decode_window(padded, nolang, task, text_prev)\n", + " return decoded.split(\">\")[-1].strip() if \"<\" in decoded else decoded\n", + "\n", + "\n", + "said_in_phones = phones(run(\"\"))\n", + "print(said_in_phones)" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## ``: the words, for contrast\n", + "\n", + "A phonetic model will transcribe. On English — where a model trained for\n", + "text has seen far more — it is the weaker choice, and what comes back below\n", + "shows it. On a language that general speech corpora barely cover, the\n", + "comparison is a different one: POWSM is competitive there, because it was\n", + "trained on a phone-annotated collection rather than on whatever happens to\n", + "be plentiful." + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "print(run(\"\"))" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## ``: phones for words you give\n", + "\n", + "Now the audio is not the only input. The words you hand it ground the\n", + "answer — the phones come out for *those* words — and it is still listening\n", + "to the audio, which is where the pronunciation comes from." + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "said = \"the sale of the hotels is part of holiday's strategy\"\n", + "\n", + "print(phones(run(\"\", text_prev=said)))" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## ``: words for phones you give\n", + "\n", + "The other direction. The prompt is phones, in the spelling the model was\n", + "trained on — each between slashes." + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "as_prompt = \"/\" + \"//\".join(said_in_phones.split()) + \"/\"\n", + "print(as_prompt[:70], \"...\")\n", + "\n", + "print(run(\"\", text_prev=as_prompt))" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## Where next\n", + "\n", + "- **From the terminal**: `espnet phonemize sample.wav` is ``, in one line\n", + "- **In the browser**: the [powsm-ctc Space](https://huggingface.co/spaces/espnet/powsm-ctc)\n", + " offers all four tasks, and `espnet demo --model espnet/powsm_ctc` is the same page\n", + "- **From an assistant**: the MCP server's `phonemize` tool\n", + "- **Alignment**, with the same kind of model: [`s2t_align_demo.ipynb`](s2t_align_demo.ipynb)\n", + "- **The encoder-decoder POWSM**, [`espnet/powsm`](https://huggingface.co/espnet/powsm),\n", + " answers the same four tasks with a search rather than a CTC pass" + ] + } + ], + "metadata": { + "colab": { + "provenance": [] + }, + "kernelspec": { + "display_name": "Python 3", + "name": "python3" + }, + "language_info": { + "name": "python" + } + }, + "nbformat": 4, + "nbformat_minor": 0 +} diff --git a/README.md b/README.md index 69d7cc3a..16831c85 100644 --- a/README.md +++ b/README.md @@ -18,6 +18,7 @@ One per task, flat in [`Demos/`](Demos), each short enough to read in a sitting. | [`asr_streaming_demo.ipynb`](Demos/asr_streaming_demo.ipynb) | [![asr_streaming_demo](https://github.com/espnet/notebook/actions/workflows/asr_streaming_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/asr_streaming_demo.yml) | Watch the words appear while the audio is still arriving | | [`st_demo.ipynb`](Demos/st_demo.ipynb) | [![st_demo](https://github.com/espnet/notebook/actions/workflows/st_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/st_demo.yml) | Translate English speech into German, French and Chinese — the same model | | [`s2t_align_demo.ipynb`](Demos/s2t_align_demo.ipynb) | [![s2t_align_demo](https://github.com/espnet/notebook/actions/workflows/s2t_align_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/s2t_align_demo.yml) | Line text up with the audio it was said in, and score how well they agree | +| [`s2t_pr_demo.ipynb`](Demos/s2t_pr_demo.ipynb) | [![s2t_pr_demo](https://github.com/espnet/notebook/actions/workflows/s2t_pr_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/s2t_pr_demo.yml) | Hear the phones, and cross between phones and words with POWSM | | [`tts_demo.ipynb`](Demos/tts_demo.ipynb) | [![tts_demo](https://github.com/espnet/notebook/actions/workflows/tts_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/tts_demo.yml) | Type a sentence, hear it spoken — one English voice, then 128 of them | | [`enh_demo.ipynb`](Demos/enh_demo.ipynb) | [![enh_demo](https://github.com/espnet/notebook/actions/workflows/enh_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/enh_demo.yml) | Pull speech out of noise, and measure how much it helped | | [`spk_demo.ipynb`](Demos/spk_demo.ipynb) | [![spk_demo](https://github.com/espnet/notebook/actions/workflows/spk_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/spk_demo.yml) | Turn a voice into a vector, and score two recordings against each other |