Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
31 changes: 31 additions & 0 deletions .github/workflows/s2t_pr_demo.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,31 @@
# Demos/s2t_pr_demo.ipynb, every Sunday. The badge on this workflow is what
# the README shows beside that notebook, which is why it has a file of its own:
# GitHub's badge is per workflow, and cannot show one job of a matrix.

name: s2t_pr_demo

on:
schedule:
# Sundays, 05:00 UTC
- cron: "0 5 * * 0"
workflow_dispatch:
pull_request:
paths:
- "Demos/s2t_pr_demo.ipynb"
- "tools/**"
- ".github/workflows/s2t_pr_demo.yml"
- ".github/workflows/_run_notebook.yml"

permissions:
contents: read

jobs:
run:
uses: ./.github/workflows/_run_notebook.yml
with:
notebook: Demos/s2t_pr_demo.ipynb
# POWSM's <g2p> and <p2g> reach the page through espnet2.bin.demo, and
# this notebook pins the release that carries them (espnet#6791 cuts
# it). Until it is on PyPI the job says so and skips. Delete this line
# when 202610.post2 is published.
allow_unreleased_pin: true
1 change: 1 addition & 0 deletions Demos/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -40,6 +40,7 @@ the answer now.
| [`asr_streaming_demo.ipynb`](asr_streaming_demo.ipynb) | [![asr_streaming_demo](https://github.com/espnet/notebook/actions/workflows/asr_streaming_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/asr_streaming_demo.yml) | Watch the words appear while the audio is still arriving |
| [`st_demo.ipynb`](st_demo.ipynb) | [![st_demo](https://github.com/espnet/notebook/actions/workflows/st_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/st_demo.yml) | Translate English speech into German, French and Chinese — the same model |
| [`s2t_align_demo.ipynb`](s2t_align_demo.ipynb) | [![s2t_align_demo](https://github.com/espnet/notebook/actions/workflows/s2t_align_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/s2t_align_demo.yml) | Line text up with the audio it was said in, and score how well they agree |
| [`s2t_pr_demo.ipynb`](s2t_pr_demo.ipynb) | [![s2t_pr_demo](https://github.com/espnet/notebook/actions/workflows/s2t_pr_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/s2t_pr_demo.yml) | Hear the phones, and cross between phones and words with POWSM |
| [`tts_demo.ipynb`](tts_demo.ipynb) | [![tts_demo](https://github.com/espnet/notebook/actions/workflows/tts_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/tts_demo.yml) | Type a sentence, hear it spoken — one English voice, then 128 of them |
| [`enh_demo.ipynb`](enh_demo.ipynb) | [![enh_demo](https://github.com/espnet/notebook/actions/workflows/enh_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/enh_demo.yml) | Pull speech out of noise, and measure how much it helped |
| [`spk_demo.ipynb`](spk_demo.ipynb) | [![spk_demo](https://github.com/espnet/notebook/actions/workflows/spk_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/spk_demo.yml) | Turn a voice into a vector, and score two recordings against each other |
Expand Down
236 changes: 236 additions & 0 deletions Demos/s2t_pr_demo.ipynb
Original file line number Diff line number Diff line change
@@ -0,0 +1,236 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# Phones, and the two ways across to words\n",
"\n",
"[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/espnet/notebook/blob/master/Demos/s2t_pr_demo.ipynb) [![s2t_pr_demo](https://github.com/espnet/notebook/actions/workflows/s2t_pr_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/s2t_pr_demo.yml)\n",
"\n",
"[POWSM](https://arxiv.org/abs/2510.24992) is a phonetic foundation model:\n",
"it hears speech and answers in phones. POWSM-CTC is the encoder-only\n",
"variant, and it does four things with one recording —\n",
"\n",
"| task | you give | it answers |\n",
"| :-- | :-- | :-- |\n",
"| `<asr>` | the audio | the words |\n",
"| `<pr>` | the audio | the phones, in IPA |\n",
"| `<g2p>` | the audio **and the words** | the phones |\n",
"| `<p2g>` | the audio **and the phones** | the words |\n",
"\n",
"— which is what this notebook runs, one after the other, on the same six\n",
"seconds of speech. CPU is enough."
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Install"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"%pip install -q \"espnet==202610.post2\" espnet_model_zoo librosa"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## A recording"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"import librosa\n",
"from IPython.display import Audio, display\n",
"\n",
"!wget -q -O sample.wav https://github.com/espnet/espnet/raw/master/test_utils/ctc_align_test.wav\n",
"speech, rate = librosa.load(\"sample.wav\", sr=16000)\n",
"display(Audio(speech, rate=rate))"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## The model\n",
"\n",
"There are two POWSMs, and this notebook runs on either: change `MODEL` below\n",
"and nothing else. `powsm_ctc` reads its window in one pass and is what\n",
"`espnet phonemize` loads; `powsm` is the encoder-decoder, which searches and\n",
"is about four times slower on a CPU. Both write the same phone set — a\n",
"diphthong is two symbols in either, by design — so what differs between them\n",
"is which phones they choose, not what they can say.\n",
"\n",
"`decode_window` is what makes the swap work: a CTC-only checkpoint is read\n",
"off its CTC head, one with a decoder is searched, and the caller does not\n",
"have to know which it is holding.\n",
"\n",
"POWSM reads a 20-second window, so a shorter clip is padded to it. The\n",
"checkpoint also carries its own spelling of \"work the language out yourself\"\n",
"- `<unk>` here, where OWSM writes `<nolang>` - so ask the model rather than\n",
"typing a symbol."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"import librosa\n",
"from espnet2.bin.s2t_inference import Speech2Text\n",
"\n",
"MODEL = \"espnet/powsm_ctc\" # or \"espnet/powsm\", the encoder-decoder\n",
"s2t = Speech2Text.from_pretrained(MODEL, device=\"cpu\")\n",
"\n",
"window = s2t.preprocessor_conf[\"speech_length\"]\n",
"nolang = s2t.no_language()\n",
"padded = librosa.util.fix_length(speech, size=rate * window)\n",
"print(f\"{window} s window, language symbol {nolang}, CTC-only: {s2t.ctc_only}\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## `<pr>`: the phones\n",
"\n",
"The task it is built for. POWSM writes each phone between slashes, so that\n",
"a phone spelled like a BPE token is still one token; spaced out is easier\n",
"to read and is what anything counting or aligning them wants."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"import re\n",
"\n",
"\n",
"def phones(decoded):\n",
" \"\"\"The phones out of POWSM's /p//h//o/ spelling.\"\"\"\n",
" return \" \".join(re.findall(r\"/([^/]+)/\", decoded))\n",
"\n",
"\n",
"def run(task, text_prev=\"<na>\"):\n",
" \"\"\"One window, decoded the way this checkpoint has to be.\"\"\"\n",
" decoded = s2t.decode_window(padded, nolang, task, text_prev)\n",
" return decoded.split(\">\")[-1].strip() if \"<\" in decoded else decoded\n",
"\n",
"\n",
"said_in_phones = phones(run(\"<pr>\"))\n",
"print(said_in_phones)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## `<asr>`: the words, for contrast\n",
"\n",
"A phonetic model will transcribe. On English — where a model trained for\n",
"text has seen far more — it is the weaker choice, and what comes back below\n",
"shows it. On a language that general speech corpora barely cover, the\n",
"comparison is a different one: POWSM is competitive there, because it was\n",
"trained on a phone-annotated collection rather than on whatever happens to\n",
"be plentiful."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"print(run(\"<asr>\"))"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## `<g2p>`: phones for words you give\n",
"\n",
"Now the audio is not the only input. The words you hand it ground the\n",
"answer — the phones come out for *those* words — and it is still listening\n",
"to the audio, which is where the pronunciation comes from."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"said = \"the sale of the hotels is part of holiday's strategy\"\n",
"\n",
"print(phones(run(\"<g2p>\", text_prev=said)))"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## `<p2g>`: words for phones you give\n",
"\n",
"The other direction. The prompt is phones, in the spelling the model was\n",
"trained on — each between slashes."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"as_prompt = \"/\" + \"//\".join(said_in_phones.split()) + \"/\"\n",
"print(as_prompt[:70], \"...\")\n",
"\n",
"print(run(\"<p2g>\", text_prev=as_prompt))"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Where next\n",
"\n",
"- **From the terminal**: `espnet phonemize sample.wav` is `<pr>`, in one line\n",
"- **In the browser**: the [powsm-ctc Space](https://huggingface.co/spaces/espnet/powsm-ctc)\n",
" offers all four tasks, and `espnet demo --model espnet/powsm_ctc` is the same page\n",
"- **From an assistant**: the MCP server's `phonemize` tool\n",
"- **Alignment**, with the same kind of model: [`s2t_align_demo.ipynb`](s2t_align_demo.ipynb)\n",
"- **The encoder-decoder POWSM**, [`espnet/powsm`](https://huggingface.co/espnet/powsm),\n",
" answers the same four tasks with a search rather than a CTC pass"
]
}
],
"metadata": {
"colab": {
"provenance": []
},
"kernelspec": {
"display_name": "Python 3",
"name": "python3"
},
"language_info": {
"name": "python"
}
},
"nbformat": 4,
"nbformat_minor": 0
}
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,7 @@ One per task, flat in [`Demos/`](Demos), each short enough to read in a sitting.
| [`asr_streaming_demo.ipynb`](Demos/asr_streaming_demo.ipynb) | [![asr_streaming_demo](https://github.com/espnet/notebook/actions/workflows/asr_streaming_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/asr_streaming_demo.yml) | Watch the words appear while the audio is still arriving |
| [`st_demo.ipynb`](Demos/st_demo.ipynb) | [![st_demo](https://github.com/espnet/notebook/actions/workflows/st_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/st_demo.yml) | Translate English speech into German, French and Chinese — the same model |
| [`s2t_align_demo.ipynb`](Demos/s2t_align_demo.ipynb) | [![s2t_align_demo](https://github.com/espnet/notebook/actions/workflows/s2t_align_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/s2t_align_demo.yml) | Line text up with the audio it was said in, and score how well they agree |
| [`s2t_pr_demo.ipynb`](Demos/s2t_pr_demo.ipynb) | [![s2t_pr_demo](https://github.com/espnet/notebook/actions/workflows/s2t_pr_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/s2t_pr_demo.yml) | Hear the phones, and cross between phones and words with POWSM |
| [`tts_demo.ipynb`](Demos/tts_demo.ipynb) | [![tts_demo](https://github.com/espnet/notebook/actions/workflows/tts_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/tts_demo.yml) | Type a sentence, hear it spoken — one English voice, then 128 of them |
| [`enh_demo.ipynb`](Demos/enh_demo.ipynb) | [![enh_demo](https://github.com/espnet/notebook/actions/workflows/enh_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/enh_demo.yml) | Pull speech out of noise, and measure how much it helped |
| [`spk_demo.ipynb`](Demos/spk_demo.ipynb) | [![spk_demo](https://github.com/espnet/notebook/actions/workflows/spk_demo.yml/badge.svg)](https://github.com/espnet/notebook/actions/workflows/spk_demo.yml) | Turn a voice into a vector, and score two recordings against each other |
Expand Down
Loading