Skip to content

A demo for POWSM's four tasks - #62

Merged
sw005320 merged 3 commits into
masterfrom
demos/powsm
Sep 22, 2026
Merged

sw005320 merged 3 commits into
masterfrom
demos/powsm

Conversation

@sw005320

Copy link
Copy Markdown
Contributor

Phone recognition has a command (espnet phonemize), an MCP tool and — with espnet#6792 — a Space. This is its notebook, and the one place where all four of POWSM's tasks sit side by side on the same six seconds of speech:

task you give it answers
<asr> the audio the words
<pr> the audio the phones, in IPA
<g2p> the audio and the words the phones
<p2g> the audio and the phones the words

The last two are why the notebook is worth having: they are the part of POWSM that no one-argument command can show, because they take something written alongside the audio. The notebook hands <g2p> the words that were said and gets phones for those words, then hands <p2g> those phones back and gets words again.

<asr> is in there for contrast, and it is visibly worse than a model built for text — which the notebook says, rather than leaving a reader to conclude the model is broken.

Two things it reads off the checkpoint rather than hard-coding, since a reader will copy this into their own script: the 20-second window (OWSM's is 30) and s2t.no_language(), which is <unk> here and <nolang> for OWSM.

Name. s2t_pr_demo.ipynb, per the rule in Demos/README.md that tools/check_layout.py enforces: <task>_<variant>_demo.ipynb, and phone recognition is a variant of s2t rather than a task of its own.

Verified with tools/run_notebook.py against espnet master: RESULT ok, 16 cells.

CI skips until the release. The notebook pins espnet==202610.post2, which carries the page and the decode_window that passes a prompt on; the job says which release it is waiting for, as #61 taught it to.

🤖 Generated with Claude Code

Phone recognition has a command, an MCP tool and, with espnet#6792, a
Space; this is its notebook, and it is the only one of the four places
where all four of POWSM's tasks can be read side by side.

<pr> hears the phones, <asr> transcribes and is visibly worse at it - which
is what a phonetic model is - and then the two that take something written
with the audio: <g2p> is given the words and answers with the phones,
<p2g> is given the phones and answers with the words. Same six seconds of
speech throughout, so the four answers can be compared.

Named s2t_pr_demo.ipynb for the rule in Demos/README.md: <task>_<variant>,
and phone recognition is a variant of s2t rather than a task of its own.

Pins espnet==202610.post2, which carries the page and the decode_window
that passes a prompt on; until that release is on PyPI the job says so and
skips, which is what allow_unreleased_pin is for.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@mergify

mergify Bot commented Sep 22, 2026

Copy link
Copy Markdown

Tick the box to add this pull request to the merge queue (same as @mergifyio queue).

  • Queue this pull request

`MODEL` at the top, and nothing else to change: `espnet/powsm_ctc` or
`espnet/powsm`. What makes the swap work is `decode_window`, which reads a
CTC-only checkpoint off its CTC head and searches one that has a decoder, so
the notebook does not have to know which it is holding - `best_path`, which
this used before, would have answered the CTC branch's own task on the
encoder-decoder and quietly returned phones for every task.

The note says what the trade is, measured rather than remembered: the
encoder-decoder's phones are finer - `pʰ`, `tʰ`, `oʊ` where the CTC writes
`p`, `t`, `o` - and it is about four times slower on a CPU.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@sw005320

Copy link
Copy Markdown
Contributor Author

Made switchable, per the CLI/MCP-on-CTC, notebook/Space-switchable split:

MODEL = "espnet/powsm_ctc"  # or "espnet/powsm", the encoder-decoder

and nothing else to change. What makes the swap actually work is that the notebook now decodes through decode_window rather than best_path: a CTC-only checkpoint is read off its CTC head, one with a decoder is searched, and the notebook does not have to know which it is holding. With best_path a reader who swapped the tag would have got the encoder-decoder's CTC branch, which answers what that branch was trained on — phones, whatever task you ask for.

The note beside it says what the trade is, measured on this file rather than remembered: the encoder-decoder's phones are finer (pʰ, tʰ, oʊ where the CTC writes p, t, o) and it is about four times slower on a CPU. The fuller table, including what happens to its text tasks and the prompt experiment, is in espnet#6792.

Re-ran after the change: RESULT ok, 16 cells.

Comment thread Demos/s2t_pr_demo.ipynb Outdated
"There are two POWSMs, and this notebook runs on either: change `MODEL` below\n",
"and nothing else. `powsm_ctc` reads its window in one pass and is what\n",
"`espnet phonemize` loads; `powsm` is the encoder-decoder, whose phones are\n",
"finer - aspiration and proper diphthongs, `pʰ` and `oʊ` where the CTC writes\n",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the phone vocab are the same for both models, not "finer". And the decoder also writes "o" and "ʊ" separately too, the model is designed to split diphthongs.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Corrected, and thank you — I had inferred an inventory difference from an output difference, which was wrong.

The section now says: "Both write the same phone set — a diphthong is two symbols in either, by design — so what differs between them is which phones they choose, not what they can say."

I had made the same claim in the POWSM Space card (espnet#6792); that is fixed too, and the row now reports what I actually observed on that file — ð ə s e ɪ l ʌ v from the decoder against d ə s e ɪ l ɔ v from the CTC head — with no explanation attached to it.

Comment thread Demos/s2t_pr_demo.ipynb Outdated
"source": [
"## `<asr>`: the words, for contrast\n",
"\n",
"A phonetic model will transcribe, and it is worse at it than a model built\n",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Worse compared to high resource languages like English, but competitive for languages that are less well presented in general datasets.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Taken — "worse at transcription" without saying where is the kind of sentence that sends a reader away with the wrong impression. The cell now reads:

A phonetic model will transcribe. On English — where a model trained for text has seen far more — it is the weaker choice, and what comes back below shows it. On a language that general speech corpora barely cover, the comparison is a different one: POWSM is competitive there, because it was trained on a phone-annotated collection rather than on whatever happens to be plentiful.

If that last clause overstates the reason, tell me and I will cut it back to the claim alone — I am describing why from the outside.

Comment thread Demos/s2t_pr_demo.ipynb Outdated
"\n",
"Now the audio is not the only input. Hand it the words that were said, and\n",
"it answers with the phones for *those* words rather than for whatever it\n",
"thought it heard."

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"whatever it thought it heard" is too much. The text input here grounds the phonemic output, but it still listens to the audio.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed. "Whatever it thought it heard" made the audio sound discarded, which is not what the task does. Now:

Now the audio is not the only input. The words you hand it ground the answer — the phones come out for those words — and it is still listening to the audio, which is where the pronunciation comes from.

That also makes the demo's own result read correctly: the phones that come back for a given text are that text as this speaker said it, which is the interesting part and what I had talked the reader out of.

@chinjouli, reviewing:

 - the two models share a phone set, and the decoder splits diphthongs by
   design, so "finer" was me reading an inventory difference out of an
   output difference. What differs is which phones each chose.
 - "worse at transcription" needed its condition: against English, where a
   model trained for text has seen far more. On a language general corpora
   barely cover, POWSM is competitive.
 - "whatever it thought it heard" made the audio sound discarded. The text
   grounds the phones; the model is still listening, and that is where the
   pronunciation comes from - which is the interesting part of the task and
   what my sentence had talked the reader out of.

RESULT ok, 16 cells, after the change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@sw005320
sw005320 merged commit ea8568a into master Sep 22, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants