A demo for POWSM's four tasks - #62
Conversation
Phone recognition has a command, an MCP tool and, with espnet#6792, a Space; this is its notebook, and it is the only one of the four places where all four of POWSM's tasks can be read side by side. <pr> hears the phones, <asr> transcribes and is visibly worse at it - which is what a phonetic model is - and then the two that take something written with the audio: <g2p> is given the words and answers with the phones, <p2g> is given the phones and answers with the words. Same six seconds of speech throughout, so the four answers can be compared. Named s2t_pr_demo.ipynb for the rule in Demos/README.md: <task>_<variant>, and phone recognition is a variant of s2t rather than a task of its own. Pins espnet==202610.post2, which carries the page and the decode_window that passes a prompt on; until that release is on PyPI the job says so and skips, which is what allow_unreleased_pin is for. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
05e53d3 to
334562c
Compare
|
Tick the box to add this pull request to the merge queue (same as
|
`MODEL` at the top, and nothing else to change: `espnet/powsm_ctc` or `espnet/powsm`. What makes the swap work is `decode_window`, which reads a CTC-only checkpoint off its CTC head and searches one that has a decoder, so the notebook does not have to know which it is holding - `best_path`, which this used before, would have answered the CTC branch's own task on the encoder-decoder and quietly returned phones for every task. The note says what the trade is, measured rather than remembered: the encoder-decoder's phones are finer - `pʰ`, `tʰ`, `oʊ` where the CTC writes `p`, `t`, `o` - and it is about four times slower on a CPU. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
Made switchable, per the CLI/MCP-on-CTC, notebook/Space-switchable split: MODEL = "espnet/powsm_ctc" # or "espnet/powsm", the encoder-decoderand nothing else to change. What makes the swap actually work is that the notebook now decodes through The note beside it says what the trade is, measured on this file rather than remembered: the encoder-decoder's phones are finer ( Re-ran after the change: |
| "There are two POWSMs, and this notebook runs on either: change `MODEL` below\n", | ||
| "and nothing else. `powsm_ctc` reads its window in one pass and is what\n", | ||
| "`espnet phonemize` loads; `powsm` is the encoder-decoder, whose phones are\n", | ||
| "finer - aspiration and proper diphthongs, `pʰ` and `oʊ` where the CTC writes\n", |
There was a problem hiding this comment.
the phone vocab are the same for both models, not "finer". And the decoder also writes "o" and "ʊ" separately too, the model is designed to split diphthongs.
There was a problem hiding this comment.
Corrected, and thank you — I had inferred an inventory difference from an output difference, which was wrong.
The section now says: "Both write the same phone set — a diphthong is two symbols in either, by design — so what differs between them is which phones they choose, not what they can say."
I had made the same claim in the POWSM Space card (espnet#6792); that is fixed too, and the row now reports what I actually observed on that file — ð ə s e ɪ l ʌ v from the decoder against d ə s e ɪ l ɔ v from the CTC head — with no explanation attached to it.
| "source": [ | ||
| "## `<asr>`: the words, for contrast\n", | ||
| "\n", | ||
| "A phonetic model will transcribe, and it is worse at it than a model built\n", |
There was a problem hiding this comment.
Worse compared to high resource languages like English, but competitive for languages that are less well presented in general datasets.
There was a problem hiding this comment.
Taken — "worse at transcription" without saying where is the kind of sentence that sends a reader away with the wrong impression. The cell now reads:
A phonetic model will transcribe. On English — where a model trained for text has seen far more — it is the weaker choice, and what comes back below shows it. On a language that general speech corpora barely cover, the comparison is a different one: POWSM is competitive there, because it was trained on a phone-annotated collection rather than on whatever happens to be plentiful.
If that last clause overstates the reason, tell me and I will cut it back to the claim alone — I am describing why from the outside.
| "\n", | ||
| "Now the audio is not the only input. Hand it the words that were said, and\n", | ||
| "it answers with the phones for *those* words rather than for whatever it\n", | ||
| "thought it heard." |
There was a problem hiding this comment.
"whatever it thought it heard" is too much. The text input here grounds the phonemic output, but it still listens to the audio.
There was a problem hiding this comment.
Fixed. "Whatever it thought it heard" made the audio sound discarded, which is not what the task does. Now:
Now the audio is not the only input. The words you hand it ground the answer — the phones come out for those words — and it is still listening to the audio, which is where the pronunciation comes from.
That also makes the demo's own result read correctly: the phones that come back for a given text are that text as this speaker said it, which is the interesting part and what I had talked the reader out of.
@chinjouli, reviewing: - the two models share a phone set, and the decoder splits diphthongs by design, so "finer" was me reading an inventory difference out of an output difference. What differs is which phones each chose. - "worse at transcription" needed its condition: against English, where a model trained for text has seen far more. On a language general corpora barely cover, POWSM is competitive. - "whatever it thought it heard" made the audio sound discarded. The text grounds the phones; the model is still listening, and that is where the pronunciation comes from - which is the interesting part of the task and what my sentence had talked the reader out of. RESULT ok, 16 cells, after the change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Phone recognition has a command (
espnet phonemize), an MCP tool and — with espnet#6792 — a Space. This is its notebook, and the one place where all four of POWSM's tasks sit side by side on the same six seconds of speech:<asr><pr><g2p><p2g>The last two are why the notebook is worth having: they are the part of POWSM that no one-argument command can show, because they take something written alongside the audio. The notebook hands
<g2p>the words that were said and gets phones for those words, then hands<p2g>those phones back and gets words again.<asr>is in there for contrast, and it is visibly worse than a model built for text — which the notebook says, rather than leaving a reader to conclude the model is broken.Two things it reads off the checkpoint rather than hard-coding, since a reader will copy this into their own script: the 20-second window (OWSM's is 30) and
s2t.no_language(), which is<unk>here and<nolang>for OWSM.Name.
s2t_pr_demo.ipynb, per the rule inDemos/README.mdthattools/check_layout.pyenforces:<task>_<variant>_demo.ipynb, and phone recognition is a variant ofs2trather than a task of its own.Verified with
tools/run_notebook.pyagainst espnet master:RESULT ok, 16 cells.CI skips until the release. The notebook pins
espnet==202610.post2, which carries the page and thedecode_windowthat passes a prompt on; the job says which release it is waiting for, as #61 taught it to.🤖 Generated with Claude Code