Skip to content

Bump vernacula-phonemizer to #1420 (a rewrite must not collapse provenance under it) - #241

Merged
christopherthompson81 merged 2 commits into
mainfrom
chore/bump-phonemizer-1420
Sep 22, 2026
Merged

christopherthompson81 merged 2 commits into
mainfrom
chore/bump-phonemizer-1420

Conversation

@christopherthompson81

Copy link
Copy Markdown
Owner

Pin ac60b3c4 → 6e2165c7, five commits. The one this repo was waiting for is #1420: normalizeRomans rewrites on \p{L}+ and runs over every language, and in a script without spaces there's no word break for that match to stop at — so it matched the whole clause and stamped its span across every token it produced.

Measured on the raw trace, not on our units

The merge added in #240 collapses these downstream either way, so counting units would have measured our own workaround instead of upstream's fix:

                                         ac60b3c4     6e2165c7
ja  rows with 2+ tokens sharing a span        14            3
cmn rows with 2+ tokens sharing a span         3            3

ja 14 → 3 reproduces upstream's figure exactly.

⚠ Mandarin was never affected, and the reason is worth keeping. The cmn trace returns one token per sentence carrying per-syllable IPA, so a whole-clause match has no second token to flatten against. The defect needed several tokens to collapse and only ja had them — so the cmn residue of 3 was never this bug.

What the reader gets

Through KokoroPhonemizer, same corpora:

        median chars per unit    median units per sentence
ja        4.33  ->  4.08              11  ->  12
cmn       1.00  ->  1.00              35  ->  35
degenerate / overlapping / out-of-range: 0 throughout

Eleven rows stop being merged into coarser units and get their real boundaries back. Small in the median because it's 11 of 123 rows — but those were the worst rows, the ones offering eight units each covering 30 of 34 characters.

The #240 merge is not inert afterwards. ja keeps 3 rows of numeral/unit overlap (83 m against 83 mです), which is a different and smaller thing. Run 4 predicted the guard would stop firing once upstream landed; Run 6 corrected that, and this measures it.

Scope — this is not a pure fix pickup

One commit is core provenance (Rewriter, Provenance, JsRegex, and the TypeScript twin). The other four are English data and lexicon work, including #1418 reverting 26 rows on re-arbitrated evidence and #1425 porting code-slot unit guards to C# — the side we run. So English behaviour changes here too, and the core change touches every language rather than just ja.

Verification

  • cold-trace gate on the new pin: 189 of 189 languages, no poisons
  • suite: 95 + 257 + 23 + 362, 0 failures

🤖 Generated with Claude Code

https://claude.ai/code/session_01KfTCTjC7FTnNm3aS2zszKE

…e provenance under it)

Pin ac60b3c4 -> 6e2165c7, five commits. The one this repo was waiting for is
#1420: normalizeRomans rewrites on \p{L}+ and runs over every language, and in a
script without spaces there is no word break for that match to stop at, so it
matched the whole clause and stamped its span across every token it produced.

Measured on the raw trace rather than on our own units, because the local merge
collapses these downstream either way and counting units would have measured the
workaround instead of the fix:

                                         ac60b3c4     6e2165c7
  ja  rows with 2+ tokens sharing a span      14            3
  cmn rows with 2+ tokens sharing a span       3            3

ja 14 -> 3 reproduces upstream's figure exactly.

MANDARIN WAS NEVER AFFECTED, and the reason is worth keeping: the cmn trace
returns ONE token per sentence carrying per-syllable IPA, so a whole-clause
match has no second token to flatten against. The defect needed several tokens
to collapse and only ja had them. The cmn residue of 3 was never this bug.

What the reader gets, through KokoroPhonemizer on the same corpora:

          median chars per unit    median units per sentence
  ja        4.33  ->  4.08              11  ->  12
  cmn       1.00  ->  1.00              35  ->  35

Eleven rows stop being merged into coarser units and get their real boundaries
back. Small in the median because it is 11 of 123 rows -- but those were the
worst rows, the ones offering eight units each covering 30 of 34 characters.

The merge added in #240 is NOT inert afterwards: ja keeps 3 rows of numeral/unit
overlap ("83 m" against "83 mです"), which is overlap rather than collapse. Run 4
predicted the guard would stop firing; Run 6 corrected that and this measures it.

Cold-trace gate re-run on the new pin: 189 of 189 languages, no poisons.
Suite: 95 + 257 + 23 + 362, 0 failures.

Scope: one commit is core provenance (Rewriter, Provenance, JsRegex and the
TypeScript twin); the rest are English data and lexicon work, including #1418
reverting 26 rows on re-arbitrated evidence and #1425 porting code-slot unit
guards to C#. So this carries English behaviour changes as well as the fix.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KfTCTjC7FTnNm3aS2zszKE
@christopherthompson81

Copy link
Copy Markdown
Owner Author

Review pass

The ja win was already measured. What this PR actually needed reviewing is the part that isn't Japanese: #1420 changes core provenance (Rewriter, Provenance, JsRegex) for every language, and our trace coverage outside ja/cmn is thin. A regression there would show as lost spans or new collapses in languages nothing here asserts on.

Measured across both pins:

              tokens   withInputSpan   rows with shared spans
en  ac60b3c4    2593    2593 (100%)             25
en  6e2165c7    2593    2593 (100%)             23
es  ac60b3c4    3280    3280 (100%)             12
es  6e2165c7    3280    3280 (100%)             12
fr  ac60b3c4    2558    2558 (100%)              6
fr  6e2165c7    2558    2558 (100%)              6

Token counts identical, span coverage 100% on both, and English shared-span rows drop 25 → 23. No spans lost and no new collapses anywhere, which is the property the core change had to preserve.

Worth noting the residual shared-span rows in en/es/fr are the legitimate kind — a numeral expansion producing several tokens from one source span. English doesn't reach FromTrace at all (it segments on whitespace), so they don't affect units here either way.

Narrowing my own scope warning

The PR body says this "carries English behaviour changes as well as the fix." True of the dictionary, but I left it vaguer than the evidence supports. Measured directly — Kokoro phonemes for every English golden row, both pins:

English golden rows whose reading differs: 0 of 114

#1418's 26-row revert and #1425's code-slot unit guards target specific lexical items (µin, CoCr, V6L 2T5) that the golden prose corpus doesn't contain.

⚠ Zero-diff here means "not exercised", not "no change" — those rows genuinely changed, and this corpus simply doesn't reach them. So the honest statement is narrower than my warning and narrower than a clean bill of health: the English changes are real, targeted at tokens absent from our measurable surface, and not detectable from this repo.

Verification

  • cold-trace gate on the new pin: 189 of 189 languages, no poisons
  • suite: 95 + 257 + 23 + 362, 0 failures
  • ja 14 → 3 collapsed rows; reader units 4.33 → 4.08 chars, 0 degenerate

Run 7 said the bump "carries English behaviour changes as well as the fix",
which is true of the dictionary and vaguer than the evidence supports.

The part the review actually had to check is that #1420 changes core provenance
for EVERY language while our trace coverage outside ja/cmn is thin. Measured
across both pins: en/es/fr token counts identical, InputSpan coverage 100% on
both, and English rows with shared spans 25 -> 23. No spans lost, no new
collapses. The residual shared-span rows are the legitimate kind -- a numeral
expansion producing several tokens from one source span -- and English does not
reach FromTrace at all, so they do not affect units here either way.

On readings: 0 of 114 English golden rows differ between the pins. #1418 and
#1425 target specific lexical items (µin, CoCr, V6L 2T5) the golden prose corpus
does not contain.

ZERO-DIFF MEANS "NOT EXERCISED", NOT "NO CHANGE". Those rows genuinely changed.
The honest statement is narrower than the warning and narrower than a clean bill
of health: real English changes, aimed at tokens absent from our measurable
surface, not detectable from this repo.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KfTCTjC7FTnNm3aS2zszKE
@christopherthompson81
christopherthompson81 merged commit d015038 into main Sep 22, 2026
1 check passed
@christopherthompson81
christopherthompson81 deleted the chore/bump-phonemizer-1420 branch September 22, 2026 18:00
christopherthompson81 added a commit that referenced this pull request Sep 22, 2026
Pin 6e2165c7 -> ed2d400e, seven commits, ALL fix(en) -- no core and no Japanese
changes, so the segmentation work in #240 and #241 is untouched and the question
is entirely what English does.

Kokoro phonemes for every golden row in the five languages this app speaks,
both pins:

  en:  1 of 114        es: 0 of 111        fr: 0 of 99
  ja:  0 of 123        cmn: 0 of 102

ONE ROW, AND IT IS THE FIX DOING WHAT IT SAYS. `re-established` read with ɹˈA --
ray, the note of the scale -- and now reads ɹˈi:

  - ... hæv bɪn ɹˈA ɪstˈæblɪʃt bᵻtwin ...
  + ... hæv bɪn ɹˈi ɪstˈæblɪʃt bᵻtwin ...

That is #1433. A hyphenated re- is ordinary English prose rather than a niche
token, which is why this batch reaches the corpus where the previous one did
not. It is also the only change with over-application risk, since a hyphenated
`re` that is not the prefix would now read ɹˈi; our corpus has one instance and
it is correct, and anything broader is upstream's test surface.

THE OTHER SIX ARE STILL INVISIBLE FROM HERE, which is the caveat from #241
rather than a clean result: TSO, prime marks as feet and inches, `.5 kg` reading
"five king", ° in a coordinate, an enumerated list lead-in -- all real changes,
aimed at tokens this prose corpus does not contain. One visible row out of seven
commits is a statement about the corpus, not about the batch.

Review compared trace health across the pins, since new token shapes are what a
lexical batch could perturb and the trace is what segmentation indexes into:
traced rows, token counts, InputSpan coverage and shared-span rows are identical
on all five languages. ja's shared-span residue is still 3, so the #240 merge
has exactly the same work to do.

Cold-trace gate on the new pin: 189 of 189 languages, no poisons.
Suite: 95 + 257 + 23 + 362, 0 failures.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KfTCTjC7FTnNm3aS2zszKE
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant