Repository navigation
Bump vernacula-phonemizer to #1420 (a rewrite must not collapse provenance under it) - #241
Conversation
…e provenance under it)
Pin ac60b3c4 -> 6e2165c7, five commits. The one this repo was waiting for is
#1420: normalizeRomans rewrites on \p{L}+ and runs over every language, and in a
script without spaces there is no word break for that match to stop at, so it
matched the whole clause and stamped its span across every token it produced.
Measured on the raw trace rather than on our own units, because the local merge
collapses these downstream either way and counting units would have measured the
workaround instead of the fix:
ac60b3c4 6e2165c7
ja rows with 2+ tokens sharing a span 14 3
cmn rows with 2+ tokens sharing a span 3 3
ja 14 -> 3 reproduces upstream's figure exactly.
MANDARIN WAS NEVER AFFECTED, and the reason is worth keeping: the cmn trace
returns ONE token per sentence carrying per-syllable IPA, so a whole-clause
match has no second token to flatten against. The defect needed several tokens
to collapse and only ja had them. The cmn residue of 3 was never this bug.
What the reader gets, through KokoroPhonemizer on the same corpora:
median chars per unit median units per sentence
ja 4.33 -> 4.08 11 -> 12
cmn 1.00 -> 1.00 35 -> 35
Eleven rows stop being merged into coarser units and get their real boundaries
back. Small in the median because it is 11 of 123 rows -- but those were the
worst rows, the ones offering eight units each covering 30 of 34 characters.
The merge added in #240 is NOT inert afterwards: ja keeps 3 rows of numeral/unit
overlap ("83 m" against "83 mです"), which is overlap rather than collapse. Run 4
predicted the guard would stop firing; Run 6 corrected that and this measures it.
Cold-trace gate re-run on the new pin: 189 of 189 languages, no poisons.
Suite: 95 + 257 + 23 + 362, 0 failures.
Scope: one commit is core provenance (Rewriter, Provenance, JsRegex and the
TypeScript twin); the rest are English data and lexicon work, including #1418
reverting 26 rows on re-arbitrated evidence and #1425 porting code-slot unit
guards to C#. So this carries English behaviour changes as well as the fix.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KfTCTjC7FTnNm3aS2zszKE
Review passThe Measured across both pins: Token counts identical, span coverage 100% on both, and English shared-span rows drop 25 → 23. No spans lost and no new collapses anywhere, which is the property the core change had to preserve. Worth noting the residual shared-span rows in Narrowing my own scope warningThe PR body says this "carries English behaviour changes as well as the fix." True of the dictionary, but I left it vaguer than the evidence supports. Measured directly — Kokoro phonemes for every English golden row, both pins: #1418's 26-row revert and #1425's code-slot unit guards target specific lexical items ( ⚠ Zero-diff here means "not exercised", not "no change" — those rows genuinely changed, and this corpus simply doesn't reach them. So the honest statement is narrower than my warning and narrower than a clean bill of health: the English changes are real, targeted at tokens absent from our measurable surface, and not detectable from this repo. Verification
|
Run 7 said the bump "carries English behaviour changes as well as the fix", which is true of the dictionary and vaguer than the evidence supports. The part the review actually had to check is that #1420 changes core provenance for EVERY language while our trace coverage outside ja/cmn is thin. Measured across both pins: en/es/fr token counts identical, InputSpan coverage 100% on both, and English rows with shared spans 25 -> 23. No spans lost, no new collapses. The residual shared-span rows are the legitimate kind -- a numeral expansion producing several tokens from one source span -- and English does not reach FromTrace at all, so they do not affect units here either way. On readings: 0 of 114 English golden rows differ between the pins. #1418 and #1425 target specific lexical items (µin, CoCr, V6L 2T5) the golden prose corpus does not contain. ZERO-DIFF MEANS "NOT EXERCISED", NOT "NO CHANGE". Those rows genuinely changed. The honest statement is narrower than the warning and narrower than a clean bill of health: real English changes, aimed at tokens absent from our measurable surface, not detectable from this repo. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KfTCTjC7FTnNm3aS2zszKE
Pin 6e2165c7 -> ed2d400e, seven commits, ALL fix(en) -- no core and no Japanese changes, so the segmentation work in #240 and #241 is untouched and the question is entirely what English does. Kokoro phonemes for every golden row in the five languages this app speaks, both pins: en: 1 of 114 es: 0 of 111 fr: 0 of 99 ja: 0 of 123 cmn: 0 of 102 ONE ROW, AND IT IS THE FIX DOING WHAT IT SAYS. `re-established` read with ɹˈA -- ray, the note of the scale -- and now reads ɹˈi: - ... hæv bɪn ɹˈA ɪstˈæblɪʃt bᵻtwin ... + ... hæv bɪn ɹˈi ɪstˈæblɪʃt bᵻtwin ... That is #1433. A hyphenated re- is ordinary English prose rather than a niche token, which is why this batch reaches the corpus where the previous one did not. It is also the only change with over-application risk, since a hyphenated `re` that is not the prefix would now read ɹˈi; our corpus has one instance and it is correct, and anything broader is upstream's test surface. THE OTHER SIX ARE STILL INVISIBLE FROM HERE, which is the caveat from #241 rather than a clean result: TSO, prime marks as feet and inches, `.5 kg` reading "five king", ° in a coordinate, an enumerated list lead-in -- all real changes, aimed at tokens this prose corpus does not contain. One visible row out of seven commits is a statement about the corpus, not about the batch. Review compared trace health across the pins, since new token shapes are what a lexical batch could perturb and the trace is what segmentation indexes into: traced rows, token counts, InputSpan coverage and shared-span rows are identical on all five languages. ja's shared-span residue is still 3, so the #240 merge has exactly the same work to do. Cold-trace gate on the new pin: 189 of 189 languages, no poisons. Suite: 95 + 257 + 23 + 362, 0 failures. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KfTCTjC7FTnNm3aS2zszKE
Pin
ac60b3c4→6e2165c7, five commits. The one this repo was waiting for is #1420:normalizeRomansrewrites on\p{L}+and runs over every language, and in a script without spaces there's no word break for that match to stop at — so it matched the whole clause and stamped its span across every token it produced.Measured on the raw trace, not on our units
The merge added in #240 collapses these downstream either way, so counting units would have measured our own workaround instead of upstream's fix:
ja14 → 3 reproduces upstream's figure exactly.⚠ Mandarin was never affected, and the reason is worth keeping. The
cmntrace returns one token per sentence carrying per-syllable IPA, so a whole-clause match has no second token to flatten against. The defect needed several tokens to collapse and onlyjahad them — so thecmnresidue of 3 was never this bug.What the reader gets
Through
KokoroPhonemizer, same corpora:Eleven rows stop being merged into coarser units and get their real boundaries back. Small in the median because it's 11 of 123 rows — but those were the worst rows, the ones offering eight units each covering 30 of 34 characters.
The #240 merge is not inert afterwards.
jakeeps 3 rows of numeral/unit overlap (83 magainst83 mです), which is a different and smaller thing. Run 4 predicted the guard would stop firing once upstream landed; Run 6 corrected that, and this measures it.Scope — this is not a pure fix pickup
One commit is core provenance (
Rewriter,Provenance,JsRegex, and the TypeScript twin). The other four are English data and lexicon work, including #1418 reverting 26 rows on re-arbitrated evidence and #1425 porting code-slot unit guards to C# — the side we run. So English behaviour changes here too, and the core change touches every language rather than justja.Verification
🤖 Generated with Claude Code
https://claude.ai/code/session_01KfTCTjC7FTnNm3aS2zszKE