Skip to content

Render the reduced de-/re-/pre- prefix vowel as ə, not ᵻ - #243

Merged
christopherthompson81 merged 2 commits into
mainfrom
fix/kokoro-prefix-schwa
Sep 23, 2026
Merged

christopherthompson81 merged 2 commits into
mainfrom
fix/kokoro-prefix-schwa

Conversation

@christopherthompson81

Copy link
Copy Markdown
Owner

Reported as "determine sounds tensed". Not a dictionary defect — our IPA is already reduced on every path — but a divergence between two spellings of the same reduced vowel, which matters because Kokoro was trained on one of them. In misaki's us_gold the prefix vowel is ⟨ə⟩ 1,057 times against ⟨ᵻ⟩ 9, so our ⟨ᵻ⟩ is effectively out of distribution for this position.

Established by synthesis, not by the distribution argument

An A/B through Kokoro itself confirmed the model renders the two differently (rms 0.05–0.08; determine comes out 50 ms longer with ⟨ᵻ⟩ against identical leading silence). A listener then preferred ⟨ə⟩ on a connected-reading sample and called the difference "pretty subtle, overall, in this reading" — which is the size of claim this change gets to make, and it's written into the code and log so nothing downstream inflates it.

The sample was built as the candidate fix rather than a blanket swap, so the listen answered the actual decision:

changed  determine  dᵻtˈɜɹmən  ->  dətˈɜɹmən
changed  describe   dᵻskɹˈIb   ->  dəskɹˈIb
changed  reduce     ɹᵻdˈus     ->  ɹədˈus
KEPT     deduce     dᵻdˈus                     (ded- carve-out)

The rule

ReduceEnglishPrefixVowel, walking groups against the group → word map. Three conditions, each earned:

  1. The source word begins de-/re-/pre- — the word, not the phonemes.
  2. Except ded- — the deduce/deduct family plus two surnames is the only place gold attests a prefix ⟨ᵻ⟩, i.e. the words with the strongest evidence in the class. Verified upstream: 0 of the 1,057 target words begin ded-, and 0 of the keep-words don't.
  3. The ⟨ᵻ⟩ must be the token's first vowel. represent is ɹˌɛpɹᵻzˈɛnt, whose ⟨ᵻ⟩ is a second syllable with nothing to do with the prefix.

It does nothing when the map is null or mismatched: a word-keyed rule that has lost the word must leave the stream alone rather than guess.

⚠ Why word-keyed is the whole point. Earlier drafts tried to express this phonologically and needed a dᵻd key plus a tie-bar guard, because dᵻd͡ʒ (degeneracy, deject) matches a bare d on half of /d͡ʒ/. All of that was manufactured by working at a layer without word identity. PhonemizeTrace is the full word → IPA trip and Phonemize already holds its map, so the carve-out is orthographic — ded- is exact, deg-/dej- are excluded by spelling, and the affricate hazard never arises.

⚠ dejected → dəʤˈɛktᵻd is the case to keep in mind: the prefix moves and the inflectional ⟨ᵻ⟩ — gold's own spelling, 1,287 entries — stays, in one word. Any rule phrased as "replace ⟨ᵻ⟩" fails exactly there.

Tests, and what they can't catch

Reverted the rule and ran them: 7 target-class cases go red, and the 10 boundary cases pass either way. That's correct and worth saying rather than glossing — boundary cases assert things are unchanged, so they can't detect the rule's absence. Their job is catching over-application, which is the failure that would actually damage a reading.

One expectation had to be weakened: dedans is rare enough to go through the neural OOV path, so its tail is dᵻdˈæns or dᵻdˈænz depending on whether that model resolved — the exact expectation was testing the model's presence, not this rule. Exact-IPA expectations for OOV words are unstable by construction, which applies to any future test in this file.

Suite: 95 + 271 + 23 + 362, 0 failures.

Not fixed by this

  • The phonemizer still emits ⟨ᵻ⟩ in prefix position, so any other consumer of the library inherits the divergence untreated. An argument for eventually fixing it upstream too — recorded so "we fixed it" isn't later read as "it is fixed". vernacula-phonemizer#1445 holds the measurements either way.
  • be- (before, become, begin) shows the same ⟨ᵻ⟩ and is out of scope: nobody has measured gold's spelling for that onset, so it's left alone rather than swept in for looking similar.

🤖 Generated with Claude Code

https://claude.ai/code/session_01KfTCTjC7FTnNm3aS2zszKE

Reported as "determine sounds tensed". It is not a dictionary defect -- our IPA
is already reduced on every path -- but a divergence between two spellings of
the same reduced vowel, and it matters because Kokoro was trained on one of
them. In misaki's us_gold the prefix vowel is ə 1,057 times against ᵻ 9 times,
so our ᵻ is effectively out of distribution for this position.

ESTABLISHED BY SYNTHESIS, NOT BY THE DISTRIBUTION ARGUMENT. An A/B through
Kokoro itself confirmed the model renders the two differently (rms 0.05-0.08,
and determine comes out 50 ms longer with ᵻ against identical leading silence).
A listener then preferred ə on a connected-reading sample and called the
difference "pretty subtle, overall, in this reading" -- which is the size of
claim this change gets to make.

The rule walks groups against the group -> word map. Three conditions:

1. The SOURCE WORD begins de-/re-/pre-. Keyed on the word, not the phonemes.
2. Except ded-. The deduce/deduct family plus two surnames is the only place
   gold attests a prefix ᵻ, so those are the words with the strongest evidence
   in the class. Verified upstream: 0 of the 1,057 target words begin ded-, and
   0 of the keep-words do not.
3. The ᵻ must be the token's FIRST VOWEL. `represent` is ɹˌɛpɹᵻzˈɛnt, whose ᵻ is
   a second syllable and nothing to do with the prefix.

And it does nothing when the map is null or mismatched: a word-keyed rule that
has lost the word must leave the stream alone rather than guess.

WHY WORD-KEYED IS THE WHOLE POINT. Earlier drafts tried to express this
phonologically and needed a `dᵻd` key plus a tie-bar guard, because dᵻd͡ʒ
(degeneracy, deject) matches a bare d on half of /d͡ʒ/. All of that was
manufactured by working at a layer without word identity. PhonemizeTrace is the
full word -> IPA trip and Phonemize already holds its map, so the carve-out is
orthographic: ded- is exact, and deg-/dej- are excluded by spelling with no
affricate hazard at all.

dejected -> dəʤˈɛktᵻd is the case to keep in mind: the prefix moves and the
inflectional ᵻ -- gold's own spelling, 1,287 entries -- stays, in one word. Any
rule phrased as "replace ᵻ" fails exactly there.

Tests were reverted and observed to fail: 7 target-class cases go red, and the
10 boundary cases pass either way. That is correct -- boundary cases assert
things are UNCHANGED, so they cannot detect the rule's absence; they catch
over-application, which is the failure that would damage a reading.

One expectation had to be weakened: `dedans` is rare enough to go through the
neural OOV path, so its tail depends on whether that model resolved. It now
asserts only the prefix vowel. Exact-IPA expectations for OOV words are unstable
by construction.

NOT FIXED BY THIS: the phonemizer still emits ᵻ in prefix position, so any other
consumer inherits the divergence untreated -- an argument for eventually fixing
it upstream too, recorded so "we fixed it" is not read as "it is fixed".
be- (before, become, begin) shows the same ᵻ and is out of scope, since nobody
has measured gold's spelling for that onset.

Suite: 95 + 271 + 23 + 362, 0 failures.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KfTCTjC7FTnNm3aS2zszKE
… map index

Four findings.

THE RULE DISAGREED WITH RENDER ABOUT WHAT ENGLISH IS. KokoroFormat.Render's
English arm is `lang is "en" or "en-GB" or "en-US" or null`; my gate omitted
null, so a call with a null language would render English phonemes and then skip
the English rule on the same call. Aligned to Render's own definition rather
than to a list written from memory.

map[group++] WAS UNBOUNDED. The count check above it should make an overrun
impossible, but this runs on every synthesised paragraph and an index bug would
take the utterance DOWN rather than degrade it. Out of range now means what it
means everywhere else in this rule: cannot identify the word, leave the token
alone.

KokoroVowels.ToCharArray() allocated on every prefix word. Static array now.

THE TESTS DID NOT CHECK THE OUTPUT WAS STILL IN VOCABULARY. audio.cpp refuses a
supplied stream carrying a symbol Kokoro has no id for -- it does not drop it,
it rejects the request -- so a rule emitting one would take out synthesis on
that backend while the ONNX path silently skipped it. Asserted per case now.

Confirmed rather than assumed that both engines get this: KokoroTts (ONNX) calls
Phonemize(text, british) and Lang(british) returns "en"/"en-GB". Had it applied
to only one, the same document would read differently per engine -- the failure
KokoroChunker exists to prevent.

AND THE CONTROL WAS INVALID, FOR THE SECOND TIME THIS WEEK. The first attempt
stashed only the uncommitted review fixes while the rule itself was already
committed, so the control ran against the rule and reported 17 green. I wrote
this trap up on #240 and walked into it again. Redone against main: 7 failed, 10
passed. Stash is the wrong instrument whenever any part of the change is
committed; reverting a named file to main does not depend on remembering which
half is in the index.

Suite: 95 + 271 + 23 + 362, 0 failures.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KfTCTjC7FTnNm3aS2zszKE
@christopherthompson81

Copy link
Copy Markdown
Owner Author

Review pass — four findings, all fixed in 7852e44

1. The rule disagreed with Render about what English is. KokoroFormat.Render's English arm is lang is "en" or "en-GB" or "en-US" or null; my gate omitted null. A call with a null language would render English phonemes and then skip the English rule on the same call. Now aligned to Render's own definition rather than to a list I wrote from memory.

2. map[group++] was unbounded. The count check above it should make an overrun impossible, but this runs on every synthesised paragraph and an index bug would take the utterance down rather than degrade it. Out of range now means what it already means everywhere else in this rule — cannot identify the word, leave the token alone.

3. KokoroVowels.ToCharArray() allocated on every prefix word. Static array.

4. The tests didn't check the output was still in vocabulary. ⚠ audio.cpp refuses a supplied stream carrying a symbol Kokoro has no id for — it doesn't drop it, it rejects the request — so a rule emitting one would take out synthesis on that backend while the ONNX path silently skipped it. Asserted per case now.

Confirmed rather than assumed: both engines get this

KokoroTts (ONNX) calls Phonemize(text, british), and Lang(british) returns "en"/"en-GB". Had the rule reached only one backend, the same document would read differently depending on the engine — the exact failure KokoroChunker exists to prevent.

⚠ And my control was invalid, for the second time this week

First attempt stashed only the uncommitted review fixes while the rule itself was already committed in 1b63a79 — so the "control" ran against the rule and reported 17 green, making the tests look vacuous when they weren't.

I wrote this exact trap up on #240 and then walked into it again. Redone against main:

7 failed, 10 passed

The #240 lesson was "check what the control actually reverted". The stronger version: stash is the wrong instrument whenever any part of the change is committed. Reverting a named file to main doesn't depend on remembering which half is in the index.

Verification

Suite: 95 + 271 + 23 + 362, 0 failures. 7 target-class tests red without the rule; the 10 boundary tests pass either way, as they should — they guard against over-application, not absence.

@christopherthompson81
christopherthompson81 merged commit cb01409 into main Sep 23, 2026
1 check passed
@christopherthompson81
christopherthompson81 deleted the fix/kokoro-prefix-schwa branch September 23, 2026 02:31
christopherthompson81 added a commit that referenced this pull request Sep 23, 2026
#244)

#243 applied the rule to de-, re- AND pre-. Counted per onset, over exactly the
words where our ᵻ is itself the first vowel:

  be-    ə 160
  de-    ə 468   i 30   ɪ 8   ᵻ 7   A 2   ɛ 1
  re-    ə 557   i 12   ɪ 2   ɛ 4   A 1   ʌ 1
  pre-   ə  32   i 46   ɪ 1

pre- RUNS THE OTHER WAY. Gold's majority there is tense i, so the lexicon
argument this rule rests on does not hold for that onset. Verified against our
own pipeline rather than taken on the counts: precede, precise, predict,
preclude, prevent, predominant, precipitate and prefer all moved ᵻ -> ə -- eight
words from one non-gold spelling to a different one.

AND IT SPLIT A PARADIGM, which the counts alone would not have shown:

  prefer      pɹᵻfˈɜɹ  -> pɹəfˈɜɹ
  preferred   pɹifˈɜɹd    pɹifˈɜɹd   (untouched: first vowel already tense i)

The listener sample never covered pre- -- determine, describe, reduce are all
de-/re-. pre- was carried along because it looks like the same class, which is
the whole error: the onsets were treated as a group because they are spelled
alike, and the evidence was never per-onset until now.

be- STAYS OUT on the opposite evidence: gold is unanimous there (ə 160, no
counterexample), stronger than either onset in the rule. It is excluded because
no listener has heard it, and this change has been driven by synthesis rather
than distribution throughout. Holding it out is the same discipline that should
have kept pre- out -- applied to the onset with the best evidence and skipped
for the one with the worst.

Exclusions are pinned by tests comparing against the render with NO rule
applied, so they keep holding if the underlying dictionary reading changes.
Against main, the five pre- cases fail and the three be- cases pass.

pre- is now untouched, not settled: gold says tense i, while this user's own
earlier ruling was that `preferred` is reduced, so "match gold" and "match the
listener" are different targets there and only a listen breaks it.

Suite: 95 + 279 + 23 + 362, 0 failures.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KfTCTjC7FTnNm3aS2zszKE
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant