Skip to content

fix: place pauses at every sentence, repair modern weight norm checkpoints - #199

Merged
thewh1teagle merged 1 commit into
mainfrom
fix/pauses-and-weightnorm
Aug 19, 2026
Merged

thewh1teagle merged 1 commit into
mainfrom
fix/pauses-and-weightnorm

Conversation

@thewh1teagle

Copy link
Copy Markdown
Owner

Three bugs found while synthesizing a Hebrew book with continuous=True, all of which failed silently.

Newlines were deleted, not turned into gaps

\n is not in the vocabulary, so it was dropped with nothing in its place and consecutive lines fused:

mimˈul?""kˈen.""ʔˈat        what the model received
mimˈul?" "kˈen." "ʔˈat      after the fix

With no word boundary there was nothing for the model to pause on, so a line of dialogue ran straight into the next. Whitespace is now collapsed to single spaces in _prepare.

Only continuous synthesis was affected: the chunker splits on \s+ and rejoins with spaces, so batched mode never saw it.

Pauses existed only at batch joins

sentence_pause and clause_pause were applied where two batches met and nowhere else, so a sentence inside a batch got only what the model produces on its own, about 0.1s, and continuous synthesis got nothing at all.

The new pauses module places them with the timings instead, after every mark wherever it falls. It tops up rather than inserts blindly: a mark is often timed slightly before the gap it causes, so it looks for the pause already there in a short window past the mark and adds only the difference.

Measured on the Hebrew book, 33 sentence marks over 79s:

before after
gap after a line of dialogue 0.12s 0.31s
median sentence pause n/a 0.48s
with sentence_pause=0.45 0.51-0.82s

Models with no duration output keep the previous batch join behaviour, since there are no timings to place anything with.

Checkpoints trained on modern torch exported to static

KModel registers weight norm as weight_g/weight_v, while checkpoints saved by recent torch store parametrizations.weight.original0/1:

model expects:   decode.0.conv1.weight_g, decode.0.conv1.weight_v
v1.0 ckpt has:   module.decode.0.conv1.weight_g, ...weight_v                    375/375 load
modern ckpt has: module.decode.0.conv1.parametrizations.weight.original0/1      235/375 load

KModel catches the failure and retries with strict=False, so the other 140 decoder tensors stayed randomly initialized and the exported model produced static, with no error at any point. verify would not have caught it either: it compares ONNX against torch, and both were noise.

export.py now rewrites those keys before loading, and is a no-op for checkpoints that do not need it. Confirmed by re-exporting the Hebrew checkpoint: spectral flatness 0.712 (noise) before, 0.109 after, and the result is indistinguishable from a known good ONNX of the same model (15.10s both, rms 0.1318 vs 0.1317).

Verification

Full suite passes, save.py and with_stream_save.py still run.

…oints

Three problems found while synthesizing a Hebrew book.

Newlines are not in the vocabulary and were dropped without replacement, so
the lines of a text ran together with no gap at all: mimul?""ken.""at. The
model then had nothing to pause on. Whitespace is now collapsed to single
spaces. This only affected continuous synthesis, the chunker already split on
whitespace and rejoined with spaces.

Pauses were only added where two batches met, so a sentence inside a batch
got whatever the model produced on its own, around 0.1s, and continuous
synthesis got none at all. The new pauses module uses the timings to top up
the gap after every mark wherever it falls: it looks for the pause the model
already made in a short window past the mark, since a mark is often timed
just before the gap it causes, and adds only the difference. Models without a
duration output keep the previous batch join behaviour.

KModel registers weight norm as weight_g/weight_v, but checkpoints trained on
recent torch store parametrizations.weight.original0/1. KModel loads those
with strict=False, so 140 of 375 decoder weights stayed randomly initialized
and the exported model emitted static with no error anywhere. export.py now
rewrites the keys first, and is a no-op for checkpoints that do not need it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@thewh1teagle
thewh1teagle merged commit d7f7b12 into main Aug 19, 2026
1 check passed
@thewh1teagle
thewh1teagle deleted the fix/pauses-and-weightnorm branch August 19, 2026 00:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant