fix: place pauses at every sentence, repair modern weight norm checkpoints - #199
Merged
Merged
Conversation
…oints Three problems found while synthesizing a Hebrew book. Newlines are not in the vocabulary and were dropped without replacement, so the lines of a text ran together with no gap at all: mimul?""ken.""at. The model then had nothing to pause on. Whitespace is now collapsed to single spaces. This only affected continuous synthesis, the chunker already split on whitespace and rejoined with spaces. Pauses were only added where two batches met, so a sentence inside a batch got whatever the model produced on its own, around 0.1s, and continuous synthesis got none at all. The new pauses module uses the timings to top up the gap after every mark wherever it falls: it looks for the pause the model already made in a short window past the mark, since a mark is often timed just before the gap it causes, and adds only the difference. Models without a duration output keep the previous batch join behaviour. KModel registers weight norm as weight_g/weight_v, but checkpoints trained on recent torch store parametrizations.weight.original0/1. KModel loads those with strict=False, so 140 of 375 decoder weights stayed randomly initialized and the exported model emitted static with no error anywhere. export.py now rewrites the keys first, and is a no-op for checkpoints that do not need it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This was referenced Aug 19, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Three bugs found while synthesizing a Hebrew book with
continuous=True, all of which failed silently.Newlines were deleted, not turned into gaps
\nis not in the vocabulary, so it was dropped with nothing in its place and consecutive lines fused:With no word boundary there was nothing for the model to pause on, so a line of dialogue ran straight into the next. Whitespace is now collapsed to single spaces in
_prepare.Only continuous synthesis was affected: the chunker splits on
\s+and rejoins with spaces, so batched mode never saw it.Pauses existed only at batch joins
sentence_pauseandclause_pausewere applied where two batches met and nowhere else, so a sentence inside a batch got only what the model produces on its own, about 0.1s, and continuous synthesis got nothing at all.The new
pausesmodule places them with the timings instead, after every mark wherever it falls. It tops up rather than inserts blindly: a mark is often timed slightly before the gap it causes, so it looks for the pause already there in a short window past the mark and adds only the difference.Measured on the Hebrew book, 33 sentence marks over 79s:
sentence_pause=0.45Models with no duration output keep the previous batch join behaviour, since there are no timings to place anything with.
Checkpoints trained on modern torch exported to static
KModelregisters weight norm asweight_g/weight_v, while checkpoints saved by recent torch storeparametrizations.weight.original0/1:KModelcatches the failure and retries withstrict=False, so the other 140 decoder tensors stayed randomly initialized and the exported model produced static, with no error at any point.verifywould not have caught it either: it compares ONNX against torch, and both were noise.export.pynow rewrites those keys before loading, and is a no-op for checkpoints that do not need it. Confirmed by re-exporting the Hebrew checkpoint: spectral flatness 0.712 (noise) before, 0.109 after, and the result is indistinguishable from a known good ONNX of the same model (15.10s both, rms 0.1318 vs 0.1317).Verification
Full suite passes,
save.pyandwith_stream_save.pystill run.