fix: only lengthen the pause a mark actually left - #200
Merged
Merged
Conversation
The pause code padded gaps that were already long enough and spliced silence into the middle of speech, which is audible as a stitch. Three causes. Silence was measured against the loudest sample rather than the loudest frame, so a real pause read as not silent and was padded regardless of its length. The gap a mark causes is rendered just before the mark's timing ends, not after it: measured over 33 marks of a Hebrew book, gaps start around 0.45s earlier and finish where the mark does. Searching only forwards found nothing every time. Searching both ways for the longest run nearby then attached the pause to the gap before the mark, lengthening the wrong silence, so the run has to be the one the mark itself sits in. Silence is now spliced inside a gap the model left rather than at the mark, and a mark with no gap at all is left alone, since the model ran through it on purpose. On that book at speed 0.8 the pass now adds 1.1s rather than 7.4s, the longest gap stays at the 0.51s the model produced instead of stretching to 0.72s, and a line of dialogue that the model gave 0.11s still opens up to the 0.25s that was asked for. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #199. The pause pass added there padded gaps that were already long enough, and spliced silence into the middle of speech, which is audible as a stitch. Found by ear on a Hebrew book at
speed=0.8, then traced to three causes.Silence was measured against the peak sample
-40dBbelow the loudest sample is far below anything a pause contains, so real gaps read as "not silent" and were topped up no matter how long they already were. Now measured against the loudest frame, which is what the gap analysis uses.The gap comes before the mark, not after
The model renders the silence a mark causes just before that mark's own timing ends. Measured over the 33 marks of the book:
Searching only forwards found nothing every time and inserted a full pause on top of an adequate one. Searching both ways for the longest run nearby then made it worse in a quieter way: for a one word line of dialogue it lengthened the gap before the word and left the one after it at 0.11s. The run has to be the one the mark itself sits in.
Silence was spliced at the mark, mid-sound
Cutting into a ringing signal to force a gap is the stitch. Silence is now spliced inside a gap the model actually left, and a mark with no gap at all is left alone, on the grounds that the model ran through it on purpose.
Effect
Same book,
speed=0.8,sentence_pause=0.25:The last row is the case #199 set out to fix, and it still works: gaps that are too short open up, gaps that are already fine are left exactly as the model made them.
Regression suite passes.