Skip to content

fix: only lengthen the pause a mark actually left - #200

Merged
thewh1teagle merged 1 commit into
mainfrom
fix/pause-placement
Aug 19, 2026
Merged

thewh1teagle merged 1 commit into
mainfrom
fix/pause-placement

Conversation

@thewh1teagle

Copy link
Copy Markdown
Owner

Follow-up to #199. The pause pass added there padded gaps that were already long enough, and spliced silence into the middle of speech, which is audible as a stitch. Found by ear on a Hebrew book at speed=0.8, then traced to three causes.

Silence was measured against the peak sample

-40dB below the loudest sample is far below anything a pause contains, so real gaps read as "not silent" and were topped up no matter how long they already were. Now measured against the loudest frame, which is what the gap analysis uses.

The gap comes before the mark, not after

The model renders the silence a mark causes just before that mark's own timing ends. Measured over the 33 marks of the book:

mark ends   6.30s | nearest gap starts   5.83s len 0.51s | offset -0.47s
mark ends   8.18s | nearest gap starts   7.70s len 0.51s | offset -0.48s
mark ends  13.70s | nearest gap starts  13.25s len 0.46s | offset -0.45s

Searching only forwards found nothing every time and inserted a full pause on top of an adequate one. Searching both ways for the longest run nearby then made it worse in a quieter way: for a one word line of dialogue it lengthened the gap before the word and left the one after it at 0.11s. The run has to be the one the mark itself sits in.

Silence was spliced at the mark, mid-sound

Cutting into a ringing signal to force a gap is the stitch. Silence is now spliced inside a gap the model actually left, and a mark with no gap at all is left alone, on the grounds that the model ran through it on purpose.

Effect

Same book, speed=0.8, sentence_pause=0.25:

before after
total added +7.4s +1.1s
longest gap 0.72s 0.51s (what the model produced)
gaps over 0.6s 12 0
short dialogue gap 0.11s 0.25s

The last row is the case #199 set out to fix, and it still works: gaps that are too short open up, gaps that are already fine are left exactly as the model made them.

Regression suite passes.

The pause code padded gaps that were already long enough and spliced silence
into the middle of speech, which is audible as a stitch. Three causes.

Silence was measured against the loudest sample rather than the loudest frame,
so a real pause read as not silent and was padded regardless of its length.

The gap a mark causes is rendered just before the mark's timing ends, not
after it: measured over 33 marks of a Hebrew book, gaps start around 0.45s
earlier and finish where the mark does. Searching only forwards found nothing
every time. Searching both ways for the longest run nearby then attached the
pause to the gap before the mark, lengthening the wrong silence, so the run
has to be the one the mark itself sits in.

Silence is now spliced inside a gap the model left rather than at the mark,
and a mark with no gap at all is left alone, since the model ran through it
on purpose.

On that book at speed 0.8 the pass now adds 1.1s rather than 7.4s, the
longest gap stays at the 0.51s the model produced instead of stretching to
0.72s, and a line of dialogue that the model gave 0.11s still opens up to the
0.25s that was asked for.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@thewh1teagle
thewh1teagle merged commit bba097d into main Aug 19, 2026
1 check passed
@thewh1teagle
thewh1teagle deleted the fix/pause-placement branch August 19, 2026 00:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant