Skip to content

fix: chunking, thread safety, GPU detection, streaming; add timestamps and continuous synthesis - #197

Merged
thewh1teagle merged 1 commit into
mainfrom
fix/chunking-timestamps-continuous
Aug 18, 2026
Merged

thewh1teagle merged 1 commit into
mainfrom
fix/chunking-timestamps-continuous

Conversation

@thewh1teagle

Copy link
Copy Markdown
Owner

Rewrites phoneme batching, fixes several long-standing defects found while triaging the open issues, and adds timestamps plus continuous (sliding window) synthesis.

Fixes

Chunker — _split_phonemes only broke on .,!?;, so a punctuation-free clause never split, and it pushed an empty first batch. Three separate user-visible failures, all reproduced against v1.0:

  • IndexError: index 510 is out of bounds on long unpunctuated text
  • the empty batch synthesized ~0.5s of audible garbage (peak 0.59) that trim did not remove
  • everything past phoneme 510 silently dropped

Batches are now balanced (smallest limit that needs no extra pass) rather than greedily filled, so no short remainder is left. A short batch is spoken differently: on one paragraph the old split [475, 66] gave a 2.17 ph/s rate gap, the new [317, 224] gives 0.98, and loudness spread drops from 0.0030 to 0.0001 rms.

Fixes #115, #184, #130, #114

Style vector — used voice[len(tokens)] where the reference implementation (and scripts/export.py) uses pack[len(phonemes) - 1], so every call picked the row for one-token-longer text, and exactly 510 tokens crashed.

Speed — the input_ids path cast speed to int32, so speed=1.5 became 1 and speed=0.9 became 0 and crashed the ONNX Loop kernel. Tensors are now built from the graph's declared input types. Note the v1.1-zh graph really does declare speed as int32; the library now warns and rounds instead of failing, and scripts/export.py no longer produces such graphs.

Fixes #155

Thread safety — espeak-ng keeps process-global state, so concurrent create() returned corrupted phonemes with no error. Measured: the same text gave 279 phonemes serially but 368/377/413/457 across 4 threads; now 24/24 concurrent calls agree. Inference stays outside the lock.

Fixes #191

GPU detection — importlib.util.find_spec("onnxruntime-gpu") tests a distribution name; a hyphen can never be a module, so it always returned None and kokoro-onnx[gpu] silently ran on CPU. Provider selection moved to session.py and looks up installed distributions.

Fixes #125, #92, #167, #159, #56, #117

Streaming — create_stream fired its worker with a bare asyncio.create_task: on any exception the sentinel was never queued and the consumer blocked forever, and the task was unreferenced and uncancellable. Failures now surface in the caller, abandoning the generator stops the producer, and the queue is bounded so the producer cannot race ahead of playback.

Fixes #168

Clear errors — empty or all-dropped input raised ValueError: need at least one array to concatenate; it now says what was wrong.

Fixes #170

Features

  • create_timed() returns (samples, sample_rate, timings) with one entry per phoneme. Fixes return timestamps for words #116, Create an SRT subtitle file along with the text output #104, Does this support transcription / word-timestamps? #165
  • create(continuous=True) synthesizes overlapping [context][keep][lookahead] windows and keeps only the middles, so neither the cold utterance onset nor the sentence-final fall lands in the output. ~1.3x the compute. On a paragraph reconstructed from 3 pieces it tracks a single-pass reference far better than hard splitting (pitch contour r 0.42 vs 0.32, energy 0.53 vs 0.21, duration drift -0.35s vs -1.28s)
  • punctuation-aware pauses between batches, defaults matched to the gaps the model produces inside a batch (sentence 0.23-0.31s, clause 0.06-0.21s). Fixes Add a bit of silence between batches of phonemes #84
  • scripts/export.py rewritten: emits waveform + duration, keeps speed a float input, embeds config.json in the graph metadata so the vocabulary travels with the model, and verifies the export against the torch model with a real voice (correlation 0.9969)

Both features need a model with a duration output: kokoro-v1.1-zh.onnx has one, and scripts/export.py produces them. Models without it work exactly as before, has_timings is False and continuous=True raises with an explanation.

Also

Dropped phonemizer-fork for upstream phonemizer>=3.4.0, which now carries both patches the fork existed for (set_data_path and backend caching, the latter as a bounded LRU covering every backend). All dependencies upgraded: onnxruntime 1.29, numpy 2.5.2, protobuf 7.35.1, ruff 0.16.3. Closes #10, #24. The memory leak in #148 is measurably gone (20 successive calls held at 1665 MB with zero growth), and #175 and #111 are resolved by the upstream phonemizer.

New examples: with_timestamps.py, with_continuous.py.

Verification

  • property tests over the chunker (bounded, non-empty, lossless) including 1200-char unbroken runs and the exact 510 boundary
  • before/after reproduction of every defect above against the released v1.0 model
  • save.py, english.py, with_blending.py, with_voice.py, with_phonemes.py, with_stream_save.py and the Hebrew custom-vocab path all still run

Behaviour change worth a minor version: audio is not bit-identical to 0.5.0. The style-row fix shifts prosody slightly and joins now carry pauses.

Rewrite phoneme batching, make concurrent use safe, and add two features
that the duration output of newly exported models makes possible.

Fixes
- chunker: split only broke on .,!?; so punctuation-free text never split,
  crashing with IndexError past 510 phonemes, and it emitted an empty first
  batch that synthesized ~0.5s of garbage. Batches are now balanced so no
  short runt is left, which also keeps speaking rate and loudness even.
- style vector: used row len(tokens) instead of len(tokens) - 1
- speed: was cast to int32 on the input_ids path, so speed=0.9 became 0 and
  crashed the model. Tensors are now built from the graph's own signature.
- espeak: phonemize is process-global state, serialize it. Concurrent calls
  returned corrupted phonemes with no error.
- providers: find_spec("onnxruntime-gpu") tests a distribution name and can
  never match a module, so kokoro-onnx[gpu] silently stayed on CPU.
- create_stream: a failing worker never queued the sentinel and the consumer
  blocked forever; the task was also unreferenced and uncancellable.
- empty input raised an opaque numpy error instead of saying what was wrong.

Features
- create_timed() reports when each phoneme is spoken
- create(continuous=True) synthesizes overlapping windows and keeps only
  their middles, so prosody runs across the joins
- punctuation-aware pauses between batches, matched to what the model
  produces inside a batch
- scripts/export.py embeds config.json in the graph, so the vocabulary
  travels with the model instead of having to match the packaged copy

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@thewh1teagle
thewh1teagle merged commit 013c2fc into main Aug 18, 2026
1 check passed
@thewh1teagle
thewh1teagle deleted the fix/chunking-timestamps-continuous branch August 18, 2026 23:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment