fix: chunking, thread safety, GPU detection, streaming; add timestamps and continuous synthesis - #197
Merged
Conversation
Rewrite phoneme batching, make concurrent use safe, and add two features
that the duration output of newly exported models makes possible.
Fixes
- chunker: split only broke on .,!?; so punctuation-free text never split,
crashing with IndexError past 510 phonemes, and it emitted an empty first
batch that synthesized ~0.5s of garbage. Batches are now balanced so no
short runt is left, which also keeps speaking rate and loudness even.
- style vector: used row len(tokens) instead of len(tokens) - 1
- speed: was cast to int32 on the input_ids path, so speed=0.9 became 0 and
crashed the model. Tensors are now built from the graph's own signature.
- espeak: phonemize is process-global state, serialize it. Concurrent calls
returned corrupted phonemes with no error.
- providers: find_spec("onnxruntime-gpu") tests a distribution name and can
never match a module, so kokoro-onnx[gpu] silently stayed on CPU.
- create_stream: a failing worker never queued the sentinel and the consumer
blocked forever; the task was also unreferenced and uncancellable.
- empty input raised an opaque numpy error instead of saying what was wrong.
Features
- create_timed() reports when each phoneme is spoken
- create(continuous=True) synthesizes overlapping windows and keeps only
their middles, so prosody runs across the joins
- punctuation-aware pauses between batches, matched to what the model
produces inside a batch
- scripts/export.py embeds config.json in the graph, so the vocabulary
travels with the model instead of having to match the packaged copy
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This was referenced Aug 18, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Rewrites phoneme batching, fixes several long-standing defects found while triaging the open issues, and adds timestamps plus continuous (sliding window) synthesis.
Fixes
Chunker —
_split_phonemesonly broke on.,!?;, so a punctuation-free clause never split, and it pushed an empty first batch. Three separate user-visible failures, all reproduced against v1.0:IndexError: index 510 is out of boundson long unpunctuated texttrimdid not removeBatches are now balanced (smallest limit that needs no extra pass) rather than greedily filled, so no short remainder is left. A short batch is spoken differently: on one paragraph the old split
[475, 66]gave a 2.17 ph/s rate gap, the new[317, 224]gives 0.98, and loudness spread drops from 0.0030 to 0.0001 rms.Fixes #115, #184, #130, #114
Style vector — used
voice[len(tokens)]where the reference implementation (andscripts/export.py) usespack[len(phonemes) - 1], so every call picked the row for one-token-longer text, and exactly 510 tokens crashed.Speed — the
input_idspath cast speed toint32, sospeed=1.5became 1 andspeed=0.9became 0 and crashed the ONNXLoopkernel. Tensors are now built from the graph's declared input types. Note the v1.1-zh graph really does declarespeedas int32; the library now warns and rounds instead of failing, andscripts/export.pyno longer produces such graphs.Fixes #155
Thread safety — espeak-ng keeps process-global state, so concurrent
create()returned corrupted phonemes with no error. Measured: the same text gave 279 phonemes serially but 368/377/413/457 across 4 threads; now 24/24 concurrent calls agree. Inference stays outside the lock.Fixes #191
GPU detection —
importlib.util.find_spec("onnxruntime-gpu")tests a distribution name; a hyphen can never be a module, so it always returnedNoneandkokoro-onnx[gpu]silently ran on CPU. Provider selection moved tosession.pyand looks up installed distributions.Fixes #125, #92, #167, #159, #56, #117
Streaming —
create_streamfired its worker with a bareasyncio.create_task: on any exception the sentinel was never queued and the consumer blocked forever, and the task was unreferenced and uncancellable. Failures now surface in the caller, abandoning the generator stops the producer, and the queue is bounded so the producer cannot race ahead of playback.Fixes #168
Clear errors — empty or all-dropped input raised
ValueError: need at least one array to concatenate; it now says what was wrong.Fixes #170
Features
create_timed()returns(samples, sample_rate, timings)with one entry per phoneme. Fixes return timestamps for words #116, Create an SRT subtitle file along with the text output #104, Does this support transcription / word-timestamps? #165create(continuous=True)synthesizes overlapping[context][keep][lookahead]windows and keeps only the middles, so neither the cold utterance onset nor the sentence-final fall lands in the output. ~1.3x the compute. On a paragraph reconstructed from 3 pieces it tracks a single-pass reference far better than hard splitting (pitch contour r 0.42 vs 0.32, energy 0.53 vs 0.21, duration drift -0.35s vs -1.28s)scripts/export.pyrewritten: emitswaveform+duration, keepsspeeda float input, embedsconfig.jsonin the graph metadata so the vocabulary travels with the model, and verifies the export against the torch model with a real voice (correlation 0.9969)Both features need a model with a duration output:
kokoro-v1.1-zh.onnxhas one, andscripts/export.pyproduces them. Models without it work exactly as before,has_timingsis False andcontinuous=Trueraises with an explanation.Also
Dropped
phonemizer-forkfor upstreamphonemizer>=3.4.0, which now carries both patches the fork existed for (set_data_pathand backend caching, the latter as a bounded LRU covering every backend). All dependencies upgraded: onnxruntime 1.29, numpy 2.5.2, protobuf 7.35.1, ruff 0.16.3. Closes #10, #24. The memory leak in #148 is measurably gone (20 successive calls held at 1665 MB with zero growth), and #175 and #111 are resolved by the upstream phonemizer.New examples:
with_timestamps.py,with_continuous.py.Verification
save.py,english.py,with_blending.py,with_voice.py,with_phonemes.py,with_stream_save.pyand the Hebrew custom-vocab path all still runBehaviour change worth a minor version: audio is not bit-identical to 0.5.0. The style-row fix shifts prosody slightly and joins now carry pauses.