Skip to content

feat: fp16 and int8 export variants, point downloads at model-files-v1.1 - #198

Merged
thewh1teagle merged 1 commit into
mainfrom
feat/release-model-variants
Aug 18, 2026
Merged

thewh1teagle merged 1 commit into
mainfrom
feat/release-model-variants

Conversation

@thewh1teagle

Copy link
Copy Markdown
Owner

Follow-up to #197. Adds quantized export variants and repoints every download URL at the model-files-v1.1 release, which now carries both models in all three precisions.

Export script

--fp16 and --int8 write reduced precision copies beside the full precision export. Both keep float inputs and outputs, so all three variants are called identically.

fp16 goes through onnxruntime's converter rather than onnxconverter_common: the latter leaves this graph with mismatched Cast types around the Loop subgraph, and the result fails to load with Type (tensor(float16)) of output arg (/Cast_7_output_0) does not match expected type (tensor(float)).

Quantized copies predict slightly different durations, so verify compares over the shared length and only holds the full precision export to identical output.

Release assets

Both models re-exported with the duration output, a float speed input and config.json embedded in the graph metadata, then quantized:

file size vs its own fp32
kokoro-v1.0.onnx 326 MB correlation 0.9945 vs torch
kokoro-v1.0.fp16.onnx 164 MB spectral 0.999, ~4x faster on CPU here
kokoro-v1.0.int8.onnx 114 MB spectral 0.916
kokoro-v1.1-zh.onnx 326 MB correlation 0.9952 vs torch
kokoro-v1.1-zh.fp16.onnx 164 MB spectral 0.994
kokoro-v1.1-zh.int8.onnx 114 MB spectral 0.874

voices-v1.0.bin is mirrored into the same release so every example downloads from one place.

All six load through Kokoro, report timings, and pick their vocabulary out of the graph (114 entries for v1.0, 171 for v1.1-zh) with no vocab_config argument.

URLs

57 references across README.md, all examples and .github/workflows/test.yml now point at model-files-v1.1. with_quant.py was still advertising kokoro-v0_19.int8.onnx and kokoro-v0_19.fp16.onnx from the original model-files release, and now points at the current v1.0 variants with corrected sizes.

Note: kokoro-v1.1-zh.onnx in that release is replaced rather than added. Same weights, but the new export takes a float speed (the published one declares int32, which is why fractional speed was broken, #155) and carries its vocabulary. Anyone pinned to that URL gets the fixed graph.

scripts/export.py takes --fp16 and --int8 and writes reduced precision copies
beside the full precision export. Both keep float inputs and outputs so every
variant is called the same way.

fp16 goes through onnxruntime's converter: onnxconverter_common leaves this
graph with mismatched Cast types around the Loop subgraph and the result
fails to load. Quantized copies predict slightly different durations, so only
the full precision export is verified for identical output.

All download URLs in the README, the examples and the test workflow now point
at model-files-v1.1, which carries both models in all three precisions, built
with duration outputs and an embedded vocabulary. with_quant.py was still
pointing at kokoro-v0_19 from the original model-files release.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@thewh1teagle
thewh1teagle merged commit eaca6a3 into main Aug 18, 2026
1 check passed
@thewh1teagle
thewh1teagle deleted the feat/release-model-variants branch August 18, 2026 23:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant