feat: fp16 and int8 export variants, point downloads at model-files-v1.1 - #198
Merged
Merged
Conversation
scripts/export.py takes --fp16 and --int8 and writes reduced precision copies beside the full precision export. Both keep float inputs and outputs so every variant is called the same way. fp16 goes through onnxruntime's converter: onnxconverter_common leaves this graph with mismatched Cast types around the Loop subgraph and the result fails to load. Quantized copies predict slightly different durations, so only the full precision export is verified for identical output. All download URLs in the README, the examples and the test workflow now point at model-files-v1.1, which carries both models in all three precisions, built with duration outputs and an embedded vocabulary. with_quant.py was still pointing at kokoro-v0_19 from the original model-files release. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #197. Adds quantized export variants and repoints every download URL at the
model-files-v1.1release, which now carries both models in all three precisions.Export script
--fp16and--int8write reduced precision copies beside the full precision export. Both keep float inputs and outputs, so all three variants are called identically.fp16 goes through onnxruntime's converter rather than
onnxconverter_common: the latter leaves this graph with mismatched Cast types around the Loop subgraph, and the result fails to load withType (tensor(float16)) of output arg (/Cast_7_output_0) does not match expected type (tensor(float)).Quantized copies predict slightly different durations, so
verifycompares over the shared length and only holds the full precision export to identical output.Release assets
Both models re-exported with the duration output, a float
speedinput andconfig.jsonembedded in the graph metadata, then quantized:voices-v1.0.binis mirrored into the same release so every example downloads from one place.All six load through
Kokoro, report timings, and pick their vocabulary out of the graph (114 entries for v1.0, 171 for v1.1-zh) with novocab_configargument.URLs
57 references across
README.md, all examples and.github/workflows/test.ymlnow point atmodel-files-v1.1.with_quant.pywas still advertisingkokoro-v0_19.int8.onnxandkokoro-v0_19.fp16.onnxfrom the originalmodel-filesrelease, and now points at the current v1.0 variants with corrected sizes.Note:
kokoro-v1.1-zh.onnxin that release is replaced rather than added. Same weights, but the new export takes a floatspeed(the published one declares int32, which is why fractional speed was broken, #155) and carries its vocabulary. Anyone pinned to that URL gets the fixed graph.