Replies: 1 comment
|
For zero-shot cloning, I would not feed the whole hour. The current inference code trims the reference to at most 12 seconds and resamples it to 24 kHz, so start with a clean 6–12 second clip with one speaker, little room noise, and an exact transcript. Try several reference clips; reference choice often matters more than adding more audio. For fine-tuning, keep clips between 1 and 30 seconds. The bundled finetuning UI loads audio as mono 24 kHz, slices on silence, and rejects clips outside that duration range. I would prepare the data in three passes:
Do not strip every breath or non-word sound mechanically. Keep natural breaths that belong to the delivery; remove unrelated noise and long dead air. I would compare three outputs on the same held-out sentences: best-reference zero-shot, clean-subset fine-tune, and all-data fine-tune. That will tell you whether training is helping rather than just making the voice more average. Also make sure you have permission to clone the speaker from these messages. |
Uh oh!
There was an error while loading. Please reload this page.
dear abby,
i'm new to this area but i setup an F5-TTS environment and ran my first model after trying elevenlabs instant voice clone with limited success.
without getting into the use case details the situation is this; i have ~150 voice messages (about an hour) of varying quality, duration, etc. from a single individual over 5 years. my challenge is that i'm not sure what i should be expecting in terms of results. i feel like it should be hard to tell it's a clone if you didn't know it's a clone. at least for a short bit of generated output, like 10 - 20 seconds.
thank you in advance,
regards,
newbie by the sea.
All reactions