Add native Coqui XTTS v2 voice cloning - #520
Conversation
Native Q8_0 synthesis demoGenerated end-to-end by this PR's xtts-v2-audiocpp-demo.mp4Independent Whisper-base.en transcription: “Hello? This is a native voice clone.” |
|
Same as usual if this passes your requirements this also needs to be merged https://huggingface.co/audio-cpp/audio.cpp-gguf/discussions/9 👍 |
Synthetic voice-cloning demo — alternate referenceThis is an explicitly labeled synthetic test using the user-supplied Spoken text: “This is a synthetic voice clone generated locally by audio dot C P P. It is not a real recording of David Attenborough.” xtts-v2-attenborough-demo.mp4 |
|
Hm that output is good but... perhaps quantization the model has lowered the voice cloning ability |
Summary
Model artifacts
Hugging Face PR: https://huggingface.co/audio-cpp/audio.cpp-gguf/discussions/9
xtts-v2-q8_0.gguf: 616,802,944 bytes, SHA-256e07fd16d7365575fc9fe3946a70e9510f73f904983459d8eef060e91e62cfd4axtts-v2-f16.gguf: 976,290,816 bytes, SHA-2561db573bd06510f9fb8513913cc4a20b3b4340d23f918dbaf7dde695dda1ce055Both packages embed the tokenizer, config, CPML license, and this model spec, and both pass the conditioning → GPT → HiFiGAN native probe.
Validation
0.999999999993, RMSE1.66e-70.99999999983, RMSE6.31e-70.9999105, identical argmax token808Hello? This is a native voice clone.for inputHello, this is a native voice clone./varvs/private/varalias failure infun_asr_nano_assets_testThe XTTS model and generated output are governed by the Coqui Public Model License 1.0.0 and are non-commercial.