Add native Coqui SpeedySpeech TTS - #519
Conversation
Native audio.cpp MP4 samplesBoth files below were generated by Default rate (4.34 s) Speaking rate 1.2 (5.89 s) The same files are included in the PR under |
38beeab to
7f50291
Compare
coqui-speedy-speech-demo-fast.mp4coqui-speedy-speech-demo.mp4 |
|
Hm this needs work the audio is too choppy |
Quality/parity updateThe original samples exposed two frontend/runtime parity issues, now fixed in commit
Replacement native audio.cpp MP4s (H.264/AAC): Default rate — 4.49 s 1.2x rate — 3.76 s Both say: “Today is a beautiful day to create natural speech on your computer.” Validation:
Updated model artifact: audio-cpp/audio.cpp-gguf PR #8. |
|
@DrewThomasson Thanks for the PR! I will review it this weekend. |
|
Don't forget to merge the gguf model addition on hugginface if it passes your tests thx |
|
Still crunchy I need to cross check with how it sounds in coqui tts |
|
@DrewThomasson Let’s just keep demos out of the repo in PRs, and maybe put them in the relevant HF model directory instead. I think the quality of Coqui XTTS v2 is fine, given that it’s pretty old. Bark needs more work and Coqui sounds robotic. Please also test longform (test case in the path test) and see if there are any VRAM management issues. The framerwork has text chunkers for chunking long text. |
|
Also Bark and Coqui XTTS v2 should go to |
Summary
Model weights are proposed separately in audio-cpp/audio.cpp-gguf PR #8. The package manifest intentionally tracks
refs/pr/8until that model PR is merged.Validation
python3 tools/check_loader_catalog_sync.pygit diff --checkSource checkpoints
tts_models/en/ljspeech/speedy-speechvocoder_models/en/ljspeech/hifigan_v2