KittenTTS module. - #1107
KittenTTS module.#1107
Conversation
…issue when processing text over 400 char. Fix an issue when we get unmapped chars.
sthibaul
left a comment
There was a problem hiding this comment.
Thanks for the nice work!
In the future, we will probably want to integrate various onnx-based voices, so code will be useful to share between modules, but your writing seems already quite well structured so that it will be convenient to do.
There are just a few changes that need to happen before we can integrate this.
|
bleh, "ubuntu-latest" doesn't seem to really be "latest", the CI is taking 24.04 rather than 26.04... Let me see that |
Could you rebase? I have bumped to 26.04 so onnxruntime is available. |
|
You need to add |
- The downloaded has been moved to a seperate package, documentation has been added on how to use the downloader. - modules dot conf file can now set the models+voices search path.
sthibaul
left a comment
There was a problem hiding this comment.
we're getting close :)
|
The CI said: This indeed should rather be a |
|
The CI also said: The xml header mentions that one should rather use |
…to the logs - Fixed depr warning with libxml content and use - Fixed potental bug that could happen if the phonemizer failed to load
|
Both of those should be fixed now. |
changed the build target name to be just kitten.
|
The conf file has been updated with comments on the configure options. |
|
Thanks! |
KittenTTS Module.
This is a speech dispatch model for running Kitten TTS. Kitten TTS is a deep learning model which provides high quality natural sounding TTS generation using models ranging from 15M to 80M parameters. Due to its small size it is able to run in real-time on CPU. The goal of this project is to integrate Kitten TTS with speech dispatch while maintaining its near real-time speech generation, with special attention being place on reading of long text's such as ebooks. To achieve this the original python code was rewrote into c and tightly integrated into a speech dispatch model.
There is a lot here so let me give a quick summary of what all is going on here.
kitten_server.c:
This file is for handling the protocol and follows some what closely to the example modules for async servers with speech dispatch handling the audio. By design I made sure very little work is done in any function in this thread. Any long running code should be handed to an async queue and ran on one of the threads in the kitten_worker.c file.
kitten_worker.c:
This is where the handling of long running tasks is done. We create two different threads here one to one to handle passing audio back to the server(since the module_tts_output_server function can block and we want the generation to continue while we are outputting audio to the server). The other thread is dedicated to the generation of audio by the model. We synchronize all this using two GasyncQueue, one for handling incoming speak requests, and the other to handle the outputted audio from our model. we also keep track of how much audio we have generated and played so that long speak commands don't run the cpu unnecessarily hard (For example I have seen that Okular will in some cases send an entire book's text in a single speak command). There is also code here for parsing ssml text using libxml.
kitten_model.c:
The handles everything we need to do to generated output from onnx using our model. The most important functions here are init_voice_style to handle loading voice styles. reload_models_and_voices: to handle changing the voice style. And kitten_speak for generating audio.
kitten_downloader.c:
This handles downloading the model+voice styles if its not already on our computer. If it download it does verify the file against a sha256, but that verification is not strongly enforce and is more of a warning.