Skip to content

KittenTTS module. - #1107

Merged
sthibaul merged 25 commits into
brailcom:masterfrom
jsett:kitten
Sep 9, 2026
Merged

sthibaul merged 25 commits into
brailcom:masterfrom
jsett:kitten

Conversation

@jsett

@jsett jsett commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

KittenTTS Module.

This is a speech dispatch model for running Kitten TTS. Kitten TTS is a deep learning model which provides high quality natural sounding TTS generation using models ranging from 15M to 80M parameters. Due to its small size it is able to run in real-time on CPU. The goal of this project is to integrate Kitten TTS with speech dispatch while maintaining its near real-time speech generation, with special attention being place on reading of long text's such as ebooks. To achieve this the original python code was rewrote into c and tightly integrated into a speech dispatch model.

There is a lot here so let me give a quick summary of what all is going on here.

kitten_server.c:

This file is for handling the protocol and follows some what closely to the example modules for async servers with speech dispatch handling the audio. By design I made sure very little work is done in any function in this thread. Any long running code should be handed to an async queue and ran on one of the threads in the kitten_worker.c file.

kitten_worker.c:

This is where the handling of long running tasks is done. We create two different threads here one to one to handle passing audio back to the server(since the module_tts_output_server function can block and we want the generation to continue while we are outputting audio to the server). The other thread is dedicated to the generation of audio by the model. We synchronize all this using two GasyncQueue, one for handling incoming speak requests, and the other to handle the outputted audio from our model. we also keep track of how much audio we have generated and played so that long speak commands don't run the cpu unnecessarily hard (For example I have seen that Okular will in some cases send an entire book's text in a single speak command). There is also code here for parsing ssml text using libxml.

kitten_model.c:

The handles everything we need to do to generated output from onnx using our model. The most important functions here are init_voice_style to handle loading voice styles. reload_models_and_voices: to handle changing the voice style. And kitten_speak for generating audio.

kitten_downloader.c:

This handles downloading the model+voice styles if its not already on our computer. If it download it does verify the file against a sha256, but that verification is not strongly enforce and is more of a warning.

@sthibaul sthibaul left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the nice work!

In the future, we will probably want to integrate various onnx-based voices, so code will be useful to share between modules, but your writing seems already quite well structured so that it will be convenient to do.

There are just a few changes that need to happen before we can integrate this.

Comment thread configure.ac Outdated
Comment thread src/modules/kitten_server.c Outdated
Comment thread src/modules/kitten_model.c
Comment thread src/modules/kitten_worker.c
Comment thread src/modules/kitten_worker.c Outdated
Comment thread src/modules/kitten.h
Comment thread src/modules/kitten.h
Comment thread src/modules/kitten_downloader.c Outdated
Comment thread src/modules/kitten_server.c
Comment thread src/modules/kitten_downloader.c Outdated
@jsett
jsett requested a review from sthibaul August 26, 2026 19:47
@sthibaul

Copy link
Copy Markdown
Collaborator

bleh, "ubuntu-latest" doesn't seem to really be "latest", the CI is taking 24.04 rather than 26.04... Let me see that

@sthibaul

Copy link
Copy Markdown
Collaborator

bleh, "ubuntu-latest" doesn't seem to really be "latest", the CI is taking 24.04 rather than 26.04... Let me see that

Could you rebase? I have bumped to 26.04 so onnxruntime is available.

@sthibaul

Copy link
Copy Markdown
Collaborator

You need to add kitten.h to EXTRA_DIST in Makefile.am so it gets shipped in the distributed tarball.

jsett added 2 commits August 31, 2026 20:40
- The downloaded has been moved to a seperate package, documentation has been added on how to use the downloader.
- modules dot conf file can now set the models+voices search path.
Comment thread configure.ac Outdated

@sthibaul sthibaul left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we're getting close :)

Comment thread src/modules/README.kitten.md Outdated
Comment thread src/modules/README.kitten.md Outdated
Comment thread src/modules/README.kitten.md Outdated
Comment thread src/modules/README.kitten.md Outdated
Comment thread .github/workflows/ci.yml Outdated
Comment thread src/modules/kitten_server.c
@sthibaul

sthibaul commented Sep 1, 2026

Copy link
Copy Markdown
Collaborator

The CI said:

kitten_worker.c:416:57: warning: format ‘%d’ expects argument of type ‘int’, but argument 3 has type ‘size_t’ {aka ‘long unsigned int’} [-Wformat=]

This indeed should rather be a %zd, to properly print a size_t

@sthibaul

sthibaul commented Sep 1, 2026

Copy link
Copy Markdown
Collaborator

The CI also said:

[8](https://github.com/brailcom/speechd/actions/runs/33558231873/job/100024779175?pr=1107#step:9:149)
kitten_worker.c: In function ‘parse_ssml_to_gqueue’:
kitten_worker.c:274:13: warning: ‘content’ is deprecated [-Wdeprecated-declarations]
  274 |             pl->text = g_string_new(buffer->content);
      |             ^~
In file included from /usr/include/libxml2/libxml/parser.h:20,
                 from kitten.h:28,
                 from kitten_worker.c:13:
/usr/include/libxml2/libxml/tree.h:115:14: note: declared here
  115 |     xmlChar *content XML_DEPRECATED_MEMBER;
      |              ^~~~~~~
kitten_worker.c:285:5: warning: ‘use’ is deprecated [-Wdeprecated-declarations]
  285 |     if (buffer->use > 0) {
      |     ^~
/usr/include/libxml2/libxml/tree.h:121:18: note: declared here
  121 |     unsigned int use XML_DEPRECATED_MEMBER;
      |                  ^~~
kitten_worker.c:289:9: warning: ‘content’ is deprecated [-Wdeprecated-declarations]
  289 |         pl->text = g_string_new(buffer->content);
      |         ^~
/usr/include/libxml2/libxml/tree.h:115:14: note: declared here
  115 |     xmlChar *content XML_DEPRECATED_MEMBER;
      |              ^~~~~~~

The xml header mentions that one should rather use xmlBufferContent and xmlBufferLength.

…to the logs

- Fixed depr warning with libxml content and use
- Fixed potental bug that could happen if the phonemizer failed to load
@jsett

jsett commented Sep 4, 2026

Copy link
Copy Markdown
Contributor Author

Both of those should be fixed now.

Comment thread src/modules/README.kitten.md Outdated
Comment thread src/modules/Makefile.am Outdated
Comment thread config/modules/kittentts.conf Outdated
changed the build target name to be just kitten.
@jsett

jsett commented Sep 7, 2026

Copy link
Copy Markdown
Contributor Author

The conf file has been updated with comments on the configure options.

@sthibaul
sthibaul merged commit 5f26f06 into brailcom:master Sep 9, 2026
7 checks passed
@sthibaul

sthibaul commented Sep 9, 2026

Copy link
Copy Markdown
Collaborator

Thanks!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants