Skip to content

VoxCPM Backend Installs CUDA 13 PyTorch Stack which results in noise rather than Speech #12356

Description

@shopsD

LocalAI version:
4.10.0

Environment, CPU architecture, OS, and Version:

Debian 13, Nvidia RTX 3060 12GB, Architecture: x86_64, VoxCPM: 2.0.3

Describe the bug
VoxCPM 2 Backend installs mismatched python environment
The installed backend is named cuda12-voxcpm, but its Python environment reports a CUDA 13 PyTorch build:

voxcpm: 2.0.3
torch: 2.14.0+cu130
torchaudio: 2.11.0+cu130
transformers: 5.17.0
torchcodec: 0.16.0
safetensors: 0.8.0
CUDA build: 13.0
CUDA avail: True
GPU: NVIDIA GeForce RTX 3060
Capability: (8, 6)
BF16: True

As a result, when generating audio in the TTS Studio, it consistently fails to generate successfully and even invoking the downloaded backend via CLI will produce an output that is just noise

Additional Note
The optimize flag should be exposed as a boolean/toggle option in the backend config

To Reproduce
backends/cuda12-voxcpm/venv/bin/python -m voxcpm.cli design
--model-path models/.artifacts/huggingface/4817c240dc5ab68c1358140aaafae1d1c2f5056a5f53bb8fe4e934f581cfeddb/snapshot
--text "Hello world."
--output /tmp/voxcpm-simple.wav
--device cuda
--no-optimize
--no-denoiser

The resulting WAV contains only synthetic noise.

This reproduces without LocalAI's /v1/audio/speech handler, so the failure appears to be in the Python dependency/runtime packaged with the backend rather than the HTTP/gRPC TTS path.

Expected behavior
cuda12-voxcpm should install a compatible CUDA 12 PyTorch/VoxCPM dependency stack and generate normal speech.

Logs
100%|...| 28/28
Badcase detected, audio_text_ratio=9.333333333333334, retrying...

100%|...| 28/28
Badcase detected, audio_text_ratio=9.333333333333334, retrying...

100%|...| 28/28
Badcase detected, audio_text_ratio=9.333333333333334, retrying...

Additional context
I then ran the same VoxCPM checkpoint and the same basic command:

python -m voxcpm.cli design
--model-path models/.artifacts/huggingface/4817c240dc5ab68c1358140aaafae1d1c2f5056a5f53bb8fe4e934f581cfeddb/snapshot
--text "Hello world."
--output /tmp/voxcpm-test-28.wav
--device cuda
--no-optimize
--no-denoiser

This produced normal intelligible speech.

Therefore:
same machine, RTX 3060, VoxCPM 2.0.3, model checkpoint & inference parameters
LocalAI backend dependency stack -> noise
PyTorch 2.8/cu128 + Transformers 4.57.3 -> correct speech

Suspected cause

The backend dependency versions appear insufficiently pinned, allowing a CUDA 12 backend installation to resolve to a CUDA 13 PyTorch wheel and very recent versions of Torch/Transformers.

This looks similar in class to #11070, where an unpinned Python TTS backend also resolved a CUDA 13 Torch wheel in a CUDA 12 LocalAI backend.

Pinning a known-working VoxCPM stack, or otherwise ensuring that the CUDA 12 backend resolves CUDA 12-compatible Torch/TorchAudio builds, appears necessary.

More Info 1

I also created a fresh Python 3.10 environment, matching the Python version used by LocalAI's VoxCPM backend, and installed VoxCPM 2.0.3 with PyTorch 2.8.0+cu128, TorchAudio 2.8.0+cu128 and Transformers 4.57.3. Using the same checkpoint and CLI command produced normal intelligible speech.
Updating those package versions inside the existing LocalAI backend venv did not resolve the issue; the existing venv continued to produce the same bad-case/noise output.
This indicates the issue is with the packaged cuda12-voxcpm environment as a whole rather than Python version or the model checkpoint.

More Info 2
requirements-cublas12.txt currently contains an unpinned torch dependency and only adds https://download.pytorch.org/whl/cu121 using --extra-index-url. On my installation this resolved cuda12-voxcpm to torch 2.14.0+cu130, despite the backend itself bundling CUDA 12.x libraries under lib/. A fresh Python 3.10 venv with pinned Torch 2.8.0+cu128 / TorchAudio 2.8.0+cu128 / Transformers 4.57.3 produces correct speech.

More Info 3
I rebuilt cuda12-voxcpm/venv from scratch using LocalAI's bundled Python 3.10 interpreter, with:
VoxCPM 2.0.3
PyTorch 2.8.0+cu128
TorchAudio 2.8.0+cu128
Transformers 4.57.3
After restarting LocalAI, VoxCPM generates intelligible speech normally through /v1/audio/speech.
This confirms that replacing the packaged backend environment with a clean environment using the pinned CUDA 12-compatible stack resolves the issue.

Attachments
Ive attached a file that is the output of this command
backends/cuda12-voxcpm/venv/bin/python -m voxcpm.cli design --model-path models/.artifacts/huggingface/4817c240dc5ab68c1358140aaafae1d1c2f5056a5f53bb8fe4e934f581cfeddb/snapshot --text "Hello world." --output /tmp/vox-test.wav --device cuda --no-optimize --no-denoiser

vox-test.wav

Showing the "noise" it generates

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions