LocalAI version:
4.10.0
Environment, CPU architecture, OS, and Version:
Debian 13, Nvidia RTX 3060 12GB, Architecture: x86_64, VoxCPM: 2.0.3
Describe the bug
VoxCPM 2 Backend installs mismatched python environment
The installed backend is named cuda12-voxcpm, but its Python environment reports a CUDA 13 PyTorch build:
voxcpm: 2.0.3
torch: 2.14.0+cu130
torchaudio: 2.11.0+cu130
transformers: 5.17.0
torchcodec: 0.16.0
safetensors: 0.8.0
CUDA build: 13.0
CUDA avail: True
GPU: NVIDIA GeForce RTX 3060
Capability: (8, 6)
BF16: True
As a result, when generating audio in the TTS Studio, it consistently fails to generate successfully and even invoking the downloaded backend via CLI will produce an output that is just noise
Additional Note
The optimize flag should be exposed as a boolean/toggle option in the backend config
To Reproduce
backends/cuda12-voxcpm/venv/bin/python -m voxcpm.cli design
--model-path models/.artifacts/huggingface/4817c240dc5ab68c1358140aaafae1d1c2f5056a5f53bb8fe4e934f581cfeddb/snapshot
--text "Hello world."
--output /tmp/voxcpm-simple.wav
--device cuda
--no-optimize
--no-denoiser
The resulting WAV contains only synthetic noise.
This reproduces without LocalAI's /v1/audio/speech handler, so the failure appears to be in the Python dependency/runtime packaged with the backend rather than the HTTP/gRPC TTS path.
Expected behavior
cuda12-voxcpm should install a compatible CUDA 12 PyTorch/VoxCPM dependency stack and generate normal speech.
Logs
100%|...| 28/28
Badcase detected, audio_text_ratio=9.333333333333334, retrying...
100%|...| 28/28
Badcase detected, audio_text_ratio=9.333333333333334, retrying...
100%|...| 28/28
Badcase detected, audio_text_ratio=9.333333333333334, retrying...
Additional context
I then ran the same VoxCPM checkpoint and the same basic command:
python -m voxcpm.cli design
--model-path models/.artifacts/huggingface/4817c240dc5ab68c1358140aaafae1d1c2f5056a5f53bb8fe4e934f581cfeddb/snapshot
--text "Hello world."
--output /tmp/voxcpm-test-28.wav
--device cuda
--no-optimize
--no-denoiser
This produced normal intelligible speech.
Therefore:
same machine, RTX 3060, VoxCPM 2.0.3, model checkpoint & inference parameters
LocalAI backend dependency stack -> noise
PyTorch 2.8/cu128 + Transformers 4.57.3 -> correct speech
Suspected cause
The backend dependency versions appear insufficiently pinned, allowing a CUDA 12 backend installation to resolve to a CUDA 13 PyTorch wheel and very recent versions of Torch/Transformers.
This looks similar in class to #11070, where an unpinned Python TTS backend also resolved a CUDA 13 Torch wheel in a CUDA 12 LocalAI backend.
Pinning a known-working VoxCPM stack, or otherwise ensuring that the CUDA 12 backend resolves CUDA 12-compatible Torch/TorchAudio builds, appears necessary.
More Info 1
I also created a fresh Python 3.10 environment, matching the Python version used by LocalAI's VoxCPM backend, and installed VoxCPM 2.0.3 with PyTorch 2.8.0+cu128, TorchAudio 2.8.0+cu128 and Transformers 4.57.3. Using the same checkpoint and CLI command produced normal intelligible speech.
Updating those package versions inside the existing LocalAI backend venv did not resolve the issue; the existing venv continued to produce the same bad-case/noise output.
This indicates the issue is with the packaged cuda12-voxcpm environment as a whole rather than Python version or the model checkpoint.
More Info 2
requirements-cublas12.txt currently contains an unpinned torch dependency and only adds https://download.pytorch.org/whl/cu121 using --extra-index-url. On my installation this resolved cuda12-voxcpm to torch 2.14.0+cu130, despite the backend itself bundling CUDA 12.x libraries under lib/. A fresh Python 3.10 venv with pinned Torch 2.8.0+cu128 / TorchAudio 2.8.0+cu128 / Transformers 4.57.3 produces correct speech.
More Info 3
I rebuilt cuda12-voxcpm/venv from scratch using LocalAI's bundled Python 3.10 interpreter, with:
VoxCPM 2.0.3
PyTorch 2.8.0+cu128
TorchAudio 2.8.0+cu128
Transformers 4.57.3
After restarting LocalAI, VoxCPM generates intelligible speech normally through /v1/audio/speech.
This confirms that replacing the packaged backend environment with a clean environment using the pinned CUDA 12-compatible stack resolves the issue.
Attachments
Ive attached a file that is the output of this command
backends/cuda12-voxcpm/venv/bin/python -m voxcpm.cli design --model-path models/.artifacts/huggingface/4817c240dc5ab68c1358140aaafae1d1c2f5056a5f53bb8fe4e934f581cfeddb/snapshot --text "Hello world." --output /tmp/vox-test.wav --device cuda --no-optimize --no-denoiser
vox-test.wav
Showing the "noise" it generates
LocalAI version:
4.10.0
Environment, CPU architecture, OS, and Version:
Debian 13, Nvidia RTX 3060 12GB, Architecture: x86_64, VoxCPM: 2.0.3
Describe the bug
VoxCPM 2 Backend installs mismatched python environment
The installed backend is named cuda12-voxcpm, but its Python environment reports a CUDA 13 PyTorch build:
voxcpm: 2.0.3
torch: 2.14.0+cu130
torchaudio: 2.11.0+cu130
transformers: 5.17.0
torchcodec: 0.16.0
safetensors: 0.8.0
CUDA build: 13.0
CUDA avail: True
GPU: NVIDIA GeForce RTX 3060
Capability: (8, 6)
BF16: True
As a result, when generating audio in the TTS Studio, it consistently fails to generate successfully and even invoking the downloaded backend via CLI will produce an output that is just noise
Additional Note
The
optimizeflag should be exposed as a boolean/toggle option in the backend configTo Reproduce
backends/cuda12-voxcpm/venv/bin/python -m voxcpm.cli design
--model-path models/.artifacts/huggingface/4817c240dc5ab68c1358140aaafae1d1c2f5056a5f53bb8fe4e934f581cfeddb/snapshot
--text "Hello world."
--output /tmp/voxcpm-simple.wav
--device cuda
--no-optimize
--no-denoiser
The resulting WAV contains only synthetic noise.
This reproduces without LocalAI's /v1/audio/speech handler, so the failure appears to be in the Python dependency/runtime packaged with the backend rather than the HTTP/gRPC TTS path.
Expected behavior
cuda12-voxcpm should install a compatible CUDA 12 PyTorch/VoxCPM dependency stack and generate normal speech.
Logs
100%|...| 28/28
Badcase detected, audio_text_ratio=9.333333333333334, retrying...
100%|...| 28/28
Badcase detected, audio_text_ratio=9.333333333333334, retrying...
100%|...| 28/28
Badcase detected, audio_text_ratio=9.333333333333334, retrying...
Additional context
I then ran the same VoxCPM checkpoint and the same basic command:
python -m voxcpm.cli design
--model-path models/.artifacts/huggingface/4817c240dc5ab68c1358140aaafae1d1c2f5056a5f53bb8fe4e934f581cfeddb/snapshot
--text "Hello world."
--output /tmp/voxcpm-test-28.wav
--device cuda
--no-optimize
--no-denoiser
This produced normal intelligible speech.
Therefore:
same machine, RTX 3060, VoxCPM 2.0.3, model checkpoint & inference parameters
LocalAI backend dependency stack -> noise
PyTorch 2.8/cu128 + Transformers 4.57.3 -> correct speech
Suspected cause
The backend dependency versions appear insufficiently pinned, allowing a CUDA 12 backend installation to resolve to a CUDA 13 PyTorch wheel and very recent versions of Torch/Transformers.
This looks similar in class to #11070, where an unpinned Python TTS backend also resolved a CUDA 13 Torch wheel in a CUDA 12 LocalAI backend.
Pinning a known-working VoxCPM stack, or otherwise ensuring that the CUDA 12 backend resolves CUDA 12-compatible Torch/TorchAudio builds, appears necessary.
More Info 1
I also created a fresh Python 3.10 environment, matching the Python version used by LocalAI's VoxCPM backend, and installed VoxCPM 2.0.3 with PyTorch 2.8.0+cu128, TorchAudio 2.8.0+cu128 and Transformers 4.57.3. Using the same checkpoint and CLI command produced normal intelligible speech.
Updating those package versions inside the existing LocalAI backend venv did not resolve the issue; the existing venv continued to produce the same bad-case/noise output.
This indicates the issue is with the packaged cuda12-voxcpm environment as a whole rather than Python version or the model checkpoint.
More Info 2
requirements-cublas12.txt currently contains an unpinned torch dependency and only adds https://download.pytorch.org/whl/cu121 using --extra-index-url. On my installation this resolved cuda12-voxcpm to torch 2.14.0+cu130, despite the backend itself bundling CUDA 12.x libraries under lib/. A fresh Python 3.10 venv with pinned Torch 2.8.0+cu128 / TorchAudio 2.8.0+cu128 / Transformers 4.57.3 produces correct speech.
More Info 3
I rebuilt cuda12-voxcpm/venv from scratch using LocalAI's bundled Python 3.10 interpreter, with:
VoxCPM 2.0.3
PyTorch 2.8.0+cu128
TorchAudio 2.8.0+cu128
Transformers 4.57.3
After restarting LocalAI, VoxCPM generates intelligible speech normally through /v1/audio/speech.
This confirms that replacing the packaged backend environment with a clean environment using the pinned CUDA 12-compatible stack resolves the issue.
Attachments
Ive attached a file that is the output of this command
backends/cuda12-voxcpm/venv/bin/python -m voxcpm.cli design --model-path models/.artifacts/huggingface/4817c240dc5ab68c1358140aaafae1d1c2f5056a5f53bb8fe4e934f581cfeddb/snapshot --text "Hello world." --output /tmp/vox-test.wav --device cuda --no-optimize --no-denoiservox-test.wav
Showing the "noise" it generates