Files
LocalAI/backend/python/qwen-tts
Ettore Di Giacinto 1dfa25e681 fix(qwen-tts): limit CUDA compiler threads
The CUDA 13 FlashAttention build still exhausts hosted-runner memory with a single ninja worker because nvcc can compile multiple threads internally. Limit nvcc to one thread for that profile and guard the setting in the backend test script.

Assisted-by: Codex:gpt-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-11 23:19:28 +00:00
..
2026-02-03 22:07:07 +01:00