mirror of
https://github.com/mudler/LocalAI.git
synced 2026-07-30 09:57:57 -04:00
feat/buun-llama-cpp-backend
5 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8d6fdf22d3 |
fix(backends): derive the protoc generator from the protobuf runtime, regenerate stubs after late installs (#11057)
* fix(backends): choose the protoc generator from the protobuf runtime, and regenerate stubs after late installs The vLLM backends still crash on startup with VersionError: Detected incompatible Protobuf Gencode/Runtime versions when loading backend.proto: gencode 7.35.0 runtime 6.33.6 despite #10735 and #10944. Three separate defects kept it alive. 1. runProtogen picked the generator from the installed *grpcio* version. grpcio-tools' version tracks grpcio, but the gencode its bundled protoc emits tracks *protobuf*, and the two move independently: grpcio-tools 1.82.1 (the version #10735 pins to, matching grpcio 1.82.1) requires protobuf>=7.35.1 and stamps gencode 7.35.0. Pinning to grpcio could therefore never constrain the gencode. Constrain the install to the protobuf already in the venv instead and let the resolver pick the newest compatible grpcio-tools. That both selects a generator the runtime accepts and stops protogen from moving the runtime under the backend's other deps. This is self-correcting, so the hardcoded GRPCIO_TOOLS_VERSION=1.78.0 escape hatch from #10944 is no longer needed and is removed. 2. The stubs were generated too early. Most branches of vllm/install.sh (and vllm-omni) install vllm *after* installRequirements, and vllm re-resolves the protobuf runtime as it lands. Stubs generated against the pre-vllm runtime can end up newer than the runtime that finally ships, which is the ROCm failure exactly. Regenerate once the dependency set is final. 3. rm -f of the .py sources left __pycache__ behind. CPython validates a .pyc against source mtime and size, both of which can be unchanged across a regeneration (the gencode triple is the same width whether it reads 7.35.0 or 6.33.5), so a stale backend_pb2.pyc could shadow the stub just written. Also fail the build when the generated stub cannot be imported, so a gencode/runtime mismatch surfaces at image build time instead of reaching users as an opaque "grpc service not ready". Verified by driving the real runProtogen through the ROCm install sequence in a venv harness: before, gencode 7.35.0 against runtime 6.33.6 (reproducing the reported error verbatim); after, gencode 6.33.5 against runtime 6.33.6 and the stub imports cleanly. Closes #10940 Closes #10718 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-4-8[1m] [Bash] [Edit] * fix(backends): regenerate protobuf stubs in the other backends that install after installRequirements Same defect as the vllm change: installRequirements generates the stubs at the end of its own run, so any backend that installs further packages afterwards can have the protobuf runtime moved out from under stubs that were already written. The gencode stamped into backend_pb2.py then exceeds the runtime that ships and the backend dies at model load with "grpc service not ready". fish-speech already had this bug and worked around the symptom: it forces protobuf>=5.29.0 after installRequirements precisely because "transitive deps (wandb, tensorboard) may downgrade protobuf to 3.x but our generated backend_pb2.py requires protobuf 5+". Regenerating after the pin addresses the cause rather than propping up the runtime to match stale stubs. Applied to the backends whose post-installRequirements step resolves a dependency graph and can therefore move protobuf: fish-speech -e . plus an explicit protobuf install vibevoice pip install . (with deps) llama-cpp-quantization gguf / GGUF_PIP_SPEC trl gguf / GGUF_PIP_SPEC Deliberately not applied to ace-step and chatterbox (both --no-deps, so the dependency graph cannot change) or voxcpm (pins setuptools only). gguf does not depend on protobuf today, but it resolves dependencies, and "this package does not touch protobuf right now" is exactly the assumption that made the earlier fix ineffective. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-4-8[1m] [Bash] [Edit] * fix(backends): resolve the protoc generator in a throwaway env so it cannot edit the backend's pinned deps Installing grpcio-tools into the backend's own venv to generate the stubs also drags its dependencies in: grpcio-tools 1.82.1 requires grpcio>=1.82.1, so a backend that pinned grpcio==1.78.1 silently shipped 1.82.1 instead. Caught by building the llama-cpp-quantization image and reading the versions back out of the artifact: before grpcio 1.82.1 (requirements.txt pins grpcio==1.78.1) after grpcio 1.78.1 grpcio-tools absent from the venv entirely Resolve the generator in a throwaway environment instead, still constrained to the protobuf the backend ships so the gencode stays compatible. The backend's dependency set is then exactly what its requirements files declared. protoc's output is plain Python and carries no dependency on the interpreter that produced it, so generating from a different env is safe; the import check still runs under the backend's python, since that is the interpreter that has to load the stubs at model load. Verified on the rebuilt image: gencode 7.35.0, runtime protobuf 7.35.1, grpcio back at its pinned 1.78.1, and the shipped stub imports cleanly against 7.35.1. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-4-8[1m] [Bash] [Edit] * fix(backends): bound the protoc generator by BOTH the installed grpcio and protobuf The generated stubs impose two independent constraints, and every fix so far, including the previous commit on this branch, satisfied one while violating the other: backend_pb2.py needs protobuf runtime >= gencode backend_pb2_grpc.py needs installed grpcio >= grpcio-tools Resolving the generator against protobuf alone picked grpcio-tools 1.82.1 for a backend holding grpcio at 1.78.1, so the gencode was fine but the gRPC stub was not: RuntimeError: The grpc package installed is at version 1.78.1, but the generated code in backend_pb2_grpc.py depends on grpcio>=1.82.1. That is also why installing grpcio-tools into the backend venv appeared to work earlier: it dragged grpcio up to match, which was load-bearing rather than the regression it looked like. Isolating the generator removed the accidental fix and exposed the missing constraint. Bound grpcio-tools from both sides instead and let the resolver find the newest version satisfying both. The protobuf ceiling makes it back off to an older generator when the runtime trails, bounding the gencode; the grpcio ceiling keeps the _grpc stub loadable. Resolved against the four real runtime pairs observed in built images: grpcio 1.78.1 / protobuf 7.35.1 -> grpcio-tools 1.78.0, gencode 6.31.1 OK grpcio 1.78.0 / protobuf 6.33.6 -> grpcio-tools 1.78.0, gencode 6.31.1 OK grpcio 1.82.1 / protobuf 6.33.6 -> grpcio-tools 1.81.1, gencode 6.33.5 OK grpcio 1.82.1 / protobuf 7.35.1 -> grpcio-tools 1.82.1, gencode 7.35.0 OK Also restore the import check to cover backend_pb2_grpc as well as backend_pb2. Narrowing it to backend_pb2 is why the image build passed while CI failed: the guard could not see the constraint that was actually broken. Verified by running the CI sequence locally for llama-cpp-quantization, the backend whose test failed: make -C backend/python/llama-cpp-quantization -> exit 0 make -C backend/python/llama-cpp-quantization test -> exit 0, OK Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-4-8[1m] [Bash] [Edit] --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
9043cbc786 |
chore(deps): bump torch CPU wheels to 2.12.1 (#10969)
* chore(deps): bump the pip group across 6 directories with 1 update Bumps the pip group with 1 update in the /backend/python/ace-step directory: torch. Bumps the pip group with 1 update in the /backend/python/llama-cpp-quantization directory: torch. Bumps the pip group with 1 update in the /backend/python/longcat-video directory: torch. Bumps the pip group with 1 update in the /backend/python/sglang directory: torch. Bumps the pip group with 1 update in the /backend/python/trl directory: torch. Bumps the pip group with 1 update in the /backend/python/vllm-omni directory: torch. Updates `torch` from 2.10.0+rocm7.0 to 2.12.1+cpu Updates `torch` from 2.10.0 to 2.12.1+cpu Updates `torch` from 2.12.1 to 2.12.1+cu130 Updates `torch` from 2.9.0 to 2.12.1+cpu Updates `torch` from 2.10.0 to 2.12.1+cpu Updates `torch` from 2.7.0 to 2.12.1+cu130 --- updated-dependencies: - dependency-name: torch dependency-version: 2.12.1+cpu dependency-type: direct:production dependency-group: pip - dependency-name: torch dependency-version: 2.12.1+cpu dependency-type: direct:production dependency-group: pip - dependency-name: torch dependency-version: 2.12.1+cu130 dependency-type: direct:production dependency-group: pip - dependency-name: torch dependency-version: 2.12.1+cpu dependency-type: direct:production dependency-group: pip - dependency-name: torch dependency-version: 2.12.1+cpu dependency-type: direct:production dependency-group: pip - dependency-name: torch dependency-version: 2.12.1+cu130 dependency-type: direct:production dependency-group: pip ... Signed-off-by: dependabot[bot] <support@github.com> * fix(deps): preserve platform-specific torch requirements Keep the 2.12.1 CPU bump only where uv resolves it cleanly, and restore ROCm, CUDA, MPS, and unrelated transformers constraints that Dependabot rewrote to incompatible wheel variants. Assisted-by: Codex:gpt-5 [uv] --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
0258f8af55 |
fix(backends): repair release CI build/test breaks (kokoros, fish-speech, llama-cpp-quantization, sglang) (#10547)
* fix(kokoros): implement new Backend RPCs to fix the build
The backend.proto grew six RPCs (SoundDetection, Depth, TokenClassify,
Score and the bidi-streaming Forward) that the kokoros gRPC service never
implemented, so the trait impl no longer satisfies `Backend`:
error[E0046]: not all trait items implemented, missing:
`sound_detection`, `depth`, `token_classify`, `score`,
`ForwardStream`, `forward`
kokoros is a TTS backend with no use for these, so add `unimplemented`
stubs (plus the `ForwardStream` associated type) matching the existing
pattern for every other unsupported RPC in this file.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-8 [Claude Code]
* fix(fish-speech): add setuptools-rust for the editable source install
install.sh installs the fish-speech source tree editable with
`--no-build-isolation`, which means the build backends of its transitive
dependencies must already be present in the venv. One of them builds a
Rust extension and its metadata step fails with:
ModuleNotFoundError: No module named 'setuptools_rust'
Add setuptools-rust to requirements.txt so installRequirements provisions
it before the editable install runs.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-8 [Claude Code]
* fix(llama-cpp-quantization): vendor convert_hf_to_gguf.py with conversion/
Upstream llama.cpp split the model-specific logic out of the single
convert_hf_to_gguf.py file into a sibling `conversion/` package, so the
script now starts with `from conversion import ...`. Downloading just the
one file therefore fails at runtime with:
ModuleNotFoundError: No module named 'conversion'
Clone the repo (reusing the clone already needed to build llama-quantize)
and copy both the script and the `conversion/` package into the backend
dir. Python puts the script's own directory on sys.path[0], so the package
resolves when it sits beside the script.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-8 [Claude Code]
* fix(sglang): pin the CPU source build to sglang v0.5.11
The CPU profile builds sgl-kernel from a `git clone` of sglang with no
ref, so it always tracks master. Recent master added CPU kernels (e.g.
mamba/fla.cpp) that fail to compile in our builder:
constexpr variable 'scale' must be initialized by a constant
static library kineto_LIBRARY-NOTFOUND not found
Pin the clone to v0.5.11, the same release the GPU path already floors on
(requirements-cublas12-after.txt). Overridable via SGLANG_VERSION so the
pin can be bumped deliberately.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-8 [Claude Code]
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
|
||
|
|
59108fbe32 |
feat: add distributed mode (#9124)
* feat: add distributed mode (experimental) Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix data races, mutexes, transactions Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactorings Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fixups Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix events and tool stream in agent chat Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * use ginkgo Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactoring and consolidation Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactoring and consolidation Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactoring and consolidation Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactoring and consolidation Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactoring and consolidation Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactoring and consolidation Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactoring and consolidation Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactoring and consolidation Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(cron): compute correctly time boundaries avoiding re-triggering Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * enhancements, refactorings Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * do not flood of healthy checks Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * do not list obvious backends as text backends Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * tests fixups Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactoring and consolidation Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * Drop redundant healthcheck Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * enhancements, refactorings Signed-off-by: Ettore Di Giacinto <mudler@localai.io> --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
f7e8d9e791 |
feat(quantization): add quantization backend (#9096)
Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |