mirror of
https://github.com/mudler/LocalAI.git
synced 2026-07-31 10:28:43 -04:00
feat(bonsai): add PrismML llama.cpp fork backend + Bonsai gallery models Adds a new `bonsai` backend that runs the PrismML fork of llama.cpp (github.com/PrismML-Eng/llama.cpp, `prism` branch), which ships the Q1_0 (1-bit) and Q2_0 (ternary / 1.58-bit) weight-quantization kernels used by the Bonsai and Ternary-Bonsai models. Stock llama.cpp cannot decode these quants. Modeled on the turboquant backend: reuses backend/cpp/llama-cpp/grpc-server.cpp against the fork's libllama via a thin wrapper Makefile, so the sub-2-bit models are served with the same OpenAI-compatible API. No grpc-server allow-list patch is needed (bonsai adds weight quants, transparent to the server, not KV-cache types), and the reused server compiles cleanly against the fork with no skew patches (validated locally via a CPU docker build; patches/ is present but empty for any future re-pin skew). Backend wiring: backend/cpp/bonsai/, .docker/bonsai-compile.sh, backend/Dockerfile.bonsai, top-level Makefile targets, backend-matrix.yml build rows (CPU, CUDA 12/13, L4T, SYCL f32/f16, Vulkan, ROCm/hipblas), backend/index.yaml meta-backend + per-platform images, and a nightly bump_deps entry tracking the `prism` branch. Gallery: 8 entries across 4 families - bonsai-8b-1bit, ternary-bonsai-8b (+g64, +pq2), bonsai-27b-1bit (vision), ternary-bonsai-27b (+pq2, +g64, vision). The 27B models wire the mmproj vision tower; the DSpark speculative drafter GGUFs are not wired (custom semi-autoregressive drafter, not a standard llama.cpp draft model). Assisted-by: Claude:claude-opus-4-8 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
49 lines
1.3 KiB
Bash
Executable File
49 lines
1.3 KiB
Bash
Executable File
#!/bin/bash
|
|
# Apply the bonsai patch series to a cloned PrismML llama.cpp (prism branch) checkout.
|
|
#
|
|
# The prism fork branched from upstream llama.cpp before a number of API changes that the
|
|
# shared backend/cpp/llama-cpp/grpc-server.cpp depends on. We carry those upstream commits
|
|
# as patch files under backend/cpp/bonsai/patches/ and apply them here so the reused
|
|
# grpc-server source compiles against the fork unmodified.
|
|
#
|
|
# Drop the corresponding patch from patches/ whenever the fork catches up with upstream —
|
|
# the build will fail fast if a patch stops applying, which is the signal to retire it.
|
|
|
|
set -euo pipefail
|
|
|
|
if [[ $# -ne 2 ]]; then
|
|
echo "usage: $0 <llama.cpp-src-dir> <patches-dir>" >&2
|
|
exit 2
|
|
fi
|
|
|
|
SRC_DIR=$1
|
|
PATCHES_DIR=$2
|
|
|
|
if [[ ! -d "$SRC_DIR" ]]; then
|
|
echo "source dir does not exist: $SRC_DIR" >&2
|
|
exit 2
|
|
fi
|
|
|
|
if [[ ! -d "$PATCHES_DIR" ]]; then
|
|
echo "no patches dir at $PATCHES_DIR, nothing to apply"
|
|
exit 0
|
|
fi
|
|
|
|
shopt -s nullglob
|
|
patches=("$PATCHES_DIR"/*.patch)
|
|
shopt -u nullglob
|
|
|
|
if [[ ${#patches[@]} -eq 0 ]]; then
|
|
echo "no .patch files in $PATCHES_DIR, nothing to apply"
|
|
exit 0
|
|
fi
|
|
|
|
cd "$SRC_DIR"
|
|
|
|
for patch in "${patches[@]}"; do
|
|
echo "==> applying $patch"
|
|
git apply --verbose "$patch"
|
|
done
|
|
|
|
echo "all bonsai patches applied successfully"
|