mirror of
https://github.com/mudler/LocalAI.git
synced 2026-05-16 20:52:08 -04:00
* ci(backend_build): plumb builder-base-image and BUILDER_TARGET build-args Adds an optional builder-base-image input. When set, BUILDER_BASE_IMAGE is forwarded as a build-arg AND BUILDER_TARGET=builder-prebuilt is set to select the variant Dockerfile's prebuilt-base stage. When empty, BUILDER_TARGET=builder-fromsource (the default) keeps the existing from-source build path. This makes the prebuilt-base optimization opt-in per matrix entry without breaking local `make backends/<name>` invocations or backends whose Dockerfile doesn't have a prebuilt path. Assisted-by: Claude:claude-opus-4-7 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * ci(llama-cpp,ik-llama-cpp,turboquant): multi-target Dockerfiles for prebuilt + from-source Restructure the three llama.cpp-derived Dockerfiles so each supports two builder paths in a single file, selected via the BUILDER_TARGET build-arg: BUILDER_TARGET=builder-fromsource (default) - Standalone build: gRPC stage + apt installs + (conditionally) CUDA/ROCm/Vulkan + compile. - Used by `make backends/llama-cpp` locally and any caller that doesn't supply a prebuilt base. BUILDER_TARGET=builder-prebuilt - FROM \${BUILDER_BASE_IMAGE} (one of quay.io/go-skynet/ci-cache: base-grpc-* shipped in PR #9737). - Skips ~25-35 min of gRPC compile + ~5-10 min of toolchain installs. - Used by CI when the matrix entry sets builder-base-image. Final FROM scratch resolves BUILDER_TARGET via an aliasing FROM stage (BuildKit doesn't support variable expansion directly in COPY --from), then COPY --from=builder pulls package output from the chosen path. BuildKit prunes the unreferenced builder, so each build only does the work for the chosen path. The compile RUN is identical between both builder stages, so it's factored into .docker/<name>-compile.sh and bind-mounted into both. ccache mount + cache-id stay per-arch / per-build-type. Local DX preserved: `make backends/llama-cpp` (no extra args) defaults to BUILDER_TARGET=builder-fromsource and works exactly as before. Assisted-by: Claude:claude-opus-4-7 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * ci(backend.yml,backend_pr.yml): forward builder-base-image from matrix Plumbs the new optional builder-base-image input from matrix into backend_build.yml. backend_build.yml derives BUILDER_TARGET from whether builder-base-image is set, so matrix entries that map to a prebuilt base get the prebuilt path; entries that don't (python/go/ rust backends) fall through to the default builder-fromsource (which their own Dockerfiles don't reference, so it's a no-op for them). Assisted-by: Claude:claude-opus-4-7 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * ci(backend-matrix): wire builder-base-image to llama-cpp variants For every entry whose Dockerfile is llama-cpp/ik-llama-cpp/turboquant, add a builder-base-image field pointing at the appropriate prebuilt quay.io/go-skynet/ci-cache:base-grpc-* tag. backend_build.yml derives BUILDER_TARGET from this field's presence: non-empty -> builder-prebuilt; empty -> builder-fromsource. So this commit alone activates the prebuilt-base path for these 23 backends in CI, while local `make backends/<name>` (no extra args) keeps the from-source path. Mapping by (build-type, arch): - '' / amd64 -> base-grpc-amd64 - '' / arm64 -> base-grpc-arm64 - cublas-12 / amd64 -> base-grpc-cuda-12-amd64 - cublas-13 / amd64 -> base-grpc-cuda-13-amd64 - cublas-13 / arm64 -> base-grpc-cuda-13-arm64 - hipblas / amd64 -> base-grpc-rocm-amd64 - vulkan / amd64 -> base-grpc-vulkan-amd64 - vulkan / arm64 -> base-grpc-vulkan-arm64 - sycl_* / amd64 -> base-grpc-intel-amd64 - cublas-12 + JetPack r36.4.0 / arm64 -> base-grpc-l4t-cuda-12-arm64 Cold-build savings expected: ~25-35 min per variant (skips the gRPC compile + toolchain install that's now in the base). Assisted-by: Claude:claude-opus-4-7 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * ci: add base-grpc-l4t-cuda-12-arm64 variant for legacy JetPack entries Two matrix entries (-nvidia-l4t-arm64-llama-cpp, -nvidia-l4t-arm64- turboquant) build against nvcr.io/nvidia/l4t-jetpack:r36.4.0 + CUDA 12 ARM64. They're distinct from -nvidia-l4t-cuda-13-arm64-* which use Ubuntu 24.04 + CUDA 13 sbsa. Add the missing JetPack-based variant to base-images.yml so those two entries' builder-base-image mapping in the previous commit resolves. Bootstrap order before merging this PR (re-run base-images.yml on this branch — 9 existing variants hit BuildKit cache, only the new l4t-cuda-12-arm64 builds cold): gh workflow run base-images.yml --ref ci/base-images-consumers Assisted-by: Claude:claude-opus-4-7 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * ci: extract base-builder install logic into .docker/install-base-deps.sh Pre-extraction, the apt + protoc + cmake + conditional CUDA/ROCm/Vulkan + gRPC install logic was duplicated across four files: - backend/Dockerfile.base-grpc-builder (CI prebuilt-base source of truth) - backend/Dockerfile.llama-cpp (builder-fromsource stage) - backend/Dockerfile.ik-llama-cpp (builder-fromsource stage) - backend/Dockerfile.turboquant (builder-fromsource stage) A bump to e.g. CUDA toolkit packages had to be made in 4 places, and drift between the prebuilt base and the variant-Dockerfile from-source path was a real concern (ik-llama-cpp's hipblas branch was already missing the rocBLAS Kernels echo that llama-cpp / turboquant / base-grpc-builder all had). Factor the install logic into a single .docker/install-base-deps.sh that reads its inputs from env vars and runs conditionally on BUILD_TYPE / CUDA_*_VERSION / TARGETARCH. Each Dockerfile now bind- mounts the script alongside .docker/apt-mirror.sh and invokes it from a single RUN step. The variant Dockerfiles' grpc-source stage is removed entirely — the script handles gRPC compile + install at /opt/grpc, and the builder-fromsource stage mirrors builder-prebuilt by copying /opt/grpc/. to /usr/local/. Result: - install-base-deps.sh: 244 lines (one source of truth) - Dockerfile.base-grpc-builder: 268 -> 98 lines - Dockerfile.llama-cpp: 361 -> 157 lines - Dockerfile.ik-llama-cpp: 348 -> 151 lines - Dockerfile.turboquant: 355 -> 154 lines - Total Dockerfile bytes: 1332 -> 560 lines (58% reduction) Bit-equivalence between prebuilt and from-source paths is now enforced by construction: both invoke the same script with the same inputs. A side-effect is that ik-llama-cpp now also gets the rocBLAS Kernels echo + clblas block parity it was previously missing. Includes the BUILD_TYPE=clblas branch (libclblast-dev) for parity even though no current CI matrix entry uses it. After this commit's force-push, base-images.yml needs to be redispatched on this branch — the Dockerfile.base-grpc-builder content shifts so the existing cache won't apply for the install layer (gRPC layer also rebuilds since it's now in the same RUN step). Assisted-by: Claude:claude-opus-4-7 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * ci(base-images): skip-drivers on JetPack l4t variant cuda-nvcc-12-0 isn't installable via apt on the JetPack r36.4.0 base image — JetPack ships CUDA preinstalled at /usr/local/cuda and its apt feed doesn't carry the cuda-nvcc-* packages from the public repositories. The original matrix entry for -nvidia-l4t-arm64-llama-cpp on master sets skip-drivers: 'true' for exactly this reason; the new base-grpc-l4t-cuda-12-arm64 base needs to match. Also forwards SKIP_DRIVERS as a build-arg from matrix into the build (was missing entirely before this commit). Caught by run 25612030775 — l4t-cuda-12-arm64 failed at: E: Package 'cuda-nvcc-12-0' has no installation candidate Assisted-by: Claude:claude-opus-4-7 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
99 lines
4.0 KiB
Docker
99 lines
4.0 KiB
Docker
# syntax=docker/dockerfile:1.7
|
|
#
|
|
# Pre-built builder base image for LocalAI's C++ backends.
|
|
#
|
|
# This Dockerfile is the source of truth for the
|
|
# `quay.io/go-skynet/ci-cache:base-grpc-*` images that
|
|
# `.github/workflows/base-images.yml` builds and pushes. The output of a
|
|
# build is a fully-prepped builder layer containing:
|
|
#
|
|
# - apt build deps (build-essential, ccache, git, make, pkg-config,
|
|
# libcurl4-openssl-dev, libssl-dev, curl, unzip, wget, ca-certificates)
|
|
# - cmake (apt or, when CMAKE_FROM_SOURCE=true, compiled from
|
|
# ${CMAKE_VERSION})
|
|
# - protoc v27.1 at /usr/local/bin/protoc
|
|
# - gRPC ${GRPC_VERSION} compiled and installed at /opt/grpc
|
|
# - Conditional CUDA toolkit (BUILD_TYPE=cublas|l4t, SKIP_DRIVERS=false)
|
|
# including the cuda-13 + arm64 cudss/nvpl special case
|
|
# - Conditional ROCm/HIP build deps (BUILD_TYPE=hipblas)
|
|
# - Conditional Vulkan SDK 1.4.335.0 (BUILD_TYPE=vulkan)
|
|
#
|
|
# Variants built by the workflow (matrix in base-images.yml):
|
|
#
|
|
# base-grpc-amd64 ubuntu:24.04, CPU-only
|
|
# base-grpc-arm64 ubuntu:24.04, CPU-only
|
|
# base-grpc-cuda-12-amd64 ubuntu:24.04 + CUDA 12.8
|
|
# base-grpc-cuda-13-amd64 ubuntu:22.04 + CUDA 13.0
|
|
# base-grpc-cuda-13-arm64 ubuntu:24.04 + CUDA 13.0 (sbsa)
|
|
# base-grpc-l4t-cuda-12-arm64 ubuntu:22.04 + CUDA 12.x (legacy JetPack)
|
|
# base-grpc-rocm-amd64 rocm/dev-ubuntu-24.04:7.2.1 + hipblas
|
|
# base-grpc-vulkan-amd64 ubuntu:24.04 + Vulkan SDK 1.4.335
|
|
# base-grpc-vulkan-arm64 ubuntu:24.04 + Vulkan SDK ARM 1.4.335
|
|
# base-grpc-intel-amd64 intel/oneapi-basekit:2025.3.2 (sycl)
|
|
#
|
|
# This is a SINGLE-stage Dockerfile by design: the final image IS the
|
|
# builder base. The intermediate gRPC compile happens inside this same
|
|
# stage so consumer Dockerfiles in PR 2 can simply
|
|
# `FROM quay.io/go-skynet/ci-cache:base-grpc-<variant>` without needing a
|
|
# COPY --from=grpc step. /opt/grpc is the canonical install prefix and
|
|
# downstream builds will add it to CMAKE_PREFIX_PATH (or copy to
|
|
# /usr/local) the same way Dockerfile.llama-cpp does today.
|
|
#
|
|
# Install logic lives in .docker/install-base-deps.sh, which is also
|
|
# bind-mounted by the variant Dockerfiles' builder-fromsource stage.
|
|
# This guarantees bit-equivalence between the prebuilt CI base and the
|
|
# from-source local-dev path — both invoke the same script with the
|
|
# same env inputs.
|
|
|
|
ARG BASE_IMAGE=ubuntu:24.04
|
|
|
|
FROM ${BASE_IMAGE}
|
|
|
|
ARG BASE_IMAGE=ubuntu:24.04
|
|
ARG BUILD_TYPE=""
|
|
ARG CUDA_MAJOR_VERSION=""
|
|
ARG CUDA_MINOR_VERSION=""
|
|
ARG CMAKE_FROM_SOURCE=false
|
|
# CUDA Toolkit 13.x compatibility: CMake 3.31.9+ fixes toolchain
|
|
# detection / arch table issues.
|
|
ARG CMAKE_VERSION=3.31.10
|
|
ARG GRPC_VERSION=v1.65.0
|
|
ARG GRPC_MAKEFLAGS="-j4 -Otarget"
|
|
ARG SKIP_DRIVERS=false
|
|
ARG TARGETARCH
|
|
ARG UBUNTU_VERSION=2404
|
|
ARG APT_MIRROR=""
|
|
ARG APT_PORTS_MIRROR=""
|
|
ARG AMDGPU_TARGETS=""
|
|
|
|
ENV BUILD_TYPE=${BUILD_TYPE} \
|
|
CUDA_MAJOR_VERSION=${CUDA_MAJOR_VERSION} \
|
|
CUDA_MINOR_VERSION=${CUDA_MINOR_VERSION} \
|
|
CMAKE_FROM_SOURCE=${CMAKE_FROM_SOURCE} \
|
|
CMAKE_VERSION=${CMAKE_VERSION} \
|
|
GRPC_VERSION=${GRPC_VERSION} \
|
|
GRPC_MAKEFLAGS=${GRPC_MAKEFLAGS} \
|
|
SKIP_DRIVERS=${SKIP_DRIVERS} \
|
|
TARGETARCH=${TARGETARCH} \
|
|
UBUNTU_VERSION=${UBUNTU_VERSION} \
|
|
APT_MIRROR=${APT_MIRROR} \
|
|
APT_PORTS_MIRROR=${APT_PORTS_MIRROR} \
|
|
AMDGPU_TARGETS=${AMDGPU_TARGETS} \
|
|
MAKEFLAGS=${GRPC_MAKEFLAGS} \
|
|
DEBIAN_FRONTEND=noninteractive
|
|
|
|
# CUDA on PATH (no-op when CUDA isn't installed)
|
|
ENV PATH=/usr/local/cuda/bin:${PATH}
|
|
# HipBLAS / ROCm on PATH (no-op when ROCm isn't installed)
|
|
ENV PATH=/opt/rocm/bin:${PATH}
|
|
|
|
WORKDIR /build
|
|
|
|
# Single RUN that delegates to .docker/install-base-deps.sh — the same
|
|
# script the variant Dockerfiles' builder-fromsource stage runs.
|
|
RUN --mount=type=bind,source=.docker/install-base-deps.sh,target=/usr/local/sbin/install-base-deps \
|
|
--mount=type=bind,source=.docker/apt-mirror.sh,target=/usr/local/sbin/apt-mirror \
|
|
bash /usr/local/sbin/install-base-deps
|
|
|
|
WORKDIR /
|