A chain is checked for nested chains when it is saved, but not when
one of its targets is later edited into a chain. Requests then served
the inner chain's config as the target, which has no backend and
triggers backend auto-detection.
Mark such a target missing so no plan picks it, and skip it without a
trip in the HTTP path when it is pinned.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
The warm list held target names as the chain lists them. For an alias
target that is the alias, but the preloader loads the alias stub (no
backend, no model) and the eviction guard compares against loaded
model names, which never include an alias. A warm alias target was
neither preloaded nor protected from eviction.
Report the model that serves each warm target instead.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
A manager snapshots a target publish when it queues it and applies
every echo the sync layer sends back. With one recovery probe, a
success moves a target from down to recovering to healthy under one
lock and queues two publishes. The echo of "recovering" then arrived
after the target was healthy, rolled it back, switched the chain away
with reason trip and restarted the dwell timer.
Tag each target snapshot with the publishing manager and ignore own
echoes. The manager already holds that state or a newer one.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Transcription and diarization checked response_format only after the
backend had run, and returned a plain error for an unknown value.
Failover counts a plain error as a target failure, so one request with
a bad response_format ran the backend on every target of a chain and
tripped all of them.
Check the format before the backend runs and answer 400.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
The proxy sent every chat and completion request with
context.Background, and the gRPC server gave rich backends no context
at all. When a client disconnected, or failover gave up on the target,
the upstream kept generating to the end, which costs tokens on a paid
or shared upstream. A silent upstream held the backend forever.
Add the optional AIModelRichContext interface. The gRPC server prefers
it and passes the call's context, like the Score and Rerank
extensions. The proxy implements it, so the upstream request ends with
the gRPC call.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
A streamed chat or completion reply never carried token counts: the
upstream LocalAI sends the usage trailer only when the request sets
stream_options.include_usage, and the proxy did not set it. Set it on
every streamed request.
Embeddings of tokenized input arrive in EmbeddingTokens with an empty
Embeddings string, so the proxy embedded an empty string. Send the
tokens as a token list instead.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Regenerate the checked-in specification now that swag can parse
OpenAIResponse again after the master merge. The output also picks up
the SystemOne and process VRAM types that master had not regenerated
yet.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
SyncedMap.Set and Delete changed memory before they wrote the Store.
When the write failed they returned the error, but the unpersisted
value stayed in memory. Callers that re-read the map then applied it
again. The failover manager rolled back a failed pin, but its periodic
pin re-sync read the pin back from the map and re-applied it on that
frontend until the database came back.
Write the Store first and change memory only after it succeeds. A
failed Set or Delete now leaves memory and peers as they were, so the
map always matches what the Store holds. This also stops a failed
create of a fine-tune, quantization or agent task from leaving an
orphan entry that the API had reported as failed.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Design specs and implementation plans are working notes and are not
kept in the tree.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
* fix(kokoros): add missing animate3_d stub to Backend trait impl
#12095 added the Animate3D RPC to backend.proto, but the kokoros
service never got a matching method. The tonic-generated Backend trait
now requires it, so kokoros fails to build with E0046 whenever the
full backend matrix runs.
Return Unimplemented, as the other unsupported RPCs do.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
* fix(kokoros): fill new Result fields with defaults
backend.proto added a metadata field to Result, so the struct literals
in the kokoros service no longer name every field and fail to compile.
Spread Default::default() into them, so later additive proto fields do
not break the build again.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
* feat(gallery): publish signed OCI fallbacks
Publish both official gallery indexes with their local base configs so
an outage of the HTTP and GitHub sources can fall back to Quay.
Keep artifact signing policies separate from backend image policies,
and expose each moving gallery tag only after its digest is signed.
Assisted-by: Codex:gpt-6
* fix(gallery): confine packaged files to selected roots
Use directory-scoped file access to reject symlink escapes during gallery packaging. Create private bundle files for the publishing runner.
Assisted-by: Codex:GPT-6
---------
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
* feat(system): report per-model DRM VRAM
Expose optional resident device memory for local backend process trees.
Deduplicate DRM clients and omit unsupported or incomplete readings.
Document accounting limits and preserve a measured zero in JSON.
Closes#11970.
Assisted-by: Codex:gpt-6
* fix(system): document trusted procfs reads
Scope G304 annotations to paths built from the fixed procfs root,
integer process IDs, and kernel directory entries. These reads accept
no user-controlled path components.
Assisted-by: Codex:GPT-6 gosec
---------
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Keep each message and reasoning item at its announced output index.
Include the answer in completed responses with reasoning or fallback
function calls, and retain reasoning supplied through backend deltas.
Add regression coverage for stream indices, final output, plain text,
and automatic tool parsing.
Assisted-by: Codex:GPT-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Partial JSON parsing heals a name-only chunk into a tool call. The
stream emits that call with empty arguments and skips later chunks.
Require complete JSON before emitting terminal tool-call events.
Preserve complete calls before an unfinished trailing call, and count
only actual tool calls. Add split-chunk regression tests and docs.
Refs #11635. The non-streaming report remains unconfirmed.
Assisted-by: Codex:GPT-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
A request with narrower allow_patterns hashes to a different CacheKey
than an already-committed broader sibling, so committedResult misses and
materializeLocked re-fetches files the sibling already holds. After the
own-tree reuseMaterializedFile miss, consult committed sibling trees
for the same Source (type+endpoint+repo+revision), re-verify the file
through verifyDownloadedFile (full SHA-256, never size-only), and
hard-link it into the writer's staging snapshot (copy fallback only on
EXDEV). Each file is matched individually against the sibling's
manifest, so a broader request can never inherit a narrower sibling's
gaps as if complete.
The sibling manifest set is loaded and source-matched once per
materialization (files indexed by path) instead of once per staged
file, so a models volume with 20 committed artifacts and a 300-file
snapshot does one manifest pass rather than ~6000 reads and JSON
parses. The sibling-reuse behavior cases live in the package's
registered Ginkgo suite so repository test conventions apply.
Refs #11047
Signed-off-by: supermario_leo <leo.stack@outlook.com>
The legacy NVIDIA device reservation requests utility without compute.
Docker derives driver capabilities from that list, leaving CUDA libraries
unavailable even when monitoring works.
Include compute in the legacy example and clarify the matching docs.
Assisted-by: Codex:GPT-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
The v4.10.0 ace-step and VibeVoice merge jobs started just after their
digest artifacts expired. Keep the small digest artifacts for seven
days so a multi-day release matrix can finish publishing its images.
Assisted-by: Codex:gpt-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Swag cannot resolve json.RawMessage in OpenAIResponse and aborts the daily
schema generation. Set its Swagger type without changing JSON encoding,
and regenerate the checked-in specifications.
Assisted-by: Codex:gpt-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
PTQ1_0 is a Prism-private GGUF type (GGML_TYPE_PTQ1_0 = 143 in the
PrismML llama.cpp fork), so stock llama-cpp cannot load it. Switch to
the bonsai backend like the existing ternary-bonsai-27b entries, and
replace the scraped Qwen3.8 description and icon.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Older Go linkers stamp pure-Go hosts with SDK metadata that disables
modern Metal APIs. Select Go 1.27 for Darwin builds and document the
backend rebuild requirement.
Assisted-by: Codex:gpt-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
The link text still said chatbot-ui, but it now pointed at the examples
repository root. Link the configurations directory, which holds the
example model config files, and describe it as such.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
The entry enables spec_type:draft-mtp, so variant ranking needs the mtp
tag. Replace the scraped model-card description, set the Swift Open
License v1.0 and link the base model repo.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
* ⬆️ Update TheTom/llama-cpp-turboquant
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* fix(turboquant): drop the upstreamed D512 patch
Upstream c0e227c guards D512 declarations, dispatch, and instances with
GGML_USE_HIP. This prevents the CUDA shared-memory overflow that our
patch addressed. The old patch now rejects the guarded source.
Remove the obsolete patch for the pinned a3d5603d revision. The remaining
patch series applies successfully, and the build-target test passes.
Assisted-by: Codex:gpt-6
---------
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Picks up mudler/LocalAGI 3ce0a08 "fix(mcp): re-dial an MCP session the
server has dropped". Agents open their MCP sessions once, when they are
created; when the MCP server restarts it forgets them and the go-sdk
client does not reconnect by itself, so the agent kept a dead session -
or, behind a server that revives unknown session IDs, a stale tool
list - until LocalAI restarted.
Only go.mod/go.sum change; core/services/agentpool builds and vets
against the new version.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>