A streamed chat or completion reply never carried token counts: the
upstream LocalAI sends the usage trailer only when the request sets
stream_options.include_usage, and the proxy did not set it. Set it on
every streamed request.
Embeddings of tokenized input arrive in EmbeddingTokens with an empty
Embeddings string, so the proxy embedded an empty string. Send the
tokens as a token list instead.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Regenerate the checked-in specification now that swag can parse
OpenAIResponse again after the master merge. The output also picks up
the SystemOne and process VRAM types that master had not regenerated
yet.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
SyncedMap.Set and Delete changed memory before they wrote the Store.
When the write failed they returned the error, but the unpersisted
value stayed in memory. Callers that re-read the map then applied it
again. The failover manager rolled back a failed pin, but its periodic
pin re-sync read the pin back from the map and re-applied it on that
frontend until the database came back.
Write the Store first and change memory only after it succeeds. A
failed Set or Delete now leaves memory and peers as they were, so the
map always matches what the Store holds. This also stops a failed
create of a fine-tune, quantization or agent task from leaving an
orphan entry that the API had reported as failed.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Design specs and implementation plans are working notes and are not
kept in the tree.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
* fix(kokoros): add missing animate3_d stub to Backend trait impl
#12095 added the Animate3D RPC to backend.proto, but the kokoros
service never got a matching method. The tonic-generated Backend trait
now requires it, so kokoros fails to build with E0046 whenever the
full backend matrix runs.
Return Unimplemented, as the other unsupported RPCs do.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
* fix(kokoros): fill new Result fields with defaults
backend.proto added a metadata field to Result, so the struct literals
in the kokoros service no longer name every field and fail to compile.
Spread Default::default() into them, so later additive proto fields do
not break the build again.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
* feat(gallery): publish signed OCI fallbacks
Publish both official gallery indexes with their local base configs so
an outage of the HTTP and GitHub sources can fall back to Quay.
Keep artifact signing policies separate from backend image policies,
and expose each moving gallery tag only after its digest is signed.
Assisted-by: Codex:gpt-6
* fix(gallery): confine packaged files to selected roots
Use directory-scoped file access to reject symlink escapes during gallery packaging. Create private bundle files for the publishing runner.
Assisted-by: Codex:GPT-6
---------
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
* feat(system): report per-model DRM VRAM
Expose optional resident device memory for local backend process trees.
Deduplicate DRM clients and omit unsupported or incomplete readings.
Document accounting limits and preserve a measured zero in JSON.
Closes#11970.
Assisted-by: Codex:gpt-6
* fix(system): document trusted procfs reads
Scope G304 annotations to paths built from the fixed procfs root,
integer process IDs, and kernel directory entries. These reads accept
no user-controlled path components.
Assisted-by: Codex:GPT-6 gosec
---------
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Keep each message and reasoning item at its announced output index.
Include the answer in completed responses with reasoning or fallback
function calls, and retain reasoning supplied through backend deltas.
Add regression coverage for stream indices, final output, plain text,
and automatic tool parsing.
Assisted-by: Codex:GPT-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Partial JSON parsing heals a name-only chunk into a tool call. The
stream emits that call with empty arguments and skips later chunks.
Require complete JSON before emitting terminal tool-call events.
Preserve complete calls before an unfinished trailing call, and count
only actual tool calls. Add split-chunk regression tests and docs.
Refs #11635. The non-streaming report remains unconfirmed.
Assisted-by: Codex:GPT-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
A request with narrower allow_patterns hashes to a different CacheKey
than an already-committed broader sibling, so committedResult misses and
materializeLocked re-fetches files the sibling already holds. After the
own-tree reuseMaterializedFile miss, consult committed sibling trees
for the same Source (type+endpoint+repo+revision), re-verify the file
through verifyDownloadedFile (full SHA-256, never size-only), and
hard-link it into the writer's staging snapshot (copy fallback only on
EXDEV). Each file is matched individually against the sibling's
manifest, so a broader request can never inherit a narrower sibling's
gaps as if complete.
The sibling manifest set is loaded and source-matched once per
materialization (files indexed by path) instead of once per staged
file, so a models volume with 20 committed artifacts and a 300-file
snapshot does one manifest pass rather than ~6000 reads and JSON
parses. The sibling-reuse behavior cases live in the package's
registered Ginkgo suite so repository test conventions apply.
Refs #11047
Signed-off-by: supermario_leo <leo.stack@outlook.com>
The legacy NVIDIA device reservation requests utility without compute.
Docker derives driver capabilities from that list, leaving CUDA libraries
unavailable even when monitoring works.
Include compute in the legacy example and clarify the matching docs.
Assisted-by: Codex:GPT-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
The v4.10.0 ace-step and VibeVoice merge jobs started just after their
digest artifacts expired. Keep the small digest artifacts for seven
days so a multi-day release matrix can finish publishing its images.
Assisted-by: Codex:gpt-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Swag cannot resolve json.RawMessage in OpenAIResponse and aborts the daily
schema generation. Set its Swagger type without changing JSON encoding,
and regenerate the checked-in specifications.
Assisted-by: Codex:gpt-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
PTQ1_0 is a Prism-private GGUF type (GGML_TYPE_PTQ1_0 = 143 in the
PrismML llama.cpp fork), so stock llama-cpp cannot load it. Switch to
the bonsai backend like the existing ternary-bonsai-27b entries, and
replace the scraped Qwen3.8 description and icon.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Older Go linkers stamp pure-Go hosts with SDK metadata that disables
modern Metal APIs. Select Go 1.27 for Darwin builds and document the
backend rebuild requirement.
Assisted-by: Codex:gpt-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
The link text still said chatbot-ui, but it now pointed at the examples
repository root. Link the configurations directory, which holds the
example model config files, and describe it as such.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
The entry enables spec_type:draft-mtp, so variant ranking needs the mtp
tag. Replace the scraped model-card description, set the Swift Open
License v1.0 and link the base model repo.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
* ⬆️ Update TheTom/llama-cpp-turboquant
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* fix(turboquant): drop the upstreamed D512 patch
Upstream c0e227c guards D512 declarations, dispatch, and instances with
GGML_USE_HIP. This prevents the CUDA shared-memory overflow that our
patch addressed. The old patch now rejects the guarded source.
Remove the obsolete patch for the pinned a3d5603d revision. The remaining
patch series applies successfully, and the build-target test passes.
Assisted-by: Codex:gpt-6
---------
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Picks up mudler/LocalAGI 3ce0a08 "fix(mcp): re-dial an MCP session the
server has dropped". Agents open their MCP sessions once, when they are
created; when the MCP server restarts it forgets them and the go-sdk
client does not reconnect by itself, so the agent kept a dead session -
or, behind a server that revives unknown session IDs, a stale tool
list - until LocalAI restarted.
Only go.mod/go.sum change; core/services/agentpool builds and vets
against the new version.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
Code scanning reports alerts on every line of a touched file. Mark the
gRPC auth env var name and the mock backend's staged-path reads as
reviewed, and log the error when evicting a model after its connection
fails instead of dropping it.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Add IQ4_XS and Q4_K_M GGUF builds with a BF16 vision projector.
Pin verified artifacts and document installation and variant selection.
Assisted-by: Codex:gpt-6
Review asked for constants instead of the stage literals ("vad",
"transcription", "llm", "tts", "sound_detection") passed to resolveStage,
stageCall and isChainStage, so the uses can be cross-checked. Add
PipelineStage* constants next to the Pipeline type in core/config: the
names match its yaml keys, and core/backend (preload roles) and the openai
realtime endpoint both need them.
Use them in realtime_model.go (stage routing and preload roles),
realtime.go (the voice_recognition preload role) and core/backend
preload.go. model_failover events take their stage from the stageChains
keys, so they now carry the constants too. The failover tests use the
constants for inputs and keep literal wire values in their event
assertions.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Review asked for constants instead of the "cloud-proxy" and
"localai-proxy" literals. Add CloudProxyBackend and LocalAIProxyBackend next
to the other backend-name constants in pkg/model (WhisperBackend,
TransformersBackend, ...), which core/config already imports, and use them
in every production check: the proxy options builder, IsRemoteProxy,
IsCloudProxyBackendPassthrough, the PII defaults, the localai-proxy backend
hook and loader warning, and the PII middleware metadata in the routes and
the in-process MCP client.
The proxy options builder now calls IsRemoteProxy() instead of repeating
the two-backend check, so the set of proxy backends is defined once.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Code scanning flagged seven issues in the new backend:
- G115 text.go: tool-call indexes and tokenize lengths come from the
upstream server as int and were cast straight to int32. Add clampInt32 so
an absurd upstream value saturates instead of wrapping.
- G115 live.go: the int16 -> uint16 cast in PCM16 encoding is a deliberate
two's-complement reinterpretation of an already clamped sample; mark it
with #nosec and say so.
- G304 proxy.go, media.go, client.go: api_key_file comes from the model
config, and the media input and output paths are files core staged or
chose for the call. None are caller-supplied. Clean the paths and add
#nosec with that reason, as core/gallery and the sound classification
endpoint already do.
- G306 media.go: write generated media 0o600. Core runs as the same user
and serves the file itself.
gosec reports 0 issues for backend/go/localai-proxy and
core/services/failover. The G104 once reported for failover/prober.go is
no longer present.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>