CachyLLaMA splits persistent prompt-cache support into separate source files. Include those implementations in the monolithic LocalAI gRPC adapter when the fork provides them so the ARM64 fallback build resolves the page-manager symbols.
Assisted-by: Codex:gpt-5 [Codex]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Avoid the CPU_ALL_VARIANTS SME matrix on ARM, where the Linux and Darwin backend toolchains fail to compile CachyLLaMA. Keep the existing variant build for x86 and use the fully linked fallback build for ARM.
Assisted-by: Codex:gpt-5 [systematic-debugging]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Add CachyLLaMA as a GGUF-compatible llama.cpp fork backend with CPU and Vulkan builds on Linux and Metal on Apple silicon. Expose it through model import, document persistent SSD cache options, and wire CI path filtering and dependency updates.
Assisted-by: Codex:gpt-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* sglang backend: pass through thinking_budget + require_reasoning
sglang's raw Engine.async_generate() API (which this backend calls
directly, bypassing sglang's own OpenAI server) supports a precise,
tokenizer-derived reasoning-length budget via
sampling_params["custom_params"]["thinking_budget"] plus
require_reasoning=True, gated behind --enable-strict-thinking. Neither
was reachable through LocalAI: this backend built sampling_params only
from a fixed field mapping (temperature, top_p, ...) with no custom_params
key, and never passed require_reasoning to async_generate at all.
- LoadModel now reads a model-level "thinking_budget" option (same
mechanism as the existing tool_parser/reasoning_parser options), and
_build_sampling_params adds it as custom_params.thinking_budget on
every request when configured.
- _new_reasoning_parser already derives, from the rendered prompt, whether
the model's chat template pre-opened a reasoning block (Qwen3-style
templates append <think> to the prompt instead of letting the model
emit it) -- the same signal sglang's own OpenAI server computes from
per-template config to decide require_reasoning. This backend has no
template manager, so it now returns that signal too and _predict
forwards it to async_generate(require_reasoning=...).
Verified against production (NVFP4, sm_121, Qwen3.6-35B-A3B) via a raw
Engine.async_generate() call bypassing this backend: 301 reasoning
tokens against a 300-token budget, clean completion, ~27s. Not yet
verified through this backend's own gRPC path end-to-end (no local
CUDA/sglang environment available here) -- existing + new unit tests in
test.py cover the pure-Python merge/passthrough logic only.
Scope note: require_reasoning is derived only from the existing
prompt-suffix heuristic, not sglang's full per-template
_get_reasoning_from_request decision tree (minimax-m3/hunyuan special
cases etc.) -- this backend has no template manager to evaluate that
tree against, and the prompt-suffix check is the one heuristic already
validated in this file (test_reasoning_parser_forced_when_template_prefills_think_tag).
Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>
* sglang backend: honour a model-level reasoning_default
A model YAML can already carry "parameters: reasoning_effort:", but that
value only reaches this backend when a *caller* sets it per request (the Go
side turns it into Metadata["enable_thinking"]). As a model-level default it
is silently dropped: a config reading "reasoning_effort: none" still produces
full reasoning on every request, so the config says one thing and the model
does another.
That gap is expensive in practice. On a self-hosted Qwen3.6-35B-A3B the
reasoning phase consumed the entire max_tokens budget before any content was
produced - 90% of code completions came back empty at max_tokens=768, and the
server log filled with "backend produced only reasoning, retrying". The
config looked like reasoning was off the whole time.
This adds "reasoning_default:off" (or ":on") on the same model-level
options: mechanism as thinking_budget. A per-request value always wins; the
default only fills in when the request is silent.
Measured on the stack above (sglang 0.5.20, NVFP4, GB10/sm_121) after
applying it:
default (nothing set) -> 0 chars reasoning, 27 tokens
"reasoning_effort": "none" -> 0 chars reasoning, 27 tokens
metadata enable_thinking=true -> capped at the 512-token thinking_budget,
541 tokens total, finish_reason stop
Tests: three cases added to backend/python/sglang/test.py covering the
default, per-request override in both directions, and the unconfigured case
(which must leave the template untouched).
Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>
* sglang backend: validate thinking_budget instead of crashing LoadModel
Addresses the review on this PR:
- `int(thinking_budget)` raised on values like "5000.0" or "abc" and took
LoadModel down. The option is now parsed by _parse_thinking_budget():
integral numbers in any spelling are accepted, anything else is ignored
with a warning on stderr.
- Zero and negative budgets are ignored with a warning instead of being
passed to sglang, where they have no defined meaning. Turning reasoning
off is what reasoning_default:off is for.
- A load-time warning when thinking_budget is set but enable_strict_thinking
is not in engine_args, since sglang then ignores the budget silently.
- Tests for integral spellings, unset, zero, negative, non-integer and the
strict-thinking warning.
Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>
* docs(sglang): explain reasoning options
Document the reasoning budget, strict-thinking requirement, and
precedence of request metadata over the model-level default.
Also note that the budget has to stay well below max_tokens (otherwise
it never triggers and the reply can end up empty), and that
POST /models/reload or a backend-only restart does not pick up changed
options; LocalAI itself has to be restarted.
Assisted-by: Codex:GPT-6
Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>
* docs(sglang): clarify configuration reloads
Distinguish rereading model configuration from updating a running backend. Keep the full LocalAI restart recommendation for changed reasoning options.
Assisted-by: Codex:GPT-6
* sglang backend: only pass require_reasoning when sglang supports it
Engine.async_generate() gained the require_reasoning keyword in sglang
0.5.13 and takes no **kwargs. The CPU profile builds v0.5.11 from source
and the other profiles only set a >=0.5.11 floor, so passing the keyword
unconditionally made every request fail with TypeError. Detect support
once at import time, as the file already does for sampling_seed.
enable_strict_thinking first appears in sglang 0.5.12; fix the comment.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
---------
Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <localai-org-maint-bot@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
The quickstart compose file still requested phi-2, which is no longer in
the gallery. Use phi-2-chat instead and fix model preload error wrapping
so discover/install failures report the real error instead of %!w(<nil>).
Keep earlier model failures when discovery fails for another model.
Document the Compose gallery default.
Fixes#11974
Signed-off-by: lei_lei <imleilei123@gmail.com>
* fix(models): fallback to application config default context size (#12202)
Honor appConfig.ContextSize in /v1/models/capabilities when model context_size is unset.
* docs(models): explain context size fallback
Describe the application default used by capability discovery and
preserve the distinction between total context and per-request limits.
Assisted-by: Codex:GPT-6
* fix(models): apply the default context size only when context_size is unset
The request path applies the application default context size only
when a model leaves context_size unset. An explicit 0 or -1 falls
through to the backend fallback. The capabilities endpoint now does
the same, so it reports the value the backend uses.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
* fix(functions): honor function_arguments_key when building the tool grammar
All four call sites of `Functions.ToJSONStructure(name, args string)` pass
`FunctionsConfig.FunctionNameKey` as *both* arguments, so
`FunctionArgumentsKey` never reaches the grammar generator.
`ToJSONStructure` writes the two properties into the same map:
property[nameKey] = FunctionName{Const: function.Name}
property[argsKey] = Argument{...}
When `nameKey == argsKey` the second assignment overwrites the first, so a
model configured with `function_name_key` gets a grammar carrying only the
arguments object -- the `{"const": "<function name>"}` constraint is gone and
the grammar can no longer express which function was called.
With `function_name_key: function`, the generated property set collapses from
{"function": {"const": "get_weather"}, "arguments": {...}}
to
{"function": {"type": "object", "properties": {...}}}
Setting only `function_arguments_key` is equally broken in the other
direction: the grammar keeps emitting `arguments` while `ParseFunctionCall`
(pkg/functions/parse.go) looks up the configured key, so the parsed call comes
back with its arguments empty.
The default configuration is unaffected -- with both keys empty
`ToJSONStructure` falls back to `name`/`arguments` for both parameters, which
is why this went unnoticed.
The existing `ToJSONStructure()` unit test already calls the helper with two
distinct keys, so only the call sites were wrong. Extend that test with a case
that keeps both custom keys distinct and asserts the two properties survive.
Signed-off-by: Anai-Guo <antai12232931@outlook.com>
* test(functions): cover configured grammar keys
Route grammar construction through FunctionsConfig so the regression test
covers the key wiring used by every endpoint.
Assisted-by: Codex:gpt-5
* chore: empty commit to trigger workflow approval
Signed-off-by: Tai An <antai12232931@outlook.com>
---------
Signed-off-by: Anai-Guo <antai12232931@outlook.com>
Signed-off-by: Tai An <antai12232931@outlook.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
fix(router): re-seed the knn corpus index when the vector store comes back empty
The corpus manager records a store as synced by file fingerprint and
embedding fingerprint. The local-store backend behind it is an in-memory
gRPC process the model loader may evict (active-backend cap, memory
pressure) or the idle watchdog may kill, and relaunch on the next
request — empty. The file is unchanged, so EnsureLoaded returned early
and the router went blind: every probe fell back with similarity 0 while
corpus/stats kept reporting the full count.
Measured on a production router (LOCALAI_MAX_ACTIVE_BACKENDS=6, four
resident models + two router stores): loading any further backend
evicted a store, and the idle watchdog killed both after 15 minutes;
/stores/find returned 0 hits against a 100-line corpus file whose stored
vectors matched fresh embeddings with cosine 1.000.
Two parts, because the knn classifier is built once and cached
(GetOrBuildClassifier), so the sync at build time is otherwise the only
one for the process lifetime:
- corpus.Manager remembers one vector it inserted (probe) and, on the
synced path, asks the live index for it. A miss means the index was
relaunched — fall through and re-seed from the file (no re-embedding).
- The router middleware wraps the knn classifier's store so every
lookup runs EnsureLoaded first; the loader gets the raw store, so its
probe never re-enters the wrapper. A sync error fails the lookup
closed, like the build-time load.
Specs: corpus package (relaunched empty store is re-seeded under an
unchanged file), middleware (relaunched index behind the cached
classifier is re-seeded instead of falling back; the spec is red without
the wrapper). The test fake now answers Search for inserted vectors.
Folds in the maintainer's follow-up (router-corpus-reseed-after-store-relaunch): reviewed and accepted.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
* feat: return 429 when backends are saturated
When backends are at capacity (per-model max_concurrent or the
process-wide --max-concurrent-backend-requests ceiling), the response
was 503. The OpenAI SDK, litellm, and most agent harnesses key on 429
for rate-limit backoff and treat 503 as a hard error.
Both saturation paths now return 429 with the existing Retry-After
header and type: "rate_limit_error" in the JSON body. The per-model
admission middleware keeps admission_rejected as the code field so
existing alerts that match on it still fire.
Non-saturation 503s are unchanged: model cold-loading (with progress
body), model-load failure cooldown, PII detector fail-closed, and
classifier unavailable. These mean "not ready" rather than "busy".
Assisted-by: AGENT:regolo/glm5.2 [TOOL]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix: return 503 when scheduler has no available nodes
When the scheduler cannot find any healthy node to serve a model —
all nodes are full and eviction cannot free a slot, or a node_selector
excludes every candidate — the error fell through to 500. A 500 tells
clients something is broken when the condition is transient and
retryable.
The router now wraps these errors with a new ErrNoAvailableNodes
sentinel. The HTTP error handler maps it to 503 via applyNoAvailableNodes,
following the same pattern as applyBackendAdmission (429). Unrelated
scheduler errors (DB timeouts, registry lookups) still return 500.
Three return sites are wrapped:
- resolveSelectorCandidates: selector matches zero healthy nodes
- scheduleNewModel eviction-busy: all models have in-flight requests
- scheduleNewModel eviction-failed: eviction itself errored
The existing scheduleAndLoad wrapper ("no available nodes: %w") preserves
the sentinel through the chain via errors.Is, as does ModelRouterAdapter.
Assisted-by: AGENT:regolo/glm5.2 [TOOL]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* test(http): use Ginkgo for admission tests
Replace forbidden testing.T calls with Ginkgo and Gomega so the lint
check accepts the admission handler tests.
Assisted-by: Codex:GPT-6 forbidigo
* fix(middleware): show the recorded status for admission rejections
The admission audit row now records 429, but the Middleware page still
printed a hard-coded 503. Read the status from the event, and update
the two package comments that still said 503.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
* fix(quantization): pin the producing backend on imported quantized models (#11875)
ImportModel hands the copied GGUF to importers.ImportLocalPath, which detects
the file format and defaults every GGUF to `backend: llama-cpp`. For a model
this service just produced with a backend stock llama.cpp cannot read, the
generated config names an engine that cannot load the file, and the import
silently registers an unloadable model. Correcting `backend:` by hand makes the
same file work.
The job record already carries the backend that served StartQuantization, so
carry it into the config instead of keeping the detected default. The gallery
publishes a quantizer as a release channel of the engine that runs its output
("llama-cpp-quantization" is llama.cpp's quantizer, whose GGUF is served by
"llama-cpp"), so the channel suffix is stripped to get the serving backend.
A backend that both quantizes and serves ("rocmfp4") carries no suffix and
passes through unchanged, as do pinned hardware variants ("rocm-rocmfp4"),
which are valid values for a config's backend field. An empty job backend
leaves the detected default in place.
Also replace the importer's generic "Fine-tuned model (GGUF)" description for
this path: the model was quantized, not fine-tuned, and the job knows the type.
Signed-off-by: Tai An <antai12232931@outlook.com>
* style: restore trailing newline in service.go for gofmt
Signed-off-by: Anai Guo <antai12232931@outlook.com>
* style(quantization): restore trailing newline in service.go
gofmt requires the file to end with a newline; the previous style commit
did not actually add it.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
---------
Signed-off-by: Tai An <antai12232931@outlook.com>
Signed-off-by: Anai Guo <antai12232931@outlook.com>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Hardcoding size/size_vram as 0 made Ollama clients treat loaded models as
free. Prefer ModelFileName+ModelPath Stat when available, and omit size_vram
(and size) when the value is unknown instead of emitting literal zeros.
Resolve each listed model by its stored ID so tagged variants use their
own weights.
Fixes#11969
Signed-off-by: lei_lei <imleilei123@gmail.com>
LocalAGI 7e0947d added allowed_tools/excluded_tools and the
required_tool_before_finish gate. Single-node agents get them through
LocalAGI's runtime, but the distributed executor drives cogito directly
and its static config meta did not list the fields, so the agent form
hid them and the worker ignored them.
The distributed config now parses the tool lists from a JSON array or a
comma/newline separated string, and the meta entries match LocalAGI's.
The executor filters the knowledge base, skill and MCP tools (MCP via
cogito.WithMCPToolFilter) before the model sees them, and re-prompts the
model when it answers before the required tool returned "ok": true, up
to the configured number of reminders.
LocalAGI keeps its filter and gate helpers unexported, so a minimal copy
lives in core/services/agents/toolpolicy.go. A spec compares the meta
entries with LocalAGI's to catch drift.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Pick up the LocalAGI PRs merged after f2a2af4:
- per-collection embedding and reranker models, locked per collection
so one agent's upload or rerank no longer stalls the others (#499)
- required_tool_before_finish: a tool the agent must call successfully
before it may answer (#495)
- allowed_tools / excluded_tools per agent, applied to MCP tools too
(#480)
Document the new agent settings. They show up in the single-node agent
form, which reads LocalAGI's config metadata; distributed mode keeps its
own field list and does not offer them yet.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Pick up the LocalAGI fixes merged since 985e5f5:
- MCP: a failed ping no longer closes borrowed sessions (#500)
- agent: nil user message guard in knowledgeBaseLookup (#485), 10m
fallback on an unparsable periodic_runs (#497), self-correction on
unknown tool calls (#481)
- core: idempotent JobResult.Finish (#491)
- sse: ordered broadcast and race-free message history (#496)
- mautrix 0.25.2 (#328), which moves several indirect deps
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
An agent with long_term_memory or summary_long_term_memory enabled but
no knowledge base crashed with a nil pointer dereference when it saved
the conversation (#11975). LocalAGI now logs a warning and skips the
write in that case (mudler/LocalAGI#501).
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
* fix(kokoros): add missing animate3_d stub to Backend trait impl
#12095 added the Animate3D RPC to backend.proto, but the kokoros
service never got a matching method. The tonic-generated Backend trait
now requires it, so kokoros fails to build with E0046 whenever the
full backend matrix runs.
Return Unimplemented, as the other unsupported RPCs do.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
* fix(kokoros): fill new Result fields with defaults
backend.proto added a metadata field to Result, so the struct literals
in the kokoros service no longer name every field and fail to compile.
Spread Default::default() into them, so later additive proto fields do
not break the build again.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
* feat(gallery): publish signed OCI fallbacks
Publish both official gallery indexes with their local base configs so
an outage of the HTTP and GitHub sources can fall back to Quay.
Keep artifact signing policies separate from backend image policies,
and expose each moving gallery tag only after its digest is signed.
Assisted-by: Codex:gpt-6
* fix(gallery): confine packaged files to selected roots
Use directory-scoped file access to reject symlink escapes during gallery packaging. Create private bundle files for the publishing runner.
Assisted-by: Codex:GPT-6
---------
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
* feat(system): report per-model DRM VRAM
Expose optional resident device memory for local backend process trees.
Deduplicate DRM clients and omit unsupported or incomplete readings.
Document accounting limits and preserve a measured zero in JSON.
Closes#11970.
Assisted-by: Codex:gpt-6
* fix(system): document trusted procfs reads
Scope G304 annotations to paths built from the fixed procfs root,
integer process IDs, and kernel directory entries. These reads accept
no user-controlled path components.
Assisted-by: Codex:GPT-6 gosec
---------
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Keep each message and reasoning item at its announced output index.
Include the answer in completed responses with reasoning or fallback
function calls, and retain reasoning supplied through backend deltas.
Add regression coverage for stream indices, final output, plain text,
and automatic tool parsing.
Assisted-by: Codex:GPT-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Partial JSON parsing heals a name-only chunk into a tool call. The
stream emits that call with empty arguments and skips later chunks.
Require complete JSON before emitting terminal tool-call events.
Preserve complete calls before an unfinished trailing call, and count
only actual tool calls. Add split-chunk regression tests and docs.
Refs #11635. The non-streaming report remains unconfirmed.
Assisted-by: Codex:GPT-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
A request with narrower allow_patterns hashes to a different CacheKey
than an already-committed broader sibling, so committedResult misses and
materializeLocked re-fetches files the sibling already holds. After the
own-tree reuseMaterializedFile miss, consult committed sibling trees
for the same Source (type+endpoint+repo+revision), re-verify the file
through verifyDownloadedFile (full SHA-256, never size-only), and
hard-link it into the writer's staging snapshot (copy fallback only on
EXDEV). Each file is matched individually against the sibling's
manifest, so a broader request can never inherit a narrower sibling's
gaps as if complete.
The sibling manifest set is loaded and source-matched once per
materialization (files indexed by path) instead of once per staged
file, so a models volume with 20 committed artifacts and a 300-file
snapshot does one manifest pass rather than ~6000 reads and JSON
parses. The sibling-reuse behavior cases live in the package's
registered Ginkgo suite so repository test conventions apply.
Refs #11047
Signed-off-by: supermario_leo <leo.stack@outlook.com>
The legacy NVIDIA device reservation requests utility without compute.
Docker derives driver capabilities from that list, leaving CUDA libraries
unavailable even when monitoring works.
Include compute in the legacy example and clarify the matching docs.
Assisted-by: Codex:GPT-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
The v4.10.0 ace-step and VibeVoice merge jobs started just after their
digest artifacts expired. Keep the small digest artifacts for seven
days so a multi-day release matrix can finish publishing its images.
Assisted-by: Codex:gpt-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Swag cannot resolve json.RawMessage in OpenAIResponse and aborts the daily
schema generation. Set its Swagger type without changing JSON encoding,
and regenerate the checked-in specifications.
Assisted-by: Codex:gpt-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
PTQ1_0 is a Prism-private GGUF type (GGML_TYPE_PTQ1_0 = 143 in the
PrismML llama.cpp fork), so stock llama-cpp cannot load it. Switch to
the bonsai backend like the existing ternary-bonsai-27b entries, and
replace the scraped Qwen3.8 description and icon.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Older Go linkers stamp pure-Go hosts with SDK metadata that disables
modern Metal APIs. Select Go 1.27 for Darwin builds and document the
backend rebuild requirement.
Assisted-by: Codex:gpt-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
The link text still said chatbot-ui, but it now pointed at the examples
repository root. Link the configurations directory, which holds the
example model config files, and describe it as such.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
The entry enables spec_type:draft-mtp, so variant ranking needs the mtp
tag. Replace the scraped model-card description, set the Swift Open
License v1.0 and link the base model repo.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
* ⬆️ Update TheTom/llama-cpp-turboquant
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* fix(turboquant): drop the upstreamed D512 patch
Upstream c0e227c guards D512 declarations, dispatch, and instances with
GGML_USE_HIP. This prevents the CUDA shared-memory overflow that our
patch addressed. The old patch now rejects the guarded source.
Remove the obsolete patch for the pinned a3d5603d revision. The remaining
patch series applies successfully, and the build-target test passes.
Assisted-by: Codex:gpt-6
---------
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>