mirror of
https://github.com/mudler/LocalAI.git
synced 2026-09-13 06:45:26 -04:00
51f906f4e03f2790f657d66ff6a2742e9ff07429
21
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8f56e4e042 |
fix(vram): persist remote probe metadata (#11487)
* fix(vram): persist remote probe metadata The startup warmer repeated remote size and GGUF metadata probes after every restart because both caches lived only in memory. Store successful HTTP probes for 24 hours so frequent restarts reuse the prior results. Bound the cache, reject invalid records, and purge it when gallery data changes. Local model files continue to bypass persistence. Assisted-by: Codex:gpt-5 * fix(vram): check temporary file cleanup The lint gate rejects the unchecked cleanup call in the persistent cache writer. Assisted-by: Codex:gpt-5.6 [golangci-lint] * fix(vram): make persistent cache optional Remote metadata probes can transfer enough data that operators need control over disk reuse and startup warming. Gallery autoload now gates both behaviors, and the runtime setting applies changes immediately. Assisted-by: Codex:gpt-5 * fix(ui): expose gallery startup pre-warm The existing gallery autoload setting also gates the startup metadata warmer. Name both effects in Settings so operators can find the requested boot control. Assisted-by: Codex:gpt-5 --------- Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
799cc9f211 |
feat: bound global admission and expose running backend traces (#11560)
feat: bound backend admission and expose running traces Add process-wide backend execution admission without blocking UI or administrative HTTP work. Represent backend operations while they are in flight, surface running traces with immediate log links, and tie streaming admission leases to the gRPC receive lifecycle. Assisted-by: OpenAI Codex: GPT-5 Signed-off-by: Richard Palethorpe <io@richiejp.com> |
||
|
|
ab52813342 |
feat(modelartifacts): support bounded parallel Hugging Face file downloads (#11162)
* feat(modelartifacts): support bounded parallel Hugging Face file downloads Closes #11114. Snapshot materialization fetched every file through the sequential executor in DownloadFilesWithContext, so a repository split into many shards spent most of its wall clock in per-file request latency rather than moving bytes. Add DownloadFilesWithConcurrency, an errgroup with SetLimit, and keep DownloadFilesWithContext as a wrapper that passes a limit of 1. That leaves the two non-artifact callers (core/gallery and the model config loader) on exactly the path they had: tasks still run in slice order, and the first failure still returns before any later task starts. Only whole files run in parallel. A single file is never split, so the .partial resume machinery and the per-file SHA check in downloadTaskWithRetry are untouched. Two details the parallel path forced: - completedBytes becomes an atomic.Int64. Several AfterDownload hooks add to it while other files' progress callbacks read it; without this the race detector reports three races on the new specs. - The caller's status callback is serialized. The sequential path gave it an implicit guarantee of never being entered twice at once, and it belongs to the caller, so the executor keeps that promise rather than pushing locking onto every caller. AfterDownload is deliberately not serialized -- it does the verify-and-promote work that parallelism exists to overlap. Manifest order needed no work: each hook already writes its own manifest.Files slot by snapshot index, so entries stay in snapshot order whatever the completion order. A spec now pins that. The default is 1, unchanged behaviour. A shared models volume is often the bottleneck rather than the link, so raising it is a deployment decision; --artifact-download-concurrency and LOCALAI_ARTIFACT_DOWNLOAD_CONCURRENCY expose it on both `run` and `models install`. Not done here, per the issue: no chunk-level parallelism within a single file, and no throughput measurements across concurrency 1/2/4/8 -- that needs a representative sharded repo and a real link. Assisted-by: Claude:claude-opus-5 go-test gofmt Signed-off-by: Adira Denis Muhando <dennisadira@gmail.com> * feat(modelartifacts): expose download concurrency in settings Follow-up to review feedback on #11162: - The CLI flag and docs no longer describe the limit as Hugging Face specific. It applies to any artifact source, as @mudler pointed out. - artifact_download_concurrency is now a persisted runtime setting and is editable from the WebUI, so it can be changed without a restart. The manager's limit becomes an atomic.Int64 behind SetDownloadConcurrency, because a live runtime setting can be updated while a materialization is already in flight. Injected materializers stay compatible through an optional setter interface, so a manager that does not implement it is simply left alone. Verified before taking this on: go build, go vet and go test -race all pass for pkg/modelartifacts, pkg/downloader and core/config. The React UI builds with vite, artifact_download_concurrency is present in the built Settings chunk, and eslint reports the same 8 pre-existing warnings on Settings.jsx as it does without the change. Implementation contributed by localai-org-maint-bot on the review thread; reviewed, verified and signed off by me. Assisted-by: Codex:gpt-5 Assisted-by: Claude:claude-opus-5 go-test vite eslint Signed-off-by: Adira Denis Muhando <dennisadira@gmail.com> --------- Signed-off-by: Adira Denis Muhando <dennisadira@gmail.com> Co-authored-by: localai-org-maint-bot <bot-opensource@localaisrl.com> |
||
|
|
d0119bf62c |
feat(chat): local-ai chat is now a terminal agent (#11291)
* chore(deps): bump cogito to v0.11 ahead of the nib harness nib is the agent harness that becomes 'local-ai chat'. It requires cogito v0.11, so pull that bump forward on its own: minimal version selection would apply it to LocalAI anyway, and both repos use cogito and cogito/clients. Landing it separately keeps the harness change reviewable. nib itself is not pinned yet. Nothing in LocalAI imports it, and 'go mod tidy' runs as a goreleaser before-hook in CI, so an unimported require line does not survive. It lands with its first importer. No LocalAI call site needed a change. Both cogito.WithMaxAttempts callers guard the argument above zero, so v0.11's new clamp is unreachable, and LocalAI's Multimedia values implement only URL(), so v0.11's new TypedMultimedia routing treats them as images exactly as v0.10 did. Binary size (cmd/local-ai): 200,301,381 -> 200,336,045 bytes (+34,664). A throwaway probe that links nib measured 201,042,243 bytes (+740,862 over the pre-change baseline). Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(chat): resolve and seed the agent state directory Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(chat): write the agent config atomically and tighten its modes Replacing config.yaml in place truncated it first, so an interrupted write would have destroyed the api_key nib keeps in the same file. Stage through a sibling temp file and rename over the target instead, and match nib's 0700 directory mode. Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(chat): probe the endpoint and classify failures Probe lists what a LocalAI endpoint advertises and separates the two failures that need different advice: nothing listening, and rejected credentials. go-openai reports a rejected key as one of two concrete types depending on the error body, and both occur against a real LocalAI. The normal error handler sends an OpenAI error envelope, which arrives as *openai.APIError; the opaque-errors handler replies with a bare status and no body, which arrives as *openai.RequestError. Classifying on only one of them misses half the cases, so the status is read from either. A cancelled probe is not reported as an unreachable server, because it learned nothing about the endpoint, and neither is a reply that could not be parsed, because something did answer. Both would otherwise send the user off to start a server that may already be running. The model list is returned verbatim and in server order. LocalAI lists whatever it finds in the models directory, including stray archives and dotfiles, and deciding which advertised ids are real belongs to whoever presents them. Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(chat): resolve the model from flag, config, or the server Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * test(chat): pin that model resolution sorts a copy of the caller's slice The sort spec asserted only on what the chooser was offered, so replacing the defensive copy with an in-place sort of req.Available still passed all 37 specs. Assert the input slice's order after the call, so the guarantee cannot be dropped silently by a later refactor. Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(chat): offer to start a server when none is reachable Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(chat): bound the server wait and pin readiness and stop semantics Set cmd.WaitDelay so a backend subprocess holding the child's stderr pipe cannot block cmd.Wait forever, which would leave exited unclosed, burn the whole shutdown grace on a clean exit, and leak the waiter goroutine. Two test gaps closed alongside it: the readiness spec now counts polls, so treating 503 as ready is observable, and Stop's single-interrupt contract is pinned by giving StartedServer interrupt/kill hooks that a spec can count. Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(chat): drive Stop through one process interface, hide exec plumbing Two independent interrupt/kill func fields plus a nil check admitted wirings no test could distinguish: the pair swapped, so a SIGKILL would strand the backends SIGINT exists to let local-ai run clean up, or kill left nil, so a wedged server never escalates. One two-method interface that *os.Process already satisfies leaves nothing to swap and nothing to nil. Also translate exec.ErrWaitDelay, whose text names an os/exec struct field, into what the user can act on. os/exec only substitutes that sentinel when the process exited without an error of its own, so no exit status is swallowed. Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(chat): replace the REPL with the built-in agent local-ai chat is now the nib agent harness compiled into the binary: tool use behind an approval gate, sub-agents, MCP, plugins, and skills, all auto-configured against the local server. The REPL goes with it. Its model listing and its 401 classifier were duplicates of the ones Probe now owns, and the classifier was the version that misreads a bare 401 with no OpenAI error envelope, so keeping either would leave the package with two divergent answers to the same question. github.com/mudler/nib lands in go.mod in this commit rather than earlier: go mod tidy runs as a goreleaser before-hook on every PR, so a require line with no importer is stripped before it reaches CI. Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(chat): split the pre-agent phase out of Run and pin it Everything before the handoff is testable and nothing after it is: once app.Run owns the terminal there is no seam left. prepare draws that line, takes interactivity as a parameter so the prompts can be driven over a pipe, and hands Run the state dir, the model, and any server it started. The questions move onto one prompter that owns its buffered reader. A fresh bufio.Reader per question reads ahead and discards what it buffered, so the model choice typed behind an answer to "start a server?" was lost and the next question saw EOF. choose answers with a list index and refuses an empty offer, so a value that was never on the list cannot reach ResolveModel, which persists it and starts every later run against it. Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(chat): bound each server check with a deadline Nothing bounded the model listing, so pointing chat at an address that accepts the connection and then never replies left the user with no output and no offer to start a server. The budget is context.WithTimeout rather than a cancel plus a timer. Probe deliberately refuses to call an endpoint unreachable on a context.Canceled, since a caller who gave up learned nothing about the server, and only honours a deadline. A cancel-based budget therefore expires as the one error that suppresses ErrUnreachable, exactly for the hung servers the offer exists to rescue. Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(chat): tell the user when their model choice cannot be saved The choice is meant to be asked for once. When saving it fails the user is silently asked again on the next run, and the only trace was an xlog.Warn: the agent runs at log level error, and a --log-level=error run swallows it entirely. ModelRequest gains Notify for exactly this class of problem, one that is worth telling the user about but not worth failing over, and the chat wiring points it at the same writer the question was asked on. Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(chat): stop a session's server when the process is signalled A server started for the session is stopped by a deferred call, and a signal skips deferred calls: a SIGTERM between the spawn and the exit left 'local-ai run' reparented to init with nothing left that knew to shut it down. Ctrl+C was already safe, but only incidentally, because the child shares this process' foreground process group. A signal handler rather than Pdeathsig on the child. Pdeathsig is Linux-only and, in Go, is delivered when the OS thread that forked exits rather than when the process does, so it can fire on a healthy parent. Setpgid would break the Ctrl+C that works today by taking the child out of the foreground group. SIGHUP joins SIGINT and SIGTERM: a terminal program whose terminal is gone has nobody left to talk to. The same context is what cancels the agent, which nib leaves to its embedder. Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(chat): only skip the server checks for work that stays local Two argument shapes were classified wrongly. Every 'mcp ...' invocation counted as management, so 'local-ai chat mcp --stdio', which serves the agent over MCP and needs a model like any other session, was handed an empty one. And --init, whose shell snippet a user pastes into an rc file long before any server exists, went the other way: it demanded a running server to print a static string. The mcp split is asked of nib's own IsMCPManageSubcommand rather than restated here, so a verb added upstream cannot drift out of this list. Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(chat): exit with the agent's status instead of reporting it twice nib writes what went wrong to stderr and returns nothing but an exit code, so returning that error unchanged had main log "Error running the application error=exit status 1" underneath the message the user had just read. The refusal to render the full-screen interface into a pipe is the one they meet in practice: it names --cli, and burying that hides the fix. ExitCodeError says "already reported, exit with this status". main honours it and prints nothing more, so a piped or redirected chat still fails a script the way it should. Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * style(chat): route interactive chatter through one writer helper The prompts and notices all write to a terminal, where a failed write is not worth failing the session over and the read that follows the question reports the real problem. say says that once instead of five discarded error returns. The command's one-line help comes along: chat is no longer "an interactive chat session". Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs(chat): record why the agent gets this process' streams Injecting them is what makes nib refuse to draw its full-screen interface into a pipe and name --cli, instead of rendering onto a terminal the caller may not own. The tradeoff is worth stating where the wiring is. Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(chat): stop the session's server on cancellation, not on the way out The deferred Stop is reached only if the agent returns, and cancelling the context does not make it: nib hands the TUI to bubbletea without the context, so what actually unwinds a running session today is bubbletea's own SIGINT and SIGTERM handler. SIGHUP has no such backstop, and registering for it removed the default disposition that used to end the process outright, so kill -HUP left a live TUI with a cancelled context and the started server still running. runSession watches the context alongside the agent and stops the server the moment it is cancelled, so the guarantee no longer depends on what the agent does with cancellation. Stop is idempotent, so the deferred call stays correct and free. The doc comment on shutdownContext described the mechanism it was supposed to work by rather than the one that does. Corrected, bubbletea's handler included. ResolveModel now checks the chooser's answer against what it offered. The shipped chooser answers by list index and cannot be wrong, but ModelChooser is exported, the answer is persisted, and every later run starts against it, so the invariant belongs at the consumer. Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * chore(chat): bump nib to v0.5.1 v0.5.1 carries four fixes that matter to 'local-ai chat': - --init now names the embedder's command, so the emitted widget invokes 'local-ai chat' rather than a bare 'nib' the user does not have. - A piped CLI session that succeeds exits 0 instead of failing with EOF. - EOF at a tool-approval prompt denies the call rather than approving it, and the session exits 3 (app.ExitCodeApprovalNoInput) so a script can tell "answered" from "refused to act" without reading stdout. Read-only tools are unaffected and still run. ExitStatus already unwraps app.ExitError, so the code propagates with no change here. - RunTUI passes the context to bubbletea and gives up bubbletea's own signal handler, which makes shutdownContext the single owner of the signal and stops a SIGHUP leaving a wedged TUI behind. Verified against a live server on 127.0.0.1:8080: the three --init shells, a piped prompt exiting 0, a denied 'touch' that left no file and exited 3, a read-only 'ls' that still ran and exited 0, and a SIGHUP that unwound a TUI running under a pty. Two comment blocks in run.go described the old TUI behavior and are now wrong, so they are corrected in the same change. No behavior change: both shutdownContext and runSession are untouched, and stopping the server on cancellation is still worth keeping independent of how promptly nib unwinds. One known gap, not addressed here. The widget --init now emits runs 'output=$(local-ai chat --height 50%)', and runAgent injects Stdout unconditionally, so under $(...) nib refuses the TUI for a non-terminal stream. This is the cost the runAgent comment already anticipated, now that the snippets no longer hardcode standalone nib. Ctrl+Space should not be documented until that is decided. Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(chat): let nib own stdout, so the Ctrl+Space widget works The widget 'local-ai chat --init' emits runs 'output=$(local-ai chat --height 50%)', which puts a pipe on stdout by construction. runAgent injected os.Stdout unconditionally, and nib refuses every mode but --cli when a stream it was handed is not a terminal, so Ctrl+Space printed "Re-run with --cli to use the injected streams" and inserted nothing. Verified against a pty before and after. nib reads a nil stream as "not injected" and falls back to the process stream, which is how an embedder asks for nib's own behavior. That is what stdout needs: the interface renders on /dev/tty but writes the chosen command to stdout even when stdout is a pipe, and that write is the whole of the shell-capture idiom. Stdin is deliberately left injected. A piped or redirected stdin really is ignored by the interface, so the refusal is the honest answer there, and it is the one users meet: 'echo q | local-ai chat' still says to re-run with --cli, once, exit 1. Nilling stdin the way stdout is nilled would delete that silently. Stderr is not gated by nib at all and is unchanged. One case does change and cannot be kept: 'local-ai chat > out.txt' from a terminal no longer refuses, because it is indistinguishable from the widget. It renders on /dev/tty and writes the capture line to the file, which is what standalone nib does. The app.Options literal moves into agentOptions so the decision is reachable from a spec rather than being a detail of a function that takes the terminal. Both sides of the asymmetry are pinned: reinstating 'Stdout: opts.Out' fails "hands nib nothing for the process stdout", and nilling stdin fails "hands the process stdin over". Also rewrites the last comments describing the pre-v0.5.1 behavior. Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs(chat): say what the stream refusal actually keys on Two comments still called it the refusal to render the interface "into a pipe". That was true when both stdin and stdout were injected, but a pipe on stdout no longer refuses, so the wording now points at precisely the case that was un-refused to make Ctrl+Space work. Only a stdin that cannot be read triggers it, and both comments now say so and name the command a user meets it with, 'echo q | local-ai chat'. The agentOptions doc also said a "file a caller chose" stays injected and refused, which reads as though 'local-ai chat > out.txt' still refuses. It does not: a shell redirect arrives as os.Stdout and is nil-ed like the widget's pipe, because the two differ only in being a regular file rather than a FIFO and nib's gate does not look at that. What stays injected is a writer an in-process caller chose for itself. Says that now, in the doc and in the spec comment that had the same ambiguity. Comments only. No behavior change. Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: local-ai chat is now the built-in terminal agent `local-ai chat` was a plain chat prompt and is now an agent that runs shell commands behind an approval gate, so the pages that described a REPL were wrong rather than merely thin. Adds a Terminal agent feature page at /features/terminal-agent covering the approval gate, piped runs and their exit codes, Ctrl+Space, model resolution, state directory, and the pass-through management commands (including the `--yes` caveat that leaves a plugin installed but disabled in a script). The three-way "looking for something else" notice becomes four-way and moves into an agentic-routing shortcode. Four hand-kept copies of the same paragraph is what produced the drift the new page would otherwise have added to; the shortcode takes `current=` so each page still marks itself, and errors the build on a name that is not one of the four. Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * website: the agent is in the binary, not a second install The nib section sold a separate tool you also install, with a GitHub link as the only way in, which is now the wrong order: the agent ships compiled into local-ai, and the standalone binary is the second reason to care rather than the first. Leads with `local-ai chat`, keeps nib as the SSH-anywhere story, and adds a docs CTA pointing at the new Terminal agent page. id="nib" is left alone because localai.io/#nib is linked from outside. The two credits on the demo clip named nib as the thing that drove the machine; they now credit the agent in LocalAI, which is the same agent. Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * website: fix the exit keys, the plugin warning, and the redirect gap Three claims on the chat-agent pages that the code does not back. try-it-out told readers to press Ctrl+D. nib has no Ctrl+D handler: the full-screen interface quits on Esc or Ctrl+C, and Ctrl+D is only an exit in --cli, where it arrives as ordinary tty EOF. That sentence had replaced the removed /exit and /quit text, so the page was left with no working way to leave a session. Document both modes, since they differ. The plugin warning said nothing tells you the install stopped short. It does: the command prints that the plugin was left disabled. What it does not do is say so in its exit code, which is 0 either way. That is the part a script cannot work around, and it is the reason to pass --yes. Overstating it in the paragraph that gives the advice only makes the advice easier to dismiss. Redirecting stdout no longer refuses; the interface goes to /dev/tty and only the yanked command reaches the file. It is what lets the Ctrl+Space widget capture a command at all, since a redirect and out=$(...) are the same thing to the stream gate. It was documented nowhere. A non-terminal stdin is still refused, and the new text says which of the two it is. Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(chat): make the CLI flags outrank the agent config file local-ai chat routed --endpoint, --model, --api-key, --trace-dir and --yolo through nib's app.Options.Defaults. Defaults are seeds: they sit beneath the config file, so the file silently undoes them. That made the flags accepted and inert, and not in an edge case, since EnsureStateDir writes base_url on the first run and the interactive picker writes model, so from the second run on the file carried a value for both. Observed against a live server: with base_url: http://127.0.0.1:9999/v1 in the config and --endpoint http://127.0.0.1:8080 on the command line, the probe hit 8080 and every agent turn posted to 9999. With model: gemma-4-e2b-it-qat-q4_0 in the config, --model lfm2.5-8b-a1b was ignored on the wire. nib v0.6.0 adds app.Options.Overrides, applied above the config file and above the bare environment block. Move the whole block there: all five values are decisions this invocation already made on the user's behalf, and a flag the config file can undo is not a flag. Nothing is left in Defaults, because LocalAI's one genuine seed, the initial base_url, is written into the config file by EnsureStateDir rather than handed to nib. Two limits come with the channel and are documented on agentOptions rather than worked around. An override can only raise a field, since nib cannot tell "set to the zero value" from "not set", so --yolo can turn approval off but nothing on the command line turns it back on over an approval_mode: auto in the file. And nib's own NIB_TRACE_DIR and NIB_YOLO are resolved after the config load and still outrank these, deliberately, upstream. The existing spec pinned that the right values reach app.Options, which they always did, which is exactly why it could not see nib discarding them. The new specs resolve the config the way app.Run resolves it, against a real config file that disagrees with every flag, and one asserts Defaults stays empty. docs/content/features/terminal-agent.md already documented --model as winning over the saved model; that was false before this change and is true now, so no docs edit was needed. Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(chat): document intentional config file read Assisted-by: Codex:gpt-5 [gosec] --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
ff299df453 |
perf(http): gzip responses, cache hashed assets, bound the trace endpoints (#11056)
Three measured HTTP-layer regressions on a live deployment, fixed together
because they all shape the bytes on the wire.
1. No compression. The server sent no Content-Encoding regardless of what
the client asked for, confirmed with curl straight at 127.0.0.1:8080 so
it was not an ingress artefact. Adds gzip middleware, on by default and
configurable via LOCALAI_DISABLE_HTTP_COMPRESSION and
LOCALAI_HTTP_COMPRESSION_MIN_LENGTH (default 1024 bytes so tiny bodies
are not wastefully wrapped). Streaming routes are skipped explicitly:
an SSE Accept header, a WebSocket upgrade, and the completion / SSE /
log-tail path prefixes, because whether a completion request streams is
decided by the request body, which the middleware runs too early to see.
Already-compressed formats (woff2, png, mp4, ...) are skipped too; gzip
made those marginally larger. Measured over the embedded React build:
JS+CSS 2815 KB raw to 808 KB gzipped (3.48x).
2. No cache headers on content-hashed assets. Vite hashes the filenames,
so a given /assets/ URL can never change content, yet they shipped with
no Cache-Control, ETag or Last-Modified, and the browser re-fetched the
whole bundle on every navigation with no conditional request available.
/assets/* now carries public, max-age=31536000, immutable. index.html
stays no-cache so a deploy is picked up, and the unhashed locale JSONs
get a short TTL rather than the immutable one.
3. Unbounded trace endpoints. /api/traces returned 21,033,606 bytes in
4.65s and /api/backend-traces 3,471,682 bytes in 1.50s, and the admin
UI polls both every few seconds. The ring buffer holds up to 1024
entries, each embedding full input_text payloads. Both list endpoints
now take limit / offset / full, default to 50 entries, and strip the
heavy fields (request and response bodies plus headers for API traces,
body and data for backend traces) unless full=true. Every trace gets a
process-lifetime ID and GET /api/traces/{id} and
/api/backend-traces/{id} serve the full record, which is what the UI
fetches when a row is expanded. The list body stays a JSON array;
paging metadata rides in X-Total-Count, X-Trace-Offset and
X-Trace-Limit. Reproducing the live shape in a test, the polled payload
goes from 21,131,097 bytes to 7,201 bytes.
Assisted-by: Claude:claude-opus-4-8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
|
||
|
|
40d35c0385 |
docs: onboarding overhaul, dedup, and error docs (#7711) (#10895)
* docs: fix CPU image tag (latest, not latest-cpu) Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: use canonical localai/localai registry in models guide Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: replace dead llama-stable backend with llama-cpp Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: correct mitm-proxy intercept config and redaction tier Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: fix text-to-audio endpoint and broken notice block Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: fix VAD example, stale FAQ, broken link, CLI list, whats-new dump Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: render advanced/reference section indexes (consolidate _index) Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: remove duplicate getting-started build/kubernetes pages Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: fold container image reference into installation/containers Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: remove stale advanced fine-tuning page (superseded by features/fine-tuning) Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: fold distribution/longcat/sound pages into their parents Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: make getting-started index accurate and complete Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: carry one concrete model through the getting-started path Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: add end-to-end 'build your first agent' walkthrough Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: add runtime errors reference; consolidate troubleshooting from FAQ Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: add agent actions catalog Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: agent-scoped MCP, skills walkthrough, agentic disambiguation Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: add concrete gallery install lines to media feature pages Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: merge installation into getting-started (URLs preserved via aliases) Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: add Operations section; move operator pages and P2P API reference Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: journey-ordered top nav and grouped feature sections Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: add docs-with-code process gate (PR template + agent instructions) Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: remove em/en dashes from documentation prose Signed-off-by: Ettore Di Giacinto <mudler@localai.io> --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
5569b2de56 |
feat(config): context_size: -1 to auto-use model's full trained context (#10752)
* feat(config): clamp negative context_size to default in EffectiveContextSize Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-4-8 [Claude Code] * feat(config): resolve context_size=-1 to model trained max with VRAM warn Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-4-8 [Claude Code] * fix(config): treat negative context_size as unset when GGUF is unparseable Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-4-8 [Claude Code] * docs(config): document context_size=-1 auto-max Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-4-8 [Claude Code] * docs(backend): drop em dashes from EffectiveContextSize comment Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-4-8 [Claude Code] --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
97175f4b5a |
feat(model): debounce model loads after a failure to stop retry-storms (#10728)
A client that keeps polling a model whose load fails (e.g. a backend that crashes deterministically on init) triggered a fresh backend start on every request: request -> load -> crash in ~10s -> 500, repeat on the next poll. Each attempt could leak GPU/CUDA state, and under LOCALAI_SINGLE_ACTIVE_BACKEND it kept stealing the active slot from healthy models. The existing loading-coalesce map only dedups *concurrent* loads, so sequential polls were never covered. Track load failures per modelID in ModelLoader. After a load fails, refuse fresh load triggers for that model until a cooldown elapses, returning a typed ModelLoadCooldownError that the HTTP layer maps to 503 with a Retry-After header. The cooldown grows exponentially per consecutive failure (base, doubling, capped at 5m) and resets on a successful load. The coalesced follower-retry of an in-flight burst bypasses the gate, so a genuinely concurrent burst still gets its one retry -- only new, independent triggers are refused, matching the report's "refuse new load-triggers" wording. Configurable via --model-load-failure-cooldown / LOCALAI_MODEL_LOAD_FAILURE_COOLDOWN (default 10s, 0 disables), plumbed through ApplicationConfig and applied unconditionally at startup. Closes #10719 Assisted-by: Claude:claude-opus-4-8 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
461ae84732 |
fix(startup): scope generated-content and upload dirs to the current user (#10698)
The `--generated-content-path` and `--upload-path` defaults were the fixed
shared locations `/tmp/generated/content` and `/tmp/localai/upload`. On any
multi-user host these collide across accounts: macOS routes `/tmp` to the
shared `/private/tmp` for every user, so whichever account starts LocalAI
first creates the parent with 0750 perms and every other account then fails
startup with:
unable to create ImageDir: "mkdir /tmp/generated/content: permission denied"
unable to create UploadDir: "mkdir /tmp/localai/upload: permission denied"
The same happens on Linux once a stale root-owned `/tmp/generated` (e.g. from
a prior `sudo` run) is left behind. This bites the desktop launcher and any
app embedding the raw binary (Wingman, nib-desktop), which start `local-ai
run` with no path flags.
Default both paths under the OS temp dir (`os.TempDir()`, honoring `$TMPDIR`;
already per-user on macOS) namespaced by the current UID
(`TMPDIR/localai-<uid>/...`), so accounts never collide while the paths stay
ephemeral. Wired via new kong vars in main.go so every consumer of the raw
binary inherits the fix. All content subdirs (audio, images) derive from
`GeneratedContentDir`, so they are fixed transitively.
As defense in depth, the launcher also anchors these two paths under its own
per-user data directory (mirroring the #10610 fix for data/config), extracted
into a testable `BuildRunArgs`.
Assisted-by: Claude:claude-opus-4-8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
|
||
|
|
8344d1c865 |
feat(cli): add interactive chat mode (#10226)
Add an opt-in `local-ai chat` command for testing chat models directly from the terminal without manually sending curl requests. The command connects to a running LocalAI server, lists available models through the existing OpenAI-compatible API, streams chat completions, and supports interactive commands such as `/models`, `/model`, `/clear`, and `/exit`. Keep `local-ai run` focused on the server lifecycle so the web UI, API clients, and multiple chat terminals can coexist against the same server. Document the new command and terminal workflow in the README and CLI docs. Tests: - go test -count=1 ./core/cli/chat - go test -count=1 ./core/cli Assisted-by: Codex:GPT-5 Signed-off-by: Ching Kao <0980124jim@gmail.com> |
||
|
|
8862e3ce60 |
feat: add node reconciler, allow to schedule to group of nodes, min/max autoscaler (#9186)
* always enable parallel requests Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat: add node reconciler, allow to schedule to group of nodes, min/max autoscaler Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * chore: move tests to ginkgo Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * chore(smart router): order by available vram Signed-off-by: Ettore Di Giacinto <mudler@localai.io> --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
aea21951a2 |
feat: add users and authentication support (#9061)
* feat(ui): add users and authentication support Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat: allow the admin user to impersonificate users Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * chore: ui improvements, disable 'Users' button in navbar when no auth is configured Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat: add OIDC support Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix: gate models Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * chore: cache requests to optimize speed Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * small UI enhancements Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * chore(ui): style improvements Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix: cover other paths by auth Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * chore: separate local auth, refactor Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * security hardening, approval mode Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix: fix tests and expectations Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * chore: update localagi/localrecall Signed-off-by: Ettore Di Giacinto <mudler@localai.io> --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
ed2c6da4bf |
fix(ui): Move routes to /app to avoid conflict with API endpoints (#8978)
Also test for regressions in HTTP GET API key exempted endpoints because this list can get out of sync with the UI routes. Also fix support for proxying on a different prefix both server and client side. Signed-off-by: Richard Palethorpe <io@richiejp.com> |
||
|
|
d200401e86 |
feat: Add --data-path CLI flag for persistent data separation (#8888)
feat: add --data-path CLI flag for persistent data separation - Add LOCALAI_DATA_PATH environment variable and --data-path CLI flag - Default data path: /data (separate from configuration directory) - Automatic migration on startup: moves agent_tasks.json, agent_jobs.json, collections/, and assets/ from old config dir to new data path - Backward compatible: preserves old behavior if LOCALAI_DATA_PATH is not set - Agent state and job directories now use DataPath with proper fallback chain - Update documentation with new flag and docker-compose example This separates mutable persistent data (collectiondb, agents, assets, skills) from configuration files, enabling better volume mounting and data persistence in containerized deployments. Signed-off-by: localai-bot <localai-bot@noreply.github.com> Co-authored-by: localai-bot <localai-bot@noreply.github.com> |
||
|
|
e45d63c86e |
fix(cli): Fix watchdog running constantly and spamming logs (#8624)
* Fix watchdog running constantly and spamming logs Signed-off-by: Andres Smith <andressmithdev@pm.me> * Update docs Signed-off-by: Andres Smith <andressmithdev@pm.me> --------- Signed-off-by: Andres Smith <andressmithdev@pm.me> |
||
|
|
64d0a96ba3 |
feat(ui): add video gen UI (#8020)
Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
c844b7ac58 |
feat: disable force eviction (#7725)
* feat: allow to set forcing backends eviction while requests are in flight Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat: try to make the request sit and retry if eviction couldn't be done Otherwise calls that in order to pass would need to shutdown other backends would just fail. In this way instead we make the request sit and retry eviction until it succeeds. The thresholds can be configured by the user. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * add tests Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * expose settings to CLI Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * Update docs Signed-off-by: Ettore Di Giacinto <mudler@localai.io> --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
fc5b9ebfcc |
feat(loader): enhance single active backend to support LRU eviction (#7535)
* feat(loader): refactor single active backend support to LRU This changeset introduces LRU management of loaded backends. Users can set now a maximum number of models to be loaded concurrently, and, when setting LocalAI in single active backend mode we set LRU to 1 for backward compatibility. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * chore: add tests Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * Update docs Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * Fixups Signed-off-by: Ettore Di Giacinto <mudler@localai.io> --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
ab022172a9 |
chore: switch from /usr/share to /var/lib for data storage (#7361)
* More appropriate place for data storing The /usr/share subtree in Linux is used for data that generally are not supposed to change. Conventional places for changeable data are usually located under /var, so /var/lib seems to be a reasonable default here. * Data paths consistency fix * Directory name consistency fix |
||
|
|
2dd42292dc |
feat(ui): runtime settings (#7320)
* feat(ui): add watchdog settings Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * Do not re-read env Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * Some refactor, move other settings to runtime (p2p) Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * Add API Keys handling Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * Allow to disable runtime settings Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * Documentation Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * Small fixups Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * show MCP toggle in index Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * Drop context default Signed-off-by: Ettore Di Giacinto <mudler@localai.io> --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
2cc4809b0d |
feat: docs revamp (#7313)
* docs Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * Small enhancements Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * Enhancements * Default to zen-dark Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fixups Signed-off-by: Ettore Di Giacinto <mudler@localai.io> --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |