Files
LocalAI/docs/content/advanced/model-configuration.md
T
4d7bdc6ff2 feat(ui): redesign the web UI around a shared kit and a calm palette (#12526)
* build(ui): vendor the shared UI kit snapshot at 0.2.0

The restyle needs the kit's tokens, motion layer and component classes.
Take a pinned snapshot instead of depending on the kit at build time,
and keep a lock file with the version and per-file checksums so a later
update shows exactly what changed. The product theme stays outside the
vendored directory.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add the ink and teal theme and bridge the old variables

Define the product colours as the shared UI kit's roles, for light and
dark, in theme-localai.css. The kit's contrast check passes on every
pair. theme.css keeps the existing --color-* and --shadow-* names but
now points each at a role, so App.css and the pages get the new palette
without edits. Radii move to the kit scale.

index.html now sets data-theme before first paint with the same rule as
ThemeContext (stored choice, otherwise dark), because the contract
layout of the theme file no longer defaults to dark by itself.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): restyle the shared chrome with the UI kit grammar

Adjust the shared classes so every page picks up the same interaction
language without per-page edits:

- Sidebar sits on the canvas and the current row lifts onto a card.
  Section labels are tracked uppercase, the badge is a soft pill, and
  the phone drawer leaves the tab order when closed.
- Buttons are flat: hover swaps the surface, press scales to .97, focus
  is a 2px ring with a 2px offset, danger is a tinted wash.
- Inputs use the card surface and the control edge; switches, tabs,
  filter chips, badges and cards follow the same rules. Cards no longer
  lift on hover; only linked or button cards react.
- Menus and popovers scale in from the trigger corner with 40px items.
  Dialogs get a veil fade and a spring settle. Toasts become pills at
  the bottom centre.
- The page transition is a 250 ms fade with a 6px rise. It fills
  backwards so a finished animation no longer leaves a transform that
  confined dialog veils to the main column.

The focus-ring test now checks the outline instead of a box shadow, and
new specs cover the theme roles, the first-paint theme and the sidebar
lift.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): move leftover hard-coded colours onto the theme roles

The YAML editor restated the old blue palette in JavaScript, and a few
pages kept literal blues, indigo and violet tints, or fallbacks that
only applied because a variable was never defined. Point them at the
theme variables so they follow light and dark and the new palette.

The status badges in the account pages built their tint by appending
"22" to a variable, which is not valid once the variable is defined, so
they had no background. Use the wash roles instead.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): raise the type scale and control size toward the kit

Body and list text moves to 15px and the rest of the scale follows the
kit's 12/13/15/17/21/32/44 steps. Page titles, section headings and
stat values are bold with tighter tracking; titles are 32px.

Buttons, inputs, selects, tabs and nav rows are 40px high with the 12px
radius, compact controls 32px. Tabs become a segmented control. The
sidebar widens to 240px (64px collapsed) and nav rows get more room.
Identifiers and counts in the split views use the mono face, and the
stat grid becomes separate inset tiles.

The Geist stack stays: it is bundled, and the thin look came from the
size, weight and negative tracking, not the face.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): separate cards, panes and floating surfaces from the canvas

Cards, the Models and Installed split panes, the composers and the
confirm dialog use a stronger card edge, the rest shadow and the 20px
radius, so they read as layers in dark as well as light. Menus and
popovers move to a float surface (the hover tone in dark) with the
float shadow.

The selected rail row gets an accent wash and a 3px accent edge. The
send buttons are a clear accent when there is something to send and a
quiet inset when not; the Home button carries data-empty for that, since
submitting an empty box does nothing. The assistant card becomes an
accent wash with a square icon.

New surfaces spec checks the pane edge, the selected row, both send
buttons and the popover in both themes. The voice library empty-state
spec now waits for the layout to settle before comparing two boxes.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): tidy the sidebar header and mark the current row with a dot

The header gives the configured horizontal logo a fixed width and
centres it in a 72px band, lined up with the nav icons. The collapsed
rail shows the configured icon logo centred, and its nav rows become
40px tiles centred in the 64px rail. The current row gets the kit's
accent dot, hidden in the rail.

The theme, language and account controls stay in the sidebar footer:
the app has no global search or command palette to put in a top bar, so
a bar would only hold controls that already have a place.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): centre the avatar in the collapsed sidebar rail

The collapsed avatar link was set to "flex: 0", which gives it a zero
flex basis; with min-width: 0 the link shrank to its padding and the
icon overflowed from the link's left edge, about 14px right of the
icon column. Use "flex: 0 0 auto" in the collapsed and tablet rail.

The footer controls now share the nav icon column in the expanded
sidebar too (6px footer padding, 40px control boxes), and the tablet
rail gets the same footer padding and hidden language code as the
collapsed one.

New spec measures the centre x of the nav icons, mark, avatar,
language, theme and collapse icons in the collapsed, expanded and
tablet states, in both themes, and asserts they agree within 1px.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): size the console and settings rails and stack settings on phones

The Operate console rail, the Settings section rail and the account tab
bar still used 13px text and the old underline tabs. They now use the
15px nav size, 40px rows and the segmented tab control. Form row labels
are 15px with 13px hints.

On a phone the Settings section rail sat beside the form and squeezed
every row into a few characters. Below 720px the rail stacks above the
content as a scrolling row and form rows wrap their control below the
label. The save button no longer carries the icon font class, which
drew a missing glyph before its label. The language menu is wide enough
to keep Bahasa Indonesia on one line.

Assisted-by: Claude Code:claude-sonnet-5-5

* build(ui): update the vendored UI kit snapshot to 0.3.0

Take the 0.3.0 snapshot: the sprite now carries the full outline icon set,
and the new icons/fa-map.json maps Font Awesome names to icon ids. The map
lets the app move off Font Awesome in the following commits. The lock file
is regenerated with the new checksums.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add an Icon component backed by the kit sprite

Icon draws an inline svg that points into the kit's outline sprite. The
sprite is inlined into the page once, so the references resolve under any
base path and in the embedded build without a request. Icons size with the
font (1em), take currentColor, hide from assistive tech unless given a
title, and spin on request. An unknown id draws a neutral circle.

FaIcon and iconFromFa resolve Font Awesome names through the kit's map,
for names that arrive at run time. iconHtml does the same for markup built
as a string. The GitHub and Apple marks are small local glyphs, as the kit
ships no brand marks.

Assisted-by: Claude Code:claude-sonnet-5-5

* refactor(ui): draw shared components and helpers with Icon

Replace the Font Awesome elements in the shared components and in the
utility modules with the Icon component. Lookup tables now hold kit icon
ids instead of class strings. Code-block copy buttons and artifact cards,
which build HTML strings, use iconHtml and a sanitizer-safe slot.

Assisted-by: Claude Code:claude-sonnet-5-5

* refactor(ui): draw model, backend and account pages with Icon

Replace the Font Awesome elements on the home, models, backends, import,
settings, login, account and users pages with the Icon component.

Assisted-by: Claude Code:claude-sonnet-5-5

* refactor(ui): draw chat, studio and recognition pages with Icon

Replace the Font Awesome elements on the chat, media generation, talk and
face and voice pages with the Icon component. The talk status table keeps
its spin and pulse states as Icon props. The connected and error states
now use a dotted circle and an alert circle, so they differ from the idle
ring by shape as well as by colour.

Assisted-by: Claude Code:claude-sonnet-5-5

* refactor(ui): draw agent, node and operate pages with Icon

Replace the Font Awesome elements on the agents, skills, collections,
jobs, fine-tune, quantize, nodes, swarm, usage, traces and activity
pages with the Icon component. Two class strings on layout elements held
leftover button and icon classes from an earlier merge; they are cleaned
up so the elements keep only their own classes.

Assisted-by: Claude Code:claude-sonnet-5-5

* style(ui): size and align icons for the svg component

Icon rules that targeted the font element now target the svg: the
descendant "i" selectors in App.css and auth.css become ".lai-icon". The
svg is 1.2em with a 2 unit line so it matches the visual size of the old
glyphs at the 12 to 16px sizes the app uses, sits on the text baseline,
and follows the context font size. Large empty-state marks get a lighter
line. Menu icons get a 16px box and the readiness badge icons keep their
20px circle with padding. Add the pulse used by the talk status.

Assisted-by: Claude Code:claude-sonnet-5-5

* build(ui): remove Font Awesome

No source file references the icon font any more. Drop the package and its
stylesheet import. The build no longer ships the solid, regular and brand
font files.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): keep focus traps off the svg use references

The dialog and drawer focus traps collect focusable elements with a
"[href]" selector. An icon's use element carries an href, so it became the
first "focusable" element and Tab at the end of the dialog stopped there
instead of wrapping to the first button. Match "a[href]" instead.

Assisted-by: Claude Code:claude-sonnet-5-5

* style(ui): keep icon sizes overridable and set the line width per svg

Give the icon base rule zero specificity so a rule that sizes one icon
(nav column, menu box, avatar, language switcher) wins whatever its order
in the file. The sprite symbols fix their own line width; the inlined copy
drops it so the width set on each svg applies, as the --lai-stroke custom
property, and large marks can use a lighter line. Pin the avatar and the
language globe to the boxes the sidebar alignment spec expects. Import the
map as JSON with an import attribute so Node can load it in the spec.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): select icons by the svg markup and cover the sprite

Specs that found icons by their Font Awesome class now select the svg by
its data-icon. The dead-icon audit checks that every svg resolves to a
sprite symbol and has a size. The class hygiene spec fails on any
remaining Font Awesome class. A new spec checks every mapped icon id has
a symbol, that the sprite is inlined once, that an icon paints at the root
and under a forwarded path prefix, and that Font Awesome names map as
documented.

Assisted-by: Claude Code:claude-sonnet-5-5

* build(ui): update the vendored UI kit snapshot to 0.4.0

Take the 0.4.0 snapshot: hub tabs with count and attention badges, the six
chart series tokens and the grid colour in the theme contract, and sample
themes on a calmer palette. The kit headers are renamed and the lock file
is regenerated with the new checksums, as for the earlier snapshots.

Assisted-by: Claude Code:claude-sonnet-5-5

* style(ui): switch the theme to the calm palette

Rewrite the LocalAI theme on the calm palette: a muted teal accent on a
near-neutral green-grey canvas, desaturated status colours, no glow and no
coloured shadows. The theme fills every role of the shared UI kit's 0.4.0
theme contract for light and dark, including the six chart series and the
grid line. The bridge in theme.css keeps the old --color-* names working,
adds the dark surface ladder (card, raised, float) and a strong edge, and
points the fixed data hues at the chart series.

Two values differ from the first sketch. The dark text on the accent fill is
#021512 instead of #04201d: it reads 5.58:1 on the fill at rest and 6.4:1 on
the hover fill, against 5.08:1 at rest for the lighter value. The light
control edge is #6b7d7a. The kit's contrast script passes for all text pairs
(4.5:1), control and focus pairs (3:1) and series colours (3:1).

Leftovers that no longer fit the palette are fixed: the usage chart takes
the six series colours in order, the audio and animation canvases fall back
to the new accent, the face box loses its glow, and two gradient fills are
now flat. The theme tests expect the new canvas colours.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): replace the console rail with hub tab bars

Build and Operate no longer open a second navigation rail beside the page.
Each is a hub: one row of the kit's hub tabs above the page, with count and
attention badges that scroll sideways on a phone. Every URL and route stays
as it was, plus a new /app/build landing page that lists the Build tools
with a line each.

Build tabs: Overview, Agents, Skills, Memory, Jobs, Fine-Tune, Quantize,
Import, Voices (recognition and library) and Faces. Operate tabs: Status,
This machine, Swarm (distributed mode only), Runtime (backends, activity,
failover), Traffic (usage, traces, middleware) and Settings (settings,
users), plus the API link. A tab that holds several pages shows a second row
of links, and a sub-page such as a node detail keeps its tab highlighted.
The feature and admin gates decide which tabs are drawn, and badges show only
values the Operate summary already has.

The sidebar lists Build and Operate under a Workspace label next to the
Create group. The voice library moves under Build and the model import page
gains the Build tab bar. The old rail styles, the rail signals and the
console config are removed, and the Operate overview docs describe the tab
bar. The specs that drove the rail now drive the tabs, and a new spec covers
the tab for each route, gating, badges and the phone layout.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild Home as a calm console

Home now opens on one command bar: the model chip shows which models are
warm, the MCP chip and attach buttons sit beside it, and Send is a solid
button with an Enter glyph. Typing "/" opens a grouped, keyboard-driven
action list built on the kit command list; every action has a destination
in the product.

Memory use folds into a one-line strip that opens into the loaded models,
with Stop per model and Stop all. It opens by itself while a model is being
staged and after a failure, and shows nodes and aggregate memory in a
cluster. The list of resident models carries no per-model size because the
API reports none.

"Jump back in" lists the conversations stored in the browser, one card per
day, with j and k to move, Enter to resume and delete with an undo toast.
First run keeps the install steps and the recommended models. The assistant
prompt is a dismissible line, the library links are one quiet row and the
API section is collapsed. Chat accepts an empty new-chat hand-off for /new.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Home console

Update the Home specs for the new structure and add specs for the slash
menu, the model chip, the memory strip (expand, stop, staging, failure,
cluster), the resume list (grouping, j/k, Enter, delete with undo), first
run, the send hand-off, a non-admin user and the phone layout.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add the fit, disk and cleanup helpers for the models page

Pure functions and hooks that the rebuilt Models page reads, with node
tests for the rules.

modelLedger turns an estimate and the memory budget into one of three
verdicts (fits, spills to CPU, over) with the headroom in bytes, and
reads the models disk from the resources reading. The disk counts as low
under 10 percent or under 20 GB free, and is absent when the server
reports none or runs as a cluster controller.

cleanupPlan ranks installed models from what the API reports: loaded,
pinned, or named by an agent, a task, a failover chain or an alias keeps
a model protected; another installed build of the same gallery model is a
duplicate; disabled models rank above idle ones. The API records no last
use or use count, so none is used. When a lookup fails, nothing is called
safe.

useModelRemoval holds a removal in the browser for an undo window and
sends the existing delete call only when the window ends. Leaving the page
drops the batch without deleting anything. The undo toast takes optional
labels so other pages can reuse it.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild Models as a ledger with a disk strip and cleanup review

Explore is one dense table. Each row carries the size, a solid memory bar
and the headroom in words ("3.7 free", "+1.5 on CPU", "0.9 over"), worked
out from the estimate at the chosen context length. Capability chips show
the server's count for each facet, search keeps its meaning and "/" jumps
to it, and a density switch (also "d") picks comfortable or compact rows.
Selection is a surface step and a check, never a rail. Arrow keys move,
Enter installs and Esc closes the inspector, which keeps the fit summary,
VRAM by context chart, variants, files, links, tags and licence. A failed
install shows its error in the row with a Retry that dismisses the old
failure first. A failed or empty listing says which it is, and a host with
no GPU is measured against memory and says so.

Installed uses the same table with state filters that carry counts, a
state per row, Load or Stop on the row, the row menu and the sort by size.
Sizes come from the files the gallery lists, so a model it does not know
shows a dash.

A strip in the header shows the free space on the models disk. It turns
amber under 10 percent or under 20 GB free, hides when the server reports
no disk or runs as a cluster controller, and opens the cleanup review.
Explore says how much an install leaves free.

The review ranks installed models as Safe to remove, Probably safe and
Your call from real facts only, lists protected models with the reason,
and says plainly that usage history is not recorded. A sticky bar shows
what a choice frees. Confirming runs a dry run that checks again and lists
what will go. Removal waits 30 seconds with an undo; nothing is deleted
before that, and leaving the page deletes nothing.

The old rail, filter band and popover styles are removed.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Models ledger, Installed table and cleanup review

Update the Models, lifecycle, cluster fit, height, search focus and
surfaces specs for the table and inspector, keeping what each one checks.

New specs, on a shared 41-model gallery stub with three machine profiles:
the fit bar and headroom words for a 24 GB card, an 8 GB laptop and a host
with no GPU; facet counts, search, "/" and Escape; selection, arrow keys,
Enter to install, density; the disk strip when normal, low and hidden; and
the states (loading, empty, offline, install failed, phone). Installed
covers filters with counts, row actions, the row menu, sizes and sort.
The cleanup specs cover grouping, protected models, the honest-data note,
the effect bar, the dry run, the undo window, a failed delete, leaving the
page, and the phone sheet.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add the placement helpers and the estimate hooks

The Placement section and the model page need the same few rules, so they
sit in plain functions that can be read and tested alone.

placement.js holds what gpu_layers, tensor_split and main_gpu mean (unset
asks for every layer and the llama.cpp engine trims it, zero is CPU only,
99999999 is the value LocalAI itself writes for all layers), the device
list taken from the resources reading, the split by free memory, the part
of an estimate that grows with context (read from two lengths, since that
term is linear), the fit states with their limit (95 percent of free
memory, and the leftover has to fit in system memory too), and a bisection
for the largest layer count whose estimate fits. The estimate returns one
total and no layer count, so the search runs over 1 to 256 and stops at
the first count that no longer changes it.

modelWalk.js keeps the order of the list a model page was opened from, in
memory and in session storage, for the previous and next buttons.

usePlacementEstimate reads /api/models/vram-estimate for a choice, again at
twice the context, and with every layer, and keeps readings for the
session. useModelPage reads a gallery entry by name, an estimate by
context size (from the model's own files when the gallery does not list
it), the builds and the loaded models. usePlacementConfig edits the four
placement keys of an installed model and saves only what changed.
useModelActions is the Load, Stop, disable, pin and remove logic of the
Installed table, shared with the model page. MemoryBar is one solid bar
with a tick at the capacity of its pool; over capacity it grows past the
tick and the tick turns red.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add the Placement section to the model editor

Run this model on: CPU only (gpu_layers: 0), Auto (the key stays unset) or
Custom. Custom takes a number, has an All layers button that writes
99999999, and shows a slider only when the estimate reports the model's
layer count, which it does not today. Context size has presets and a
number field because the KV cache follows it. With two or more GPUs there
is a split (written as percentages, with a button that takes them from the
free memory of each card) and a main GPU.

A bar per GPU and one for system memory show what other programs use, the
model's weights and working memory, and the part that grows with context,
with the room left or how far over it is. Under them a verdict in plain
words: Fits in GPU, Spills to CPU, Too many layers for the GPU, Runs on CPU
only, No GPU found, Not enough memory. It says "slower" and never a
multiplier, because the estimate has none. Fit it for me asks the estimate
for the largest layer count that fits the free GPU memory and says what it
set, with Undo; it is hidden when the estimate is unavailable or the host
has no GPU. Loading shows skeletons, an unavailable estimate shows a note
with Retry, and a server that schedules onto other machines shows no bars,
because its device list is the controller's.

The editor shows the section for an installed model, with a link in its
section rail. Auto sends null for the key, since a patch only merges, and
a null read back opens as Auto. The docs describe the section and what each
mode writes.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): open a model on its own page

A model has an address, /app/models/<name>, for an installed model and a
gallery entry alike. Open it from the arrow at the end of a row, a double
click, "o" on the selected row, the inspector's Open details button, or a
tap on a phone. The title block holds the main action: Install with a
chevron that chooses the build, or Load and Stop with a menu (disable, pin,
edit configuration, logs, delete with a confirm). A strip answers whether
it fits, what it does and what installing leaves free.

Tabs: Overview (about, a memory bar, state, the pages it opens in, and the
agents, tasks, chains and aliases that name it); Fit and memory (verdict,
context sizes, the bar split into weights and context, and memory by
context against the limit, with a data table); Variants and files (builds
with size and fit, install any, the files of the chosen build). For an
installed model also Usage and history, which says what the API does not
record instead of drawing an empty chart, Configuration, which is the
Placement section with the file it writes and a link to the full editor,
and Logs, the backend log viewer without its page. Keys 1 to 6 switch
tabs, [ ] and j k walk the list the page was opened from, Esc or Backspace
go back.

The list stays mounted behind the page, so Back finds its view, search,
filters, selection and scroll as they were, and focus returns to the row's
arrow. The page covers loading, an unknown name with the closest matches,
the gallery being out of reach, an install in progress with Cancel, and a
failed install with Retry.

The docs describe the page and its keys.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the model page and the Placement section

New specs for the model page: reaching it from Explore, Installed, a
double click, "o", a pasted link and a phone tap; the walker and Back with
the search, a filter, the selection, the Installed view and the scroll
kept, and no second read of the gallery; the title block, the answer strip,
tabs by click, keys and arrows; Fit and memory, builds and files with the
install call each one makes; an installed model's actions, used-by,
the honest usage tab, configuration and logs; loading, an unknown name,
offline, an install in flight and a failed one; and the phone.

New specs for Placement: every mode and the keys it writes, the slider
only when a layer count exists, the context presets, the bars and every
verdict, two GPUs, no GPU, a cluster, an unread machine, a loading and an
unavailable estimate, Fit it for me and Undo, and the section in the model
editor with its save.

The phone tap on a row now opens the page, so the two phone specs that
expected the inspector as the page check the page and keep the inspector
check for a window between a phone and a desk.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): link Studio results and open workspaces from a prompt

Each workspace now records the result it was made from (parentId and an
edge kind such as take, animate or to-3d) and reads a prompt, model, size,
count and source from the query string, so one page can hand work to
another. A source result is fetched from the server's own output file and
becomes the start image, the picture for 3D, or the audio file. A note on
the page says when the source loaded or could not be loaded.

Diarization had no history; it now keeps the file name, the model and a
speaker count, never the recording. Prompts are cut at 2000 characters
when stored. The pure helpers (type suggestion, grouping, lineage layout,
favourites, clearing) have node tests.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild the Studio front page as a composer with your work

The front page is a prompt box with a chip per type, a type suggestion
from the words, starters, and the options each workspace accepts. Generate
opens the workspace with those filled in. A type with no model is a dashed
chip that shows a gallery model, its size, memory need and an Install
button only when picked; the typed words stay while it installs.

Under it, Your work lists results from every workspace as a masonry with
filters, counts, favourites and a Clear history action. Results made from
each other stack into a project tile and open as a lineage board with a
dock for running a new take or branching to the next step; steps the
destination cannot start from yet are disabled with the reason.

The docs describe the page, what is stored in the browser, and the query
parameters a workspace accepts.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Studio composer, your work and the lineage view

Specs for the type suggestion, the keys, hand-off to each workspace, the
install path for a missing model, the masonry filters, favourites and
clearing, stacking, the lineage board, new take and branch, steps that
are disabled with a reason, and the phone layout. Existing Studio specs
move from lanes to chips with the same intent.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add the shared Studio workspace frame and move Images onto it

The seven Studio workspaces get one layout: a row of type tabs, a compose
card (optional sources as chips, a prompt with starters, a model chip,
the essential options as chips, an Advanced fold that names what is
inside, the memory the model needs, and one action with the reason when
it cannot run), a run area, and a strip of recent results of the type.

The run area shows a job card with the time that has passed and an
indeterminate bar, because these endpoints report no phase or percentage;
a failure with what the server said and one action; or the result with a
toolbar: Favourite (the list the front page keeps), Download, Use in (the
hand-off targets, disabled with the reason when a destination cannot
start from the result), Re-run with edits (the take's values go back in
the form, changed fields are outlined and listed) and Lineage. A type
with no model shows the install note from the front page.

Images is the first workspace on the frame. It keeps its size, count,
steps, seed, negative prompt, source image and reference images, and its
history writes, including the parent link and edge of a hand-off run.
useMediaHistory.addEntry now returns the id of the entry it stored. The
docs describe the workspace page.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): move Video onto the workspace frame

Video keeps its size list, duration, frame rate, steps, seed, CFG scale,
frame count, negative prompt, start and end image and avatar audio.
The start and end image are source chips, the avatar audio opens the
recording and paste input from a chip, and the rest sit in the Advanced
fold. A start image from a hand-off shows as a chip with its picture.
Results play in the video player with the shared toolbar.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): move TTS onto the workspace frame

TTS keeps the saved-voice picker for cloning models, the typed voice for
the others, the voice library deep link, and the delivery instructions,
which now sit in the Advanced fold. The result is the waveform player
with the words under it. The stored entry also keeps the voice id so
Re-run with edits can select the same saved voice.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): move Sound onto the workspace frame

Sound keeps its Simple and Advanced modes and every field of both: the
description, instrumental, vocal language, caption, lyrics, BPM,
duration, key, language, time signature and think mode. The mode switch,
instrumental and duration are in the compose card, the rest in a More
options fold. The stored entry keeps all of the fields, so Re-run with
edits restores the form as it was.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): move Transform onto the workspace frame

Transform keeps its audio and reference inputs with upload and record,
the echo test, the key=value parameters (now in the Advanced fold), the
input and output spectra and the three waveform players. The audio that
was chosen shows before the run, waiting to be transformed. Re-run with
edits puts back the model and parameters and fetches the audio and
reference the server kept for that run.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): move 3D onto the workspace frame

3D keeps the picture input with paste and webcam, the animation
operations a model declares, quality and background, the shape and
material steps, guidance and seed, the GLB and animation viewers, the
remesh control and the download. A 3D result now has a title from the
motion prompt when it has no label, so the strip and the front page name
animation results by what was asked. Re-run with edits is shown disabled
with the reason, because only a small thumbnail of the picture is kept.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): move Diarization onto the workspace frame

Diarization keeps its model and recording inputs, the option to prepare
speakers to remember, the clean-speech previews, naming and remembering a
speaker, and the history entry with only the file name, model and
counts. The result now shows a timeline with one lane per speaker, the
talk time of each speaker, and the segments with their start time and
text. RTTM, SRT (only when the run has text) and JSON are built in the
browser from the result. The helpers for talk time, axis ticks and the
two text formats have node tests.

Assisted-by: Claude Code:claude-sonnet-5-5

* refactor(ui): remove the styles and lists the old workspace layout used

Nothing renders the two-column workbench, the control column, the old
history lists, the generation progress tiles, the TTS voice picker or the
result echo any more. The inline-style baseline drops with them.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the workspace frame and one run per type

Specs for the type tabs, the compose card and its reason when Generate
cannot run, starters, the Advanced fold, the job card with no invented
progress, a failed run and its one action, the install note, the strip
with its favourites filter, Use in with its disabled steps, Lineage, the
parent link, Re-run with edits and its list of changes, deleting and
clearing, and the hand-off note. One run through each of Video, TTS,
Sound, Transform, 3D and Diarization, the phone layout of all seven, and
reduced motion. Existing Studio specs move from the old control column to
the compose card with the same intent.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild the chat thread with raised user turns, prose replies and one-line activity

Your messages are raised blocks on the right at a 760 px measure and the
model's replies are plain prose under its name and a warm or not loaded
dot. Reasoning, tool calls and their results fold into one quiet line
that opens inline into steps. Code blocks carry a Copy button and a
Canvas button that opens that block in the canvas, image attachments are
thumbnails that open in the lightbox, and files are chips. Per-message
actions show on hover, on focus and on the last turn, and a turn takes
focus so the arrow keys and C, E, R and B work. A failed reply keeps the
text written so far and shows the reason with one Retry action.

The Agent chat page keeps the older rules: the new styles are scoped to
the chat page and use their own class names.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): use the Home command bar as the Chat composer

Chat now ends in the same object as Home: the model chip, the MCP chip,
a Canvas chip, the message box with attach buttons, a solid Send and the
hint line, with the slash menu on the kit command list. The slash menu
lists what Chat can do today (switch model, new chat, conversations,
manage mode, canvas, find, settings, export, clear). While a reply is
streaming Send becomes Stop, which Esc also presses, and Up in an empty
box edits your last message. Attached images show as thumbnails and a
line under the bar carries the speed and the token count.

HomeComposer takes optional props for this (extra chips, its own slash
list, Stop, paste, a stricter Enter); Home passes none of them.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): open conversations from a Ctrl K menu with day groups, undo and a slim header

The conversations list opens as a centred menu on Ctrl or Cmd K. It
groups chats by day like the Home resume list, shows the model that
answered and the time, searches names and message text, and moves with
the arrow keys. Enter opens a chat, F2 renames it and Delete removes it.
Removing a chat hides the row and shows the kit undo toast; the chat is
deleted for good only when the undo time ends. Rename, duplicate, copy
and export are on each row, as before.

The header is one slim bar: the Chats button, the chat name (click to
rename), a context meter when the context size is known, settings and a
More menu with rename, duplicate, copy, export, model info, keyboard
shortcuts and clear. A dialog lists the shortcuts the page answers to.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): show loaded state, capabilities and fit in the Chat model switcher

The model chip in Chat opens the same list as Home, grouped as Loaded
now and Installed. Each row says warm or not loaded and marks models that
understand images. When the list opens, the page reads the host memory
once and asks the server to estimate each listed model at the chat's
context size (up to twelve, three at a time), then shows what the model
needs and whether it fits: free memory, how much would run on the CPU, or
how far over the machine it is. A model with no estimate shows no fit
text, and no load time is shown because the API does not report one. A
memory bar closes the list.

The picker takes the model list from the page when it has one, and
useModels can skip its own request.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): move chat settings into a sheet and add find in chat, jump to latest and a wider canvas

Settings open as a kit sheet: the system prompt, temperature, top P and
top K (each says "model default" until it is changed and has a Reset),
the context size with quick sizes and a note that it only drives the
meter, Manage mode and Focus mode, the model info for admins with its
Edit config button, and Clear conversation behind a confirmation. The old
slide-out drawer and the model info panel are gone.

Ctrl or Cmd Shift F (or the search button, or /find) opens a search bar
over the thread. It marks matches in the messages already on the page,
shows "n of m" and steps with Enter and Shift+Enter. Nothing is sent to
the server. Jump to latest is a pill above the composer. Esc stops a
reply, then closes the search, then closes the canvas.

The canvas panel gets the kit look: tabs, a Code and Preview switch, Copy
and Download, a full-page layout on narrow windows, and translated
labels. The Agent chat page shares it and gets the same look.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add the empty, no-model, loading and phone states to Chat

An empty chat opens with the composer under one line, starters to try,
whether the model is loaded, and the Jump back in list: the same rows as
Home, read from the chats the page holds. With no chat model installed,
an install card offers the starter models for this hardware, the gallery
and import, and the composer stays so the text is not lost.

While a reply waits for a model, a load card shows what the page knows:
the phase the server names, the node, the bytes and the time left when the
server reports them, and a progress bar. A model that is just not loaded
yet gets a plain note, with no invented phases or estimates. The foot
warns when the context is nearly full.

On a phone the header drops its labels, the model list and the settings
open as sheets from the bottom, per-message actions stay in view and the
canvas takes the whole page.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild Talk as a calm voice page over what the connection really does

Talk is one stage and one transcript. The stage has the pipeline chip,
the voice and language chips, an outline orb that follows the real
microphone and playback levels, a heading and a sentence for the current
state, and the controls. The transcript lists You, Reply, Tool and Result
lines and can be copied. Session settings (instructions, voice, language,
tools, Manage mode and the pipeline's parts) open in a sheet.

The states are the ones the code reaches: no pipeline model, idle,
connecting, listening, thinking (also while a tool runs), speaking, an
interrupted reply (the server cancelled it; a note marks the cut), a
blocked microphone, a link that failed during a session, and any other
error with its reason and a link to the traces. Push to talk and
hands-free are not on the page, so they are not shown. Diagnostics keep
their waveform, spectrum and stats, drawn in theme colours.

The page text moves into the talk namespace, and the old Talk and
visualizer styles and the inline-style count go down with the rebuild.

Assisted-by: Claude Code:claude-sonnet-5-5

* refactor(ui): remove the chat styles and strings the rebuilt page replaced

The settings drawer, the model info panel, the bubble avatars, the
conversation menu popover, the context bar, the recent strip, the
staging bar, the file badges and the focus-mode rules have no user now.
Their rules, the Chat page's focus class and seven unused empty-state
strings are removed. The Agent chat page keeps the shared message,
sidebar and input rules it still renders with.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the rebuilt Chat and Talk pages

Add a Chat page under features (thread, message actions and keys, the
message box and its slash actions, the model list with loaded state and
fit, conversations on Ctrl K, settings, find, canvas and the empty,
no-model and loading states) and a Talk section to the realtime API page
with the states the page shows. Manage mode now turns on from the chat
settings or /assistant, and the client MCP steps point at the MCP chip.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): settle the rough edges of the new Chat and Talk pages

The undo toast sat under the conversations menu, so the Undo button could
not be pressed while the menu was open; the menu, the sheets and the
fullscreen canvas now stay below the toast layer. Esc in a rename box
saved the text through the blur that follows it; it now cancels. The
image viewer closed on Esc only when the page did not re-render on the
same key, so its key listener is registered once and reads the latest
handlers. Keys on a focused message no longer type their letter into the
editor they open, "/" from outside a text field starts a command as it
does on Home, and Esc leaves the page's own dialogs alone.

Code in the canvas is highlighted for languages that have no preview.
The conversations menu drops its key hints on a phone so Clear all
stays in view. Talk hides Test tone while connecting and calls a server
error "Something went wrong", since the call can still be open.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the rebuilt Chat and Talk pages

Specs for the thread layout and the activity fold, code blocks, image
thumbnails and the viewer, per-message actions and their keys, a failed
reply with its one Retry, Stop and Esc while streaming, the composer and
every slash action, the conversations menu (groups, search, resume,
rename, delete with undo that ends by itself, one chat left), the model
switcher with loaded state, vision and fit text from stubbed estimates,
the settings sheet, the canvas panel, find in chat, Jump to latest, the
empty, no-model and loading states, the phone layout and reduced motion.
Talk is driven over a fake WebRTC link through idle, connecting,
listening, thinking, speaking, interrupted, blocked, lost, error and no
pipeline, its settings sheet and its phone layout. Node tests cover the
message text helpers and the conversation grouping.

The existing chat specs move to the new structure with the same intent:
the transcript spec now describes the raised turn and the prose reply, and
the render smoke accepts Talk's own header.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): keep a bounded run log for agents in the browser

The server keeps no run history for an agent, so a run is one task and
the events until the agent answers, written to browser storage while the
page watches the stream: up to 50 runs per agent, task, step and answer
text only. Stored chats from the earlier agent chat page read as runs
with stable ids. A run still marked running five minutes after its last
event reads as stopped. Helpers read an agent's config into chips, build
the list of changed fields against the saved config, hide secret values
and offer starting points.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild the Agents area around runs

The Agents page shows what needs a look (work in flight, a run that
failed in the last day), then each agent with its model, attached memory
and skills, and a strip of its last 14 runs. An agent has its own page: model,
tools, memory, skills, instructions, a task box and its runs. A run has
an address, shows the thread while it works (steps folded into one line,
the tool in use, the answer as it arrives) and settles into a report
about a second and a half after the agent answers: task, outcome,
follow-ups, evidence and steps, with wide tables opening wider on demand.
A failure says in plain words what happened and offers Run again.

Create and edit fold into sections with a ready mark and a one-line
summary, start from a template or an optional model-written draft, and
open a preview sheet with the config as saved and the changes against
the saved agent. Status becomes a quiet panel in the same language, and
the old chat link opens the agent page.

There is no Stop, approval, steer, version or dry-run control, because
the agent API has no call behind them.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Agents launcher, agent page, runs and editor

Specs for the Now strip and run strip, search and the empty state, the
agent page, starting a run, the live thread, settling into the report,
the run address across a reload and for a run from another browser,
follow-ups with their history, failures, the folding editor with ready
marks, templates, the preview sheet with hidden secrets and changes, the
status page, and the phone, 1440 and 2560 layouts.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe runs and the new agent create flow

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add helpers for tasks, schedules and job outcomes

Reads a cron expression the way the server does (five fields or an @
shortcut), checks it, and puts the common shapes in words. The next run
is left out on purpose, because the schedule follows the server clock,
which the browser cannot read. Also groups jobs by day, sums the last
seven days, and gives each job one outcome line from its result or
error. A rerun call starts a new job with the same parameters and media.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild the Jobs area around runs

The Jobs page opens with one sentence about the last seven days, then
the tasks (model, schedule in words, last 14 jobs, enabled switch, Run
now) and a run history grouped by day. Each row has an outcome sentence
and opens to the error or the start of the result with one next action.
Deleting a task waits 30 seconds with an undo button.

A task opens as a page with its recent runs, its prompt with the gaps
marked and its schedule. The task form folds into sections, takes a
schedule as a preset or a checked cron expression, warns about prompt
gaps the schedule does not fill, and has a preview sheet. A job opens
as a document: task, outcome, delivery and the recorded steps; a failed
job says what happened and offers Run again.

Run now now sends attached media through the job call, which is the
only one that takes it. "Clear History" only ever cancelled running
jobs, so it is now called Stop running jobs. Webhook headers of a saved
task show as JSON instead of [object Object].

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Jobs page, task pages and job pages

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the Jobs page and the task form

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add helpers that say who uses a skill or a collection

An agent loads a skill when skills are on and the skill is in its selection (an empty selection means every skill). It reads the one collection that carries its own name, when its knowledge base is on. The helpers derive that from the saved agent configs, build the config that adds or removes a skill or a collection, and estimate tokens as characters divided by four. Removing the last selected skill switches skills off, because an empty selection would mean every skill.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild Skills and Memory as one library

Skills and collections sit in a list with an open item beside it. Each row says who uses it, read from the saved agent configs, or says it is not used yet. Chat reads neither, so it is never named. An item opens in a pane with a Used by strip (names link to the agent, a small x removes it, with undo) and an Add to menu that shows what the addition costs. A collection can be added only to the agent that carries its name.

The Memory pane searches the collection alone and shows ranked passages with scores, lists web sources with their refresh interval and the files, shows the server message when an upload fails, and names the endpoints and where files stay. The Simulate a message sheet runs a collection search and shows an agent's skills with a token estimate. It runs no model. The collection details route now opens the same page.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): show use and cost in the agent form pickers

Each skill in the agent form says which other agents use it and what it adds to every message, with a total for the selection. The memory section names the collection the agent reads.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Skills and Memory libraries

Specs for the used-by lines (including an agent that uses every skill), the filters, search, add to agent, remove with undo, the last-skill case, an unreadable agent list, the empty states, git repositories, the Memory question box, sources, uploads that fail, the Simulate sheet with the parts the API can run, the agent form hints and the phone layout.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the Skills and Memory libraries

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add helpers for the Operate status page and backend rows

Pure functions for the parts that need rules. They work out the memory
pool the page measures (GPU memory, system memory, or the workers of a
cluster that are answering), which pools are too full, the headline and
the four ledger rows, the geometry of the capacity chart, and what
removing a backend would leave without a runtime (models name their
backend, and a meta backend names the concrete one it points at). A
second set says what a backend row states: installing, queued, removing,
failed, update available, current or absent.

LocalAI keeps no memory history, so the chart reads a bounded buffer of
readings the page took itself and says so. A reading with no total is
dropped rather than drawn as zero.

Two hooks are shared by the pages that need them. One retries a failed
operation after moving the failure into the record. The other holds a
cancel for an undo window, because the server cannot take a cancel back.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild Operate Status, This machine, Backends, Activity and Logs

Status opens with one sentence ("2 things need you", or "Everything is
running") and four rows: Needs you, Capacity, Running now and Recent
failures. A row with a problem opens by itself and holds the button that
deals with it: Update a backend, Retry or Dismiss a failed operation,
Unload a model. A quiet row stays one line. A new installation gets a
first-run screen, a cluster sums the memory of the workers that are
answering, and a page still waiting for an answer says so. The chart
under the rows is drawn from readings the page took while it was open
and is labelled that way, because LocalAI keeps no memory history.

This machine leads with GPU memory as one bar, then host memory split by
running model, then VRAM, RAM, CPU and disk with a bar each. The running
models become a kit table with the same menu and stop dialog.

Backends is one list with Installed and Catalog views. A row says what
the backend is doing (a progress bar with Cancel, Queued, Failed with
Retry, Update 1.2.0, Current), carries the one button that matters, and
opens in place. Removing a backend names the models and the meta
backends that would stop working. Check for updates, Update all, From
URL and a first-run recommendation for llama-cpp are in the header.

Activity keeps its three sections as quiet rows. Cancel waits eight
seconds with an undo toast, because the server cannot take a cancel
back; a cancelled install can be started again from the record. Logs
gets a process list, a picker, stream and text filters, Follow and
Times switches, and a Clear with an undo window.

Not shown, because the API has no data for them: GPU temperature and
power, a size per backend, an earlier version to roll back to, a
dependency lookup beyond the models and meta backends that name a
backend, and models that failed to load.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover Operate Status, This machine, Backends, Activity and Logs

New specs for the Status headline and ledger (healthy, needs attention,
one thing, a full memory pool alone, loading, first run, cluster), its
actions (Update, Retry, Dismiss, Unload with its dialog), the capacity
chart built from readings taken while the page is open and bounded, the
phone layout, no coloured edge on a row, and reduced motion.

The Backends specs cover the two views, install progress with Cancel and
its undo window, Retry on a failed install, Update, Update all, Check for
updates, removal with the models and meta backends it would break,
Install from URL, the first-run recommendation, a cluster, and a phone.
Activity gains cancel with undo, Cancel now, a second cancel, leaving the
page, progress, and starting a cancelled install again. Logs covers the
stream and text filters, Follow, Times, Export, Clear with undo, the
process picker and list. This machine covers the GPU strip, several GPUs,
no GPU and Add a machine.

Existing specs keep their intent and follow the new structure: rows open
in place instead of in a pane, Update replaces Upgrade, the notice spec
now pins that an update is a row state and not a banner or a rail, and a
cancel waits for its undo window.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the rebuilt Operate Status, Backends and Activity pages

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add helpers for the Swarm pages

Pure functions for what the pages work out from the cluster API: a node's
state in words, which nodes a placement rule may use, what a rule would
ask for, what a drain or a lost node would leave without service, the
nodes a bulk backend update reaches, and the join commands for a worker,
a peer instance and a memory shard. Hooks read the roster, the loaded
replicas and the rules.

Everything runs in the browser from data the page already holds, and
says when it cannot see free memory or disk.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild the Swarm hub: nodes, node page, placement rules, failover

Nodes is a sortable table with comfortable and compact rows, a Needs
attention filter by reason, a map of the cluster that is not drawn on a
phone, the running models, and a bulk backend update for the nodes that
drifted. A node is a page: state, vitals, a drain preview computed from
the loaded replicas and the rules, tabs for models, backends, logs and
capacity and labels, and Remove that asks for the node's name.

Placement rules are written as sentences, show where each model is
loaded now, and edit in a side sheet with a preview of the nodes a draft
could use. Deleting a rule waits a few seconds so it can be taken back.
Failover keeps its chains, adds what the router does when a worker stops
answering and a per-node preview of what would stop. Add a node covers a
registered worker, a peer instance and a memory shard, with a command to
copy and a live line that says when the machine arrived. P2P keeps its
page in the same vocabulary, and the node logs page follows the local
logs page.

Previews are labelled as worked out in the browser. Per-GPU readings and
node events are not drawn because the API does not return them. Failover
moves to Swarm when distributed mode is on. Legacy fleet components and
their styles are removed.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Swarm hub

New specs for adding a node (each join method, the command, copy,
waiting and found, approve, a single install, P2P, a phone), placement
rules (sentences, where models are loaded, the preview matrix, the sheet
and its preview, delete with undo) and failover on a cluster. Node
detail covers its tabs, the drain preview and its dialog, resume, remove
with the typed name, a node that stopped answering, and unload.

The nodes specs follow the new structure and keep their intent: the
table, filters, grouping, pagination, bulk actions, the map, and running
models with stop, logs and the loading, error and empty states. The
scheduling, failover, P2P, hub and smoke specs follow the renames.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the rebuilt Swarm pages

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add helpers for the Traffic pages

Pure functions for what the pages work out from the usage ledger, the
trace summary, the trace buffers and the resources reading: the shared
time window, grouping, sorting and filtering of usage rows, chart series
and axes that start at zero, the overview figures, per-model statistics,
the state of a trace and the words for a failure, the backend operations
that ran during a request, CSV export, the Prometheus metric list and
scrape config, and a bounded buffer of host readings.

A figure whose source cannot say is null, never zero. The trace summary
call takes the window in hours, and a helper reads /metrics with its
status.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild the Traffic hub: overview, usage, models, host, traces, middleware

Traffic opens on an overview: five figures (requests, failed, p95, tokens
in and out) and three charts, each naming its source. A second row of
links reaches Usage, Models, GPU and host, Traces, Middleware and
Prometheus, and one time window is shared by the first three.

Usage groups by model, user or API key, filters, sorts, opens a row on its
own chart, exports the rows it holds as CSV or JSON in the browser, and
keeps the opt-in cost estimate and the quota forecast. A user who is not
an admin sees only their own numbers. Models joins the ledger, the
backend-operation buffer and the loaded models. GPU and host shows the
current reading and two charts of readings taken since the page opened.

Traces gets filters, a settings strip and an explained off state. An API
request is a page: the error LocalAI recorded, a timeline with the backend
operations that ran meanwhile, and bodies that stay closed until revealed.
Middleware draws the pipeline as five steps and shows the rules of the
selected step. Prometheus documents /metrics, checks it against the
server and gives a scrape config to copy.

Alerts is not built: LocalAI has no alert rules. Per-model latency
percentiles, GPU utilisation and compare with the previous period are not
drawn because the API does not return them. Legacy usage, trace and
middleware styles and the usage source components are removed.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Traffic hub

New specs for the overview (figures, charts with a data table and arrow
key readout, failed and first-run and tracing-off states, the shared
window, a phone), usage (group by, filters, sort, export, cost, quotas, a
non-admin, empty and loading), models, GPU and host (snapshot, the
since-opened labelling, a cluster), the traces list, a trace page (the
real error, the timeline, reveal, no headers, a trace that left the
buffer), Prometheus and the Middleware pipeline, with shared fixtures.

The usage, traces, middleware, hub and smoke specs follow the new
structure and keep their intent.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the Traffic hub

Add an operations page for the Traffic tab: which record each page reads,
what it leaves out and why, the trace page and its reveal, the GPU and
host readings kept since the page opened, and the Prometheus endpoint.
Link it from the operations index, the tracing page and the middleware
page.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): let metric names wrap in the Prometheus table on a phone

The long metric names pushed the type and "on this server" columns out of
view. Names now wrap inside the table.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild Settings with groups by intent, search, a pending bar and history

The fifteen sections become eight groups by intent: memory and models,
speed and defaults, backends and galleries, access and security,
debugging and traces, agents and responses, swarm and sharing, look and
feel. Search covers names, descriptions, keys and the old section name, and
says where a result used to be.

Edits wait in a bar with Discard, Show diff and Apply. The diff lists old and
new values and the checks the browser can make: durations parse the way Go
parses them, a GPU memory budget is one the server accepts, a gallery box
holds JSON, and warnings repeat what the handler and the field text say.
Apply sends only the changed keys. Undo saves the previous values again; it is
a new save, not a rollback. History lists the changes applied from this
browser, since LocalAI keeps no settings log, and Revert stages the old value.

A value is marked as changed only where the built-in default is known from
the CLI defaults. A row says "Applies now" or "Needs restart" only where the
handler or the docs say so.

Three things were wrong before and are fixed with the rebuild: the gallery
boxes and the shared API keys box were sent under names the server ignores,
the "Enable CSRF Protection" switch showed the disable flag the wrong way
round, and every save restarted peer-to-peer networking because every field
was sent.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild Users and keys, Account, sign-in, invite and the 404 page

Users and keys is a tabbed page under the Settings tab: people, invites and
API keys. The people table filters by state and role, sorts, approves or
disables (disabling offers an undo that sets the status back), and opens a
side sheet for one person's features, model allow-list and limits. Role,
password reset and delete sit in the row menu; delete asks for the name.
Invites choose a lifetime of 1, 7 or 30 days and show the link once. API
keys can be created with a lifetime, are shown once in full, can be paused,
and are revoked after a ten second undo window in which nothing is sent.
LocalAI lists keys only to their owner, so the tab shows the signed-in
person's own keys and says so.

Account has Profile, Security, API keys and Usage. Usage shows the last 30
days, tokens by model and the limits an admin set. The Security tab now
shows for a GitHub or SSO account and says the password is not theirs to
change.

Sign-in asks for one field per step and draws a provider button only for a
provider /api/auth/status lists. It has the notice for a sign-up that waits
for approval, the first-admin screen, the key-only screen and the invite
page. An address outside the app now gets the 404 page too, which names the
address and lists the places the sidebar lists, with the same gates.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover Settings, Users and keys, Account, sign-in and the 404 page

Settings: groups, search by name, key and old section, the changed marker
only where a default is known, apply hints, the pending bar and diff, the
checks, apply sending only changed keys, undo as a second save, discard,
history, the CSRF inversion and the gallery and API key wire forms, and the
phone layout.

Users and keys: the table, filters, sort, approve, disable with undo, the row
menu, the access sheet, invites, key creation with a one-time reveal, the
ten second revoke with undo and with a page leave, and the non-admin redirect.
Account, each sign-in variant (error, pending, first admin, key-only, invite,
provider buttons) and the 404 page have specs too. Fixtures are shared with
the screenshot scripts. Existing specs follow the new structure.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the rebuilt Settings, Users and keys, Account and sign-in pages

Runtime settings: the eight groups and where each old section went, search,
the pending bar, the diff and its checks, apply, undo, the history, and which
settings show a default or an apply note and why. Authentication: the
sign-in screen variants, the Account tabs, key lifetimes, the one-time key
reveal, the revoke undo window, and the fact that keys are listed only to
their owner.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): let the Settings undo toast stand alone and read back a generated P2P token

The saved message and the undo toast sat on the same spot at the bottom of the
page. The undo toast now carries the saved message.

A new P2P token is made by the server when the page sends 0. The page reads
it back after the save so the field shows the token and not the placeholder.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): fit the users table, API keys and Account figures on a phone

On a phone the users table dropped its Role and Status columns off the screen
edge with the row actions. The role and state now sit under the name, so the
actions stay in view. API key rows no longer put the key icon on a line of its
own, and the three Account figures keep one row.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the phone users table, reduced motion and the empty Account state

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): drop the apply note from three settings the save handler does not mention

Size-aware eviction, automatic backend upgrades and development backends said
Applies now, but nothing in the handler or the docs says when they take effect.
A row now carries a note only where the code or the docs say so.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild Voices and Faces as one identity family

Voices is one page with three tabs: Speakers (voiceprints for recognising
who is speaking), Speech voices (the text-to-speech reference library,
kept apart because it is a different store) and From a recording (a link
into the diarization workspace). Faces uses the same layout.

Who is this and Same person? give the answer in a sentence with the real
distance and cut-off, a word for how far inside the cut-off it sits, and
a distance scale with the cut-off drawn on it. The cut-off slider re-reads
the answer in the browser; the identify call sends the cut-off, and verify
uses the threshold the model returns. The old confidence percentage is
gone because it is not a probability.

The server has no list call, so the people list stays in the browser and
the page says so. After a search that asked for more people than it got
back, a saved person the server did not return is marked, and people the
server returned that the browser does not know are listed. Nothing is
claimed from a short or cut-off search.

Enrolling is a sheet: sample, name, labels, permission. A copy of the
sample in the browser is opt-in, and an administrator can also keep the
recording as a speech voice in the same step. Removing a person waits ten
seconds behind an Undo toast and sends nothing before then.

Errors say what happened (no face found, model missing, call failed), a
blocked or missing microphone is explained, and a missing model or a
missing permission renders a page that says what turns the feature on
instead of a redirect. Analyze, detect and raw embedding move under
More tools, with attribute guesses off by default.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Voices and Faces pages

Specs for who is this (match, no match, working, failed, missing model),
the cut-off slider, same person, a blocked, allowed and insecure
microphone, the registry notes and the not-on-the-server marks, the
enrol sheet and its opt-in copy, delete with undo on a fake clock, the
disabled and no-permission states, the phone layout, reduced motion and
Faces. Existing library and diarization specs follow the new structure
and keep their intent. Node tests cover the distance words, scale
layout, stored list and error mapping.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the Voices and Faces pages

Add a WebUI section to the voice and face recognition pages: the two
tools, the cut-off, what the people list is and why it can be stale, the
undo window, and what is stored where. Point the Voice Library and
Fish Audio notes at Build, Voices, Speech voices.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild the Build landing, Fine-tune, Quantize, Import and Explorer

The Build landing says what each tool is for and what it needs from the
machine: the installed backend, the GPU memory, RAM and disk the server
reports, and a job that is running or the newest one when it failed. A
tool that cannot run says why and what enables it.

Fine-tune and Quantize share one page: set up, a check list that is
redrawn as the form changes, a run view with progress, stages and a log,
and a result with real next steps (export, import, chat, Models). The
checks state only what the server reports. A job needs no estimate the
server cannot make, so none is invented. Stop on a fine-tuning job asks
whether to keep a checkpoint, a failed job shows the server's message,
and a memory failure offers two changes that are applied to a copy of
the setup.

Import is a guided flow: source, review, import, done. The server
returns no preview before an import starts, so the review reads the
spelling of the source, prints the request the form will send and runs
the checks that can be made early. The estimate that arrives when the
import starts is set against free memory and disk. The ambiguity picker
and the Write YAML tab stay.

Explorer shows what GET /networks returns and lists a swarm with POST
/network/add, with a join sheet that carries the token and commands.
Build tools the account may not use say so instead of redirecting.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Build landing, the tool pages, Import and Explorer

Specs for the landing (a tool ready, missing a backend, with no GPU, a
running or failed job, a feature switched off, a member without admin,
phone, reduced motion), the shared tool pattern for Fine-tune and
Quantize (set up, live checks, start request, running with progress,
chart and log, the stop choice, failure with the server message, finish
with next steps, earlier jobs, the account-disabled page, phone), Import
(source detection, review, checks, ambiguity, running with the estimate
against free memory, done, Write YAML, phone) and Explorer (list, join,
list a swarm, empty, not an explorer, retry, phone).

Existing specs follow the new structure and keep their intent. Node
tests cover the machine facts, tool status, checks, log lines, source
detection, the import request and the join commands.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the Build tool pages, the import flow and the Explorer

Fine-tuning and quantization now describe the set up, check, run and
result steps and what the check list can and cannot say. The import
section explains the review step and why the size and memory appear only
after the import starts. The distributed page describes the Explorer
list, the join sheet and what listing a swarm publishes.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): turn the hardware recommendations into a "Best for this machine" shelf

The shelf in the Models inspector put five columns into a 400 px pane,
so long model ids wrapped letter by letter underneath the size and the
memory figures. Each row now stacks the tag, the id and the size and
memory facts beside one Install button, and the id wraps inside its own
column.

Once a model is installed the shelf narrows to the best fit and keeps
the others behind a "N more that fit" toggle. Specs cover the ranking,
the layout, the narrowing and the install request against a gallery
fixture that carries the 4K estimate the shelf sizes against.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): quiet the Studio tab markers and say their state in words

The type tabs drew a saturated green dot for every modality that has a
model. The dot now uses a text colour, filled when a model is installed
and hollow when none is, and each tab carries "(model installed)" or
"(no model installed)" as hidden text so the state is not only a colour.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): stack the model editor empty state and drop its section hues

"No fields configured" sat in a flex row, so the icon, the title and
the text ran together. It now uses the stacked empty-state layout. The
section icons took a different status colour each (amber, red, green);
they now share one quiet colour, with the accent on the current section.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): keep the Home memory sentence whole on a phone and drop side rails

On a 390 px screen the memory strip clipped "2 models loaded" to make
room for the figure. The sentence now takes the first line and the
figure and device wrap under it.

The sweep also removed coloured left rails from the editor section
rail, the skill editor list, the install strip and the audio transform
notice (now an outlined note), plus unused chat rules that carried
rails and two glow animations that nothing referenced.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): line up hub pages, list the model templates, and stop clipped text

Medium-width pages inside Build and Operate were centred while the tab
bar above them was flush left, so the title started 60 px right of the
first tab. They now start at the bar's edge.

Add Model offered nine templates as a grid of identical cards with chip
clouds and inline styles. It is now one list of rows, each with the
field names it fills in on a single muted line.

Two clipped strings are fixed: the Studio voice field cut its
placeholder mid-word, and the phone job list ended the schedule line in
an ellipsis. The recommendation shelf also separates size and memory
with a dot, and the docs describe the shelf.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): keep the hidden Studio tab state inside its tab

The hidden state text added to each type tab was absolutely positioned
against the page, so on a phone it sat outside the scrolling tab row and
widened the page by hundreds of pixels. The tab is now the containing
block.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the on-disk sizes on the Installed table, model page and cleanup sheet

The fixtures stub GET /api/models/storage. The default report is empty,
so existing specs keep the gallery estimates. makeStorage() builds a
report from files and the models that use them, the way the server
does, and storageSpec() is a models directory with shared and missing
files.

New specs cover the Size column and its shared line, the fallback when
the call fails or the user is not an admin, the files list on the model
page, a missing file, the bytes a removal frees with shared files, and
the cleanup findings. Node tests cover the storage helpers and the
batch arithmetic.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): wait for the page before pressing keys and ticking the clock

Two specs failed in loaded full runs and passed alone. The Alt+1 to
Alt+7 spec pressed a key before the composer had armed its key
handler. The capacity chart spec advanced the fake clock before the
poller had mounted, so it counted fewer readings than it expected.

Both now wait for the page to mount. The key spec retries a press that
lands during a re-render, and the clock spec advances in small steps and
polls for the row count.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): hide the Installed footer when the storage report is empty

An empty report from the storage call made the footer read "0.0 GB on
disk" next to sizes taken from the gallery estimate. An empty report
says nothing about the disk, so the footer now shows only the model
count. A spec covers it.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): keep Explore pane actions inside the pane

The inspector actions sat in a non-wrapping flex row beside the title,
so the buttons ran past the pane edge once it got narrow. The row now
takes its own line and wraps.

The primary action (Install, Retry, Open) comes first. Manage
installation becomes a ghost button, and Open details moves to the end
of the row, so one action stands out and the others are quiet. No
action or test id is removed.

Add a spec that checks, in light and dark at several widths and with a
pane forced to 320 px, that every action stays inside the pane box and
that the pane keeps its inner padding.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* refactor(ui): drop the type chips from the Studio composer

The Studio tabs and the composer's type chips listed the same seven
modes, so the page said the same thing twice. Keep the tabs as the one
place to switch modes.

The composer now shows the type it will open as a small label in its
header. The type suggestion from the typed words stays as the quiet
hint line under the prompt, and Alt+1 to Alt+7 still pick a type. The
composer root carries data-type, data-types and data-missing so tests
can read the state.

Specs pick a type through a shared Alt+digit helper and read a
missing model from the tab dot instead of a chip. Remove the unused
chip locale strings and CSS.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* refactor(ui): remove the section crumb above page titles

Page headers drew a small uppercase crumb with a short rule before it
above the title. On the hub pages it repeated the hub name, so Build
sat above a heading that also said Build.

PageHeader now renders only the title, the supporting line and the
actions. Drop the eyebrow prop, the route-derived section name, its CSS
and the unused section helper, and remove the explicit eyebrow props
from the pages that passed one. Pages stay reachable through the
sidebar and the hub tab bar.

Add a spec that checks several pages show their title with nothing
ahead of it in the header.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* refactor(ui): remove left accent rails from tiles, rows and quotes

Several surfaces marked state with a coloured strip on the left edge.
Replace each one with a cue that is not a rail:

- Stat cards lose the strip; the icon and value still carry the colour.
- The highlighted card is a raised surface with a firmer edge.
- The selected rail row is an accent wash with a hairline outline.
- The status stripe on rail items is a small status dot.
- The active failover row is a tinted row.
- Quotes in markdown and chat prose are italic instead of barred.
- The variant detail panel has a full hairline border.

Add a spec that walks the main routes in light and dark and fails on a
left border thicker than 1px, a sideways inset shadow, a narrow
absolute strip in ::before or ::after, or a narrow tall child pinned to
a left edge.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-08 22:03:07 +02:00

58 KiB
Raw Blame History

+++ disableToc = false title = "Model Configuration" weight = 23 url = '/advanced/model-configuration' +++

LocalAI uses YAML configuration files to define model parameters, templates, and behavior. This page provides a complete reference for all available configuration options.

Configuration scopes and precedence

[CLI flags and environment variables]({{% relref "reference/cli-reference" %}}) configure the LocalAI server process. Model YAML files configure one model, while supported fields in an API request can override that model's defaults for that request. For example, a request containing temperature overrides the model YAML parameters.temperature only for that request.

Precedence is setting-specific rather than one universal ordering. For the overlapping threads setting, an explicit nonzero server --threads value is applied after model YAML and therefore wins over the YAML threads value. Most server flags have no model YAML equivalent, so consult the relevant reference for the scope of each setting.

Overview

Model configuration files allow you to:

  • Define default parameters (temperature, top_p, etc.)
  • Configure prompt templates
  • Specify backend settings
  • Set up function calling
  • Configure GPU and memory options
  • And much more

Configuration File Locations

You can create model configuration files in several ways:

  1. Individual YAML files in the models directory (e.g., models/gpt-3.5-turbo.yaml)
  2. Single config file with multiple models using --models-config-file or LOCALAI_MODELS_CONFIG_FILE
  3. Remote URLs - specify a URL to a YAML configuration file at startup

Example: Basic Configuration

name: gpt-3.5-turbo
parameters:
  model: luna-ai-llama2-uncensored.ggmlv3.q5_K_M.bin
  temperature: 0.3

context_size: 512
threads: 10
backend: llama-cpp

template:
  completion: completion
  chat: chat

Example: Multiple Models in One File

When using --models-config-file, you can define multiple models as a list:

- name: model1
  parameters:
    model: model1.bin
  context_size: 512
  backend: llama-cpp

- name: model2
  parameters:
    model: model2.bin
  context_size: 1024
  backend: llama-cpp

LocalAI changes only config files that are inside the models directory. If the file from --models-config-file is outside the models directory, you cannot view, edit, pin, enable or disable its models from the web UI or the model admin API. Edit the file directly, then restart LocalAI.

Core Configuration Fields

Basic Model Settings

Field Type Description Example
name string Model name, used to identify the model in API calls gpt-3.5-turbo
backend string Backend to use (e.g. llama-cpp, vllm, diffusers, whisper) llama-cpp
description string Human-readable description of the model A conversational AI model
usage string Usage instructions or notes Best for general conversation

Model File and Downloads

Field Type Description
parameters.model string Path to the model file (relative to models directory) or URL
download_files array List of files to download. Each entry has filename, uri, and optional sha256

Example:

parameters:
  model: my-model.gguf

download_files:
  - filename: my-model.gguf
    uri: https://example.com/model.gguf
    sha256: abc123...

Model artifacts

The artifacts section makes installation of a Hugging Face model eager and repeatable. LocalAI resolves the requested revision to an immutable commit, downloads the selected repository files, and commits the complete snapshot before the model installation succeeds.

artifacts:
  - name: model
    target: model
    source:
      type: huggingface
      repo: Qwen/Qwen3-ASR-1.7B
      revision: main
      token_env: HF_TOKEN
    resolved:
      endpoint: https://huggingface.co
      revision: 0123456789abcdef0123456789abcdef01234567
      cache_key: 0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef

parameters:
  model: Qwen/Qwen3-ASR-1.7B

Declare source when authoring a configuration. LocalAI owns the resolved block and writes it after installation; do not choose its values manually. For a public repository, omit token_env. For a private or gated repository, set it to HF_TOKEN and provide that environment variable to the LocalAI controller.

Field Meaning
name Logical artifact name; model for the initial primary artifact
target Binding target; only model is supported initially
source.type huggingface
source.repo owner/repository or hf://owner/repository
source.revision Branch, tag, or commit; defaults to main and resolves to a commit
source.token_env Empty or HF_TOKEN; the secret value is never persisted
source.allow_patterns Optional slash-separated glob allow-list
source.ignore_patterns Optional slash-separated glob deny-list
resolved Installer-owned immutable endpoint, revision, and cache key

Managed installation finishes only after every selected file is committed locally. parameters.model remains the logical repository ID. Once resolved.cache_key is present, LocalAI derives .artifacts/huggingface/<cache-key>/snapshot as the runtime ModelFile. Configurations without artifacts keep the existing lazy repository-ID behavior.

The initially migrated backend families are transformers and its aliases, diffusers, qwen-asr, fish-speech, nemo, voxcpm, qwen-tts, liquid-audio, vllm, vllm-omni, and sglang. Automatic imports add artifact declarations only for this set. Compatible external backends may opt in by declaring the artifact explicitly.

Parameters Section

The parameters section contains all OpenAI-compatible request parameters and model-specific options.

OpenAI-Compatible Parameters

These settings will be used as defaults for all the API calls to the model.

Field Type Default Description
temperature float 0.9 Sampling temperature (0.0-2.0). Higher values make output more random
top_p float 0.95 Nucleus sampling: consider tokens with top_p probability mass
top_k int 40 Consider only the top K most likely tokens
max_tokens int 0 Maximum number of tokens to generate (0 = unlimited)
frequency_penalty float 0.0 Penalty for token frequency (-2.0 to 2.0)
presence_penalty float 0.0 Penalty for token presence (-2.0 to 2.0)
repeat_penalty float 1.1 Penalty for repeating tokens
repeat_last_n int 64 Number of previous tokens to consider for repeat penalty
seed int -1 Random seed (omit for random)
echo bool false Echo back the prompt in the response
n int 1 Number of completions to generate
logprobs bool/int false Return log probabilities of tokens
top_logprobs int 0 Number of top logprobs to return per token (0-20)
logit_bias map {} Map of token IDs to bias values (-100 to 100)
typical_p float 1.0 Typical sampling parameter
tfz float 1.0 Tail free z parameter
keep int 0 Number of tokens to keep from the prompt

{{% notice note %}} The DS4 backend preserves its legacy behavior for omitted or non-positive max_tokens values by generating at most 256 tokens. Set max_tokens to a positive value when you need a specific DS4 output limit. After processing the prompt, DS4 clamps that limit to the available context space and reserves one context slot for safe generation. {{% /notice %}}

Language and Translation

Field Type Description
language string Language code for transcription/translation
translate bool Whether to translate audio transcription

Custom Parameters

Field Type Description
batch int Batch size for processing
ignore_eos bool Ignore end-of-sequence tokens
negative_prompt string Negative prompt for image generation
rope_freq_base float32 RoPE frequency base
rope_freq_scale float32 RoPE frequency scale
negative_prompt_scale float32 Scale for negative prompt
tokenizer string Tokenizer to use (RWKV)

LLM Configuration

These settings apply to most LLM backends (llama.cpp, vLLM, etc.):

Performance Settings

Field Type Default Description
threads int processor count Number of threads for parallel computation. A per-model value overrides the server-wide --threads/LOCALAI_THREADS setting
context_size int 512 Maximum context size in tokens. Set to -1 to auto-use the model's full trained context from GGUF metadata (raw max, no VRAM capping; a warning is logged if it may not fit detected VRAM).
f16 bool false Enable 16-bit floating point precision (GPU acceleration)
gpu_layers int 99999999 Number of layers to offload to GPU. The default requests all layers; 0 keeps model layers on CPU. See mixed CPU/GPU inference.

Memory Management

Field Type Default Description
mmap bool true Use memory mapping for model loading (faster, less RAM)
mmlock bool false Lock model in memory (prevents swapping)
low_vram bool false Use minimal VRAM mode
no_kv_offloading bool false Disable KV cache offloading

GPU Configuration

Field Type Description
tensor_split string Comma-separated GPU memory allocation (e.g., "0.8,0.2" for 80%/20%)
main_gpu string Main GPU identifier for multi-GPU setups
cuda bool Explicitly enable/disable CUDA

Placement

The model page's Configuration tab and the model editor have a Placement section for the keys above. It edits gpu_layers, tensor_split, main_gpu and context_size and nothing else.

  • CPU only writes gpu_layers: 0.
  • Auto leaves gpu_layers unset. LocalAI then asks the backend for every layer (99999999), and the llama-cpp backend lowers that to what fits in the free device memory unless the model turns off fit_params (see GPU auto-fit settings). Other backends read an unset value as their own default.
  • Custom writes the number you enter. All layers writes 99999999, the value LocalAI uses when the key is unset. A slider appears when the memory estimate reports the model's layer count (block_count); today it does not, so you enter a number.
  • With two or more GPUs, Split across GPUs writes tensor_split as percentages (for example 65,35; llama.cpp reads them as proportions) and Main GPU writes main_gpu. Split by free memory sets the shares from the memory each card has free now.

The bars show, for each GPU and for system memory, what other programs use, what this model needs (weights and working memory, and the KV cache that grows with context_size) and what is left. They come from the device list in /api/resources and from POST /api/models/vram-estimate, which returns one total for the chosen context size and layer count. The per-device split of that total follows tensor_split, so it is an approximation. The page says what the choice means in words ("Fits in GPU", "Spills to CPU", "Too many layers for the GPU") and does not state a speed. A usable limit is 95 percent of the free memory.

Fit it for me searches for the largest gpu_layers whose estimate fits the free GPU memory, by asking the estimate endpoint, and Undo puts the previous value back. It is hidden when the estimate is unavailable (the model file is missing, still downloading or in a format the estimate does not cover) and when the host has no GPU. A server that schedules models onto other machines shows no per-device bars, because the devices listed belong to the controller.

Saving from the model page sends only the keys that changed. A key set back to Auto is sent as null. The patch endpoint merges values into the file and cannot delete a key, so the file then reads gpu_layers: null, which LocalAI treats as unset. To remove the line itself, edit the YAML.

Mixed CPU/GPU inference

The llama-cpp backend can run one GGUF model across CPU and GPU, using both system RAM and GPU VRAM. Use a GPU-capable build of the backend for your hardware. A CPU-only build cannot offload layers to the GPU.

Offload some model layers

Set gpu_layers to a positive number smaller than the model's layer count. The remaining layers run on CPU. Merge these settings into your existing model YAML, keeping its model path, template, and other options:

backend: llama-cpp
gpu_layers: 12
context_size: 4096

The value 12 is an example, not a memory estimate. Reload the model after changing its configuration. Check the backend startup log for the number of layers offloaded and the CPU/GPU buffer sizes. Increase gpu_layers if VRAM has room; reduce it if loading runs out of GPU memory. Set gpu_layers: 0 to keep all model layers on CPU.

Keep MoE experts on CPU

For a mixture-of-experts (MoE) model, you can keep expert weights in system RAM while offloading other tensors to the GPU:

backend: llama-cpp
gpu_layers: 99999999
context_size: 4096
options:
  - cpu_moe:true

Append cpu_moe:true to any existing options list instead of replacing that list. This option applies to the main model's expert weights. To keep experts from only the first 12 layers on CPU, replace cpu_moe:true with n_cpu_moe:12. Use one of these options at a time.

CPU execution and data transfers can reduce generation speed compared with a model that fits entirely on GPU. RAM and VRAM do not form one interchangeable allocation pool. Leave memory for the KV cache, compute buffers, the operating system, and other processes. Reduce context_size if the KV cache consumes too much memory. The GPU auto-fit settings provide a separate way to let llama.cpp choose the allocation.

Sampling and Generation

Field Type Default Description
mirostat int 0 Mirostat sampling mode (0=disabled, 1=Mirostat, 2=Mirostat 2.0)
mirostat_tau float 5.0 Mirostat target entropy
mirostat_eta float 0.1 Mirostat learning rate

LoRA Configuration

Field Type Description
lora_adapter string Path to LoRA adapter file
lora_base string Base model for LoRA
lora_scale float32 LoRA scale factor
lora_adapters array Multiple LoRA adapters
lora_scales array Scales for multiple LoRA adapters

Advanced Options

Field Type Description
no_mulmatq bool Disable matrix multiplication queuing
draft_model string Draft model GGUF file for speculative decoding (see Speculative Decoding)
n_draft int32 Maximum number of draft tokens per speculative step (default: 16)
quantization string Quantization format
load_format string Model load format
numa bool Enable NUMA (Non-Uniform Memory Access)
rms_norm_eps float32 RMS normalization epsilon
ngqa int32 Natural question generation parameter
rope_scaling string RoPE scaling configuration
type string Model type/architecture
grammar string Grammar file path for constrained generation

YARN Configuration

YARN (Yet Another RoPE extensioN) settings for context extension:

Field Type Description
yarn_ext_factor float32 YARN extension factor
yarn_attn_factor float32 YARN attention factor
yarn_beta_fast float32 YARN beta fast parameter
yarn_beta_slow float32 YARN beta slow parameter

Speculative Decoding

Speculative decoding speeds up text generation by predicting multiple tokens ahead and verifying them in a single forward pass. The output is identical to normal decoding - only faster. This feature is only available with the llama-cpp backend.

There are two approaches:

Draft Model Speculative Decoding

Uses a smaller, faster model from the same model family to draft candidate tokens, which the main model then verifies. Requires a separate GGUF file for the draft model.

name: my-model
backend: llama-cpp
parameters:
  model: large-model.gguf
draft_model: small-draft-model.gguf
n_draft: 8
options:
  - spec_p_min:0.8
  - draft_gpu_layers:99

N-gram Self-Speculative Decoding

Uses patterns from the token history to predict future tokens - no extra model required. Works well for repetitive or structured output (code, JSON, lists).

name: my-model
backend: llama-cpp
parameters:
  model: my-model.gguf
options:
  - spec_type:ngram_simple
  - spec_n_max:16

Speculative Decoding Options

These are set via the options: array in the model configuration (format: key:value):

Common options

Option Type Default Description
spec_type / speculative_type string none Speculative decoding type, or comma-separated list to chain multiple (see table below)
spec_n_max / draft_max int 16 Maximum number of tokens to draft per step
spec_n_min / draft_min int 0 Minimum draft tokens required to use speculation
spec_p_min / draft_p_min float 0.75 Minimum probability threshold for greedy acceptance
spec_p_split float 0.1 Split probability for tree-based branching

Draft-model options (apply when spec_type=draft, i.e. a draft_model is configured)

Option Type Default Description
draft_gpu_layers int -1 GPU layers for the draft model (-1 = use default)
draft_threads / spec_draft_threads int same as main Threads used by the draft model (<= 0 = hardware concurrency)
draft_threads_batch / spec_draft_threads_batch int same as draft_threads Threads used by the draft model during batch / prompt processing
draft_cache_type_k / spec_draft_cache_type_k string f16 KV cache K data type for the draft model (same values as cache_type_k)
draft_cache_type_v / spec_draft_cache_type_v string f16 KV cache V data type for the draft model
draft_cpu_moe / spec_draft_cpu_moe bool false Keep all MoE expert weights of the draft model on CPU
draft_n_cpu_moe / spec_draft_n_cpu_moe int 0 Keep MoE expert weights of the first N draft-model layers on CPU
draft_override_tensor / spec_draft_override_tensor string "" Comma-separated <tensor regex>=<buffer type> overrides for the draft model
draft_ctx_size int (ignored) Deprecated upstream: the draft now shares the target context size. Accepted for backward compatibility but has no effect.

ngram_simple options (used when spec_type includes ngram_simple)

Option Type Default Description
spec_ngram_size_n / ngram_size_n int 12 N-gram lookup size
spec_ngram_size_m / ngram_size_m int 48 M-gram proposal size
spec_ngram_min_hits / ngram_min_hits int 1 Minimum hits for accepting n-gram proposals

ngram_mod options (used when spec_type includes ngram_mod)

Option Type Default Description
spec_ngram_mod_n_min int 48 Minimum number of ngram tokens to use
spec_ngram_mod_n_max int 64 Maximum number of ngram tokens to use
spec_ngram_mod_n_match int 24 Ngram lookup length

ngram_map_k options (used when spec_type includes ngram_map_k)

Option Type Default Description
spec_ngram_map_k_size_n int 12 N-gram lookup size
spec_ngram_map_k_size_m int 48 M-gram proposal size
spec_ngram_map_k_min_hits int 1 Minimum hits for accepting proposals

ngram_map_k4v options (used when spec_type includes ngram_map_k4v)

Option Type Default Description
spec_ngram_map_k4v_size_n int 12 N-gram lookup size
spec_ngram_map_k4v_size_m int 48 M-gram proposal size
spec_ngram_map_k4v_min_hits int 1 Minimum hits for accepting proposals

ngram_cache lookup files

Option Type Default Description
spec_lookup_cache_static / lookup_cache_static string "" Path to a static ngram lookup cache file
spec_lookup_cache_dynamic / lookup_cache_dynamic string "" Path to a dynamic ngram lookup cache file (updated by generation)

Speculative Type Values

The canonical names match upstream llama.cpp (dash-separated). For backward compatibility LocalAI also accepts the underscore-separated forms and the bare draft / eagle3 aliases.

Type Aliases accepted Description
none No speculative decoding (default)
draft-simple draft, draft_simple Draft model-based speculation (auto-set when draft_model is configured)
draft-eagle3 eagle3, draft_eagle3 EAGLE3 draft model architecture
draft-mtp draft_mtp Multi-Token Prediction. Reuses the target model's embedded MTP head; no separate draft GGUF required (draft_model can be omitted).
ngram-simple ngram_simple Simple self-speculative using token history
ngram-map-k ngram_map_k N-gram with key-only map
ngram-map-k4v ngram_map_k4v N-gram with keys and 4 m-gram values
ngram-mod ngram_mod Modified n-gram speculation
ngram-cache ngram_cache 3-level n-gram cache

Multiple types can be chained by passing a comma-separated list to spec_type (e.g. spec_type:ngram-simple,ngram-mod). The runtime tries them in order and accepts the first proposal that meets the acceptance criteria.

{{% notice note %}} The current LocalAI llama.cpp backend supports speculative decoding with multimodal models that load an mmproj, including MTP. LocalAI passes both configurations to llama.cpp and does not disable speculation merely because an mmproj is present. Upstream llama.cpp removed the former general multimodal/speculative restriction in ggml-org/llama.cpp#19493; ggml-org/llama.cpp#22673 later added MTP support and explicitly documented its compatibility with vision input.

Compatibility still depends on the installed backend version and the target/draft model architecture. Check the backend logs for successful projector loading and speculative-context initialization, then look for the draft acceptance statistics line and its accepted / generated counts. A representative run with zero accepted draft tokens receives no speculative speedup and can indicate that the model or settings need tuning. {{% /notice %}}

Multi-Token Prediction (MTP)

draft-mtp enables Multi-Token Prediction (ggml-org/llama.cpp#22673). MTP uses a small prediction head trained into the target model: the head runs alongside the main forward pass and proposes the next few tokens, which the target then verifies in a single batched step. Upstream reports ~1.85x-2.1x token throughput at ~72-82% draft acceptance on Qwen3.6 27B / 35B A3B.

Auto-detection (default). When a GGUF declares an MTP head (the upstream <arch>.nextn_predict_layers metadata key, set by convert_hf_to_gguf.py for Qwen3.5/3.6 family models and similar), LocalAI auto-enables MTP with the following defaults:

options:
  - spec_type:draft-mtp
  - spec_n_max:6
  - spec_p_min:0.75

Detection runs both at import time (the /import-model UI / POST /models/import-uri flow range-fetches the GGUF header and writes the options into the generated YAML before you save it) and at load time (every llama-cpp model start re-checks the local header and appends the options if spec_type isn't already set). To opt out, set an explicit spec_type: / speculative_type: in your YAML - auto-detection always preserves the user value, including spec_type:none.

Two ways to load the MTP head:

  1. Embedded in the target GGUF (the recommended path for LocalAI, and what auto-detection assumes). When spec_type includes draft-mtp and draft_model is empty, the backend builds the MTP draft context directly from the target model's weights. The GGUF must have been converted with the MTP tensors included.
  2. Separate mtp-*.gguf sibling file. If you point draft_model at the separate MTP-head GGUF that ships next to the main weights on HuggingFace, the backend will load it as a draft model. Note: upstream's -hf auto-discovery of mtp-*.gguf siblings is not wired into LocalAI's gRPC layer - you need to download the sibling file and configure draft_model explicitly.

Manual override knobs (overlap with the auto-detect defaults above):

Option Recommended Notes
spec_type draft-mtp Activates MTP. Can be chained with other types (see below).
spec_n_max / draft_max 2-6 Number of draft tokens per step. Upstream's PR suggests 2-3 for the tightest acceptance window; LocalAI's auto-default is 6 to favour throughput on models with high acceptance.
spec_p_min 0.75 Pinned because upstream marks the current default with a "change to 0.0f" TODO; locking it here keeps acceptance thresholds stable across future llama.cpp bumps.
mmproj_use_gpu true for vision MTP does not require disabling the projector. Keep mmproj configured for image input; set this option to false to keep the projector on CPU when VRAM is tight. Remove mmproj only for text-only use when vision is not needed.

Minimal config (override-only, since auto-detection already covers this for MTP-capable GGUFs):

name: qwen3-mtp
backend: llama-cpp
parameters:
  model: qwen3-27b-with-mtp.gguf
options:
  - spec_type:draft-mtp
  - spec_n_max:3

With vision enabled:

name: qwen3-vision-mtp
backend: llama-cpp
known_usecases:
  - chat
  - vision
parameters:
  model: qwen3-with-mtp.gguf
mmproj: mmproj-qwen3.gguf
options:
  - spec_type:draft-mtp
  - spec_n_max:3
  - spec_p_min:0.75

With a separate MTP head file:

name: qwen3-mtp
backend: llama-cpp
parameters:
  model: qwen3-27b.gguf
  draft_model: qwen3-27b-mtp-head.gguf
options:
  - spec_type:draft-mtp
  - spec_n_max:3

Chaining MTP with n-gram fallback (experimental, from the PR's usage notes - useful when MTP acceptance drops on highly repetitive output):

options:
  - spec_type:draft-mtp,ngram-mod
  - spec_n_max:3
  - spec_ngram_mod_n_match:24

Pre-converted GGUFs with MTP heads are published on the ggml-org HuggingFace org (initially Qwen3.6 27B and Qwen3.6 35B A3B).

Reasoning Models (DeepSeek-R1, Qwen3, etc.)

These load-time options control how the backend parses <think> reasoning blocks and how much budget the model is allowed for thinking. They are set per model via the options: array. For how reasoning is returned alongside tool calls and survives the tool-result round trip, see [Interleaved Thinking with Tool Calls]({{%relref "features/interleaved-thinking" %}}).

Option Type Default Description
reasoning_format string deepseek Parser for reasoning/thinking blocks. One of none, auto, deepseek, deepseek-legacy (alias deepseek_legacy).
enable_reasoning / reasoning_budget int -1 Reasoning budget in tokens: -1 unlimited, 0 disabled, >0 token cap for the thinking section.
prefill_assistant bool true When false, the trailing assistant message is not pre-filled by the chat template.

{{% notice note %}} This is the load-time reasoning configuration. The orthogonal per-request enable_thinking chat-template kwarg toggles thinking on/off per call without restarting the model. It can be driven either by the YAML reasoning.disable field (model default) or per request via the OpenAI reasoning_effort field on /v1/chat/completions:

  • reasoning_effort: "none" disables thinking for that request (enable_thinking=false) - useful to run a single reasoning model like Qwen3 for low-latency tasks while still enabling reasoning on other requests.
  • reasoning_effort: "minimal" | "low" | "medium" | "high" enables thinking, unless the model config explicitly set reasoning.disable: true (an operator's explicit disable wins and is never re-enabled by a request). {{% /notice %}}

reasoning_effort as a chat-template kwarg

reasoning_effort is also forwarded to the backend as a chat_template_kwarg, so models whose jinja chat template keys on it - e.g. gpt-oss (Harmony) or LFM2.5 - honor the level, not just the on/off enable_thinking flag. This matters for models that ignore enable_thinking entirely (LFM2.5 keeps emitting <think> for enable_thinking=false, but respects reasoning_effort).

Set a per-model default in the config so every request inherits it (a per-request reasoning_effort still overrides):

name: my-model
reasoning_effort: none   # none | minimal | low | medium | high

For [realtime pipelines]({{%relref "features/openai-realtime" %}}), set it on the pipeline so it applies to the pipeline's LLM without editing that model's own config:

name: gpt-realtime
pipeline:
  llm: lfm2.5
  reasoning_effort: none   # overrides the LLM model's own reasoning_effort

Custom chat_template_kwargs

Some jinja chat templates expose extra variables beyond enable_thinking / reasoning_effort (for example Qwen3's preserve_thinking). Set arbitrary key/values in the model config and they are forwarded to the backend's chat_template_kwargs as-is, so you don't need a dedicated server option per template variable:

name: qwen3
chat_template_kwargs:
  preserve_thinking: true

You can also override (or add) any of these per request through the OpenAI metadata field on /v1/chat/completions. Values are strings; "true" / "false" are coerced to booleans, anything else is passed through as a string:

{
  "model": "qwen3",
  "messages": [{"role": "user", "content": "hi"}],
  "metadata": { "preserve_thinking": "true", "enable_thinking": "false" }
}

Per-request metadata overrides the model config defaults and the reasoning-config levers, and (for enable_thinking / reasoning_effort) takes effect across every backend that reads them, not just llama.cpp. Typed (non-boolean) values are only supported through the model YAML chat_template_kwargs, where YAML preserves the type.

Multimodal Backend Options

Option Type Default Description
mmproj_use_gpu / mmproj_offload bool true Set false to keep the multimodal projector on CPU (saves VRAM at cost of speed).
image_min_tokens int -1 Minimum vision tokens per image. -1 keeps the model default.
image_max_tokens int -1 Maximum vision tokens per image. -1 keeps the model default.

Embedding & Reranking Backend Options

Option Type Default Description
pooling_type / pooling string auto Pooling strategy for embeddings: none, mean, cls, last, rank. Reranking automatically uses rank.
embd_normalize / embedding_normalize int 2 Normalization: -1 none, 0 max-abs, 1 taxicab, 2 Euclidean (L2), >2 p-norm.

Other Backend Tuning Options

These llama.cpp options are passed through the options: array.

Option Type Default Description
n_ubatch / ubatch int same as batch Physical batch size. Decouple from n_batch when an embedding/rerank workload needs a different value.
threads_batch / n_threads_batch int same as threads Threads used during prompt processing. <= 0 means hardware_concurrency().
direct_io / use_direct_io bool false Open the model with O_DIRECT (faster cold loads on NVMe; ignored if not supported).
verbosity int 3 llama.cpp internal log verbosity threshold. Higher = more verbose.
device / devices string all devices Select the llama.cpp backend devices to use. Repeat the option or pass a comma-separated list; unlisted devices are excluded. Use the names reported by llama-server --list-devices / --list-devices.
override_tensor / tensor_buft_overrides string "" Per-tensor buffer-type overrides for the main model. Format: <tensor regex>=<buffer type>,<tensor regex>=<buffer type>,.... Mirrors the existing draft_override_tensor syntax for the draft model.
cpu_moe bool false Keep all MoE expert weights of the main model on CPU (upstream --cpu-moe). Frees VRAM on large MoE models (DeepSeek, Qwen3 *-A3B).
n_cpu_moe int 0 Keep MoE expert weights of the first N main-model layers on CPU (upstream --n-cpu-moe).

Generic option passthrough

Any options: entry whose name starts with - is forwarded verbatim to upstream llama.cpp's own llama-server argument parser. This means any flag the bundled llama.cpp supports works without LocalAI needing a dedicated option, even ones added after your LocalAI version was built. See the upstream server flags reference.

Format mirrors the rest of the array - --flag for a boolean, or --flag:value for a flag that takes a value. Everything after the first : is the value, so embedded colons (e.g. host:port) are preserved:

options:
  - "--cpu-moe"                 # boolean flag
  - "--n-cpu-moe:4"             # flag with a value
  - "--override-tensor:exps=CPU"
  - "devices:CUDA1,CUDA2,CUDA3" # skip CUDA0, e.g. a display GPU

Notes:

  • Precedence: passthrough flags are applied last, so an explicit flag overrides the LocalAI option it maps to (e.g. --ctx-size:8192 overrides context_size).
  • Power-user territory: an invalid flag or value is rejected by the upstream parser exactly as it would be by llama-server, which can fail model loading. Prefer the named options above when one exists.
  • Flags that would terminate the process (such as --help, --usage, --version, --license, --list-devices, --cache-list, and --completion*) are ignored.

Prompt Caching

The recommended way to enable prompt caching for the llama-cpp backend is the server-side prompt cache controlled by cache_ram / kv_unified / cache_idle_slots in the options: array (see [llama.cpp backend options]({{%relref "features/text-generation#server-side-prompt-cache-repeated-system-prompts" %}})). It's on by default since LocalAI v4.3 and is what gives repeated system prompts a near-zero prefill on the second call.

The fields below come from upstream llama.cpp's CLI completion tool and are passed through to the gRPC backend for compatibility, but the gRPC server itself does not consume them: keep them empty unless you're targeting a non-llama-cpp backend that reads them.

Field Type Description
prompt_cache_path string (legacy / unused by llama-cpp gRPC server) Path to a file-backed prompt cache for upstream's CLI completion tool.
prompt_cache_all bool (legacy / unused by llama-cpp gRPC server)
prompt_cache_ro bool (legacy / unused by llama-cpp gRPC server)

Text Processing

Field Type Description
stopwords array Words or phrases that stop generation
cutstrings array Strings to cut from responses
trimspace array Strings to trim whitespace from
trimsuffix array Suffixes to trim from responses
extract_regex array Regular expressions to extract content

System Prompt

Field Type Description
system_prompt string Default system prompt for the model

vLLM-Specific Configuration

These options apply when using the vllm backend:

Field Type Description
gpu_memory_utilization float32 GPU memory utilization (0.0-1.0, default 0.9)
trust_remote_code bool Trust and execute remote code
enforce_eager bool Force eager execution mode
swap_space int Swap space in GB
max_model_len int Maximum model length
tensor_parallel_size int Tensor parallelism size
disable_log_stats bool Disable logging statistics
dtype string Data type (e.g., float16, bfloat16)
flash_attention string Flash attention configuration
cache_type_k string Key cache quantization type. Maps to llama.cpp's -ctk. Accepted values for llama.cpp-family backends (llama-cpp, ik-llama-cpp, turboquant): f16, f32, q8_0, q4_0, q4_1, q5_0, q5_1. The turboquant backend additionally accepts turbo2, turbo3, turbo4 - the fork's TurboQuant KV-cache schemes. turbo3/turbo4 auto-enable flash_attention.
cache_type_v string Value cache quantization type. Maps to llama.cpp's -ctv. Same accepted values as cache_type_k. Note: any quantized V cache requires flash_attention to be enabled.
limit_mm_per_prompt object Limit multimodal content per prompt: {image: int, video: int, audio: int}

Template Configuration

Templates use Go templates with Sprig functions.

Field Type Description
template.chat string Template for chat completion endpoint
template.chat_message string Template for individual chat messages
template.completion string Template for text completion
template.edit string Template for edit operations
template.function string Template for function/tool calls
template.multimodal string Template for multimodal interactions
template.reply_prefix string Prefix to add to model replies
template.use_tokenizer_template bool Use tokenizer's built-in template (vLLM/transformers)
template.system_messages_after_first string What to do with system-role messages that appear after the leading system block: merge folds them into the first system message, user forwards them as user-role turns at their position. Unset keeps them as-is. Needed for tokenizer templates that reject late system turns (e.g. Qwen3.8) while agent frameworks append instructions mid-conversation.
template.join_chat_messages_by_character string Character to join chat messages (default: \n)

Template Variables

Templating supports sprig functions.

Following are common variables available in templates:

  • {{.Input}} - User input
  • {{.Instruction}} - Instruction for edit operations
  • {{.System}} - System message
  • {{.Prompt}} - Full prompt
  • {{.Functions}} - Function definitions (for function calling)
  • {{.FunctionCall}} - Function call result

Example Template

template:
  chat: |
    {{.System}}
    {{range .Messages}}
    {{if eq .Role "user"}}User: {{.Content}}{{end}}
    {{if eq .Role "assistant"}}Assistant: {{.Content}}{{end}}
    {{end}}
    Assistant:

Function Calling Configuration

Configure how the model handles function/tool calls:

Field Type Default Description
function.disable_no_action bool false Disable the no-action behavior
function.no_action_function_name string answer Name of the no-action function
function.no_action_description_name string Description for no-action function
function.function_name_key string name JSON key for function name
function.function_arguments_key string arguments JSON key for function arguments
function.response_regex array Named regex patterns to extract function calls
function.argument_regex array Named regex to extract function arguments
function.argument_regex_key_name string key Named regex capture for argument key
function.argument_regex_value_name string value Named regex capture for argument value
function.json_regex_match array Regex patterns to match JSON in tool mode
function.replace_function_results array Replace function call results with patterns
function.replace_llm_results array Replace LLM results with patterns
function.capture_llm_results array Capture LLM results as text (e.g., for "thinking" blocks)

Grammar Configuration

Field Type Default Description
function.grammar.disable bool false Completely disable grammar enforcement
function.grammar.parallel_calls bool false Allow parallel function calls
function.grammar.mixed_mode bool false Allow mixed-mode grammar enforcing
function.grammar.no_mixed_free_string bool false Disallow free strings in mixed mode
function.grammar.disable_parallel_new_lines bool false Disable parallel processing for new lines
function.grammar.prefix string Prefix to add before grammar rules
function.grammar.expect_strings_after_json bool false Expect strings after JSON data

Diffusers Configuration

For image generation models using the diffusers backend:

Field Type Description
diffusers.cuda bool Force CUDA. By default the backend auto-detects and uses CUDA when a compatible GPU is present (ROCm builds included). Pin the CPU with options: ["device:cpu"]
diffusers.pipeline_type string Pipeline type (e.g., stable-diffusion, stable-diffusion-xl)
diffusers.scheduler_type string Scheduler type (e.g., euler, ddpm)
diffusers.original_config_file string Local path or URL to the original configuration for loading a single-file checkpoint
diffusers.enable_parameters string Comma-separated parameters to enable
diffusers.cfg_scale float32 Classifier-free guidance scale
diffusers.img2img bool Enable image-to-image transformation
diffusers.clip_skip int Number of CLIP layers to skip
diffusers.clip_model string CLIP model to use
diffusers.clip_subfolder string CLIP model subfolder
diffusers.control_net string ControlNet model to use
step int Number of diffusion steps

TTS Configuration

For text-to-speech models:

Field Type Description
tts.voice string Default backend voice ID, speaker name, or reference path. A request voice takes precedence.
tts.audio_path string Default reference-audio path for cloning backends. A request voice or saved Voice Library profile takes precedence.
tts.voice_cloning bool Optional Voice Library capability override. Omit for automatic backend/variant detection; true opts in a verified custom-named variant and false rejects saved profile references.

For example, a custom-named model on a known cloning backend can declare support explicitly while retaining a model-wide reference fallback:

name: private-voice-model
backend: qwen3-tts-cpp
parameters:
  model: private/qwen-talker-base.gguf
known_usecases:
  - tts
tts:
  voice_cloning: true
  audio_path: voices/default-reference.wav

tts.voice_cloning: true only overrides model-variant detection. It cannot enable cloning on a backend that does not implement LocalAI's reference-audio contract.

Roles Configuration

Map conversation roles to specific strings:

roles:
  user: "### Instruction:"
  assistant: "### Response:"
  system: "### System Instruction:"

Feature Flags

Enable or disable experimental features:

feature_flags:
  feature_name: true
  another_feature: false

MCP Configuration

Model Context Protocol (MCP) configuration:

Field Type Description
mcp.remote string YAML string defining remote MCP servers
mcp.stdio string YAML string defining STDIO MCP servers

Agent Configuration

Agent/autonomous agent configuration:

Field Type Description
agent.max_attempts int Maximum number of attempts
agent.max_iterations int Maximum number of iterations
agent.enable_reasoning bool Enable reasoning capabilities
agent.enable_planning bool Enable planning capabilities
agent.enable_mcp_prompts bool Enable MCP prompts
agent.enable_plan_re_evaluator bool Enable plan re-evaluation

Reasoning Configuration

Configure how reasoning tags are extracted and processed from model output. Reasoning tags are used by models like DeepSeek, Command-R, and others to include internal reasoning steps in their responses.

Field Type Default Description
reasoning.disable bool false When true, disables reasoning extraction entirely. The original content is returned without any processing.
reasoning.disable_reasoning_tag_prefill bool false When true, disables automatic prepending of thinking start tokens. Use this when your model already includes reasoning tags in its output format.
reasoning.strip_reasoning_only bool false When true, extracts and removes reasoning tags from content but discards the reasoning text. Useful when you want to clean reasoning tags from output without storing the reasoning content.
reasoning.thinking_start_tokens array [] List of custom thinking start tokens to detect in prompts. Custom tokens are checked before default tokens.
reasoning.tag_pairs array [] List of custom tag pairs for reasoning extraction. Each entry has start and end fields. Custom pairs are checked before default pairs.

Reasoning Tag Formats

The reasoning extraction supports multiple tag formats used by different models:

  • <thinking>...</thinking> - General thinking tag
  • <think>...</think> - DeepSeek, Granite, ExaOne, GLM models
  • <|START_THINKING|>...<|END_THINKING|> - Command-R models
  • <|inner_prefix|>...<|inner_suffix|> - Apertus models
  • <seed:think>...</seed:think> - Seed models
  • <|think|>...<|end|><|begin|>assistant<|content|> - Solar Open models
  • [THINK]...[/THINK] - Magistral models

Examples

Disable reasoning extraction:

reasoning:
  disable: true

Extract reasoning but don't prepend tags:

reasoning:
  disable_reasoning_tag_prefill: true

Strip reasoning tags without storing reasoning content:

reasoning:
  strip_reasoning_only: true

Complete example with reasoning configuration:

name: deepseek-model
backend: llama-cpp
parameters:
  model: deepseek.gguf

reasoning:
  disable: false
  disable_reasoning_tag_prefill: false
  strip_reasoning_only: false

Example with custom tokens and tag pairs:

name: custom-reasoning-model
backend: llama-cpp
parameters:
  model: custom.gguf

reasoning:
  thinking_start_tokens:
    - "<custom:think>"
    - "<my:reasoning>"
  tag_pairs:
    - start: "<custom:think>"
      end: "</custom:think>"
    - start: "<my:reasoning>"
      end: "</my:reasoning>"

Note: Custom tokens and tag pairs are checked before the default ones, giving them priority. This allows you to override default behavior or add support for new reasoning tag formats.

Per-Request Override via Metadata

The reasoning.disable setting from model configuration can be overridden on a per-request basis using the metadata field in the OpenAI chat completion request. This allows you to enable or disable thinking for individual requests without changing the model configuration.

The metadata field accepts a map[string]string that is forwarded to the backend. The enable_thinking key controls thinking behavior:

# Enable thinking for a single request (overrides model config)
curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3",
    "messages": [{"role": "user", "content": "Explain quantum computing"}],
    "metadata": {"enable_thinking": "true"}
  }'

# Disable thinking for a single request (overrides model config)
curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3",
    "messages": [{"role": "user", "content": "Hello"}],
    "metadata": {"enable_thinking": "false"}
  }'

Priority order:

  1. Request-level metadata.enable_thinking (highest priority)
  2. Model config reasoning.disable (fallback)
  3. Auto-detected from model template (default)

Pipeline Configuration

Define pipelines for audio-to-audio processing and the [Realtime API]({{%relref "features/openai-realtime" %}}):

Field Type Description
pipeline.tts string TTS model name
pipeline.llm string LLM model name
pipeline.transcription string Transcription model name
pipeline.vad string Voice activity detection model name
pipeline.turn_detection object Realtime turn-detection defaults. Keys: type (server_vad/semantic_vad), eagerness (low/medium/high/auto), retranscribe, vad_window_sec (widen the per-tick VAD scan window; values below the automatic floor are ignored). See [Realtime turn detection]({{%relref "features/openai-realtime" %}})
pipeline.classifier object Realtime classifier mode: prefill-scored option selection instead of generation. Keys: enabled, threshold, normalization (raw/mean), history_items, fallback (mode: none/reply/generate, reply), options (list of id, description, reply, tool {name, arguments}), address (wake-word gate: names, mode: ignore/reply, reply), model (optional separate scoring config). See [Realtime classifier mode]({{%relref "features/openai-realtime#classifier-mode-localai-extension" %}})

gRPC Configuration

Backend gRPC communication settings. These control the readiness handshake between LocalAI and a freshly spawned backend process - LocalAI polls the backend's Health gRPC method up to grpc.attempts times, sleeping grpc.attempts_sleep_time seconds between polls, before giving up and terminating the backend as unresponsive.

Field Type Default Description
grpc.attempts int 20 Number of health-check attempts before the backend is killed as unresponsive
grpc.attempts_sleep_time int 2 Sleep time between health-check attempts (seconds)

Total load window ≈ grpc.attempts × (grpc.attempts_sleep_time + per-call gRPC dial timeout). The default of 20 × 2 s ≈ 40 s is fine for typical backends but is too short for large models that need substantial time to become gRPC-ready after the process starts - for example NVFP4 / FP8 models whose shard loading and CUDA-graph capture can take several minutes, or slow storage backends. If the backend keeps getting killed while still legitimately loading (visible as exitCode=120 + rpc error: code = Canceled desc = context canceled in the LocalAI log, while the backend's own stderr shows continued forward progress), raise these values.

Example configuration for a model that needs up to ~10 minutes to become gRPC-ready (large NVFP4 model, cold shard load + CUDA-graph capture):

grpc:
  attempts: 140
  attempts_sleep_time: 5

This gives a ~700 s window while keeping health-check polling frequent enough to detect real backend crashes quickly. The values only affect the initial readiness handshake - inference-request timeouts and the watchdog are unchanged.

Overrides

Override model configuration values at runtime (llama.cpp):

overrides:
  - "qwen3moe.expert_used_count=int:10"
  - "some_key=string:value"

Format: KEY=TYPE:VALUE where TYPE is int, float, string, or bool.

Known Use Cases

Specify which endpoints this model supports:

known_usecases:
  - chat
  - completion
  - embeddings

Available flags: chat, completion, edit, embeddings, rerank, image, transcript, tts, sound_generation, tokenize, vad, video, detection, score, token_classify, decisions, llm (combination of CHAT, COMPLETION, EDIT).

decisions marks a model as a decision model for the [Decisions API]({{% relref "features/decisions" %}}) (POST /v1/systemone). It is never guessed, and a model that declares it is not listed as a chat, completion or embeddings model.

token_classify marks a model as a token-classification (NER) provider for the PII filter (e.g. an openai-privacy-filter GGUF). Declare it explicitly together with embeddings: true (the classifier loads via TOKEN_CLS pooling). It runs on the dedicated privacy-filter backend (backend/cpp/privacy-filter), a standalone GGML engine for the openai-privacy-filter family - separate from llama-cpp, which no longer carries the token-classification path.

Known input and output modalities

Use known_input_modalities and known_output_modalities when a use case does not fully describe a model's I/O. For example, both text-to-video and audio-driven avatar models use the video use case, but only the avatar model accepts audio:

known_usecases:
  - video
known_input_modalities:
  - text
  - image
  - audio
known_output_modalities:
  - video

Valid modality values are text, image, audio, and video. Explicit values are combined with modalities LocalAI can infer from the model use cases and configuration. The resulting canonical, de-duplicated lists are exposed by GET /v1/models/capabilities.

PII filtering

PII redaction is NER-based and runs on the request (input) side. It has two halves:

  • Detector models are token_classify models that carry the detection policy in a top-level pii_detection: block. The policy is defined once, on the model itself:

    name: privacy-filter-multilingual
    backend: llama-cpp
    embeddings: true
    known_usecases:
      - token_classify
    pii_detection:
      min_score: 0.5            # drop detections below this confidence
      default_action: mask      # mask | block | allow - applied to any detected
                                # group with no explicit entry (empty = mask)
      entity_actions:           # which PII to block vs mask vs allow-log
        PASSWORD: block
        CREDITCARD: block
        EMAIL: mask
    
  • Consuming models opt in and reference one or more detectors by name - no per-consumer policy:

    name: my-assistant
    pii:
      enabled: true             # default: off for local backends, on for cloud-proxy
      detectors:
        - privacy-filter-multilingual
    

Multiple detectors union their detections; overlapping spans resolve to the strongest action (block > mask > allow). A configured detector that can't be loaded fails the request closed (HTTP 503) rather than silently skipping the check. Detections are audited at /api/pii/events (hash-prefix only, never the raw value).

The earlier regex pattern tier (pii.patterns, the global pattern catalogue, --pii-config, and the /api/pii/patterns admin endpoints) has been removed, along with response/streaming-side redaction. Those keys now no-op with a startup warning; migrate to pii.detectors + a detector's pii_detection block.

Environment Variables Configuration

Model configurations can specify environment variables passed to the backend process:

name: vllm-model
backend: vllm
parameters:
  model: my-vllm-model

env:
  VLLM_WORKER_MULTIPROC_METHOD: "spawn"
  VLLM_CACHE_DIR: "/tmp/vllm_cache"
  CUDA_VISIBLE_DEVICES: "0,1"

Environment variables are appended to the system environment variables and will override any conflicting system variables with the same name.

Complete Example

Here's a comprehensive example combining many options:

name: my-llm-model
description: A high-performance LLM model
backend: llama-cpp

parameters:
  model: my-model.gguf
  temperature: 0.7
  top_p: 0.9
  top_k: 40
  max_tokens: 2048

context_size: 4096
threads: 8
f16: true
gpu_layers: 35

system_prompt: "You are a helpful AI assistant."

template:
  chat: |
    {{.System}}
    {{range .Messages}}
    {{if eq .Role "user"}}User: {{.Content}}
    {{else if eq .Role "assistant"}}Assistant: {{.Content}}
    {{end}}
    {{end}}
    Assistant:

roles:
  user: "User:"
  assistant: "Assistant:"
  system: "System:"

stopwords:
  - "\n\nUser:"
  - "\n\nHuman:"

prompt_cache_path: "cache/my-model"
prompt_cache_all: true

function:
  grammar:
    parallel_calls: true
    mixed_mode: false

feature_flags:
  experimental_feature: true
  • See [Advanced Usage]({{%relref "advanced/advanced-usage" %}}) for other configuration options
  • See [Prompt Templates]({{%relref "advanced/advanced-usage#prompt-templates" %}}) for template examples
  • See [CLI Reference]({{%relref "reference/cli-reference" %}}) for command-line options

GPU Auto-Fit Mode

Note: By default, LocalAI sets gpu_layers to a very large value (99999999), which effectively disables llama-cpp's auto-fit functionality. This is intentional to work with LocalAI's VRAM-based model unloading mechanism.

To enable llama-cpp's auto-fit mode, set gpu_layers: -1 in your model configuration. However, be aware of the following:

  1. Trade-off: Enabling auto-fit conflicts with LocalAI's built-in VRAM threshold-based unloading. Auto-fit attempts to fit all tensors into GPU memory automatically, while LocalAI's unloading mechanism removes models when VRAM usage exceeds thresholds.

  2. Known Issues: Setting gpu_layers: -1 may trigger tensor_buft_override buffer errors in some configurations, particularly when the model exceeds available GPU memory.

  3. Recommendation:

    • Use the default settings for most use cases (LocalAI manages VRAM automatically)
    • Only enable gpu_layers: -1 if you understand the implications and have tested on your specific hardware
    • Monitor VRAM usage carefully when using auto-fit mode

This is a known limitation being tracked in issue #8562. A future implementation may provide a runtime toggle or custom logic to reconcile auto-fit with threshold-based unloading.