feat(ui): redesign the web UI around a shared kit and a calm palette (#12526)

* build(ui): vendor the shared UI kit snapshot at 0.2.0

The restyle needs the kit's tokens, motion layer and component classes.
Take a pinned snapshot instead of depending on the kit at build time,
and keep a lock file with the version and per-file checksums so a later
update shows exactly what changed. The product theme stays outside the
vendored directory.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add the ink and teal theme and bridge the old variables

Define the product colours as the shared UI kit's roles, for light and
dark, in theme-localai.css. The kit's contrast check passes on every
pair. theme.css keeps the existing --color-* and --shadow-* names but
now points each at a role, so App.css and the pages get the new palette
without edits. Radii move to the kit scale.

index.html now sets data-theme before first paint with the same rule as
ThemeContext (stored choice, otherwise dark), because the contract
layout of the theme file no longer defaults to dark by itself.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): restyle the shared chrome with the UI kit grammar

Adjust the shared classes so every page picks up the same interaction
language without per-page edits:

- Sidebar sits on the canvas and the current row lifts onto a card.
  Section labels are tracked uppercase, the badge is a soft pill, and
  the phone drawer leaves the tab order when closed.
- Buttons are flat: hover swaps the surface, press scales to .97, focus
  is a 2px ring with a 2px offset, danger is a tinted wash.
- Inputs use the card surface and the control edge; switches, tabs,
  filter chips, badges and cards follow the same rules. Cards no longer
  lift on hover; only linked or button cards react.
- Menus and popovers scale in from the trigger corner with 40px items.
  Dialogs get a veil fade and a spring settle. Toasts become pills at
  the bottom centre.
- The page transition is a 250 ms fade with a 6px rise. It fills
  backwards so a finished animation no longer leaves a transform that
  confined dialog veils to the main column.

The focus-ring test now checks the outline instead of a box shadow, and
new specs cover the theme roles, the first-paint theme and the sidebar
lift.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): move leftover hard-coded colours onto the theme roles

The YAML editor restated the old blue palette in JavaScript, and a few
pages kept literal blues, indigo and violet tints, or fallbacks that
only applied because a variable was never defined. Point them at the
theme variables so they follow light and dark and the new palette.

The status badges in the account pages built their tint by appending
"22" to a variable, which is not valid once the variable is defined, so
they had no background. Use the wash roles instead.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): raise the type scale and control size toward the kit

Body and list text moves to 15px and the rest of the scale follows the
kit's 12/13/15/17/21/32/44 steps. Page titles, section headings and
stat values are bold with tighter tracking; titles are 32px.

Buttons, inputs, selects, tabs and nav rows are 40px high with the 12px
radius, compact controls 32px. Tabs become a segmented control. The
sidebar widens to 240px (64px collapsed) and nav rows get more room.
Identifiers and counts in the split views use the mono face, and the
stat grid becomes separate inset tiles.

The Geist stack stays: it is bundled, and the thin look came from the
size, weight and negative tracking, not the face.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): separate cards, panes and floating surfaces from the canvas

Cards, the Models and Installed split panes, the composers and the
confirm dialog use a stronger card edge, the rest shadow and the 20px
radius, so they read as layers in dark as well as light. Menus and
popovers move to a float surface (the hover tone in dark) with the
float shadow.

The selected rail row gets an accent wash and a 3px accent edge. The
send buttons are a clear accent when there is something to send and a
quiet inset when not; the Home button carries data-empty for that, since
submitting an empty box does nothing. The assistant card becomes an
accent wash with a square icon.

New surfaces spec checks the pane edge, the selected row, both send
buttons and the popover in both themes. The voice library empty-state
spec now waits for the layout to settle before comparing two boxes.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): tidy the sidebar header and mark the current row with a dot

The header gives the configured horizontal logo a fixed width and
centres it in a 72px band, lined up with the nav icons. The collapsed
rail shows the configured icon logo centred, and its nav rows become
40px tiles centred in the 64px rail. The current row gets the kit's
accent dot, hidden in the rail.

The theme, language and account controls stay in the sidebar footer:
the app has no global search or command palette to put in a top bar, so
a bar would only hold controls that already have a place.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): centre the avatar in the collapsed sidebar rail

The collapsed avatar link was set to "flex: 0", which gives it a zero
flex basis; with min-width: 0 the link shrank to its padding and the
icon overflowed from the link's left edge, about 14px right of the
icon column. Use "flex: 0 0 auto" in the collapsed and tablet rail.

The footer controls now share the nav icon column in the expanded
sidebar too (6px footer padding, 40px control boxes), and the tablet
rail gets the same footer padding and hidden language code as the
collapsed one.

New spec measures the centre x of the nav icons, mark, avatar,
language, theme and collapse icons in the collapsed, expanded and
tablet states, in both themes, and asserts they agree within 1px.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): size the console and settings rails and stack settings on phones

The Operate console rail, the Settings section rail and the account tab
bar still used 13px text and the old underline tabs. They now use the
15px nav size, 40px rows and the segmented tab control. Form row labels
are 15px with 13px hints.

On a phone the Settings section rail sat beside the form and squeezed
every row into a few characters. Below 720px the rail stacks above the
content as a scrolling row and form rows wrap their control below the
label. The save button no longer carries the icon font class, which
drew a missing glyph before its label. The language menu is wide enough
to keep Bahasa Indonesia on one line.

Assisted-by: Claude Code:claude-sonnet-5-5

* build(ui): update the vendored UI kit snapshot to 0.3.0

Take the 0.3.0 snapshot: the sprite now carries the full outline icon set,
and the new icons/fa-map.json maps Font Awesome names to icon ids. The map
lets the app move off Font Awesome in the following commits. The lock file
is regenerated with the new checksums.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add an Icon component backed by the kit sprite

Icon draws an inline svg that points into the kit's outline sprite. The
sprite is inlined into the page once, so the references resolve under any
base path and in the embedded build without a request. Icons size with the
font (1em), take currentColor, hide from assistive tech unless given a
title, and spin on request. An unknown id draws a neutral circle.

FaIcon and iconFromFa resolve Font Awesome names through the kit's map,
for names that arrive at run time. iconHtml does the same for markup built
as a string. The GitHub and Apple marks are small local glyphs, as the kit
ships no brand marks.

Assisted-by: Claude Code:claude-sonnet-5-5

* refactor(ui): draw shared components and helpers with Icon

Replace the Font Awesome elements in the shared components and in the
utility modules with the Icon component. Lookup tables now hold kit icon
ids instead of class strings. Code-block copy buttons and artifact cards,
which build HTML strings, use iconHtml and a sanitizer-safe slot.

Assisted-by: Claude Code:claude-sonnet-5-5

* refactor(ui): draw model, backend and account pages with Icon

Replace the Font Awesome elements on the home, models, backends, import,
settings, login, account and users pages with the Icon component.

Assisted-by: Claude Code:claude-sonnet-5-5

* refactor(ui): draw chat, studio and recognition pages with Icon

Replace the Font Awesome elements on the chat, media generation, talk and
face and voice pages with the Icon component. The talk status table keeps
its spin and pulse states as Icon props. The connected and error states
now use a dotted circle and an alert circle, so they differ from the idle
ring by shape as well as by colour.

Assisted-by: Claude Code:claude-sonnet-5-5

* refactor(ui): draw agent, node and operate pages with Icon

Replace the Font Awesome elements on the agents, skills, collections,
jobs, fine-tune, quantize, nodes, swarm, usage, traces and activity
pages with the Icon component. Two class strings on layout elements held
leftover button and icon classes from an earlier merge; they are cleaned
up so the elements keep only their own classes.

Assisted-by: Claude Code:claude-sonnet-5-5

* style(ui): size and align icons for the svg component

Icon rules that targeted the font element now target the svg: the
descendant "i" selectors in App.css and auth.css become ".lai-icon". The
svg is 1.2em with a 2 unit line so it matches the visual size of the old
glyphs at the 12 to 16px sizes the app uses, sits on the text baseline,
and follows the context font size. Large empty-state marks get a lighter
line. Menu icons get a 16px box and the readiness badge icons keep their
20px circle with padding. Add the pulse used by the talk status.

Assisted-by: Claude Code:claude-sonnet-5-5

* build(ui): remove Font Awesome

No source file references the icon font any more. Drop the package and its
stylesheet import. The build no longer ships the solid, regular and brand
font files.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): keep focus traps off the svg use references

The dialog and drawer focus traps collect focusable elements with a
"[href]" selector. An icon's use element carries an href, so it became the
first "focusable" element and Tab at the end of the dialog stopped there
instead of wrapping to the first button. Match "a[href]" instead.

Assisted-by: Claude Code:claude-sonnet-5-5

* style(ui): keep icon sizes overridable and set the line width per svg

Give the icon base rule zero specificity so a rule that sizes one icon
(nav column, menu box, avatar, language switcher) wins whatever its order
in the file. The sprite symbols fix their own line width; the inlined copy
drops it so the width set on each svg applies, as the --lai-stroke custom
property, and large marks can use a lighter line. Pin the avatar and the
language globe to the boxes the sidebar alignment spec expects. Import the
map as JSON with an import attribute so Node can load it in the spec.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): select icons by the svg markup and cover the sprite

Specs that found icons by their Font Awesome class now select the svg by
its data-icon. The dead-icon audit checks that every svg resolves to a
sprite symbol and has a size. The class hygiene spec fails on any
remaining Font Awesome class. A new spec checks every mapped icon id has
a symbol, that the sprite is inlined once, that an icon paints at the root
and under a forwarded path prefix, and that Font Awesome names map as
documented.

Assisted-by: Claude Code:claude-sonnet-5-5

* build(ui): update the vendored UI kit snapshot to 0.4.0

Take the 0.4.0 snapshot: hub tabs with count and attention badges, the six
chart series tokens and the grid colour in the theme contract, and sample
themes on a calmer palette. The kit headers are renamed and the lock file
is regenerated with the new checksums, as for the earlier snapshots.

Assisted-by: Claude Code:claude-sonnet-5-5

* style(ui): switch the theme to the calm palette

Rewrite the LocalAI theme on the calm palette: a muted teal accent on a
near-neutral green-grey canvas, desaturated status colours, no glow and no
coloured shadows. The theme fills every role of the shared UI kit's 0.4.0
theme contract for light and dark, including the six chart series and the
grid line. The bridge in theme.css keeps the old --color-* names working,
adds the dark surface ladder (card, raised, float) and a strong edge, and
points the fixed data hues at the chart series.

Two values differ from the first sketch. The dark text on the accent fill is
#021512 instead of #04201d: it reads 5.58:1 on the fill at rest and 6.4:1 on
the hover fill, against 5.08:1 at rest for the lighter value. The light
control edge is #6b7d7a. The kit's contrast script passes for all text pairs
(4.5:1), control and focus pairs (3:1) and series colours (3:1).

Leftovers that no longer fit the palette are fixed: the usage chart takes
the six series colours in order, the audio and animation canvases fall back
to the new accent, the face box loses its glow, and two gradient fills are
now flat. The theme tests expect the new canvas colours.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): replace the console rail with hub tab bars

Build and Operate no longer open a second navigation rail beside the page.
Each is a hub: one row of the kit's hub tabs above the page, with count and
attention badges that scroll sideways on a phone. Every URL and route stays
as it was, plus a new /app/build landing page that lists the Build tools
with a line each.

Build tabs: Overview, Agents, Skills, Memory, Jobs, Fine-Tune, Quantize,
Import, Voices (recognition and library) and Faces. Operate tabs: Status,
This machine, Swarm (distributed mode only), Runtime (backends, activity,
failover), Traffic (usage, traces, middleware) and Settings (settings,
users), plus the API link. A tab that holds several pages shows a second row
of links, and a sub-page such as a node detail keeps its tab highlighted.
The feature and admin gates decide which tabs are drawn, and badges show only
values the Operate summary already has.

The sidebar lists Build and Operate under a Workspace label next to the
Create group. The voice library moves under Build and the model import page
gains the Build tab bar. The old rail styles, the rail signals and the
console config are removed, and the Operate overview docs describe the tab
bar. The specs that drove the rail now drive the tabs, and a new spec covers
the tab for each route, gating, badges and the phone layout.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild Home as a calm console

Home now opens on one command bar: the model chip shows which models are
warm, the MCP chip and attach buttons sit beside it, and Send is a solid
button with an Enter glyph. Typing "/" opens a grouped, keyboard-driven
action list built on the kit command list; every action has a destination
in the product.

Memory use folds into a one-line strip that opens into the loaded models,
with Stop per model and Stop all. It opens by itself while a model is being
staged and after a failure, and shows nodes and aggregate memory in a
cluster. The list of resident models carries no per-model size because the
API reports none.

"Jump back in" lists the conversations stored in the browser, one card per
day, with j and k to move, Enter to resume and delete with an undo toast.
First run keeps the install steps and the recommended models. The assistant
prompt is a dismissible line, the library links are one quiet row and the
API section is collapsed. Chat accepts an empty new-chat hand-off for /new.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Home console

Update the Home specs for the new structure and add specs for the slash
menu, the model chip, the memory strip (expand, stop, staging, failure,
cluster), the resume list (grouping, j/k, Enter, delete with undo), first
run, the send hand-off, a non-admin user and the phone layout.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add the fit, disk and cleanup helpers for the models page

Pure functions and hooks that the rebuilt Models page reads, with node
tests for the rules.

modelLedger turns an estimate and the memory budget into one of three
verdicts (fits, spills to CPU, over) with the headroom in bytes, and
reads the models disk from the resources reading. The disk counts as low
under 10 percent or under 20 GB free, and is absent when the server
reports none or runs as a cluster controller.

cleanupPlan ranks installed models from what the API reports: loaded,
pinned, or named by an agent, a task, a failover chain or an alias keeps
a model protected; another installed build of the same gallery model is a
duplicate; disabled models rank above idle ones. The API records no last
use or use count, so none is used. When a lookup fails, nothing is called
safe.

useModelRemoval holds a removal in the browser for an undo window and
sends the existing delete call only when the window ends. Leaving the page
drops the batch without deleting anything. The undo toast takes optional
labels so other pages can reuse it.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild Models as a ledger with a disk strip and cleanup review

Explore is one dense table. Each row carries the size, a solid memory bar
and the headroom in words ("3.7 free", "+1.5 on CPU", "0.9 over"), worked
out from the estimate at the chosen context length. Capability chips show
the server's count for each facet, search keeps its meaning and "/" jumps
to it, and a density switch (also "d") picks comfortable or compact rows.
Selection is a surface step and a check, never a rail. Arrow keys move,
Enter installs and Esc closes the inspector, which keeps the fit summary,
VRAM by context chart, variants, files, links, tags and licence. A failed
install shows its error in the row with a Retry that dismisses the old
failure first. A failed or empty listing says which it is, and a host with
no GPU is measured against memory and says so.

Installed uses the same table with state filters that carry counts, a
state per row, Load or Stop on the row, the row menu and the sort by size.
Sizes come from the files the gallery lists, so a model it does not know
shows a dash.

A strip in the header shows the free space on the models disk. It turns
amber under 10 percent or under 20 GB free, hides when the server reports
no disk or runs as a cluster controller, and opens the cleanup review.
Explore says how much an install leaves free.

The review ranks installed models as Safe to remove, Probably safe and
Your call from real facts only, lists protected models with the reason,
and says plainly that usage history is not recorded. A sticky bar shows
what a choice frees. Confirming runs a dry run that checks again and lists
what will go. Removal waits 30 seconds with an undo; nothing is deleted
before that, and leaving the page deletes nothing.

The old rail, filter band and popover styles are removed.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Models ledger, Installed table and cleanup review

Update the Models, lifecycle, cluster fit, height, search focus and
surfaces specs for the table and inspector, keeping what each one checks.

New specs, on a shared 41-model gallery stub with three machine profiles:
the fit bar and headroom words for a 24 GB card, an 8 GB laptop and a host
with no GPU; facet counts, search, "/" and Escape; selection, arrow keys,
Enter to install, density; the disk strip when normal, low and hidden; and
the states (loading, empty, offline, install failed, phone). Installed
covers filters with counts, row actions, the row menu, sizes and sort.
The cleanup specs cover grouping, protected models, the honest-data note,
the effect bar, the dry run, the undo window, a failed delete, leaving the
page, and the phone sheet.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add the placement helpers and the estimate hooks

The Placement section and the model page need the same few rules, so they
sit in plain functions that can be read and tested alone.

placement.js holds what gpu_layers, tensor_split and main_gpu mean (unset
asks for every layer and the llama.cpp engine trims it, zero is CPU only,
99999999 is the value LocalAI itself writes for all layers), the device
list taken from the resources reading, the split by free memory, the part
of an estimate that grows with context (read from two lengths, since that
term is linear), the fit states with their limit (95 percent of free
memory, and the leftover has to fit in system memory too), and a bisection
for the largest layer count whose estimate fits. The estimate returns one
total and no layer count, so the search runs over 1 to 256 and stops at
the first count that no longer changes it.

modelWalk.js keeps the order of the list a model page was opened from, in
memory and in session storage, for the previous and next buttons.

usePlacementEstimate reads /api/models/vram-estimate for a choice, again at
twice the context, and with every layer, and keeps readings for the
session. useModelPage reads a gallery entry by name, an estimate by
context size (from the model's own files when the gallery does not list
it), the builds and the loaded models. usePlacementConfig edits the four
placement keys of an installed model and saves only what changed.
useModelActions is the Load, Stop, disable, pin and remove logic of the
Installed table, shared with the model page. MemoryBar is one solid bar
with a tick at the capacity of its pool; over capacity it grows past the
tick and the tick turns red.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add the Placement section to the model editor

Run this model on: CPU only (gpu_layers: 0), Auto (the key stays unset) or
Custom. Custom takes a number, has an All layers button that writes
99999999, and shows a slider only when the estimate reports the model's
layer count, which it does not today. Context size has presets and a
number field because the KV cache follows it. With two or more GPUs there
is a split (written as percentages, with a button that takes them from the
free memory of each card) and a main GPU.

A bar per GPU and one for system memory show what other programs use, the
model's weights and working memory, and the part that grows with context,
with the room left or how far over it is. Under them a verdict in plain
words: Fits in GPU, Spills to CPU, Too many layers for the GPU, Runs on CPU
only, No GPU found, Not enough memory. It says "slower" and never a
multiplier, because the estimate has none. Fit it for me asks the estimate
for the largest layer count that fits the free GPU memory and says what it
set, with Undo; it is hidden when the estimate is unavailable or the host
has no GPU. Loading shows skeletons, an unavailable estimate shows a note
with Retry, and a server that schedules onto other machines shows no bars,
because its device list is the controller's.

The editor shows the section for an installed model, with a link in its
section rail. Auto sends null for the key, since a patch only merges, and
a null read back opens as Auto. The docs describe the section and what each
mode writes.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): open a model on its own page

A model has an address, /app/models/<name>, for an installed model and a
gallery entry alike. Open it from the arrow at the end of a row, a double
click, "o" on the selected row, the inspector's Open details button, or a
tap on a phone. The title block holds the main action: Install with a
chevron that chooses the build, or Load and Stop with a menu (disable, pin,
edit configuration, logs, delete with a confirm). A strip answers whether
it fits, what it does and what installing leaves free.

Tabs: Overview (about, a memory bar, state, the pages it opens in, and the
agents, tasks, chains and aliases that name it); Fit and memory (verdict,
context sizes, the bar split into weights and context, and memory by
context against the limit, with a data table); Variants and files (builds
with size and fit, install any, the files of the chosen build). For an
installed model also Usage and history, which says what the API does not
record instead of drawing an empty chart, Configuration, which is the
Placement section with the file it writes and a link to the full editor,
and Logs, the backend log viewer without its page. Keys 1 to 6 switch
tabs, [ ] and j k walk the list the page was opened from, Esc or Backspace
go back.

The list stays mounted behind the page, so Back finds its view, search,
filters, selection and scroll as they were, and focus returns to the row's
arrow. The page covers loading, an unknown name with the closest matches,
the gallery being out of reach, an install in progress with Cancel, and a
failed install with Retry.

The docs describe the page and its keys.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the model page and the Placement section

New specs for the model page: reaching it from Explore, Installed, a
double click, "o", a pasted link and a phone tap; the walker and Back with
the search, a filter, the selection, the Installed view and the scroll
kept, and no second read of the gallery; the title block, the answer strip,
tabs by click, keys and arrows; Fit and memory, builds and files with the
install call each one makes; an installed model's actions, used-by,
the honest usage tab, configuration and logs; loading, an unknown name,
offline, an install in flight and a failed one; and the phone.

New specs for Placement: every mode and the keys it writes, the slider
only when a layer count exists, the context presets, the bars and every
verdict, two GPUs, no GPU, a cluster, an unread machine, a loading and an
unavailable estimate, Fit it for me and Undo, and the section in the model
editor with its save.

The phone tap on a row now opens the page, so the two phone specs that
expected the inspector as the page check the page and keep the inspector
check for a window between a phone and a desk.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): link Studio results and open workspaces from a prompt

Each workspace now records the result it was made from (parentId and an
edge kind such as take, animate or to-3d) and reads a prompt, model, size,
count and source from the query string, so one page can hand work to
another. A source result is fetched from the server's own output file and
becomes the start image, the picture for 3D, or the audio file. A note on
the page says when the source loaded or could not be loaded.

Diarization had no history; it now keeps the file name, the model and a
speaker count, never the recording. Prompts are cut at 2000 characters
when stored. The pure helpers (type suggestion, grouping, lineage layout,
favourites, clearing) have node tests.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild the Studio front page as a composer with your work

The front page is a prompt box with a chip per type, a type suggestion
from the words, starters, and the options each workspace accepts. Generate
opens the workspace with those filled in. A type with no model is a dashed
chip that shows a gallery model, its size, memory need and an Install
button only when picked; the typed words stay while it installs.

Under it, Your work lists results from every workspace as a masonry with
filters, counts, favourites and a Clear history action. Results made from
each other stack into a project tile and open as a lineage board with a
dock for running a new take or branching to the next step; steps the
destination cannot start from yet are disabled with the reason.

The docs describe the page, what is stored in the browser, and the query
parameters a workspace accepts.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Studio composer, your work and the lineage view

Specs for the type suggestion, the keys, hand-off to each workspace, the
install path for a missing model, the masonry filters, favourites and
clearing, stacking, the lineage board, new take and branch, steps that
are disabled with a reason, and the phone layout. Existing Studio specs
move from lanes to chips with the same intent.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add the shared Studio workspace frame and move Images onto it

The seven Studio workspaces get one layout: a row of type tabs, a compose
card (optional sources as chips, a prompt with starters, a model chip,
the essential options as chips, an Advanced fold that names what is
inside, the memory the model needs, and one action with the reason when
it cannot run), a run area, and a strip of recent results of the type.

The run area shows a job card with the time that has passed and an
indeterminate bar, because these endpoints report no phase or percentage;
a failure with what the server said and one action; or the result with a
toolbar: Favourite (the list the front page keeps), Download, Use in (the
hand-off targets, disabled with the reason when a destination cannot
start from the result), Re-run with edits (the take's values go back in
the form, changed fields are outlined and listed) and Lineage. A type
with no model shows the install note from the front page.

Images is the first workspace on the frame. It keeps its size, count,
steps, seed, negative prompt, source image and reference images, and its
history writes, including the parent link and edge of a hand-off run.
useMediaHistory.addEntry now returns the id of the entry it stored. The
docs describe the workspace page.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): move Video onto the workspace frame

Video keeps its size list, duration, frame rate, steps, seed, CFG scale,
frame count, negative prompt, start and end image and avatar audio.
The start and end image are source chips, the avatar audio opens the
recording and paste input from a chip, and the rest sit in the Advanced
fold. A start image from a hand-off shows as a chip with its picture.
Results play in the video player with the shared toolbar.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): move TTS onto the workspace frame

TTS keeps the saved-voice picker for cloning models, the typed voice for
the others, the voice library deep link, and the delivery instructions,
which now sit in the Advanced fold. The result is the waveform player
with the words under it. The stored entry also keeps the voice id so
Re-run with edits can select the same saved voice.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): move Sound onto the workspace frame

Sound keeps its Simple and Advanced modes and every field of both: the
description, instrumental, vocal language, caption, lyrics, BPM,
duration, key, language, time signature and think mode. The mode switch,
instrumental and duration are in the compose card, the rest in a More
options fold. The stored entry keeps all of the fields, so Re-run with
edits restores the form as it was.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): move Transform onto the workspace frame

Transform keeps its audio and reference inputs with upload and record,
the echo test, the key=value parameters (now in the Advanced fold), the
input and output spectra and the three waveform players. The audio that
was chosen shows before the run, waiting to be transformed. Re-run with
edits puts back the model and parameters and fetches the audio and
reference the server kept for that run.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): move 3D onto the workspace frame

3D keeps the picture input with paste and webcam, the animation
operations a model declares, quality and background, the shape and
material steps, guidance and seed, the GLB and animation viewers, the
remesh control and the download. A 3D result now has a title from the
motion prompt when it has no label, so the strip and the front page name
animation results by what was asked. Re-run with edits is shown disabled
with the reason, because only a small thumbnail of the picture is kept.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): move Diarization onto the workspace frame

Diarization keeps its model and recording inputs, the option to prepare
speakers to remember, the clean-speech previews, naming and remembering a
speaker, and the history entry with only the file name, model and
counts. The result now shows a timeline with one lane per speaker, the
talk time of each speaker, and the segments with their start time and
text. RTTM, SRT (only when the run has text) and JSON are built in the
browser from the result. The helpers for talk time, axis ticks and the
two text formats have node tests.

Assisted-by: Claude Code:claude-sonnet-5-5

* refactor(ui): remove the styles and lists the old workspace layout used

Nothing renders the two-column workbench, the control column, the old
history lists, the generation progress tiles, the TTS voice picker or the
result echo any more. The inline-style baseline drops with them.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the workspace frame and one run per type

Specs for the type tabs, the compose card and its reason when Generate
cannot run, starters, the Advanced fold, the job card with no invented
progress, a failed run and its one action, the install note, the strip
with its favourites filter, Use in with its disabled steps, Lineage, the
parent link, Re-run with edits and its list of changes, deleting and
clearing, and the hand-off note. One run through each of Video, TTS,
Sound, Transform, 3D and Diarization, the phone layout of all seven, and
reduced motion. Existing Studio specs move from the old control column to
the compose card with the same intent.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild the chat thread with raised user turns, prose replies and one-line activity

Your messages are raised blocks on the right at a 760 px measure and the
model's replies are plain prose under its name and a warm or not loaded
dot. Reasoning, tool calls and their results fold into one quiet line
that opens inline into steps. Code blocks carry a Copy button and a
Canvas button that opens that block in the canvas, image attachments are
thumbnails that open in the lightbox, and files are chips. Per-message
actions show on hover, on focus and on the last turn, and a turn takes
focus so the arrow keys and C, E, R and B work. A failed reply keeps the
text written so far and shows the reason with one Retry action.

The Agent chat page keeps the older rules: the new styles are scoped to
the chat page and use their own class names.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): use the Home command bar as the Chat composer

Chat now ends in the same object as Home: the model chip, the MCP chip,
a Canvas chip, the message box with attach buttons, a solid Send and the
hint line, with the slash menu on the kit command list. The slash menu
lists what Chat can do today (switch model, new chat, conversations,
manage mode, canvas, find, settings, export, clear). While a reply is
streaming Send becomes Stop, which Esc also presses, and Up in an empty
box edits your last message. Attached images show as thumbnails and a
line under the bar carries the speed and the token count.

HomeComposer takes optional props for this (extra chips, its own slash
list, Stop, paste, a stricter Enter); Home passes none of them.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): open conversations from a Ctrl K menu with day groups, undo and a slim header

The conversations list opens as a centred menu on Ctrl or Cmd K. It
groups chats by day like the Home resume list, shows the model that
answered and the time, searches names and message text, and moves with
the arrow keys. Enter opens a chat, F2 renames it and Delete removes it.
Removing a chat hides the row and shows the kit undo toast; the chat is
deleted for good only when the undo time ends. Rename, duplicate, copy
and export are on each row, as before.

The header is one slim bar: the Chats button, the chat name (click to
rename), a context meter when the context size is known, settings and a
More menu with rename, duplicate, copy, export, model info, keyboard
shortcuts and clear. A dialog lists the shortcuts the page answers to.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): show loaded state, capabilities and fit in the Chat model switcher

The model chip in Chat opens the same list as Home, grouped as Loaded
now and Installed. Each row says warm or not loaded and marks models that
understand images. When the list opens, the page reads the host memory
once and asks the server to estimate each listed model at the chat's
context size (up to twelve, three at a time), then shows what the model
needs and whether it fits: free memory, how much would run on the CPU, or
how far over the machine it is. A model with no estimate shows no fit
text, and no load time is shown because the API does not report one. A
memory bar closes the list.

The picker takes the model list from the page when it has one, and
useModels can skip its own request.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): move chat settings into a sheet and add find in chat, jump to latest and a wider canvas

Settings open as a kit sheet: the system prompt, temperature, top P and
top K (each says "model default" until it is changed and has a Reset),
the context size with quick sizes and a note that it only drives the
meter, Manage mode and Focus mode, the model info for admins with its
Edit config button, and Clear conversation behind a confirmation. The old
slide-out drawer and the model info panel are gone.

Ctrl or Cmd Shift F (or the search button, or /find) opens a search bar
over the thread. It marks matches in the messages already on the page,
shows "n of m" and steps with Enter and Shift+Enter. Nothing is sent to
the server. Jump to latest is a pill above the composer. Esc stops a
reply, then closes the search, then closes the canvas.

The canvas panel gets the kit look: tabs, a Code and Preview switch, Copy
and Download, a full-page layout on narrow windows, and translated
labels. The Agent chat page shares it and gets the same look.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add the empty, no-model, loading and phone states to Chat

An empty chat opens with the composer under one line, starters to try,
whether the model is loaded, and the Jump back in list: the same rows as
Home, read from the chats the page holds. With no chat model installed,
an install card offers the starter models for this hardware, the gallery
and import, and the composer stays so the text is not lost.

While a reply waits for a model, a load card shows what the page knows:
the phase the server names, the node, the bytes and the time left when the
server reports them, and a progress bar. A model that is just not loaded
yet gets a plain note, with no invented phases or estimates. The foot
warns when the context is nearly full.

On a phone the header drops its labels, the model list and the settings
open as sheets from the bottom, per-message actions stay in view and the
canvas takes the whole page.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild Talk as a calm voice page over what the connection really does

Talk is one stage and one transcript. The stage has the pipeline chip,
the voice and language chips, an outline orb that follows the real
microphone and playback levels, a heading and a sentence for the current
state, and the controls. The transcript lists You, Reply, Tool and Result
lines and can be copied. Session settings (instructions, voice, language,
tools, Manage mode and the pipeline's parts) open in a sheet.

The states are the ones the code reaches: no pipeline model, idle,
connecting, listening, thinking (also while a tool runs), speaking, an
interrupted reply (the server cancelled it; a note marks the cut), a
blocked microphone, a link that failed during a session, and any other
error with its reason and a link to the traces. Push to talk and
hands-free are not on the page, so they are not shown. Diagnostics keep
their waveform, spectrum and stats, drawn in theme colours.

The page text moves into the talk namespace, and the old Talk and
visualizer styles and the inline-style count go down with the rebuild.

Assisted-by: Claude Code:claude-sonnet-5-5

* refactor(ui): remove the chat styles and strings the rebuilt page replaced

The settings drawer, the model info panel, the bubble avatars, the
conversation menu popover, the context bar, the recent strip, the
staging bar, the file badges and the focus-mode rules have no user now.
Their rules, the Chat page's focus class and seven unused empty-state
strings are removed. The Agent chat page keeps the shared message,
sidebar and input rules it still renders with.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the rebuilt Chat and Talk pages

Add a Chat page under features (thread, message actions and keys, the
message box and its slash actions, the model list with loaded state and
fit, conversations on Ctrl K, settings, find, canvas and the empty,
no-model and loading states) and a Talk section to the realtime API page
with the states the page shows. Manage mode now turns on from the chat
settings or /assistant, and the client MCP steps point at the MCP chip.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): settle the rough edges of the new Chat and Talk pages

The undo toast sat under the conversations menu, so the Undo button could
not be pressed while the menu was open; the menu, the sheets and the
fullscreen canvas now stay below the toast layer. Esc in a rename box
saved the text through the blur that follows it; it now cancels. The
image viewer closed on Esc only when the page did not re-render on the
same key, so its key listener is registered once and reads the latest
handlers. Keys on a focused message no longer type their letter into the
editor they open, "/" from outside a text field starts a command as it
does on Home, and Esc leaves the page's own dialogs alone.

Code in the canvas is highlighted for languages that have no preview.
The conversations menu drops its key hints on a phone so Clear all
stays in view. Talk hides Test tone while connecting and calls a server
error "Something went wrong", since the call can still be open.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the rebuilt Chat and Talk pages

Specs for the thread layout and the activity fold, code blocks, image
thumbnails and the viewer, per-message actions and their keys, a failed
reply with its one Retry, Stop and Esc while streaming, the composer and
every slash action, the conversations menu (groups, search, resume,
rename, delete with undo that ends by itself, one chat left), the model
switcher with loaded state, vision and fit text from stubbed estimates,
the settings sheet, the canvas panel, find in chat, Jump to latest, the
empty, no-model and loading states, the phone layout and reduced motion.
Talk is driven over a fake WebRTC link through idle, connecting,
listening, thinking, speaking, interrupted, blocked, lost, error and no
pipeline, its settings sheet and its phone layout. Node tests cover the
message text helpers and the conversation grouping.

The existing chat specs move to the new structure with the same intent:
the transcript spec now describes the raised turn and the prose reply, and
the render smoke accepts Talk's own header.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): keep a bounded run log for agents in the browser

The server keeps no run history for an agent, so a run is one task and
the events until the agent answers, written to browser storage while the
page watches the stream: up to 50 runs per agent, task, step and answer
text only. Stored chats from the earlier agent chat page read as runs
with stable ids. A run still marked running five minutes after its last
event reads as stopped. Helpers read an agent's config into chips, build
the list of changed fields against the saved config, hide secret values
and offer starting points.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild the Agents area around runs

The Agents page shows what needs a look (work in flight, a run that
failed in the last day), then each agent with its model, attached memory
and skills, and a strip of its last 14 runs. An agent has its own page: model,
tools, memory, skills, instructions, a task box and its runs. A run has
an address, shows the thread while it works (steps folded into one line,
the tool in use, the answer as it arrives) and settles into a report
about a second and a half after the agent answers: task, outcome,
follow-ups, evidence and steps, with wide tables opening wider on demand.
A failure says in plain words what happened and offers Run again.

Create and edit fold into sections with a ready mark and a one-line
summary, start from a template or an optional model-written draft, and
open a preview sheet with the config as saved and the changes against
the saved agent. Status becomes a quiet panel in the same language, and
the old chat link opens the agent page.

There is no Stop, approval, steer, version or dry-run control, because
the agent API has no call behind them.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Agents launcher, agent page, runs and editor

Specs for the Now strip and run strip, search and the empty state, the
agent page, starting a run, the live thread, settling into the report,
the run address across a reload and for a run from another browser,
follow-ups with their history, failures, the folding editor with ready
marks, templates, the preview sheet with hidden secrets and changes, the
status page, and the phone, 1440 and 2560 layouts.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe runs and the new agent create flow

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add helpers for tasks, schedules and job outcomes

Reads a cron expression the way the server does (five fields or an @
shortcut), checks it, and puts the common shapes in words. The next run
is left out on purpose, because the schedule follows the server clock,
which the browser cannot read. Also groups jobs by day, sums the last
seven days, and gives each job one outcome line from its result or
error. A rerun call starts a new job with the same parameters and media.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild the Jobs area around runs

The Jobs page opens with one sentence about the last seven days, then
the tasks (model, schedule in words, last 14 jobs, enabled switch, Run
now) and a run history grouped by day. Each row has an outcome sentence
and opens to the error or the start of the result with one next action.
Deleting a task waits 30 seconds with an undo button.

A task opens as a page with its recent runs, its prompt with the gaps
marked and its schedule. The task form folds into sections, takes a
schedule as a preset or a checked cron expression, warns about prompt
gaps the schedule does not fill, and has a preview sheet. A job opens
as a document: task, outcome, delivery and the recorded steps; a failed
job says what happened and offers Run again.

Run now now sends attached media through the job call, which is the
only one that takes it. "Clear History" only ever cancelled running
jobs, so it is now called Stop running jobs. Webhook headers of a saved
task show as JSON instead of [object Object].

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Jobs page, task pages and job pages

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the Jobs page and the task form

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add helpers that say who uses a skill or a collection

An agent loads a skill when skills are on and the skill is in its selection (an empty selection means every skill). It reads the one collection that carries its own name, when its knowledge base is on. The helpers derive that from the saved agent configs, build the config that adds or removes a skill or a collection, and estimate tokens as characters divided by four. Removing the last selected skill switches skills off, because an empty selection would mean every skill.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild Skills and Memory as one library

Skills and collections sit in a list with an open item beside it. Each row says who uses it, read from the saved agent configs, or says it is not used yet. Chat reads neither, so it is never named. An item opens in a pane with a Used by strip (names link to the agent, a small x removes it, with undo) and an Add to menu that shows what the addition costs. A collection can be added only to the agent that carries its name.

The Memory pane searches the collection alone and shows ranked passages with scores, lists web sources with their refresh interval and the files, shows the server message when an upload fails, and names the endpoints and where files stay. The Simulate a message sheet runs a collection search and shows an agent's skills with a token estimate. It runs no model. The collection details route now opens the same page.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): show use and cost in the agent form pickers

Each skill in the agent form says which other agents use it and what it adds to every message, with a total for the selection. The memory section names the collection the agent reads.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Skills and Memory libraries

Specs for the used-by lines (including an agent that uses every skill), the filters, search, add to agent, remove with undo, the last-skill case, an unreadable agent list, the empty states, git repositories, the Memory question box, sources, uploads that fail, the Simulate sheet with the parts the API can run, the agent form hints and the phone layout.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the Skills and Memory libraries

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add helpers for the Operate status page and backend rows

Pure functions for the parts that need rules. They work out the memory
pool the page measures (GPU memory, system memory, or the workers of a
cluster that are answering), which pools are too full, the headline and
the four ledger rows, the geometry of the capacity chart, and what
removing a backend would leave without a runtime (models name their
backend, and a meta backend names the concrete one it points at). A
second set says what a backend row states: installing, queued, removing,
failed, update available, current or absent.

LocalAI keeps no memory history, so the chart reads a bounded buffer of
readings the page took itself and says so. A reading with no total is
dropped rather than drawn as zero.

Two hooks are shared by the pages that need them. One retries a failed
operation after moving the failure into the record. The other holds a
cancel for an undo window, because the server cannot take a cancel back.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild Operate Status, This machine, Backends, Activity and Logs

Status opens with one sentence ("2 things need you", or "Everything is
running") and four rows: Needs you, Capacity, Running now and Recent
failures. A row with a problem opens by itself and holds the button that
deals with it: Update a backend, Retry or Dismiss a failed operation,
Unload a model. A quiet row stays one line. A new installation gets a
first-run screen, a cluster sums the memory of the workers that are
answering, and a page still waiting for an answer says so. The chart
under the rows is drawn from readings the page took while it was open
and is labelled that way, because LocalAI keeps no memory history.

This machine leads with GPU memory as one bar, then host memory split by
running model, then VRAM, RAM, CPU and disk with a bar each. The running
models become a kit table with the same menu and stop dialog.

Backends is one list with Installed and Catalog views. A row says what
the backend is doing (a progress bar with Cancel, Queued, Failed with
Retry, Update 1.2.0, Current), carries the one button that matters, and
opens in place. Removing a backend names the models and the meta
backends that would stop working. Check for updates, Update all, From
URL and a first-run recommendation for llama-cpp are in the header.

Activity keeps its three sections as quiet rows. Cancel waits eight
seconds with an undo toast, because the server cannot take a cancel
back; a cancelled install can be started again from the record. Logs
gets a process list, a picker, stream and text filters, Follow and
Times switches, and a Clear with an undo window.

Not shown, because the API has no data for them: GPU temperature and
power, a size per backend, an earlier version to roll back to, a
dependency lookup beyond the models and meta backends that name a
backend, and models that failed to load.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover Operate Status, This machine, Backends, Activity and Logs

New specs for the Status headline and ledger (healthy, needs attention,
one thing, a full memory pool alone, loading, first run, cluster), its
actions (Update, Retry, Dismiss, Unload with its dialog), the capacity
chart built from readings taken while the page is open and bounded, the
phone layout, no coloured edge on a row, and reduced motion.

The Backends specs cover the two views, install progress with Cancel and
its undo window, Retry on a failed install, Update, Update all, Check for
updates, removal with the models and meta backends it would break,
Install from URL, the first-run recommendation, a cluster, and a phone.
Activity gains cancel with undo, Cancel now, a second cancel, leaving the
page, progress, and starting a cancelled install again. Logs covers the
stream and text filters, Follow, Times, Export, Clear with undo, the
process picker and list. This machine covers the GPU strip, several GPUs,
no GPU and Add a machine.

Existing specs keep their intent and follow the new structure: rows open
in place instead of in a pane, Update replaces Upgrade, the notice spec
now pins that an update is a row state and not a banner or a rail, and a
cancel waits for its undo window.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the rebuilt Operate Status, Backends and Activity pages

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add helpers for the Swarm pages

Pure functions for what the pages work out from the cluster API: a node's
state in words, which nodes a placement rule may use, what a rule would
ask for, what a drain or a lost node would leave without service, the
nodes a bulk backend update reaches, and the join commands for a worker,
a peer instance and a memory shard. Hooks read the roster, the loaded
replicas and the rules.

Everything runs in the browser from data the page already holds, and
says when it cannot see free memory or disk.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild the Swarm hub: nodes, node page, placement rules, failover

Nodes is a sortable table with comfortable and compact rows, a Needs
attention filter by reason, a map of the cluster that is not drawn on a
phone, the running models, and a bulk backend update for the nodes that
drifted. A node is a page: state, vitals, a drain preview computed from
the loaded replicas and the rules, tabs for models, backends, logs and
capacity and labels, and Remove that asks for the node's name.

Placement rules are written as sentences, show where each model is
loaded now, and edit in a side sheet with a preview of the nodes a draft
could use. Deleting a rule waits a few seconds so it can be taken back.
Failover keeps its chains, adds what the router does when a worker stops
answering and a per-node preview of what would stop. Add a node covers a
registered worker, a peer instance and a memory shard, with a command to
copy and a live line that says when the machine arrived. P2P keeps its
page in the same vocabulary, and the node logs page follows the local
logs page.

Previews are labelled as worked out in the browser. Per-GPU readings and
node events are not drawn because the API does not return them. Failover
moves to Swarm when distributed mode is on. Legacy fleet components and
their styles are removed.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Swarm hub

New specs for adding a node (each join method, the command, copy,
waiting and found, approve, a single install, P2P, a phone), placement
rules (sentences, where models are loaded, the preview matrix, the sheet
and its preview, delete with undo) and failover on a cluster. Node
detail covers its tabs, the drain preview and its dialog, resume, remove
with the typed name, a node that stopped answering, and unload.

The nodes specs follow the new structure and keep their intent: the
table, filters, grouping, pagination, bulk actions, the map, and running
models with stop, logs and the loading, error and empty states. The
scheduling, failover, P2P, hub and smoke specs follow the renames.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the rebuilt Swarm pages

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add helpers for the Traffic pages

Pure functions for what the pages work out from the usage ledger, the
trace summary, the trace buffers and the resources reading: the shared
time window, grouping, sorting and filtering of usage rows, chart series
and axes that start at zero, the overview figures, per-model statistics,
the state of a trace and the words for a failure, the backend operations
that ran during a request, CSV export, the Prometheus metric list and
scrape config, and a bounded buffer of host readings.

A figure whose source cannot say is null, never zero. The trace summary
call takes the window in hours, and a helper reads /metrics with its
status.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild the Traffic hub: overview, usage, models, host, traces, middleware

Traffic opens on an overview: five figures (requests, failed, p95, tokens
in and out) and three charts, each naming its source. A second row of
links reaches Usage, Models, GPU and host, Traces, Middleware and
Prometheus, and one time window is shared by the first three.

Usage groups by model, user or API key, filters, sorts, opens a row on its
own chart, exports the rows it holds as CSV or JSON in the browser, and
keeps the opt-in cost estimate and the quota forecast. A user who is not
an admin sees only their own numbers. Models joins the ledger, the
backend-operation buffer and the loaded models. GPU and host shows the
current reading and two charts of readings taken since the page opened.

Traces gets filters, a settings strip and an explained off state. An API
request is a page: the error LocalAI recorded, a timeline with the backend
operations that ran meanwhile, and bodies that stay closed until revealed.
Middleware draws the pipeline as five steps and shows the rules of the
selected step. Prometheus documents /metrics, checks it against the
server and gives a scrape config to copy.

Alerts is not built: LocalAI has no alert rules. Per-model latency
percentiles, GPU utilisation and compare with the previous period are not
drawn because the API does not return them. Legacy usage, trace and
middleware styles and the usage source components are removed.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Traffic hub

New specs for the overview (figures, charts with a data table and arrow
key readout, failed and first-run and tracing-off states, the shared
window, a phone), usage (group by, filters, sort, export, cost, quotas, a
non-admin, empty and loading), models, GPU and host (snapshot, the
since-opened labelling, a cluster), the traces list, a trace page (the
real error, the timeline, reveal, no headers, a trace that left the
buffer), Prometheus and the Middleware pipeline, with shared fixtures.

The usage, traces, middleware, hub and smoke specs follow the new
structure and keep their intent.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the Traffic hub

Add an operations page for the Traffic tab: which record each page reads,
what it leaves out and why, the trace page and its reveal, the GPU and
host readings kept since the page opened, and the Prometheus endpoint.
Link it from the operations index, the tracing page and the middleware
page.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): let metric names wrap in the Prometheus table on a phone

The long metric names pushed the type and "on this server" columns out of
view. Names now wrap inside the table.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild Settings with groups by intent, search, a pending bar and history

The fifteen sections become eight groups by intent: memory and models,
speed and defaults, backends and galleries, access and security,
debugging and traces, agents and responses, swarm and sharing, look and
feel. Search covers names, descriptions, keys and the old section name, and
says where a result used to be.

Edits wait in a bar with Discard, Show diff and Apply. The diff lists old and
new values and the checks the browser can make: durations parse the way Go
parses them, a GPU memory budget is one the server accepts, a gallery box
holds JSON, and warnings repeat what the handler and the field text say.
Apply sends only the changed keys. Undo saves the previous values again; it is
a new save, not a rollback. History lists the changes applied from this
browser, since LocalAI keeps no settings log, and Revert stages the old value.

A value is marked as changed only where the built-in default is known from
the CLI defaults. A row says "Applies now" or "Needs restart" only where the
handler or the docs say so.

Three things were wrong before and are fixed with the rebuild: the gallery
boxes and the shared API keys box were sent under names the server ignores,
the "Enable CSRF Protection" switch showed the disable flag the wrong way
round, and every save restarted peer-to-peer networking because every field
was sent.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild Users and keys, Account, sign-in, invite and the 404 page

Users and keys is a tabbed page under the Settings tab: people, invites and
API keys. The people table filters by state and role, sorts, approves or
disables (disabling offers an undo that sets the status back), and opens a
side sheet for one person's features, model allow-list and limits. Role,
password reset and delete sit in the row menu; delete asks for the name.
Invites choose a lifetime of 1, 7 or 30 days and show the link once. API
keys can be created with a lifetime, are shown once in full, can be paused,
and are revoked after a ten second undo window in which nothing is sent.
LocalAI lists keys only to their owner, so the tab shows the signed-in
person's own keys and says so.

Account has Profile, Security, API keys and Usage. Usage shows the last 30
days, tokens by model and the limits an admin set. The Security tab now
shows for a GitHub or SSO account and says the password is not theirs to
change.

Sign-in asks for one field per step and draws a provider button only for a
provider /api/auth/status lists. It has the notice for a sign-up that waits
for approval, the first-admin screen, the key-only screen and the invite
page. An address outside the app now gets the 404 page too, which names the
address and lists the places the sidebar lists, with the same gates.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover Settings, Users and keys, Account, sign-in and the 404 page

Settings: groups, search by name, key and old section, the changed marker
only where a default is known, apply hints, the pending bar and diff, the
checks, apply sending only changed keys, undo as a second save, discard,
history, the CSRF inversion and the gallery and API key wire forms, and the
phone layout.

Users and keys: the table, filters, sort, approve, disable with undo, the row
menu, the access sheet, invites, key creation with a one-time reveal, the
ten second revoke with undo and with a page leave, and the non-admin redirect.
Account, each sign-in variant (error, pending, first admin, key-only, invite,
provider buttons) and the 404 page have specs too. Fixtures are shared with
the screenshot scripts. Existing specs follow the new structure.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the rebuilt Settings, Users and keys, Account and sign-in pages

Runtime settings: the eight groups and where each old section went, search,
the pending bar, the diff and its checks, apply, undo, the history, and which
settings show a default or an apply note and why. Authentication: the
sign-in screen variants, the Account tabs, key lifetimes, the one-time key
reveal, the revoke undo window, and the fact that keys are listed only to
their owner.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): let the Settings undo toast stand alone and read back a generated P2P token

The saved message and the undo toast sat on the same spot at the bottom of the
page. The undo toast now carries the saved message.

A new P2P token is made by the server when the page sends 0. The page reads
it back after the save so the field shows the token and not the placeholder.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): fit the users table, API keys and Account figures on a phone

On a phone the users table dropped its Role and Status columns off the screen
edge with the row actions. The role and state now sit under the name, so the
actions stay in view. API key rows no longer put the key icon on a line of its
own, and the three Account figures keep one row.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the phone users table, reduced motion and the empty Account state

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): drop the apply note from three settings the save handler does not mention

Size-aware eviction, automatic backend upgrades and development backends said
Applies now, but nothing in the handler or the docs says when they take effect.
A row now carries a note only where the code or the docs say so.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild Voices and Faces as one identity family

Voices is one page with three tabs: Speakers (voiceprints for recognising
who is speaking), Speech voices (the text-to-speech reference library,
kept apart because it is a different store) and From a recording (a link
into the diarization workspace). Faces uses the same layout.

Who is this and Same person? give the answer in a sentence with the real
distance and cut-off, a word for how far inside the cut-off it sits, and
a distance scale with the cut-off drawn on it. The cut-off slider re-reads
the answer in the browser; the identify call sends the cut-off, and verify
uses the threshold the model returns. The old confidence percentage is
gone because it is not a probability.

The server has no list call, so the people list stays in the browser and
the page says so. After a search that asked for more people than it got
back, a saved person the server did not return is marked, and people the
server returned that the browser does not know are listed. Nothing is
claimed from a short or cut-off search.

Enrolling is a sheet: sample, name, labels, permission. A copy of the
sample in the browser is opt-in, and an administrator can also keep the
recording as a speech voice in the same step. Removing a person waits ten
seconds behind an Undo toast and sends nothing before then.

Errors say what happened (no face found, model missing, call failed), a
blocked or missing microphone is explained, and a missing model or a
missing permission renders a page that says what turns the feature on
instead of a redirect. Analyze, detect and raw embedding move under
More tools, with attribute guesses off by default.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Voices and Faces pages

Specs for who is this (match, no match, working, failed, missing model),
the cut-off slider, same person, a blocked, allowed and insecure
microphone, the registry notes and the not-on-the-server marks, the
enrol sheet and its opt-in copy, delete with undo on a fake clock, the
disabled and no-permission states, the phone layout, reduced motion and
Faces. Existing library and diarization specs follow the new structure
and keep their intent. Node tests cover the distance words, scale
layout, stored list and error mapping.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the Voices and Faces pages

Add a WebUI section to the voice and face recognition pages: the two
tools, the cut-off, what the people list is and why it can be stale, the
undo window, and what is stored where. Point the Voice Library and
Fish Audio notes at Build, Voices, Speech voices.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild the Build landing, Fine-tune, Quantize, Import and Explorer

The Build landing says what each tool is for and what it needs from the
machine: the installed backend, the GPU memory, RAM and disk the server
reports, and a job that is running or the newest one when it failed. A
tool that cannot run says why and what enables it.

Fine-tune and Quantize share one page: set up, a check list that is
redrawn as the form changes, a run view with progress, stages and a log,
and a result with real next steps (export, import, chat, Models). The
checks state only what the server reports. A job needs no estimate the
server cannot make, so none is invented. Stop on a fine-tuning job asks
whether to keep a checkpoint, a failed job shows the server's message,
and a memory failure offers two changes that are applied to a copy of
the setup.

Import is a guided flow: source, review, import, done. The server
returns no preview before an import starts, so the review reads the
spelling of the source, prints the request the form will send and runs
the checks that can be made early. The estimate that arrives when the
import starts is set against free memory and disk. The ambiguity picker
and the Write YAML tab stay.

Explorer shows what GET /networks returns and lists a swarm with POST
/network/add, with a join sheet that carries the token and commands.
Build tools the account may not use say so instead of redirecting.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Build landing, the tool pages, Import and Explorer

Specs for the landing (a tool ready, missing a backend, with no GPU, a
running or failed job, a feature switched off, a member without admin,
phone, reduced motion), the shared tool pattern for Fine-tune and
Quantize (set up, live checks, start request, running with progress,
chart and log, the stop choice, failure with the server message, finish
with next steps, earlier jobs, the account-disabled page, phone), Import
(source detection, review, checks, ambiguity, running with the estimate
against free memory, done, Write YAML, phone) and Explorer (list, join,
list a swarm, empty, not an explorer, retry, phone).

Existing specs follow the new structure and keep their intent. Node
tests cover the machine facts, tool status, checks, log lines, source
detection, the import request and the join commands.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the Build tool pages, the import flow and the Explorer

Fine-tuning and quantization now describe the set up, check, run and
result steps and what the check list can and cannot say. The import
section explains the review step and why the size and memory appear only
after the import starts. The distributed page describes the Explorer
list, the join sheet and what listing a swarm publishes.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): turn the hardware recommendations into a "Best for this machine" shelf

The shelf in the Models inspector put five columns into a 400 px pane,
so long model ids wrapped letter by letter underneath the size and the
memory figures. Each row now stacks the tag, the id and the size and
memory facts beside one Install button, and the id wraps inside its own
column.

Once a model is installed the shelf narrows to the best fit and keeps
the others behind a "N more that fit" toggle. Specs cover the ranking,
the layout, the narrowing and the install request against a gallery
fixture that carries the 4K estimate the shelf sizes against.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): quiet the Studio tab markers and say their state in words

The type tabs drew a saturated green dot for every modality that has a
model. The dot now uses a text colour, filled when a model is installed
and hollow when none is, and each tab carries "(model installed)" or
"(no model installed)" as hidden text so the state is not only a colour.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): stack the model editor empty state and drop its section hues

"No fields configured" sat in a flex row, so the icon, the title and
the text ran together. It now uses the stacked empty-state layout. The
section icons took a different status colour each (amber, red, green);
they now share one quiet colour, with the accent on the current section.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): keep the Home memory sentence whole on a phone and drop side rails

On a 390 px screen the memory strip clipped "2 models loaded" to make
room for the figure. The sentence now takes the first line and the
figure and device wrap under it.

The sweep also removed coloured left rails from the editor section
rail, the skill editor list, the install strip and the audio transform
notice (now an outlined note), plus unused chat rules that carried
rails and two glow animations that nothing referenced.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): line up hub pages, list the model templates, and stop clipped text

Medium-width pages inside Build and Operate were centred while the tab
bar above them was flush left, so the title started 60 px right of the
first tab. They now start at the bar's edge.

Add Model offered nine templates as a grid of identical cards with chip
clouds and inline styles. It is now one list of rows, each with the
field names it fills in on a single muted line.

Two clipped strings are fixed: the Studio voice field cut its
placeholder mid-word, and the phone job list ended the schedule line in
an ellipsis. The recommendation shelf also separates size and memory
with a dot, and the docs describe the shelf.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): keep the hidden Studio tab state inside its tab

The hidden state text added to each type tab was absolutely positioned
against the page, so on a phone it sat outside the scrolling tab row and
widened the page by hundreds of pixels. The tab is now the containing
block.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the on-disk sizes on the Installed table, model page and cleanup sheet

The fixtures stub GET /api/models/storage. The default report is empty,
so existing specs keep the gallery estimates. makeStorage() builds a
report from files and the models that use them, the way the server
does, and storageSpec() is a models directory with shared and missing
files.

New specs cover the Size column and its shared line, the fallback when
the call fails or the user is not an admin, the files list on the model
page, a missing file, the bytes a removal frees with shared files, and
the cleanup findings. Node tests cover the storage helpers and the
batch arithmetic.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): wait for the page before pressing keys and ticking the clock

Two specs failed in loaded full runs and passed alone. The Alt+1 to
Alt+7 spec pressed a key before the composer had armed its key
handler. The capacity chart spec advanced the fake clock before the
poller had mounted, so it counted fewer readings than it expected.

Both now wait for the page to mount. The key spec retries a press that
lands during a re-render, and the clock spec advances in small steps and
polls for the row count.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): hide the Installed footer when the storage report is empty

An empty report from the storage call made the footer read "0.0 GB on
disk" next to sizes taken from the gallery estimate. An empty report
says nothing about the disk, so the footer now shows only the model
count. A spec covers it.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): keep Explore pane actions inside the pane

The inspector actions sat in a non-wrapping flex row beside the title,
so the buttons ran past the pane edge once it got narrow. The row now
takes its own line and wraps.

The primary action (Install, Retry, Open) comes first. Manage
installation becomes a ghost button, and Open details moves to the end
of the row, so one action stands out and the others are quiet. No
action or test id is removed.

Add a spec that checks, in light and dark at several widths and with a
pane forced to 320 px, that every action stays inside the pane box and
that the pane keeps its inner padding.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* refactor(ui): drop the type chips from the Studio composer

The Studio tabs and the composer's type chips listed the same seven
modes, so the page said the same thing twice. Keep the tabs as the one
place to switch modes.

The composer now shows the type it will open as a small label in its
header. The type suggestion from the typed words stays as the quiet
hint line under the prompt, and Alt+1 to Alt+7 still pick a type. The
composer root carries data-type, data-types and data-missing so tests
can read the state.

Specs pick a type through a shared Alt+digit helper and read a
missing model from the tab dot instead of a chip. Remove the unused
chip locale strings and CSS.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* refactor(ui): remove the section crumb above page titles

Page headers drew a small uppercase crumb with a short rule before it
above the title. On the hub pages it repeated the hub name, so Build
sat above a heading that also said Build.

PageHeader now renders only the title, the supporting line and the
actions. Drop the eyebrow prop, the route-derived section name, its CSS
and the unused section helper, and remove the explicit eyebrow props
from the pages that passed one. Pages stay reachable through the
sidebar and the hub tab bar.

Add a spec that checks several pages show their title with nothing
ahead of it in the header.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* refactor(ui): remove left accent rails from tiles, rows and quotes

Several surfaces marked state with a coloured strip on the left edge.
Replace each one with a cue that is not a rail:

- Stat cards lose the strip; the icon and value still carry the colour.
- The highlighted card is a raised surface with a firmer edge.
- The selected rail row is an accent wash with a hairline outline.
- The status stripe on rail items is a small status dot.
- The active failover row is a tinted row.
- Quotes in markdown and chat prose are italic instead of barred.
- The variant detail panel has a full hairline border.

Add a spec that walks the main routes in light and dark and fails on a
left border thicker than 1px, a sideways inset shadow, a narrow
absolute strip in ::before or ::after, or a narrow tall child pinned to
a left edge.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
authored and GitHub committed 2026-10-08 22:03:07 +02:00
1 parent 66d046edee
commit 4d7bdc6ff2
579 files changed
+77619 -30833

No files matched your search

@@ -246,6 +246,48 @@ These settings apply to most LLM backends (llama.cpp, vLLM, etc.):
| `main_gpu` | string | Main GPU identifier for multi-GPU setups |
| `cuda` | bool | Explicitly enable/disable CUDA |
### Placement
The model page's **Configuration** tab and the model editor have a **Placement**
section for the keys above. It edits `gpu_layers`, `tensor_split`, `main_gpu`
and `context_size` and nothing else.
- **CPU only** writes `gpu_layers: 0`.
- **Auto** leaves `gpu_layers` unset. LocalAI then asks the backend for every
layer (`99999999`), and the `llama-cpp` backend lowers that to what fits in the
free device memory unless the model turns off `fit_params` (see [GPU auto-fit settings](#gpu-auto-fit-mode)). Other backends read
an unset value as their own default.
- **Custom** writes the number you enter. **All layers** writes `99999999`, the
value LocalAI uses when the key is unset. A slider appears when the memory
estimate reports the model's layer count (`block_count`); today it does not, so
you enter a number.
- With two or more GPUs, **Split across GPUs** writes `tensor_split` as
percentages (for example `65,35`; llama.cpp reads them as proportions) and
**Main GPU** writes `main_gpu`. **Split by free memory** sets the shares from
the memory each card has free now.
The bars show, for each GPU and for system memory, what other programs use, what
this model needs (weights and working memory, and the KV cache that grows with
`context_size`) and what is left. They come from the device list in
`/api/resources` and from `POST /api/models/vram-estimate`, which returns one
total for the chosen context size and layer count. The per-device split of that
total follows `tensor_split`, so it is an approximation. The page says what the
choice means in words ("Fits in GPU", "Spills to CPU", "Too many layers for the
GPU") and does not state a speed. A usable limit is 95 percent of the free
memory.
**Fit it for me** searches for the largest `gpu_layers` whose estimate fits the
free GPU memory, by asking the estimate endpoint, and **Undo** puts the previous
value back. It is hidden when the estimate is unavailable (the model file is
missing, still downloading or in a format the estimate does not cover) and when
the host has no GPU. A server that schedules models onto other machines shows no
per-device bars, because the devices listed belong to the controller.
Saving from the model page sends only the keys that changed. A key set back to
Auto is sent as `null`. The patch endpoint merges values into the file and cannot
delete a key, so the file then reads `gpu_layers: null`, which LocalAI treats as
unset. To remove the line itself, edit the YAML.
### Mixed CPU/GPU inference
The `llama-cpp` backend can run one GGUF model across CPU and GPU, using both system RAM and GPU VRAM.
+28 -2
View File
@@ -43,7 +43,24 @@ LOCALAI_DISABLE_AGENTS=true
1. Navigate to the **Agents** page in the web UI
2. Click **Create Agent** or import one from the [Agent Hub](https://agenthub.localai.io)
3. Configure the agent's name, model, system prompt, and actions
4. Save and start chatting
4. Open **Preview** to read the configuration that will be saved and, when editing, what differs from the saved agent. Secret values are hidden in the preview.
5. Save, then give the agent a task from its page
The form folds into sections. Each section shows a mark when it is ready and one line that says what it holds. **Start from** offers a few starting points. **Draft** asks the model you chose to write a name, a description and instructions from one sentence; it is optional, nothing is saved until you save, and the draft may be wrong.
### Running a task
Open an agent from the Agents page to see its model, tools, memory, skills and instructions, and a box to give it a task. A task starts a **run** with its own address, `/app/agents/<name>/runs/<id>`. While the agent works the page shows the thread: the steps folded into one line ("Worked 26 s, 5 steps"), the tool in use and the answer as it arrives. About a second and a half after the agent answers, the page settles into a report: the task, the outcome, the evidence (what each tool returned) and the steps. Follow-ups sit under the outcome and carry the earlier turns. Wide tables and code open wider on demand.
The server keeps no run history. Runs are recorded in the browser that watched them, up to 50 per agent, and the record holds task text, step text and answers. A run link opens only in the browser that recorded it, **Clear run record** on the agent page removes the record, and the last 14 runs of each agent appear as a strip on the Agents page. A run still marked as working five minutes after its last event reads as stopped. Durations are measured by the browser between sending the task and receiving the answer. The page does not show tokens or per-step timings from the server, and it has no Stop or approval control, because the agent API has neither; **Pause** on the agent stops it taking new work.
### Jobs and scheduled tasks
The **Jobs** tab lists tasks: a prompt a model runs for you, on a schedule or whenever you start it. The page opens with one sentence about the last 7 days, built from the jobs the server returned (how many ran, how many finished, failed, were cancelled or are still going, and which task failed last time). Each task shows its model, its schedule in plain words with the cron expression under it, its last 14 jobs and an enabled switch; **Run now** asks for the values the prompt uses (`{{.name}}` gaps) and any media to attach, and the menu holds Edit and Delete. Delete waits 30 seconds with an Undo button, and nothing is deleted on the server until that time ends.
The run history is grouped by day, with one sentence for each job (the first line of its result, or its error). A row opens to show the error or the start of the result and one next action, **Run again** for a finished job or **Cancel** for one that is still going. Run again starts a new job with the same parameters and media, because the API has no retry call. A job opens as a document: the task as that run sent it, the outcome, whether the webhook was delivered, and the steps the server recorded.
The task form takes a schedule as presets (hourly, daily, weekdays) with a time, or as a custom cron expression of five fields (minute, hour, day, month, weekday) or an `@hourly` style shortcut. The form checks the expression and says what it means in words. Times follow the clock of the machine that runs LocalAI, which the browser cannot read, so the page shows no next run. The page also shows no tokens, because the jobs API returns none, and durations come from the job's own start and end times.
### Importing an Agent
@@ -252,6 +269,15 @@ curl http://localhost:8080/api/agents/skills
If a skill you expect is missing, confirm LocalAI was started with `LOCALAI_AGENT_POOL_ENABLE_SKILLS=true`.
### The Skills and Memory libraries
The **Skills** and **Memory** pages work as libraries. A skill or a collection exists on its own and needs no agent. Each row says who uses it, read from the saved configuration of your agents: "Used by research-assistant +2", or a quiet "Not used yet". An agent uses a skill when `enable_skills` is on and the skill is in `selected_skills` (an empty selection means every skill). An agent reads the one collection that carries its own name, when `enable_kb` is on. Chat does not read skills or collections, so it is never listed as a user.
- **Add to...** on a skill opens a menu of agents. Each row shows an estimate of what the skill adds to every message: the characters of its content divided by four. With `skills_mode` set to `tools` the content is read only when the model asks, so nothing is added up front. Removing the last selected skill from an agent switches skills off for it, because an empty selection would mean every skill. Removing a skill or a collection from an agent can be undone for a few seconds.
- **Add to...** on a collection offers only the agent with the same name, and turns its knowledge base on.
- On a collection, **Try a question** searches that collection alone and shows the passages that come back with their scores (`POST /api/agents/collections/{name}/search`, with `max_results`). **Add source** uploads a file or adds a URL with a refresh interval in minutes. A failed upload shows the message the server returned, and can be retried.
- **Simulate a message** shows what a context would load. Pick a collection on its own, or one of your agents: you see the skills the agent has on, an estimate of the tokens it adds around the message, and the passages that its own collection returns for that message. Nothing is sent to a model, so no answer is shown. A skill or collection that is not added anywhere can be tried in the sheet for that test only, and is never saved.
## API Endpoints
All agent endpoints are grouped under `/api/agents/`:
@@ -388,7 +414,7 @@ curl -X POST http://localhost:8080/api/agents/my-agent/chat \
}'
```
The web UI does this for you: each conversation in the agent chat sends only its own visible turns. **New Chat** and switching conversations therefore continue from that conversation alone, and **Clear** starts the conversation over without history. A request without `history` is answered without earlier context; the server keeps no web chat history of its own. In distributed mode (NATS) the history is not forwarded yet.
The web UI does this for you: a run sends only its own turns as history, so a follow-up sees the task and the answers before it, and a new run starts without history. A request without `history` is answered without earlier context; the server keeps no web chat history of its own. In distributed mode (NATS) the history is not forwarded yet.
Listen to real-time events via SSE:
+1 -1
View File
@@ -238,7 +238,7 @@ voice conversion from the same weights.
## Family notes
- **Fish Audio voice cloning**: save a reference clip with its transcript in the
Voice Library, then select **Use in Text to Speech**. The backend accepts
Voice Library (**Build → Voices → Speech voices**), then select **Use in Text to Speech**. The backend accepts
`params.ref_text` as an alias for `params.reference_text` in both ordinary and
streaming speech requests. If you supply both parameters, `reference_text`
takes precedence. For direct requests with a reference file in `voice`, supply
+20 -3
View File
@@ -174,7 +174,7 @@ The **first user** to sign in is automatically assigned the admin role. Addition
### Invite Links
Admins can generate single-use, time-limited invite links from the **Users → Invites** tab in the web UI, or via the API:
Admins can generate single-use, time-limited invite links from **Settings → Users and keys → Invites** in the web UI (choose 1, 7 or 30 days), or via the API:
```bash
# Create an invite link (default: expires in 7 days)
@@ -192,7 +192,7 @@ curl -X DELETE http://localhost:8080/api/auth/admin/invites/<invite-id> \
-H "Authorization: Bearer <admin-key>"
```
Share the invite URL (`/invite/<code>`) with the user. When they open it, the registration form is pre-filled with the invite code. LocalAI validates the code only when the user submits registration. Invite codes are single-use - once consumed, they cannot be reused. Expired or used invites are rejected.
Share the invite URL (`/invite/<code>`) with the user. The link is shown once, when you create it: the invite list keeps only the first characters of the code, so a link you did not copy can only be revoked and made again. When the user opens it, the registration form is pre-filled with the invite code. LocalAI validates the code only when the user submits registration. Invite codes are single-use - once consumed, they cannot be reused. Expired or used invites are rejected.
For GitHub OAuth, the invite code is passed as a query parameter to the login URL (`/api/auth/github/login?invite_code=<code>`) and stored in a cookie during the OAuth flow.
@@ -236,10 +236,23 @@ When authentication is enabled, the following endpoints require admin role:
When auth is enabled, the React UI sidebar dynamically shows/hides sections based on the user's role:
- **All users see**: Home, Chat, Images, Video, TTS, Sound, Talk, Usage, API docs link
- **Admins also see**: Models, the Build console (Agents, Skills, Memory, Jobs, Training, Recognition), and the Operate console (Backends, Activity, Nodes, Usage, Traces, Users, Middleware, Settings)
- **Admins also see**: Models, the Build console (Agents, Skills, Memory, Jobs, Training, Recognition), and the Operate console (Backends, Activity, Nodes, Traffic, Settings, and Users and keys)
Admin-only pages are also protected at the router level - navigating directly to an admin URL redirects non-admin users to the home page.
### Sign-in, invite and account pages
The sign-in screen offers only what `GET /api/auth/status` reports:
- A GitHub or SSO button appears only when that provider is configured. With `LOCALAI_DISABLE_LOCAL_AUTH=true` there is no email form.
- Email and password sign-in asks for one field at a time: the email, then the password. Both are sent in one call, so the server still never says whether an email exists.
- Registration asks for the email, then the name, password and (when the registration mode allows one) an invite code. A sign-up that waits for approval returns to the sign-in screen with a notice that an admin must approve it.
- On a server with no users yet, the first screen is "Create the admin account".
- On a server protected only by legacy API keys, the screen asks for the key and nothing else.
- An invite link (`/invite/<code>`) opens registration with the code filled in and read-only.
The **Account** page has four tabs: Profile (name and picture address), Security (change a local password; for a GitHub or SSO account it says the password is managed by the provider), API keys, and Usage (your requests, tokens by model, and the limits an admin set on you, for the last 30 days). **Settings → Users and keys** has the people, the invites and your own API keys.
### GitHub OAuth Setup
1. Create a GitHub OAuth App at **Settings → Developer settings → OAuth Apps → New OAuth App**
@@ -315,6 +328,10 @@ curl -X PATCH http://localhost:8080/api/auth/api-keys/<key-id> \
The key list returns `disabled` and, when set, `pausedUntil` for each key. The Account page in the web UI has a Pause and Resume button for each key.
A key can be created with a lifetime: send `expiresIn` (`30d`, `90d` or `1y`) or `expiresAt` (RFC 3339). Without either, the server applies `LOCALAI_DEFAULT_API_KEY_EXPIRY` when it is set, and the key does not expire otherwise. The full key is returned once, in the response to the create call; the list returns only its prefix, the creation time, `lastUsed` and `expiresAt`.
In the web UI, the **API keys** tab of the Account page and of **Settings → Users and keys** shows the signed-in person's own keys: a new key is shown once with a Copy button, and the lifetime list sends `expiresIn`. Revoking a key waits ten seconds in the browser before it sends the `DELETE`, and Undo inside that time sends nothing; leaving the page sends a revoke that is still waiting. LocalAI lists a person's keys only to that person, so an admin cannot list or revoke other users' keys, in the UI or through the API. The Usage page can group requests by API key for everyone.
### Auth API Endpoints
| Method | Endpoint | Description | Auth Required |
+35 -10
View File
@@ -16,27 +16,52 @@ For the complete list of backends, the model families they support, and their ac
## Managing Backends in the UI
The **Operate → Backends** page is the canonical home for the complete backend
lifecycle:
The **Operate → Runtime → Backends** page is the canonical home for the
complete backend lifecycle. It is one list with two views, **Installed** and
**Catalog**, each with a count:
1. **Catalog** browses configured galleries, searches by name or description,
filters by capability, and installs a backend. Catalog is the default view.
2. **Installed** shows the runtimes present on the host or cluster. Search and
filter by user, system, update, or offline-node state, then select a backend
to inspect its version, source, node placement, and lifecycle actions.
On a new installation with nothing installed it recommends llama-cpp, which
runs most text models, with a single install button.
2. **Installed** shows the runtimes present on the host or cluster, those with
an update first. Search and filter by user, system, update, or offline-node
state.
3. Variant and development builds remain opt-in refinements. Target-node links
compose with the current view and selection instead of opening a separate
management page.
The current view, search, filter, selected backend, and target node are stored
in the URL. Browser Back and shared links therefore restore the same state.
Each row shows the backend, its version and its state: **Current**, **Update
1.2.0**, **Not installed**, **Queued**, **Removing**, or a progress bar while it
installs, with **Cancel** in the row. A failed install says why, with **Retry**.
The one action that matters sits in the row (**Install**, **Update**), and the
chevron opens the row for the rest. **Update all** starts every pending update,
and **Check for updates** asks the server to check now instead of waiting for
its next scheduled check. **From URL** installs a backend from an OCI image, a
URL or a path.
The current view, search, filter, open row, and target node are stored in the
URL. Browser Back and shared links therefore restore the same state.
**Removing** a backend asks first. The dialog names the configured models that
ask for that backend by name, and any installed meta backend that points at it
(for example `llama-cpp` pointing at a hardware-specific build), because those
stop working without it. The catalog does not report a size per backend, and
LocalAI keeps no earlier version of a backend, so there is no size column and
no rollback.
**Operate → Runtime → Logs** (`/app/backend-logs`) lists the backend processes
that have printed something. Open one to read its output live, filter by stream
or text, follow the end, show or hide times, export the lines as JSON, or clear
them. Clearing hides the lines at once and wipes them on the server after 6
seconds, unless you press **Undo**.
Installs run in the background. The strip at the top of the app follows the
current one, and **Operate → Activity** lists everything in flight, what needs
attention, and what has finished, and is where a running install is cancelled
or a failed one retried. See [Activity]({{% relref "operations/activity" %}}).
attention, and what has finished, and is where a running install is paused,
cancelled or a failed one retried. See [Activity]({{% relref "operations/activity" %}}).
Each selected backend displays:
Each open row displays:
- Backend name and description
- Type of models it supports
- Installation status
+73
View File
@@ -0,0 +1,73 @@
+++
disableToc = false
title = "Chat"
weight = 48
url = "/features/chat/"
+++
Chat is the conversation page of the web UI. It keeps one thread per conversation, in your browser, and sends it to a chat model of your choice. The message box is the same one Home uses, so a message typed on Home opens here with the same model and the same tools.
## The thread
Your messages are the raised blocks on the right. The model's replies are plain text under its name, with a dot that is filled when the server holds the model in memory and hollow when it is not loaded yet. Code blocks have **Copy** and **Canvas** buttons. Images you attach appear as thumbnails that open in a viewer, and other files appear as chips.
Reasoning and tool calls fold into one quiet line (for example "Thought · read_file"). Click it to open the steps with their arguments and results. While a reply is arriving the line is open and shimmers, and it folds when the answer starts.
If a reply fails, the text written so far is kept, the reason is shown, and **Retry** asks again from your last message. **View traces** opens the backend traces.
### Actions on a message
Hover a message, focus it, or look at the last one (on a phone every message shows them): **Copy**, **Edit**, **Regenerate** and **Branch from here**. Edit saves the new text in place and does not send anything. Regenerate drops the answer and everything after it and asks again. Branch starts a new chat with the history up to that answer.
| Key | Action |
|---|---|
| `Up`, `Down` | Move between messages (a message must have focus) |
| `C`, `E`, `R`, `B` | Copy, edit, regenerate or branch the focused message |
| `Up` in an empty box | Edit your last message |
| `Esc` | Stop the reply that is streaming; then close the search; then close the canvas |
## The message box
The model chip lists the chat models. **Loaded now** holds the models in memory, **Installed** the rest. When the list opens, LocalAI reads the host memory once and asks the server to estimate each listed model at this chat's context size (up to twelve models). A row shows what the model needs and whether it fits: free memory, how much would run on the CPU, or how far over the machine it is. A model with no estimate shows no fit text. The server reports no load time, so none is shown.
The **MCP** chip opens the server and client tool lists. **Canvas** turns code blocks into cards that open in a side panel. Type `/` for the actions the page has:
| Action | What it does |
|---|---|
| `/model` | Open the model list |
| `/new` | Start an empty chat on the same model |
| `/chats` | Open the conversations list |
| `/assistant` | Turn Manage mode on or off (admin only) |
| `/canvas` | Turn Canvas on or off |
| `/find` | Search this chat |
| `/settings` | Open the chat settings |
| `/export` | Download the chat as Markdown |
| `/clear` | Remove every message, after a confirmation |
Enter sends, Shift+Enter adds a line. Paste an image to attach it. Under the box, the line shows the speed while a reply streams and the token count of the chat.
## Conversations
Press `Ctrl+K` (`Cmd+K`) or **Chats** to open the list of conversations, grouped by day like **Jump back in** on Home. Type to search names and message text. Arrow keys move, `Enter` opens, `F2` renames and `Delete` removes. Each row also offers rename, duplicate, copy and export. Removing a chat hides it and shows an **Undo** toast; the chat is deleted for good when the toast goes away. The name in the header can be renamed with a click.
The history is stored in this browser (`localai_chats_data`), not on the server.
## Settings
The sliders button opens the chat settings. They apply from the next message.
- **System prompt.** Sent before every message in this chat. Empty means the model's own default.
- **Sampling.** Temperature, top P and top K. Each says "model default" until you change it, and has a **Reset**.
- **Context window.** The size drives the meter in the header and the warning when the context is nearly full. It is not sent to the model. Admins get it filled in from the model's configuration.
- **Behaviour.** Manage mode (admin) and Focus mode, which collapses the app sidebar while a conversation is open.
- **Model info** (admin). Backend, model file, context size, threads, GPU layers and an **Edit config** button.
## Find, canvas and long threads
`Ctrl+Shift+F` (`Cmd+Shift+F`) searches this chat: it marks the matches in the messages that are loaded in the page, shows "n of m" and steps with `Enter` and `Shift+Enter`. Nothing is sent to the server. When you scroll away from the end of a long thread, **Jump to latest** brings you back.
With Canvas on, code blocks become cards. The canvas panel opens beside the thread with tabs, a Code and Preview switch for HTML, SVG and Markdown, Copy and Download. On a narrow window it takes the whole page.
## When there is nothing to chat with
An empty chat shows a few starters, whether the model is loaded and your recent conversations. With no chat model installed, the page offers the starter models for the hardware, the gallery and import, and keeps what you typed. While a model is loading on a worker or being staged, the reply waits behind a card that names the phase the server reports, the node, the bytes and the time left when the server gives them, and then sends by itself. Stop cancels the wait.
+16 -14
View File
@@ -559,9 +559,17 @@ When the SmartRouter needs to free capacity, it can unload models with zero in-f
### Managing nodes in the WebUI
Open **Operate → Nodes** to inspect fleet health, filter or select workers, and view running models across the cluster. The **Running models** view groups replicas by model. Its **View logs…** action opens logs directly when there is one placement; when a model has several placements, it opens the model inspector so you can choose all logs for one node or the logs for one replica.
With distributed mode on, **Operate → Swarm** holds the cluster: **Nodes**, **Placement rules**, **Failover** and **P2P**. A single-node install shows **This machine** instead and never draws these pages.
Open a node's full details for node-scoped work: viewing replica logs, unloading a model, managing installed backends, changing replica capacity, or editing scheduling labels. Diagnostic actions are listed before destructive actions in row menus.
**Nodes** lists every worker in a sortable table: name, role, state in words, GPU or system memory, loaded models, last heartbeat and version. Switch between comfortable and compact rows, between **List**, **Map** and **Running models**, and filter to **Needs attention** (waiting for approval, not answering, or low GPU memory, system memory or models disk). The **Map** draws this instance, the message bus and database, and each worker; a dashed line is a worker that gets no traffic. It is not drawn on a phone. When a backend has a newer version, an **Update** button on the page sends the upgrade to the nodes that differ from the rest of the cluster, or to the nodes you selected.
The **Running models** view groups replicas by model. Its **View logs…** action opens logs directly when there is one placement. When a model has several, it opens the row so you can choose the logs of one replica.
**Add a node** (`/app/nodes/add`) explains how a machine joins: a registered worker, a peer instance, or a memory shard. It prints the command to run on the new machine with a **Copy** button, and updates the page when the machine appears. On a single-node install it starts with the command that turns distributed mode on.
Open a node for node-scoped work. The page shows its state, VRAM, RAM, models disk, CPU and in-flight requests, and has tabs for **Models** (replica logs, unload), **Backends** (upgrade, delete), **Logs**, and **Capacity and labels** (replica capacity, labels). **Drain…** shows what the drain would change before it does anything; **Remove…** asks for the node's name. A node that stopped answering says so and shows the last figures it reported.
The "what happens if I drain this node" list is a preview worked out in the browser from the node list, the loaded replicas and the placement rules. The server does not compute it, and the scheduler also weighs free memory and disk when it loads a model, so the preview never claims a model will fit.
## Node Management API
@@ -601,17 +609,11 @@ Used by the WebUI and admin API consumers. Requires admin authentication.
| `PUT` | `/api/nodes/:id/vram-budget` | Set a VRAM budget for a worker (`{"value":"80%"}`) |
| `DELETE` | `/api/nodes/:id/vram-budget` | Clear a worker's VRAM budget (revert to all detected VRAM) |
The **Nodes** page in the React WebUI is a fleet operations dashboard. Its health band and VRAM, RAM, CPU, and models-disk gauges aggregate the single `GET /api/nodes` response and identify how many workers do not report each metric. The attention queue isolates pending, impaired, or low-capacity workers without double-counting the headline affected-node total.
The **Nodes** page reads `GET /api/nodes` every five seconds and renders 50 workers at a time. Bulk drain, resume and remove run with bounded concurrency, so the page stays usable for fleets with thousands of registrations. Selecting the visible page or a group does not discard selections elsewhere; selections are removed only when a later poll confirms the worker no longer exists.
The fleet table supports search, status and type filters, label or type grouping, sortable columns, and selection across filters. It renders 50 workers at a time and bulk drain, resume, and remove operations run with bounded concurrency, so the page remains usable for fleets with thousands of registrations. Selecting the visible page or a group does not discard selections elsewhere; selections are removed only when a later poll confirms the worker no longer exists.
The list never fetches backend inventory, and the **Running models** and **Map** views stay lazy: the first use of either makes one controller database request that is kept until the page is left. A node's backends are read only when its page opens.
Selecting a row opens an in-context inspector with health, labels, capacity, model activity, and heartbeat details. Backend inventory is fetched only for the open inspector. The inspector links to the dedicated node detail page at `/app/nodes/:id`, where model, backend, label, capacity, CPU utilization and load, and models-disk management remain available. Model scheduling lives on its own **Scheduling** page.
The workbench's **Running models** tab shows the current loaded replicas on healthy workers. It stays lazy: opening the Nodes page does not query model inventory, and the first activation makes one controller database request that is retained until the page is left. The view groups replicas by model, reports their worker spread, active requests, backend types, and most recent use, and renders 50 models per page for large fleets. Loading, empty, and query-failure states are shown in place; a failed query can be retried.
Use a model row's actions menu to stop that model across the fleet. LocalAI sends one controller shutdown request for the model, which stops all loaded placements; the browser does not contact workers individually. The dashboard refreshes the running-model inventory after both successful and failed shutdown attempts because a failed request can still have stopped some replicas.
Opening a model reveals its replica placement without another request. Replicas on the same worker remain individually visible with their process addresses and workload. From there, select a known worker to move into its node inspector, then return to the model with **Back to model**. That worker transition is the only point in this flow that requests backend inventory, preserving the Nodes page's no-prefetch behavior.
Use a model row's actions menu in **Running models** to stop that model across the fleet. LocalAI sends one controller shutdown request for the model, which stops all loaded placements; the browser does not contact workers individually. The view refreshes the running-model inventory after both successful and failed shutdown attempts because a failed request can still have stopped some replicas.
### Model sizing in the WebUI
@@ -790,7 +792,7 @@ curl -X POST http://frontend:8080/api/nodes/<node-id>/approve \
-H "Authorization: Bearer <admin-token>"
```
The **Nodes** page in the WebUI also shows pending nodes with an **Approve** button.
The **Nodes** page in the WebUI shows pending nodes with an **Approve** button, and so does the node's own page and the **Add a node** page once the machine has registered.
To skip manual approval and let nodes join immediately, set `--auto-approve-nodes` (or `LOCALAI_AUTO_APPROVE_NODES=true`) on the frontend. This is convenient for development and trusted environments.
@@ -1136,7 +1138,7 @@ local-ai worker \
## Model Scheduling
Model scheduling controls where models are placed and how many replicas are maintained. In the React WebUI it has its own **Scheduling** page (a top-level nav item, separate from the Nodes page). It combines two optional features:
Model scheduling controls where models are placed and how many replicas are maintained. In the React WebUI it has its own **Placement rules** page (**Operate → Swarm → Placement rules**, at `/app/scheduling`). Each rule is written as a sentence, shows the nodes the model is loaded on now, and opens in a side sheet that previews which nodes the draft rule could use. The preview is worked out in the browser from the node list and labels; the scheduler also checks free memory and disk when it loads a model. A deleted rule can be taken back for a few seconds. A rule combines two optional features:
### Node Selectors
@@ -1210,7 +1212,7 @@ curl -X POST http://frontend:8080/api/nodes/scheduling \
This makes an alias a stable deployment slot: the placement policy belongs to
the slot, and the model filling it can change without rewriting the rule. The
WebUI lists aliases in the model picker on the **Scheduling** page, tagged with
WebUI lists aliases in the model picker on the **Placement rules** page, tagged with
the model each one resolves to.
Each frontend resolves the alias from its own copy of the model configs, and a
@@ -33,11 +33,13 @@ LocalAI supports two modes of distributed inferencing via p2p:
A list of global instances shared by the community is available at [explorer.localai.io](https://explorer.localai.io).
A LocalAI server started with `local-ai explorer` serves the same list at `/explorer`. Each swarm shows its name, description, the types of its clusters, how many workers are online and a shortened token. **How to join** shows the whole token and, for a federated cluster, the Docker and command-line commands that start a node on it. **List a swarm** adds a swarm to the list. Listing publishes its token, so anyone who sees the list can use the swarm's workers. Only swarms with at least one online worker are listed. On a server that is not in explorer mode, the page says so.
## Usage
Starting LocalAI with `--p2p` generates a shared token for connecting multiple instances: and that's all you need to create AI clusters, eliminating the need for intricate network setups.
Simply navigate to the "Swarm" section in the WebUI and follow the on-screen instructions.
Navigate to **Operate → Swarm → P2P** in the WebUI, or open **Add a node** and choose a peer instance or a memory shard, and follow the on-screen instructions.
For fully shared instances, initiate LocalAI with --p2p --federated and adhere to the Swarm section's guidance. This feature, while still experimental, offers a tech preview quality experience.
@@ -63,7 +65,7 @@ local-ai federated
To see all the available options, run `local-ai federated --help`.
The instructions are displayed in the "Swarm" section of the WebUI, guiding you through the process of connecting multiple instances.
The instructions are on the **P2P** page of the Swarm hub and in **Add a node**, guiding you through the process of connecting multiple instances.
### Workers mode
@@ -79,7 +81,7 @@ To connect multiple workers to a single LocalAI instance, start first a server i
local-ai run --p2p
```
And navigate the WebUI to the "Swarm" section to see the instructions to connect multiple workers to the network.
And open **Operate → Swarm → P2P** to see the network token and the instructions to connect multiple workers to the network.
![346663124-1d2324fd-8b55-4fa2-9856-721a467969c2](https://github.com/user-attachments/assets/b8cadddf-a467-49cf-a1ed-8850de95366d)
@@ -396,6 +396,12 @@ The recommended default `threshold` for `/v1/face/verify` and
Pass `threshold` explicitly when switching engines - the per-engine
default only fires when the field is omitted.
## The WebUI page
**Build → Faces** follows the same layout as the Voices page. **Who is this** looks up a photo against the people you enrolled (`POST /v1/face/identify`, cut-off 0.35 by default, adjustable on the scale). **Same person?** compares two photos (`POST /v1/face/verify`), optionally with the liveness check, and draws the face the model found in each photo. **Enrol a person** takes a photo, a name, optional labels and a permission tick; a copy of the photo stays in the browser only if you tick it.
The people list is kept in the browser because the server has no list call. After a search, a saved person the server did not return is marked "not on the server", and a server restart empties the server's index. Removing a person waits ten seconds behind an Undo toast before `POST /v1/face/forget` is sent. The server matches one face per photo. Detecting faces, attribute guesses (off by default, often wrong) and the raw embedding are under **More tools**. With no face model installed the page says what is missing and offers gallery models to install.
## Related features
- [Object Detection](/features/object-detection/) - generic bounding-box
+9 -5
View File
@@ -203,12 +203,16 @@ curl -X POST http://localhost:8080/api/fine-tuning/jobs \
## Web UI
When fine-tuning is enabled, a "Fine-Tune" page appears in the sidebar under the Agents section. The UI provides:
**Build → Fine-Tune** is one page that follows the job from set-up to result. A line at the top shows where you are: set up, check, run, result.
1. **Job Configuration** - Select backend, model, training method, adapter type, and hyperparameters
2. **Dataset Upload** - Upload local datasets or reference HuggingFace datasets
3. **Training Monitor** - Real-time loss chart, progress bar, metrics display
4. **Export** - Export trained models in various formats
1. **Set up** - Choose the base model, the dataset (a Hugging Face id or an uploaded file), the kind of training (an adapter, or the full model) and the epochs, batch size and learning rate. Method, backend, adapter settings, optimizer, evaluation, reward functions (GRPO) and extra options are under **More options**.
2. **Check before you start** - A list of what is known before the job starts, redrawn as you change the form: whether the model and dataset are set, whether a fine-tuning backend is installed, how much GPU memory (or RAM, when there is no GPU) is free, how much space is free on the models disk, and whether a token is set. LocalAI does not estimate how much memory a job needs or how long it takes, because the server cannot know that before the job starts, so the page says so. A missing model or dataset blocks the start. A warning, such as a missing backend or no GPU, does not.
3. **Run** - Percent, step, epoch, the server's time estimate and tokens per second, the stages the job passes through, a chart of loss (with the evaluation loss as hollow dots), learning rate and gradient norm, and a log of what the job reported since the page opened. **Stop** asks whether to keep a checkpoint.
4. **Result** - A failed job shows the server's message. When the message says memory ran out, the page offers two changes (a batch size of 1 and gradient checkpointing) that are applied to a copy of the setup and start nothing until you press Start. A finished job lists its checkpoints and exports the result as a model, then links to a chat with it and to Models.
Earlier jobs are listed below the form. **Reuse** puts a job's setup back in the form.
A user needs the fine-tuning permission. Without it, the page says the account cannot fine-tune and who can change that.
## Dataset Formats
+2 -2
View File
@@ -13,9 +13,9 @@ The same MCP server is published as a Go package and can also be served over **s
## Enabling the assistant in chat
Open the chat UI as an **admin** user and pick a chat-capable model in the model selector. The header shows a **Manage** toggle - flip it on, and a `Manage mode` badge appears next to the chat title. Starter chips ("What is installed?", "Install a chat model", "Show system status", "Update a backend") help you get going.
Open the chat UI as an **admin** user and pick a chat-capable model with the model chip above the message box. Open **Chat settings** (the sliders button in the header) and turn on **Manage mode**, or type `/assistant` in the message box. A shield icon appears next to the chat title while the mode is on. In a new chat in Manage mode, starter chips ("What is installed?", "Install a chat model", "Show system status", "Update a backend") help you get going.
The home page also exposes a **Manage by chat** CTA that opens a fresh chat already in Manage mode.
The home page shows a one-line **Manage LocalAI by chatting** prompt that opens a fresh chat already in Manage mode. You can dismiss it. The same action stays available as `/assistant` in the command bar and in the Library row.
Once on, try:
+1 -1
View File
@@ -562,7 +562,7 @@ In addition to server-side MCP (where the backend connects to MCP servers), Loca
### How It Works
1. **Add servers in the UI**: Click **MCP** in the chat header, open the **Client** tab, and add MCP server URLs
1. **Add servers in the UI**: Click the **MCP** chip above the message box, open the **Client** tab, and add MCP server URLs
2. **Browser connects directly**: The browser uses the MCP TypeScript SDK (`StreamableHTTPClientTransport` or `SSEClientTransport`) to connect to MCP servers
3. **Tool discovery**: Connected servers' tools are sent as `tools` in the chat request body
4. **Browser-side execution**: When the LLM calls a client-side tool, the browser executes it against the MCP server and sends the result back in a follow-up request
+10 -5
View File
@@ -246,11 +246,16 @@ continue. On a single LocalAI instance, a restart removes the pin. In
and a table of every target with its kind, warm flag, status, last probe
time and last error. An admin sees a **Pin** button on each target and an
**Unpin** action for the chain, both behind a confirmation dialog.
- The **Failover** page (`/app/failover`, admin only, linked from the
console navigation) lists every chain with its status pill, active target,
a small pill per target, and time since the last switch. It links each
chain name to its model editor page and shows an empty state linking to
the failover template when no chains exist yet.
- The **Failover** page (`/app/failover`, admin only) lists every chain with
its status pill, active target, a small pill per target, and time since the
last switch. It links each chain name to its model editor page and shows an
empty state linking to the failover template when no chains exist yet. It
sits under **Operate → Runtime** on a single install and under **Operate →
Swarm** when distributed mode is on. On a cluster it also describes what the
router and the health monitor do when a worker stops answering, and shows a
preview of what would stop if a chosen node went away, worked out in the
browser from the loaded replicas and the placement rules. The server does not
compute that preview.
- The Installed Models list badges a model that belongs to a chain with
`chain → <active target>`, next to the alias badge.
- All of the above update live from the same event stream as
+105 -11
View File
@@ -28,16 +28,110 @@ GPT and text generation models might have a license which is not permissive for
Open **Models** in the WebUI. It is the canonical page for a model's complete
lifecycle and has two views:
- **Explore** browses configured galleries, compares hardware fit and variants,
and installs models. This is the default view.
- **Installed** lists local model configurations and their running, idle,
disabled, pinned, and distributed state. Select a model to load or stop it,
edit its configuration, open a supported use case, inspect backend logs, or
remove it.
- **Explore** browses configured galleries and installs models. It is one
dense table: each row shows the model's size, a bar for the memory it needs at
the chosen context length, and the headroom in words ("3.7 free", "+1.5 on
CPU", "0.9 over"). Capability chips show how many models each one matches.
Select a row to open the details beside the table: the fit on this machine,
VRAM by context length, variants, files, links, tags and licence. This is the
default view.
- **Installed** lists local model configurations in the same table, with their
running, idle, disabled, pinned, and distributed state and their size on disk.
Load or stop a model from its row, or select it to edit its configuration,
open a supported use case, inspect backend logs, or remove it.
Both views use the same model selection and store the view, search, filter, and
selection in the URL. Installing from Explore does not move you away from the
catalog; the entry updates in place when the operation finishes.
Both views store the view, search, filter, and selection in the URL. Installing
from Explore does not move you away from the catalog; the entry updates in place
when the operation finishes.
### A model's own page
Every model also has a page of its own at `/app/models/<name>`, so it can be
linked. Open it with the arrow at the end of a row, a double click on the row,
or `o` on the selected row. On a phone, a tap on a row opens it. The details
beside the table stay as the quick look.
The title block names the model and holds the main action: **Install**, with a
chevron to choose which build to install, or **Load** and **Stop** with a menu
for an installed model (disable, pin, edit configuration, logs, delete). A strip
under it answers three questions: whether the model fits this machine, what it
does, and what installing leaves free on the models disk (or its state, when it
is installed). The tabs are:
- **Overview**: the description, backend, licence, largest context, tags, links
and a memory bar. For an installed model it also shows its state, the pages it
opens in, and which agents, agent tasks, failover chains and aliases name it.
- **Fit and memory**: the verdict in words, a context size selector, the memory
bar split into weights and the part that grows with context, and a chart of the
memory needed at each context size against the memory this machine offers.
When a model has several builds, pick the build to see its own figures.
- **Variants and files**: the builds with their size and fit, and the files the
chosen build downloads. Install any build from its row.
- **Usage and history**, **Configuration** and **Logs**, for installed models.
LocalAI does not record requests, timings, loads or configuration changes for
each model, so the usage tab lists what it cannot show yet instead of an empty
chart. The same tab lists the files the model uses on disk, with their size,
the other models that use them, and any file the configuration names that is
not on disk (see [Disk and cleanup](#disk-and-cleanup)). Configuration holds the [Placement](/advanced/model-configuration/#placement)
section. Logs is the backend log viewer of the Operate section.
The page reads the same lists as the table, so it works for a model the gallery
does not list (it has no variants or files tab) and for a gallery model that is
not installed (it has no usage, configuration or logs tab). If the gallery cannot
be reached, an installed model keeps working and Install says why it is off.
Going back with `Esc`, `Backspace` or the **Models** button returns to the list
with its view, search, filters, selection and scroll as you left them. The
previous and next buttons step through the rows of the list you came from, in
the order you saw them.
### Keyboard
On the Models page, `/` jumps to the search field, the up and down arrows move
the selection, `Enter` installs the selected model in Explore, `d` switches
between comfortable and compact rows, `o` opens the selected model's page, and
`Esc` closes the details.
On a model's page, `1` to `6` switch tabs, `[` and `]` (or `k` and `j`) step to
the previous and next model of the list, and `Esc` goes back to the list.
### Disk and cleanup
When the server reports the disk that holds the models directory, a strip in the
page header shows how much of it is free. It turns amber when less than 10
percent, or less than 20 GB, is free. In Explore, the details of a model say how
much disk an install leaves free. The strip is hidden when the disk cannot be
read, and on a distributed controller, where the models live on the workers.
Select the strip to open the cleanup review. LocalAI does not record when a
model was last used or how often, so the review says so and ranks installed
models only by what it can see:
- **Safe to remove**: another build of the same gallery model is installed, and
the build LocalAI would pick on this host is the one that stays.
- **Probably safe**: disabled, unused, and available in the gallery to download
again.
- **Your call**: not loaded and not used by anything, but with nothing more
known. A model that is not in the gallery cannot be downloaded again, and the
review says so.
- **Protected**: loaded, pinned, or named by an agent, an agent task, a failover
chain or an alias. These are never suggested. If an agent or task cannot be
read, nothing is marked safe.
For an admin, sizes are what each model uses on disk, read from
`GET /api/models/storage`. A file that another installed model also uses stays
on disk when you remove one of the two, so the review counts only the files a
model does not share. It says which models share files, and the amount freed
by a selection counts a shared file only when every model that uses it is in
the selection. A configuration that names files that are not on disk (a download
that did not finish, or files removed by hand) is listed as a finding. If the
report cannot be read, for example because the user is not an admin, sizes fall
back to the sizes of the files the gallery lists. A model that is not in the
gallery then shows no size and is not counted in what a removal frees. Before you
confirm, the review checks again and lists what will go, why, and how much it
frees. Removal then waits 30 seconds, during which you can undo it; the delete
request is sent only when that time ends. If you leave the page during the wait,
nothing is deleted.
## Cyber-Ornith 1.5 9B
@@ -171,9 +265,9 @@ This removal does not delete previously installed models. Remove that configurat
When browsing the gallery or importing a model by URI, LocalAI can show **estimated download size** and **estimated VRAM** for models.
- **Where they appear**: In the model gallery table (Size / VRAM column), in the model detail modal, and after starting an import from URI (in the success message).
- **Where they appear**: In the model gallery table (Size and Fit columns), in the model inspector beside it, and after starting an import from URI (in the success message).
- **How they are computed**: GGUF models use file size (HTTP HEAD or local stat) and optional GGUF metadata (HTTP Range) for KV cache and overhead; other formats use Hugging Face file sizes and optional config when available. If metadata is unavailable, a size-only heuristic is used. GGUF metadata lengths that exceed the file size are rejected before allocation; these files also use the size-only estimate.
- **Hardware fit indicator**: When your system reports GPU or RAM capacity, the gallery shows whether the estimated VRAM fits (green) or may not fit (red) using a 95% headroom rule.
- **Hardware fit indicator**: When your system reports GPU or RAM capacity, each row shows whether the estimated memory at the chosen context length fits, using a 95% headroom rule. A model that is too big for the GPU but would run from system RAM is marked as spilling to the CPU, with how much; a model too big for both is marked as over by the shortfall.
- Estimates are best-effort and may be missing if the server does not support HEAD/Range or the request times out.
## Useful Links and resources
+16
View File
@@ -645,3 +645,19 @@ This event is a LocalAI extension to the OpenAI Realtime API and is server-emitt
- [Realtime voice assistant demo (Go)](https://github.com/localai-org/localai-realtime-demo): a minimal Go client for the Realtime (WebSocket) API with a full talk-back voice loop and an example tool call. Ships a `docker compose` setup that brings up a realtime-capable LocalAI for you.
- [Realtime voice assistant example (Python)](https://github.com/mudler/LocalAI-examples/tree/main/realtime): thin-client architecture (Silero VAD on the client, heavy lifting on LocalAI), suited to running the client on a Raspberry Pi.
## Talk in the web UI
The **Talk** page of the web UI is a client for this API over WebRTC. Pick a pipeline model with the chip at the top, then press **Start session**. The microphone streams while the session runs. The server detects when you stop talking and answers on its own, and you can speak over a reply to interrupt it. The page has no push-to-talk or hands-free switch.
The heading under the orb says what is happening: connecting, listening, thinking (also while a tool runs), speaking, or interrupted after you cut a reply off. The orb follows the real microphone level while it listens and the reply's audio while it speaks. These states have their own screen:
- **Microphone blocked.** The browser did not give the page the microphone. Allow it in the address bar and press **Try again**.
- **Connection lost.** The WebRTC link failed during a session. The transcript stays and **Reconnect** starts a new session.
- **Something went wrong.** The server reported an error or the call could not be set up. The reason is shown, with a link to the traces.
- **Talk needs a pipeline model.** No pipeline model exists yet. The page links to the model editor with the pipeline template and to the gallery.
The transcript shows what you said, the replies and, in Manage mode, the tool calls and results. **Copy** puts it on the clipboard. It is not saved.
The sliders button opens the session settings: the instructions, the voice, the transcription language, client-side MCP tools, Manage mode (admin, fixed once a session is open) and the parts of the selected pipeline. The gauge button, shown during a session, adds audio diagnostics (waveform, spectrum and WebRTC statistics).
+9 -5
View File
@@ -142,12 +142,16 @@ The UI also supports entering a custom quantization type string for any format s
## Web UI
A "Quantize" page appears in the sidebar under the Tools section. The UI provides:
**Build → Quantize** uses the same page pattern as fine-tuning: set up, check, run, result.
1. **Job Configuration** - Select model, quantization type (dropdown with presets or custom input), backend, and HuggingFace token
2. **Progress Monitor** - Real-time progress bar and log output via SSE
3. **Jobs List** - View all quantization jobs with status, stop/delete actions
4. **Output** - Download the quantized GGUF file or import it directly into LocalAI for immediate use
1. **Set up** - Choose the model and the quantization type (a preset, or a custom type). The backend and the Hugging Face token are under **More options**.
2. **Check before you start** - Whether a model and type are set, whether a quantization backend is installed, the free RAM and the free space on the models disk. The server reports no model size before the job starts, so the page does not estimate one.
3. **Run** - Percent, the stage (downloading, converting, quantizing), and a log of what the job reported since the page opened. **Stop** ends the job.
4. **Result** - A failed job shows the server's message. A finished job shows the output file, downloads it, or imports it into LocalAI under a name you choose, then links to a chat with it and to Models.
Earlier jobs are listed below the form, with **Reuse** and **Delete** (after a confirmation).
A user needs the quantization permission. Without it, the page says the account cannot quantize.
## Architecture
+32 -1
View File
@@ -9,7 +9,38 @@ LocalAI provides a web-based interface for managing application settings at runt
## Accessing Runtime Settings
Navigate to the **Settings** page from the management interface at `http://localhost:8080/manage`. The settings page provides a comprehensive interface for configuring various aspects of LocalAI.
Open **Operate → Settings** in the web UI (`/app/settings`). The page lists every setting below under eight groups by intent, and keeps your edits in a pending bar until you apply them.
### Groups
The groups reorganise the fields the page used to show under fifteen sections. The mapping:
| Group | Settings it holds | Used to be under |
|---|---|---|
| Memory and models | Watchdog (idle and busy checks, timeouts, interval), eviction (force when busy, largest first, retries, retry interval), free memory automatically and its threshold, models kept loaded, GPU memory budget | Watchdog, Memory Reclaimer, Backend Management, Performance |
| Speed and defaults | Default threads, default context size, artifact download concurrency, F16 | Performance |
| Backends and galleries | Automatic backend upgrades, development backends, gallery loading on boot, the persistent VRAM cache, the model and backend gallery lists | Backend Management, Galleries |
| Access and security | CORS and allowed origins, CSRF protection, shared API keys | API & CORS, API Keys |
| Debugging and traces | Verbose debug logging, API traces and their limits, backend logging | Performance, Tracing |
| Agents and responses | Agent job history, the agent pool, the LocalAI Assistant, the response store TTL | Agent Jobs, Agent Pool, LocalAI Assistant, Open Responses |
| Swarm and sharing | P2P token, network ID and federated mode, the disk headroom check | P2P Network, Distributed |
| Look and feel | Instance name, tagline, the three logos | Branding |
Search at the top of the page matches a setting's name, description and key, the group, and the old section name, so a search for "Watchdog" still finds the watchdog settings and each result says where it used to be.
### Editing, applying and undoing
Edits are not sent as you type. A bar at the bottom of the page counts them and offers **Discard**, **Show diff** and **Apply**. The diff lists each old and new value, and runs the checks the browser can make: a duration parses the way the server parses it, a GPU memory budget is one the server accepts, a gallery list is valid JSON with a `url` in each entry, and warnings repeat what the server says (a restart is needed, an empty P2P token stops P2P, evicting while busy can interrupt requests, CSRF protection off). A check that would make the server refuse the value stops **Apply**.
**Apply** sends only the settings that changed. After it, **Undo** (for ten seconds) saves the previous values again. This is a new save, not a rollback: anything that changed in between stays changed.
A setting shows **Changed** and its built-in default, with a **Reset** link, only when the default is known (the defaults of the `local-ai run` flags). Settings without a stated default, such as threads and the context size, never show the marker. An environment variable or flag can change the value the server starts with, so the marker compares with the built-in default, not with the startup value.
A row says **Applies now** when the save handler applies the setting immediately (the watchdog and the memory reclaimer restart with the new values; so do the P2P stack and the agent job service when their settings change), and **Needs restart** for the agent pool settings. A setting without either note is saved, and the code does not say when it takes effect.
**History** lists the settings you applied from this browser, newest first, up to the last 50. LocalAI keeps no settings log of its own, so changes made elsewhere do not appear, and a secret (the P2P token, shared API keys, the agent database URL) is listed as changed without its value. **Revert** puts the old value back as an edit.
`POST /api/settings` accepts a partial body and merges it over the saved settings, so the page sends only the keys it changes. In the request, the `csrf` field carries the **disable** flag: the page shows the inverse as "CSRF protection". The galleries are sent as `galleries` and `backend_galleries`, and the shared API keys as `api_keys`, each as a JSON list.
## Available Settings
+74
View File
@@ -0,0 +1,74 @@
+++
disableToc = false
title = "Studio"
weight = 49
url = "/features/studio/"
+++
Studio is the part of the web UI where you make images, video, 3D objects, speech, sound and transformed audio, and find what you made before. It opens on a front page with a prompt box and your own work. Each type still has its own workspace page (Images, Video, 3D, TTS, Sound, Transform, Diarization) where a run happens.
## Make something
Type what you want in the box under **What do you want to make?**, or pick a type first. The box suggests a type from your words ("Sounds like Video") and never switches by itself. The suggestion is a short list of keywords in the browser. No model is called.
| Key | Action |
|---|---|
| `/` | Focus the prompt box |
| `Alt+1` to `Alt+7` | Pick a type, in the order of the chips |
| `Ctrl+Enter` (`Cmd+Enter`) | Generate |
| `Alt+Enter` | Take the suggested type |
| `Esc` | Close the details in a lineage view |
**Generate** opens the workspace for the chosen type with the prompt, the model, and the options that workspace has already filled in: size and count for Images, size for Video. The run itself happens on the workspace page, as before. The front page does not run anything.
3D, Transform and Diarization start from a file. Use the **Start from** list to pick one of your earlier results as the file, or add a file on the next page.
### A type with no model
A type with no installed model is a dashed chip. Picking it shows one model from the gallery, its download size and memory need where the gallery knows them, the memory free now, and an **Install** button. Nothing installs until you press it, and the words you typed stay in the box while it installs. When the gallery gives no size, the note says so.
## Your work
**Your work** lists the results this browser has a record of, newest first, with filters for each type and counts. Each tile shows a thumbnail (a waveform drawing for audio, one bar per speaker for diarization), the prompt, the model and the age. The star marks a favourite.
Results made from each other stack into one project tile. Open a tile to see the lineage.
### Where the history is kept
The history lives in your browser, not on the server. Each workspace writes its own list when a result is produced: images, video, TTS, sound and audio transform in browser storage (up to 100 entries each), 3D in IndexedDB (up to 20 entries, with the model file), and diarization in browser storage with the file name, the model and the speaker count only. It never keeps the recording or the transcript. A prompt longer than 2000 characters is cut when it is stored. Favourites are a list of up to 500 ids.
Results made before this version have no link to the result they came from, so they appear as single tiles. The files themselves are on the server; if the server has cleaned its output folder, the tile shows the type icon instead of a picture.
**Clear history** removes all of these lists and the favourites from this browser after a confirmation. It does not delete files on the server.
## Lineage
When a workspace is opened from a result (for example **Animate** on a picture), the new result records the id of the result it came from and how: take, animate, to 3D, variation, transform or who spoke. The lineage view draws those links as a board: prompts or files on the left, results next to them, the path through the selected result drawn heavier, and one dashed suggested next step chosen from the installed models.
The dock under the board acts on the selected result:
- **Run as a new take** opens the same workspace with the prompt, model and size filled in. It is disabled when the original input was not kept (3D, diarization).
- **Branch from here** opens a draft: choose the next step, edit the words, and **Open in** the workspace with the result as its starting point.
A step is listed but disabled, with the reason, when the destination page cannot start from that kind of result yet. Today a video cannot start a sound and a 3D object cannot start a clip.
Arrow keys walk the board, `B` branches and `Esc` closes the draft, then the details, then the view.
## The workspace page
All seven workspaces share one layout, under a row of tabs, one per type:
- **Compose card.** Optional sources as dashed chips (a start image, reference images, an end image, avatar audio), or a drop area where the run cannot start without a file (a recording for Diarization, audio for Transform, a picture for 3D). Then the prompt, with starters while it is empty, a model chip, the essential options as chips (size and count, duration and frame rate, voice, mode), and an **Advanced** fold that names what is inside when it is closed. Below that: the memory the model needs and whether it fits, when the gallery knows it, and one button. When the button cannot run, the reason is next to it. With no model for the type, the install note from the front page shows in the card.
- **Run area.** While a request is out, a job card shows the model and options and the time that has passed. The server reports no phase and no percentage on these endpoints, so the bar is indeterminate. A failed run shows what the server said, says the prompt and settings are still there, and offers **Try again**. A finished run shows the result in a viewer for its type (picture grid, video player, waveform player, 3D viewer, spectrograms, or a timeline of speakers) with a toolbar:
- **Favourite**, the same list the front page keeps.
- **Download**.
- **Use in** lists where the result can go next. A step is disabled, with the reason, when the destination cannot start from it, or when no model for it is installed it says so.
- **Re-run with edits** puts the result's values back in the form. The fields you then change are outlined and listed under **Changed from this take**, and the button reads **Run again**. It is disabled when the original input was not kept (3D, diarization).
- **Lineage** opens the lineage view of the front page for this result.
- **Recent results.** A strip of the results of this type from the same history, with an All and a Favourites filter. Click a tile to show it above, and click it again to go back to the latest.
Diarization lists who spoke when as one lane per speaker, the talk time of each speaker, and the segments with their text when the model returned it. **RTTM**, **SRT** (only when there is text) and **JSON** are built in the browser from the result in hand.
## What a workspace accepts from the front page
The workspaces read these query parameters, so a link of your own works too: `prompt`, `model`, `size`, `n` (count), `from` (the id of the result it starts from) and `edge` (how). For example `/app/studio/video?prompt=Slow%20push-in&from=<id>&edge=animate`.
+2 -2
View File
@@ -80,9 +80,9 @@ this control may ignore it.
## Voice Library
Administrators can manage reusable voice-cloning references from **Operate → Voice Library** in the LocalAI WebUI. The library replaces per-model filesystem and YAML setup for supported cloning backends:
Administrators can manage reusable voice-cloning references from **Build → Voices → Speech voices** in the LocalAI WebUI. (Speakers, the first tab of the same page, is a separate feature: voiceprints for recognising who is speaking. See [Voice recognition](/features/voice-recognition/).) The library replaces per-model filesystem and YAML setup for supported cloning backends:
1. Select **Create voice** and upload or record a clear reference clip.
1. Select **Add a speech voice** and upload or record a clear reference clip. A speech voice can also be made while enrolling a speaker, from the enrol sheet.
2. Enter the exact words spoken in the clip. Add more audio/transcript pairs when the personality needs more examples.
3. Confirm that you have permission to clone the voice, then save the profile.
4. Open **Text to Speech**, choose a model marked **Cloning ready**, and select the saved voice.
+6
View File
@@ -22,3 +22,9 @@ Each history remains independently bounded by `tracing_max_items`. When a
history reaches that limit, LocalAI removes its oldest records from memory and
disk. The existing clear actions on the Traces page remove both the in-memory
history and its persisted records.
Opening an API request on the Traces page shows the request as its own page, at
`/app/traces/<id>`: the status, the error LocalAI recorded, a timeline with the
backend operations that ran while the request was open, and the request and
response bodies. Bodies stay closed until revealed, and request headers are not
listed. See [Traffic]({{% relref "operations/traffic" %}}).
@@ -557,6 +557,21 @@ convention the Whisper / Voxtral transcription backends use.
Pass `threshold` explicitly when switching recognizers - the per-model
default only applies when omitted.
## The WebUI page
**Build → Voices** has three tabs. **Speakers** is this feature. **Speech voices** is the [Voice Library](/features/text-to-audio/#voice-library) for text-to-speech, a different store with its own recordings and transcripts. **From a recording** links to the diarization workspace, where a speaker can be named and remembered.
On Speakers:
- **Who is this** matches a clip against the people you enrolled (`POST /v1/voice/identify`). The answer is a sentence with the real distance and the cut-off, a word for how far inside the cut-off it sits (strong, likely, close call), and a scale with the cut-off drawn on it. The cut-off slider re-reads the answer in the browser; it does not call the server again. The page sends the cut-off in the request, 0.25 by default.
- **Same person?** compares two clips (`POST /v1/voice/verify`) and uses the threshold the model returns. The word is not a probability, and the page says so.
- **Enrol a speaker** opens a sheet: a recording, a name, optional labels, and a permission tick. An administrator can also keep the recording as a speech voice in the same step. Keeping a copy of the recording in the browser is off unless you tick it.
- **Known speakers** is a list kept in the browser. The server has no list call, so the page cannot check it by itself. After a search it marks a saved person the server did not return as "not on the server" (only when the search asked for more people than it got back, so a short answer is never read as proof), and lists anyone the server returned that this browser has no record of. A server restart empties the server's index; enrol again, or use **Re-enrol from saved copy** if you kept one.
- Removing a person waits ten seconds behind an **Undo** toast. Nothing is sent to the server until the time ends; Undo cancels it.
- If no speaker-recognition model is installed, the tool is replaced by a note that says so, with models from the gallery to install. Users without the Voice recognition permission see a page that says it is off for their account.
The clip you test with is sent to the model on the server and is not kept. Analyze (age, gender and emotion guesses, all off by default) and the raw embedding are under **More tools**.
## Related features
- [Face Recognition](/features/face-recognition/) - the image analog;
+1 -1
View File
@@ -50,7 +50,7 @@ Setting a per-agent model always overrides the default.
## Send a message
Open the new agent from the Agents page and type a message in its chat box, for example `Hello, what can you do?`. The agent replies in the chat panel within a few seconds. When the agent decides to use the action you configured, you will see the tool call and its result appear inline before the final answer, streamed live as the agent works.
Open the new agent from the Agents page, type a task in its task box, for example `Hello, what can you do?`, and press **Start run**. The run page shows the agent working and then its answer, within a few seconds. When the agent decides to use the action you configured, the step count and the tool in use show live, and the report lists what the tool returned.
That is a complete agent: a model, a system prompt, and one tool, all running inside your LocalAI process.
+42 -23
View File
@@ -19,7 +19,7 @@ This section covers everything you need to know about installing and configuring
The Model Gallery is the simplest way to install models. It provides pre-configured models ready to use.
GPU recommendations require a memory estimate within 95% of the detected model memory budget at a 4096-token context. If none of the sampled candidates fit, the recommendation section is hidden. You can still browse the gallery and check individual models at your intended context size. The Home page also omits static GPU suggestions when no fitting recommendation is available.
GPU recommendations require a memory estimate within 95% of the detected model memory budget at a 4096-token context. If none of the sampled candidates fit, the recommendation section is hidden. You can still browse the gallery and check individual models at your intended context size. On the Models page the recommendations appear as a "Best for this machine" list in the pane beside the table while no model is selected. Once you have installed a model, the list shows only the best fit and offers the others behind "more that fit". The Home page also omits static GPU suggestions when no fitting recommendation is available.
### Via WebUI
@@ -35,10 +35,13 @@ For more details, refer to the [Gallery Documentation]({{% relref "features/mode
The same Models page owns the complete lifecycle. Switch to **Installed** to
search local configurations, filter them by running, idle, disabled, pinned,
or distributed state, and open a model's runtime controls. Load, stop, edit,
pin, disable, inspect backend logs, and remove actions stay with the selected
model. The current view, search, filter, and selection are stored in the URL so
links and browser history preserve your place.
or distributed state, and open a model's runtime controls. Load and stop are on
each row; edit, pin, disable, backend logs, and remove are in the row menu and in
the details of the selected model. The current view, search, filter, and
selection are stored in the URL so links and browser history preserve your place.
The disk strip in the page header shows the free space on the models disk and
opens a review of what can be removed to free more (see
[Model gallery]({{% relref "features/model-gallery" %}})).
### Via CLI
@@ -73,15 +76,25 @@ Visit [models.localai.io](https://models.localai.io) to browse all available mod
## Method 1.5: Import Models via WebUI
The WebUI import page takes either a source to resolve or a configuration to
write. Both live on the same page, behind the two tabs in its header.
The WebUI import page (**Build → Import**) takes either a source to resolve or a
configuration to write. Both live on the same page, behind the two tabs in its
header. From a source, it is a short guided flow: source, review, import, done.
### From a source
1. Open the LocalAI WebUI at `http://localhost:8080`
2. Click "Import Model"
3. Paste the source into the **Source** field (e.g. `https://huggingface.co/Qwen/Qwen3-VL-8B-Instruct-GGUF`)
4. Press Enter, or click **Import**
2. Open **Build**, then **Import**
3. Paste the source into the **Source** field (e.g. `https://huggingface.co/Qwen/Qwen3-VL-8B-Instruct-GGUF`), or start from one of the examples
4. Review what the page says, then press Enter or click **Import**
While you type, the page reads the spelling of the source (Hugging Face
repository, direct URL, configuration file, OCI image, Ollama model, or a path
on the host) and shows the request the form will send. It also runs the checks
that can be made before the import starts: whether the name is already in use,
whether the chosen backend is installed, and how much disk and memory are free.
The server returns no preview of the model configuration, the size or the memory
a model needs before the import starts, so the page does not show them. They
appear once the import starts, next to the free memory and disk.
The **What you can paste** panel beside the field lists every accepted scheme:
`huggingface://`, `hf://`, a full Hugging Face URL, any direct `https://` URL,
@@ -105,8 +118,10 @@ vision-language models and `mlx-audio` for text-to-speech models; other MLX
repositories use `mlx`. An explicit backend selection in the import form always
overrides this automatic routing.
Once the import starts, the page reports the current phase, the bytes
transferred and a progress bar until the model is ready.
Once the import starts, the page reports the download size and the memory the
model needs against what is free, then the current phase, the bytes transferred
and a progress bar. When the model is ready, the page names it and links to a
chat with it and to Models.
### Writing YAML
@@ -403,19 +418,23 @@ local-ai models list
### Disk Usage
The **Installed** tab shows how much disk the installed configurations use: a
summary strip with the total on disk and the portion shared across models, and
a per-model size in the selected model's detail panel. A file referenced by
more than one configuration is counted once in the total; the per-model size
notes the shared portion, since deleting that model frees only its exclusive
files. A reference whose file is not on disk — a download that never finished,
or a file removed by hand — is shown as a warning on the model.
In the WebUI, the **Installed** tab of **Models** shows how much disk each
installed configuration uses. The **Size** column holds the size on disk, and
for a model that shares files with others it adds the shared part ("1.2 GB
shared"). A file referenced by more than one configuration is counted once in
the total under the table, which also gives the shared total. The inspector
beside the table names the models a model shares files with, and a model page
lists its files under **Usage and history**, with each file's size, the other
models that use it, and whether it is missing. Removing a model frees only its
exclusive files; the files it shares stay while another configuration uses
them. A reference whose file is not on disk (a download that never finished, or
a file removed by hand) is marked on the model and in the cleanup review.
While no model is selected, the detail pane lists every referenced file with
its size, which models reference it, and whether it is shared or missing;
clicking a model name selects it.
Only an admin can read the report. For anyone else, or if the read fails, the
Size column shows the size of the files the gallery lists, and "size unknown"
for a model the gallery does not list.
The report behind both views is available directly (admin only):
The report behind these views is available directly (admin only):
```bash
curl http://localhost:8080/api/models/storage
+1
View File
@@ -10,6 +10,7 @@ This section collects the operator-facing concerns for running LocalAI in produc
## Pages
- [Traffic]({{% relref "operations/traffic" %}}) - usage, failed requests, models, GPU and host, traces and Prometheus in one tab.
- [Middleware: PII filtering and intelligent routing]({{% relref "operations/middleware" %}}) - per-model PII redaction and policy-based request routing.
- [Cloud passthrough proxy]({{% relref "operations/cloud-proxy" %}}) - forward requests to OpenAI, Anthropic, or any compatible provider.
- [MITM proxy for Claude Code / Codex CLI]({{% relref "operations/mitm-proxy" %}}) - redact PII from cloud-AI traffic without LocalAI holding API keys.
+25 -9
View File
@@ -58,7 +58,7 @@ each shown only when it has something in it.
### In progress
One card per running or queued operation, each naming what is being done and to
One row per running or queued operation, each naming what is being done and to
what. A percentage and progress bar appear only once the operation is running
and has reported progress, so a queued operation and a removal show none.
Artifact-backed gallery models also report the phase, downloaded and total
@@ -88,7 +88,8 @@ worker finishes.
### Needs attention
Model and backend operations that failed and have not been acknowledged yet,
each card carrying the error returned by the installer. Cluster staging never
each row carrying the error returned by the installer, with **Retry** and
**Dismiss**. Cluster staging never
appears here: a staging job reports no error to the page, so a staging failure
has to be read from the logs.
@@ -96,7 +97,11 @@ has to be read from the logs.
What finished, newest first, one row each: the name, what happened
(`installed in 1m 12s`, `removed`, `cancelled`, or `failed:` with the error),
the time of day it finished, and a link into Models or Backends. Model and
the time of day it finished, and a link into Models or Backends. A cancelled
install also has **Start again**, which installs the same target again: a
download that was paused continues where it stopped, and one that was cancelled
starts over. The record does not say which of the two it was, so the button
says what it does in both cases. A cancelled removal has no such button. Model and
backend installs and removals are recorded; cluster staging is not, so a
staging run leaves nothing behind here once it finishes.
@@ -110,10 +115,13 @@ finished fan-out install is filed under Models or Backends, not under Cluster.
## Cancelling, retrying and dismissing
These are on the operation cards. **Cancel** and **Retry** are labelled
buttons; dismissing is the **X** at the end of a failed card. The strip has no
cancel button; the page is the only place work is stopped or restarted.
These are on the operation rows. **Pause**, **Cancel**, **Retry** and
**Dismiss** are labelled buttons. The strip has no cancel button; the page is
the only place work is stopped or restarted. The Backends page also shows the
progress of a backend install in the backend's own row, with its **Cancel**.
- **Pause** stops a running download and keeps the bytes already fetched, so
installing the same model or backend again continues from there.
- **Cancel** is offered while an operation is queued, whatever it is, and while
an install is running. It is not offered once a removal has started: a
removal in progress cannot be interrupted, so the window to call one off is
@@ -121,13 +129,21 @@ cancel button; the page is the only place work is stopped or restarted.
anything is touched. For artifact-backed gallery models, cancelling an active
download leaves its partial files in place so a later install resumes rather
than starting over. A cancelled operation leaves the live sections
immediately and is not held on the strip the way a completed one is; it
appears in the record as `cancelled`.
and is not held on the strip the way a completed one is; it appears in the
record as `cancelled`.
**Cancel waits for 8 seconds before it is sent.** The row says "Cancelling
unless you undo" and a toast offers **Undo**. The server cannot take a cancel
back, so the wait in the browser is the whole undo: while it runs nothing has
been stopped and the download carries on. Closing the toast cancels at once,
asking to cancel a second download ends the first one's wait, and leaving the
page sends a cancel that is still waiting. A job that finishes during the wait
is left alone.
- **Retry** is offered on a failed model or backend install. It acknowledges
the failure, which moves it into the record, and installs the same target
again. It is not offered on a failed removal, which is not restarted by
reinstalling.
- **Dismiss**, the **X** on a failed card, acknowledges the failure without
- **Dismiss** acknowledges the failure without
retrying. The operation moves into the record with a `failed` outcome; it is
not deleted. This is why the same failure can be found either under **Needs
attention** or in the **Record**, depending on whether it has been
+3 -1
View File
@@ -18,7 +18,9 @@ a single client-facing model name fans out across multiple downstream
targets.
Both are inspected and configured from the same admin page
(`/app/middleware`), backed by the same REST surface (`/api/middleware/*`,
(`/app/middleware`), which draws the order a request passes through as five
steps (Proxy, Admission, Filtering, Routing, Model) and shows only the rules of
the step you select, backed by the same REST surface (`/api/middleware/*`,
`/api/pii/*`, `/api/router/*`) and the same MCP tools.
## Request lifecycle
+95 -51
View File
@@ -6,23 +6,66 @@ weight = 1
`/app/operate` is the front door to the Operate console. It answers one
question — is anything wrong — without you having to open four other pages.
## Needs attention
The page opens with one sentence: **Everything is running.**, **2 things need
you.** or, when nothing needs a decision but a row has something to read,
**Running, with 1 thing to look at.** On a new installation, with no backend, no
model and nothing loaded, it says **Nothing is running yet.** and offers
**Install a backend** and **Browse models**.
The block the page exists for. It lists only things that want a decision:
Under the sentence are four rows. A row with a problem opens by itself and holds
the button that deals with it. A row with nothing to say stays one line. Click a
row to open or close it.
- a backend with an update available
- an operation that failed
- a node reporting unhealthy
## Needs you
**When nothing needs attention it says so in one line and renders nothing
else.** There is no green panel: a status page that shouts when everything is
fine teaches you to stop reading it.
Only things that want a decision:
## Headline totals
- a backend with an update available, with an **Update** button
- an operation that failed, with **Retry** and **Dismiss**
- a node that is not answering, with a link to the nodes
Requests, failed requests and p95 latency over the last 24 hours, each with a
sparkline of the trend. These come from `GET /api/traces/summary`, which counts
the trace buffer server-side:
**Update** starts the same update as the Backends page. **Retry** installs the
failed model or backend again, after moving the failure into the Activity
record, and **Dismiss** moves it there without retrying. When nothing needs you
the row says so in one line. There is no green panel: a status page that shouts
when everything is fine teaches you to stop reading it.
The attention count on the **Status** tab is the number of items in this row.
## Capacity
One line says how full GPU memory is (system memory on a host without a GPU),
and how much room the models disk has left. On a cluster it adds the memory of
the nodes that are answering, and the row lists each node's memory when opened.
The row opens by itself when memory is 90% full or more, or when the models disk
has less than a tenth free or less than 20 GB. It then lists the models loaded
on this machine, each with an **Unload** button. Unloading asks first, stops the
backend process of that model and frees its memory. The next request that uses
the model loads it again.
Under the rows, a bar shows the memory in use now. **LocalAI keeps no memory
history**, so the chart under the bar is drawn from readings the page took
itself, one with each Operate summary poll (every 15 seconds), while Operate is
open. It starts with the second reading, keeps at most 240 readings (about an
hour), and is labelled "Since you opened Operate". The axis starts at zero, the
capacity is a labelled line, and a data table sits behind the chart. Leaving
Operate empties it. Temperature and power are not shown because the resources
endpoint does not report them.
## Running now
Operations in progress, and the models loaded on this machine. On a single-node
install the row lists the heaviest five models: backend, resident memory, CPU
share and uptime, with **View logs** and **Stop model…** in the row menu. **Open
this machine** leads to the full list. With distributed mode on, models run on
workers rather than on the controller, so the row links to the Swarm page
instead.
## Recent failures
Requests and failed requests over the last 24 hours, and p95 latency, counted
server-side by `GET /api/traces/summary`:
```bash
curl http://localhost:8080/api/traces/summary?hours=24 \
@@ -42,39 +85,30 @@ curl http://localhost:8080/api/traces/summary?hours=24 \
`hours` defaults to 24 and is capped at 168. Only 5xx responses and transport
errors count as failures — a 4xx is the caller getting it wrong, not the
installation being unhealthy. `p95_ms` is a nearest-rank percentile, not the
slowest request.
slowest request. The endpoint exists so a dashboard wanting three numbers does
not fetch the whole trace list to count it. An installation that has recorded
nothing says so rather than showing zeroes dressed as telemetry. The row opens
when at least one request failed, and links to Traces and Usage.
The endpoint exists so a dashboard wanting three numbers does not fetch the
whole trace list to count it. An installation that has served nothing yet says
so rather than showing three zeroes dressed as telemetry.
## Host capacity
The overview also shows the host's current RAM or GPU capacity, utilization,
and model storage. Loading, unavailable, and empty states are explicit. This
uses the same 15-second Operate summary poll as the rail and attention data, so
opening the overview does not start a second resource poller.
## Running now
On a single-node install the overview lists the models loaded on this machine,
heaviest first, up to five. Each row shows the backend, resident memory, CPU
share and uptime, with **View logs** and **Stop model…** in the row menu.
**Open this machine** leads to the full list.
With distributed mode on, models run on workers rather than on the controller,
so this section links to **Operate → Nodes → Running models** instead.
The page does not report models that failed to load, because LocalAI records no
load failures to read.
## This machine
On a single-node install, **Operate → This machine** (`/app/nodes`) shows the
host and everything loaded on it:
- **Capacity gauges** for VRAM, RAM, CPU and the models disk, the same gauges
the Nodes page draws for a cluster. A host without a GPU says so rather than
showing an empty VRAM gauge.
- **A memory bar** splitting host RAM by running model, so you can see which
model is holding memory.
- **GPU memory** as one bar, with what is free and a note that the next model
must fit in that, or a loaded model has to be stopped first. With several GPUs
each one gets its own line. A host without a GPU says so instead of drawing
an empty bar.
- **Models in host memory**: one bar split by running model, so you can see
which model is holding memory. LocalAI reports the resident memory of each
backend process, not the GPU memory of each model, so the split is of host
RAM and no model is shown with a GPU size.
- **The machine**: VRAM, RAM, CPU and the models disk, each with a bar, a
percentage and what is left. A reading the host does not report is shown as
"No data", never as zero.
- **Running models**: search, sort by memory, CPU or uptime, open a model's
logs, or stop it. Stopping asks for confirmation; the model loads again on
its next request.
@@ -84,9 +118,10 @@ per-model readings come from the `process` block of
[`GET /system`]({{% relref "reference/system-info" %}}); the host CPU and disk
readings come from the `cpu` and `disk` fields of `GET /api/resources`.
**Add machines** reveals the command to start LocalAI in distributed mode. Once
distributed mode is on, the same route becomes the Nodes page and the rail
entry moves to the Cluster group.
**Add a machine** reveals the command to start LocalAI in distributed mode, and
a link to the steps for adding a worker. Once distributed mode is on, the same
route becomes the Nodes page and the **This machine** tab gives way to the
**Swarm** tab.
Models and backends no longer live under a nested Host page. Use **Models →
Installed** for model runtime and configuration actions, and **Operate →
@@ -97,15 +132,24 @@ Old `/app/manage` bookmarks remain supported. They redirect with replace
semantics to the matching Installed Models or Installed Backends view while
preserving legacy search, filter, selection, variant, and development flags.
## The rail
## The tab bar
The Operate rail groups its destinations under four headings —
Runtime, Cluster, Observability and Administration — and shows a live value
beside several of them: pending backend updates, running operations, healthy
node count, request volume and error count. Host capacity lives on the overview
instead of appearing as a separate destination.
Operate has one row of tabs above the page: **Status** (this overview),
**This machine**, **Swarm** (only with distributed mode on), **Runtime**,
**Traffic** and **Settings**. A tab that holds several pages shows a second row
of links under the bar: Runtime holds Backends, Activity, Logs and (on a single
install) Failover; Traffic holds Usage, Traces and Middleware; Settings holds
Settings and Users (with authentication on); Swarm holds Nodes, Placement rules,
Failover and P2P. Adding a node opens from the Nodes page. Every page keeps its
own URL, and a page such as a node detail keeps its tab highlighted. On a phone
the bar scrolls sideways. The **API** link at the end of the bar opens the
API documentation.
Those values are **orientation, not an alarm**. The rail only exists on Operate
routes and can be collapsed, so anything urgent also appears in Needs attention
and on the operations badge attached to the sidebar entry, which is always
visible.
Several tabs carry a live value: pending backend updates and running operations
on Runtime, the healthy node count on Swarm, running models on This machine, the
attention count on Status and the error count on Traffic. Host capacity lives on
the overview instead of appearing as a separate destination.
Those values are **orientation, not an alarm**. The bar only exists on Operate
routes, so anything urgent also appears in the Needs you row and on the
operations badge attached to the sidebar entry, which is always visible.
+98
View File
@@ -0,0 +1,98 @@
+++
title = "Traffic"
weight = 3
+++
The **Traffic** tab of the Operate console answers three questions: how much is
the server used, what failed, and how are the machine and the models doing. It
opens on an overview and has a second row of links: Overview, Usage, Models,
GPU and host, Traces, Middleware and Prometheus. One time window (24 hours, 7
days, 30 days or all time) is shared by the Overview, Usage and Models pages.
Every figure comes from a record LocalAI really keeps, and each page says which.
Where LocalAI keeps no record, the page leaves the figure out and does not draw
a zero.
| Record | What it holds | Pages that read it |
|---|---|---|
| Usage ledger | Requests and tokens, by model, user and API key, in hourly, daily or monthly buckets | Overview, Usage, Models |
| Trace buffer | The most recent API requests, with status, error and duration. Off until tracing is on | Overview (failed requests, p95), Traces |
| Backend-operation buffer | Loads and runs of each model, with their errors | Models, a trace |
| Resources reading | Memory per GPU, system memory, CPU share, models disk | GPU and host |
## Overview
Five figures in a row: requests, failed requests, latency p95, tokens in and
tokens out. Under them, three charts: requests per bucket, failed requests (the
trace buffer, in twelve columns) and tokens in and out per bucket. Each chart has
a data table behind it, and you can read a value with the arrow keys. A table of
the five busiest models closes the page.
- **Failed** counts a transport error or a 5xx answer. A 4xx answer is the caller
being refused, and LocalAI does not count it as a failure.
- **p95** is the 95th percentile of request duration in the trace buffer. LocalAI
does not compute a p50 or a p99, so there are none.
- With tracing off, the failed and p95 figures say so and link to the setting.
The ledger figures do not depend on tracing.
- The trace summary reports at most 7 days. For the 30-day and all-time windows,
failed requests and latency cover the last 7 days, and the page says so.
## Usage
The ledger as a table, grouped by model, by user (admins) or by API key (when
authentication is on). Rows sort by any column, can be searched, and can be
filtered by model. A row opens in place on its own chart. A user who is not an
admin sees only their own numbers and their quotas, and the page tells them
whether the current pace stays inside each quota.
**Export CSV** and **Export JSON** save the rows the table holds. The export runs
in the browser. The ledger does not record status, endpoint or node, so those are
not groups. Estimated cost is optional: you type a price per million tokens, it
stays in your browser, and LocalAI has no price of its own.
## Models
One row per model: requests and tokens from the ledger, failed operations and the
mean operation time from the backend-operation buffer, and the memory the backend
process holds now. That is the resident host memory of the process, when the
server reports it. GPU memory per model is not reported, and neither is a
per-model latency percentile or a memory history.
## GPU and host
The current reading, refreshed every 5 seconds: memory per GPU, system memory,
the CPU share and the models disk. LocalAI does not report GPU utilisation or
temperature. Two charts show readings the page took itself, **since you opened
this page**: the memory pool and the CPU share. They are kept in the browser, at
most 240 readings, and leaving the page empties them. On a cluster the page also
lists every node with its memory and CPU share.
## Traces
The recent API requests and backend operations, with the tracing settings above
the list. Filter by failed or slow requests, search, sort, export and clear.
Opening an API request shows its page: the status and the error LocalAI recorded,
a timeline of the request with the backend operations that ran while it was open,
and the request and response bodies. Bodies stay closed until you press
**Reveal**, because they can hold prompts and personal data. Request headers are
never listed. The two buffers share no request id, so operations are matched on
time and model, and the page says so. A trace leaves the buffer when newer
requests push it out; its page then says it is no longer there.
With tracing off, the page explains what is lost and offers **Turn on tracing**.
You can also start LocalAI with `LOCALAI_ENABLE_TRACING=true`.
## Middleware
The order a request passes through: Proxy, Admission, Filtering, Routing, Model.
The server fixes that order. Selecting a step shows only its rules. See
[Middleware]({{% relref "operations/middleware" %}}) for what each step does.
## Prometheus
`GET /metrics` is admin only and returns the Prometheus text format. The page
shows the URL and a scrape configuration with a copy button, checks the endpoint
against the running server, and lists the metrics LocalAI can export. Only
`api_call` is always present: a histogram of request duration in seconds by HTTP
method and route. The others appear while the feature behind them runs. Metrics
are off when LocalAI runs with `LOCALAI_DISABLE_METRICS_ENDPOINT=true`.