mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-10 07:47:29 -04:00
feat(ui): redesign the web UI around a shared kit and a calm palette (#12526)
* build(ui): vendor the shared UI kit snapshot at 0.2.0 The restyle needs the kit's tokens, motion layer and component classes. Take a pinned snapshot instead of depending on the kit at build time, and keep a lock file with the version and per-file checksums so a later update shows exactly what changed. The product theme stays outside the vendored directory. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the ink and teal theme and bridge the old variables Define the product colours as the shared UI kit's roles, for light and dark, in theme-localai.css. The kit's contrast check passes on every pair. theme.css keeps the existing --color-* and --shadow-* names but now points each at a role, so App.css and the pages get the new palette without edits. Radii move to the kit scale. index.html now sets data-theme before first paint with the same rule as ThemeContext (stored choice, otherwise dark), because the contract layout of the theme file no longer defaults to dark by itself. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): restyle the shared chrome with the UI kit grammar Adjust the shared classes so every page picks up the same interaction language without per-page edits: - Sidebar sits on the canvas and the current row lifts onto a card. Section labels are tracked uppercase, the badge is a soft pill, and the phone drawer leaves the tab order when closed. - Buttons are flat: hover swaps the surface, press scales to .97, focus is a 2px ring with a 2px offset, danger is a tinted wash. - Inputs use the card surface and the control edge; switches, tabs, filter chips, badges and cards follow the same rules. Cards no longer lift on hover; only linked or button cards react. - Menus and popovers scale in from the trigger corner with 40px items. Dialogs get a veil fade and a spring settle. Toasts become pills at the bottom centre. - The page transition is a 250 ms fade with a 6px rise. It fills backwards so a finished animation no longer leaves a transform that confined dialog veils to the main column. The focus-ring test now checks the outline instead of a box shadow, and new specs cover the theme roles, the first-paint theme and the sidebar lift. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): move leftover hard-coded colours onto the theme roles The YAML editor restated the old blue palette in JavaScript, and a few pages kept literal blues, indigo and violet tints, or fallbacks that only applied because a variable was never defined. Point them at the theme variables so they follow light and dark and the new palette. The status badges in the account pages built their tint by appending "22" to a variable, which is not valid once the variable is defined, so they had no background. Use the wash roles instead. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): raise the type scale and control size toward the kit Body and list text moves to 15px and the rest of the scale follows the kit's 12/13/15/17/21/32/44 steps. Page titles, section headings and stat values are bold with tighter tracking; titles are 32px. Buttons, inputs, selects, tabs and nav rows are 40px high with the 12px radius, compact controls 32px. Tabs become a segmented control. The sidebar widens to 240px (64px collapsed) and nav rows get more room. Identifiers and counts in the split views use the mono face, and the stat grid becomes separate inset tiles. The Geist stack stays: it is bundled, and the thin look came from the size, weight and negative tracking, not the face. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): separate cards, panes and floating surfaces from the canvas Cards, the Models and Installed split panes, the composers and the confirm dialog use a stronger card edge, the rest shadow and the 20px radius, so they read as layers in dark as well as light. Menus and popovers move to a float surface (the hover tone in dark) with the float shadow. The selected rail row gets an accent wash and a 3px accent edge. The send buttons are a clear accent when there is something to send and a quiet inset when not; the Home button carries data-empty for that, since submitting an empty box does nothing. The assistant card becomes an accent wash with a square icon. New surfaces spec checks the pane edge, the selected row, both send buttons and the popover in both themes. The voice library empty-state spec now waits for the layout to settle before comparing two boxes. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): tidy the sidebar header and mark the current row with a dot The header gives the configured horizontal logo a fixed width and centres it in a 72px band, lined up with the nav icons. The collapsed rail shows the configured icon logo centred, and its nav rows become 40px tiles centred in the 64px rail. The current row gets the kit's accent dot, hidden in the rail. The theme, language and account controls stay in the sidebar footer: the app has no global search or command palette to put in a top bar, so a bar would only hold controls that already have a place. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): centre the avatar in the collapsed sidebar rail The collapsed avatar link was set to "flex: 0", which gives it a zero flex basis; with min-width: 0 the link shrank to its padding and the icon overflowed from the link's left edge, about 14px right of the icon column. Use "flex: 0 0 auto" in the collapsed and tablet rail. The footer controls now share the nav icon column in the expanded sidebar too (6px footer padding, 40px control boxes), and the tablet rail gets the same footer padding and hidden language code as the collapsed one. New spec measures the centre x of the nav icons, mark, avatar, language, theme and collapse icons in the collapsed, expanded and tablet states, in both themes, and asserts they agree within 1px. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): size the console and settings rails and stack settings on phones The Operate console rail, the Settings section rail and the account tab bar still used 13px text and the old underline tabs. They now use the 15px nav size, 40px rows and the segmented tab control. Form row labels are 15px with 13px hints. On a phone the Settings section rail sat beside the form and squeezed every row into a few characters. Below 720px the rail stacks above the content as a scrolling row and form rows wrap their control below the label. The save button no longer carries the icon font class, which drew a missing glyph before its label. The language menu is wide enough to keep Bahasa Indonesia on one line. Assisted-by: Claude Code:claude-sonnet-5-5 * build(ui): update the vendored UI kit snapshot to 0.3.0 Take the 0.3.0 snapshot: the sprite now carries the full outline icon set, and the new icons/fa-map.json maps Font Awesome names to icon ids. The map lets the app move off Font Awesome in the following commits. The lock file is regenerated with the new checksums. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add an Icon component backed by the kit sprite Icon draws an inline svg that points into the kit's outline sprite. The sprite is inlined into the page once, so the references resolve under any base path and in the embedded build without a request. Icons size with the font (1em), take currentColor, hide from assistive tech unless given a title, and spin on request. An unknown id draws a neutral circle. FaIcon and iconFromFa resolve Font Awesome names through the kit's map, for names that arrive at run time. iconHtml does the same for markup built as a string. The GitHub and Apple marks are small local glyphs, as the kit ships no brand marks. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): draw shared components and helpers with Icon Replace the Font Awesome elements in the shared components and in the utility modules with the Icon component. Lookup tables now hold kit icon ids instead of class strings. Code-block copy buttons and artifact cards, which build HTML strings, use iconHtml and a sanitizer-safe slot. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): draw model, backend and account pages with Icon Replace the Font Awesome elements on the home, models, backends, import, settings, login, account and users pages with the Icon component. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): draw chat, studio and recognition pages with Icon Replace the Font Awesome elements on the chat, media generation, talk and face and voice pages with the Icon component. The talk status table keeps its spin and pulse states as Icon props. The connected and error states now use a dotted circle and an alert circle, so they differ from the idle ring by shape as well as by colour. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): draw agent, node and operate pages with Icon Replace the Font Awesome elements on the agents, skills, collections, jobs, fine-tune, quantize, nodes, swarm, usage, traces and activity pages with the Icon component. Two class strings on layout elements held leftover button and icon classes from an earlier merge; they are cleaned up so the elements keep only their own classes. Assisted-by: Claude Code:claude-sonnet-5-5 * style(ui): size and align icons for the svg component Icon rules that targeted the font element now target the svg: the descendant "i" selectors in App.css and auth.css become ".lai-icon". The svg is 1.2em with a 2 unit line so it matches the visual size of the old glyphs at the 12 to 16px sizes the app uses, sits on the text baseline, and follows the context font size. Large empty-state marks get a lighter line. Menu icons get a 16px box and the readiness badge icons keep their 20px circle with padding. Add the pulse used by the talk status. Assisted-by: Claude Code:claude-sonnet-5-5 * build(ui): remove Font Awesome No source file references the icon font any more. Drop the package and its stylesheet import. The build no longer ships the solid, regular and brand font files. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): keep focus traps off the svg use references The dialog and drawer focus traps collect focusable elements with a "[href]" selector. An icon's use element carries an href, so it became the first "focusable" element and Tab at the end of the dialog stopped there instead of wrapping to the first button. Match "a[href]" instead. Assisted-by: Claude Code:claude-sonnet-5-5 * style(ui): keep icon sizes overridable and set the line width per svg Give the icon base rule zero specificity so a rule that sizes one icon (nav column, menu box, avatar, language switcher) wins whatever its order in the file. The sprite symbols fix their own line width; the inlined copy drops it so the width set on each svg applies, as the --lai-stroke custom property, and large marks can use a lighter line. Pin the avatar and the language globe to the boxes the sidebar alignment spec expects. Import the map as JSON with an import attribute so Node can load it in the spec. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): select icons by the svg markup and cover the sprite Specs that found icons by their Font Awesome class now select the svg by its data-icon. The dead-icon audit checks that every svg resolves to a sprite symbol and has a size. The class hygiene spec fails on any remaining Font Awesome class. A new spec checks every mapped icon id has a symbol, that the sprite is inlined once, that an icon paints at the root and under a forwarded path prefix, and that Font Awesome names map as documented. Assisted-by: Claude Code:claude-sonnet-5-5 * build(ui): update the vendored UI kit snapshot to 0.4.0 Take the 0.4.0 snapshot: hub tabs with count and attention badges, the six chart series tokens and the grid colour in the theme contract, and sample themes on a calmer palette. The kit headers are renamed and the lock file is regenerated with the new checksums, as for the earlier snapshots. Assisted-by: Claude Code:claude-sonnet-5-5 * style(ui): switch the theme to the calm palette Rewrite the LocalAI theme on the calm palette: a muted teal accent on a near-neutral green-grey canvas, desaturated status colours, no glow and no coloured shadows. The theme fills every role of the shared UI kit's 0.4.0 theme contract for light and dark, including the six chart series and the grid line. The bridge in theme.css keeps the old --color-* names working, adds the dark surface ladder (card, raised, float) and a strong edge, and points the fixed data hues at the chart series. Two values differ from the first sketch. The dark text on the accent fill is #021512 instead of #04201d: it reads 5.58:1 on the fill at rest and 6.4:1 on the hover fill, against 5.08:1 at rest for the lighter value. The light control edge is #6b7d7a. The kit's contrast script passes for all text pairs (4.5:1), control and focus pairs (3:1) and series colours (3:1). Leftovers that no longer fit the palette are fixed: the usage chart takes the six series colours in order, the audio and animation canvases fall back to the new accent, the face box loses its glow, and two gradient fills are now flat. The theme tests expect the new canvas colours. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): replace the console rail with hub tab bars Build and Operate no longer open a second navigation rail beside the page. Each is a hub: one row of the kit's hub tabs above the page, with count and attention badges that scroll sideways on a phone. Every URL and route stays as it was, plus a new /app/build landing page that lists the Build tools with a line each. Build tabs: Overview, Agents, Skills, Memory, Jobs, Fine-Tune, Quantize, Import, Voices (recognition and library) and Faces. Operate tabs: Status, This machine, Swarm (distributed mode only), Runtime (backends, activity, failover), Traffic (usage, traces, middleware) and Settings (settings, users), plus the API link. A tab that holds several pages shows a second row of links, and a sub-page such as a node detail keeps its tab highlighted. The feature and admin gates decide which tabs are drawn, and badges show only values the Operate summary already has. The sidebar lists Build and Operate under a Workspace label next to the Create group. The voice library moves under Build and the model import page gains the Build tab bar. The old rail styles, the rail signals and the console config are removed, and the Operate overview docs describe the tab bar. The specs that drove the rail now drive the tabs, and a new spec covers the tab for each route, gating, badges and the phone layout. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Home as a calm console Home now opens on one command bar: the model chip shows which models are warm, the MCP chip and attach buttons sit beside it, and Send is a solid button with an Enter glyph. Typing "/" opens a grouped, keyboard-driven action list built on the kit command list; every action has a destination in the product. Memory use folds into a one-line strip that opens into the loaded models, with Stop per model and Stop all. It opens by itself while a model is being staged and after a failure, and shows nodes and aggregate memory in a cluster. The list of resident models carries no per-model size because the API reports none. "Jump back in" lists the conversations stored in the browser, one card per day, with j and k to move, Enter to resume and delete with an undo toast. First run keeps the install steps and the recommended models. The assistant prompt is a dismissible line, the library links are one quiet row and the API section is collapsed. Chat accepts an empty new-chat hand-off for /new. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Home console Update the Home specs for the new structure and add specs for the slash menu, the model chip, the memory strip (expand, stop, staging, failure, cluster), the resume list (grouping, j/k, Enter, delete with undo), first run, the send hand-off, a non-admin user and the phone layout. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the fit, disk and cleanup helpers for the models page Pure functions and hooks that the rebuilt Models page reads, with node tests for the rules. modelLedger turns an estimate and the memory budget into one of three verdicts (fits, spills to CPU, over) with the headroom in bytes, and reads the models disk from the resources reading. The disk counts as low under 10 percent or under 20 GB free, and is absent when the server reports none or runs as a cluster controller. cleanupPlan ranks installed models from what the API reports: loaded, pinned, or named by an agent, a task, a failover chain or an alias keeps a model protected; another installed build of the same gallery model is a duplicate; disabled models rank above idle ones. The API records no last use or use count, so none is used. When a lookup fails, nothing is called safe. useModelRemoval holds a removal in the browser for an undo window and sends the existing delete call only when the window ends. Leaving the page drops the batch without deleting anything. The undo toast takes optional labels so other pages can reuse it. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Models as a ledger with a disk strip and cleanup review Explore is one dense table. Each row carries the size, a solid memory bar and the headroom in words ("3.7 free", "+1.5 on CPU", "0.9 over"), worked out from the estimate at the chosen context length. Capability chips show the server's count for each facet, search keeps its meaning and "/" jumps to it, and a density switch (also "d") picks comfortable or compact rows. Selection is a surface step and a check, never a rail. Arrow keys move, Enter installs and Esc closes the inspector, which keeps the fit summary, VRAM by context chart, variants, files, links, tags and licence. A failed install shows its error in the row with a Retry that dismisses the old failure first. A failed or empty listing says which it is, and a host with no GPU is measured against memory and says so. Installed uses the same table with state filters that carry counts, a state per row, Load or Stop on the row, the row menu and the sort by size. Sizes come from the files the gallery lists, so a model it does not know shows a dash. A strip in the header shows the free space on the models disk. It turns amber under 10 percent or under 20 GB free, hides when the server reports no disk or runs as a cluster controller, and opens the cleanup review. Explore says how much an install leaves free. The review ranks installed models as Safe to remove, Probably safe and Your call from real facts only, lists protected models with the reason, and says plainly that usage history is not recorded. A sticky bar shows what a choice frees. Confirming runs a dry run that checks again and lists what will go. Removal waits 30 seconds with an undo; nothing is deleted before that, and leaving the page deletes nothing. The old rail, filter band and popover styles are removed. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Models ledger, Installed table and cleanup review Update the Models, lifecycle, cluster fit, height, search focus and surfaces specs for the table and inspector, keeping what each one checks. New specs, on a shared 41-model gallery stub with three machine profiles: the fit bar and headroom words for a 24 GB card, an 8 GB laptop and a host with no GPU; facet counts, search, "/" and Escape; selection, arrow keys, Enter to install, density; the disk strip when normal, low and hidden; and the states (loading, empty, offline, install failed, phone). Installed covers filters with counts, row actions, the row menu, sizes and sort. The cleanup specs cover grouping, protected models, the honest-data note, the effect bar, the dry run, the undo window, a failed delete, leaving the page, and the phone sheet. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the placement helpers and the estimate hooks The Placement section and the model page need the same few rules, so they sit in plain functions that can be read and tested alone. placement.js holds what gpu_layers, tensor_split and main_gpu mean (unset asks for every layer and the llama.cpp engine trims it, zero is CPU only, 99999999 is the value LocalAI itself writes for all layers), the device list taken from the resources reading, the split by free memory, the part of an estimate that grows with context (read from two lengths, since that term is linear), the fit states with their limit (95 percent of free memory, and the leftover has to fit in system memory too), and a bisection for the largest layer count whose estimate fits. The estimate returns one total and no layer count, so the search runs over 1 to 256 and stops at the first count that no longer changes it. modelWalk.js keeps the order of the list a model page was opened from, in memory and in session storage, for the previous and next buttons. usePlacementEstimate reads /api/models/vram-estimate for a choice, again at twice the context, and with every layer, and keeps readings for the session. useModelPage reads a gallery entry by name, an estimate by context size (from the model's own files when the gallery does not list it), the builds and the loaded models. usePlacementConfig edits the four placement keys of an installed model and saves only what changed. useModelActions is the Load, Stop, disable, pin and remove logic of the Installed table, shared with the model page. MemoryBar is one solid bar with a tick at the capacity of its pool; over capacity it grows past the tick and the tick turns red. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the Placement section to the model editor Run this model on: CPU only (gpu_layers: 0), Auto (the key stays unset) or Custom. Custom takes a number, has an All layers button that writes 99999999, and shows a slider only when the estimate reports the model's layer count, which it does not today. Context size has presets and a number field because the KV cache follows it. With two or more GPUs there is a split (written as percentages, with a button that takes them from the free memory of each card) and a main GPU. A bar per GPU and one for system memory show what other programs use, the model's weights and working memory, and the part that grows with context, with the room left or how far over it is. Under them a verdict in plain words: Fits in GPU, Spills to CPU, Too many layers for the GPU, Runs on CPU only, No GPU found, Not enough memory. It says "slower" and never a multiplier, because the estimate has none. Fit it for me asks the estimate for the largest layer count that fits the free GPU memory and says what it set, with Undo; it is hidden when the estimate is unavailable or the host has no GPU. Loading shows skeletons, an unavailable estimate shows a note with Retry, and a server that schedules onto other machines shows no bars, because its device list is the controller's. The editor shows the section for an installed model, with a link in its section rail. Auto sends null for the key, since a patch only merges, and a null read back opens as Auto. The docs describe the section and what each mode writes. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): open a model on its own page A model has an address, /app/models/<name>, for an installed model and a gallery entry alike. Open it from the arrow at the end of a row, a double click, "o" on the selected row, the inspector's Open details button, or a tap on a phone. The title block holds the main action: Install with a chevron that chooses the build, or Load and Stop with a menu (disable, pin, edit configuration, logs, delete with a confirm). A strip answers whether it fits, what it does and what installing leaves free. Tabs: Overview (about, a memory bar, state, the pages it opens in, and the agents, tasks, chains and aliases that name it); Fit and memory (verdict, context sizes, the bar split into weights and context, and memory by context against the limit, with a data table); Variants and files (builds with size and fit, install any, the files of the chosen build). For an installed model also Usage and history, which says what the API does not record instead of drawing an empty chart, Configuration, which is the Placement section with the file it writes and a link to the full editor, and Logs, the backend log viewer without its page. Keys 1 to 6 switch tabs, [ ] and j k walk the list the page was opened from, Esc or Backspace go back. The list stays mounted behind the page, so Back finds its view, search, filters, selection and scroll as they were, and focus returns to the row's arrow. The page covers loading, an unknown name with the closest matches, the gallery being out of reach, an install in progress with Cancel, and a failed install with Retry. The docs describe the page and its keys. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the model page and the Placement section New specs for the model page: reaching it from Explore, Installed, a double click, "o", a pasted link and a phone tap; the walker and Back with the search, a filter, the selection, the Installed view and the scroll kept, and no second read of the gallery; the title block, the answer strip, tabs by click, keys and arrows; Fit and memory, builds and files with the install call each one makes; an installed model's actions, used-by, the honest usage tab, configuration and logs; loading, an unknown name, offline, an install in flight and a failed one; and the phone. New specs for Placement: every mode and the keys it writes, the slider only when a layer count exists, the context presets, the bars and every verdict, two GPUs, no GPU, a cluster, an unread machine, a loading and an unavailable estimate, Fit it for me and Undo, and the section in the model editor with its save. The phone tap on a row now opens the page, so the two phone specs that expected the inspector as the page check the page and keep the inspector check for a window between a phone and a desk. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): link Studio results and open workspaces from a prompt Each workspace now records the result it was made from (parentId and an edge kind such as take, animate or to-3d) and reads a prompt, model, size, count and source from the query string, so one page can hand work to another. A source result is fetched from the server's own output file and becomes the start image, the picture for 3D, or the audio file. A note on the page says when the source loaded or could not be loaded. Diarization had no history; it now keeps the file name, the model and a speaker count, never the recording. Prompts are cut at 2000 characters when stored. The pure helpers (type suggestion, grouping, lineage layout, favourites, clearing) have node tests. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Studio front page as a composer with your work The front page is a prompt box with a chip per type, a type suggestion from the words, starters, and the options each workspace accepts. Generate opens the workspace with those filled in. A type with no model is a dashed chip that shows a gallery model, its size, memory need and an Install button only when picked; the typed words stay while it installs. Under it, Your work lists results from every workspace as a masonry with filters, counts, favourites and a Clear history action. Results made from each other stack into a project tile and open as a lineage board with a dock for running a new take or branching to the next step; steps the destination cannot start from yet are disabled with the reason. The docs describe the page, what is stored in the browser, and the query parameters a workspace accepts. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Studio composer, your work and the lineage view Specs for the type suggestion, the keys, hand-off to each workspace, the install path for a missing model, the masonry filters, favourites and clearing, stacking, the lineage board, new take and branch, steps that are disabled with a reason, and the phone layout. Existing Studio specs move from lanes to chips with the same intent. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the shared Studio workspace frame and move Images onto it The seven Studio workspaces get one layout: a row of type tabs, a compose card (optional sources as chips, a prompt with starters, a model chip, the essential options as chips, an Advanced fold that names what is inside, the memory the model needs, and one action with the reason when it cannot run), a run area, and a strip of recent results of the type. The run area shows a job card with the time that has passed and an indeterminate bar, because these endpoints report no phase or percentage; a failure with what the server said and one action; or the result with a toolbar: Favourite (the list the front page keeps), Download, Use in (the hand-off targets, disabled with the reason when a destination cannot start from the result), Re-run with edits (the take's values go back in the form, changed fields are outlined and listed) and Lineage. A type with no model shows the install note from the front page. Images is the first workspace on the frame. It keeps its size, count, steps, seed, negative prompt, source image and reference images, and its history writes, including the parent link and edge of a hand-off run. useMediaHistory.addEntry now returns the id of the entry it stored. The docs describe the workspace page. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move Video onto the workspace frame Video keeps its size list, duration, frame rate, steps, seed, CFG scale, frame count, negative prompt, start and end image and avatar audio. The start and end image are source chips, the avatar audio opens the recording and paste input from a chip, and the rest sit in the Advanced fold. A start image from a hand-off shows as a chip with its picture. Results play in the video player with the shared toolbar. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move TTS onto the workspace frame TTS keeps the saved-voice picker for cloning models, the typed voice for the others, the voice library deep link, and the delivery instructions, which now sit in the Advanced fold. The result is the waveform player with the words under it. The stored entry also keeps the voice id so Re-run with edits can select the same saved voice. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move Sound onto the workspace frame Sound keeps its Simple and Advanced modes and every field of both: the description, instrumental, vocal language, caption, lyrics, BPM, duration, key, language, time signature and think mode. The mode switch, instrumental and duration are in the compose card, the rest in a More options fold. The stored entry keeps all of the fields, so Re-run with edits restores the form as it was. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move Transform onto the workspace frame Transform keeps its audio and reference inputs with upload and record, the echo test, the key=value parameters (now in the Advanced fold), the input and output spectra and the three waveform players. The audio that was chosen shows before the run, waiting to be transformed. Re-run with edits puts back the model and parameters and fetches the audio and reference the server kept for that run. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move 3D onto the workspace frame 3D keeps the picture input with paste and webcam, the animation operations a model declares, quality and background, the shape and material steps, guidance and seed, the GLB and animation viewers, the remesh control and the download. A 3D result now has a title from the motion prompt when it has no label, so the strip and the front page name animation results by what was asked. Re-run with edits is shown disabled with the reason, because only a small thumbnail of the picture is kept. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move Diarization onto the workspace frame Diarization keeps its model and recording inputs, the option to prepare speakers to remember, the clean-speech previews, naming and remembering a speaker, and the history entry with only the file name, model and counts. The result now shows a timeline with one lane per speaker, the talk time of each speaker, and the segments with their start time and text. RTTM, SRT (only when the run has text) and JSON are built in the browser from the result. The helpers for talk time, axis ticks and the two text formats have node tests. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): remove the styles and lists the old workspace layout used Nothing renders the two-column workbench, the control column, the old history lists, the generation progress tiles, the TTS voice picker or the result echo any more. The inline-style baseline drops with them. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the workspace frame and one run per type Specs for the type tabs, the compose card and its reason when Generate cannot run, starters, the Advanced fold, the job card with no invented progress, a failed run and its one action, the install note, the strip with its favourites filter, Use in with its disabled steps, Lineage, the parent link, Re-run with edits and its list of changes, deleting and clearing, and the hand-off note. One run through each of Video, TTS, Sound, Transform, 3D and Diarization, the phone layout of all seven, and reduced motion. Existing Studio specs move from the old control column to the compose card with the same intent. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the chat thread with raised user turns, prose replies and one-line activity Your messages are raised blocks on the right at a 760 px measure and the model's replies are plain prose under its name and a warm or not loaded dot. Reasoning, tool calls and their results fold into one quiet line that opens inline into steps. Code blocks carry a Copy button and a Canvas button that opens that block in the canvas, image attachments are thumbnails that open in the lightbox, and files are chips. Per-message actions show on hover, on focus and on the last turn, and a turn takes focus so the arrow keys and C, E, R and B work. A failed reply keeps the text written so far and shows the reason with one Retry action. The Agent chat page keeps the older rules: the new styles are scoped to the chat page and use their own class names. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): use the Home command bar as the Chat composer Chat now ends in the same object as Home: the model chip, the MCP chip, a Canvas chip, the message box with attach buttons, a solid Send and the hint line, with the slash menu on the kit command list. The slash menu lists what Chat can do today (switch model, new chat, conversations, manage mode, canvas, find, settings, export, clear). While a reply is streaming Send becomes Stop, which Esc also presses, and Up in an empty box edits your last message. Attached images show as thumbnails and a line under the bar carries the speed and the token count. HomeComposer takes optional props for this (extra chips, its own slash list, Stop, paste, a stricter Enter); Home passes none of them. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): open conversations from a Ctrl K menu with day groups, undo and a slim header The conversations list opens as a centred menu on Ctrl or Cmd K. It groups chats by day like the Home resume list, shows the model that answered and the time, searches names and message text, and moves with the arrow keys. Enter opens a chat, F2 renames it and Delete removes it. Removing a chat hides the row and shows the kit undo toast; the chat is deleted for good only when the undo time ends. Rename, duplicate, copy and export are on each row, as before. The header is one slim bar: the Chats button, the chat name (click to rename), a context meter when the context size is known, settings and a More menu with rename, duplicate, copy, export, model info, keyboard shortcuts and clear. A dialog lists the shortcuts the page answers to. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): show loaded state, capabilities and fit in the Chat model switcher The model chip in Chat opens the same list as Home, grouped as Loaded now and Installed. Each row says warm or not loaded and marks models that understand images. When the list opens, the page reads the host memory once and asks the server to estimate each listed model at the chat's context size (up to twelve, three at a time), then shows what the model needs and whether it fits: free memory, how much would run on the CPU, or how far over the machine it is. A model with no estimate shows no fit text, and no load time is shown because the API does not report one. A memory bar closes the list. The picker takes the model list from the page when it has one, and useModels can skip its own request. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move chat settings into a sheet and add find in chat, jump to latest and a wider canvas Settings open as a kit sheet: the system prompt, temperature, top P and top K (each says "model default" until it is changed and has a Reset), the context size with quick sizes and a note that it only drives the meter, Manage mode and Focus mode, the model info for admins with its Edit config button, and Clear conversation behind a confirmation. The old slide-out drawer and the model info panel are gone. Ctrl or Cmd Shift F (or the search button, or /find) opens a search bar over the thread. It marks matches in the messages already on the page, shows "n of m" and steps with Enter and Shift+Enter. Nothing is sent to the server. Jump to latest is a pill above the composer. Esc stops a reply, then closes the search, then closes the canvas. The canvas panel gets the kit look: tabs, a Code and Preview switch, Copy and Download, a full-page layout on narrow windows, and translated labels. The Agent chat page shares it and gets the same look. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the empty, no-model, loading and phone states to Chat An empty chat opens with the composer under one line, starters to try, whether the model is loaded, and the Jump back in list: the same rows as Home, read from the chats the page holds. With no chat model installed, an install card offers the starter models for this hardware, the gallery and import, and the composer stays so the text is not lost. While a reply waits for a model, a load card shows what the page knows: the phase the server names, the node, the bytes and the time left when the server reports them, and a progress bar. A model that is just not loaded yet gets a plain note, with no invented phases or estimates. The foot warns when the context is nearly full. On a phone the header drops its labels, the model list and the settings open as sheets from the bottom, per-message actions stay in view and the canvas takes the whole page. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Talk as a calm voice page over what the connection really does Talk is one stage and one transcript. The stage has the pipeline chip, the voice and language chips, an outline orb that follows the real microphone and playback levels, a heading and a sentence for the current state, and the controls. The transcript lists You, Reply, Tool and Result lines and can be copied. Session settings (instructions, voice, language, tools, Manage mode and the pipeline's parts) open in a sheet. The states are the ones the code reaches: no pipeline model, idle, connecting, listening, thinking (also while a tool runs), speaking, an interrupted reply (the server cancelled it; a note marks the cut), a blocked microphone, a link that failed during a session, and any other error with its reason and a link to the traces. Push to talk and hands-free are not on the page, so they are not shown. Diagnostics keep their waveform, spectrum and stats, drawn in theme colours. The page text moves into the talk namespace, and the old Talk and visualizer styles and the inline-style count go down with the rebuild. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): remove the chat styles and strings the rebuilt page replaced The settings drawer, the model info panel, the bubble avatars, the conversation menu popover, the context bar, the recent strip, the staging bar, the file badges and the focus-mode rules have no user now. Their rules, the Chat page's focus class and seven unused empty-state strings are removed. The Agent chat page keeps the shared message, sidebar and input rules it still renders with. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the rebuilt Chat and Talk pages Add a Chat page under features (thread, message actions and keys, the message box and its slash actions, the model list with loaded state and fit, conversations on Ctrl K, settings, find, canvas and the empty, no-model and loading states) and a Talk section to the realtime API page with the states the page shows. Manage mode now turns on from the chat settings or /assistant, and the client MCP steps point at the MCP chip. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): settle the rough edges of the new Chat and Talk pages The undo toast sat under the conversations menu, so the Undo button could not be pressed while the menu was open; the menu, the sheets and the fullscreen canvas now stay below the toast layer. Esc in a rename box saved the text through the blur that follows it; it now cancels. The image viewer closed on Esc only when the page did not re-render on the same key, so its key listener is registered once and reads the latest handlers. Keys on a focused message no longer type their letter into the editor they open, "/" from outside a text field starts a command as it does on Home, and Esc leaves the page's own dialogs alone. Code in the canvas is highlighted for languages that have no preview. The conversations menu drops its key hints on a phone so Clear all stays in view. Talk hides Test tone while connecting and calls a server error "Something went wrong", since the call can still be open. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the rebuilt Chat and Talk pages Specs for the thread layout and the activity fold, code blocks, image thumbnails and the viewer, per-message actions and their keys, a failed reply with its one Retry, Stop and Esc while streaming, the composer and every slash action, the conversations menu (groups, search, resume, rename, delete with undo that ends by itself, one chat left), the model switcher with loaded state, vision and fit text from stubbed estimates, the settings sheet, the canvas panel, find in chat, Jump to latest, the empty, no-model and loading states, the phone layout and reduced motion. Talk is driven over a fake WebRTC link through idle, connecting, listening, thinking, speaking, interrupted, blocked, lost, error and no pipeline, its settings sheet and its phone layout. Node tests cover the message text helpers and the conversation grouping. The existing chat specs move to the new structure with the same intent: the transcript spec now describes the raised turn and the prose reply, and the render smoke accepts Talk's own header. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): keep a bounded run log for agents in the browser The server keeps no run history for an agent, so a run is one task and the events until the agent answers, written to browser storage while the page watches the stream: up to 50 runs per agent, task, step and answer text only. Stored chats from the earlier agent chat page read as runs with stable ids. A run still marked running five minutes after its last event reads as stopped. Helpers read an agent's config into chips, build the list of changed fields against the saved config, hide secret values and offer starting points. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Agents area around runs The Agents page shows what needs a look (work in flight, a run that failed in the last day), then each agent with its model, attached memory and skills, and a strip of its last 14 runs. An agent has its own page: model, tools, memory, skills, instructions, a task box and its runs. A run has an address, shows the thread while it works (steps folded into one line, the tool in use, the answer as it arrives) and settles into a report about a second and a half after the agent answers: task, outcome, follow-ups, evidence and steps, with wide tables opening wider on demand. A failure says in plain words what happened and offers Run again. Create and edit fold into sections with a ready mark and a one-line summary, start from a template or an optional model-written draft, and open a preview sheet with the config as saved and the changes against the saved agent. Status becomes a quiet panel in the same language, and the old chat link opens the agent page. There is no Stop, approval, steer, version or dry-run control, because the agent API has no call behind them. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Agents launcher, agent page, runs and editor Specs for the Now strip and run strip, search and the empty state, the agent page, starting a run, the live thread, settling into the report, the run address across a reload and for a run from another browser, follow-ups with their history, failures, the folding editor with ready marks, templates, the preview sheet with hidden secrets and changes, the status page, and the phone, 1440 and 2560 layouts. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe runs and the new agent create flow Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add helpers for tasks, schedules and job outcomes Reads a cron expression the way the server does (five fields or an @ shortcut), checks it, and puts the common shapes in words. The next run is left out on purpose, because the schedule follows the server clock, which the browser cannot read. Also groups jobs by day, sums the last seven days, and gives each job one outcome line from its result or error. A rerun call starts a new job with the same parameters and media. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Jobs area around runs The Jobs page opens with one sentence about the last seven days, then the tasks (model, schedule in words, last 14 jobs, enabled switch, Run now) and a run history grouped by day. Each row has an outcome sentence and opens to the error or the start of the result with one next action. Deleting a task waits 30 seconds with an undo button. A task opens as a page with its recent runs, its prompt with the gaps marked and its schedule. The task form folds into sections, takes a schedule as a preset or a checked cron expression, warns about prompt gaps the schedule does not fill, and has a preview sheet. A job opens as a document: task, outcome, delivery and the recorded steps; a failed job says what happened and offers Run again. Run now now sends attached media through the job call, which is the only one that takes it. "Clear History" only ever cancelled running jobs, so it is now called Stop running jobs. Webhook headers of a saved task show as JSON instead of [object Object]. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Jobs page, task pages and job pages Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the Jobs page and the task form Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add helpers that say who uses a skill or a collection An agent loads a skill when skills are on and the skill is in its selection (an empty selection means every skill). It reads the one collection that carries its own name, when its knowledge base is on. The helpers derive that from the saved agent configs, build the config that adds or removes a skill or a collection, and estimate tokens as characters divided by four. Removing the last selected skill switches skills off, because an empty selection would mean every skill. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Skills and Memory as one library Skills and collections sit in a list with an open item beside it. Each row says who uses it, read from the saved agent configs, or says it is not used yet. Chat reads neither, so it is never named. An item opens in a pane with a Used by strip (names link to the agent, a small x removes it, with undo) and an Add to menu that shows what the addition costs. A collection can be added only to the agent that carries its name. The Memory pane searches the collection alone and shows ranked passages with scores, lists web sources with their refresh interval and the files, shows the server message when an upload fails, and names the endpoints and where files stay. The Simulate a message sheet runs a collection search and shows an agent's skills with a token estimate. It runs no model. The collection details route now opens the same page. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): show use and cost in the agent form pickers Each skill in the agent form says which other agents use it and what it adds to every message, with a total for the selection. The memory section names the collection the agent reads. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Skills and Memory libraries Specs for the used-by lines (including an agent that uses every skill), the filters, search, add to agent, remove with undo, the last-skill case, an unreadable agent list, the empty states, git repositories, the Memory question box, sources, uploads that fail, the Simulate sheet with the parts the API can run, the agent form hints and the phone layout. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the Skills and Memory libraries Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add helpers for the Operate status page and backend rows Pure functions for the parts that need rules. They work out the memory pool the page measures (GPU memory, system memory, or the workers of a cluster that are answering), which pools are too full, the headline and the four ledger rows, the geometry of the capacity chart, and what removing a backend would leave without a runtime (models name their backend, and a meta backend names the concrete one it points at). A second set says what a backend row states: installing, queued, removing, failed, update available, current or absent. LocalAI keeps no memory history, so the chart reads a bounded buffer of readings the page took itself and says so. A reading with no total is dropped rather than drawn as zero. Two hooks are shared by the pages that need them. One retries a failed operation after moving the failure into the record. The other holds a cancel for an undo window, because the server cannot take a cancel back. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Operate Status, This machine, Backends, Activity and Logs Status opens with one sentence ("2 things need you", or "Everything is running") and four rows: Needs you, Capacity, Running now and Recent failures. A row with a problem opens by itself and holds the button that deals with it: Update a backend, Retry or Dismiss a failed operation, Unload a model. A quiet row stays one line. A new installation gets a first-run screen, a cluster sums the memory of the workers that are answering, and a page still waiting for an answer says so. The chart under the rows is drawn from readings the page took while it was open and is labelled that way, because LocalAI keeps no memory history. This machine leads with GPU memory as one bar, then host memory split by running model, then VRAM, RAM, CPU and disk with a bar each. The running models become a kit table with the same menu and stop dialog. Backends is one list with Installed and Catalog views. A row says what the backend is doing (a progress bar with Cancel, Queued, Failed with Retry, Update 1.2.0, Current), carries the one button that matters, and opens in place. Removing a backend names the models and the meta backends that would stop working. Check for updates, Update all, From URL and a first-run recommendation for llama-cpp are in the header. Activity keeps its three sections as quiet rows. Cancel waits eight seconds with an undo toast, because the server cannot take a cancel back; a cancelled install can be started again from the record. Logs gets a process list, a picker, stream and text filters, Follow and Times switches, and a Clear with an undo window. Not shown, because the API has no data for them: GPU temperature and power, a size per backend, an earlier version to roll back to, a dependency lookup beyond the models and meta backends that name a backend, and models that failed to load. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover Operate Status, This machine, Backends, Activity and Logs New specs for the Status headline and ledger (healthy, needs attention, one thing, a full memory pool alone, loading, first run, cluster), its actions (Update, Retry, Dismiss, Unload with its dialog), the capacity chart built from readings taken while the page is open and bounded, the phone layout, no coloured edge on a row, and reduced motion. The Backends specs cover the two views, install progress with Cancel and its undo window, Retry on a failed install, Update, Update all, Check for updates, removal with the models and meta backends it would break, Install from URL, the first-run recommendation, a cluster, and a phone. Activity gains cancel with undo, Cancel now, a second cancel, leaving the page, progress, and starting a cancelled install again. Logs covers the stream and text filters, Follow, Times, Export, Clear with undo, the process picker and list. This machine covers the GPU strip, several GPUs, no GPU and Add a machine. Existing specs keep their intent and follow the new structure: rows open in place instead of in a pane, Update replaces Upgrade, the notice spec now pins that an update is a row state and not a banner or a rail, and a cancel waits for its undo window. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the rebuilt Operate Status, Backends and Activity pages Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add helpers for the Swarm pages Pure functions for what the pages work out from the cluster API: a node's state in words, which nodes a placement rule may use, what a rule would ask for, what a drain or a lost node would leave without service, the nodes a bulk backend update reaches, and the join commands for a worker, a peer instance and a memory shard. Hooks read the roster, the loaded replicas and the rules. Everything runs in the browser from data the page already holds, and says when it cannot see free memory or disk. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Swarm hub: nodes, node page, placement rules, failover Nodes is a sortable table with comfortable and compact rows, a Needs attention filter by reason, a map of the cluster that is not drawn on a phone, the running models, and a bulk backend update for the nodes that drifted. A node is a page: state, vitals, a drain preview computed from the loaded replicas and the rules, tabs for models, backends, logs and capacity and labels, and Remove that asks for the node's name. Placement rules are written as sentences, show where each model is loaded now, and edit in a side sheet with a preview of the nodes a draft could use. Deleting a rule waits a few seconds so it can be taken back. Failover keeps its chains, adds what the router does when a worker stops answering and a per-node preview of what would stop. Add a node covers a registered worker, a peer instance and a memory shard, with a command to copy and a live line that says when the machine arrived. P2P keeps its page in the same vocabulary, and the node logs page follows the local logs page. Previews are labelled as worked out in the browser. Per-GPU readings and node events are not drawn because the API does not return them. Failover moves to Swarm when distributed mode is on. Legacy fleet components and their styles are removed. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Swarm hub New specs for adding a node (each join method, the command, copy, waiting and found, approve, a single install, P2P, a phone), placement rules (sentences, where models are loaded, the preview matrix, the sheet and its preview, delete with undo) and failover on a cluster. Node detail covers its tabs, the drain preview and its dialog, resume, remove with the typed name, a node that stopped answering, and unload. The nodes specs follow the new structure and keep their intent: the table, filters, grouping, pagination, bulk actions, the map, and running models with stop, logs and the loading, error and empty states. The scheduling, failover, P2P, hub and smoke specs follow the renames. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the rebuilt Swarm pages Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add helpers for the Traffic pages Pure functions for what the pages work out from the usage ledger, the trace summary, the trace buffers and the resources reading: the shared time window, grouping, sorting and filtering of usage rows, chart series and axes that start at zero, the overview figures, per-model statistics, the state of a trace and the words for a failure, the backend operations that ran during a request, CSV export, the Prometheus metric list and scrape config, and a bounded buffer of host readings. A figure whose source cannot say is null, never zero. The trace summary call takes the window in hours, and a helper reads /metrics with its status. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Traffic hub: overview, usage, models, host, traces, middleware Traffic opens on an overview: five figures (requests, failed, p95, tokens in and out) and three charts, each naming its source. A second row of links reaches Usage, Models, GPU and host, Traces, Middleware and Prometheus, and one time window is shared by the first three. Usage groups by model, user or API key, filters, sorts, opens a row on its own chart, exports the rows it holds as CSV or JSON in the browser, and keeps the opt-in cost estimate and the quota forecast. A user who is not an admin sees only their own numbers. Models joins the ledger, the backend-operation buffer and the loaded models. GPU and host shows the current reading and two charts of readings taken since the page opened. Traces gets filters, a settings strip and an explained off state. An API request is a page: the error LocalAI recorded, a timeline with the backend operations that ran meanwhile, and bodies that stay closed until revealed. Middleware draws the pipeline as five steps and shows the rules of the selected step. Prometheus documents /metrics, checks it against the server and gives a scrape config to copy. Alerts is not built: LocalAI has no alert rules. Per-model latency percentiles, GPU utilisation and compare with the previous period are not drawn because the API does not return them. Legacy usage, trace and middleware styles and the usage source components are removed. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Traffic hub New specs for the overview (figures, charts with a data table and arrow key readout, failed and first-run and tracing-off states, the shared window, a phone), usage (group by, filters, sort, export, cost, quotas, a non-admin, empty and loading), models, GPU and host (snapshot, the since-opened labelling, a cluster), the traces list, a trace page (the real error, the timeline, reveal, no headers, a trace that left the buffer), Prometheus and the Middleware pipeline, with shared fixtures. The usage, traces, middleware, hub and smoke specs follow the new structure and keep their intent. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the Traffic hub Add an operations page for the Traffic tab: which record each page reads, what it leaves out and why, the trace page and its reveal, the GPU and host readings kept since the page opened, and the Prometheus endpoint. Link it from the operations index, the tracing page and the middleware page. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): let metric names wrap in the Prometheus table on a phone The long metric names pushed the type and "on this server" columns out of view. Names now wrap inside the table. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Settings with groups by intent, search, a pending bar and history The fifteen sections become eight groups by intent: memory and models, speed and defaults, backends and galleries, access and security, debugging and traces, agents and responses, swarm and sharing, look and feel. Search covers names, descriptions, keys and the old section name, and says where a result used to be. Edits wait in a bar with Discard, Show diff and Apply. The diff lists old and new values and the checks the browser can make: durations parse the way Go parses them, a GPU memory budget is one the server accepts, a gallery box holds JSON, and warnings repeat what the handler and the field text say. Apply sends only the changed keys. Undo saves the previous values again; it is a new save, not a rollback. History lists the changes applied from this browser, since LocalAI keeps no settings log, and Revert stages the old value. A value is marked as changed only where the built-in default is known from the CLI defaults. A row says "Applies now" or "Needs restart" only where the handler or the docs say so. Three things were wrong before and are fixed with the rebuild: the gallery boxes and the shared API keys box were sent under names the server ignores, the "Enable CSRF Protection" switch showed the disable flag the wrong way round, and every save restarted peer-to-peer networking because every field was sent. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Users and keys, Account, sign-in, invite and the 404 page Users and keys is a tabbed page under the Settings tab: people, invites and API keys. The people table filters by state and role, sorts, approves or disables (disabling offers an undo that sets the status back), and opens a side sheet for one person's features, model allow-list and limits. Role, password reset and delete sit in the row menu; delete asks for the name. Invites choose a lifetime of 1, 7 or 30 days and show the link once. API keys can be created with a lifetime, are shown once in full, can be paused, and are revoked after a ten second undo window in which nothing is sent. LocalAI lists keys only to their owner, so the tab shows the signed-in person's own keys and says so. Account has Profile, Security, API keys and Usage. Usage shows the last 30 days, tokens by model and the limits an admin set. The Security tab now shows for a GitHub or SSO account and says the password is not theirs to change. Sign-in asks for one field per step and draws a provider button only for a provider /api/auth/status lists. It has the notice for a sign-up that waits for approval, the first-admin screen, the key-only screen and the invite page. An address outside the app now gets the 404 page too, which names the address and lists the places the sidebar lists, with the same gates. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover Settings, Users and keys, Account, sign-in and the 404 page Settings: groups, search by name, key and old section, the changed marker only where a default is known, apply hints, the pending bar and diff, the checks, apply sending only changed keys, undo as a second save, discard, history, the CSRF inversion and the gallery and API key wire forms, and the phone layout. Users and keys: the table, filters, sort, approve, disable with undo, the row menu, the access sheet, invites, key creation with a one-time reveal, the ten second revoke with undo and with a page leave, and the non-admin redirect. Account, each sign-in variant (error, pending, first admin, key-only, invite, provider buttons) and the 404 page have specs too. Fixtures are shared with the screenshot scripts. Existing specs follow the new structure. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the rebuilt Settings, Users and keys, Account and sign-in pages Runtime settings: the eight groups and where each old section went, search, the pending bar, the diff and its checks, apply, undo, the history, and which settings show a default or an apply note and why. Authentication: the sign-in screen variants, the Account tabs, key lifetimes, the one-time key reveal, the revoke undo window, and the fact that keys are listed only to their owner. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): let the Settings undo toast stand alone and read back a generated P2P token The saved message and the undo toast sat on the same spot at the bottom of the page. The undo toast now carries the saved message. A new P2P token is made by the server when the page sends 0. The page reads it back after the save so the field shows the token and not the placeholder. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): fit the users table, API keys and Account figures on a phone On a phone the users table dropped its Role and Status columns off the screen edge with the row actions. The role and state now sit under the name, so the actions stay in view. API key rows no longer put the key icon on a line of its own, and the three Account figures keep one row. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the phone users table, reduced motion and the empty Account state Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): drop the apply note from three settings the save handler does not mention Size-aware eviction, automatic backend upgrades and development backends said Applies now, but nothing in the handler or the docs says when they take effect. A row now carries a note only where the code or the docs say so. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Voices and Faces as one identity family Voices is one page with three tabs: Speakers (voiceprints for recognising who is speaking), Speech voices (the text-to-speech reference library, kept apart because it is a different store) and From a recording (a link into the diarization workspace). Faces uses the same layout. Who is this and Same person? give the answer in a sentence with the real distance and cut-off, a word for how far inside the cut-off it sits, and a distance scale with the cut-off drawn on it. The cut-off slider re-reads the answer in the browser; the identify call sends the cut-off, and verify uses the threshold the model returns. The old confidence percentage is gone because it is not a probability. The server has no list call, so the people list stays in the browser and the page says so. After a search that asked for more people than it got back, a saved person the server did not return is marked, and people the server returned that the browser does not know are listed. Nothing is claimed from a short or cut-off search. Enrolling is a sheet: sample, name, labels, permission. A copy of the sample in the browser is opt-in, and an administrator can also keep the recording as a speech voice in the same step. Removing a person waits ten seconds behind an Undo toast and sends nothing before then. Errors say what happened (no face found, model missing, call failed), a blocked or missing microphone is explained, and a missing model or a missing permission renders a page that says what turns the feature on instead of a redirect. Analyze, detect and raw embedding move under More tools, with attribute guesses off by default. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Voices and Faces pages Specs for who is this (match, no match, working, failed, missing model), the cut-off slider, same person, a blocked, allowed and insecure microphone, the registry notes and the not-on-the-server marks, the enrol sheet and its opt-in copy, delete with undo on a fake clock, the disabled and no-permission states, the phone layout, reduced motion and Faces. Existing library and diarization specs follow the new structure and keep their intent. Node tests cover the distance words, scale layout, stored list and error mapping. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the Voices and Faces pages Add a WebUI section to the voice and face recognition pages: the two tools, the cut-off, what the people list is and why it can be stale, the undo window, and what is stored where. Point the Voice Library and Fish Audio notes at Build, Voices, Speech voices. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Build landing, Fine-tune, Quantize, Import and Explorer The Build landing says what each tool is for and what it needs from the machine: the installed backend, the GPU memory, RAM and disk the server reports, and a job that is running or the newest one when it failed. A tool that cannot run says why and what enables it. Fine-tune and Quantize share one page: set up, a check list that is redrawn as the form changes, a run view with progress, stages and a log, and a result with real next steps (export, import, chat, Models). The checks state only what the server reports. A job needs no estimate the server cannot make, so none is invented. Stop on a fine-tuning job asks whether to keep a checkpoint, a failed job shows the server's message, and a memory failure offers two changes that are applied to a copy of the setup. Import is a guided flow: source, review, import, done. The server returns no preview before an import starts, so the review reads the spelling of the source, prints the request the form will send and runs the checks that can be made early. The estimate that arrives when the import starts is set against free memory and disk. The ambiguity picker and the Write YAML tab stay. Explorer shows what GET /networks returns and lists a swarm with POST /network/add, with a join sheet that carries the token and commands. Build tools the account may not use say so instead of redirecting. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Build landing, the tool pages, Import and Explorer Specs for the landing (a tool ready, missing a backend, with no GPU, a running or failed job, a feature switched off, a member without admin, phone, reduced motion), the shared tool pattern for Fine-tune and Quantize (set up, live checks, start request, running with progress, chart and log, the stop choice, failure with the server message, finish with next steps, earlier jobs, the account-disabled page, phone), Import (source detection, review, checks, ambiguity, running with the estimate against free memory, done, Write YAML, phone) and Explorer (list, join, list a swarm, empty, not an explorer, retry, phone). Existing specs follow the new structure and keep their intent. Node tests cover the machine facts, tool status, checks, log lines, source detection, the import request and the join commands. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the Build tool pages, the import flow and the Explorer Fine-tuning and quantization now describe the set up, check, run and result steps and what the check list can and cannot say. The import section explains the review step and why the size and memory appear only after the import starts. The distributed page describes the Explorer list, the join sheet and what listing a swarm publishes. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): turn the hardware recommendations into a "Best for this machine" shelf The shelf in the Models inspector put five columns into a 400 px pane, so long model ids wrapped letter by letter underneath the size and the memory figures. Each row now stacks the tag, the id and the size and memory facts beside one Install button, and the id wraps inside its own column. Once a model is installed the shelf narrows to the best fit and keeps the others behind a "N more that fit" toggle. Specs cover the ranking, the layout, the narrowing and the install request against a gallery fixture that carries the 4K estimate the shelf sizes against. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): quiet the Studio tab markers and say their state in words The type tabs drew a saturated green dot for every modality that has a model. The dot now uses a text colour, filled when a model is installed and hollow when none is, and each tab carries "(model installed)" or "(no model installed)" as hidden text so the state is not only a colour. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): stack the model editor empty state and drop its section hues "No fields configured" sat in a flex row, so the icon, the title and the text ran together. It now uses the stacked empty-state layout. The section icons took a different status colour each (amber, red, green); they now share one quiet colour, with the accent on the current section. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): keep the Home memory sentence whole on a phone and drop side rails On a 390 px screen the memory strip clipped "2 models loaded" to make room for the figure. The sentence now takes the first line and the figure and device wrap under it. The sweep also removed coloured left rails from the editor section rail, the skill editor list, the install strip and the audio transform notice (now an outlined note), plus unused chat rules that carried rails and two glow animations that nothing referenced. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): line up hub pages, list the model templates, and stop clipped text Medium-width pages inside Build and Operate were centred while the tab bar above them was flush left, so the title started 60 px right of the first tab. They now start at the bar's edge. Add Model offered nine templates as a grid of identical cards with chip clouds and inline styles. It is now one list of rows, each with the field names it fills in on a single muted line. Two clipped strings are fixed: the Studio voice field cut its placeholder mid-word, and the phone job list ended the schedule line in an ellipsis. The recommendation shelf also separates size and memory with a dot, and the docs describe the shelf. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): keep the hidden Studio tab state inside its tab The hidden state text added to each type tab was absolutely positioned against the page, so on a phone it sat outside the scrolling tab row and widened the page by hundreds of pixels. The tab is now the containing block. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the on-disk sizes on the Installed table, model page and cleanup sheet The fixtures stub GET /api/models/storage. The default report is empty, so existing specs keep the gallery estimates. makeStorage() builds a report from files and the models that use them, the way the server does, and storageSpec() is a models directory with shared and missing files. New specs cover the Size column and its shared line, the fallback when the call fails or the user is not an admin, the files list on the model page, a missing file, the bytes a removal frees with shared files, and the cleanup findings. Node tests cover the storage helpers and the batch arithmetic. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): wait for the page before pressing keys and ticking the clock Two specs failed in loaded full runs and passed alone. The Alt+1 to Alt+7 spec pressed a key before the composer had armed its key handler. The capacity chart spec advanced the fake clock before the poller had mounted, so it counted fewer readings than it expected. Both now wait for the page to mount. The key spec retries a press that lands during a re-render, and the clock spec advances in small steps and polls for the row count. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): hide the Installed footer when the storage report is empty An empty report from the storage call made the footer read "0.0 GB on disk" next to sizes taken from the gallery estimate. An empty report says nothing about the disk, so the footer now shows only the model count. A spec covers it. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): keep Explore pane actions inside the pane The inspector actions sat in a non-wrapping flex row beside the title, so the buttons ran past the pane edge once it got narrow. The row now takes its own line and wraps. The primary action (Install, Retry, Open) comes first. Manage installation becomes a ghost button, and Open details moves to the end of the row, so one action stands out and the others are quiet. No action or test id is removed. Add a spec that checks, in light and dark at several widths and with a pane forced to 320 px, that every action stays inside the pane box and that the pane keeps its inner padding. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * refactor(ui): drop the type chips from the Studio composer The Studio tabs and the composer's type chips listed the same seven modes, so the page said the same thing twice. Keep the tabs as the one place to switch modes. The composer now shows the type it will open as a small label in its header. The type suggestion from the typed words stays as the quiet hint line under the prompt, and Alt+1 to Alt+7 still pick a type. The composer root carries data-type, data-types and data-missing so tests can read the state. Specs pick a type through a shared Alt+digit helper and read a missing model from the tab dot instead of a chip. Remove the unused chip locale strings and CSS. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * refactor(ui): remove the section crumb above page titles Page headers drew a small uppercase crumb with a short rule before it above the title. On the hub pages it repeated the hub name, so Build sat above a heading that also said Build. PageHeader now renders only the title, the supporting line and the actions. Drop the eyebrow prop, the route-derived section name, its CSS and the unused section helper, and remove the explicit eyebrow props from the pages that passed one. Pages stay reachable through the sidebar and the hub tab bar. Add a spec that checks several pages show their title with nothing ahead of it in the header. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * refactor(ui): remove left accent rails from tiles, rows and quotes Several surfaces marked state with a coloured strip on the left edge. Replace each one with a cue that is not a rail: - Stat cards lose the strip; the icon and value still carry the colour. - The highlighted card is a raised surface with a firmer edge. - The selected rail row is an accent wash with a hairline outline. - The status stripe on rail items is a small status dot. - The active failover row is a tinted row. - Quotes in markdown and chat prose are italic instead of barred. - The variant detail panel has a full hairline border. Add a spec that walks the main routes in light and dark and fails on a left border thicker than 1px, a sideways inset shadow, a narrow absolute strip in ::before or ::after, or a narrow tall child pinned to a left edge. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
579 files changed
+77619
-30833
No files matched your search
@@ -246,6 +246,48 @@ These settings apply to most LLM backends (llama.cpp, vLLM, etc.):
|
||||
| `main_gpu` | string | Main GPU identifier for multi-GPU setups |
|
||||
| `cuda` | bool | Explicitly enable/disable CUDA |
|
||||
|
||||
### Placement
|
||||
|
||||
The model page's **Configuration** tab and the model editor have a **Placement**
|
||||
section for the keys above. It edits `gpu_layers`, `tensor_split`, `main_gpu`
|
||||
and `context_size` and nothing else.
|
||||
|
||||
- **CPU only** writes `gpu_layers: 0`.
|
||||
- **Auto** leaves `gpu_layers` unset. LocalAI then asks the backend for every
|
||||
layer (`99999999`), and the `llama-cpp` backend lowers that to what fits in the
|
||||
free device memory unless the model turns off `fit_params` (see [GPU auto-fit settings](#gpu-auto-fit-mode)). Other backends read
|
||||
an unset value as their own default.
|
||||
- **Custom** writes the number you enter. **All layers** writes `99999999`, the
|
||||
value LocalAI uses when the key is unset. A slider appears when the memory
|
||||
estimate reports the model's layer count (`block_count`); today it does not, so
|
||||
you enter a number.
|
||||
- With two or more GPUs, **Split across GPUs** writes `tensor_split` as
|
||||
percentages (for example `65,35`; llama.cpp reads them as proportions) and
|
||||
**Main GPU** writes `main_gpu`. **Split by free memory** sets the shares from
|
||||
the memory each card has free now.
|
||||
|
||||
The bars show, for each GPU and for system memory, what other programs use, what
|
||||
this model needs (weights and working memory, and the KV cache that grows with
|
||||
`context_size`) and what is left. They come from the device list in
|
||||
`/api/resources` and from `POST /api/models/vram-estimate`, which returns one
|
||||
total for the chosen context size and layer count. The per-device split of that
|
||||
total follows `tensor_split`, so it is an approximation. The page says what the
|
||||
choice means in words ("Fits in GPU", "Spills to CPU", "Too many layers for the
|
||||
GPU") and does not state a speed. A usable limit is 95 percent of the free
|
||||
memory.
|
||||
|
||||
**Fit it for me** searches for the largest `gpu_layers` whose estimate fits the
|
||||
free GPU memory, by asking the estimate endpoint, and **Undo** puts the previous
|
||||
value back. It is hidden when the estimate is unavailable (the model file is
|
||||
missing, still downloading or in a format the estimate does not cover) and when
|
||||
the host has no GPU. A server that schedules models onto other machines shows no
|
||||
per-device bars, because the devices listed belong to the controller.
|
||||
|
||||
Saving from the model page sends only the keys that changed. A key set back to
|
||||
Auto is sent as `null`. The patch endpoint merges values into the file and cannot
|
||||
delete a key, so the file then reads `gpu_layers: null`, which LocalAI treats as
|
||||
unset. To remove the line itself, edit the YAML.
|
||||
|
||||
### Mixed CPU/GPU inference
|
||||
|
||||
The `llama-cpp` backend can run one GGUF model across CPU and GPU, using both system RAM and GPU VRAM.
|
||||
|
||||
@@ -43,7 +43,24 @@ LOCALAI_DISABLE_AGENTS=true
|
||||
1. Navigate to the **Agents** page in the web UI
|
||||
2. Click **Create Agent** or import one from the [Agent Hub](https://agenthub.localai.io)
|
||||
3. Configure the agent's name, model, system prompt, and actions
|
||||
4. Save and start chatting
|
||||
4. Open **Preview** to read the configuration that will be saved and, when editing, what differs from the saved agent. Secret values are hidden in the preview.
|
||||
5. Save, then give the agent a task from its page
|
||||
|
||||
The form folds into sections. Each section shows a mark when it is ready and one line that says what it holds. **Start from** offers a few starting points. **Draft** asks the model you chose to write a name, a description and instructions from one sentence; it is optional, nothing is saved until you save, and the draft may be wrong.
|
||||
|
||||
### Running a task
|
||||
|
||||
Open an agent from the Agents page to see its model, tools, memory, skills and instructions, and a box to give it a task. A task starts a **run** with its own address, `/app/agents/<name>/runs/<id>`. While the agent works the page shows the thread: the steps folded into one line ("Worked 26 s, 5 steps"), the tool in use and the answer as it arrives. About a second and a half after the agent answers, the page settles into a report: the task, the outcome, the evidence (what each tool returned) and the steps. Follow-ups sit under the outcome and carry the earlier turns. Wide tables and code open wider on demand.
|
||||
|
||||
The server keeps no run history. Runs are recorded in the browser that watched them, up to 50 per agent, and the record holds task text, step text and answers. A run link opens only in the browser that recorded it, **Clear run record** on the agent page removes the record, and the last 14 runs of each agent appear as a strip on the Agents page. A run still marked as working five minutes after its last event reads as stopped. Durations are measured by the browser between sending the task and receiving the answer. The page does not show tokens or per-step timings from the server, and it has no Stop or approval control, because the agent API has neither; **Pause** on the agent stops it taking new work.
|
||||
|
||||
### Jobs and scheduled tasks
|
||||
|
||||
The **Jobs** tab lists tasks: a prompt a model runs for you, on a schedule or whenever you start it. The page opens with one sentence about the last 7 days, built from the jobs the server returned (how many ran, how many finished, failed, were cancelled or are still going, and which task failed last time). Each task shows its model, its schedule in plain words with the cron expression under it, its last 14 jobs and an enabled switch; **Run now** asks for the values the prompt uses (`{{.name}}` gaps) and any media to attach, and the menu holds Edit and Delete. Delete waits 30 seconds with an Undo button, and nothing is deleted on the server until that time ends.
|
||||
|
||||
The run history is grouped by day, with one sentence for each job (the first line of its result, or its error). A row opens to show the error or the start of the result and one next action, **Run again** for a finished job or **Cancel** for one that is still going. Run again starts a new job with the same parameters and media, because the API has no retry call. A job opens as a document: the task as that run sent it, the outcome, whether the webhook was delivered, and the steps the server recorded.
|
||||
|
||||
The task form takes a schedule as presets (hourly, daily, weekdays) with a time, or as a custom cron expression of five fields (minute, hour, day, month, weekday) or an `@hourly` style shortcut. The form checks the expression and says what it means in words. Times follow the clock of the machine that runs LocalAI, which the browser cannot read, so the page shows no next run. The page also shows no tokens, because the jobs API returns none, and durations come from the job's own start and end times.
|
||||
|
||||
### Importing an Agent
|
||||
|
||||
@@ -252,6 +269,15 @@ curl http://localhost:8080/api/agents/skills
|
||||
|
||||
If a skill you expect is missing, confirm LocalAI was started with `LOCALAI_AGENT_POOL_ENABLE_SKILLS=true`.
|
||||
|
||||
### The Skills and Memory libraries
|
||||
|
||||
The **Skills** and **Memory** pages work as libraries. A skill or a collection exists on its own and needs no agent. Each row says who uses it, read from the saved configuration of your agents: "Used by research-assistant +2", or a quiet "Not used yet". An agent uses a skill when `enable_skills` is on and the skill is in `selected_skills` (an empty selection means every skill). An agent reads the one collection that carries its own name, when `enable_kb` is on. Chat does not read skills or collections, so it is never listed as a user.
|
||||
|
||||
- **Add to...** on a skill opens a menu of agents. Each row shows an estimate of what the skill adds to every message: the characters of its content divided by four. With `skills_mode` set to `tools` the content is read only when the model asks, so nothing is added up front. Removing the last selected skill from an agent switches skills off for it, because an empty selection would mean every skill. Removing a skill or a collection from an agent can be undone for a few seconds.
|
||||
- **Add to...** on a collection offers only the agent with the same name, and turns its knowledge base on.
|
||||
- On a collection, **Try a question** searches that collection alone and shows the passages that come back with their scores (`POST /api/agents/collections/{name}/search`, with `max_results`). **Add source** uploads a file or adds a URL with a refresh interval in minutes. A failed upload shows the message the server returned, and can be retried.
|
||||
- **Simulate a message** shows what a context would load. Pick a collection on its own, or one of your agents: you see the skills the agent has on, an estimate of the tokens it adds around the message, and the passages that its own collection returns for that message. Nothing is sent to a model, so no answer is shown. A skill or collection that is not added anywhere can be tried in the sheet for that test only, and is never saved.
|
||||
|
||||
## API Endpoints
|
||||
|
||||
All agent endpoints are grouped under `/api/agents/`:
|
||||
@@ -388,7 +414,7 @@ curl -X POST http://localhost:8080/api/agents/my-agent/chat \
|
||||
}'
|
||||
```
|
||||
|
||||
The web UI does this for you: each conversation in the agent chat sends only its own visible turns. **New Chat** and switching conversations therefore continue from that conversation alone, and **Clear** starts the conversation over without history. A request without `history` is answered without earlier context; the server keeps no web chat history of its own. In distributed mode (NATS) the history is not forwarded yet.
|
||||
The web UI does this for you: a run sends only its own turns as history, so a follow-up sees the task and the answers before it, and a new run starts without history. A request without `history` is answered without earlier context; the server keeps no web chat history of its own. In distributed mode (NATS) the history is not forwarded yet.
|
||||
|
||||
Listen to real-time events via SSE:
|
||||
|
||||
|
||||
@@ -238,7 +238,7 @@ voice conversion from the same weights.
|
||||
## Family notes
|
||||
|
||||
- **Fish Audio voice cloning**: save a reference clip with its transcript in the
|
||||
Voice Library, then select **Use in Text to Speech**. The backend accepts
|
||||
Voice Library (**Build → Voices → Speech voices**), then select **Use in Text to Speech**. The backend accepts
|
||||
`params.ref_text` as an alias for `params.reference_text` in both ordinary and
|
||||
streaming speech requests. If you supply both parameters, `reference_text`
|
||||
takes precedence. For direct requests with a reference file in `voice`, supply
|
||||
|
||||
@@ -174,7 +174,7 @@ The **first user** to sign in is automatically assigned the admin role. Addition
|
||||
|
||||
### Invite Links
|
||||
|
||||
Admins can generate single-use, time-limited invite links from the **Users → Invites** tab in the web UI, or via the API:
|
||||
Admins can generate single-use, time-limited invite links from **Settings → Users and keys → Invites** in the web UI (choose 1, 7 or 30 days), or via the API:
|
||||
|
||||
```bash
|
||||
# Create an invite link (default: expires in 7 days)
|
||||
@@ -192,7 +192,7 @@ curl -X DELETE http://localhost:8080/api/auth/admin/invites/<invite-id> \
|
||||
-H "Authorization: Bearer <admin-key>"
|
||||
```
|
||||
|
||||
Share the invite URL (`/invite/<code>`) with the user. When they open it, the registration form is pre-filled with the invite code. LocalAI validates the code only when the user submits registration. Invite codes are single-use - once consumed, they cannot be reused. Expired or used invites are rejected.
|
||||
Share the invite URL (`/invite/<code>`) with the user. The link is shown once, when you create it: the invite list keeps only the first characters of the code, so a link you did not copy can only be revoked and made again. When the user opens it, the registration form is pre-filled with the invite code. LocalAI validates the code only when the user submits registration. Invite codes are single-use - once consumed, they cannot be reused. Expired or used invites are rejected.
|
||||
|
||||
For GitHub OAuth, the invite code is passed as a query parameter to the login URL (`/api/auth/github/login?invite_code=<code>`) and stored in a cookie during the OAuth flow.
|
||||
|
||||
@@ -236,10 +236,23 @@ When authentication is enabled, the following endpoints require admin role:
|
||||
When auth is enabled, the React UI sidebar dynamically shows/hides sections based on the user's role:
|
||||
|
||||
- **All users see**: Home, Chat, Images, Video, TTS, Sound, Talk, Usage, API docs link
|
||||
- **Admins also see**: Models, the Build console (Agents, Skills, Memory, Jobs, Training, Recognition), and the Operate console (Backends, Activity, Nodes, Usage, Traces, Users, Middleware, Settings)
|
||||
- **Admins also see**: Models, the Build console (Agents, Skills, Memory, Jobs, Training, Recognition), and the Operate console (Backends, Activity, Nodes, Traffic, Settings, and Users and keys)
|
||||
|
||||
Admin-only pages are also protected at the router level - navigating directly to an admin URL redirects non-admin users to the home page.
|
||||
|
||||
### Sign-in, invite and account pages
|
||||
|
||||
The sign-in screen offers only what `GET /api/auth/status` reports:
|
||||
|
||||
- A GitHub or SSO button appears only when that provider is configured. With `LOCALAI_DISABLE_LOCAL_AUTH=true` there is no email form.
|
||||
- Email and password sign-in asks for one field at a time: the email, then the password. Both are sent in one call, so the server still never says whether an email exists.
|
||||
- Registration asks for the email, then the name, password and (when the registration mode allows one) an invite code. A sign-up that waits for approval returns to the sign-in screen with a notice that an admin must approve it.
|
||||
- On a server with no users yet, the first screen is "Create the admin account".
|
||||
- On a server protected only by legacy API keys, the screen asks for the key and nothing else.
|
||||
- An invite link (`/invite/<code>`) opens registration with the code filled in and read-only.
|
||||
|
||||
The **Account** page has four tabs: Profile (name and picture address), Security (change a local password; for a GitHub or SSO account it says the password is managed by the provider), API keys, and Usage (your requests, tokens by model, and the limits an admin set on you, for the last 30 days). **Settings → Users and keys** has the people, the invites and your own API keys.
|
||||
|
||||
### GitHub OAuth Setup
|
||||
|
||||
1. Create a GitHub OAuth App at **Settings → Developer settings → OAuth Apps → New OAuth App**
|
||||
@@ -315,6 +328,10 @@ curl -X PATCH http://localhost:8080/api/auth/api-keys/<key-id> \
|
||||
|
||||
The key list returns `disabled` and, when set, `pausedUntil` for each key. The Account page in the web UI has a Pause and Resume button for each key.
|
||||
|
||||
A key can be created with a lifetime: send `expiresIn` (`30d`, `90d` or `1y`) or `expiresAt` (RFC 3339). Without either, the server applies `LOCALAI_DEFAULT_API_KEY_EXPIRY` when it is set, and the key does not expire otherwise. The full key is returned once, in the response to the create call; the list returns only its prefix, the creation time, `lastUsed` and `expiresAt`.
|
||||
|
||||
In the web UI, the **API keys** tab of the Account page and of **Settings → Users and keys** shows the signed-in person's own keys: a new key is shown once with a Copy button, and the lifetime list sends `expiresIn`. Revoking a key waits ten seconds in the browser before it sends the `DELETE`, and Undo inside that time sends nothing; leaving the page sends a revoke that is still waiting. LocalAI lists a person's keys only to that person, so an admin cannot list or revoke other users' keys, in the UI or through the API. The Usage page can group requests by API key for everyone.
|
||||
|
||||
### Auth API Endpoints
|
||||
|
||||
| Method | Endpoint | Description | Auth Required |
|
||||
|
||||
@@ -16,27 +16,52 @@ For the complete list of backends, the model families they support, and their ac
|
||||
|
||||
## Managing Backends in the UI
|
||||
|
||||
The **Operate → Backends** page is the canonical home for the complete backend
|
||||
lifecycle:
|
||||
The **Operate → Runtime → Backends** page is the canonical home for the
|
||||
complete backend lifecycle. It is one list with two views, **Installed** and
|
||||
**Catalog**, each with a count:
|
||||
|
||||
1. **Catalog** browses configured galleries, searches by name or description,
|
||||
filters by capability, and installs a backend. Catalog is the default view.
|
||||
2. **Installed** shows the runtimes present on the host or cluster. Search and
|
||||
filter by user, system, update, or offline-node state, then select a backend
|
||||
to inspect its version, source, node placement, and lifecycle actions.
|
||||
On a new installation with nothing installed it recommends llama-cpp, which
|
||||
runs most text models, with a single install button.
|
||||
2. **Installed** shows the runtimes present on the host or cluster, those with
|
||||
an update first. Search and filter by user, system, update, or offline-node
|
||||
state.
|
||||
3. Variant and development builds remain opt-in refinements. Target-node links
|
||||
compose with the current view and selection instead of opening a separate
|
||||
management page.
|
||||
|
||||
The current view, search, filter, selected backend, and target node are stored
|
||||
in the URL. Browser Back and shared links therefore restore the same state.
|
||||
Each row shows the backend, its version and its state: **Current**, **Update
|
||||
1.2.0**, **Not installed**, **Queued**, **Removing**, or a progress bar while it
|
||||
installs, with **Cancel** in the row. A failed install says why, with **Retry**.
|
||||
The one action that matters sits in the row (**Install**, **Update**), and the
|
||||
chevron opens the row for the rest. **Update all** starts every pending update,
|
||||
and **Check for updates** asks the server to check now instead of waiting for
|
||||
its next scheduled check. **From URL** installs a backend from an OCI image, a
|
||||
URL or a path.
|
||||
|
||||
The current view, search, filter, open row, and target node are stored in the
|
||||
URL. Browser Back and shared links therefore restore the same state.
|
||||
|
||||
**Removing** a backend asks first. The dialog names the configured models that
|
||||
ask for that backend by name, and any installed meta backend that points at it
|
||||
(for example `llama-cpp` pointing at a hardware-specific build), because those
|
||||
stop working without it. The catalog does not report a size per backend, and
|
||||
LocalAI keeps no earlier version of a backend, so there is no size column and
|
||||
no rollback.
|
||||
|
||||
**Operate → Runtime → Logs** (`/app/backend-logs`) lists the backend processes
|
||||
that have printed something. Open one to read its output live, filter by stream
|
||||
or text, follow the end, show or hide times, export the lines as JSON, or clear
|
||||
them. Clearing hides the lines at once and wipes them on the server after 6
|
||||
seconds, unless you press **Undo**.
|
||||
|
||||
Installs run in the background. The strip at the top of the app follows the
|
||||
current one, and **Operate → Activity** lists everything in flight, what needs
|
||||
attention, and what has finished, and is where a running install is cancelled
|
||||
or a failed one retried. See [Activity]({{% relref "operations/activity" %}}).
|
||||
attention, and what has finished, and is where a running install is paused,
|
||||
cancelled or a failed one retried. See [Activity]({{% relref "operations/activity" %}}).
|
||||
|
||||
Each selected backend displays:
|
||||
Each open row displays:
|
||||
- Backend name and description
|
||||
- Type of models it supports
|
||||
- Installation status
|
||||
|
||||
@@ -0,0 +1,73 @@
|
||||
+++
|
||||
disableToc = false
|
||||
title = "Chat"
|
||||
weight = 48
|
||||
url = "/features/chat/"
|
||||
+++
|
||||
|
||||
Chat is the conversation page of the web UI. It keeps one thread per conversation, in your browser, and sends it to a chat model of your choice. The message box is the same one Home uses, so a message typed on Home opens here with the same model and the same tools.
|
||||
|
||||
## The thread
|
||||
|
||||
Your messages are the raised blocks on the right. The model's replies are plain text under its name, with a dot that is filled when the server holds the model in memory and hollow when it is not loaded yet. Code blocks have **Copy** and **Canvas** buttons. Images you attach appear as thumbnails that open in a viewer, and other files appear as chips.
|
||||
|
||||
Reasoning and tool calls fold into one quiet line (for example "Thought · read_file"). Click it to open the steps with their arguments and results. While a reply is arriving the line is open and shimmers, and it folds when the answer starts.
|
||||
|
||||
If a reply fails, the text written so far is kept, the reason is shown, and **Retry** asks again from your last message. **View traces** opens the backend traces.
|
||||
|
||||
### Actions on a message
|
||||
|
||||
Hover a message, focus it, or look at the last one (on a phone every message shows them): **Copy**, **Edit**, **Regenerate** and **Branch from here**. Edit saves the new text in place and does not send anything. Regenerate drops the answer and everything after it and asks again. Branch starts a new chat with the history up to that answer.
|
||||
|
||||
| Key | Action |
|
||||
|---|---|
|
||||
| `Up`, `Down` | Move between messages (a message must have focus) |
|
||||
| `C`, `E`, `R`, `B` | Copy, edit, regenerate or branch the focused message |
|
||||
| `Up` in an empty box | Edit your last message |
|
||||
| `Esc` | Stop the reply that is streaming; then close the search; then close the canvas |
|
||||
|
||||
## The message box
|
||||
|
||||
The model chip lists the chat models. **Loaded now** holds the models in memory, **Installed** the rest. When the list opens, LocalAI reads the host memory once and asks the server to estimate each listed model at this chat's context size (up to twelve models). A row shows what the model needs and whether it fits: free memory, how much would run on the CPU, or how far over the machine it is. A model with no estimate shows no fit text. The server reports no load time, so none is shown.
|
||||
|
||||
The **MCP** chip opens the server and client tool lists. **Canvas** turns code blocks into cards that open in a side panel. Type `/` for the actions the page has:
|
||||
|
||||
| Action | What it does |
|
||||
|---|---|
|
||||
| `/model` | Open the model list |
|
||||
| `/new` | Start an empty chat on the same model |
|
||||
| `/chats` | Open the conversations list |
|
||||
| `/assistant` | Turn Manage mode on or off (admin only) |
|
||||
| `/canvas` | Turn Canvas on or off |
|
||||
| `/find` | Search this chat |
|
||||
| `/settings` | Open the chat settings |
|
||||
| `/export` | Download the chat as Markdown |
|
||||
| `/clear` | Remove every message, after a confirmation |
|
||||
|
||||
Enter sends, Shift+Enter adds a line. Paste an image to attach it. Under the box, the line shows the speed while a reply streams and the token count of the chat.
|
||||
|
||||
## Conversations
|
||||
|
||||
Press `Ctrl+K` (`Cmd+K`) or **Chats** to open the list of conversations, grouped by day like **Jump back in** on Home. Type to search names and message text. Arrow keys move, `Enter` opens, `F2` renames and `Delete` removes. Each row also offers rename, duplicate, copy and export. Removing a chat hides it and shows an **Undo** toast; the chat is deleted for good when the toast goes away. The name in the header can be renamed with a click.
|
||||
|
||||
The history is stored in this browser (`localai_chats_data`), not on the server.
|
||||
|
||||
## Settings
|
||||
|
||||
The sliders button opens the chat settings. They apply from the next message.
|
||||
|
||||
- **System prompt.** Sent before every message in this chat. Empty means the model's own default.
|
||||
- **Sampling.** Temperature, top P and top K. Each says "model default" until you change it, and has a **Reset**.
|
||||
- **Context window.** The size drives the meter in the header and the warning when the context is nearly full. It is not sent to the model. Admins get it filled in from the model's configuration.
|
||||
- **Behaviour.** Manage mode (admin) and Focus mode, which collapses the app sidebar while a conversation is open.
|
||||
- **Model info** (admin). Backend, model file, context size, threads, GPU layers and an **Edit config** button.
|
||||
|
||||
## Find, canvas and long threads
|
||||
|
||||
`Ctrl+Shift+F` (`Cmd+Shift+F`) searches this chat: it marks the matches in the messages that are loaded in the page, shows "n of m" and steps with `Enter` and `Shift+Enter`. Nothing is sent to the server. When you scroll away from the end of a long thread, **Jump to latest** brings you back.
|
||||
|
||||
With Canvas on, code blocks become cards. The canvas panel opens beside the thread with tabs, a Code and Preview switch for HTML, SVG and Markdown, Copy and Download. On a narrow window it takes the whole page.
|
||||
|
||||
## When there is nothing to chat with
|
||||
|
||||
An empty chat shows a few starters, whether the model is loaded and your recent conversations. With no chat model installed, the page offers the starter models for the hardware, the gallery and import, and keeps what you typed. While a model is loading on a worker or being staged, the reply waits behind a card that names the phase the server reports, the node, the bytes and the time left when the server gives them, and then sends by itself. Stop cancels the wait.
|
||||
@@ -559,9 +559,17 @@ When the SmartRouter needs to free capacity, it can unload models with zero in-f
|
||||
|
||||
### Managing nodes in the WebUI
|
||||
|
||||
Open **Operate → Nodes** to inspect fleet health, filter or select workers, and view running models across the cluster. The **Running models** view groups replicas by model. Its **View logs…** action opens logs directly when there is one placement; when a model has several placements, it opens the model inspector so you can choose all logs for one node or the logs for one replica.
|
||||
With distributed mode on, **Operate → Swarm** holds the cluster: **Nodes**, **Placement rules**, **Failover** and **P2P**. A single-node install shows **This machine** instead and never draws these pages.
|
||||
|
||||
Open a node's full details for node-scoped work: viewing replica logs, unloading a model, managing installed backends, changing replica capacity, or editing scheduling labels. Diagnostic actions are listed before destructive actions in row menus.
|
||||
**Nodes** lists every worker in a sortable table: name, role, state in words, GPU or system memory, loaded models, last heartbeat and version. Switch between comfortable and compact rows, between **List**, **Map** and **Running models**, and filter to **Needs attention** (waiting for approval, not answering, or low GPU memory, system memory or models disk). The **Map** draws this instance, the message bus and database, and each worker; a dashed line is a worker that gets no traffic. It is not drawn on a phone. When a backend has a newer version, an **Update** button on the page sends the upgrade to the nodes that differ from the rest of the cluster, or to the nodes you selected.
|
||||
|
||||
The **Running models** view groups replicas by model. Its **View logs…** action opens logs directly when there is one placement. When a model has several, it opens the row so you can choose the logs of one replica.
|
||||
|
||||
**Add a node** (`/app/nodes/add`) explains how a machine joins: a registered worker, a peer instance, or a memory shard. It prints the command to run on the new machine with a **Copy** button, and updates the page when the machine appears. On a single-node install it starts with the command that turns distributed mode on.
|
||||
|
||||
Open a node for node-scoped work. The page shows its state, VRAM, RAM, models disk, CPU and in-flight requests, and has tabs for **Models** (replica logs, unload), **Backends** (upgrade, delete), **Logs**, and **Capacity and labels** (replica capacity, labels). **Drain…** shows what the drain would change before it does anything; **Remove…** asks for the node's name. A node that stopped answering says so and shows the last figures it reported.
|
||||
|
||||
The "what happens if I drain this node" list is a preview worked out in the browser from the node list, the loaded replicas and the placement rules. The server does not compute it, and the scheduler also weighs free memory and disk when it loads a model, so the preview never claims a model will fit.
|
||||
|
||||
## Node Management API
|
||||
|
||||
@@ -601,17 +609,11 @@ Used by the WebUI and admin API consumers. Requires admin authentication.
|
||||
| `PUT` | `/api/nodes/:id/vram-budget` | Set a VRAM budget for a worker (`{"value":"80%"}`) |
|
||||
| `DELETE` | `/api/nodes/:id/vram-budget` | Clear a worker's VRAM budget (revert to all detected VRAM) |
|
||||
|
||||
The **Nodes** page in the React WebUI is a fleet operations dashboard. Its health band and VRAM, RAM, CPU, and models-disk gauges aggregate the single `GET /api/nodes` response and identify how many workers do not report each metric. The attention queue isolates pending, impaired, or low-capacity workers without double-counting the headline affected-node total.
|
||||
The **Nodes** page reads `GET /api/nodes` every five seconds and renders 50 workers at a time. Bulk drain, resume and remove run with bounded concurrency, so the page stays usable for fleets with thousands of registrations. Selecting the visible page or a group does not discard selections elsewhere; selections are removed only when a later poll confirms the worker no longer exists.
|
||||
|
||||
The fleet table supports search, status and type filters, label or type grouping, sortable columns, and selection across filters. It renders 50 workers at a time and bulk drain, resume, and remove operations run with bounded concurrency, so the page remains usable for fleets with thousands of registrations. Selecting the visible page or a group does not discard selections elsewhere; selections are removed only when a later poll confirms the worker no longer exists.
|
||||
The list never fetches backend inventory, and the **Running models** and **Map** views stay lazy: the first use of either makes one controller database request that is kept until the page is left. A node's backends are read only when its page opens.
|
||||
|
||||
Selecting a row opens an in-context inspector with health, labels, capacity, model activity, and heartbeat details. Backend inventory is fetched only for the open inspector. The inspector links to the dedicated node detail page at `/app/nodes/:id`, where model, backend, label, capacity, CPU utilization and load, and models-disk management remain available. Model scheduling lives on its own **Scheduling** page.
|
||||
|
||||
The workbench's **Running models** tab shows the current loaded replicas on healthy workers. It stays lazy: opening the Nodes page does not query model inventory, and the first activation makes one controller database request that is retained until the page is left. The view groups replicas by model, reports their worker spread, active requests, backend types, and most recent use, and renders 50 models per page for large fleets. Loading, empty, and query-failure states are shown in place; a failed query can be retried.
|
||||
|
||||
Use a model row's actions menu to stop that model across the fleet. LocalAI sends one controller shutdown request for the model, which stops all loaded placements; the browser does not contact workers individually. The dashboard refreshes the running-model inventory after both successful and failed shutdown attempts because a failed request can still have stopped some replicas.
|
||||
|
||||
Opening a model reveals its replica placement without another request. Replicas on the same worker remain individually visible with their process addresses and workload. From there, select a known worker to move into its node inspector, then return to the model with **Back to model**. That worker transition is the only point in this flow that requests backend inventory, preserving the Nodes page's no-prefetch behavior.
|
||||
Use a model row's actions menu in **Running models** to stop that model across the fleet. LocalAI sends one controller shutdown request for the model, which stops all loaded placements; the browser does not contact workers individually. The view refreshes the running-model inventory after both successful and failed shutdown attempts because a failed request can still have stopped some replicas.
|
||||
|
||||
### Model sizing in the WebUI
|
||||
|
||||
@@ -790,7 +792,7 @@ curl -X POST http://frontend:8080/api/nodes/<node-id>/approve \
|
||||
-H "Authorization: Bearer <admin-token>"
|
||||
```
|
||||
|
||||
The **Nodes** page in the WebUI also shows pending nodes with an **Approve** button.
|
||||
The **Nodes** page in the WebUI shows pending nodes with an **Approve** button, and so does the node's own page and the **Add a node** page once the machine has registered.
|
||||
|
||||
To skip manual approval and let nodes join immediately, set `--auto-approve-nodes` (or `LOCALAI_AUTO_APPROVE_NODES=true`) on the frontend. This is convenient for development and trusted environments.
|
||||
|
||||
@@ -1136,7 +1138,7 @@ local-ai worker \
|
||||
|
||||
## Model Scheduling
|
||||
|
||||
Model scheduling controls where models are placed and how many replicas are maintained. In the React WebUI it has its own **Scheduling** page (a top-level nav item, separate from the Nodes page). It combines two optional features:
|
||||
Model scheduling controls where models are placed and how many replicas are maintained. In the React WebUI it has its own **Placement rules** page (**Operate → Swarm → Placement rules**, at `/app/scheduling`). Each rule is written as a sentence, shows the nodes the model is loaded on now, and opens in a side sheet that previews which nodes the draft rule could use. The preview is worked out in the browser from the node list and labels; the scheduler also checks free memory and disk when it loads a model. A deleted rule can be taken back for a few seconds. A rule combines two optional features:
|
||||
|
||||
### Node Selectors
|
||||
|
||||
@@ -1210,7 +1212,7 @@ curl -X POST http://frontend:8080/api/nodes/scheduling \
|
||||
|
||||
This makes an alias a stable deployment slot: the placement policy belongs to
|
||||
the slot, and the model filling it can change without rewriting the rule. The
|
||||
WebUI lists aliases in the model picker on the **Scheduling** page, tagged with
|
||||
WebUI lists aliases in the model picker on the **Placement rules** page, tagged with
|
||||
the model each one resolves to.
|
||||
|
||||
Each frontend resolves the alias from its own copy of the model configs, and a
|
||||
|
||||
@@ -33,11 +33,13 @@ LocalAI supports two modes of distributed inferencing via p2p:
|
||||
|
||||
A list of global instances shared by the community is available at [explorer.localai.io](https://explorer.localai.io).
|
||||
|
||||
A LocalAI server started with `local-ai explorer` serves the same list at `/explorer`. Each swarm shows its name, description, the types of its clusters, how many workers are online and a shortened token. **How to join** shows the whole token and, for a federated cluster, the Docker and command-line commands that start a node on it. **List a swarm** adds a swarm to the list. Listing publishes its token, so anyone who sees the list can use the swarm's workers. Only swarms with at least one online worker are listed. On a server that is not in explorer mode, the page says so.
|
||||
|
||||
## Usage
|
||||
|
||||
Starting LocalAI with `--p2p` generates a shared token for connecting multiple instances: and that's all you need to create AI clusters, eliminating the need for intricate network setups.
|
||||
|
||||
Simply navigate to the "Swarm" section in the WebUI and follow the on-screen instructions.
|
||||
Navigate to **Operate → Swarm → P2P** in the WebUI, or open **Add a node** and choose a peer instance or a memory shard, and follow the on-screen instructions.
|
||||
|
||||
For fully shared instances, initiate LocalAI with --p2p --federated and adhere to the Swarm section's guidance. This feature, while still experimental, offers a tech preview quality experience.
|
||||
|
||||
@@ -63,7 +65,7 @@ local-ai federated
|
||||
|
||||
To see all the available options, run `local-ai federated --help`.
|
||||
|
||||
The instructions are displayed in the "Swarm" section of the WebUI, guiding you through the process of connecting multiple instances.
|
||||
The instructions are on the **P2P** page of the Swarm hub and in **Add a node**, guiding you through the process of connecting multiple instances.
|
||||
|
||||
### Workers mode
|
||||
|
||||
@@ -79,7 +81,7 @@ To connect multiple workers to a single LocalAI instance, start first a server i
|
||||
local-ai run --p2p
|
||||
```
|
||||
|
||||
And navigate the WebUI to the "Swarm" section to see the instructions to connect multiple workers to the network.
|
||||
And open **Operate → Swarm → P2P** to see the network token and the instructions to connect multiple workers to the network.
|
||||
|
||||

|
||||
|
||||
|
||||
@@ -396,6 +396,12 @@ The recommended default `threshold` for `/v1/face/verify` and
|
||||
Pass `threshold` explicitly when switching engines - the per-engine
|
||||
default only fires when the field is omitted.
|
||||
|
||||
## The WebUI page
|
||||
|
||||
**Build → Faces** follows the same layout as the Voices page. **Who is this** looks up a photo against the people you enrolled (`POST /v1/face/identify`, cut-off 0.35 by default, adjustable on the scale). **Same person?** compares two photos (`POST /v1/face/verify`), optionally with the liveness check, and draws the face the model found in each photo. **Enrol a person** takes a photo, a name, optional labels and a permission tick; a copy of the photo stays in the browser only if you tick it.
|
||||
|
||||
The people list is kept in the browser because the server has no list call. After a search, a saved person the server did not return is marked "not on the server", and a server restart empties the server's index. Removing a person waits ten seconds behind an Undo toast before `POST /v1/face/forget` is sent. The server matches one face per photo. Detecting faces, attribute guesses (off by default, often wrong) and the raw embedding are under **More tools**. With no face model installed the page says what is missing and offers gallery models to install.
|
||||
|
||||
## Related features
|
||||
|
||||
- [Object Detection](/features/object-detection/) - generic bounding-box
|
||||
|
||||
@@ -203,12 +203,16 @@ curl -X POST http://localhost:8080/api/fine-tuning/jobs \
|
||||
|
||||
## Web UI
|
||||
|
||||
When fine-tuning is enabled, a "Fine-Tune" page appears in the sidebar under the Agents section. The UI provides:
|
||||
**Build → Fine-Tune** is one page that follows the job from set-up to result. A line at the top shows where you are: set up, check, run, result.
|
||||
|
||||
1. **Job Configuration** - Select backend, model, training method, adapter type, and hyperparameters
|
||||
2. **Dataset Upload** - Upload local datasets or reference HuggingFace datasets
|
||||
3. **Training Monitor** - Real-time loss chart, progress bar, metrics display
|
||||
4. **Export** - Export trained models in various formats
|
||||
1. **Set up** - Choose the base model, the dataset (a Hugging Face id or an uploaded file), the kind of training (an adapter, or the full model) and the epochs, batch size and learning rate. Method, backend, adapter settings, optimizer, evaluation, reward functions (GRPO) and extra options are under **More options**.
|
||||
2. **Check before you start** - A list of what is known before the job starts, redrawn as you change the form: whether the model and dataset are set, whether a fine-tuning backend is installed, how much GPU memory (or RAM, when there is no GPU) is free, how much space is free on the models disk, and whether a token is set. LocalAI does not estimate how much memory a job needs or how long it takes, because the server cannot know that before the job starts, so the page says so. A missing model or dataset blocks the start. A warning, such as a missing backend or no GPU, does not.
|
||||
3. **Run** - Percent, step, epoch, the server's time estimate and tokens per second, the stages the job passes through, a chart of loss (with the evaluation loss as hollow dots), learning rate and gradient norm, and a log of what the job reported since the page opened. **Stop** asks whether to keep a checkpoint.
|
||||
4. **Result** - A failed job shows the server's message. When the message says memory ran out, the page offers two changes (a batch size of 1 and gradient checkpointing) that are applied to a copy of the setup and start nothing until you press Start. A finished job lists its checkpoints and exports the result as a model, then links to a chat with it and to Models.
|
||||
|
||||
Earlier jobs are listed below the form. **Reuse** puts a job's setup back in the form.
|
||||
|
||||
A user needs the fine-tuning permission. Without it, the page says the account cannot fine-tune and who can change that.
|
||||
|
||||
## Dataset Formats
|
||||
|
||||
|
||||
@@ -13,9 +13,9 @@ The same MCP server is published as a Go package and can also be served over **s
|
||||
|
||||
## Enabling the assistant in chat
|
||||
|
||||
Open the chat UI as an **admin** user and pick a chat-capable model in the model selector. The header shows a **Manage** toggle - flip it on, and a `Manage mode` badge appears next to the chat title. Starter chips ("What is installed?", "Install a chat model", "Show system status", "Update a backend") help you get going.
|
||||
Open the chat UI as an **admin** user and pick a chat-capable model with the model chip above the message box. Open **Chat settings** (the sliders button in the header) and turn on **Manage mode**, or type `/assistant` in the message box. A shield icon appears next to the chat title while the mode is on. In a new chat in Manage mode, starter chips ("What is installed?", "Install a chat model", "Show system status", "Update a backend") help you get going.
|
||||
|
||||
The home page also exposes a **Manage by chat** CTA that opens a fresh chat already in Manage mode.
|
||||
The home page shows a one-line **Manage LocalAI by chatting** prompt that opens a fresh chat already in Manage mode. You can dismiss it. The same action stays available as `/assistant` in the command bar and in the Library row.
|
||||
|
||||
Once on, try:
|
||||
|
||||
|
||||
@@ -562,7 +562,7 @@ In addition to server-side MCP (where the backend connects to MCP servers), Loca
|
||||
|
||||
### How It Works
|
||||
|
||||
1. **Add servers in the UI**: Click **MCP** in the chat header, open the **Client** tab, and add MCP server URLs
|
||||
1. **Add servers in the UI**: Click the **MCP** chip above the message box, open the **Client** tab, and add MCP server URLs
|
||||
2. **Browser connects directly**: The browser uses the MCP TypeScript SDK (`StreamableHTTPClientTransport` or `SSEClientTransport`) to connect to MCP servers
|
||||
3. **Tool discovery**: Connected servers' tools are sent as `tools` in the chat request body
|
||||
4. **Browser-side execution**: When the LLM calls a client-side tool, the browser executes it against the MCP server and sends the result back in a follow-up request
|
||||
|
||||
@@ -246,11 +246,16 @@ continue. On a single LocalAI instance, a restart removes the pin. In
|
||||
and a table of every target with its kind, warm flag, status, last probe
|
||||
time and last error. An admin sees a **Pin** button on each target and an
|
||||
**Unpin** action for the chain, both behind a confirmation dialog.
|
||||
- The **Failover** page (`/app/failover`, admin only, linked from the
|
||||
console navigation) lists every chain with its status pill, active target,
|
||||
a small pill per target, and time since the last switch. It links each
|
||||
chain name to its model editor page and shows an empty state linking to
|
||||
the failover template when no chains exist yet.
|
||||
- The **Failover** page (`/app/failover`, admin only) lists every chain with
|
||||
its status pill, active target, a small pill per target, and time since the
|
||||
last switch. It links each chain name to its model editor page and shows an
|
||||
empty state linking to the failover template when no chains exist yet. It
|
||||
sits under **Operate → Runtime** on a single install and under **Operate →
|
||||
Swarm** when distributed mode is on. On a cluster it also describes what the
|
||||
router and the health monitor do when a worker stops answering, and shows a
|
||||
preview of what would stop if a chosen node went away, worked out in the
|
||||
browser from the loaded replicas and the placement rules. The server does not
|
||||
compute that preview.
|
||||
- The Installed Models list badges a model that belongs to a chain with
|
||||
`chain → <active target>`, next to the alias badge.
|
||||
- All of the above update live from the same event stream as
|
||||
|
||||
@@ -28,16 +28,110 @@ GPT and text generation models might have a license which is not permissive for
|
||||
Open **Models** in the WebUI. It is the canonical page for a model's complete
|
||||
lifecycle and has two views:
|
||||
|
||||
- **Explore** browses configured galleries, compares hardware fit and variants,
|
||||
and installs models. This is the default view.
|
||||
- **Installed** lists local model configurations and their running, idle,
|
||||
disabled, pinned, and distributed state. Select a model to load or stop it,
|
||||
edit its configuration, open a supported use case, inspect backend logs, or
|
||||
remove it.
|
||||
- **Explore** browses configured galleries and installs models. It is one
|
||||
dense table: each row shows the model's size, a bar for the memory it needs at
|
||||
the chosen context length, and the headroom in words ("3.7 free", "+1.5 on
|
||||
CPU", "0.9 over"). Capability chips show how many models each one matches.
|
||||
Select a row to open the details beside the table: the fit on this machine,
|
||||
VRAM by context length, variants, files, links, tags and licence. This is the
|
||||
default view.
|
||||
- **Installed** lists local model configurations in the same table, with their
|
||||
running, idle, disabled, pinned, and distributed state and their size on disk.
|
||||
Load or stop a model from its row, or select it to edit its configuration,
|
||||
open a supported use case, inspect backend logs, or remove it.
|
||||
|
||||
Both views use the same model selection and store the view, search, filter, and
|
||||
selection in the URL. Installing from Explore does not move you away from the
|
||||
catalog; the entry updates in place when the operation finishes.
|
||||
Both views store the view, search, filter, and selection in the URL. Installing
|
||||
from Explore does not move you away from the catalog; the entry updates in place
|
||||
when the operation finishes.
|
||||
|
||||
### A model's own page
|
||||
|
||||
Every model also has a page of its own at `/app/models/<name>`, so it can be
|
||||
linked. Open it with the arrow at the end of a row, a double click on the row,
|
||||
or `o` on the selected row. On a phone, a tap on a row opens it. The details
|
||||
beside the table stay as the quick look.
|
||||
|
||||
The title block names the model and holds the main action: **Install**, with a
|
||||
chevron to choose which build to install, or **Load** and **Stop** with a menu
|
||||
for an installed model (disable, pin, edit configuration, logs, delete). A strip
|
||||
under it answers three questions: whether the model fits this machine, what it
|
||||
does, and what installing leaves free on the models disk (or its state, when it
|
||||
is installed). The tabs are:
|
||||
|
||||
- **Overview**: the description, backend, licence, largest context, tags, links
|
||||
and a memory bar. For an installed model it also shows its state, the pages it
|
||||
opens in, and which agents, agent tasks, failover chains and aliases name it.
|
||||
- **Fit and memory**: the verdict in words, a context size selector, the memory
|
||||
bar split into weights and the part that grows with context, and a chart of the
|
||||
memory needed at each context size against the memory this machine offers.
|
||||
When a model has several builds, pick the build to see its own figures.
|
||||
- **Variants and files**: the builds with their size and fit, and the files the
|
||||
chosen build downloads. Install any build from its row.
|
||||
- **Usage and history**, **Configuration** and **Logs**, for installed models.
|
||||
LocalAI does not record requests, timings, loads or configuration changes for
|
||||
each model, so the usage tab lists what it cannot show yet instead of an empty
|
||||
chart. The same tab lists the files the model uses on disk, with their size,
|
||||
the other models that use them, and any file the configuration names that is
|
||||
not on disk (see [Disk and cleanup](#disk-and-cleanup)). Configuration holds the [Placement](/advanced/model-configuration/#placement)
|
||||
section. Logs is the backend log viewer of the Operate section.
|
||||
|
||||
The page reads the same lists as the table, so it works for a model the gallery
|
||||
does not list (it has no variants or files tab) and for a gallery model that is
|
||||
not installed (it has no usage, configuration or logs tab). If the gallery cannot
|
||||
be reached, an installed model keeps working and Install says why it is off.
|
||||
|
||||
Going back with `Esc`, `Backspace` or the **Models** button returns to the list
|
||||
with its view, search, filters, selection and scroll as you left them. The
|
||||
previous and next buttons step through the rows of the list you came from, in
|
||||
the order you saw them.
|
||||
|
||||
### Keyboard
|
||||
|
||||
On the Models page, `/` jumps to the search field, the up and down arrows move
|
||||
the selection, `Enter` installs the selected model in Explore, `d` switches
|
||||
between comfortable and compact rows, `o` opens the selected model's page, and
|
||||
`Esc` closes the details.
|
||||
|
||||
On a model's page, `1` to `6` switch tabs, `[` and `]` (or `k` and `j`) step to
|
||||
the previous and next model of the list, and `Esc` goes back to the list.
|
||||
|
||||
### Disk and cleanup
|
||||
|
||||
When the server reports the disk that holds the models directory, a strip in the
|
||||
page header shows how much of it is free. It turns amber when less than 10
|
||||
percent, or less than 20 GB, is free. In Explore, the details of a model say how
|
||||
much disk an install leaves free. The strip is hidden when the disk cannot be
|
||||
read, and on a distributed controller, where the models live on the workers.
|
||||
|
||||
Select the strip to open the cleanup review. LocalAI does not record when a
|
||||
model was last used or how often, so the review says so and ranks installed
|
||||
models only by what it can see:
|
||||
|
||||
- **Safe to remove**: another build of the same gallery model is installed, and
|
||||
the build LocalAI would pick on this host is the one that stays.
|
||||
- **Probably safe**: disabled, unused, and available in the gallery to download
|
||||
again.
|
||||
- **Your call**: not loaded and not used by anything, but with nothing more
|
||||
known. A model that is not in the gallery cannot be downloaded again, and the
|
||||
review says so.
|
||||
- **Protected**: loaded, pinned, or named by an agent, an agent task, a failover
|
||||
chain or an alias. These are never suggested. If an agent or task cannot be
|
||||
read, nothing is marked safe.
|
||||
|
||||
For an admin, sizes are what each model uses on disk, read from
|
||||
`GET /api/models/storage`. A file that another installed model also uses stays
|
||||
on disk when you remove one of the two, so the review counts only the files a
|
||||
model does not share. It says which models share files, and the amount freed
|
||||
by a selection counts a shared file only when every model that uses it is in
|
||||
the selection. A configuration that names files that are not on disk (a download
|
||||
that did not finish, or files removed by hand) is listed as a finding. If the
|
||||
report cannot be read, for example because the user is not an admin, sizes fall
|
||||
back to the sizes of the files the gallery lists. A model that is not in the
|
||||
gallery then shows no size and is not counted in what a removal frees. Before you
|
||||
confirm, the review checks again and lists what will go, why, and how much it
|
||||
frees. Removal then waits 30 seconds, during which you can undo it; the delete
|
||||
request is sent only when that time ends. If you leave the page during the wait,
|
||||
nothing is deleted.
|
||||
|
||||
## Cyber-Ornith 1.5 9B
|
||||
|
||||
@@ -171,9 +265,9 @@ This removal does not delete previously installed models. Remove that configurat
|
||||
|
||||
When browsing the gallery or importing a model by URI, LocalAI can show **estimated download size** and **estimated VRAM** for models.
|
||||
|
||||
- **Where they appear**: In the model gallery table (Size / VRAM column), in the model detail modal, and after starting an import from URI (in the success message).
|
||||
- **Where they appear**: In the model gallery table (Size and Fit columns), in the model inspector beside it, and after starting an import from URI (in the success message).
|
||||
- **How they are computed**: GGUF models use file size (HTTP HEAD or local stat) and optional GGUF metadata (HTTP Range) for KV cache and overhead; other formats use Hugging Face file sizes and optional config when available. If metadata is unavailable, a size-only heuristic is used. GGUF metadata lengths that exceed the file size are rejected before allocation; these files also use the size-only estimate.
|
||||
- **Hardware fit indicator**: When your system reports GPU or RAM capacity, the gallery shows whether the estimated VRAM fits (green) or may not fit (red) using a 95% headroom rule.
|
||||
- **Hardware fit indicator**: When your system reports GPU or RAM capacity, each row shows whether the estimated memory at the chosen context length fits, using a 95% headroom rule. A model that is too big for the GPU but would run from system RAM is marked as spilling to the CPU, with how much; a model too big for both is marked as over by the shortfall.
|
||||
- Estimates are best-effort and may be missing if the server does not support HEAD/Range or the request times out.
|
||||
|
||||
## Useful Links and resources
|
||||
|
||||
@@ -645,3 +645,19 @@ This event is a LocalAI extension to the OpenAI Realtime API and is server-emitt
|
||||
|
||||
- [Realtime voice assistant demo (Go)](https://github.com/localai-org/localai-realtime-demo): a minimal Go client for the Realtime (WebSocket) API with a full talk-back voice loop and an example tool call. Ships a `docker compose` setup that brings up a realtime-capable LocalAI for you.
|
||||
- [Realtime voice assistant example (Python)](https://github.com/mudler/LocalAI-examples/tree/main/realtime): thin-client architecture (Silero VAD on the client, heavy lifting on LocalAI), suited to running the client on a Raspberry Pi.
|
||||
|
||||
## Talk in the web UI
|
||||
|
||||
The **Talk** page of the web UI is a client for this API over WebRTC. Pick a pipeline model with the chip at the top, then press **Start session**. The microphone streams while the session runs. The server detects when you stop talking and answers on its own, and you can speak over a reply to interrupt it. The page has no push-to-talk or hands-free switch.
|
||||
|
||||
The heading under the orb says what is happening: connecting, listening, thinking (also while a tool runs), speaking, or interrupted after you cut a reply off. The orb follows the real microphone level while it listens and the reply's audio while it speaks. These states have their own screen:
|
||||
|
||||
- **Microphone blocked.** The browser did not give the page the microphone. Allow it in the address bar and press **Try again**.
|
||||
- **Connection lost.** The WebRTC link failed during a session. The transcript stays and **Reconnect** starts a new session.
|
||||
- **Something went wrong.** The server reported an error or the call could not be set up. The reason is shown, with a link to the traces.
|
||||
- **Talk needs a pipeline model.** No pipeline model exists yet. The page links to the model editor with the pipeline template and to the gallery.
|
||||
|
||||
The transcript shows what you said, the replies and, in Manage mode, the tool calls and results. **Copy** puts it on the clipboard. It is not saved.
|
||||
|
||||
The sliders button opens the session settings: the instructions, the voice, the transcription language, client-side MCP tools, Manage mode (admin, fixed once a session is open) and the parts of the selected pipeline. The gauge button, shown during a session, adds audio diagnostics (waveform, spectrum and WebRTC statistics).
|
||||
|
||||
@@ -142,12 +142,16 @@ The UI also supports entering a custom quantization type string for any format s
|
||||
|
||||
## Web UI
|
||||
|
||||
A "Quantize" page appears in the sidebar under the Tools section. The UI provides:
|
||||
**Build → Quantize** uses the same page pattern as fine-tuning: set up, check, run, result.
|
||||
|
||||
1. **Job Configuration** - Select model, quantization type (dropdown with presets or custom input), backend, and HuggingFace token
|
||||
2. **Progress Monitor** - Real-time progress bar and log output via SSE
|
||||
3. **Jobs List** - View all quantization jobs with status, stop/delete actions
|
||||
4. **Output** - Download the quantized GGUF file or import it directly into LocalAI for immediate use
|
||||
1. **Set up** - Choose the model and the quantization type (a preset, or a custom type). The backend and the Hugging Face token are under **More options**.
|
||||
2. **Check before you start** - Whether a model and type are set, whether a quantization backend is installed, the free RAM and the free space on the models disk. The server reports no model size before the job starts, so the page does not estimate one.
|
||||
3. **Run** - Percent, the stage (downloading, converting, quantizing), and a log of what the job reported since the page opened. **Stop** ends the job.
|
||||
4. **Result** - A failed job shows the server's message. A finished job shows the output file, downloads it, or imports it into LocalAI under a name you choose, then links to a chat with it and to Models.
|
||||
|
||||
Earlier jobs are listed below the form, with **Reuse** and **Delete** (after a confirmation).
|
||||
|
||||
A user needs the quantization permission. Without it, the page says the account cannot quantize.
|
||||
|
||||
## Architecture
|
||||
|
||||
|
||||
@@ -9,7 +9,38 @@ LocalAI provides a web-based interface for managing application settings at runt
|
||||
|
||||
## Accessing Runtime Settings
|
||||
|
||||
Navigate to the **Settings** page from the management interface at `http://localhost:8080/manage`. The settings page provides a comprehensive interface for configuring various aspects of LocalAI.
|
||||
Open **Operate → Settings** in the web UI (`/app/settings`). The page lists every setting below under eight groups by intent, and keeps your edits in a pending bar until you apply them.
|
||||
|
||||
### Groups
|
||||
|
||||
The groups reorganise the fields the page used to show under fifteen sections. The mapping:
|
||||
|
||||
| Group | Settings it holds | Used to be under |
|
||||
|---|---|---|
|
||||
| Memory and models | Watchdog (idle and busy checks, timeouts, interval), eviction (force when busy, largest first, retries, retry interval), free memory automatically and its threshold, models kept loaded, GPU memory budget | Watchdog, Memory Reclaimer, Backend Management, Performance |
|
||||
| Speed and defaults | Default threads, default context size, artifact download concurrency, F16 | Performance |
|
||||
| Backends and galleries | Automatic backend upgrades, development backends, gallery loading on boot, the persistent VRAM cache, the model and backend gallery lists | Backend Management, Galleries |
|
||||
| Access and security | CORS and allowed origins, CSRF protection, shared API keys | API & CORS, API Keys |
|
||||
| Debugging and traces | Verbose debug logging, API traces and their limits, backend logging | Performance, Tracing |
|
||||
| Agents and responses | Agent job history, the agent pool, the LocalAI Assistant, the response store TTL | Agent Jobs, Agent Pool, LocalAI Assistant, Open Responses |
|
||||
| Swarm and sharing | P2P token, network ID and federated mode, the disk headroom check | P2P Network, Distributed |
|
||||
| Look and feel | Instance name, tagline, the three logos | Branding |
|
||||
|
||||
Search at the top of the page matches a setting's name, description and key, the group, and the old section name, so a search for "Watchdog" still finds the watchdog settings and each result says where it used to be.
|
||||
|
||||
### Editing, applying and undoing
|
||||
|
||||
Edits are not sent as you type. A bar at the bottom of the page counts them and offers **Discard**, **Show diff** and **Apply**. The diff lists each old and new value, and runs the checks the browser can make: a duration parses the way the server parses it, a GPU memory budget is one the server accepts, a gallery list is valid JSON with a `url` in each entry, and warnings repeat what the server says (a restart is needed, an empty P2P token stops P2P, evicting while busy can interrupt requests, CSRF protection off). A check that would make the server refuse the value stops **Apply**.
|
||||
|
||||
**Apply** sends only the settings that changed. After it, **Undo** (for ten seconds) saves the previous values again. This is a new save, not a rollback: anything that changed in between stays changed.
|
||||
|
||||
A setting shows **Changed** and its built-in default, with a **Reset** link, only when the default is known (the defaults of the `local-ai run` flags). Settings without a stated default, such as threads and the context size, never show the marker. An environment variable or flag can change the value the server starts with, so the marker compares with the built-in default, not with the startup value.
|
||||
|
||||
A row says **Applies now** when the save handler applies the setting immediately (the watchdog and the memory reclaimer restart with the new values; so do the P2P stack and the agent job service when their settings change), and **Needs restart** for the agent pool settings. A setting without either note is saved, and the code does not say when it takes effect.
|
||||
|
||||
**History** lists the settings you applied from this browser, newest first, up to the last 50. LocalAI keeps no settings log of its own, so changes made elsewhere do not appear, and a secret (the P2P token, shared API keys, the agent database URL) is listed as changed without its value. **Revert** puts the old value back as an edit.
|
||||
|
||||
`POST /api/settings` accepts a partial body and merges it over the saved settings, so the page sends only the keys it changes. In the request, the `csrf` field carries the **disable** flag: the page shows the inverse as "CSRF protection". The galleries are sent as `galleries` and `backend_galleries`, and the shared API keys as `api_keys`, each as a JSON list.
|
||||
|
||||
## Available Settings
|
||||
|
||||
|
||||
@@ -0,0 +1,74 @@
|
||||
+++
|
||||
disableToc = false
|
||||
title = "Studio"
|
||||
weight = 49
|
||||
url = "/features/studio/"
|
||||
+++
|
||||
|
||||
Studio is the part of the web UI where you make images, video, 3D objects, speech, sound and transformed audio, and find what you made before. It opens on a front page with a prompt box and your own work. Each type still has its own workspace page (Images, Video, 3D, TTS, Sound, Transform, Diarization) where a run happens.
|
||||
|
||||
## Make something
|
||||
|
||||
Type what you want in the box under **What do you want to make?**, or pick a type first. The box suggests a type from your words ("Sounds like Video") and never switches by itself. The suggestion is a short list of keywords in the browser. No model is called.
|
||||
|
||||
| Key | Action |
|
||||
|---|---|
|
||||
| `/` | Focus the prompt box |
|
||||
| `Alt+1` to `Alt+7` | Pick a type, in the order of the chips |
|
||||
| `Ctrl+Enter` (`Cmd+Enter`) | Generate |
|
||||
| `Alt+Enter` | Take the suggested type |
|
||||
| `Esc` | Close the details in a lineage view |
|
||||
|
||||
**Generate** opens the workspace for the chosen type with the prompt, the model, and the options that workspace has already filled in: size and count for Images, size for Video. The run itself happens on the workspace page, as before. The front page does not run anything.
|
||||
|
||||
3D, Transform and Diarization start from a file. Use the **Start from** list to pick one of your earlier results as the file, or add a file on the next page.
|
||||
|
||||
### A type with no model
|
||||
|
||||
A type with no installed model is a dashed chip. Picking it shows one model from the gallery, its download size and memory need where the gallery knows them, the memory free now, and an **Install** button. Nothing installs until you press it, and the words you typed stay in the box while it installs. When the gallery gives no size, the note says so.
|
||||
|
||||
## Your work
|
||||
|
||||
**Your work** lists the results this browser has a record of, newest first, with filters for each type and counts. Each tile shows a thumbnail (a waveform drawing for audio, one bar per speaker for diarization), the prompt, the model and the age. The star marks a favourite.
|
||||
|
||||
Results made from each other stack into one project tile. Open a tile to see the lineage.
|
||||
|
||||
### Where the history is kept
|
||||
|
||||
The history lives in your browser, not on the server. Each workspace writes its own list when a result is produced: images, video, TTS, sound and audio transform in browser storage (up to 100 entries each), 3D in IndexedDB (up to 20 entries, with the model file), and diarization in browser storage with the file name, the model and the speaker count only. It never keeps the recording or the transcript. A prompt longer than 2000 characters is cut when it is stored. Favourites are a list of up to 500 ids.
|
||||
|
||||
Results made before this version have no link to the result they came from, so they appear as single tiles. The files themselves are on the server; if the server has cleaned its output folder, the tile shows the type icon instead of a picture.
|
||||
|
||||
**Clear history** removes all of these lists and the favourites from this browser after a confirmation. It does not delete files on the server.
|
||||
|
||||
## Lineage
|
||||
|
||||
When a workspace is opened from a result (for example **Animate** on a picture), the new result records the id of the result it came from and how: take, animate, to 3D, variation, transform or who spoke. The lineage view draws those links as a board: prompts or files on the left, results next to them, the path through the selected result drawn heavier, and one dashed suggested next step chosen from the installed models.
|
||||
|
||||
The dock under the board acts on the selected result:
|
||||
|
||||
- **Run as a new take** opens the same workspace with the prompt, model and size filled in. It is disabled when the original input was not kept (3D, diarization).
|
||||
- **Branch from here** opens a draft: choose the next step, edit the words, and **Open in** the workspace with the result as its starting point.
|
||||
|
||||
A step is listed but disabled, with the reason, when the destination page cannot start from that kind of result yet. Today a video cannot start a sound and a 3D object cannot start a clip.
|
||||
|
||||
Arrow keys walk the board, `B` branches and `Esc` closes the draft, then the details, then the view.
|
||||
|
||||
## The workspace page
|
||||
|
||||
All seven workspaces share one layout, under a row of tabs, one per type:
|
||||
|
||||
- **Compose card.** Optional sources as dashed chips (a start image, reference images, an end image, avatar audio), or a drop area where the run cannot start without a file (a recording for Diarization, audio for Transform, a picture for 3D). Then the prompt, with starters while it is empty, a model chip, the essential options as chips (size and count, duration and frame rate, voice, mode), and an **Advanced** fold that names what is inside when it is closed. Below that: the memory the model needs and whether it fits, when the gallery knows it, and one button. When the button cannot run, the reason is next to it. With no model for the type, the install note from the front page shows in the card.
|
||||
- **Run area.** While a request is out, a job card shows the model and options and the time that has passed. The server reports no phase and no percentage on these endpoints, so the bar is indeterminate. A failed run shows what the server said, says the prompt and settings are still there, and offers **Try again**. A finished run shows the result in a viewer for its type (picture grid, video player, waveform player, 3D viewer, spectrograms, or a timeline of speakers) with a toolbar:
|
||||
- **Favourite**, the same list the front page keeps.
|
||||
- **Download**.
|
||||
- **Use in** lists where the result can go next. A step is disabled, with the reason, when the destination cannot start from it, or when no model for it is installed it says so.
|
||||
- **Re-run with edits** puts the result's values back in the form. The fields you then change are outlined and listed under **Changed from this take**, and the button reads **Run again**. It is disabled when the original input was not kept (3D, diarization).
|
||||
- **Lineage** opens the lineage view of the front page for this result.
|
||||
- **Recent results.** A strip of the results of this type from the same history, with an All and a Favourites filter. Click a tile to show it above, and click it again to go back to the latest.
|
||||
|
||||
Diarization lists who spoke when as one lane per speaker, the talk time of each speaker, and the segments with their text when the model returned it. **RTTM**, **SRT** (only when there is text) and **JSON** are built in the browser from the result in hand.
|
||||
|
||||
## What a workspace accepts from the front page
|
||||
|
||||
The workspaces read these query parameters, so a link of your own works too: `prompt`, `model`, `size`, `n` (count), `from` (the id of the result it starts from) and `edge` (how). For example `/app/studio/video?prompt=Slow%20push-in&from=<id>&edge=animate`.
|
||||
@@ -80,9 +80,9 @@ this control may ignore it.
|
||||
|
||||
## Voice Library
|
||||
|
||||
Administrators can manage reusable voice-cloning references from **Operate → Voice Library** in the LocalAI WebUI. The library replaces per-model filesystem and YAML setup for supported cloning backends:
|
||||
Administrators can manage reusable voice-cloning references from **Build → Voices → Speech voices** in the LocalAI WebUI. (Speakers, the first tab of the same page, is a separate feature: voiceprints for recognising who is speaking. See [Voice recognition](/features/voice-recognition/).) The library replaces per-model filesystem and YAML setup for supported cloning backends:
|
||||
|
||||
1. Select **Create voice** and upload or record a clear reference clip.
|
||||
1. Select **Add a speech voice** and upload or record a clear reference clip. A speech voice can also be made while enrolling a speaker, from the enrol sheet.
|
||||
2. Enter the exact words spoken in the clip. Add more audio/transcript pairs when the personality needs more examples.
|
||||
3. Confirm that you have permission to clone the voice, then save the profile.
|
||||
4. Open **Text to Speech**, choose a model marked **Cloning ready**, and select the saved voice.
|
||||
|
||||
@@ -22,3 +22,9 @@ Each history remains independently bounded by `tracing_max_items`. When a
|
||||
history reaches that limit, LocalAI removes its oldest records from memory and
|
||||
disk. The existing clear actions on the Traces page remove both the in-memory
|
||||
history and its persisted records.
|
||||
|
||||
Opening an API request on the Traces page shows the request as its own page, at
|
||||
`/app/traces/<id>`: the status, the error LocalAI recorded, a timeline with the
|
||||
backend operations that ran while the request was open, and the request and
|
||||
response bodies. Bodies stay closed until revealed, and request headers are not
|
||||
listed. See [Traffic]({{% relref "operations/traffic" %}}).
|
||||
@@ -557,6 +557,21 @@ convention the Whisper / Voxtral transcription backends use.
|
||||
Pass `threshold` explicitly when switching recognizers - the per-model
|
||||
default only applies when omitted.
|
||||
|
||||
## The WebUI page
|
||||
|
||||
**Build → Voices** has three tabs. **Speakers** is this feature. **Speech voices** is the [Voice Library](/features/text-to-audio/#voice-library) for text-to-speech, a different store with its own recordings and transcripts. **From a recording** links to the diarization workspace, where a speaker can be named and remembered.
|
||||
|
||||
On Speakers:
|
||||
|
||||
- **Who is this** matches a clip against the people you enrolled (`POST /v1/voice/identify`). The answer is a sentence with the real distance and the cut-off, a word for how far inside the cut-off it sits (strong, likely, close call), and a scale with the cut-off drawn on it. The cut-off slider re-reads the answer in the browser; it does not call the server again. The page sends the cut-off in the request, 0.25 by default.
|
||||
- **Same person?** compares two clips (`POST /v1/voice/verify`) and uses the threshold the model returns. The word is not a probability, and the page says so.
|
||||
- **Enrol a speaker** opens a sheet: a recording, a name, optional labels, and a permission tick. An administrator can also keep the recording as a speech voice in the same step. Keeping a copy of the recording in the browser is off unless you tick it.
|
||||
- **Known speakers** is a list kept in the browser. The server has no list call, so the page cannot check it by itself. After a search it marks a saved person the server did not return as "not on the server" (only when the search asked for more people than it got back, so a short answer is never read as proof), and lists anyone the server returned that this browser has no record of. A server restart empties the server's index; enrol again, or use **Re-enrol from saved copy** if you kept one.
|
||||
- Removing a person waits ten seconds behind an **Undo** toast. Nothing is sent to the server until the time ends; Undo cancels it.
|
||||
- If no speaker-recognition model is installed, the tool is replaced by a note that says so, with models from the gallery to install. Users without the Voice recognition permission see a page that says it is off for their account.
|
||||
|
||||
The clip you test with is sent to the model on the server and is not kept. Analyze (age, gender and emotion guesses, all off by default) and the raw embedding are under **More tools**.
|
||||
|
||||
## Related features
|
||||
|
||||
- [Face Recognition](/features/face-recognition/) - the image analog;
|
||||
|
||||
@@ -50,7 +50,7 @@ Setting a per-agent model always overrides the default.
|
||||
|
||||
## Send a message
|
||||
|
||||
Open the new agent from the Agents page and type a message in its chat box, for example `Hello, what can you do?`. The agent replies in the chat panel within a few seconds. When the agent decides to use the action you configured, you will see the tool call and its result appear inline before the final answer, streamed live as the agent works.
|
||||
Open the new agent from the Agents page, type a task in its task box, for example `Hello, what can you do?`, and press **Start run**. The run page shows the agent working and then its answer, within a few seconds. When the agent decides to use the action you configured, the step count and the tool in use show live, and the report lists what the tool returned.
|
||||
|
||||
That is a complete agent: a model, a system prompt, and one tool, all running inside your LocalAI process.
|
||||
|
||||
|
||||
@@ -19,7 +19,7 @@ This section covers everything you need to know about installing and configuring
|
||||
|
||||
The Model Gallery is the simplest way to install models. It provides pre-configured models ready to use.
|
||||
|
||||
GPU recommendations require a memory estimate within 95% of the detected model memory budget at a 4096-token context. If none of the sampled candidates fit, the recommendation section is hidden. You can still browse the gallery and check individual models at your intended context size. The Home page also omits static GPU suggestions when no fitting recommendation is available.
|
||||
GPU recommendations require a memory estimate within 95% of the detected model memory budget at a 4096-token context. If none of the sampled candidates fit, the recommendation section is hidden. You can still browse the gallery and check individual models at your intended context size. On the Models page the recommendations appear as a "Best for this machine" list in the pane beside the table while no model is selected. Once you have installed a model, the list shows only the best fit and offers the others behind "more that fit". The Home page also omits static GPU suggestions when no fitting recommendation is available.
|
||||
|
||||
### Via WebUI
|
||||
|
||||
@@ -35,10 +35,13 @@ For more details, refer to the [Gallery Documentation]({{% relref "features/mode
|
||||
|
||||
The same Models page owns the complete lifecycle. Switch to **Installed** to
|
||||
search local configurations, filter them by running, idle, disabled, pinned,
|
||||
or distributed state, and open a model's runtime controls. Load, stop, edit,
|
||||
pin, disable, inspect backend logs, and remove actions stay with the selected
|
||||
model. The current view, search, filter, and selection are stored in the URL so
|
||||
links and browser history preserve your place.
|
||||
or distributed state, and open a model's runtime controls. Load and stop are on
|
||||
each row; edit, pin, disable, backend logs, and remove are in the row menu and in
|
||||
the details of the selected model. The current view, search, filter, and
|
||||
selection are stored in the URL so links and browser history preserve your place.
|
||||
The disk strip in the page header shows the free space on the models disk and
|
||||
opens a review of what can be removed to free more (see
|
||||
[Model gallery]({{% relref "features/model-gallery" %}})).
|
||||
|
||||
### Via CLI
|
||||
|
||||
@@ -73,15 +76,25 @@ Visit [models.localai.io](https://models.localai.io) to browse all available mod
|
||||
|
||||
## Method 1.5: Import Models via WebUI
|
||||
|
||||
The WebUI import page takes either a source to resolve or a configuration to
|
||||
write. Both live on the same page, behind the two tabs in its header.
|
||||
The WebUI import page (**Build → Import**) takes either a source to resolve or a
|
||||
configuration to write. Both live on the same page, behind the two tabs in its
|
||||
header. From a source, it is a short guided flow: source, review, import, done.
|
||||
|
||||
### From a source
|
||||
|
||||
1. Open the LocalAI WebUI at `http://localhost:8080`
|
||||
2. Click "Import Model"
|
||||
3. Paste the source into the **Source** field (e.g. `https://huggingface.co/Qwen/Qwen3-VL-8B-Instruct-GGUF`)
|
||||
4. Press Enter, or click **Import**
|
||||
2. Open **Build**, then **Import**
|
||||
3. Paste the source into the **Source** field (e.g. `https://huggingface.co/Qwen/Qwen3-VL-8B-Instruct-GGUF`), or start from one of the examples
|
||||
4. Review what the page says, then press Enter or click **Import**
|
||||
|
||||
While you type, the page reads the spelling of the source (Hugging Face
|
||||
repository, direct URL, configuration file, OCI image, Ollama model, or a path
|
||||
on the host) and shows the request the form will send. It also runs the checks
|
||||
that can be made before the import starts: whether the name is already in use,
|
||||
whether the chosen backend is installed, and how much disk and memory are free.
|
||||
The server returns no preview of the model configuration, the size or the memory
|
||||
a model needs before the import starts, so the page does not show them. They
|
||||
appear once the import starts, next to the free memory and disk.
|
||||
|
||||
The **What you can paste** panel beside the field lists every accepted scheme:
|
||||
`huggingface://`, `hf://`, a full Hugging Face URL, any direct `https://` URL,
|
||||
@@ -105,8 +118,10 @@ vision-language models and `mlx-audio` for text-to-speech models; other MLX
|
||||
repositories use `mlx`. An explicit backend selection in the import form always
|
||||
overrides this automatic routing.
|
||||
|
||||
Once the import starts, the page reports the current phase, the bytes
|
||||
transferred and a progress bar until the model is ready.
|
||||
Once the import starts, the page reports the download size and the memory the
|
||||
model needs against what is free, then the current phase, the bytes transferred
|
||||
and a progress bar. When the model is ready, the page names it and links to a
|
||||
chat with it and to Models.
|
||||
|
||||
### Writing YAML
|
||||
|
||||
@@ -403,19 +418,23 @@ local-ai models list
|
||||
|
||||
### Disk Usage
|
||||
|
||||
The **Installed** tab shows how much disk the installed configurations use: a
|
||||
summary strip with the total on disk and the portion shared across models, and
|
||||
a per-model size in the selected model's detail panel. A file referenced by
|
||||
more than one configuration is counted once in the total; the per-model size
|
||||
notes the shared portion, since deleting that model frees only its exclusive
|
||||
files. A reference whose file is not on disk — a download that never finished,
|
||||
or a file removed by hand — is shown as a warning on the model.
|
||||
In the WebUI, the **Installed** tab of **Models** shows how much disk each
|
||||
installed configuration uses. The **Size** column holds the size on disk, and
|
||||
for a model that shares files with others it adds the shared part ("1.2 GB
|
||||
shared"). A file referenced by more than one configuration is counted once in
|
||||
the total under the table, which also gives the shared total. The inspector
|
||||
beside the table names the models a model shares files with, and a model page
|
||||
lists its files under **Usage and history**, with each file's size, the other
|
||||
models that use it, and whether it is missing. Removing a model frees only its
|
||||
exclusive files; the files it shares stay while another configuration uses
|
||||
them. A reference whose file is not on disk (a download that never finished, or
|
||||
a file removed by hand) is marked on the model and in the cleanup review.
|
||||
|
||||
While no model is selected, the detail pane lists every referenced file with
|
||||
its size, which models reference it, and whether it is shared or missing;
|
||||
clicking a model name selects it.
|
||||
Only an admin can read the report. For anyone else, or if the read fails, the
|
||||
Size column shows the size of the files the gallery lists, and "size unknown"
|
||||
for a model the gallery does not list.
|
||||
|
||||
The report behind both views is available directly (admin only):
|
||||
The report behind these views is available directly (admin only):
|
||||
|
||||
```bash
|
||||
curl http://localhost:8080/api/models/storage
|
||||
|
||||
@@ -10,6 +10,7 @@ This section collects the operator-facing concerns for running LocalAI in produc
|
||||
|
||||
## Pages
|
||||
|
||||
- [Traffic]({{% relref "operations/traffic" %}}) - usage, failed requests, models, GPU and host, traces and Prometheus in one tab.
|
||||
- [Middleware: PII filtering and intelligent routing]({{% relref "operations/middleware" %}}) - per-model PII redaction and policy-based request routing.
|
||||
- [Cloud passthrough proxy]({{% relref "operations/cloud-proxy" %}}) - forward requests to OpenAI, Anthropic, or any compatible provider.
|
||||
- [MITM proxy for Claude Code / Codex CLI]({{% relref "operations/mitm-proxy" %}}) - redact PII from cloud-AI traffic without LocalAI holding API keys.
|
||||
|
||||
@@ -58,7 +58,7 @@ each shown only when it has something in it.
|
||||
|
||||
### In progress
|
||||
|
||||
One card per running or queued operation, each naming what is being done and to
|
||||
One row per running or queued operation, each naming what is being done and to
|
||||
what. A percentage and progress bar appear only once the operation is running
|
||||
and has reported progress, so a queued operation and a removal show none.
|
||||
Artifact-backed gallery models also report the phase, downloaded and total
|
||||
@@ -88,7 +88,8 @@ worker finishes.
|
||||
### Needs attention
|
||||
|
||||
Model and backend operations that failed and have not been acknowledged yet,
|
||||
each card carrying the error returned by the installer. Cluster staging never
|
||||
each row carrying the error returned by the installer, with **Retry** and
|
||||
**Dismiss**. Cluster staging never
|
||||
appears here: a staging job reports no error to the page, so a staging failure
|
||||
has to be read from the logs.
|
||||
|
||||
@@ -96,7 +97,11 @@ has to be read from the logs.
|
||||
|
||||
What finished, newest first, one row each: the name, what happened
|
||||
(`installed in 1m 12s`, `removed`, `cancelled`, or `failed:` with the error),
|
||||
the time of day it finished, and a link into Models or Backends. Model and
|
||||
the time of day it finished, and a link into Models or Backends. A cancelled
|
||||
install also has **Start again**, which installs the same target again: a
|
||||
download that was paused continues where it stopped, and one that was cancelled
|
||||
starts over. The record does not say which of the two it was, so the button
|
||||
says what it does in both cases. A cancelled removal has no such button. Model and
|
||||
backend installs and removals are recorded; cluster staging is not, so a
|
||||
staging run leaves nothing behind here once it finishes.
|
||||
|
||||
@@ -110,10 +115,13 @@ finished fan-out install is filed under Models or Backends, not under Cluster.
|
||||
|
||||
## Cancelling, retrying and dismissing
|
||||
|
||||
These are on the operation cards. **Cancel** and **Retry** are labelled
|
||||
buttons; dismissing is the **X** at the end of a failed card. The strip has no
|
||||
cancel button; the page is the only place work is stopped or restarted.
|
||||
These are on the operation rows. **Pause**, **Cancel**, **Retry** and
|
||||
**Dismiss** are labelled buttons. The strip has no cancel button; the page is
|
||||
the only place work is stopped or restarted. The Backends page also shows the
|
||||
progress of a backend install in the backend's own row, with its **Cancel**.
|
||||
|
||||
- **Pause** stops a running download and keeps the bytes already fetched, so
|
||||
installing the same model or backend again continues from there.
|
||||
- **Cancel** is offered while an operation is queued, whatever it is, and while
|
||||
an install is running. It is not offered once a removal has started: a
|
||||
removal in progress cannot be interrupted, so the window to call one off is
|
||||
@@ -121,13 +129,21 @@ cancel button; the page is the only place work is stopped or restarted.
|
||||
anything is touched. For artifact-backed gallery models, cancelling an active
|
||||
download leaves its partial files in place so a later install resumes rather
|
||||
than starting over. A cancelled operation leaves the live sections
|
||||
immediately and is not held on the strip the way a completed one is; it
|
||||
appears in the record as `cancelled`.
|
||||
and is not held on the strip the way a completed one is; it appears in the
|
||||
record as `cancelled`.
|
||||
|
||||
**Cancel waits for 8 seconds before it is sent.** The row says "Cancelling
|
||||
unless you undo" and a toast offers **Undo**. The server cannot take a cancel
|
||||
back, so the wait in the browser is the whole undo: while it runs nothing has
|
||||
been stopped and the download carries on. Closing the toast cancels at once,
|
||||
asking to cancel a second download ends the first one's wait, and leaving the
|
||||
page sends a cancel that is still waiting. A job that finishes during the wait
|
||||
is left alone.
|
||||
- **Retry** is offered on a failed model or backend install. It acknowledges
|
||||
the failure, which moves it into the record, and installs the same target
|
||||
again. It is not offered on a failed removal, which is not restarted by
|
||||
reinstalling.
|
||||
- **Dismiss**, the **X** on a failed card, acknowledges the failure without
|
||||
- **Dismiss** acknowledges the failure without
|
||||
retrying. The operation moves into the record with a `failed` outcome; it is
|
||||
not deleted. This is why the same failure can be found either under **Needs
|
||||
attention** or in the **Record**, depending on whether it has been
|
||||
|
||||
@@ -18,7 +18,9 @@ a single client-facing model name fans out across multiple downstream
|
||||
targets.
|
||||
|
||||
Both are inspected and configured from the same admin page
|
||||
(`/app/middleware`), backed by the same REST surface (`/api/middleware/*`,
|
||||
(`/app/middleware`), which draws the order a request passes through as five
|
||||
steps (Proxy, Admission, Filtering, Routing, Model) and shows only the rules of
|
||||
the step you select, backed by the same REST surface (`/api/middleware/*`,
|
||||
`/api/pii/*`, `/api/router/*`) and the same MCP tools.
|
||||
|
||||
## Request lifecycle
|
||||
|
||||
@@ -6,23 +6,66 @@ weight = 1
|
||||
`/app/operate` is the front door to the Operate console. It answers one
|
||||
question — is anything wrong — without you having to open four other pages.
|
||||
|
||||
## Needs attention
|
||||
The page opens with one sentence: **Everything is running.**, **2 things need
|
||||
you.** or, when nothing needs a decision but a row has something to read,
|
||||
**Running, with 1 thing to look at.** On a new installation, with no backend, no
|
||||
model and nothing loaded, it says **Nothing is running yet.** and offers
|
||||
**Install a backend** and **Browse models**.
|
||||
|
||||
The block the page exists for. It lists only things that want a decision:
|
||||
Under the sentence are four rows. A row with a problem opens by itself and holds
|
||||
the button that deals with it. A row with nothing to say stays one line. Click a
|
||||
row to open or close it.
|
||||
|
||||
- a backend with an update available
|
||||
- an operation that failed
|
||||
- a node reporting unhealthy
|
||||
## Needs you
|
||||
|
||||
**When nothing needs attention it says so in one line and renders nothing
|
||||
else.** There is no green panel: a status page that shouts when everything is
|
||||
fine teaches you to stop reading it.
|
||||
Only things that want a decision:
|
||||
|
||||
## Headline totals
|
||||
- a backend with an update available, with an **Update** button
|
||||
- an operation that failed, with **Retry** and **Dismiss**
|
||||
- a node that is not answering, with a link to the nodes
|
||||
|
||||
Requests, failed requests and p95 latency over the last 24 hours, each with a
|
||||
sparkline of the trend. These come from `GET /api/traces/summary`, which counts
|
||||
the trace buffer server-side:
|
||||
**Update** starts the same update as the Backends page. **Retry** installs the
|
||||
failed model or backend again, after moving the failure into the Activity
|
||||
record, and **Dismiss** moves it there without retrying. When nothing needs you
|
||||
the row says so in one line. There is no green panel: a status page that shouts
|
||||
when everything is fine teaches you to stop reading it.
|
||||
|
||||
The attention count on the **Status** tab is the number of items in this row.
|
||||
|
||||
## Capacity
|
||||
|
||||
One line says how full GPU memory is (system memory on a host without a GPU),
|
||||
and how much room the models disk has left. On a cluster it adds the memory of
|
||||
the nodes that are answering, and the row lists each node's memory when opened.
|
||||
|
||||
The row opens by itself when memory is 90% full or more, or when the models disk
|
||||
has less than a tenth free or less than 20 GB. It then lists the models loaded
|
||||
on this machine, each with an **Unload** button. Unloading asks first, stops the
|
||||
backend process of that model and frees its memory. The next request that uses
|
||||
the model loads it again.
|
||||
|
||||
Under the rows, a bar shows the memory in use now. **LocalAI keeps no memory
|
||||
history**, so the chart under the bar is drawn from readings the page took
|
||||
itself, one with each Operate summary poll (every 15 seconds), while Operate is
|
||||
open. It starts with the second reading, keeps at most 240 readings (about an
|
||||
hour), and is labelled "Since you opened Operate". The axis starts at zero, the
|
||||
capacity is a labelled line, and a data table sits behind the chart. Leaving
|
||||
Operate empties it. Temperature and power are not shown because the resources
|
||||
endpoint does not report them.
|
||||
|
||||
## Running now
|
||||
|
||||
Operations in progress, and the models loaded on this machine. On a single-node
|
||||
install the row lists the heaviest five models: backend, resident memory, CPU
|
||||
share and uptime, with **View logs** and **Stop model…** in the row menu. **Open
|
||||
this machine** leads to the full list. With distributed mode on, models run on
|
||||
workers rather than on the controller, so the row links to the Swarm page
|
||||
instead.
|
||||
|
||||
## Recent failures
|
||||
|
||||
Requests and failed requests over the last 24 hours, and p95 latency, counted
|
||||
server-side by `GET /api/traces/summary`:
|
||||
|
||||
```bash
|
||||
curl http://localhost:8080/api/traces/summary?hours=24 \
|
||||
@@ -42,39 +85,30 @@ curl http://localhost:8080/api/traces/summary?hours=24 \
|
||||
`hours` defaults to 24 and is capped at 168. Only 5xx responses and transport
|
||||
errors count as failures — a 4xx is the caller getting it wrong, not the
|
||||
installation being unhealthy. `p95_ms` is a nearest-rank percentile, not the
|
||||
slowest request.
|
||||
slowest request. The endpoint exists so a dashboard wanting three numbers does
|
||||
not fetch the whole trace list to count it. An installation that has recorded
|
||||
nothing says so rather than showing zeroes dressed as telemetry. The row opens
|
||||
when at least one request failed, and links to Traces and Usage.
|
||||
|
||||
The endpoint exists so a dashboard wanting three numbers does not fetch the
|
||||
whole trace list to count it. An installation that has served nothing yet says
|
||||
so rather than showing three zeroes dressed as telemetry.
|
||||
|
||||
## Host capacity
|
||||
|
||||
The overview also shows the host's current RAM or GPU capacity, utilization,
|
||||
and model storage. Loading, unavailable, and empty states are explicit. This
|
||||
uses the same 15-second Operate summary poll as the rail and attention data, so
|
||||
opening the overview does not start a second resource poller.
|
||||
|
||||
## Running now
|
||||
|
||||
On a single-node install the overview lists the models loaded on this machine,
|
||||
heaviest first, up to five. Each row shows the backend, resident memory, CPU
|
||||
share and uptime, with **View logs** and **Stop model…** in the row menu.
|
||||
**Open this machine** leads to the full list.
|
||||
|
||||
With distributed mode on, models run on workers rather than on the controller,
|
||||
so this section links to **Operate → Nodes → Running models** instead.
|
||||
The page does not report models that failed to load, because LocalAI records no
|
||||
load failures to read.
|
||||
|
||||
## This machine
|
||||
|
||||
On a single-node install, **Operate → This machine** (`/app/nodes`) shows the
|
||||
host and everything loaded on it:
|
||||
|
||||
- **Capacity gauges** for VRAM, RAM, CPU and the models disk, the same gauges
|
||||
the Nodes page draws for a cluster. A host without a GPU says so rather than
|
||||
showing an empty VRAM gauge.
|
||||
- **A memory bar** splitting host RAM by running model, so you can see which
|
||||
model is holding memory.
|
||||
- **GPU memory** as one bar, with what is free and a note that the next model
|
||||
must fit in that, or a loaded model has to be stopped first. With several GPUs
|
||||
each one gets its own line. A host without a GPU says so instead of drawing
|
||||
an empty bar.
|
||||
- **Models in host memory**: one bar split by running model, so you can see
|
||||
which model is holding memory. LocalAI reports the resident memory of each
|
||||
backend process, not the GPU memory of each model, so the split is of host
|
||||
RAM and no model is shown with a GPU size.
|
||||
- **The machine**: VRAM, RAM, CPU and the models disk, each with a bar, a
|
||||
percentage and what is left. A reading the host does not report is shown as
|
||||
"No data", never as zero.
|
||||
- **Running models**: search, sort by memory, CPU or uptime, open a model's
|
||||
logs, or stop it. Stopping asks for confirmation; the model loads again on
|
||||
its next request.
|
||||
@@ -84,9 +118,10 @@ per-model readings come from the `process` block of
|
||||
[`GET /system`]({{% relref "reference/system-info" %}}); the host CPU and disk
|
||||
readings come from the `cpu` and `disk` fields of `GET /api/resources`.
|
||||
|
||||
**Add machines** reveals the command to start LocalAI in distributed mode. Once
|
||||
distributed mode is on, the same route becomes the Nodes page and the rail
|
||||
entry moves to the Cluster group.
|
||||
**Add a machine** reveals the command to start LocalAI in distributed mode, and
|
||||
a link to the steps for adding a worker. Once distributed mode is on, the same
|
||||
route becomes the Nodes page and the **This machine** tab gives way to the
|
||||
**Swarm** tab.
|
||||
|
||||
Models and backends no longer live under a nested Host page. Use **Models →
|
||||
Installed** for model runtime and configuration actions, and **Operate →
|
||||
@@ -97,15 +132,24 @@ Old `/app/manage` bookmarks remain supported. They redirect with replace
|
||||
semantics to the matching Installed Models or Installed Backends view while
|
||||
preserving legacy search, filter, selection, variant, and development flags.
|
||||
|
||||
## The rail
|
||||
## The tab bar
|
||||
|
||||
The Operate rail groups its destinations under four headings —
|
||||
Runtime, Cluster, Observability and Administration — and shows a live value
|
||||
beside several of them: pending backend updates, running operations, healthy
|
||||
node count, request volume and error count. Host capacity lives on the overview
|
||||
instead of appearing as a separate destination.
|
||||
Operate has one row of tabs above the page: **Status** (this overview),
|
||||
**This machine**, **Swarm** (only with distributed mode on), **Runtime**,
|
||||
**Traffic** and **Settings**. A tab that holds several pages shows a second row
|
||||
of links under the bar: Runtime holds Backends, Activity, Logs and (on a single
|
||||
install) Failover; Traffic holds Usage, Traces and Middleware; Settings holds
|
||||
Settings and Users (with authentication on); Swarm holds Nodes, Placement rules,
|
||||
Failover and P2P. Adding a node opens from the Nodes page. Every page keeps its
|
||||
own URL, and a page such as a node detail keeps its tab highlighted. On a phone
|
||||
the bar scrolls sideways. The **API** link at the end of the bar opens the
|
||||
API documentation.
|
||||
|
||||
Those values are **orientation, not an alarm**. The rail only exists on Operate
|
||||
routes and can be collapsed, so anything urgent also appears in Needs attention
|
||||
and on the operations badge attached to the sidebar entry, which is always
|
||||
visible.
|
||||
Several tabs carry a live value: pending backend updates and running operations
|
||||
on Runtime, the healthy node count on Swarm, running models on This machine, the
|
||||
attention count on Status and the error count on Traffic. Host capacity lives on
|
||||
the overview instead of appearing as a separate destination.
|
||||
|
||||
Those values are **orientation, not an alarm**. The bar only exists on Operate
|
||||
routes, so anything urgent also appears in the Needs you row and on the
|
||||
operations badge attached to the sidebar entry, which is always visible.
|
||||
@@ -0,0 +1,98 @@
|
||||
+++
|
||||
title = "Traffic"
|
||||
weight = 3
|
||||
+++
|
||||
|
||||
The **Traffic** tab of the Operate console answers three questions: how much is
|
||||
the server used, what failed, and how are the machine and the models doing. It
|
||||
opens on an overview and has a second row of links: Overview, Usage, Models,
|
||||
GPU and host, Traces, Middleware and Prometheus. One time window (24 hours, 7
|
||||
days, 30 days or all time) is shared by the Overview, Usage and Models pages.
|
||||
|
||||
Every figure comes from a record LocalAI really keeps, and each page says which.
|
||||
Where LocalAI keeps no record, the page leaves the figure out and does not draw
|
||||
a zero.
|
||||
|
||||
| Record | What it holds | Pages that read it |
|
||||
|---|---|---|
|
||||
| Usage ledger | Requests and tokens, by model, user and API key, in hourly, daily or monthly buckets | Overview, Usage, Models |
|
||||
| Trace buffer | The most recent API requests, with status, error and duration. Off until tracing is on | Overview (failed requests, p95), Traces |
|
||||
| Backend-operation buffer | Loads and runs of each model, with their errors | Models, a trace |
|
||||
| Resources reading | Memory per GPU, system memory, CPU share, models disk | GPU and host |
|
||||
|
||||
## Overview
|
||||
|
||||
Five figures in a row: requests, failed requests, latency p95, tokens in and
|
||||
tokens out. Under them, three charts: requests per bucket, failed requests (the
|
||||
trace buffer, in twelve columns) and tokens in and out per bucket. Each chart has
|
||||
a data table behind it, and you can read a value with the arrow keys. A table of
|
||||
the five busiest models closes the page.
|
||||
|
||||
- **Failed** counts a transport error or a 5xx answer. A 4xx answer is the caller
|
||||
being refused, and LocalAI does not count it as a failure.
|
||||
- **p95** is the 95th percentile of request duration in the trace buffer. LocalAI
|
||||
does not compute a p50 or a p99, so there are none.
|
||||
- With tracing off, the failed and p95 figures say so and link to the setting.
|
||||
The ledger figures do not depend on tracing.
|
||||
- The trace summary reports at most 7 days. For the 30-day and all-time windows,
|
||||
failed requests and latency cover the last 7 days, and the page says so.
|
||||
|
||||
## Usage
|
||||
|
||||
The ledger as a table, grouped by model, by user (admins) or by API key (when
|
||||
authentication is on). Rows sort by any column, can be searched, and can be
|
||||
filtered by model. A row opens in place on its own chart. A user who is not an
|
||||
admin sees only their own numbers and their quotas, and the page tells them
|
||||
whether the current pace stays inside each quota.
|
||||
|
||||
**Export CSV** and **Export JSON** save the rows the table holds. The export runs
|
||||
in the browser. The ledger does not record status, endpoint or node, so those are
|
||||
not groups. Estimated cost is optional: you type a price per million tokens, it
|
||||
stays in your browser, and LocalAI has no price of its own.
|
||||
|
||||
## Models
|
||||
|
||||
One row per model: requests and tokens from the ledger, failed operations and the
|
||||
mean operation time from the backend-operation buffer, and the memory the backend
|
||||
process holds now. That is the resident host memory of the process, when the
|
||||
server reports it. GPU memory per model is not reported, and neither is a
|
||||
per-model latency percentile or a memory history.
|
||||
|
||||
## GPU and host
|
||||
|
||||
The current reading, refreshed every 5 seconds: memory per GPU, system memory,
|
||||
the CPU share and the models disk. LocalAI does not report GPU utilisation or
|
||||
temperature. Two charts show readings the page took itself, **since you opened
|
||||
this page**: the memory pool and the CPU share. They are kept in the browser, at
|
||||
most 240 readings, and leaving the page empties them. On a cluster the page also
|
||||
lists every node with its memory and CPU share.
|
||||
|
||||
## Traces
|
||||
|
||||
The recent API requests and backend operations, with the tracing settings above
|
||||
the list. Filter by failed or slow requests, search, sort, export and clear.
|
||||
Opening an API request shows its page: the status and the error LocalAI recorded,
|
||||
a timeline of the request with the backend operations that ran while it was open,
|
||||
and the request and response bodies. Bodies stay closed until you press
|
||||
**Reveal**, because they can hold prompts and personal data. Request headers are
|
||||
never listed. The two buffers share no request id, so operations are matched on
|
||||
time and model, and the page says so. A trace leaves the buffer when newer
|
||||
requests push it out; its page then says it is no longer there.
|
||||
|
||||
With tracing off, the page explains what is lost and offers **Turn on tracing**.
|
||||
You can also start LocalAI with `LOCALAI_ENABLE_TRACING=true`.
|
||||
|
||||
## Middleware
|
||||
|
||||
The order a request passes through: Proxy, Admission, Filtering, Routing, Model.
|
||||
The server fixes that order. Selecting a step shows only its rules. See
|
||||
[Middleware]({{% relref "operations/middleware" %}}) for what each step does.
|
||||
|
||||
## Prometheus
|
||||
|
||||
`GET /metrics` is admin only and returns the Prometheus text format. The page
|
||||
shows the URL and a scrape configuration with a copy button, checks the endpoint
|
||||
against the running server, and lists the metrics LocalAI can export. Only
|
||||
`api_call` is always present: a histogram of request duration in seconds by HTTP
|
||||
method and route. The others appear while the feature behind them runs. Metrics
|
||||
are off when LocalAI runs with `LOCALAI_DISABLE_METRICS_ENDPOINT=true`.
|
||||
Reference in new issue
Block a user