mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-09 22:54:42 -04:00
* build(ui): vendor the shared UI kit snapshot at 0.2.0 The restyle needs the kit's tokens, motion layer and component classes. Take a pinned snapshot instead of depending on the kit at build time, and keep a lock file with the version and per-file checksums so a later update shows exactly what changed. The product theme stays outside the vendored directory. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the ink and teal theme and bridge the old variables Define the product colours as the shared UI kit's roles, for light and dark, in theme-localai.css. The kit's contrast check passes on every pair. theme.css keeps the existing --color-* and --shadow-* names but now points each at a role, so App.css and the pages get the new palette without edits. Radii move to the kit scale. index.html now sets data-theme before first paint with the same rule as ThemeContext (stored choice, otherwise dark), because the contract layout of the theme file no longer defaults to dark by itself. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): restyle the shared chrome with the UI kit grammar Adjust the shared classes so every page picks up the same interaction language without per-page edits: - Sidebar sits on the canvas and the current row lifts onto a card. Section labels are tracked uppercase, the badge is a soft pill, and the phone drawer leaves the tab order when closed. - Buttons are flat: hover swaps the surface, press scales to .97, focus is a 2px ring with a 2px offset, danger is a tinted wash. - Inputs use the card surface and the control edge; switches, tabs, filter chips, badges and cards follow the same rules. Cards no longer lift on hover; only linked or button cards react. - Menus and popovers scale in from the trigger corner with 40px items. Dialogs get a veil fade and a spring settle. Toasts become pills at the bottom centre. - The page transition is a 250 ms fade with a 6px rise. It fills backwards so a finished animation no longer leaves a transform that confined dialog veils to the main column. The focus-ring test now checks the outline instead of a box shadow, and new specs cover the theme roles, the first-paint theme and the sidebar lift. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): move leftover hard-coded colours onto the theme roles The YAML editor restated the old blue palette in JavaScript, and a few pages kept literal blues, indigo and violet tints, or fallbacks that only applied because a variable was never defined. Point them at the theme variables so they follow light and dark and the new palette. The status badges in the account pages built their tint by appending "22" to a variable, which is not valid once the variable is defined, so they had no background. Use the wash roles instead. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): raise the type scale and control size toward the kit Body and list text moves to 15px and the rest of the scale follows the kit's 12/13/15/17/21/32/44 steps. Page titles, section headings and stat values are bold with tighter tracking; titles are 32px. Buttons, inputs, selects, tabs and nav rows are 40px high with the 12px radius, compact controls 32px. Tabs become a segmented control. The sidebar widens to 240px (64px collapsed) and nav rows get more room. Identifiers and counts in the split views use the mono face, and the stat grid becomes separate inset tiles. The Geist stack stays: it is bundled, and the thin look came from the size, weight and negative tracking, not the face. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): separate cards, panes and floating surfaces from the canvas Cards, the Models and Installed split panes, the composers and the confirm dialog use a stronger card edge, the rest shadow and the 20px radius, so they read as layers in dark as well as light. Menus and popovers move to a float surface (the hover tone in dark) with the float shadow. The selected rail row gets an accent wash and a 3px accent edge. The send buttons are a clear accent when there is something to send and a quiet inset when not; the Home button carries data-empty for that, since submitting an empty box does nothing. The assistant card becomes an accent wash with a square icon. New surfaces spec checks the pane edge, the selected row, both send buttons and the popover in both themes. The voice library empty-state spec now waits for the layout to settle before comparing two boxes. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): tidy the sidebar header and mark the current row with a dot The header gives the configured horizontal logo a fixed width and centres it in a 72px band, lined up with the nav icons. The collapsed rail shows the configured icon logo centred, and its nav rows become 40px tiles centred in the 64px rail. The current row gets the kit's accent dot, hidden in the rail. The theme, language and account controls stay in the sidebar footer: the app has no global search or command palette to put in a top bar, so a bar would only hold controls that already have a place. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): centre the avatar in the collapsed sidebar rail The collapsed avatar link was set to "flex: 0", which gives it a zero flex basis; with min-width: 0 the link shrank to its padding and the icon overflowed from the link's left edge, about 14px right of the icon column. Use "flex: 0 0 auto" in the collapsed and tablet rail. The footer controls now share the nav icon column in the expanded sidebar too (6px footer padding, 40px control boxes), and the tablet rail gets the same footer padding and hidden language code as the collapsed one. New spec measures the centre x of the nav icons, mark, avatar, language, theme and collapse icons in the collapsed, expanded and tablet states, in both themes, and asserts they agree within 1px. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): size the console and settings rails and stack settings on phones The Operate console rail, the Settings section rail and the account tab bar still used 13px text and the old underline tabs. They now use the 15px nav size, 40px rows and the segmented tab control. Form row labels are 15px with 13px hints. On a phone the Settings section rail sat beside the form and squeezed every row into a few characters. Below 720px the rail stacks above the content as a scrolling row and form rows wrap their control below the label. The save button no longer carries the icon font class, which drew a missing glyph before its label. The language menu is wide enough to keep Bahasa Indonesia on one line. Assisted-by: Claude Code:claude-sonnet-5-5 * build(ui): update the vendored UI kit snapshot to 0.3.0 Take the 0.3.0 snapshot: the sprite now carries the full outline icon set, and the new icons/fa-map.json maps Font Awesome names to icon ids. The map lets the app move off Font Awesome in the following commits. The lock file is regenerated with the new checksums. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add an Icon component backed by the kit sprite Icon draws an inline svg that points into the kit's outline sprite. The sprite is inlined into the page once, so the references resolve under any base path and in the embedded build without a request. Icons size with the font (1em), take currentColor, hide from assistive tech unless given a title, and spin on request. An unknown id draws a neutral circle. FaIcon and iconFromFa resolve Font Awesome names through the kit's map, for names that arrive at run time. iconHtml does the same for markup built as a string. The GitHub and Apple marks are small local glyphs, as the kit ships no brand marks. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): draw shared components and helpers with Icon Replace the Font Awesome elements in the shared components and in the utility modules with the Icon component. Lookup tables now hold kit icon ids instead of class strings. Code-block copy buttons and artifact cards, which build HTML strings, use iconHtml and a sanitizer-safe slot. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): draw model, backend and account pages with Icon Replace the Font Awesome elements on the home, models, backends, import, settings, login, account and users pages with the Icon component. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): draw chat, studio and recognition pages with Icon Replace the Font Awesome elements on the chat, media generation, talk and face and voice pages with the Icon component. The talk status table keeps its spin and pulse states as Icon props. The connected and error states now use a dotted circle and an alert circle, so they differ from the idle ring by shape as well as by colour. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): draw agent, node and operate pages with Icon Replace the Font Awesome elements on the agents, skills, collections, jobs, fine-tune, quantize, nodes, swarm, usage, traces and activity pages with the Icon component. Two class strings on layout elements held leftover button and icon classes from an earlier merge; they are cleaned up so the elements keep only their own classes. Assisted-by: Claude Code:claude-sonnet-5-5 * style(ui): size and align icons for the svg component Icon rules that targeted the font element now target the svg: the descendant "i" selectors in App.css and auth.css become ".lai-icon". The svg is 1.2em with a 2 unit line so it matches the visual size of the old glyphs at the 12 to 16px sizes the app uses, sits on the text baseline, and follows the context font size. Large empty-state marks get a lighter line. Menu icons get a 16px box and the readiness badge icons keep their 20px circle with padding. Add the pulse used by the talk status. Assisted-by: Claude Code:claude-sonnet-5-5 * build(ui): remove Font Awesome No source file references the icon font any more. Drop the package and its stylesheet import. The build no longer ships the solid, regular and brand font files. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): keep focus traps off the svg use references The dialog and drawer focus traps collect focusable elements with a "[href]" selector. An icon's use element carries an href, so it became the first "focusable" element and Tab at the end of the dialog stopped there instead of wrapping to the first button. Match "a[href]" instead. Assisted-by: Claude Code:claude-sonnet-5-5 * style(ui): keep icon sizes overridable and set the line width per svg Give the icon base rule zero specificity so a rule that sizes one icon (nav column, menu box, avatar, language switcher) wins whatever its order in the file. The sprite symbols fix their own line width; the inlined copy drops it so the width set on each svg applies, as the --lai-stroke custom property, and large marks can use a lighter line. Pin the avatar and the language globe to the boxes the sidebar alignment spec expects. Import the map as JSON with an import attribute so Node can load it in the spec. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): select icons by the svg markup and cover the sprite Specs that found icons by their Font Awesome class now select the svg by its data-icon. The dead-icon audit checks that every svg resolves to a sprite symbol and has a size. The class hygiene spec fails on any remaining Font Awesome class. A new spec checks every mapped icon id has a symbol, that the sprite is inlined once, that an icon paints at the root and under a forwarded path prefix, and that Font Awesome names map as documented. Assisted-by: Claude Code:claude-sonnet-5-5 * build(ui): update the vendored UI kit snapshot to 0.4.0 Take the 0.4.0 snapshot: hub tabs with count and attention badges, the six chart series tokens and the grid colour in the theme contract, and sample themes on a calmer palette. The kit headers are renamed and the lock file is regenerated with the new checksums, as for the earlier snapshots. Assisted-by: Claude Code:claude-sonnet-5-5 * style(ui): switch the theme to the calm palette Rewrite the LocalAI theme on the calm palette: a muted teal accent on a near-neutral green-grey canvas, desaturated status colours, no glow and no coloured shadows. The theme fills every role of the shared UI kit's 0.4.0 theme contract for light and dark, including the six chart series and the grid line. The bridge in theme.css keeps the old --color-* names working, adds the dark surface ladder (card, raised, float) and a strong edge, and points the fixed data hues at the chart series. Two values differ from the first sketch. The dark text on the accent fill is #021512 instead of #04201d: it reads 5.58:1 on the fill at rest and 6.4:1 on the hover fill, against 5.08:1 at rest for the lighter value. The light control edge is #6b7d7a. The kit's contrast script passes for all text pairs (4.5:1), control and focus pairs (3:1) and series colours (3:1). Leftovers that no longer fit the palette are fixed: the usage chart takes the six series colours in order, the audio and animation canvases fall back to the new accent, the face box loses its glow, and two gradient fills are now flat. The theme tests expect the new canvas colours. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): replace the console rail with hub tab bars Build and Operate no longer open a second navigation rail beside the page. Each is a hub: one row of the kit's hub tabs above the page, with count and attention badges that scroll sideways on a phone. Every URL and route stays as it was, plus a new /app/build landing page that lists the Build tools with a line each. Build tabs: Overview, Agents, Skills, Memory, Jobs, Fine-Tune, Quantize, Import, Voices (recognition and library) and Faces. Operate tabs: Status, This machine, Swarm (distributed mode only), Runtime (backends, activity, failover), Traffic (usage, traces, middleware) and Settings (settings, users), plus the API link. A tab that holds several pages shows a second row of links, and a sub-page such as a node detail keeps its tab highlighted. The feature and admin gates decide which tabs are drawn, and badges show only values the Operate summary already has. The sidebar lists Build and Operate under a Workspace label next to the Create group. The voice library moves under Build and the model import page gains the Build tab bar. The old rail styles, the rail signals and the console config are removed, and the Operate overview docs describe the tab bar. The specs that drove the rail now drive the tabs, and a new spec covers the tab for each route, gating, badges and the phone layout. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Home as a calm console Home now opens on one command bar: the model chip shows which models are warm, the MCP chip and attach buttons sit beside it, and Send is a solid button with an Enter glyph. Typing "/" opens a grouped, keyboard-driven action list built on the kit command list; every action has a destination in the product. Memory use folds into a one-line strip that opens into the loaded models, with Stop per model and Stop all. It opens by itself while a model is being staged and after a failure, and shows nodes and aggregate memory in a cluster. The list of resident models carries no per-model size because the API reports none. "Jump back in" lists the conversations stored in the browser, one card per day, with j and k to move, Enter to resume and delete with an undo toast. First run keeps the install steps and the recommended models. The assistant prompt is a dismissible line, the library links are one quiet row and the API section is collapsed. Chat accepts an empty new-chat hand-off for /new. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Home console Update the Home specs for the new structure and add specs for the slash menu, the model chip, the memory strip (expand, stop, staging, failure, cluster), the resume list (grouping, j/k, Enter, delete with undo), first run, the send hand-off, a non-admin user and the phone layout. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the fit, disk and cleanup helpers for the models page Pure functions and hooks that the rebuilt Models page reads, with node tests for the rules. modelLedger turns an estimate and the memory budget into one of three verdicts (fits, spills to CPU, over) with the headroom in bytes, and reads the models disk from the resources reading. The disk counts as low under 10 percent or under 20 GB free, and is absent when the server reports none or runs as a cluster controller. cleanupPlan ranks installed models from what the API reports: loaded, pinned, or named by an agent, a task, a failover chain or an alias keeps a model protected; another installed build of the same gallery model is a duplicate; disabled models rank above idle ones. The API records no last use or use count, so none is used. When a lookup fails, nothing is called safe. useModelRemoval holds a removal in the browser for an undo window and sends the existing delete call only when the window ends. Leaving the page drops the batch without deleting anything. The undo toast takes optional labels so other pages can reuse it. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Models as a ledger with a disk strip and cleanup review Explore is one dense table. Each row carries the size, a solid memory bar and the headroom in words ("3.7 free", "+1.5 on CPU", "0.9 over"), worked out from the estimate at the chosen context length. Capability chips show the server's count for each facet, search keeps its meaning and "/" jumps to it, and a density switch (also "d") picks comfortable or compact rows. Selection is a surface step and a check, never a rail. Arrow keys move, Enter installs and Esc closes the inspector, which keeps the fit summary, VRAM by context chart, variants, files, links, tags and licence. A failed install shows its error in the row with a Retry that dismisses the old failure first. A failed or empty listing says which it is, and a host with no GPU is measured against memory and says so. Installed uses the same table with state filters that carry counts, a state per row, Load or Stop on the row, the row menu and the sort by size. Sizes come from the files the gallery lists, so a model it does not know shows a dash. A strip in the header shows the free space on the models disk. It turns amber under 10 percent or under 20 GB free, hides when the server reports no disk or runs as a cluster controller, and opens the cleanup review. Explore says how much an install leaves free. The review ranks installed models as Safe to remove, Probably safe and Your call from real facts only, lists protected models with the reason, and says plainly that usage history is not recorded. A sticky bar shows what a choice frees. Confirming runs a dry run that checks again and lists what will go. Removal waits 30 seconds with an undo; nothing is deleted before that, and leaving the page deletes nothing. The old rail, filter band and popover styles are removed. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Models ledger, Installed table and cleanup review Update the Models, lifecycle, cluster fit, height, search focus and surfaces specs for the table and inspector, keeping what each one checks. New specs, on a shared 41-model gallery stub with three machine profiles: the fit bar and headroom words for a 24 GB card, an 8 GB laptop and a host with no GPU; facet counts, search, "/" and Escape; selection, arrow keys, Enter to install, density; the disk strip when normal, low and hidden; and the states (loading, empty, offline, install failed, phone). Installed covers filters with counts, row actions, the row menu, sizes and sort. The cleanup specs cover grouping, protected models, the honest-data note, the effect bar, the dry run, the undo window, a failed delete, leaving the page, and the phone sheet. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the placement helpers and the estimate hooks The Placement section and the model page need the same few rules, so they sit in plain functions that can be read and tested alone. placement.js holds what gpu_layers, tensor_split and main_gpu mean (unset asks for every layer and the llama.cpp engine trims it, zero is CPU only, 99999999 is the value LocalAI itself writes for all layers), the device list taken from the resources reading, the split by free memory, the part of an estimate that grows with context (read from two lengths, since that term is linear), the fit states with their limit (95 percent of free memory, and the leftover has to fit in system memory too), and a bisection for the largest layer count whose estimate fits. The estimate returns one total and no layer count, so the search runs over 1 to 256 and stops at the first count that no longer changes it. modelWalk.js keeps the order of the list a model page was opened from, in memory and in session storage, for the previous and next buttons. usePlacementEstimate reads /api/models/vram-estimate for a choice, again at twice the context, and with every layer, and keeps readings for the session. useModelPage reads a gallery entry by name, an estimate by context size (from the model's own files when the gallery does not list it), the builds and the loaded models. usePlacementConfig edits the four placement keys of an installed model and saves only what changed. useModelActions is the Load, Stop, disable, pin and remove logic of the Installed table, shared with the model page. MemoryBar is one solid bar with a tick at the capacity of its pool; over capacity it grows past the tick and the tick turns red. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the Placement section to the model editor Run this model on: CPU only (gpu_layers: 0), Auto (the key stays unset) or Custom. Custom takes a number, has an All layers button that writes 99999999, and shows a slider only when the estimate reports the model's layer count, which it does not today. Context size has presets and a number field because the KV cache follows it. With two or more GPUs there is a split (written as percentages, with a button that takes them from the free memory of each card) and a main GPU. A bar per GPU and one for system memory show what other programs use, the model's weights and working memory, and the part that grows with context, with the room left or how far over it is. Under them a verdict in plain words: Fits in GPU, Spills to CPU, Too many layers for the GPU, Runs on CPU only, No GPU found, Not enough memory. It says "slower" and never a multiplier, because the estimate has none. Fit it for me asks the estimate for the largest layer count that fits the free GPU memory and says what it set, with Undo; it is hidden when the estimate is unavailable or the host has no GPU. Loading shows skeletons, an unavailable estimate shows a note with Retry, and a server that schedules onto other machines shows no bars, because its device list is the controller's. The editor shows the section for an installed model, with a link in its section rail. Auto sends null for the key, since a patch only merges, and a null read back opens as Auto. The docs describe the section and what each mode writes. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): open a model on its own page A model has an address, /app/models/<name>, for an installed model and a gallery entry alike. Open it from the arrow at the end of a row, a double click, "o" on the selected row, the inspector's Open details button, or a tap on a phone. The title block holds the main action: Install with a chevron that chooses the build, or Load and Stop with a menu (disable, pin, edit configuration, logs, delete with a confirm). A strip answers whether it fits, what it does and what installing leaves free. Tabs: Overview (about, a memory bar, state, the pages it opens in, and the agents, tasks, chains and aliases that name it); Fit and memory (verdict, context sizes, the bar split into weights and context, and memory by context against the limit, with a data table); Variants and files (builds with size and fit, install any, the files of the chosen build). For an installed model also Usage and history, which says what the API does not record instead of drawing an empty chart, Configuration, which is the Placement section with the file it writes and a link to the full editor, and Logs, the backend log viewer without its page. Keys 1 to 6 switch tabs, [ ] and j k walk the list the page was opened from, Esc or Backspace go back. The list stays mounted behind the page, so Back finds its view, search, filters, selection and scroll as they were, and focus returns to the row's arrow. The page covers loading, an unknown name with the closest matches, the gallery being out of reach, an install in progress with Cancel, and a failed install with Retry. The docs describe the page and its keys. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the model page and the Placement section New specs for the model page: reaching it from Explore, Installed, a double click, "o", a pasted link and a phone tap; the walker and Back with the search, a filter, the selection, the Installed view and the scroll kept, and no second read of the gallery; the title block, the answer strip, tabs by click, keys and arrows; Fit and memory, builds and files with the install call each one makes; an installed model's actions, used-by, the honest usage tab, configuration and logs; loading, an unknown name, offline, an install in flight and a failed one; and the phone. New specs for Placement: every mode and the keys it writes, the slider only when a layer count exists, the context presets, the bars and every verdict, two GPUs, no GPU, a cluster, an unread machine, a loading and an unavailable estimate, Fit it for me and Undo, and the section in the model editor with its save. The phone tap on a row now opens the page, so the two phone specs that expected the inspector as the page check the page and keep the inspector check for a window between a phone and a desk. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): link Studio results and open workspaces from a prompt Each workspace now records the result it was made from (parentId and an edge kind such as take, animate or to-3d) and reads a prompt, model, size, count and source from the query string, so one page can hand work to another. A source result is fetched from the server's own output file and becomes the start image, the picture for 3D, or the audio file. A note on the page says when the source loaded or could not be loaded. Diarization had no history; it now keeps the file name, the model and a speaker count, never the recording. Prompts are cut at 2000 characters when stored. The pure helpers (type suggestion, grouping, lineage layout, favourites, clearing) have node tests. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Studio front page as a composer with your work The front page is a prompt box with a chip per type, a type suggestion from the words, starters, and the options each workspace accepts. Generate opens the workspace with those filled in. A type with no model is a dashed chip that shows a gallery model, its size, memory need and an Install button only when picked; the typed words stay while it installs. Under it, Your work lists results from every workspace as a masonry with filters, counts, favourites and a Clear history action. Results made from each other stack into a project tile and open as a lineage board with a dock for running a new take or branching to the next step; steps the destination cannot start from yet are disabled with the reason. The docs describe the page, what is stored in the browser, and the query parameters a workspace accepts. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Studio composer, your work and the lineage view Specs for the type suggestion, the keys, hand-off to each workspace, the install path for a missing model, the masonry filters, favourites and clearing, stacking, the lineage board, new take and branch, steps that are disabled with a reason, and the phone layout. Existing Studio specs move from lanes to chips with the same intent. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the shared Studio workspace frame and move Images onto it The seven Studio workspaces get one layout: a row of type tabs, a compose card (optional sources as chips, a prompt with starters, a model chip, the essential options as chips, an Advanced fold that names what is inside, the memory the model needs, and one action with the reason when it cannot run), a run area, and a strip of recent results of the type. The run area shows a job card with the time that has passed and an indeterminate bar, because these endpoints report no phase or percentage; a failure with what the server said and one action; or the result with a toolbar: Favourite (the list the front page keeps), Download, Use in (the hand-off targets, disabled with the reason when a destination cannot start from the result), Re-run with edits (the take's values go back in the form, changed fields are outlined and listed) and Lineage. A type with no model shows the install note from the front page. Images is the first workspace on the frame. It keeps its size, count, steps, seed, negative prompt, source image and reference images, and its history writes, including the parent link and edge of a hand-off run. useMediaHistory.addEntry now returns the id of the entry it stored. The docs describe the workspace page. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move Video onto the workspace frame Video keeps its size list, duration, frame rate, steps, seed, CFG scale, frame count, negative prompt, start and end image and avatar audio. The start and end image are source chips, the avatar audio opens the recording and paste input from a chip, and the rest sit in the Advanced fold. A start image from a hand-off shows as a chip with its picture. Results play in the video player with the shared toolbar. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move TTS onto the workspace frame TTS keeps the saved-voice picker for cloning models, the typed voice for the others, the voice library deep link, and the delivery instructions, which now sit in the Advanced fold. The result is the waveform player with the words under it. The stored entry also keeps the voice id so Re-run with edits can select the same saved voice. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move Sound onto the workspace frame Sound keeps its Simple and Advanced modes and every field of both: the description, instrumental, vocal language, caption, lyrics, BPM, duration, key, language, time signature and think mode. The mode switch, instrumental and duration are in the compose card, the rest in a More options fold. The stored entry keeps all of the fields, so Re-run with edits restores the form as it was. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move Transform onto the workspace frame Transform keeps its audio and reference inputs with upload and record, the echo test, the key=value parameters (now in the Advanced fold), the input and output spectra and the three waveform players. The audio that was chosen shows before the run, waiting to be transformed. Re-run with edits puts back the model and parameters and fetches the audio and reference the server kept for that run. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move 3D onto the workspace frame 3D keeps the picture input with paste and webcam, the animation operations a model declares, quality and background, the shape and material steps, guidance and seed, the GLB and animation viewers, the remesh control and the download. A 3D result now has a title from the motion prompt when it has no label, so the strip and the front page name animation results by what was asked. Re-run with edits is shown disabled with the reason, because only a small thumbnail of the picture is kept. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move Diarization onto the workspace frame Diarization keeps its model and recording inputs, the option to prepare speakers to remember, the clean-speech previews, naming and remembering a speaker, and the history entry with only the file name, model and counts. The result now shows a timeline with one lane per speaker, the talk time of each speaker, and the segments with their start time and text. RTTM, SRT (only when the run has text) and JSON are built in the browser from the result. The helpers for talk time, axis ticks and the two text formats have node tests. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): remove the styles and lists the old workspace layout used Nothing renders the two-column workbench, the control column, the old history lists, the generation progress tiles, the TTS voice picker or the result echo any more. The inline-style baseline drops with them. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the workspace frame and one run per type Specs for the type tabs, the compose card and its reason when Generate cannot run, starters, the Advanced fold, the job card with no invented progress, a failed run and its one action, the install note, the strip with its favourites filter, Use in with its disabled steps, Lineage, the parent link, Re-run with edits and its list of changes, deleting and clearing, and the hand-off note. One run through each of Video, TTS, Sound, Transform, 3D and Diarization, the phone layout of all seven, and reduced motion. Existing Studio specs move from the old control column to the compose card with the same intent. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the chat thread with raised user turns, prose replies and one-line activity Your messages are raised blocks on the right at a 760 px measure and the model's replies are plain prose under its name and a warm or not loaded dot. Reasoning, tool calls and their results fold into one quiet line that opens inline into steps. Code blocks carry a Copy button and a Canvas button that opens that block in the canvas, image attachments are thumbnails that open in the lightbox, and files are chips. Per-message actions show on hover, on focus and on the last turn, and a turn takes focus so the arrow keys and C, E, R and B work. A failed reply keeps the text written so far and shows the reason with one Retry action. The Agent chat page keeps the older rules: the new styles are scoped to the chat page and use their own class names. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): use the Home command bar as the Chat composer Chat now ends in the same object as Home: the model chip, the MCP chip, a Canvas chip, the message box with attach buttons, a solid Send and the hint line, with the slash menu on the kit command list. The slash menu lists what Chat can do today (switch model, new chat, conversations, manage mode, canvas, find, settings, export, clear). While a reply is streaming Send becomes Stop, which Esc also presses, and Up in an empty box edits your last message. Attached images show as thumbnails and a line under the bar carries the speed and the token count. HomeComposer takes optional props for this (extra chips, its own slash list, Stop, paste, a stricter Enter); Home passes none of them. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): open conversations from a Ctrl K menu with day groups, undo and a slim header The conversations list opens as a centred menu on Ctrl or Cmd K. It groups chats by day like the Home resume list, shows the model that answered and the time, searches names and message text, and moves with the arrow keys. Enter opens a chat, F2 renames it and Delete removes it. Removing a chat hides the row and shows the kit undo toast; the chat is deleted for good only when the undo time ends. Rename, duplicate, copy and export are on each row, as before. The header is one slim bar: the Chats button, the chat name (click to rename), a context meter when the context size is known, settings and a More menu with rename, duplicate, copy, export, model info, keyboard shortcuts and clear. A dialog lists the shortcuts the page answers to. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): show loaded state, capabilities and fit in the Chat model switcher The model chip in Chat opens the same list as Home, grouped as Loaded now and Installed. Each row says warm or not loaded and marks models that understand images. When the list opens, the page reads the host memory once and asks the server to estimate each listed model at the chat's context size (up to twelve, three at a time), then shows what the model needs and whether it fits: free memory, how much would run on the CPU, or how far over the machine it is. A model with no estimate shows no fit text, and no load time is shown because the API does not report one. A memory bar closes the list. The picker takes the model list from the page when it has one, and useModels can skip its own request. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move chat settings into a sheet and add find in chat, jump to latest and a wider canvas Settings open as a kit sheet: the system prompt, temperature, top P and top K (each says "model default" until it is changed and has a Reset), the context size with quick sizes and a note that it only drives the meter, Manage mode and Focus mode, the model info for admins with its Edit config button, and Clear conversation behind a confirmation. The old slide-out drawer and the model info panel are gone. Ctrl or Cmd Shift F (or the search button, or /find) opens a search bar over the thread. It marks matches in the messages already on the page, shows "n of m" and steps with Enter and Shift+Enter. Nothing is sent to the server. Jump to latest is a pill above the composer. Esc stops a reply, then closes the search, then closes the canvas. The canvas panel gets the kit look: tabs, a Code and Preview switch, Copy and Download, a full-page layout on narrow windows, and translated labels. The Agent chat page shares it and gets the same look. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the empty, no-model, loading and phone states to Chat An empty chat opens with the composer under one line, starters to try, whether the model is loaded, and the Jump back in list: the same rows as Home, read from the chats the page holds. With no chat model installed, an install card offers the starter models for this hardware, the gallery and import, and the composer stays so the text is not lost. While a reply waits for a model, a load card shows what the page knows: the phase the server names, the node, the bytes and the time left when the server reports them, and a progress bar. A model that is just not loaded yet gets a plain note, with no invented phases or estimates. The foot warns when the context is nearly full. On a phone the header drops its labels, the model list and the settings open as sheets from the bottom, per-message actions stay in view and the canvas takes the whole page. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Talk as a calm voice page over what the connection really does Talk is one stage and one transcript. The stage has the pipeline chip, the voice and language chips, an outline orb that follows the real microphone and playback levels, a heading and a sentence for the current state, and the controls. The transcript lists You, Reply, Tool and Result lines and can be copied. Session settings (instructions, voice, language, tools, Manage mode and the pipeline's parts) open in a sheet. The states are the ones the code reaches: no pipeline model, idle, connecting, listening, thinking (also while a tool runs), speaking, an interrupted reply (the server cancelled it; a note marks the cut), a blocked microphone, a link that failed during a session, and any other error with its reason and a link to the traces. Push to talk and hands-free are not on the page, so they are not shown. Diagnostics keep their waveform, spectrum and stats, drawn in theme colours. The page text moves into the talk namespace, and the old Talk and visualizer styles and the inline-style count go down with the rebuild. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): remove the chat styles and strings the rebuilt page replaced The settings drawer, the model info panel, the bubble avatars, the conversation menu popover, the context bar, the recent strip, the staging bar, the file badges and the focus-mode rules have no user now. Their rules, the Chat page's focus class and seven unused empty-state strings are removed. The Agent chat page keeps the shared message, sidebar and input rules it still renders with. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the rebuilt Chat and Talk pages Add a Chat page under features (thread, message actions and keys, the message box and its slash actions, the model list with loaded state and fit, conversations on Ctrl K, settings, find, canvas and the empty, no-model and loading states) and a Talk section to the realtime API page with the states the page shows. Manage mode now turns on from the chat settings or /assistant, and the client MCP steps point at the MCP chip. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): settle the rough edges of the new Chat and Talk pages The undo toast sat under the conversations menu, so the Undo button could not be pressed while the menu was open; the menu, the sheets and the fullscreen canvas now stay below the toast layer. Esc in a rename box saved the text through the blur that follows it; it now cancels. The image viewer closed on Esc only when the page did not re-render on the same key, so its key listener is registered once and reads the latest handlers. Keys on a focused message no longer type their letter into the editor they open, "/" from outside a text field starts a command as it does on Home, and Esc leaves the page's own dialogs alone. Code in the canvas is highlighted for languages that have no preview. The conversations menu drops its key hints on a phone so Clear all stays in view. Talk hides Test tone while connecting and calls a server error "Something went wrong", since the call can still be open. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the rebuilt Chat and Talk pages Specs for the thread layout and the activity fold, code blocks, image thumbnails and the viewer, per-message actions and their keys, a failed reply with its one Retry, Stop and Esc while streaming, the composer and every slash action, the conversations menu (groups, search, resume, rename, delete with undo that ends by itself, one chat left), the model switcher with loaded state, vision and fit text from stubbed estimates, the settings sheet, the canvas panel, find in chat, Jump to latest, the empty, no-model and loading states, the phone layout and reduced motion. Talk is driven over a fake WebRTC link through idle, connecting, listening, thinking, speaking, interrupted, blocked, lost, error and no pipeline, its settings sheet and its phone layout. Node tests cover the message text helpers and the conversation grouping. The existing chat specs move to the new structure with the same intent: the transcript spec now describes the raised turn and the prose reply, and the render smoke accepts Talk's own header. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): keep a bounded run log for agents in the browser The server keeps no run history for an agent, so a run is one task and the events until the agent answers, written to browser storage while the page watches the stream: up to 50 runs per agent, task, step and answer text only. Stored chats from the earlier agent chat page read as runs with stable ids. A run still marked running five minutes after its last event reads as stopped. Helpers read an agent's config into chips, build the list of changed fields against the saved config, hide secret values and offer starting points. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Agents area around runs The Agents page shows what needs a look (work in flight, a run that failed in the last day), then each agent with its model, attached memory and skills, and a strip of its last 14 runs. An agent has its own page: model, tools, memory, skills, instructions, a task box and its runs. A run has an address, shows the thread while it works (steps folded into one line, the tool in use, the answer as it arrives) and settles into a report about a second and a half after the agent answers: task, outcome, follow-ups, evidence and steps, with wide tables opening wider on demand. A failure says in plain words what happened and offers Run again. Create and edit fold into sections with a ready mark and a one-line summary, start from a template or an optional model-written draft, and open a preview sheet with the config as saved and the changes against the saved agent. Status becomes a quiet panel in the same language, and the old chat link opens the agent page. There is no Stop, approval, steer, version or dry-run control, because the agent API has no call behind them. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Agents launcher, agent page, runs and editor Specs for the Now strip and run strip, search and the empty state, the agent page, starting a run, the live thread, settling into the report, the run address across a reload and for a run from another browser, follow-ups with their history, failures, the folding editor with ready marks, templates, the preview sheet with hidden secrets and changes, the status page, and the phone, 1440 and 2560 layouts. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe runs and the new agent create flow Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add helpers for tasks, schedules and job outcomes Reads a cron expression the way the server does (five fields or an @ shortcut), checks it, and puts the common shapes in words. The next run is left out on purpose, because the schedule follows the server clock, which the browser cannot read. Also groups jobs by day, sums the last seven days, and gives each job one outcome line from its result or error. A rerun call starts a new job with the same parameters and media. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Jobs area around runs The Jobs page opens with one sentence about the last seven days, then the tasks (model, schedule in words, last 14 jobs, enabled switch, Run now) and a run history grouped by day. Each row has an outcome sentence and opens to the error or the start of the result with one next action. Deleting a task waits 30 seconds with an undo button. A task opens as a page with its recent runs, its prompt with the gaps marked and its schedule. The task form folds into sections, takes a schedule as a preset or a checked cron expression, warns about prompt gaps the schedule does not fill, and has a preview sheet. A job opens as a document: task, outcome, delivery and the recorded steps; a failed job says what happened and offers Run again. Run now now sends attached media through the job call, which is the only one that takes it. "Clear History" only ever cancelled running jobs, so it is now called Stop running jobs. Webhook headers of a saved task show as JSON instead of [object Object]. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Jobs page, task pages and job pages Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the Jobs page and the task form Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add helpers that say who uses a skill or a collection An agent loads a skill when skills are on and the skill is in its selection (an empty selection means every skill). It reads the one collection that carries its own name, when its knowledge base is on. The helpers derive that from the saved agent configs, build the config that adds or removes a skill or a collection, and estimate tokens as characters divided by four. Removing the last selected skill switches skills off, because an empty selection would mean every skill. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Skills and Memory as one library Skills and collections sit in a list with an open item beside it. Each row says who uses it, read from the saved agent configs, or says it is not used yet. Chat reads neither, so it is never named. An item opens in a pane with a Used by strip (names link to the agent, a small x removes it, with undo) and an Add to menu that shows what the addition costs. A collection can be added only to the agent that carries its name. The Memory pane searches the collection alone and shows ranked passages with scores, lists web sources with their refresh interval and the files, shows the server message when an upload fails, and names the endpoints and where files stay. The Simulate a message sheet runs a collection search and shows an agent's skills with a token estimate. It runs no model. The collection details route now opens the same page. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): show use and cost in the agent form pickers Each skill in the agent form says which other agents use it and what it adds to every message, with a total for the selection. The memory section names the collection the agent reads. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Skills and Memory libraries Specs for the used-by lines (including an agent that uses every skill), the filters, search, add to agent, remove with undo, the last-skill case, an unreadable agent list, the empty states, git repositories, the Memory question box, sources, uploads that fail, the Simulate sheet with the parts the API can run, the agent form hints and the phone layout. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the Skills and Memory libraries Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add helpers for the Operate status page and backend rows Pure functions for the parts that need rules. They work out the memory pool the page measures (GPU memory, system memory, or the workers of a cluster that are answering), which pools are too full, the headline and the four ledger rows, the geometry of the capacity chart, and what removing a backend would leave without a runtime (models name their backend, and a meta backend names the concrete one it points at). A second set says what a backend row states: installing, queued, removing, failed, update available, current or absent. LocalAI keeps no memory history, so the chart reads a bounded buffer of readings the page took itself and says so. A reading with no total is dropped rather than drawn as zero. Two hooks are shared by the pages that need them. One retries a failed operation after moving the failure into the record. The other holds a cancel for an undo window, because the server cannot take a cancel back. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Operate Status, This machine, Backends, Activity and Logs Status opens with one sentence ("2 things need you", or "Everything is running") and four rows: Needs you, Capacity, Running now and Recent failures. A row with a problem opens by itself and holds the button that deals with it: Update a backend, Retry or Dismiss a failed operation, Unload a model. A quiet row stays one line. A new installation gets a first-run screen, a cluster sums the memory of the workers that are answering, and a page still waiting for an answer says so. The chart under the rows is drawn from readings the page took while it was open and is labelled that way, because LocalAI keeps no memory history. This machine leads with GPU memory as one bar, then host memory split by running model, then VRAM, RAM, CPU and disk with a bar each. The running models become a kit table with the same menu and stop dialog. Backends is one list with Installed and Catalog views. A row says what the backend is doing (a progress bar with Cancel, Queued, Failed with Retry, Update 1.2.0, Current), carries the one button that matters, and opens in place. Removing a backend names the models and the meta backends that would stop working. Check for updates, Update all, From URL and a first-run recommendation for llama-cpp are in the header. Activity keeps its three sections as quiet rows. Cancel waits eight seconds with an undo toast, because the server cannot take a cancel back; a cancelled install can be started again from the record. Logs gets a process list, a picker, stream and text filters, Follow and Times switches, and a Clear with an undo window. Not shown, because the API has no data for them: GPU temperature and power, a size per backend, an earlier version to roll back to, a dependency lookup beyond the models and meta backends that name a backend, and models that failed to load. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover Operate Status, This machine, Backends, Activity and Logs New specs for the Status headline and ledger (healthy, needs attention, one thing, a full memory pool alone, loading, first run, cluster), its actions (Update, Retry, Dismiss, Unload with its dialog), the capacity chart built from readings taken while the page is open and bounded, the phone layout, no coloured edge on a row, and reduced motion. The Backends specs cover the two views, install progress with Cancel and its undo window, Retry on a failed install, Update, Update all, Check for updates, removal with the models and meta backends it would break, Install from URL, the first-run recommendation, a cluster, and a phone. Activity gains cancel with undo, Cancel now, a second cancel, leaving the page, progress, and starting a cancelled install again. Logs covers the stream and text filters, Follow, Times, Export, Clear with undo, the process picker and list. This machine covers the GPU strip, several GPUs, no GPU and Add a machine. Existing specs keep their intent and follow the new structure: rows open in place instead of in a pane, Update replaces Upgrade, the notice spec now pins that an update is a row state and not a banner or a rail, and a cancel waits for its undo window. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the rebuilt Operate Status, Backends and Activity pages Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add helpers for the Swarm pages Pure functions for what the pages work out from the cluster API: a node's state in words, which nodes a placement rule may use, what a rule would ask for, what a drain or a lost node would leave without service, the nodes a bulk backend update reaches, and the join commands for a worker, a peer instance and a memory shard. Hooks read the roster, the loaded replicas and the rules. Everything runs in the browser from data the page already holds, and says when it cannot see free memory or disk. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Swarm hub: nodes, node page, placement rules, failover Nodes is a sortable table with comfortable and compact rows, a Needs attention filter by reason, a map of the cluster that is not drawn on a phone, the running models, and a bulk backend update for the nodes that drifted. A node is a page: state, vitals, a drain preview computed from the loaded replicas and the rules, tabs for models, backends, logs and capacity and labels, and Remove that asks for the node's name. Placement rules are written as sentences, show where each model is loaded now, and edit in a side sheet with a preview of the nodes a draft could use. Deleting a rule waits a few seconds so it can be taken back. Failover keeps its chains, adds what the router does when a worker stops answering and a per-node preview of what would stop. Add a node covers a registered worker, a peer instance and a memory shard, with a command to copy and a live line that says when the machine arrived. P2P keeps its page in the same vocabulary, and the node logs page follows the local logs page. Previews are labelled as worked out in the browser. Per-GPU readings and node events are not drawn because the API does not return them. Failover moves to Swarm when distributed mode is on. Legacy fleet components and their styles are removed. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Swarm hub New specs for adding a node (each join method, the command, copy, waiting and found, approve, a single install, P2P, a phone), placement rules (sentences, where models are loaded, the preview matrix, the sheet and its preview, delete with undo) and failover on a cluster. Node detail covers its tabs, the drain preview and its dialog, resume, remove with the typed name, a node that stopped answering, and unload. The nodes specs follow the new structure and keep their intent: the table, filters, grouping, pagination, bulk actions, the map, and running models with stop, logs and the loading, error and empty states. The scheduling, failover, P2P, hub and smoke specs follow the renames. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the rebuilt Swarm pages Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add helpers for the Traffic pages Pure functions for what the pages work out from the usage ledger, the trace summary, the trace buffers and the resources reading: the shared time window, grouping, sorting and filtering of usage rows, chart series and axes that start at zero, the overview figures, per-model statistics, the state of a trace and the words for a failure, the backend operations that ran during a request, CSV export, the Prometheus metric list and scrape config, and a bounded buffer of host readings. A figure whose source cannot say is null, never zero. The trace summary call takes the window in hours, and a helper reads /metrics with its status. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Traffic hub: overview, usage, models, host, traces, middleware Traffic opens on an overview: five figures (requests, failed, p95, tokens in and out) and three charts, each naming its source. A second row of links reaches Usage, Models, GPU and host, Traces, Middleware and Prometheus, and one time window is shared by the first three. Usage groups by model, user or API key, filters, sorts, opens a row on its own chart, exports the rows it holds as CSV or JSON in the browser, and keeps the opt-in cost estimate and the quota forecast. A user who is not an admin sees only their own numbers. Models joins the ledger, the backend-operation buffer and the loaded models. GPU and host shows the current reading and two charts of readings taken since the page opened. Traces gets filters, a settings strip and an explained off state. An API request is a page: the error LocalAI recorded, a timeline with the backend operations that ran meanwhile, and bodies that stay closed until revealed. Middleware draws the pipeline as five steps and shows the rules of the selected step. Prometheus documents /metrics, checks it against the server and gives a scrape config to copy. Alerts is not built: LocalAI has no alert rules. Per-model latency percentiles, GPU utilisation and compare with the previous period are not drawn because the API does not return them. Legacy usage, trace and middleware styles and the usage source components are removed. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Traffic hub New specs for the overview (figures, charts with a data table and arrow key readout, failed and first-run and tracing-off states, the shared window, a phone), usage (group by, filters, sort, export, cost, quotas, a non-admin, empty and loading), models, GPU and host (snapshot, the since-opened labelling, a cluster), the traces list, a trace page (the real error, the timeline, reveal, no headers, a trace that left the buffer), Prometheus and the Middleware pipeline, with shared fixtures. The usage, traces, middleware, hub and smoke specs follow the new structure and keep their intent. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the Traffic hub Add an operations page for the Traffic tab: which record each page reads, what it leaves out and why, the trace page and its reveal, the GPU and host readings kept since the page opened, and the Prometheus endpoint. Link it from the operations index, the tracing page and the middleware page. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): let metric names wrap in the Prometheus table on a phone The long metric names pushed the type and "on this server" columns out of view. Names now wrap inside the table. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Settings with groups by intent, search, a pending bar and history The fifteen sections become eight groups by intent: memory and models, speed and defaults, backends and galleries, access and security, debugging and traces, agents and responses, swarm and sharing, look and feel. Search covers names, descriptions, keys and the old section name, and says where a result used to be. Edits wait in a bar with Discard, Show diff and Apply. The diff lists old and new values and the checks the browser can make: durations parse the way Go parses them, a GPU memory budget is one the server accepts, a gallery box holds JSON, and warnings repeat what the handler and the field text say. Apply sends only the changed keys. Undo saves the previous values again; it is a new save, not a rollback. History lists the changes applied from this browser, since LocalAI keeps no settings log, and Revert stages the old value. A value is marked as changed only where the built-in default is known from the CLI defaults. A row says "Applies now" or "Needs restart" only where the handler or the docs say so. Three things were wrong before and are fixed with the rebuild: the gallery boxes and the shared API keys box were sent under names the server ignores, the "Enable CSRF Protection" switch showed the disable flag the wrong way round, and every save restarted peer-to-peer networking because every field was sent. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Users and keys, Account, sign-in, invite and the 404 page Users and keys is a tabbed page under the Settings tab: people, invites and API keys. The people table filters by state and role, sorts, approves or disables (disabling offers an undo that sets the status back), and opens a side sheet for one person's features, model allow-list and limits. Role, password reset and delete sit in the row menu; delete asks for the name. Invites choose a lifetime of 1, 7 or 30 days and show the link once. API keys can be created with a lifetime, are shown once in full, can be paused, and are revoked after a ten second undo window in which nothing is sent. LocalAI lists keys only to their owner, so the tab shows the signed-in person's own keys and says so. Account has Profile, Security, API keys and Usage. Usage shows the last 30 days, tokens by model and the limits an admin set. The Security tab now shows for a GitHub or SSO account and says the password is not theirs to change. Sign-in asks for one field per step and draws a provider button only for a provider /api/auth/status lists. It has the notice for a sign-up that waits for approval, the first-admin screen, the key-only screen and the invite page. An address outside the app now gets the 404 page too, which names the address and lists the places the sidebar lists, with the same gates. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover Settings, Users and keys, Account, sign-in and the 404 page Settings: groups, search by name, key and old section, the changed marker only where a default is known, apply hints, the pending bar and diff, the checks, apply sending only changed keys, undo as a second save, discard, history, the CSRF inversion and the gallery and API key wire forms, and the phone layout. Users and keys: the table, filters, sort, approve, disable with undo, the row menu, the access sheet, invites, key creation with a one-time reveal, the ten second revoke with undo and with a page leave, and the non-admin redirect. Account, each sign-in variant (error, pending, first admin, key-only, invite, provider buttons) and the 404 page have specs too. Fixtures are shared with the screenshot scripts. Existing specs follow the new structure. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the rebuilt Settings, Users and keys, Account and sign-in pages Runtime settings: the eight groups and where each old section went, search, the pending bar, the diff and its checks, apply, undo, the history, and which settings show a default or an apply note and why. Authentication: the sign-in screen variants, the Account tabs, key lifetimes, the one-time key reveal, the revoke undo window, and the fact that keys are listed only to their owner. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): let the Settings undo toast stand alone and read back a generated P2P token The saved message and the undo toast sat on the same spot at the bottom of the page. The undo toast now carries the saved message. A new P2P token is made by the server when the page sends 0. The page reads it back after the save so the field shows the token and not the placeholder. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): fit the users table, API keys and Account figures on a phone On a phone the users table dropped its Role and Status columns off the screen edge with the row actions. The role and state now sit under the name, so the actions stay in view. API key rows no longer put the key icon on a line of its own, and the three Account figures keep one row. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the phone users table, reduced motion and the empty Account state Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): drop the apply note from three settings the save handler does not mention Size-aware eviction, automatic backend upgrades and development backends said Applies now, but nothing in the handler or the docs says when they take effect. A row now carries a note only where the code or the docs say so. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Voices and Faces as one identity family Voices is one page with three tabs: Speakers (voiceprints for recognising who is speaking), Speech voices (the text-to-speech reference library, kept apart because it is a different store) and From a recording (a link into the diarization workspace). Faces uses the same layout. Who is this and Same person? give the answer in a sentence with the real distance and cut-off, a word for how far inside the cut-off it sits, and a distance scale with the cut-off drawn on it. The cut-off slider re-reads the answer in the browser; the identify call sends the cut-off, and verify uses the threshold the model returns. The old confidence percentage is gone because it is not a probability. The server has no list call, so the people list stays in the browser and the page says so. After a search that asked for more people than it got back, a saved person the server did not return is marked, and people the server returned that the browser does not know are listed. Nothing is claimed from a short or cut-off search. Enrolling is a sheet: sample, name, labels, permission. A copy of the sample in the browser is opt-in, and an administrator can also keep the recording as a speech voice in the same step. Removing a person waits ten seconds behind an Undo toast and sends nothing before then. Errors say what happened (no face found, model missing, call failed), a blocked or missing microphone is explained, and a missing model or a missing permission renders a page that says what turns the feature on instead of a redirect. Analyze, detect and raw embedding move under More tools, with attribute guesses off by default. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Voices and Faces pages Specs for who is this (match, no match, working, failed, missing model), the cut-off slider, same person, a blocked, allowed and insecure microphone, the registry notes and the not-on-the-server marks, the enrol sheet and its opt-in copy, delete with undo on a fake clock, the disabled and no-permission states, the phone layout, reduced motion and Faces. Existing library and diarization specs follow the new structure and keep their intent. Node tests cover the distance words, scale layout, stored list and error mapping. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the Voices and Faces pages Add a WebUI section to the voice and face recognition pages: the two tools, the cut-off, what the people list is and why it can be stale, the undo window, and what is stored where. Point the Voice Library and Fish Audio notes at Build, Voices, Speech voices. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Build landing, Fine-tune, Quantize, Import and Explorer The Build landing says what each tool is for and what it needs from the machine: the installed backend, the GPU memory, RAM and disk the server reports, and a job that is running or the newest one when it failed. A tool that cannot run says why and what enables it. Fine-tune and Quantize share one page: set up, a check list that is redrawn as the form changes, a run view with progress, stages and a log, and a result with real next steps (export, import, chat, Models). The checks state only what the server reports. A job needs no estimate the server cannot make, so none is invented. Stop on a fine-tuning job asks whether to keep a checkpoint, a failed job shows the server's message, and a memory failure offers two changes that are applied to a copy of the setup. Import is a guided flow: source, review, import, done. The server returns no preview before an import starts, so the review reads the spelling of the source, prints the request the form will send and runs the checks that can be made early. The estimate that arrives when the import starts is set against free memory and disk. The ambiguity picker and the Write YAML tab stay. Explorer shows what GET /networks returns and lists a swarm with POST /network/add, with a join sheet that carries the token and commands. Build tools the account may not use say so instead of redirecting. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Build landing, the tool pages, Import and Explorer Specs for the landing (a tool ready, missing a backend, with no GPU, a running or failed job, a feature switched off, a member without admin, phone, reduced motion), the shared tool pattern for Fine-tune and Quantize (set up, live checks, start request, running with progress, chart and log, the stop choice, failure with the server message, finish with next steps, earlier jobs, the account-disabled page, phone), Import (source detection, review, checks, ambiguity, running with the estimate against free memory, done, Write YAML, phone) and Explorer (list, join, list a swarm, empty, not an explorer, retry, phone). Existing specs follow the new structure and keep their intent. Node tests cover the machine facts, tool status, checks, log lines, source detection, the import request and the join commands. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the Build tool pages, the import flow and the Explorer Fine-tuning and quantization now describe the set up, check, run and result steps and what the check list can and cannot say. The import section explains the review step and why the size and memory appear only after the import starts. The distributed page describes the Explorer list, the join sheet and what listing a swarm publishes. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): turn the hardware recommendations into a "Best for this machine" shelf The shelf in the Models inspector put five columns into a 400 px pane, so long model ids wrapped letter by letter underneath the size and the memory figures. Each row now stacks the tag, the id and the size and memory facts beside one Install button, and the id wraps inside its own column. Once a model is installed the shelf narrows to the best fit and keeps the others behind a "N more that fit" toggle. Specs cover the ranking, the layout, the narrowing and the install request against a gallery fixture that carries the 4K estimate the shelf sizes against. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): quiet the Studio tab markers and say their state in words The type tabs drew a saturated green dot for every modality that has a model. The dot now uses a text colour, filled when a model is installed and hollow when none is, and each tab carries "(model installed)" or "(no model installed)" as hidden text so the state is not only a colour. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): stack the model editor empty state and drop its section hues "No fields configured" sat in a flex row, so the icon, the title and the text ran together. It now uses the stacked empty-state layout. The section icons took a different status colour each (amber, red, green); they now share one quiet colour, with the accent on the current section. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): keep the Home memory sentence whole on a phone and drop side rails On a 390 px screen the memory strip clipped "2 models loaded" to make room for the figure. The sentence now takes the first line and the figure and device wrap under it. The sweep also removed coloured left rails from the editor section rail, the skill editor list, the install strip and the audio transform notice (now an outlined note), plus unused chat rules that carried rails and two glow animations that nothing referenced. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): line up hub pages, list the model templates, and stop clipped text Medium-width pages inside Build and Operate were centred while the tab bar above them was flush left, so the title started 60 px right of the first tab. They now start at the bar's edge. Add Model offered nine templates as a grid of identical cards with chip clouds and inline styles. It is now one list of rows, each with the field names it fills in on a single muted line. Two clipped strings are fixed: the Studio voice field cut its placeholder mid-word, and the phone job list ended the schedule line in an ellipsis. The recommendation shelf also separates size and memory with a dot, and the docs describe the shelf. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): keep the hidden Studio tab state inside its tab The hidden state text added to each type tab was absolutely positioned against the page, so on a phone it sat outside the scrolling tab row and widened the page by hundreds of pixels. The tab is now the containing block. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the on-disk sizes on the Installed table, model page and cleanup sheet The fixtures stub GET /api/models/storage. The default report is empty, so existing specs keep the gallery estimates. makeStorage() builds a report from files and the models that use them, the way the server does, and storageSpec() is a models directory with shared and missing files. New specs cover the Size column and its shared line, the fallback when the call fails or the user is not an admin, the files list on the model page, a missing file, the bytes a removal frees with shared files, and the cleanup findings. Node tests cover the storage helpers and the batch arithmetic. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): wait for the page before pressing keys and ticking the clock Two specs failed in loaded full runs and passed alone. The Alt+1 to Alt+7 spec pressed a key before the composer had armed its key handler. The capacity chart spec advanced the fake clock before the poller had mounted, so it counted fewer readings than it expected. Both now wait for the page to mount. The key spec retries a press that lands during a re-render, and the clock spec advances in small steps and polls for the row count. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): hide the Installed footer when the storage report is empty An empty report from the storage call made the footer read "0.0 GB on disk" next to sizes taken from the gallery estimate. An empty report says nothing about the disk, so the footer now shows only the model count. A spec covers it. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): keep Explore pane actions inside the pane The inspector actions sat in a non-wrapping flex row beside the title, so the buttons ran past the pane edge once it got narrow. The row now takes its own line and wraps. The primary action (Install, Retry, Open) comes first. Manage installation becomes a ghost button, and Open details moves to the end of the row, so one action stands out and the others are quiet. No action or test id is removed. Add a spec that checks, in light and dark at several widths and with a pane forced to 320 px, that every action stays inside the pane box and that the pane keeps its inner padding. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * refactor(ui): drop the type chips from the Studio composer The Studio tabs and the composer's type chips listed the same seven modes, so the page said the same thing twice. Keep the tabs as the one place to switch modes. The composer now shows the type it will open as a small label in its header. The type suggestion from the typed words stays as the quiet hint line under the prompt, and Alt+1 to Alt+7 still pick a type. The composer root carries data-type, data-types and data-missing so tests can read the state. Specs pick a type through a shared Alt+digit helper and read a missing model from the tab dot instead of a chip. Remove the unused chip locale strings and CSS. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * refactor(ui): remove the section crumb above page titles Page headers drew a small uppercase crumb with a short rule before it above the title. On the hub pages it repeated the hub name, so Build sat above a heading that also said Build. PageHeader now renders only the title, the supporting line and the actions. Drop the eyebrow prop, the route-derived section name, its CSS and the unused section helper, and remove the explicit eyebrow props from the pages that passed one. Pages stay reachable through the sidebar and the hub tab bar. Add a spec that checks several pages show their title with nothing ahead of it in the header. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * refactor(ui): remove left accent rails from tiles, rows and quotes Several surfaces marked state with a coloured strip on the left edge. Replace each one with a cue that is not a rail: - Stat cards lose the strip; the icon and value still carry the colour. - The highlighted card is a raised surface with a firmer edge. - The selected rail row is an accent wash with a hairline outline. - The status stripe on rail items is a small status dot. - The active failover row is a tinted row. - Quotes in markdown and chat prose are italic instead of barred. - The variant detail panel has a full hairline border. Add a spec that walks the main routes in light and dark and fails on a left border thicker than 1px, a sideways inset shadow, a narrow absolute strip in ::before or ::after, or a narrow tall child pinned to a left edge. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
1491 lines
107 KiB
Markdown
1491 lines
107 KiB
Markdown
+++
|
||
disableToc = false
|
||
title = "Distributed Mode"
|
||
weight = 71
|
||
url = "/features/distributed-mode/"
|
||
+++
|
||
|
||
Distributed mode enables horizontal scaling of LocalAI across multiple machines using **PostgreSQL** for state and node registry, and **NATS** for real-time coordination. Unlike the [P2P/federation approach]({{% relref "features/distributed_inferencing" %}}), distributed mode is designed for production deployments and Kubernetes environments where you need centralized management, health monitoring, and deterministic routing.
|
||
|
||
{{% notice note %}}
|
||
Distributed mode requires authentication enabled with a **PostgreSQL** database - SQLite is not supported. This is because the node registry, job store, and other distributed state are stored in PostgreSQL tables.
|
||
{{% /notice %}}
|
||
|
||
## Architecture Overview
|
||
|
||

|
||
|
||
**Frontends** are stateless LocalAI instances that receive API requests and route them to worker nodes via the **SmartRouter**. All frontends share state through PostgreSQL and coordinate via NATS.
|
||
|
||
**Workers** are generic processes that self-register with a frontend. They don't have a fixed backend type - the SmartRouter dynamically installs the required backend via NATS `backend.install` events when a model request arrives.
|
||
|
||
### Scheduling Algorithm
|
||
|
||

|
||
|
||
The SmartRouter uses **idle-first** scheduling with **preemptive eviction**:
|
||
1. If the model is already loaded on a node → use it (per-model gRPC address)
|
||
2. Drop any node without room to **store** the model on its models filesystem (see [Disk headroom](#disk-headroom))
|
||
3. If no node has the model → prefer nodes with enough free VRAM
|
||
4. Fall back to idle nodes (zero models), then least-loaded nodes
|
||
5. If no node has capacity → **evict the least-recently-used model with zero in-flight requests** to free a node
|
||
6. If all models are busy → wait (with timeout) for a model to become idle, then evict
|
||
7. Send `backend.install` NATS event with backend name + model ID → worker starts a new gRPC process on a dynamic port
|
||
8. SmartRouter calls gRPC `LoadModel` on the model-specific port, records in DB
|
||
|
||
Each model gets its own gRPC backend process, so a single worker can serve multiple models simultaneously (e.g., a chat model and an embedding model).
|
||
|
||
## Prerequisites
|
||
|
||
- **PostgreSQL** (with pgvector extension recommended for RAG) - used for node registry, job store, auth, and shared state
|
||
- **NATS** server - used for real-time backend lifecycle events and file staging
|
||
- All services must be on the same network (or reachable via configured URLs)
|
||
|
||
## Quick Start with Docker Compose
|
||
|
||
The easiest way to try distributed mode locally is with the provided Docker Compose file:
|
||
|
||
```bash
|
||
docker compose -f docker-compose.distributed.yaml up
|
||
```
|
||
|
||
This starts PostgreSQL, NATS, a LocalAI frontend, and one worker node. When you send an inference request, the SmartRouter automatically installs the needed backend on the worker and loads the model. See the file for details on adding GPU support, shared volumes, and additional workers.
|
||
|
||
{{% notice tip %}}
|
||
Use `docker-compose.distributed.yaml` for quick local testing. For production, deploy PostgreSQL and NATS as managed services and run frontends/workers on separate hosts.
|
||
{{% /notice %}}
|
||
|
||
## Frontend Configuration
|
||
|
||
The frontend is a standard LocalAI instance with distributed mode enabled. These flags are added to the `local-ai run` command:
|
||
|
||
| Flag | Env Var | Default | Description |
|
||
|------|---------|---------|-------------|
|
||
| `--distributed` | `LOCALAI_DISTRIBUTED` | `false` | Enable distributed mode |
|
||
| `--instance-id` | `LOCALAI_INSTANCE_ID` | auto UUID | Unique instance ID for this frontend |
|
||
| `--nats-url` | `LOCALAI_NATS_URL` | *(required)* | NATS server URL (e.g., `nats://localhost:4222`) |
|
||
| `--registration-token` | `LOCALAI_REGISTRATION_TOKEN` | *(empty)* | Token that workers must provide to register |
|
||
| `--registration-require-auth` | `LOCALAI_REGISTRATION_REQUIRE_AUTH` | `false` | Fail startup when distributed mode is enabled but the registration token is empty (node endpoints and worker file-transfer would otherwise be unauthenticated) |
|
||
| `--distributed-require-auth` | `LOCALAI_DISTRIBUTED_REQUIRE_AUTH` | `false` | **Umbrella switch.** Implies both `--nats-require-auth` and `--registration-require-auth` - one knob to lock down the NATS bus *and* the registration/file-transfer layer. Set this in production instead of the two granular flags. |
|
||
| `--auto-approve-nodes` | `LOCALAI_AUTO_APPROVE_NODES` | `false` | Auto-approve new worker nodes (skip admin approval) |
|
||
| `--distributed-shared-models` | `LOCALAI_DISTRIBUTED_SHARED_MODELS` | `false` | Assert that every node mounts the **same** models directory at the **same** path (a shared volume). When `true`, the router skips file staging entirely and workers load models directly from the shared path instead of re-downloading them. See [Shared models directory](#shared-models-directory). |
|
||
| `--distributed-disk-headroom-check` | `LOCALAI_DISTRIBUTED_DISK_HEADROOM_CHECK` | `true` | Reject worker nodes that lack free space to store the model, at scheduling time rather than partway through staging. When `false`, node selection ignores free disk; the check still runs and warns when it would have rejected every node. Also toggleable at runtime via the `distributed_disk_headroom_check` setting. See [Disk headroom](#disk-headroom). |
|
||
| `--auth` | `LOCALAI_AUTH` | `false` | **Must be `true`** for distributed mode |
|
||
| `--auth-database-url` | `LOCALAI_AUTH_DATABASE_URL` | *(required)* | PostgreSQL connection URL |
|
||
| `--backend-install-timeout` | `LOCALAI_NATS_BACKEND_INSTALL_TIMEOUT` | `15m` | How long the frontend waits for a worker to acknowledge a backend install before considering the request stalled. Raise it when workers pull large backend images over slow links. If a worker takes longer than this, the operation shows as "still installing in background" in the admin UI and clears once the worker finishes. |
|
||
| `--backend-upgrade-timeout` | `LOCALAI_NATS_BACKEND_UPGRADE_TIMEOUT` | `15m` | Same as the install timeout, applied to backend upgrades (force-reinstall). |
|
||
| `--model-load-timeout` | `LOCALAI_NATS_MODEL_LOAD_TIMEOUT` | *(derived from checkpoint size)* | Pins the deadline for the `LoadModel` gRPC call the frontend issues to a worker. Leave it unset: by default the deadline is **derived from the checkpoint's on-disk size** (see below), which is what the worker actually spends its load time reading. Set it only to pin a specific budget — the value is then used verbatim, including when it is *shorter* than the derived one, so an operator who wants fast failure gets it. |
|
||
| *(env only)* | `LOCALAI_MODEL_LOAD_WAIT` | `60s` | How long an inference request waits for a model that is still cold-loading onto a worker before it is answered with `503`, a `Retry-After` header and live staging progress. The request is served the moment the model becomes ready, so a model already most of the way staged needs no client retry. Set to `0` to wait as long as the load takes — only safe when no ingress or load balancer with an idle timeout sits in front. See [Requests for a model that is still loading](#requests-for-a-model-that-is-still-loading). |
|
||
| `--node-heartbeat-checkpoint` | `LOCALAI_NODE_HEARTBEAT_CHECKPOINT` | `60s` | Minimum gap between **durable** heartbeat writes for a worker node. A beat that only carries a fresher timestamp is kept in memory until this interval elapses instead of being written to PostgreSQL; every reported field is compared against the value last written rather than merely tested for presence, so a node's first beat, a changed total VRAM / total disk / GPU vendor, and a free VRAM / RAM / disk reading that has moved more than 256 MiB from the written value all still write immediately, and a node that is not active is never suppressed. Set it below the worker's `--heartbeat-interval` to restore a write per beat. See [Heartbeat writes and stale-node detection](#heartbeat-writes-and-stale-node-detection). |
|
||
| `--stale-node-threshold` | `LOCALAI_STALE_NODE_THRESHOLD` | `5m` | How long a node may go without a **durable** heartbeat before the health monitor marks it `offline`. Because `--node-heartbeat-checkpoint` holds back a beat that only carries a fresher timestamp, this has to stay comfortably wider than that interval: raising the checkpoint without raising this marks healthy, beating nodes offline. Neither the per-model gRPC health check nor request-time failure reads `last_heartbeat`, so neither is affected by this knob. See [Heartbeat writes and stale-node detection](#heartbeat-writes-and-stale-node-detection). |
|
||
| `--model-config-resync-interval` | `LOCALAI_MODEL_CONFIG_RESYNC_INTERVAL` | `30s` | How often each frontend compares its model configs with the shared models directory, to apply a change whose NATS message it missed. A frontend that missed a message serves the old config for at most this long. See [Model configs across frontends](#model-configs-across-frontends). |
|
||
| `--expose-node-header` | `LOCALAI_EXPOSE_NODE_HEADER` | `false` | When enabled, inference responses carry an `X-LocalAI-Node` header with the ID of the worker node that served the request. Coverage spans the OpenAI-compatible endpoints (chat completions, completions, embeddings, audio transcriptions, audio speech / TTS, image generations, image inpainting), the Jina rerank endpoint (`/v1/rerank`), the VAD endpoints (`/v1/vad`, `/vad`), and the Anthropic Messages (`/v1/messages`) and Ollama (`/api/chat`, `/api/generate`, `/api/embed`) shims. Useful for debugging, observability and load-balancer attribution. Off by default: the node ID reveals internal cluster topology and should not be exposed on a public endpoint. Best-effort: under heavy concurrency for the same model across multiple replicas, the header may reflect a recent routing decision rather than this exact request's. Acceptable for observability and debugging. |
|
||
|
||
### The model load deadline scales with the checkpoint
|
||
|
||
The `LoadModel` deadline starts *after* the backend is installed and the model files are staged, so it covers only the worker backend's own checkpoint read and pipeline init. That work is proportional to the bytes on disk, which makes any fixed deadline a model-size cliff rather than a timeout: a 70 GB video checkpoint on a Jetson Thor worker failed reproducibly against the old fixed 5m default (`rpc error: code = DeadlineExceeded` after 953.5s of wall clock, roughly 11m of which was backend install and staging), and simply raising the constant would only move the cliff to the next larger model while making a genuinely wedged *small* model hang for the whole inflated duration.
|
||
|
||
So the deadline is derived per model:
|
||
|
||
```
|
||
budget = 5m + 20s per GiB of checkpoint, capped at 6h
|
||
```
|
||
|
||
| Checkpoint | Derived budget |
|
||
|---|---|
|
||
| 2 GB | 5m40s |
|
||
| 70 GB | 28m20s |
|
||
| 600 GB | 3h25m |
|
||
|
||
The per-GiB rate is deliberately pessimistic — it corresponds to reading weights at about 54 MB/s, below what any supported storage sustains — because the two errors are not symmetric: a budget that is too long costs only *failure latency* on a load that was going to fail anyway, while a budget that is too short causes a guaranteed false failure on a load that was perfectly healthy.
|
||
|
||
The size is measured from the model files on the frontend's disk, over the same set of paths that get staged to the worker. If those files are not present locally — a backend handed a bare HuggingFace repo id fetches its own weights on the worker — there is nothing to measure and the budget stays at the plain 5m default. Pin `LOCALAI_NATS_MODEL_LOAD_TIMEOUT` for those models if their load is slow.
|
||
|
||
When the budget *is* exceeded, the error names the budget, the checkpoint size it was derived from, and the knob that overrides it, instead of surfacing a bare `context deadline exceeded`.
|
||
|
||
### The cold-load lock ceiling
|
||
|
||
The router also bounds how long a single cold load may hold the per-model advisory lock, so a worker that dies mid-install cannot pin every other replica's request for that model. That bound is **derived**, not configured, and it is based on *progress* rather than on wall-clock time.
|
||
|
||
The load starts with a base budget of `max(backend-install-timeout + model-load-timeout + 5m, 25m)` — with the defaults, `15m + 5m + 5m = 25m`. That budget covers the steps that report no progress: node selection, backend install, and the remote `LoadModel` call. Raising either timeout widens it in step, so a longer load deadline is never clipped.
|
||
|
||
That base is the hold's *starting* budget, not its maximum. Because the derived load budget above can exceed it — a 70 GB checkpoint's 28m20s against a 25m base — the hold is widened again as the router enters the load phase, by the derived budget plus the same 5m of slack. Without that step the ceiling would cancel a load that was still comfortably inside its own deadline.
|
||
|
||
While **model files are staging**, however, the deadline extends every time staging does real work, and expires only once staging has been silent for a 5-minute stall window. Real work means uploaded bytes, and also the resumable-upload verify phase: when a shard is already present on the worker from an earlier attempt, the frontend HEADs it and hashes the local copy to confirm it matches, then skips the transfer. That phase uploads nothing at all — on a 70 GB model resuming with 56 GB already staged it ran for six-plus consecutive minutes at ~45s per shard — so hashing counts as progress too. Otherwise a resumed transfer would be mistaken for a wedged one. Staging time is a function of checkpoint size and available bandwidth, not a constant: a 70 GB model at 26 MB/s needs about 45 minutes, and a 600 GB checkpoint needs hours. A fixed ceiling would therefore be a model-size cliff — every increase just moves the cliff to the next larger model. Extending on progress means a large model transfers for as long as it legitimately needs, while a worker that dies mid-transfer still releases the lock within the stall window.
|
||
|
||
An absolute cap of 24h ends the hold even if progress keeps arriving, so a degenerate peer trickling a few bytes at a time cannot pin the lock forever. No configuration is needed for either value; both are sized well above any legitimate transfer.
|
||
|
||
### Requests for a model that is still loading
|
||
|
||
A cold load in distributed mode is a long-running background job: install the backend, stage multi-GB model files to the worker, then load the checkpoint. Staging a 35.7 GB GGUF onto a fresh worker takes roughly twenty minutes on a fast LAN — far longer than any HTTP request can be held open.
|
||
|
||
So the load does **not** run on the request. The first request for an unloaded model claims a durable **model load job** — that claim takes milliseconds and is the only part that holds the per-model advisory lock — and the job then runs in the background on the frontend replica that claimed it. Every other request for the same model, on any replica, attaches to that job as a waiter:
|
||
|
||
- It is **served the moment the model is ready**, with no client-side retry. A model already 90% staged usually needs no second request.
|
||
- It never starts a duplicate load and never blocks on the database lock. (Before this split, concurrent requests blocked on `pg_advisory_lock` for the whole load and were killed by the PostgreSQL role's `statement_timeout` — `SQLSTATE 57014` — so from the operator's seat the model simply never loaded.)
|
||
- If the load fails, the waiter gets the *real* cause (`worker out of disk`), not an anonymous timeout.
|
||
- If the client disconnects, the load keeps going. It belongs to the job record, not to the request.
|
||
- The owner of a job holds a 30 second lease, which it renews on every heartbeat. The database clock decides whether a lease has expired, so a frontend with a wrong clock cannot expire a live lease or keep a dead one. If an owner cannot renew for a whole lease, it stops its own load.
|
||
- If an owner dies, its job is marked failed once the lease runs out, with or without a new request. A request that arrives then reads the cause. The model is held for a 2.5 minute stop window, because remote work may still run. After that the next request starts a new attempt. No manual cleanup is needed.
|
||
- A failure that is known to have ended the remote work (an error from the backend, or a failure before a node was chosen) is kept for 15 seconds only, so every waiter reads the same cause and the next request can retry.
|
||
- Each attempt has its own generation. If a job is replaced, the old owner notices at its next heartbeat and stops its load. Its late writes to the job and to the replica table are rejected.
|
||
- If the job table cannot be read, a cold load fails instead of running without a job. Models that are already loaded keep serving, because routing to a loaded replica does not read the job table.
|
||
|
||
When the wait budget (`LOCALAI_MODEL_LOAD_WAIT`, default `60s`) runs out, the request is answered with `503`, a `Retry-After` header, and a body that says exactly where the load is:
|
||
|
||
```json
|
||
{
|
||
"error": {
|
||
"message": "model Qwen3.6-27B-MTP-GGUF is staging on node nvidia-thor (41%, ETA ~11m)",
|
||
"type": "model_loading",
|
||
"code": "model_loading"
|
||
},
|
||
"loading": {
|
||
"model": "Qwen3.6-27B-MTP-GGUF",
|
||
"state": "staging",
|
||
"node": "nvidia-thor",
|
||
"progress": 41.2,
|
||
"bytes_sent": 14730000000,
|
||
"total_bytes": 35776484480,
|
||
"file_index": 1,
|
||
"total_files": 2,
|
||
"eta_seconds": 660
|
||
}
|
||
}
|
||
```
|
||
|
||
The `error` envelope keeps OpenAI clients working unchanged; `loading` is additive, so a client that understands it renders progress instead of an error. `eta_seconds` is derived from the job's own observed transfer rate and is **omitted rather than guessed** until enough bytes have moved for that rate to mean anything — a confidently wrong ETA on a twenty-minute wait is worse than none. `state` is one of `pending` (choosing a node), `installing`, `staging` (transferring files) or `loading` (the worker is reading the checkpoint).
|
||
|
||
The chat UI renders this state inline and retries automatically once the model reports ready. Poll `GET /api/models/{id}/load-status` for the same `loading` object at any time.
|
||
|
||
{{% notice note %}}
|
||
A frontend replica that dies mid-load does not wedge the model: the job row carries a lease, and a job whose lease ran out is failed and then released. The lease is renewed on a timer, not on byte progress, because a checkpoint load legitimately transfers zero bytes for many minutes.
|
||
{{% /notice %}}
|
||
|
||
#### The worker bounds the work it runs
|
||
|
||
The frontend owns the job row, but the real work runs in a backend process on a worker. The worker therefore watches each load too. A load is an **operation** named by the job's generation:
|
||
|
||
- The install request carries the operation id and the longest the load may run, as a duration, so a worker clock that is wrong changes nothing. The backend starts in its own process group.
|
||
- The frontend renews the operation every few seconds and completes it when the load finishes. The worker kills the whole process group when no renewal arrives for 90 seconds, or when the deadline passes. A backend that already reports `READY` is never killed: a lost completion message must not destroy a model that serves.
|
||
- A stop names the operation, the process key and, when known, the address and process instance. The worker refuses unless they all match its own records. There is no fallback to "any running backend". A stop for a load that already finished leaves the serving model alone.
|
||
- A worker that is killed cannot stop its backends. It records each backend's process group in a small file under its data directory, and the next worker kills every group listed there before it serves (Linux; the start time of the leader guards against a recycled pid). A new incarnation, reported on the next heartbeat, then confirms the failed loads on that node. The worker also skips a gRPC port that something already listens on, and a readiness answer only counts while the worker's own backend process is alive.
|
||
- If the worker answers a renewal with "unknown operation" three times in a row, the frontend fails the load at once instead of waiting for the load budget. The worker lost the operation, so the work is gone.
|
||
|
||
When a load fails after remote work may have started (a timeout, a cancel, a lost lease), the owner stops the operation immediately. An acknowledged stop shortens the hold to the 15 second report window. A silent worker keeps the hold at the stop window (2.5 minutes), and the reconciler retries the stop on every pass until the worker answers or the window ends. Nothing needs manual cleanup.
|
||
|
||
| Setting | Value | Meaning |
|
||
|---------|-------|---------|
|
||
| Lease TTL | 30 s | How long a job's lease lasts after each renewal |
|
||
| Worker kill TTL | 90 s | No renewal for this long: the worker kills the operation |
|
||
| Stop window | 150 s | How long a failed load holds the model if the worker never confirms |
|
||
| Report window | 15 s | How long a failure with confirmed-ended work is kept |
|
||
|
||
A model held by a failed job answers `503` with `Retry-After` set to the seconds until the hold ends, and the real cause in the body.
|
||
|
||
#### Cancelling a load
|
||
|
||
`POST /api/models/{id}/load-cancel` (admin only) cancels one load attempt. The body names the exact attempt, as `GET /api/models/{id}/load-status` reports it:
|
||
|
||
```json
|
||
{"job_id": "0b6e4a3c-5c1d-4d52-8f0a-0f3c9e0b8f11"}
|
||
```
|
||
|
||
| Status | Meaning |
|
||
|--------|---------|
|
||
| `200` | `state: stopped` (the worker confirmed) or `state: gone` (no such load any more) |
|
||
| `202` | `state: stopping`. The cancel is recorded and the stop is pending. The model is released after `retry_after` seconds regardless. |
|
||
| `400` | The body is not `{"job_id": "..."}` |
|
||
| `404` | Unknown model, or the server is not distributed |
|
||
| `409` | A different attempt is current. The body carries its `current_job_id`. |
|
||
|
||
The call is idempotent. A repeat retries the stop and never extends the hold. A load that has not been placed on a node yet can be cancelled too. Unloading a model on a node, draining a node, and removing a node all cancel the loads placed there through the same stop path, and an unload still unloads the loaded replicas. The replica rows of a cancelled attempt are removed as soon as the worker confirms the stop. The `cancel_model_load` tool of the assistant calls the same service.
|
||
|
||
`load-status` also reports `job_id`, `lease_expires_in`, `cancel_requested`, `last_error`, `stopping`, `stop_deadline` and `retry_after`. A database error is a `503`, never an empty answer.
|
||
|
||
#### Rolling upgrades
|
||
|
||
Upgrade the frontends first. A worker that predates operations ignores the new request fields and does not report `reports_operations`. The frontend then treats the node as legacy: it cannot confirm a stop, so a failed load holds the model for the 45 minute load deadline, as it did before leases existed, and never longer. For such a node the stop, including a cancel, is sent by exact process address, never by model name. If the address is not known, no stop is claimed, and the model is held for the 45 minutes. A new worker that gets an install from an older frontend tracks it as an anonymous operation: it kills it at its deadline only, never for missing renewals.
|
||
|
||
### NATS JWT authentication (recommended for production)
|
||
|
||
By default, NATS connections are anonymous: any client that can reach port `4222` may publish control-plane subjects such as `nodes.<id>.backend.install`. Enable JWT auth to scope workers to their own node subjects and give the frontend a dedicated service credential.
|
||
|
||
| Flag | Env Var | Description |
|
||
|------|---------|-------------|
|
||
| `--nats-account-seed` | `LOCALAI_NATS_ACCOUNT_SEED` | Account signing seed (`SU...`). The frontend mints a per-node user JWT at registration (`nats_jwt` in the register response). |
|
||
| `--nats-service-jwt` | `LOCALAI_NATS_SERVICE_JWT` | User JWT for the frontend (and optional fallback for agent workers) to publish install/upgrade and related subjects. |
|
||
| `--nats-service-seed` | `LOCALAI_NATS_SERVICE_SEED` | User signing seed (`SU...`) paired with the service JWT. |
|
||
| `--nats-worker-jwt-ttl` | `LOCALAI_NATS_WORKER_JWT_TTL` | Lifetime of minted worker JWTs (default `24h`). |
|
||
| `--nats-require-auth` | `LOCALAI_NATS_REQUIRE_AUTH` | Fail startup if JWT credentials are missing when distributed mode is enabled. |
|
||
|
||
### NATS TLS / mTLS (optional)
|
||
|
||
Use `tls://` in `--nats-url` / `LOCALAI_NATS_URL` for encrypted transport. When the server uses a private CA or requires client certificates, set:
|
||
|
||
| Flag | Env Var | Description |
|
||
|------|---------|-------------|
|
||
| `--nats-tls-ca` | `LOCALAI_NATS_TLS_CA` | PEM file to verify the NATS server (private CA) |
|
||
| `--nats-tls-cert` | `LOCALAI_NATS_TLS_CERT` | Client certificate for NATS mTLS |
|
||
| `--nats-tls-key` | `LOCALAI_NATS_TLS_KEY` | Client private key (required with `--nats-tls-cert`) |
|
||
|
||
The same env vars apply to backend workers and `local-ai agent-worker`. If the server cert is already trusted by the OS, `tls://` alone is enough.
|
||
|
||
**Worker register response** (when minting is enabled and the node is approved):
|
||
|
||
```json
|
||
{
|
||
"id": "…",
|
||
"nats_jwt": "eyJ…",
|
||
"nats_user_seed": "SU…"
|
||
}
|
||
```
|
||
|
||
Workers connect with that JWT and seed automatically (shown once; store securely). Override with `LOCALAI_NATS_JWT` / `LOCALAI_NATS_USER_SEED` if needed. Set `LOCALAI_NATS_REQUIRE_AUTH=true` on workers when the bus requires credentials.
|
||
|
||
When `LOCALAI_NATS_REQUIRE_AUTH=true` and no static credentials are provided, a worker that registers while still **pending admin approval** keeps re-registering (with backoff) until an admin approves it and the frontend mints its JWT - it does not start unauthenticated. This retry is **bounded**: if the node is never approved (or no credentials are minted) after a large number of attempts, the worker exits non-zero so the failure is visible (a crash-looping or failed worker) rather than hanging silently. Minted worker JWTs are also **refreshed automatically** before they expire (the worker re-registers at ~75% of the JWT lifetime), so long-running workers survive past `LOCALAI_NATS_WORKER_JWT_TTL`; the NATS connection picks up the new JWT on its next reconnect. If refresh fails persistently, the worker exits (to restart and re-acquire) rather than drifting toward an expired, unrenewable JWT. Statically configured (`LOCALAI_NATS_JWT`) and service (`LOCALAI_NATS_SERVICE_JWT`) credentials are used as-is and not refreshed.
|
||
|
||
Generate operator/account material with [`scripts/nats-auth-setup.sh`](https://github.com/mudler/LocalAI/blob/master/scripts/nats-auth-setup.sh) (requires [nsc](https://docs.nats.io/running-a-nats-service/configuration/securing_nats/auth_intro/nsc)). Configure the NATS server with account resolver JWTs before enabling `LOCALAI_NATS_REQUIRE_AUTH`.
|
||
|
||
{{% notice note %}}
|
||
`LOCALAI_AUTH` (HTTP users/sessions) and NATS JWTs are separate: end-user API keys do not connect to NATS. HTTP registration still uses `LOCALAI_REGISTRATION_TOKEN`.
|
||
{{% /notice %}}
|
||
|
||
### Optional: S3 Object Storage
|
||
|
||
For multi-host deployments where workers don't share a filesystem, S3-compatible storage enables distributed file transfer (model files, configs):
|
||
|
||
| Flag | Env Var | Default | Description |
|
||
|------|---------|---------|-------------|
|
||
| `--storage-url` | `LOCALAI_STORAGE_URL` | *(empty)* | S3 endpoint URL (e.g., `http://minio:9000`) |
|
||
| `--storage-bucket` | `LOCALAI_STORAGE_BUCKET` | `localai` | S3 bucket name |
|
||
| `--storage-region` | `LOCALAI_STORAGE_REGION` | `us-east-1` | S3 region |
|
||
| `--storage-access-key` | `LOCALAI_STORAGE_ACCESS_KEY` | *(empty)* | S3 access key |
|
||
| `--storage-secret-key` | `LOCALAI_STORAGE_SECRET_KEY` | *(empty)* | S3 secret key |
|
||
|
||
When S3 is not configured, model files are transferred directly from the frontend to workers via **HTTP** - no shared filesystem needed. Each worker runs a small HTTP file transfer server alongside the gRPC backend process. This is the default and works out of the box.
|
||
|
||
For high-throughput or very large model files, S3 can be more efficient since it avoids streaming through the frontend.
|
||
|
||
### Shared models directory
|
||
|
||
If every node (frontend and workers) mounts the **same** models directory at the **same** path - for example a shared volume or network filesystem, as shown in the "Shared Volume Mode" section of `docker-compose.distributed.yaml` - the model files are already present on each worker at their canonical path. In that case staging is wasted work: it copies files that already exist into a per-model subdirectory the worker then loads from, which shows up as a re-download of a model you already have.
|
||
|
||
Set `LOCALAI_DISTRIBUTED_SHARED_MODELS=true` (or `--distributed-shared-models`) on the frontend to skip staging entirely. The router then leaves the model's absolute paths untouched and the worker loads them directly from the shared volume.
|
||
|
||
This flag is a contract you assert: all nodes must mount identical paths. Leave it off (the default) when workers have independent models directories - the frontend stages files to them over HTTP (or S3) as described above.
|
||
|
||
### Which files are staged
|
||
|
||
The frontend stages the files that the model config names (`parameters.model`, `mmproj`, draft model, LoRA adapters and similar fields). It also stages every other file that the model declares:
|
||
|
||
- The `files:` of the gallery entry or `/import-model` import that installed the model. LocalAI records these in `._gallery_<name>.yaml` next to the model config.
|
||
- The `download_files:` of the model config.
|
||
|
||
A backend can read files that the config does not name. For example, llama.cpp opens all shards of a split GGUF (`<name>-00002-of-00004.gguf` and the rest) from the directory of the first shard. The worker cannot see the frontend's models directory, so it gets only the files that the frontend stages.
|
||
|
||
If you write a model config by hand and the model has files like these, list them under `download_files:`. If you do not, the worker gets only the first shard and the load fails with `failed to load GGUF split`.
|
||
|
||
The file sizes used for the load deadline and for the disk headroom check include all of these files.
|
||
|
||
### Model artifact staging
|
||
|
||
For managed Hugging Face artifacts, the controller resolves the repository and
|
||
downloads every selected file. Workers receive the committed snapshot through
|
||
the existing directory stager. They never receive `HF_TOKEN` and do not contact
|
||
Hugging Face for managed artifacts.
|
||
|
||
With `LOCALAI_DISTRIBUTED_SHARED_MODELS` enabled, workers use the shared
|
||
absolute snapshot path and skip transfer. Otherwise, the controller stages the
|
||
complete snapshot tree to each worker before loading the backend.
|
||
|
||
{{% notice warning %}}
|
||
Every controller and worker must have enough disk space for its own snapshot
|
||
copy unless shared-models mode is enabled. Account for temporary partial files
|
||
during installation as well as the committed snapshot.
|
||
{{% /notice %}}
|
||
|
||
{{% notice warning %}}
|
||
The worker HTTP file transfer server is authenticated by `LOCALAI_REGISTRATION_TOKEN`. If the token is **empty**, the server **fails open** - anyone who can reach the port gets read/write access to the worker's models/staging/data directories (a remote model-poisoning / exfiltration vector). The worker logs a loud warning at startup in this case. Always set `LOCALAI_REGISTRATION_TOKEN` in distributed mode, and set `LOCALAI_DISTRIBUTED_REQUIRE_AUTH=true` (frontend **and** workers) to make a missing token *or* missing NATS credentials a hard startup error rather than a silent fail-open. Firewall the file-transfer port (gRPC base − 1) so only the frontend can reach it.
|
||
{{% /notice %}}
|
||
|
||
### Watching Backend Installs
|
||
|
||
While a worker downloads a backend, the admin operations strip at the top
|
||
of the UI shows real-time progress: a percentage, and, when the install
|
||
targets several workers, a roll-up of how far the fan-out has got,
|
||
`2 of 5 nodes done`. Per-file byte counts are not on the strip; they are in
|
||
the per-node detail below.
|
||
|
||
The per-node detail is on the **Operate → Activity** page
|
||
([Activity]({{% relref "operations/activity" %}})). When an install targets
|
||
more than one worker, an **N nodes** tag appears on the operation card, with
|
||
one row per worker showing:
|
||
|
||
- A status pill: **Queued** (gray), **Downloading** (blue), **Worker busy**
|
||
(yellow), **Done** (green), or **Failed** (red).
|
||
- The file currently being downloaded with current/total bytes and percentage.
|
||
- A thin per-node progress bar.
|
||
- Any error returned by the worker.
|
||
|
||
The yellow **Worker busy** pill means the worker took longer than
|
||
`--backend-install-timeout` to acknowledge but is most likely still
|
||
working in the background. The admin UI clears it as soon as the worker
|
||
finishes; no action is required from the operator.
|
||
|
||
If a worker is running an older LocalAI release that does not report
|
||
progress, its row in the breakdown will still show terminal status
|
||
(queued / done / failed / worker busy) but no per-file progress.
|
||
|
||
The **Record** on that page - what model and backend installs and removals have
|
||
finished - is read from PostgreSQL rather than from the replica's memory. Every
|
||
replica reports the same record, it survives restarts, a replica added by a
|
||
scale-out or a rolling deploy reports it in full, and **Clear history** clears
|
||
it for every replica.
|
||
|
||
## Worker Configuration
|
||
|
||
Workers are started with the `worker` subcommand. Each worker is generic - it doesn't need a backend type at startup:
|
||
|
||
```bash
|
||
local-ai worker \
|
||
--register-to http://frontend:8080 \
|
||
--registration-token changeme \
|
||
--nats-url nats://nats:4222
|
||
```
|
||
|
||
| Flag | Env Var | Default | Description |
|
||
|------|---------|---------|-------------|
|
||
| `--addr` | `LOCALAI_SERVE_ADDR` | `0.0.0.0:50051` | gRPC listen address |
|
||
| `--grpc-max-port` | `LOCALAI_GRPC_MAX_PORT` | `65535` | Highest port the worker may assign to a backend gRPC process. Each backend gets its own port, allocated upward from the base port, so the width of `[base port, this]` caps how many backends this worker can run at once (see [Backend gRPC port range](#backend-grpc-port-range)) |
|
||
| `--advertise-addr` | `LOCALAI_ADVERTISE_ADDR` | *(auto)* | Address the frontend uses to reach this node (see below) |
|
||
| `--http-addr` | `LOCALAI_HTTP_ADDR` | gRPC port - 1 | HTTP file transfer server bind address |
|
||
| `--advertise-http-addr` | `LOCALAI_ADVERTISE_HTTP_ADDR` | *(auto)* | HTTP address the frontend uses for file transfer |
|
||
| `--ephemeral-staging-byte-limit` | `LOCALAI_EPHEMERAL_STAGING_BYTE_LIMIT` | `0` (automatic) | Maximum bytes held by request-input staging across the worker's HTTP staging directory and S3 cache. Automatic mode uses the smaller of 10 GiB and 10% of filesystem capacity. |
|
||
| `--ephemeral-staging-min-free-bytes` | `LOCALAI_EPHEMERAL_STAGING_MIN_FREE_BYTES` | `0` (automatic) | Free filesystem space preserved while staging request inputs. Automatic mode uses the larger of 1 GiB and 5% of filesystem capacity. |
|
||
| `--register-to` | `LOCALAI_REGISTER_TO` | *(required)* | Frontend URL for self-registration |
|
||
| `--node-name` | `LOCALAI_NODE_NAME` | hostname | Human-readable node name |
|
||
| `--registration-token` | `LOCALAI_REGISTRATION_TOKEN` | *(empty)* | Token to authenticate with the frontend |
|
||
| `--registration-require-auth` | `LOCALAI_REGISTRATION_REQUIRE_AUTH` | `false` | Refuse to start the HTTP file-transfer server when no registration token is set (it would otherwise fail open) |
|
||
| `--distributed-require-auth` | `LOCALAI_DISTRIBUTED_REQUIRE_AUTH` | `false` | Umbrella switch implying both `--registration-require-auth` and `--nats-require-auth` |
|
||
| `--heartbeat-interval` | `LOCALAI_HEARTBEAT_INTERVAL` | `10s` | Interval between heartbeat pings |
|
||
| `--nats-url` | `LOCALAI_NATS_URL` | *(required)* | NATS URL for backend installation and file staging |
|
||
| `--nats-jwt` | `LOCALAI_NATS_JWT` | *(empty)* | Optional override for the `nats_jwt` returned at registration |
|
||
| `--nats-user-seed` | `LOCALAI_NATS_USER_SEED` | *(empty)* | Optional override for `nats_user_seed` from registration |
|
||
| `--nats-require-auth` | `LOCALAI_NATS_REQUIRE_AUTH` | `false` | Require NATS JWT+seed (from registration or env) |
|
||
| `--nats-tls-ca` | `LOCALAI_NATS_TLS_CA` | *(empty)* | PEM file for NATS server CA |
|
||
| `--nats-tls-cert` | `LOCALAI_NATS_TLS_CERT` | *(empty)* | Client certificate for NATS mTLS |
|
||
| `--nats-tls-key` | `LOCALAI_NATS_TLS_KEY` | *(empty)* | Client private key for NATS mTLS |
|
||
| `--backends-path` | `LOCALAI_BACKENDS_PATH` | `./backends` | Path to backend binaries |
|
||
| `--models-path` | `LOCALAI_MODELS_PATH` | `./models` | Path to model files |
|
||
| `--vram-budget` | `LOCALAI_VRAM_BUDGET` | *(empty)* | Cap the VRAM this node advertises for model placement, as a percentage (e.g. `80%`) or an absolute amount (e.g. `12GB`). Empty uses all detected VRAM. See [Per-node VRAM budget](#per-node-vram-budget). |
|
||
|
||
{{% notice tip %}}
|
||
**Advertise address:** The `--addr` flag is the local bind address for gRPC. The `--advertise-addr` is the address the frontend stores and uses to reach the worker via gRPC. If not set, the worker auto-derives it by replacing `0.0.0.0` with the OS hostname (which in Docker is the container ID, resolvable via Docker DNS). Set `--advertise-addr` explicitly when the auto-detected hostname is not routable from the frontend (e.g., in Kubernetes, use the pod's service DNS name).
|
||
|
||
**HTTP file transfer:** Each worker also runs a small HTTP server for file transfer (model files, configs). By default it listens on the gRPC base port - 1 (e.g., if gRPC base is 50051, HTTP is on 50050). gRPC ports grow upward from the base port as additional models are loaded. Set `--advertise-http-addr` if the auto-detected address is not routable from the frontend.
|
||
{{% /notice %}}
|
||
|
||
### Ephemeral request-input storage
|
||
|
||
Workers reserve local capacity before accepting per-request audio, image, and other ephemeral inputs. The limit covers both direct HTTP staging and the worker's S3 download cache. A request is rejected before inference when accepting its input would exceed the byte limit or the configured free-space headroom. One request-scoped cleanup operation releases all exact input keys and their reservations after inference, while a one-hour recovery sweep removes abandoned files after crashes. The sweep runs at startup and every 15 minutes, preserves active requests, and considers the newest file in each request directory.
|
||
|
||
Set both capacity variables to positive byte counts when a worker needs fixed limits. Leaving either value at zero selects its filesystem-based default. These settings apply only below the two `ephemeral` roots; model, data, and configuration files are excluded.
|
||
|
||
### Worker Health Probes
|
||
|
||
The worker's HTTP server (base port - 1, default 50050) exposes two unauthenticated probes:
|
||
|
||
| Endpoint | Meaning |
|
||
|----------|---------|
|
||
| `/healthz` | **Liveness.** 200 whenever the process is up and serving. Deliberately independent of readiness, so a brief NATS outage does not trigger a restart storm across every worker. |
|
||
| `/readyz` | **Readiness.** 200 only when the worker is registered, its NATS connection is live, *and* every backend process it is currently serving answers a short TCP dial on its gRPC address; 503 otherwise. A worker holding no backends is ready, because idle is a healthy state, and so is one whose backends are still starting up. |
|
||
|
||
`/readyz` reports something the frontend cannot see on its own. The node registry's `status` and `last_heartbeat` are driven by an HTTP heartbeat to the frontend, which is a different network path from NATS — a worker can keep heartbeating while its NATS link is dead, and so appear `healthy` in the registry while being unable to receive any work. The local probe closes that gap.
|
||
|
||
The same applies to the data path. A worker can hold a live NATS link while the backend processes it believes it is running have died, so it reports healthy while every load routed to it fails. `/readyz` therefore also dials the recorded gRPC address of each backend the worker is serving, and a worker whose backend port refuses connections drops out of rotation instead of absorbing work it cannot serve.
|
||
|
||
Only backends in the middle of their lifecycle are dialled. A backend that is still starting is skipped until its gRPC server has answered a health check, which can take 10 to 15 seconds on a slow node, and a backend that is stopping is skipped from the moment shutdown begins. Neither a cold start nor an ordinary shutdown makes a worker report 503, so a Kubernetes `readinessProbe` at the usual 10s period does not pull a worker out of rotation every time it loads a model.
|
||
|
||
The container image's `HEALTHCHECK` detects worker mode and probes this endpoint automatically; no `HEALTHCHECK_ENDPOINT` override is needed. Set `HEALTHCHECK_ENDPOINT` only to pin an explicit URL.
|
||
|
||
### Worker Address Configuration
|
||
|
||
The simplest way to configure a worker's network address is with a single variable:
|
||
|
||
| Variable | Description |
|
||
|----------|-------------|
|
||
| `LOCALAI_ADDR` | Reachable address of this worker (`host:port`). The port is used as the base for gRPC backend processes, and `port-1` for the HTTP file transfer server. |
|
||
|
||
**Example:**
|
||
```yaml
|
||
environment:
|
||
LOCALAI_ADDR: "192.168.1.100:50051"
|
||
LOCALAI_NATS_URL: "nats://frontend:4222"
|
||
LOCALAI_REGISTER_TO: "http://frontend:8080"
|
||
LOCALAI_REGISTRATION_TOKEN: "my-secret"
|
||
```
|
||
|
||
For advanced networking scenarios (NAT, load balancers, separate gRPC/HTTP ports), the following override variables are available:
|
||
|
||
| Variable | Description | Default |
|
||
|----------|-------------|---------|
|
||
| `LOCALAI_SERVE_ADDR` | gRPC base port bind address | `0.0.0.0:50051` |
|
||
| `LOCALAI_GRPC_MAX_PORT` | Highest port assignable to a backend gRPC process | `65535` |
|
||
| `LOCALAI_HTTP_ADDR` | HTTP file transfer bind address | `0.0.0.0:{gRPC port - 1}` |
|
||
| `LOCALAI_ADVERTISE_ADDR` | Public gRPC address (if different from `LOCALAI_ADDR`) | Derived from `LOCALAI_ADDR` |
|
||
| `LOCALAI_ADVERTISE_HTTP_ADDR` | Public HTTP address (if different from gRPC host) | Derived from advertise host + HTTP port |
|
||
|
||
### Backend gRPC port range
|
||
|
||
Every backend process a worker starts listens on its own gRPC port, allocated
|
||
upward from the worker's base port (`LOCALAI_SERVE_ADDR`, default `50051`).
|
||
`LOCALAI_GRPC_MAX_PORT` sets the top of that range. The width of
|
||
`[base port, LOCALAI_GRPC_MAX_PORT]` is therefore a hard cap on how many
|
||
backend processes one worker can run concurrently.
|
||
|
||
Set it when the worker shares a host with other services and you need to keep
|
||
the rest of the ephemeral range clear, or when you want a worker's backend
|
||
count bounded explicitly rather than by whatever the host happens to allow:
|
||
|
||
```bash
|
||
# Confine this worker's backends to 50051-50150 (100 concurrent backends).
|
||
LOCALAI_SERVE_ADDR=0.0.0.0:50051
|
||
LOCALAI_GRPC_MAX_PORT=50150
|
||
```
|
||
|
||
Leave it unset (the default) and the worker may use anything up to 65535.
|
||
|
||
Budget headroom above your real concurrency. When a backend stops, its port is
|
||
held briefly before it can be reused, so a worker with heavy start/stop churn
|
||
has more ports tied up than it has running backends at any instant. If the
|
||
range does fill, backend starts fail with:
|
||
|
||
```
|
||
no free gRPC port in range: 50051-50150 is fully consumed by 100 running
|
||
backend(s) and 12 port(s) still in quarantine; raise LOCALAI_GRPC_MAX_PORT to
|
||
widen the range
|
||
```
|
||
|
||
Raise `LOCALAI_GRPC_MAX_PORT` (or reduce how many models you schedule onto that
|
||
worker). A value above 65535 is clamped, and a value below the base port is
|
||
ignored in favour of the full range, so a typo degrades the setting rather than
|
||
wedging every backend start on the node.
|
||
|
||
### NVIDIA GPU support
|
||
|
||
When running workers in a container, two runtime settings affect how VRAM
|
||
usage is reported back to the frontend:
|
||
|
||
- **`NVIDIA_DRIVER_CAPABILITIES` must include `utility`.** Without it, the
|
||
NVML library (and therefore `nvidia-smi`) is not available inside the
|
||
container. CUDA compute still works, but the worker cannot query free VRAM
|
||
and the Nodes page will show the node as fully used. Set
|
||
`NVIDIA_DRIVER_CAPABILITIES=compute,utility` when using the NVIDIA runtime.
|
||
For Docker Compose with `driver: nvidia`, use
|
||
`capabilities: [gpu, compute, utility]` on the device reservation.
|
||
Docker derives driver capabilities from this reservation, so include `compute`
|
||
for CUDA libraries such as `libcuda.so.1`. The `utility` capability alone
|
||
enables monitoring but does not provide CUDA libraries.
|
||
|
||
- **Run the container with `init: true` (or `docker run --init`).** The
|
||
worker process becomes PID 1 in the container and cannot reap zombies on
|
||
its own. Without an init, `nvidia-smi` calls can fail intermittently with
|
||
`waitid: no child processes`, which briefly clears free-VRAM metrics.
|
||
|
||
**Unified memory devices (Jetson, DGX Spark / GB10, Thor):** these SoCs
|
||
share one physical RAM between CPU and GPU. LocalAI detects them via
|
||
`/sys/devices/soc0/family` and `/sys/devices/soc0/soc_id` (no `nvidia-smi`
|
||
required) and reports system-RAM figures as VRAM. Free VRAM therefore tracks
|
||
`MemAvailable` in `/proc/meminfo`. Workers report RAM metrics independently
|
||
from VRAM on every registration and heartbeat. On unified-memory nodes, the
|
||
available RAM and available VRAM values should therefore track each other
|
||
closely; on discrete-GPU nodes they can change independently.
|
||
|
||
### CPU telemetry
|
||
|
||
Backend workers report host-wide CPU telemetry in `GET /api/nodes` and
|
||
`GET /api/nodes/:id`:
|
||
|
||
| Field | Meaning |
|
||
|-------|---------|
|
||
| `cpu_logical_cores` | Logical processor count, sampled at registration |
|
||
| `cpu_usage_percent` | Utilization across the whole host, clamped to `0..100` |
|
||
| `cpu_load_1` | One-minute system load average |
|
||
|
||
Utilization and load are sampled at registration and again at each worker
|
||
heartbeat (every 10 seconds by default). The frontend persists heartbeat
|
||
samples on the normal heartbeat checkpoint cadence; CPU movement alone does
|
||
not force an extra database write. If a sample fails, the worker omits all CPU
|
||
fields and the frontend keeps the last successful reading.
|
||
|
||
Workers from releases that predate CPU reporting remain compatible. Their
|
||
`cpu_logical_cores` value is zero, which means unknown rather than a zero-core
|
||
machine. Fleet capacity excludes those workers from CPU totals and reports
|
||
them as unknown. The dashboard derives available CPU as idle logical-core
|
||
equivalents:
|
||
|
||
```
|
||
idle cores = cpu_logical_cores * (1 - cpu_usage_percent / 100)
|
||
```
|
||
|
||
### Node Labels
|
||
|
||
Workers can declare labels at startup for scheduling constraints:
|
||
|
||
| Variable | Description | Example |
|
||
|----------|-------------|---------|
|
||
| `LOCALAI_NODE_LABELS` | Comma-separated `key=value` labels | `tier=premium,gpu=a100,zone=us-east` |
|
||
|
||
Labels can also be managed via the admin API (see [Label Management API](#label-management-api) below).
|
||
|
||
The system automatically applies hardware-detected labels on registration:
|
||
- `gpu.vendor` -- GPU vendor (nvidia, amd, intel, vulkan)
|
||
- `gpu.vram` -- GPU VRAM bucket (8GB, 16GB, 24GB, 48GB, 80GB+)
|
||
- `node.name` -- The node's registered name
|
||
|
||
### How Workers Operate
|
||
|
||
Workers start as generic processes with no backend installed. When the SmartRouter needs to load a model on a worker, it sends a NATS `backend.install` event with the backend name and model ID. The worker:
|
||
|
||
1. Installs the backend from the gallery (if not already installed)
|
||
2. Starts a **new gRPC backend process on a dynamic port** (each model gets its own process)
|
||
3. Replies with the allocated gRPC address
|
||
4. The SmartRouter calls `LoadModel` via direct gRPC to that address
|
||
|
||
Workers can run **multiple models concurrently** - each model gets its own gRPC process on a separate port. For example, an embedding model on port 50051 and a chat model on port 50052 can run simultaneously on the same worker.
|
||
|
||
When the SmartRouter needs to free capacity, it can unload models with zero in-flight requests without affecting other models on the same worker.
|
||
|
||
### Managing nodes in the WebUI
|
||
|
||
With distributed mode on, **Operate → Swarm** holds the cluster: **Nodes**, **Placement rules**, **Failover** and **P2P**. A single-node install shows **This machine** instead and never draws these pages.
|
||
|
||
**Nodes** lists every worker in a sortable table: name, role, state in words, GPU or system memory, loaded models, last heartbeat and version. Switch between comfortable and compact rows, between **List**, **Map** and **Running models**, and filter to **Needs attention** (waiting for approval, not answering, or low GPU memory, system memory or models disk). The **Map** draws this instance, the message bus and database, and each worker; a dashed line is a worker that gets no traffic. It is not drawn on a phone. When a backend has a newer version, an **Update** button on the page sends the upgrade to the nodes that differ from the rest of the cluster, or to the nodes you selected.
|
||
|
||
The **Running models** view groups replicas by model. Its **View logs…** action opens logs directly when there is one placement. When a model has several, it opens the row so you can choose the logs of one replica.
|
||
|
||
**Add a node** (`/app/nodes/add`) explains how a machine joins: a registered worker, a peer instance, or a memory shard. It prints the command to run on the new machine with a **Copy** button, and updates the page when the machine appears. On a single-node install it starts with the command that turns distributed mode on.
|
||
|
||
Open a node for node-scoped work. The page shows its state, VRAM, RAM, models disk, CPU and in-flight requests, and has tabs for **Models** (replica logs, unload), **Backends** (upgrade, delete), **Logs**, and **Capacity and labels** (replica capacity, labels). **Drain…** shows what the drain would change before it does anything; **Remove…** asks for the node's name. A node that stopped answering says so and shows the last figures it reported.
|
||
|
||
The "what happens if I drain this node" list is a preview worked out in the browser from the node list, the loaded replicas and the placement rules. The server does not compute it, and the scheduler also weighs free memory and disk when it loads a model, so the preview never claims a model will fit.
|
||
|
||
## Node Management API
|
||
|
||
The API is split into two prefixes with distinct auth:
|
||
|
||
### `/api/node/` - Node self-service
|
||
|
||
Used by workers themselves (registration, heartbeat, etc.). Authenticated via the registration token, exempt from global auth.
|
||
|
||
| Method | Path | Description |
|
||
|--------|------|-------------|
|
||
| `POST` | `/api/node/register` | Register a new worker |
|
||
| `POST` | `/api/node/:id/heartbeat` | Update heartbeat timestamp |
|
||
| `POST` | `/api/node/:id/drain` | Mark self as draining |
|
||
| `GET` | `/api/node/:id/models` | Query own loaded models |
|
||
| `DELETE` | `/api/node/:id` | Deregister self |
|
||
|
||
### `/api/nodes/` - Admin management
|
||
|
||
Used by the WebUI and admin API consumers. Requires admin authentication.
|
||
|
||
| Method | Path | Description |
|
||
|--------|------|-------------|
|
||
| `GET` | `/api/nodes` | List all registered workers |
|
||
| `GET` | `/api/nodes/:id` | Get a single worker by ID |
|
||
| `GET` | `/api/nodes/:id/models` | List models loaded on a worker |
|
||
| `GET` | `/api/nodes/models` | List loaded model replicas on healthy workers |
|
||
| `DELETE` | `/api/nodes/:id` | Admin-delete a worker |
|
||
| `POST` | `/api/nodes/:id/drain` | Admin-drain a worker |
|
||
| `POST` | `/api/nodes/:id/approve` | Approve a pending worker node |
|
||
| `POST` | `/api/nodes/:id/backends/install` | Install a backend on a worker |
|
||
| `POST` | `/api/nodes/:id/backends/upgrade` | Upgrade (force-reinstall) a backend on a worker |
|
||
| `POST` | `/api/nodes/:id/backends/delete` | Delete a backend from a worker |
|
||
| `POST` | `/api/nodes/:id/models/unload` | Unload a model from a worker. Cancels a load of that model on the worker first. |
|
||
| `POST` | `/api/models/:id/load-cancel` | Cancel one load attempt (`{"job_id": "..."}`) |
|
||
| `POST` | `/api/nodes/:id/models/delete` | Delete model files from a worker |
|
||
| `PUT` | `/api/nodes/:id/vram-budget` | Set a VRAM budget for a worker (`{"value":"80%"}`) |
|
||
| `DELETE` | `/api/nodes/:id/vram-budget` | Clear a worker's VRAM budget (revert to all detected VRAM) |
|
||
|
||
The **Nodes** page reads `GET /api/nodes` every five seconds and renders 50 workers at a time. Bulk drain, resume and remove run with bounded concurrency, so the page stays usable for fleets with thousands of registrations. Selecting the visible page or a group does not discard selections elsewhere; selections are removed only when a later poll confirms the worker no longer exists.
|
||
|
||
The list never fetches backend inventory, and the **Running models** and **Map** views stay lazy: the first use of either makes one controller database request that is kept until the page is left. A node's backends are read only when its page opens.
|
||
|
||
Use a model row's actions menu in **Running models** to stop that model across the fleet. LocalAI sends one controller shutdown request for the model, which stops all loaded placements; the browser does not contact workers individually. The view refreshes the running-model inventory after both successful and failed shutdown attempts because a failed request can still have stopped some replicas.
|
||
|
||
### Model sizing in the WebUI
|
||
|
||
The model gallery answers "will this model run here" against the cluster, not
|
||
against the frontend. A distributed frontend is usually a GPU-less pod, so
|
||
sizing models against its own memory would report that a fleet of GPU workers
|
||
can only run the smallest CPU build.
|
||
|
||
The budget is the **largest single healthy backend node**, not the sum of the
|
||
fleet: a model loads into one node, so four 16GB workers do not add up to a home
|
||
for a 40GB model. A node's operator-set VRAM budget caps its contribution, since
|
||
the scheduler would refuse a load above that ceiling anyway, and a GPU node wins
|
||
over a CPU node holding more system RAM. The gallery names the node its verdict
|
||
belongs to ("Fits on dgx-01").
|
||
|
||
`GET /api/resources` and `GET /api/models` carry this as an additional `cluster`
|
||
object; their existing `aggregate` and `ram*` fields keep reporting the
|
||
frontend's own hardware, which is what the resource monitor shows. The object is
|
||
absent in single-node mode, and also whenever the registry cannot be read, in
|
||
which case every sizing surface falls back to the local host:
|
||
|
||
```json
|
||
{
|
||
"cluster": {
|
||
"enabled": true,
|
||
"node_id": "a1b2c3",
|
||
"node_name": "dgx-01",
|
||
"total_memory": 85899345920,
|
||
"is_gpu": true,
|
||
"node_count": 4
|
||
}
|
||
}
|
||
```
|
||
|
||
Variant selection (`GET /api/models/variants/:id`) uses the same reading, and
|
||
judges backend compatibility against the union of the capabilities present in
|
||
the cluster, so a CUDA-only build is offered when any worker can run it.
|
||
|
||
### Model configs across frontends
|
||
|
||
Every frontend keeps its own in-memory copy of the model configs in the shared models directory. When a frontend installs, edits, toggles or deletes a model, it writes the change to the directory and publishes a message on NATS. The other frontends reload the directory when they receive it.
|
||
|
||
A gallery install or delete publishes this message as soon as the new config is in place, before the frontend preloads model files. The preload can take minutes on a large models directory, and other frontends do not wait for it. If the preload fails, the operation reports the error, but the config change stays applied on every frontend.
|
||
|
||
NATS keeps no history of these messages. A frontend that is disconnected when a message is published never receives it. To recover, each frontend also reloads the models directory:
|
||
|
||
- every `--model-config-resync-interval` (default `30s`), when a config file changed since its last pass, and
|
||
- after each NATS reconnect.
|
||
|
||
The pass is the same reconcile that a NATS message triggers, so it is idempotent. Only models whose file changed get a new [configuration revision](#model-configuration-revisions). A pass over an unchanged directory reads the config files and does nothing else.
|
||
|
||
Distributed-state mode: each frontend derives this state from the shared directory, so there is no leader and nothing to replicate. The only per-frontend memory is a hash of the config files from its last pass, which only saves work. The consequence is a bounded delay: a frontend that missed a message serves the previous config for at most one interval.
|
||
|
||
Two limits apply:
|
||
|
||
- Models loaded with `--config-file` exist only on the frontend that loaded them. A reload of the models directory keeps them, and a config-file model wins over a directory file with the same name, as it does at startup.
|
||
- The reload is strict: if any config file in the directory does not parse, the frontend keeps its current configs and logs the error once. It retries on each pass until the file is fixed. This keeps a half-written file from looking like a deleted model.
|
||
|
||
### Model configuration revisions
|
||
|
||
Distributed mode assigns a `config_revision` to each validated model configuration. It hashes the persisted semantic configuration, including fields such as `context_size` and parallel settings. YAML formatting, comments, and map order do not change it.
|
||
|
||
The first request for a model establishes its current revision and replay information. The replica reconciler uses only replay information that matches the current revision. This lets `min_replicas` recover after an ordinary worker failure without restoring an old configuration.
|
||
|
||
When you save a valid model edit, LocalAI makes replicas from the old revision ineligible immediately. New requests cannot route to those replicas. This rule applies to raw YAML edits, structured patches, renames, disabled models, and changes from another frontend.
|
||
|
||
The edit response includes these fields:
|
||
|
||
- `config_revision` identifies the saved semantic configuration.
|
||
- `pending_cleanup` counts old replicas that still need cleanup when the response returns.
|
||
|
||
LocalAI sends an acknowledged stop request for each exact backend process. If a worker or NATS is unreachable, LocalAI keeps the replica in the `unloading` state and retries with durable backoff. The saved edit remains successful while cleanup is pending.
|
||
|
||
Workers must support the exact model-stop protocol. Upgrade all workers before you rely on revision cleanup. An older worker cannot acknowledge the request, so its stale replica remains `unloading` until cleanup succeeds or the worker re-registers.
|
||
|
||
Worker re-registration removes stale live-replica rows, but it preserves the current model revision and matching replay information. A temporary worker outage therefore does not make an old revision routable. The reconciler can restore the current revision after the worker becomes healthy.
|
||
|
||
The responses from `GET /api/node/:id/models` and `GET /api/nodes/:id/models` include these replica fields:
|
||
|
||
| Field | Meaning |
|
||
|-------|---------|
|
||
| `config_revision` | Hash of the persisted semantic model configuration that created the replica. Routable replicas match the current revision. |
|
||
| `effective_options_hash` | Hash of the final node-specific load options after defaults and file staging have been applied. Different hashes can be valid on heterogeneous workers when `config_revision` matches. |
|
||
| `state` | Replica lifecycle state, such as `staging`, `loading`, `loaded`, or `unloading`. Only eligible `loaded` replicas receive requests. |
|
||
| `cleanup_error` | Last exact-stop error. This field appears while cleanup is pending. |
|
||
| `cleanup_next_retry_at` | Time of the next durable cleanup attempt. This field appears after a failed attempt. |
|
||
|
||
`model.unload` releases model memory inside a running backend. It does not replace the exact process stop that configuration cleanup requires. The `backend.stop` operation remains an administrative backend operation.
|
||
|
||
#### `backend.stop` is acknowledged
|
||
|
||
`backend.stop` is request-reply. The worker answers with what it terminated, so the controller can tell a stop that worked from one that matched nothing or failed outright.
|
||
|
||
This matters for `POST /api/nodes/:id/models/unload`, which stops the backend after unloading the model. The stop used to be fire-and-forget, so the endpoint answered `200` as soon as the message left the frontend — including when the backend was still running and still holding its VRAM. It now returns an error when the worker reports that the stop failed.
|
||
|
||
Two outcomes are deliberately **not** errors:
|
||
|
||
- **Nothing matched.** The worker reports an empty stopped-process list, logged as `backend.stop matched no running process`. Stopping a backend that is not running leaves the caller in the state it asked for, and eviction and cleanup paths stop already-gone models routinely.
|
||
- **No answer.** A worker built before this reply performs the stop and never responds. The controller waits 15 seconds, logs `Worker did not acknowledge backend.stop`, and assumes delivery, so a fleet mid-upgrade keeps working. A transport failure is reported rather than assumed.
|
||
|
||
{{% notice note %}}
|
||
On a mixed fleet, every stop against a worker that predates the reply costs the full 15-second wait before falling back. Upgrading the workers removes the delay.
|
||
{{% /notice %}}
|
||
|
||
### Per-node VRAM budget
|
||
|
||
Each worker advertises its detected VRAM, and the SmartRouter uses that number when picking a node with enough free memory. You can cap the VRAM a node offers for placement so it never gets scheduled beyond a chosen limit, leaving headroom for other workloads on that machine.
|
||
|
||
There are two ways to set the cap:
|
||
|
||
- **At the worker:** start it with `--vram-budget` / `LOCALAI_VRAM_BUDGET` (see [Worker Configuration](#worker-configuration)).
|
||
- **From the frontend, live:** set it per node in the **node capacity editor** on the node detail page, or via the admin API:
|
||
|
||
```bash
|
||
# Cap node placement at 80% of its detected VRAM
|
||
curl -X PUT http://frontend:8080/api/nodes/<node-id>/vram-budget \
|
||
-H "Authorization: Bearer <admin-token>" \
|
||
-H "Content-Type: application/json" \
|
||
-d '{"value":"80%"}'
|
||
|
||
# Or an absolute amount
|
||
curl -X PUT http://frontend:8080/api/nodes/<node-id>/vram-budget \
|
||
-H "Authorization: Bearer <admin-token>" \
|
||
-d '{"value":"12GB"}'
|
||
|
||
# Clear the budget (revert to all detected VRAM)
|
||
curl -X DELETE http://frontend:8080/api/nodes/<node-id>/vram-budget \
|
||
-H "Authorization: Bearer <admin-token>"
|
||
```
|
||
|
||
The value accepts the same formats as the standalone budget: a percentage (`80%`) or an absolute amount (`12GB`, `12GiB`, `12000MB`, or raw bytes). It is a **hard ceiling**: the node's advertised VRAM becomes `min(detected, budget)`, so a budget can only lower the number, never raise it above the hardware. An admin-set node budget is **sticky across worker restarts**: it is stored in the node registry and reapplied when the worker re-registers, so it wins over whatever the worker reports on reconnect. For the underlying semantics and the standalone equivalent, see [VRAM Budget]({{%relref "advanced/vram-management#vram-budget-allocation-ceiling" %}}).
|
||
|
||
### Disk headroom
|
||
|
||
Model weights are **staged onto the worker's disk** before the backend loads them, so a node needs free space as well as free VRAM. Each worker reports the capacity of the filesystem backing its **models directory** (`--models-path`), not the root filesystem, on registration and on every heartbeat. Those figures appear as `total_disk` and `available_disk` in the nodes API and as **Models disk free** on the node detail page.
|
||
|
||
Before placing a model, the SmartRouter removes any node whose models filesystem cannot hold it. The requirement is derived from the model's actual on-disk size plus a small margin (5%, at least 1 GiB), rather than a fixed percentage of the node's disk — a fixed threshold would take a small-but-usable node out of rotation for models it could comfortably store. When the model's size cannot be determined locally (a bare HuggingFace repo id that the worker fetches itself), the node only has to clear a 2 GiB floor.
|
||
|
||
If **no** node has enough space, the request fails immediately with a capacity error naming the requirement and each node's free space, for example:
|
||
|
||
```
|
||
scheduling longcat-video-avatar-1.5: no node has enough free disk for the model:
|
||
need 73.5 GB free on the models filesystem, but nvidia-thor has 0 B free of 937.0 GB
|
||
```
|
||
|
||
This is deliberately a scheduling-time verdict. Without it, a worker with a full disk still reported `status: healthy`, accepted the staging request, transferred tens of gigabytes and only then failed with `no space left on device` — minutes after a decision that could never have succeeded.
|
||
|
||
Workers that predate this feature (or whose disk reading fails) report `total_disk` as `0`. Such nodes are treated as *unknown*, not *full*, and stay in rotation, so a rolling upgrade never empties the candidate pool. A full disk is distinguishable because it reports a non-zero `total_disk` with `available_disk` at `0`.
|
||
|
||
Low disk does **not** mark a node `unhealthy`. Disk is compared per model rather than against a global threshold, so a node that is too small for one model remains a valid target for smaller ones. The check is also skipped entirely in [shared-models mode](#shared-models-directory), where nothing is staged to the worker at all.
|
||
|
||
#### Turning the check off
|
||
|
||
The check is **on by default**. To disable it, start the frontend with `--distributed-disk-headroom-check=false` / `LOCALAI_DISTRIBUTED_DISK_HEADROOM_CHECK=false`, or toggle **Settings → Distributed → Disk headroom check** in the WebUI (`distributed_disk_headroom_check` via `POST /api/settings`). The runtime setting takes effect on the next placement, with no restart; the env/CLI flag only sets the value LocalAI boots with, and both write the same underlying value, so the last change wins.
|
||
|
||
Disabling means **warn, do not block**. Node selection goes back to ignoring free disk (the pre-check behaviour), but the check still runs, and when it would have rejected *every* node it logs a warning naming the shortfall:
|
||
|
||
```
|
||
WARN No node has room to store this model, but the disk-headroom check is DISABLED;
|
||
scheduling anyway — staging will most likely fail with ENOSPC
|
||
model=longcat-video-avatar-1.5 knob=distributed-disk-headroom-check
|
||
```
|
||
|
||
The alternative — skipping the check outright — was rejected because it reproduces the condition that made the original bug expensive: the cluster was doing something that could not work and said nothing about it. The escape hatch exists for setups where the size estimate is wrong (deduplicating or compressing filesystems, a backend that fetches its own weights rather than using the staged copy), and in exactly those cases the operator needs to see what LocalAI thought was wrong. Disabling is logged once at startup as well.
|
||
|
||
The **LocalAI Assistant** can also set a node budget conversationally through the `set_node_vram_budget` MCP tool.
|
||
|
||
## Node Approval
|
||
|
||
By default, new worker nodes start in **pending** status and must be approved by an admin before they can receive traffic. This prevents unknown machines from joining the cluster.
|
||
|
||
To approve a pending node via the API:
|
||
|
||
```bash
|
||
curl -X POST http://frontend:8080/api/nodes/<node-id>/approve \
|
||
-H "Authorization: Bearer <admin-token>"
|
||
```
|
||
|
||
The **Nodes** page in the WebUI shows pending nodes with an **Approve** button, and so does the node's own page and the **Add a node** page once the machine has registered.
|
||
|
||
To skip manual approval and let nodes join immediately, set `--auto-approve-nodes` (or `LOCALAI_AUTO_APPROVE_NODES=true`) on the frontend. This is convenient for development and trusted environments.
|
||
|
||
## Node Statuses
|
||
|
||
| Status | Meaning |
|
||
|--------|---------|
|
||
| `pending` | Node registered but waiting for admin approval (when `--auto-approve-nodes` is `false`) |
|
||
| `healthy` | Node is active and responding to heartbeats |
|
||
| `unhealthy` | Node has missed heartbeats beyond the threshold (detected by the HealthMonitor) |
|
||
| `offline` | Node is temporarily offline (graceful shutdown or stale heartbeat). The node row is preserved so re-registration restores the previous approval status without requiring re-approval |
|
||
| `draining` | Node is shutting down gracefully - no new requests are routed to it, existing in-flight requests are allowed to complete |
|
||
|
||
### Heartbeat writes and stale-node detection
|
||
|
||
Workers beat every `--heartbeat-interval` (default `10s`). Writing each beat straight
|
||
to PostgreSQL means roughly 52,000 `UPDATE`s a day against a table that holds one row
|
||
per node. With autovacuum healthy that is merely wasteful. With autovacuum blocked --
|
||
by a long-lived idle transaction, for example -- the dead tuples accumulate, and a
|
||
six-row table has been observed growing to 460 MB, at which point scanning it cost
|
||
867 ms and the queries that place models began timing out.
|
||
|
||
So the frontend **checkpoints** the write. A beat that carries nothing but a fresher
|
||
timestamp is held in memory until `--node-heartbeat-checkpoint` (default `60s`) has
|
||
elapsed since that node's last durable write. These beats still reach the database
|
||
without waiting:
|
||
|
||
- the node's first beat after the frontend starts, or after it was seen offline
|
||
- any beat from a node that is not active (`pending`, `offline`), because such a node
|
||
recovers only when the health monitor sees a fresh timestamp
|
||
- a GPU vendor, total VRAM or total disk that **differs** from the stored value, since
|
||
those are hardware facts and a change to one is a real event
|
||
- a free VRAM, free RAM or free disk reading that has moved more than 256 MiB, because
|
||
the scheduler places against those figures
|
||
|
||
CPU utilization and load follow the scheduled checkpoint instead of making a
|
||
heartbeat material. They are dashboard observations and do not affect model
|
||
placement, so persisting every fluctuation would defeat write suppression.
|
||
|
||
Every figure is compared against the value **last written**, not against the previous
|
||
beat. A worker reports its disk capacity on every single beat, so testing whether a
|
||
field is merely *present* would make every real beat look like a change and suppress
|
||
nothing. Measuring from the written value also means a reading that walks away in
|
||
sub-256 MiB steps still writes once the total distance crosses the threshold, rather
|
||
than drifting arbitrarily far from the figure the scheduler is reading.
|
||
|
||
The consequence is that `last_heartbeat` is up to one checkpoint interval behind
|
||
reality **by design**. The stale-node threshold therefore defaults to **5 minutes**
|
||
(it was 60 seconds before checkpointing existed): the health monitor waits that long
|
||
without a fresh timestamp before it marks a node `offline`. It is configurable with
|
||
`--stale-node-threshold` / `LOCALAI_STALE_NODE_THRESHOLD`, and an operator who widens
|
||
`--node-heartbeat-checkpoint` must widen this to match, or the beats that checkpointing
|
||
suppresses will read as a dead node.
|
||
|
||
{{% notice note %}}
|
||
Marking a node `offline` from a stale heartbeat now takes up to five minutes. This is
|
||
the slowest of the three ways a dead worker is noticed, not the only one. The per-model
|
||
gRPC health check still probes each loaded model on the health-monitor interval
|
||
(default `15s`) and removes replicas whose backend has died, and a request routed to a
|
||
gone worker still fails and is retried elsewhere at request time. Neither of those
|
||
paths reads `last_heartbeat`, so neither is slowed by this change.
|
||
{{% /notice %}}
|
||
|
||
To go back to a durable write per beat -- on a database with plenty of write headroom,
|
||
or while debugging heartbeat delivery -- set `LOCALAI_NODE_HEARTBEAT_CHECKPOINT` to a
|
||
value below the worker's heartbeat interval, for example `1s`.
|
||
|
||
### Operations: do not share a database with the vector store
|
||
|
||
Give the control plane a PostgreSQL **database of its own**. Sharing one with the
|
||
agent vector store, or with anything else that holds long transactions, is the fastest
|
||
way to reproduce the 460 MB node registry described above.
|
||
|
||
PostgreSQL computes the removable-tuple cutoff **per database**, not per table. One
|
||
transaction left open anywhere in the database -- a stalled embedding batch, an idle
|
||
`BEGIN` from a connection pool, an abandoned `psql` session -- pins that cutoff for
|
||
**every** table in it. Autovacuum still runs, finds nothing it is allowed to reclaim,
|
||
and moves on. The node registry is six rows rewritten tens of thousands of times a
|
||
day, so it is the table that pays: it bloats into hundreds of megabytes, a sequential
|
||
scan starts costing the best part of a second, and model placement begins timing out
|
||
while the vector store that caused it looks perfectly healthy.
|
||
|
||
Concretely, these two must point at different databases:
|
||
|
||
| Variable | What it holds |
|
||
|----------|---------------|
|
||
| `LOCALAI_AUTH_DATABASE_URL` | Auth **and the distributed control plane** - nodes, replicas, load jobs |
|
||
| `LOCALAI_AGENT_POOL_DATABASE_URL` | Agent collections and their embeddings |
|
||
|
||
Different databases on the same PostgreSQL server is enough; they do not need separate
|
||
servers. Different *schemas* in one database is **not** enough, because the cutoff is
|
||
per database.
|
||
|
||
To detect it before placement starts failing, watch the
|
||
`localai_control_plane_oldest_xmin_age` gauge, exported on the frontend's OpenTelemetry
|
||
meter. It reports how many transactions have elapsed
|
||
since the oldest snapshot still held open against the control plane's database. Under
|
||
normal load it stays small and flat. A line that climbs without coming back down means
|
||
something is holding a transaction open and autovacuum has stopped reclaiming the node
|
||
registry; find it with:
|
||
|
||
```sql
|
||
SELECT pid, state, age(backend_xmin) AS xmin_age, query
|
||
FROM pg_stat_activity
|
||
WHERE backend_xmin IS NOT NULL
|
||
ORDER BY age(backend_xmin) DESC
|
||
LIMIT 5;
|
||
```
|
||
|
||
Grant `pg_read_all_stats` to the role LocalAI connects as (`GRANT pg_read_all_stats TO
|
||
localai;`), or make it a superuser. PostgreSQL blanks `backend_xmin` and `xact_start` in
|
||
`pg_stat_activity` for sessions owned by **other** roles, so without that grant both the
|
||
gauge and the query above see only LocalAI's own sessions -- and the transaction that
|
||
wedges the horizon is typically the co-located vector store connecting as a different
|
||
role, which is exactly the case they exist to catch.
|
||
|
||
## Agent collections across frontends
|
||
|
||
Collections (the `/api/agents/collections` routes) are **shared state**: every frontend shows the same list. With the `postgres` vector engine, the database is the source of truth for which collections exist. Each frontend keeps the collections it has opened in memory, but only as a cache:
|
||
|
||
- Listing a collection set reads it from the database.
|
||
- A request for a collection that this frontend has not opened yet checks the database before it answers `404`, and opens the collection if it exists. A collection created through one frontend is therefore usable through any other one at once.
|
||
- A frontend re-checks a cached collection at most every 5 seconds, so a collection removed through another frontend stops being served within that time.
|
||
- Creating or resetting a collection publishes an event on the message bus, so the other frontends re-check immediately.
|
||
|
||
This needs `LOCALAI_AGENT_POOL_VECTOR_ENGINE=postgres` and `LOCALAI_AGENT_POOL_DATABASE_URL`. With another vector engine there is no shared registry, and each frontend keeps its own list. Per-user collections (authentication enabled) are not yet coordinated this way.
|
||
|
||
## Agent Workers
|
||
|
||
Agent workers are dedicated processes for executing agent chats and MCP CI jobs. Unlike backend workers (which run gRPC model inference), agent workers use cogito to orchestrate multi-step conversations with tool calls.
|
||
|
||
```bash
|
||
local-ai agent-worker \
|
||
--register-to http://frontend:8080 \
|
||
--nats-url nats://nats:4222 \
|
||
--registration-token changeme
|
||
```
|
||
|
||
Agent workers:
|
||
- Execute agent chat messages dispatched via NATS
|
||
- Run MCP CI jobs (with access to MCP servers via docker)
|
||
- Handle MCP tool discovery and execution requests from the frontend
|
||
- Get auto-provisioned API keys during registration for calling the inference API
|
||
|
||
`LOCALAI_AGENT_SUBJECT` (default `agent.execute`) must be a subject that LocalAI serves. Use the `agent` root, for example `agent.execute`. The worker refuses to start with a subject whose root LocalAI does not serve (for example `tenant-a.agent.execute`) or with a `>` wildcard, because no message is carried on those subjects.
|
||
|
||
In the docker-compose setup, the agent worker mounts the Docker socket so it can run MCP stdio servers (e.g., `docker run` commands):
|
||
|
||
```yaml
|
||
agent-worker-1:
|
||
command: agent-worker
|
||
volumes:
|
||
- /var/run/docker.sock:/var/run/docker.sock
|
||
```
|
||
|
||
## MCP in Distributed Mode
|
||
|
||
MCP servers configured in model configs work in distributed mode. The frontend routes MCP operations through NATS to agent workers:
|
||
|
||
- **MCP discovery** (`GET /v1/mcp/servers/:model`): routed to agent workers which create sessions and return server info
|
||
- **MCP tool execution** (during `/v1/chat/completions`): tool calls are routed to agent workers via NATS request-reply
|
||
- **MCP CI jobs**: executed entirely on agent workers with access to docker for stdio-based MCP servers
|
||
|
||
## vLLM Multi-Node (Data-Parallel)
|
||
|
||
A single vLLM model can span multiple GPU nodes via data parallelism: the head node serves the OpenAI API and runs the local DP ranks, follower nodes run vanilla `vllm serve --headless` and speak ZMQ directly to the head. LocalAI's role is starting the follower processes and surfacing them in the admin UI; the cross-rank tensor traffic is vLLM's own.
|
||
|
||
This mode is **operator-launched** - the head config and each follower's invocation must agree on the topology (`data_parallel_size`, `data_parallel_size_local`, `data_parallel_address`, `data_parallel_rpc_port`). The SmartRouter does not place follower ranks automatically.
|
||
|
||
### Head node configuration
|
||
|
||
The head runs the existing single-node vLLM gRPC backend. Set `engine_args` to publish the DP topology vLLM expects:
|
||
|
||
```yaml
|
||
backend: vllm
|
||
parameters:
|
||
model: moonshotai/Kimi-K2.6-Instruct
|
||
engine_args:
|
||
data_parallel_size: 4 # total ranks across all nodes
|
||
data_parallel_size_local: 2 # ranks on the head node
|
||
data_parallel_address: 10.0.0.1 # head's reachable IP
|
||
data_parallel_rpc_port: 32100 # any free port; followers connect here
|
||
enable_expert_parallel: true # for MoE models
|
||
```
|
||
|
||
The head will start its 2 local ranks, listen on `10.0.0.1:32100`, and wait for the remaining 2 ranks to handshake.
|
||
|
||
### Follower nodes
|
||
|
||
Each follower runs `local-ai p2p-worker vllm` with matching topology, an explicit start rank, and the head's address:
|
||
|
||
```bash
|
||
local-ai p2p-worker vllm \
|
||
moonshotai/Kimi-K2.6-Instruct \
|
||
--data-parallel-size 4 \
|
||
--data-parallel-size-local 2 \
|
||
--start-rank 2 \
|
||
--master-addr 10.0.0.1 \
|
||
--master-port 32100 \
|
||
--register-to http://frontend:8080 \
|
||
--registration-token changeme
|
||
```
|
||
|
||
`--register-to` is optional but recommended - it makes the follower visible in the admin UI as an `agent`-type node tagged with `node.role=vllm-follower`. Without it the worker just runs vLLM and exits silently when vLLM does. The role label discourages SmartRouter from placing other models on the follower; pair it with model selectors like `{"!node.role":"vllm-follower"}` if you also run regular LocalAI models on the same fleet.
|
||
|
||
### Worked example: 2-node Kimi-K2.6 deployment
|
||
|
||
Two A100 nodes (`10.0.0.1`, `10.0.0.2`), 8 GPUs total, `data_parallel_size=8` with 4 ranks per node:
|
||
|
||
```yaml
|
||
# /models/kimi.yaml on the head (10.0.0.1)
|
||
name: kimi-k2-6
|
||
backend: vllm
|
||
parameters:
|
||
model: moonshotai/Kimi-K2.6-Instruct
|
||
engine_args:
|
||
data_parallel_size: 8
|
||
data_parallel_size_local: 4
|
||
data_parallel_address: 10.0.0.1
|
||
data_parallel_rpc_port: 32100
|
||
enable_expert_parallel: true
|
||
all2all_backend: deepep_high_throughput
|
||
```
|
||
|
||
```bash
|
||
# On 10.0.0.2 (follower)
|
||
local-ai p2p-worker vllm moonshotai/Kimi-K2.6-Instruct \
|
||
--data-parallel-size 8 --data-parallel-size-local 4 --start-rank 4 \
|
||
--master-addr 10.0.0.1 --master-port 32100 \
|
||
--register-to http://10.0.0.1:8080 --registration-token changeme
|
||
```
|
||
|
||
A `curl http://10.0.0.1:8080/v1/chat/completions ...` against the head will then dispatch across all 8 ranks.
|
||
|
||
### Intel Arc / XPU notes
|
||
|
||
vLLM XPU supports DP (`vllm/platforms/xpu.py:198` handles `world_size_across_dp > 1`; ranks bind to `xpu:{local_rank}` in `xpu_worker.py:62`, with xccl as the collective backend). Each rank still needs a distinct discrete GPU - the iGPU on a hybrid host is not a viable second device.
|
||
|
||
Older XE-HPG GPUs (e.g. Arc A770) need to bypass the cutlass attention path:
|
||
|
||
```yaml
|
||
engine_args:
|
||
attention_backend: TRITON_ATTN
|
||
```
|
||
|
||
`docker-compose.vllm-multinode.intel.yaml` at the repo root is the Intel equivalent of `docker-compose.vllm-multinode.yaml` - uses `/dev/dri` passthrough, `ZE_AFFINITY_MASK` to pin each rank to one device, and `latest-gpu-intel` images. Run via `./tests/e2e/vllm-multinode/smoke.sh --intel`.
|
||
|
||
### Caveats
|
||
|
||
- **Tensor parallel within a node only.** vLLM v1 does not support TP across nodes; combine `tensor_parallel_size` (within a node, via `engine_args`) with `data_parallel_size` (across nodes).
|
||
- **Followers don't host LocalAI gRPC.** The follower process is vanilla vLLM, so `/api/backend-logs/<modelId>` does not stream follower output. Use `journalctl` / `kubectl logs` / compose logs for the follower's stderr.
|
||
- **Network reachability.** The head's `data_parallel_rpc_port` plus a range of ZMQ ports (typically `data_parallel_rpc_port..+N`) must be reachable from every follower. Open them in your firewall / security group.
|
||
- **Topology must match exactly.** A mismatch in `--data-parallel-size` between head and any follower will hang the handshake. Check the head's vLLM logs for `waiting for N DP ranks` if startup stalls.
|
||
|
||
## ds4 Layer-Split Distributed Inference
|
||
|
||
The ds4 backend (DeepSeek V4 Flash) supports **layer-parallel** distributed inference: a single model that is too large for one machine is split by transformer layer across several machines. Each machine must have the GGUF present locally, but loads **only its own slice** of the layers. This lets you run a model whose weights exceed any single host's memory.
|
||
|
||
This is **not** routed through the SmartRouter: it is a model-internal split, configured manually (Phase 1). It is unrelated to the NATS/PostgreSQL distributed mode described above.
|
||
|
||
### Topology
|
||
|
||

|
||
|
||
ds4 uses a **coordinator/worker** split:
|
||
|
||
- The **coordinator** owns tokenization, sampling, the prompt, and a low layer range (e.g. `0:19`). It is LocalAI's ds4 backend and **listens** on a host/port. Workers dial into it.
|
||
- One or more **workers** own higher layer ranges (e.g. `20:output`). Each worker loads only its slice and **dials the coordinator** to register the range it can serve. The last worker normally owns the output head.
|
||
- Activations flow through the connected slices and back to the coordinator. The route is "ready" only once the coordinator plus all connected workers cover every layer.
|
||
|
||
This dial direction is the **inverse** of the llama.cpp RPC model, where the main server dials *out* to a list of `rpc-server` workers. With ds4 the **workers dial in** to the coordinator.
|
||
|
||
### Coordinator setup
|
||
|
||
The coordinator is a normal LocalAI ds4 model whose YAML carries distributed `options:`:
|
||
|
||
```yaml
|
||
name: ds4flash
|
||
backend: ds4
|
||
options:
|
||
- "ds4_role:coordinator"
|
||
- "ds4_layers:0:19"
|
||
- "ds4_listen:0.0.0.0:1234"
|
||
```
|
||
|
||
| Option | Meaning |
|
||
|--------|---------|
|
||
| `ds4_role:coordinator` | Enables distributed coordinator mode. Without `ds4_role`, the backend behaves as a normal single-node ds4 model. |
|
||
| `ds4_layers:0:19` | The coordinator's own layer slice (inclusive). |
|
||
| `ds4_listen:0.0.0.0:1234` | Address that workers dial into. |
|
||
| `ds4_route_timeout:60` | Optional. Seconds the coordinator waits for the worker route to form before returning an error on a request. Defaults to 60. |
|
||
|
||
{{% notice warning %}}
|
||
Worker↔coordinator traffic is **plaintext and unauthenticated**: there is no TLS or auth on this channel. Bind `ds4_listen` to an address on a trusted/private network only; using `0.0.0.0` exposes the coordinator on every interface. Run the layer split exclusively over a network you control.
|
||
{{% /notice %}}
|
||
|
||
Once the model is loaded, the coordinator serves requests exactly like a single-node ds4 model: generation goes through the ordinary inference path and is transparently routed across the layer slices.
|
||
|
||
### Worker setup
|
||
|
||
On each worker machine (with the GGUF present locally), start a worker pointed at the coordinator:
|
||
|
||
```bash
|
||
local-ai worker ds4-distributed -- \
|
||
--role worker \
|
||
--model /models/ds4flash.gguf \
|
||
--layers 20:output \
|
||
--coordinator <coordinator-host> 1234
|
||
```
|
||
|
||
`local-ai worker ds4-distributed` resolves the ds4 backend and execs the packaged `ds4-worker` binary, passing everything after `--` straight through.
|
||
|
||
### Layer-range semantics
|
||
|
||
- Ranges are **inclusive**: `0:19` is layers 0 through 19.
|
||
- `N:output` means layer N through the final layer **plus the output head**. The last worker normally owns the output head.
|
||
- The coordinator and all connected workers together **must cover every layer**. Until they do, the coordinator returns a gRPC `UNAVAILABLE` error on inference requests (so a worker that starts slightly after the coordinator is tolerated: once it connects and the route is complete, requests succeed). The wait is tunable via `ds4_route_timeout`.
|
||
|
||
{{% notice note %}}
|
||
ds4 layer-split inference is **manual setup** in this release (Phase 1): you place the coordinator config and launch each worker yourself, and the layer ranges must be partitioned by hand so they cover the whole model. P2P auto-discovery of the coordinator is planned for a later phase.
|
||
{{% /notice %}}
|
||
|
||
## Scaling
|
||
|
||
**Adding worker capacity:** Start additional `worker` instances pointing to the same frontend. They self-register automatically:
|
||
|
||
```bash
|
||
# Additional workers - no backend type needed
|
||
local-ai worker \
|
||
--register-to http://frontend:8080 \
|
||
--node-name worker-2 \
|
||
--nats-url nats://nats:4222 \
|
||
--registration-token changeme
|
||
|
||
local-ai worker \
|
||
--register-to http://frontend:8080 \
|
||
--node-name worker-3 \
|
||
--nats-url nats://nats:4222 \
|
||
--registration-token changeme
|
||
```
|
||
|
||
**Multiple frontend replicas:** Run multiple LocalAI frontends behind a load balancer. Since all state is in PostgreSQL and coordination is via NATS, frontends are fully stateless and interchangeable.
|
||
|
||
## Model Scheduling
|
||
|
||
Model scheduling controls where models are placed and how many replicas are maintained. In the React WebUI it has its own **Placement rules** page (**Operate → Swarm → Placement rules**, at `/app/scheduling`). Each rule is written as a sentence, shows the nodes the model is loaded on now, and opens in a side sheet that previews which nodes the draft rule could use. The preview is worked out in the browser from the node list and labels; the scheduler also checks free memory and disk when it loads a model. A deleted rule can be taken back for a few seconds. A rule combines two optional features:
|
||
|
||
### Node Selectors
|
||
|
||
Pin models to nodes with specific labels. Only nodes matching **all** selector labels are eligible:
|
||
|
||
```bash
|
||
# Only schedule on NVIDIA nodes in the us-east zone
|
||
curl -X POST http://frontend:8080/api/nodes/scheduling \
|
||
-H "Content-Type: application/json" \
|
||
-d '{"model_name": "llama3", "node_selector": {"gpu.vendor": "nvidia", "zone": "us-east"}}'
|
||
```
|
||
|
||
Without a node selector, models can schedule on any healthy node (default behavior).
|
||
|
||
In the WebUI, the node selector field completes what you type against the labels
|
||
your cluster actually reports: start typing a key and the matching label keys
|
||
appear inline, then the value field offers only the values that key takes. A key
|
||
no node reports yet is still accepted as typed, so you can write a rule before
|
||
labelling the nodes for it.
|
||
|
||
### Replica Auto-Scaling
|
||
|
||
Control the number of model replicas across the cluster:
|
||
|
||
| Field | Description |
|
||
|-------|-------------|
|
||
| `min_replicas` | Minimum replicas to maintain (0 = no minimum, single replica default) |
|
||
| `max_replicas` | Maximum replicas allowed (0 = unlimited) |
|
||
|
||
Auto-scaling is **only active** when `min_replicas > 0` or `max_replicas > 0`.
|
||
|
||
```bash
|
||
# Scale llama3 between 2 and 4 replicas on NVIDIA nodes
|
||
curl -X POST http://frontend:8080/api/nodes/scheduling \
|
||
-H "Content-Type: application/json" \
|
||
-d '{
|
||
"model_name": "llama3",
|
||
"node_selector": {"gpu.vendor": "nvidia"},
|
||
"min_replicas": 2,
|
||
"max_replicas": 4
|
||
}'
|
||
```
|
||
|
||
The **Replica Reconciler** runs as a background process on the frontend:
|
||
- **Scale up**: Adds replicas when all existing replicas are busy (have in-flight requests)
|
||
- **Scale down**: Removes idle replicas after 5 minutes of inactivity
|
||
- **Maintain minimum**: Ensures `min_replicas` are always loaded (recovers from node failures)
|
||
- **Eviction protection**: Models with auto-scaling enabled are never evicted below `min_replicas`
|
||
- **Restart-safe**: Per-model load metadata (backend type + `ModelOptions`) is persisted in the `model_load_infos` PostgreSQL table on the first successful dispatch, so a frontend restart or rolling upgrade does not require a fresh inference request to repopulate state before the reconciler can scale up replacement replicas.
|
||
|
||
All fields are optional and composable:
|
||
- Node selector only: pin model to matching nodes, single replica
|
||
- Replicas only: auto-scale across all nodes
|
||
- Both: auto-scale on matching nodes only
|
||
|
||
### Scheduling a model alias
|
||
|
||
`model_name` accepts a [model alias](/features/model-aliases/) as well as a
|
||
model. A rule keyed by an alias governs whatever model that alias currently
|
||
points at, and keeps governing it after you repoint the alias:
|
||
|
||
```bash
|
||
# "production" is an alias for llama3
|
||
curl -X POST http://frontend:8080/api/nodes/scheduling \
|
||
-H "Content-Type: application/json" \
|
||
-d '{"model_name": "production", "node_selector": {"tier": "gpu"}, "min_replicas": 2}'
|
||
|
||
# Repoint the alias at a new model: the rule follows, llama4 now runs
|
||
# two replicas on the GPU tier and llama3 falls back to on-demand placement.
|
||
```
|
||
|
||
This makes an alias a stable deployment slot: the placement policy belongs to
|
||
the slot, and the model filling it can change without rewriting the rule. The
|
||
WebUI lists aliases in the model picker on the **Placement rules** page, tagged with
|
||
the model each one resolves to.
|
||
|
||
Each frontend resolves the alias from its own copy of the model configs, and a
|
||
frontend that has not yet reloaded a repointed alias still resolves it the old
|
||
way (see [Model configs across frontends](#model-configs-across-frontends)).
|
||
The rule's stored target therefore follows the alias only through frontends
|
||
whose copy of the alias config matches the
|
||
[configuration revision](#model-configuration-revisions) the cluster accepted.
|
||
A frontend that is behind uses the stored target for the replica reconciler and
|
||
does not write it, so two frontends cannot overwrite the rule's target against
|
||
each other, and a frontend that is behind cannot reload the model the alias
|
||
used to point at.
|
||
|
||
Two constraints follow from replicas being shared. A single load of `llama3`
|
||
serves both `production` and any request that names `llama3` directly, so only
|
||
one rule can decide where it runs: a rule whose target is already governed by
|
||
another rule is rejected with `409 Conflict` naming the rule that has it. And a
|
||
rule keyed by an alias that resolves to nothing (its target was deleted, or it
|
||
points at another alias) is rejected, since it would govern nothing loadable.
|
||
|
||
A rule can still end up inert if the pair is created some other way, for example
|
||
by a declarative seed or by repointing an alias onto a model that already has a
|
||
rule. The rule that governs is the one keyed by the model's own name, or failing
|
||
that the oldest one; the rest are listed as **Shadowed** in the WebUI and carry
|
||
`"shadowed": true` in `GET /api/nodes/scheduling`.
|
||
|
||
### Declarative per-model scheduling (unattended installs)
|
||
|
||
In distributed mode you can declare per-model scheduling at startup, instead of
|
||
using the WebUI/API. Config is **authoritative**: it is re-applied on every boot
|
||
and overwrites the listed models (models not listed are left untouched).
|
||
|
||
| Variable | Description |
|
||
|----------|-------------|
|
||
| `LOCALAI_MODEL_SCHEDULING` | Inline JSON list of scheduling entries |
|
||
| `LOCALAI_MODEL_SCHEDULING_CONFIG` | Path to a YAML file with the same list |
|
||
|
||
Entry fields: `model_name` (required), `node_selector` (a label map; **omit it to
|
||
match every node**), and then **one of two replica modes** (they are mutually
|
||
exclusive):
|
||
|
||
- **`replicas: all`** - static spread: place exactly **one replica on every
|
||
matching node**, proactively, regardless of load, and keep it in sync as nodes
|
||
join and leave. Use this for "run model X everywhere (with this label)".
|
||
- **`min_replicas` / `max_replicas`** - elastic auto-scaling: keep at least
|
||
`min_replicas` running, and burst **up to** `max_replicas` only when all
|
||
replicas are busy, scaling back down to the minimum when idle. `max_replicas: 0`
|
||
means **no upper bound** (grow to cluster capacity). To enable this mode you
|
||
must set `min_replicas >= 1` or `max_replicas >= 1` - an entry with only
|
||
`max_replicas: 0` (and no `replicas: all`) does nothing.
|
||
|
||
Net effect at a glance:
|
||
|
||
| Config | Behavior |
|
||
|--------|----------|
|
||
| `replicas: all` | One replica per matching node, placed immediately, tracks join/leave |
|
||
| `min_replicas: 1, max_replicas: 0` | Always >=1, bursts to cluster capacity under load, back to 1 when idle |
|
||
| `min_replicas: 2, max_replicas: 4` | Always >=2, bursts to at most 4 under load |
|
||
|
||
`node_selector` constrains which nodes a model may use; with no selector the
|
||
model may use **all** healthy nodes. So "spread model X across all nodes" is just
|
||
`replicas: all` with no `node_selector`. `replicas: all` targets one replica per
|
||
matching node; with the default per-node cap of one replica per model this lands
|
||
exactly one on each node (see the note below about `LOCALAI_MAX_REPLICAS_PER_MODEL`).
|
||
|
||
YAML example (`scheduling.yaml`):
|
||
|
||
```yaml
|
||
# One replica on every GPU-labelled node (static spread, tracks join/leave):
|
||
- model_name: gpt-oss
|
||
node_selector:
|
||
tier: gpu
|
||
replicas: all
|
||
|
||
# One replica on EVERY node in the cluster (no selector = all nodes):
|
||
- model_name: embeddings
|
||
replicas: all
|
||
|
||
# Elastic on CPU nodes: always >=1, burst to capacity under load, 0 = no cap:
|
||
- model_name: whisper
|
||
node_selector:
|
||
tier: cpu
|
||
min_replicas: 1
|
||
max_replicas: 0
|
||
```
|
||
|
||
```bash
|
||
LOCALAI_DISTRIBUTED=true \
|
||
LOCALAI_MODEL_SCHEDULING_CONFIG=/etc/localai/scheduling.yaml \
|
||
local-ai run
|
||
```
|
||
|
||
Inline equivalent:
|
||
|
||
```bash
|
||
LOCALAI_MODEL_SCHEDULING='[{"model_name":"gpt-oss","node_selector":{"tier":"gpu"},"replicas":"all"}]'
|
||
```
|
||
|
||
Notes:
|
||
|
||
- Because the config is authoritative, each listed model's **entire** scheduling
|
||
row is replaced on every boot, including the optional prefix-cache routing
|
||
overrides (`route_policy`, `balance_abs_threshold`, `balance_rel_threshold`,
|
||
`min_prefix_match`). For a model you manage via this config, set those fields
|
||
here too if you need non-default values; values set only through the API are
|
||
reset on the next restart. Models not listed in the config are never touched.
|
||
- `replicas: all` places one replica per matching node by relying on the default
|
||
per-node cap of one replica per model. If you raise `LOCALAI_MAX_REPLICAS_PER_MODEL`
|
||
on a worker above 1, the target count can be met by stacking replicas on fewer
|
||
nodes rather than spreading one to each.
|
||
|
||
## Label Management API
|
||
|
||
| Method | Path | Description |
|
||
|--------|------|-------------|
|
||
| `GET` | `/api/nodes/:id/labels` | Get labels for a node |
|
||
| `PUT` | `/api/nodes/:id/labels` | Replace all labels (JSON object) |
|
||
| `PATCH` | `/api/nodes/:id/labels` | Merge labels (add/update) |
|
||
| `DELETE` | `/api/nodes/:id/labels/:key` | Remove a single label |
|
||
|
||
## Scheduling API
|
||
|
||
| Method | Path | Description |
|
||
|--------|------|-------------|
|
||
| `GET` | `/api/nodes/scheduling` | List all scheduling configs |
|
||
| `GET` | `/api/nodes/scheduling/:model` | Get config for a model |
|
||
| `POST` | `/api/nodes/scheduling` | Create/update config |
|
||
| `DELETE` | `/api/nodes/scheduling/:model` | Remove config |
|
||
|
||
## Comparison with P2P
|
||
|
||
| | P2P / Federation | Distributed Mode |
|
||
|---|---|---|
|
||
| **Discovery** | Automatic via libp2p token | Self-registration to frontend URL |
|
||
| **State storage** | In-memory / ledger | PostgreSQL |
|
||
| **Coordination** | Gossip protocol | NATS messaging |
|
||
| **Node management** | Automatic | REST API + WebUI |
|
||
| **Health monitoring** | Peer heartbeats | Centralized HealthMonitor |
|
||
| **Backend management** | Manual per node | Dynamic via NATS backend.install |
|
||
| **Best for** | Ad-hoc clusters, community sharing | Production, Kubernetes, managed infrastructure |
|
||
| **Setup complexity** | Minimal (share a token) | Requires PostgreSQL + NATS |
|
||
|
||
## Troubleshooting
|
||
|
||
**Worker not registering:**
|
||
- Verify the frontend URL is reachable from the worker (`curl http://frontend:8080/api/node/register`)
|
||
- Check that `--registration-token` matches on both frontend and worker
|
||
- Ensure auth is enabled on the frontend (`LOCALAI_AUTH=true`)
|
||
|
||
**NATS connection errors:**
|
||
- Confirm NATS is running and reachable (`nats-server --signal ldm` or check port 4222)
|
||
- Check that `--nats-url` uses the correct hostname/IP from the worker's network perspective
|
||
|
||
**PostgreSQL connection errors:**
|
||
- Verify the connection URL format: `postgresql://user:password@host:5432/dbname?sslmode=disable`
|
||
- Ensure the database exists and the user has CREATE TABLE permissions (for auto-migration)
|
||
- Check that pgvector extension is installed if using RAG features
|
||
|
||
**Node shows as unhealthy or offline:**
|
||
- The HealthMonitor marks nodes offline when heartbeats are missed. Check network connectivity between worker and frontend.
|
||
- Verify `--heartbeat-interval` is not set too high
|
||
- Offline nodes automatically restore to healthy when they re-register (no re-approval needed)
|
||
|
||
**InsightFace reports a missing MiniFASNet file after staging:**
|
||
- Gallery models such as `insightface-buffalo-m` use a virtual primary name and load their files through options. The frontend derives the worker's model directory from successfully staged companion files or directories, so relative options resolve inside the model's staging directory.
|
||
- If logs show matching hashes for the staged files but InsightFace still reports a bare filename such as `MiniFASNetV2.onnx` as missing, upgrade the frontend to include this path-resolution fix. Re-uploading the same files does not correct the directory passed to the backend.
|
||
|
||
**Backend not installing:**
|
||
- Check the worker logs for `backend.install` events
|
||
|
||
**Model staging repeatedly fails with HTTP 416 after all bytes have arrived:**
|
||
- An interrupted upload can leave a full-size file marked as unfinished (`.sha256.target`). On retry, the worker verifies the file's SHA-256 and finalizes it if it matches, without rewriting the model. Corrupt content fails integrity validation and is removed.
|
||
- Upgrade the affected worker to get this recovery behavior. Older workers can repeatedly reject retries from byte zero with `Content-Range start 0 does not match current file size`. File size alone is not proof that an upload is valid.
|
||
|
||
**Requests still report an old context size or another old load option:**
|
||
- Query `/api/nodes/:id/models` for every worker that hosts the model.
|
||
- Confirm that every routable replica has `state: loaded` and the same current `config_revision`.
|
||
- Treat a different `effective_options_hash` as diagnostic information. Node-specific defaults can cause valid differences.
|
||
- Check `cleanup_error` and `cleanup_next_retry_at` on replicas in the `unloading` state.
|
||
- Check connectivity to the worker and NATS when cleanup reports a timeout or no responder.
|
||
- Upgrade the worker when it does not support the exact model-stop request.
|
||
- Stop and restart the stale backend only as an operational recovery action. LocalAI keeps it non-routable while durable cleanup is pending.
|
||
|
||
**A model cannot be scheduled on a node that looks free (`no replica slot ... all models busy, cannot evict`):**
|
||
- A replica row in `staging` or `loading` holds its slot: slot allocation counts every state except `unloading`. If a worker drops out mid-transfer, that row never reaches `loaded`, and eviction only ever considers `loaded` replicas, so on a node with one replica slot per model the model became unschedulable there.
|
||
- The reconciler now reclaims a replica row stuck before serving when no load job is still driving it, and the freed slot is immediately reusable.
|
||
- Each replica row names the load attempt that made it. Liveness is decided by that attempt's job lease, not by elapsed time. Staging a large checkpoint legitimately runs for a long time without touching the replica row, so a transfer whose owner still renews its lease is never reclaimed however long it takes. A row is reclaimed when its attempt has no job: the job was released after a failure, or another attempt replaced it.
|
||
- `Reconciler: reclaimed a replica slot held by a load nobody is driving` names each row reclaimed this way.
|
||
|
||
**A request fails with `nats: no responders available for request`:**
|
||
- The chosen worker was not subscribed on the bus when the frontend tried to install the backend on it. A node's status comes from its HTTP heartbeat, which is a separate channel: a worker that stops stays `healthy` until that heartbeat ages out.
|
||
- The scheduler now checks that a node still answers on the bus before it commits to it, marks one that does not as unhealthy, and picks another. A request should therefore see this only when no reachable node is left.
|
||
- Only a no-responders answer counts as absent. A worker that answers slowly stays eligible, because excluding it would cost capacity that is really there.
|
||
- Check the worker process is running and its NATS connection is up. `Scheduled node is not answering on the bus` in the frontend log names each node demoted this way.
|
||
|
||
**A worker fills its own disk over time:**
|
||
- A request that carries a file (an image, an audio clip, a video) stages that file below the worker's HTTP staging or S3 cache `ephemeral/` directory. The frontend releases each request-owned input when inference finishes, and the worker reserves capacity before accepting it.
|
||
- A one-hour recovery sweep runs at startup and every 15 minutes to reclaim inputs left by interrupted requests. It preserves active reservations and uses the newest file timestamp in each request directory.
|
||
- Releases before request-owned cleanup existed can leave a legacy backlog. Delete the affected `ephemeral/` directory once, as the user the worker runs as; capacity admission and recovery cleanup keep new staging bounded.
|
||
- Staged **model** files are not touched by this. They live beside the ephemeral directory and are not per-request scratch.
|
||
- A worker whose volume is genuinely full reports `creating backend process state directory under ...: no space left on device` when a backend starts.
|
||
|
||
**Requests fail with `stale model config revision` although nobody edited the model:**
|
||
- A model's stored revision must describe its persisted configuration. Releases before this fix also hashed the per-request prediction parameters, so the first request after a restart pinned the revision to its own `temperature`, `top_p`, `stop` and similar values. Every later request that sent different values was then rejected.
|
||
- Upgrade the frontend replicas first. After the upgrade the revision is stamped when the configuration is loaded, so it no longer depends on the request body.
|
||
- Each frontend now reconciles the stored revisions against the configuration on disk at startup, and republishes any that disagree, so a drifted revision heals on the next restart. Only models that actually drifted are republished, because republishing quarantines the replicas loaded under the old revision.
|
||
- A model that has never been served has no stored revision and is left alone; its first request establishes one.
|
||
- On a release without that reconciliation, clear the row once per affected model so the next request establishes the correct revision: `DELETE FROM model_config_states WHERE model_name = '<model>';` Saving any edit through the API or the WebUI has the same effect.
|
||
|
||
**Port conflicts on workers:**
|
||
- Each model gets its own gRPC process on an incrementing port (50051, 50052, ...)
|
||
- The HTTP file transfer server runs on the base port - 1 (default: 50050)
|
||
- Ensure the port range is not blocked by firewalls or used by other services
|
||
- Verify the backend gallery configuration is correct
|
||
- The worker needs network access to download backends from the gallery
|
||
|
||
## Routing pipeline
|
||
|
||
Loaded replicas are selected through a filter, scorer, and picker pipeline.
|
||
The initial pipeline applies the load guard as an eligibility filter, scores
|
||
eligible replicas using prefix-cache affinity and cold-placement order, then
|
||
picks the highest score with a deterministic node/replica tie-break.
|
||
|
||
Per-model scheduling fields configure the initial pipeline:
|
||
|
||
- `route_policy` enables `prefix_cache` scoring or selects the
|
||
`round_robin` floor.
|
||
- `balance_abs_threshold` and `balance_rel_threshold` configure the load
|
||
eligibility filter.
|
||
- `min_prefix_match` controls when prefix affinity contributes the highest
|
||
score.
|
||
- `scorer_weights` enables or weights named scorers. The initial scorer is
|
||
`prefix_cache`; set `scorer_weights: {prefix_cache: 0}` to disable its
|
||
contribution while retaining the load filter and deterministic picker.
|
||
|
||
The pipeline accepts additional independently weighted scorers and alternate
|
||
pickers without coupling them to `SmartRouter`. This is the extension point for
|
||
queue depth, precise KV utilization, latency, and fairness signals.
|
||
|
||
## Roadmap: Routing and Caching Enhancements
|
||
|
||
The scheduling algorithm supports **prefix-cache-aware** routing: bias each request toward the replica that already holds the relevant KV/prefix cache (multi-turn conversations and shared system prompts), so backends reuse cache instead of recomputing it. A router-side radix tree maps prompt-prefix hashes to nodes, with longest-prefix match, a load guard that preserves round-robin behavior under imbalance, and NATS sync across frontends. It is purely a routing-layer hint (no backend changes) and never routes worse than round-robin.
|
||
|
||
When the load guard must route away from a warm replica, the frontend emits a forced-disturb event. These events and the reset sent after a successful pressure-triggered scale-up are broadcast over NATS, so the rolling autoscale threshold is cluster-wide rather than per frontend. The Prometheus counter `localai_prefix_cache_forced_disturb_total{model="..."}` records events at their originating frontend; sum it across frontend replicas to inspect cluster pressure without counting the NATS copies.
|
||
|
||
Backends can report exact KV-cache residency on the `prefixcache.residency`
|
||
NATS subject. The JSON event contract is:
|
||
|
||
```json
|
||
{
|
||
"operation": "store",
|
||
"model": "model-name",
|
||
"node_id": "worker-id",
|
||
"replica": 0,
|
||
"chain": [1203053429005847826, 15485907386658061715]
|
||
}
|
||
```
|
||
|
||
`operation` is `store`, `remove`, or `clear`. `store` adds the announced
|
||
shallow-to-deep chain for one model replica, `remove` removes only that exact
|
||
announced chain, and `clear` removes all reported residency for that model
|
||
replica (and may omit `chain`). Producers must generate the chain with exactly
|
||
the same windowing and hashing algorithm as the router; hashes from a different
|
||
chain algorithm are not compatible and will never match requests correctly.
|
||
Reported events populate the exact-residency provider, but the guessed provider
|
||
remains the routing default until a backend producer is available.
|
||
|
||
Further enhancements, surfaced from a survey of SGLang, vLLM production-stack, Ray Serve, llm-d, AIBrix, and NVIDIA Dynamo, are tracked under the routing roadmap epic ([#10063](https://github.com/mudler/LocalAI/issues/10063)):
|
||
|
||
- **Reported/precise KV-event mode** ([#10064](https://github.com/mudler/LocalAI/issues/10064)): subscribe to actual backend KV-cache events for exact residency instead of inferring it from routing history.
|
||
- **Multi-tier cache-overlap scoring** ([#10065](https://github.com/mudler/LocalAI/issues/10065)): credit GPU/CPU/disk cache tiers separately.
|
||
- **Load-shaping** ([#10067](https://github.com/mudler/LocalAI/issues/10067)): anti-herding (softmax/temperature) and dispatch-time freshness.
|
||
- **Prefill/decode disaggregation routing** ([#10068](https://github.com/mudler/LocalAI/issues/10068)): route prefill and decode to separate pools with KV transfer.
|
||
- **Per-user fairness (VTC)** ([#10069](https://github.com/mudler/LocalAI/issues/10069)): balance per-user token usage against pod load.
|
||
- **Minor tuning + MCP parity** ([#10070](https://github.com/mudler/LocalAI/issues/10070)): per-model TTL override, probabilistic LRU updates, and MCP scheduling-config tool parity.
|