mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-10 15:52:29 -04:00
65623166f0d546c8dbe82f319a44aefc39adde9a
100
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
65623166f0 |
fix(ui): keep model actions visible on phones
Let model descriptions wrap without losing the install action offscreen. Give page table rules precedence over the shared table styles. Separate background facet probes from navigation assertions and wait for asynchronous size sorting and masonry layout. Fix the chat fixture clock so runs near midnight keep the expected conversation groups and order. Assisted-by: Codex:gpt-6 |
||
|
|
8db1d165dd |
feat(ui): add moderated group conversations
Let users give installed models individual turns or ordered rounds in a shared conversation. Attribute completed responses and exclude interrupted output from subsequent prompts. Add streaming and cancellation tests, browser coverage for CI, and a guide for the session-only page. Assisted-by: Codex:GPT-6 |
||
|
|
4d7bdc6ff2 |
feat(ui): redesign the web UI around a shared kit and a calm palette (#12526)
* build(ui): vendor the shared UI kit snapshot at 0.2.0 The restyle needs the kit's tokens, motion layer and component classes. Take a pinned snapshot instead of depending on the kit at build time, and keep a lock file with the version and per-file checksums so a later update shows exactly what changed. The product theme stays outside the vendored directory. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the ink and teal theme and bridge the old variables Define the product colours as the shared UI kit's roles, for light and dark, in theme-localai.css. The kit's contrast check passes on every pair. theme.css keeps the existing --color-* and --shadow-* names but now points each at a role, so App.css and the pages get the new palette without edits. Radii move to the kit scale. index.html now sets data-theme before first paint with the same rule as ThemeContext (stored choice, otherwise dark), because the contract layout of the theme file no longer defaults to dark by itself. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): restyle the shared chrome with the UI kit grammar Adjust the shared classes so every page picks up the same interaction language without per-page edits: - Sidebar sits on the canvas and the current row lifts onto a card. Section labels are tracked uppercase, the badge is a soft pill, and the phone drawer leaves the tab order when closed. - Buttons are flat: hover swaps the surface, press scales to .97, focus is a 2px ring with a 2px offset, danger is a tinted wash. - Inputs use the card surface and the control edge; switches, tabs, filter chips, badges and cards follow the same rules. Cards no longer lift on hover; only linked or button cards react. - Menus and popovers scale in from the trigger corner with 40px items. Dialogs get a veil fade and a spring settle. Toasts become pills at the bottom centre. - The page transition is a 250 ms fade with a 6px rise. It fills backwards so a finished animation no longer leaves a transform that confined dialog veils to the main column. The focus-ring test now checks the outline instead of a box shadow, and new specs cover the theme roles, the first-paint theme and the sidebar lift. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): move leftover hard-coded colours onto the theme roles The YAML editor restated the old blue palette in JavaScript, and a few pages kept literal blues, indigo and violet tints, or fallbacks that only applied because a variable was never defined. Point them at the theme variables so they follow light and dark and the new palette. The status badges in the account pages built their tint by appending "22" to a variable, which is not valid once the variable is defined, so they had no background. Use the wash roles instead. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): raise the type scale and control size toward the kit Body and list text moves to 15px and the rest of the scale follows the kit's 12/13/15/17/21/32/44 steps. Page titles, section headings and stat values are bold with tighter tracking; titles are 32px. Buttons, inputs, selects, tabs and nav rows are 40px high with the 12px radius, compact controls 32px. Tabs become a segmented control. The sidebar widens to 240px (64px collapsed) and nav rows get more room. Identifiers and counts in the split views use the mono face, and the stat grid becomes separate inset tiles. The Geist stack stays: it is bundled, and the thin look came from the size, weight and negative tracking, not the face. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): separate cards, panes and floating surfaces from the canvas Cards, the Models and Installed split panes, the composers and the confirm dialog use a stronger card edge, the rest shadow and the 20px radius, so they read as layers in dark as well as light. Menus and popovers move to a float surface (the hover tone in dark) with the float shadow. The selected rail row gets an accent wash and a 3px accent edge. The send buttons are a clear accent when there is something to send and a quiet inset when not; the Home button carries data-empty for that, since submitting an empty box does nothing. The assistant card becomes an accent wash with a square icon. New surfaces spec checks the pane edge, the selected row, both send buttons and the popover in both themes. The voice library empty-state spec now waits for the layout to settle before comparing two boxes. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): tidy the sidebar header and mark the current row with a dot The header gives the configured horizontal logo a fixed width and centres it in a 72px band, lined up with the nav icons. The collapsed rail shows the configured icon logo centred, and its nav rows become 40px tiles centred in the 64px rail. The current row gets the kit's accent dot, hidden in the rail. The theme, language and account controls stay in the sidebar footer: the app has no global search or command palette to put in a top bar, so a bar would only hold controls that already have a place. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): centre the avatar in the collapsed sidebar rail The collapsed avatar link was set to "flex: 0", which gives it a zero flex basis; with min-width: 0 the link shrank to its padding and the icon overflowed from the link's left edge, about 14px right of the icon column. Use "flex: 0 0 auto" in the collapsed and tablet rail. The footer controls now share the nav icon column in the expanded sidebar too (6px footer padding, 40px control boxes), and the tablet rail gets the same footer padding and hidden language code as the collapsed one. New spec measures the centre x of the nav icons, mark, avatar, language, theme and collapse icons in the collapsed, expanded and tablet states, in both themes, and asserts they agree within 1px. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): size the console and settings rails and stack settings on phones The Operate console rail, the Settings section rail and the account tab bar still used 13px text and the old underline tabs. They now use the 15px nav size, 40px rows and the segmented tab control. Form row labels are 15px with 13px hints. On a phone the Settings section rail sat beside the form and squeezed every row into a few characters. Below 720px the rail stacks above the content as a scrolling row and form rows wrap their control below the label. The save button no longer carries the icon font class, which drew a missing glyph before its label. The language menu is wide enough to keep Bahasa Indonesia on one line. Assisted-by: Claude Code:claude-sonnet-5-5 * build(ui): update the vendored UI kit snapshot to 0.3.0 Take the 0.3.0 snapshot: the sprite now carries the full outline icon set, and the new icons/fa-map.json maps Font Awesome names to icon ids. The map lets the app move off Font Awesome in the following commits. The lock file is regenerated with the new checksums. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add an Icon component backed by the kit sprite Icon draws an inline svg that points into the kit's outline sprite. The sprite is inlined into the page once, so the references resolve under any base path and in the embedded build without a request. Icons size with the font (1em), take currentColor, hide from assistive tech unless given a title, and spin on request. An unknown id draws a neutral circle. FaIcon and iconFromFa resolve Font Awesome names through the kit's map, for names that arrive at run time. iconHtml does the same for markup built as a string. The GitHub and Apple marks are small local glyphs, as the kit ships no brand marks. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): draw shared components and helpers with Icon Replace the Font Awesome elements in the shared components and in the utility modules with the Icon component. Lookup tables now hold kit icon ids instead of class strings. Code-block copy buttons and artifact cards, which build HTML strings, use iconHtml and a sanitizer-safe slot. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): draw model, backend and account pages with Icon Replace the Font Awesome elements on the home, models, backends, import, settings, login, account and users pages with the Icon component. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): draw chat, studio and recognition pages with Icon Replace the Font Awesome elements on the chat, media generation, talk and face and voice pages with the Icon component. The talk status table keeps its spin and pulse states as Icon props. The connected and error states now use a dotted circle and an alert circle, so they differ from the idle ring by shape as well as by colour. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): draw agent, node and operate pages with Icon Replace the Font Awesome elements on the agents, skills, collections, jobs, fine-tune, quantize, nodes, swarm, usage, traces and activity pages with the Icon component. Two class strings on layout elements held leftover button and icon classes from an earlier merge; they are cleaned up so the elements keep only their own classes. Assisted-by: Claude Code:claude-sonnet-5-5 * style(ui): size and align icons for the svg component Icon rules that targeted the font element now target the svg: the descendant "i" selectors in App.css and auth.css become ".lai-icon". The svg is 1.2em with a 2 unit line so it matches the visual size of the old glyphs at the 12 to 16px sizes the app uses, sits on the text baseline, and follows the context font size. Large empty-state marks get a lighter line. Menu icons get a 16px box and the readiness badge icons keep their 20px circle with padding. Add the pulse used by the talk status. Assisted-by: Claude Code:claude-sonnet-5-5 * build(ui): remove Font Awesome No source file references the icon font any more. Drop the package and its stylesheet import. The build no longer ships the solid, regular and brand font files. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): keep focus traps off the svg use references The dialog and drawer focus traps collect focusable elements with a "[href]" selector. An icon's use element carries an href, so it became the first "focusable" element and Tab at the end of the dialog stopped there instead of wrapping to the first button. Match "a[href]" instead. Assisted-by: Claude Code:claude-sonnet-5-5 * style(ui): keep icon sizes overridable and set the line width per svg Give the icon base rule zero specificity so a rule that sizes one icon (nav column, menu box, avatar, language switcher) wins whatever its order in the file. The sprite symbols fix their own line width; the inlined copy drops it so the width set on each svg applies, as the --lai-stroke custom property, and large marks can use a lighter line. Pin the avatar and the language globe to the boxes the sidebar alignment spec expects. Import the map as JSON with an import attribute so Node can load it in the spec. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): select icons by the svg markup and cover the sprite Specs that found icons by their Font Awesome class now select the svg by its data-icon. The dead-icon audit checks that every svg resolves to a sprite symbol and has a size. The class hygiene spec fails on any remaining Font Awesome class. A new spec checks every mapped icon id has a symbol, that the sprite is inlined once, that an icon paints at the root and under a forwarded path prefix, and that Font Awesome names map as documented. Assisted-by: Claude Code:claude-sonnet-5-5 * build(ui): update the vendored UI kit snapshot to 0.4.0 Take the 0.4.0 snapshot: hub tabs with count and attention badges, the six chart series tokens and the grid colour in the theme contract, and sample themes on a calmer palette. The kit headers are renamed and the lock file is regenerated with the new checksums, as for the earlier snapshots. Assisted-by: Claude Code:claude-sonnet-5-5 * style(ui): switch the theme to the calm palette Rewrite the LocalAI theme on the calm palette: a muted teal accent on a near-neutral green-grey canvas, desaturated status colours, no glow and no coloured shadows. The theme fills every role of the shared UI kit's 0.4.0 theme contract for light and dark, including the six chart series and the grid line. The bridge in theme.css keeps the old --color-* names working, adds the dark surface ladder (card, raised, float) and a strong edge, and points the fixed data hues at the chart series. Two values differ from the first sketch. The dark text on the accent fill is #021512 instead of #04201d: it reads 5.58:1 on the fill at rest and 6.4:1 on the hover fill, against 5.08:1 at rest for the lighter value. The light control edge is #6b7d7a. The kit's contrast script passes for all text pairs (4.5:1), control and focus pairs (3:1) and series colours (3:1). Leftovers that no longer fit the palette are fixed: the usage chart takes the six series colours in order, the audio and animation canvases fall back to the new accent, the face box loses its glow, and two gradient fills are now flat. The theme tests expect the new canvas colours. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): replace the console rail with hub tab bars Build and Operate no longer open a second navigation rail beside the page. Each is a hub: one row of the kit's hub tabs above the page, with count and attention badges that scroll sideways on a phone. Every URL and route stays as it was, plus a new /app/build landing page that lists the Build tools with a line each. Build tabs: Overview, Agents, Skills, Memory, Jobs, Fine-Tune, Quantize, Import, Voices (recognition and library) and Faces. Operate tabs: Status, This machine, Swarm (distributed mode only), Runtime (backends, activity, failover), Traffic (usage, traces, middleware) and Settings (settings, users), plus the API link. A tab that holds several pages shows a second row of links, and a sub-page such as a node detail keeps its tab highlighted. The feature and admin gates decide which tabs are drawn, and badges show only values the Operate summary already has. The sidebar lists Build and Operate under a Workspace label next to the Create group. The voice library moves under Build and the model import page gains the Build tab bar. The old rail styles, the rail signals and the console config are removed, and the Operate overview docs describe the tab bar. The specs that drove the rail now drive the tabs, and a new spec covers the tab for each route, gating, badges and the phone layout. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Home as a calm console Home now opens on one command bar: the model chip shows which models are warm, the MCP chip and attach buttons sit beside it, and Send is a solid button with an Enter glyph. Typing "/" opens a grouped, keyboard-driven action list built on the kit command list; every action has a destination in the product. Memory use folds into a one-line strip that opens into the loaded models, with Stop per model and Stop all. It opens by itself while a model is being staged and after a failure, and shows nodes and aggregate memory in a cluster. The list of resident models carries no per-model size because the API reports none. "Jump back in" lists the conversations stored in the browser, one card per day, with j and k to move, Enter to resume and delete with an undo toast. First run keeps the install steps and the recommended models. The assistant prompt is a dismissible line, the library links are one quiet row and the API section is collapsed. Chat accepts an empty new-chat hand-off for /new. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Home console Update the Home specs for the new structure and add specs for the slash menu, the model chip, the memory strip (expand, stop, staging, failure, cluster), the resume list (grouping, j/k, Enter, delete with undo), first run, the send hand-off, a non-admin user and the phone layout. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the fit, disk and cleanup helpers for the models page Pure functions and hooks that the rebuilt Models page reads, with node tests for the rules. modelLedger turns an estimate and the memory budget into one of three verdicts (fits, spills to CPU, over) with the headroom in bytes, and reads the models disk from the resources reading. The disk counts as low under 10 percent or under 20 GB free, and is absent when the server reports none or runs as a cluster controller. cleanupPlan ranks installed models from what the API reports: loaded, pinned, or named by an agent, a task, a failover chain or an alias keeps a model protected; another installed build of the same gallery model is a duplicate; disabled models rank above idle ones. The API records no last use or use count, so none is used. When a lookup fails, nothing is called safe. useModelRemoval holds a removal in the browser for an undo window and sends the existing delete call only when the window ends. Leaving the page drops the batch without deleting anything. The undo toast takes optional labels so other pages can reuse it. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Models as a ledger with a disk strip and cleanup review Explore is one dense table. Each row carries the size, a solid memory bar and the headroom in words ("3.7 free", "+1.5 on CPU", "0.9 over"), worked out from the estimate at the chosen context length. Capability chips show the server's count for each facet, search keeps its meaning and "/" jumps to it, and a density switch (also "d") picks comfortable or compact rows. Selection is a surface step and a check, never a rail. Arrow keys move, Enter installs and Esc closes the inspector, which keeps the fit summary, VRAM by context chart, variants, files, links, tags and licence. A failed install shows its error in the row with a Retry that dismisses the old failure first. A failed or empty listing says which it is, and a host with no GPU is measured against memory and says so. Installed uses the same table with state filters that carry counts, a state per row, Load or Stop on the row, the row menu and the sort by size. Sizes come from the files the gallery lists, so a model it does not know shows a dash. A strip in the header shows the free space on the models disk. It turns amber under 10 percent or under 20 GB free, hides when the server reports no disk or runs as a cluster controller, and opens the cleanup review. Explore says how much an install leaves free. The review ranks installed models as Safe to remove, Probably safe and Your call from real facts only, lists protected models with the reason, and says plainly that usage history is not recorded. A sticky bar shows what a choice frees. Confirming runs a dry run that checks again and lists what will go. Removal waits 30 seconds with an undo; nothing is deleted before that, and leaving the page deletes nothing. The old rail, filter band and popover styles are removed. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Models ledger, Installed table and cleanup review Update the Models, lifecycle, cluster fit, height, search focus and surfaces specs for the table and inspector, keeping what each one checks. New specs, on a shared 41-model gallery stub with three machine profiles: the fit bar and headroom words for a 24 GB card, an 8 GB laptop and a host with no GPU; facet counts, search, "/" and Escape; selection, arrow keys, Enter to install, density; the disk strip when normal, low and hidden; and the states (loading, empty, offline, install failed, phone). Installed covers filters with counts, row actions, the row menu, sizes and sort. The cleanup specs cover grouping, protected models, the honest-data note, the effect bar, the dry run, the undo window, a failed delete, leaving the page, and the phone sheet. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the placement helpers and the estimate hooks The Placement section and the model page need the same few rules, so they sit in plain functions that can be read and tested alone. placement.js holds what gpu_layers, tensor_split and main_gpu mean (unset asks for every layer and the llama.cpp engine trims it, zero is CPU only, 99999999 is the value LocalAI itself writes for all layers), the device list taken from the resources reading, the split by free memory, the part of an estimate that grows with context (read from two lengths, since that term is linear), the fit states with their limit (95 percent of free memory, and the leftover has to fit in system memory too), and a bisection for the largest layer count whose estimate fits. The estimate returns one total and no layer count, so the search runs over 1 to 256 and stops at the first count that no longer changes it. modelWalk.js keeps the order of the list a model page was opened from, in memory and in session storage, for the previous and next buttons. usePlacementEstimate reads /api/models/vram-estimate for a choice, again at twice the context, and with every layer, and keeps readings for the session. useModelPage reads a gallery entry by name, an estimate by context size (from the model's own files when the gallery does not list it), the builds and the loaded models. usePlacementConfig edits the four placement keys of an installed model and saves only what changed. useModelActions is the Load, Stop, disable, pin and remove logic of the Installed table, shared with the model page. MemoryBar is one solid bar with a tick at the capacity of its pool; over capacity it grows past the tick and the tick turns red. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the Placement section to the model editor Run this model on: CPU only (gpu_layers: 0), Auto (the key stays unset) or Custom. Custom takes a number, has an All layers button that writes 99999999, and shows a slider only when the estimate reports the model's layer count, which it does not today. Context size has presets and a number field because the KV cache follows it. With two or more GPUs there is a split (written as percentages, with a button that takes them from the free memory of each card) and a main GPU. A bar per GPU and one for system memory show what other programs use, the model's weights and working memory, and the part that grows with context, with the room left or how far over it is. Under them a verdict in plain words: Fits in GPU, Spills to CPU, Too many layers for the GPU, Runs on CPU only, No GPU found, Not enough memory. It says "slower" and never a multiplier, because the estimate has none. Fit it for me asks the estimate for the largest layer count that fits the free GPU memory and says what it set, with Undo; it is hidden when the estimate is unavailable or the host has no GPU. Loading shows skeletons, an unavailable estimate shows a note with Retry, and a server that schedules onto other machines shows no bars, because its device list is the controller's. The editor shows the section for an installed model, with a link in its section rail. Auto sends null for the key, since a patch only merges, and a null read back opens as Auto. The docs describe the section and what each mode writes. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): open a model on its own page A model has an address, /app/models/<name>, for an installed model and a gallery entry alike. Open it from the arrow at the end of a row, a double click, "o" on the selected row, the inspector's Open details button, or a tap on a phone. The title block holds the main action: Install with a chevron that chooses the build, or Load and Stop with a menu (disable, pin, edit configuration, logs, delete with a confirm). A strip answers whether it fits, what it does and what installing leaves free. Tabs: Overview (about, a memory bar, state, the pages it opens in, and the agents, tasks, chains and aliases that name it); Fit and memory (verdict, context sizes, the bar split into weights and context, and memory by context against the limit, with a data table); Variants and files (builds with size and fit, install any, the files of the chosen build). For an installed model also Usage and history, which says what the API does not record instead of drawing an empty chart, Configuration, which is the Placement section with the file it writes and a link to the full editor, and Logs, the backend log viewer without its page. Keys 1 to 6 switch tabs, [ ] and j k walk the list the page was opened from, Esc or Backspace go back. The list stays mounted behind the page, so Back finds its view, search, filters, selection and scroll as they were, and focus returns to the row's arrow. The page covers loading, an unknown name with the closest matches, the gallery being out of reach, an install in progress with Cancel, and a failed install with Retry. The docs describe the page and its keys. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the model page and the Placement section New specs for the model page: reaching it from Explore, Installed, a double click, "o", a pasted link and a phone tap; the walker and Back with the search, a filter, the selection, the Installed view and the scroll kept, and no second read of the gallery; the title block, the answer strip, tabs by click, keys and arrows; Fit and memory, builds and files with the install call each one makes; an installed model's actions, used-by, the honest usage tab, configuration and logs; loading, an unknown name, offline, an install in flight and a failed one; and the phone. New specs for Placement: every mode and the keys it writes, the slider only when a layer count exists, the context presets, the bars and every verdict, two GPUs, no GPU, a cluster, an unread machine, a loading and an unavailable estimate, Fit it for me and Undo, and the section in the model editor with its save. The phone tap on a row now opens the page, so the two phone specs that expected the inspector as the page check the page and keep the inspector check for a window between a phone and a desk. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): link Studio results and open workspaces from a prompt Each workspace now records the result it was made from (parentId and an edge kind such as take, animate or to-3d) and reads a prompt, model, size, count and source from the query string, so one page can hand work to another. A source result is fetched from the server's own output file and becomes the start image, the picture for 3D, or the audio file. A note on the page says when the source loaded or could not be loaded. Diarization had no history; it now keeps the file name, the model and a speaker count, never the recording. Prompts are cut at 2000 characters when stored. The pure helpers (type suggestion, grouping, lineage layout, favourites, clearing) have node tests. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Studio front page as a composer with your work The front page is a prompt box with a chip per type, a type suggestion from the words, starters, and the options each workspace accepts. Generate opens the workspace with those filled in. A type with no model is a dashed chip that shows a gallery model, its size, memory need and an Install button only when picked; the typed words stay while it installs. Under it, Your work lists results from every workspace as a masonry with filters, counts, favourites and a Clear history action. Results made from each other stack into a project tile and open as a lineage board with a dock for running a new take or branching to the next step; steps the destination cannot start from yet are disabled with the reason. The docs describe the page, what is stored in the browser, and the query parameters a workspace accepts. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Studio composer, your work and the lineage view Specs for the type suggestion, the keys, hand-off to each workspace, the install path for a missing model, the masonry filters, favourites and clearing, stacking, the lineage board, new take and branch, steps that are disabled with a reason, and the phone layout. Existing Studio specs move from lanes to chips with the same intent. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the shared Studio workspace frame and move Images onto it The seven Studio workspaces get one layout: a row of type tabs, a compose card (optional sources as chips, a prompt with starters, a model chip, the essential options as chips, an Advanced fold that names what is inside, the memory the model needs, and one action with the reason when it cannot run), a run area, and a strip of recent results of the type. The run area shows a job card with the time that has passed and an indeterminate bar, because these endpoints report no phase or percentage; a failure with what the server said and one action; or the result with a toolbar: Favourite (the list the front page keeps), Download, Use in (the hand-off targets, disabled with the reason when a destination cannot start from the result), Re-run with edits (the take's values go back in the form, changed fields are outlined and listed) and Lineage. A type with no model shows the install note from the front page. Images is the first workspace on the frame. It keeps its size, count, steps, seed, negative prompt, source image and reference images, and its history writes, including the parent link and edge of a hand-off run. useMediaHistory.addEntry now returns the id of the entry it stored. The docs describe the workspace page. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move Video onto the workspace frame Video keeps its size list, duration, frame rate, steps, seed, CFG scale, frame count, negative prompt, start and end image and avatar audio. The start and end image are source chips, the avatar audio opens the recording and paste input from a chip, and the rest sit in the Advanced fold. A start image from a hand-off shows as a chip with its picture. Results play in the video player with the shared toolbar. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move TTS onto the workspace frame TTS keeps the saved-voice picker for cloning models, the typed voice for the others, the voice library deep link, and the delivery instructions, which now sit in the Advanced fold. The result is the waveform player with the words under it. The stored entry also keeps the voice id so Re-run with edits can select the same saved voice. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move Sound onto the workspace frame Sound keeps its Simple and Advanced modes and every field of both: the description, instrumental, vocal language, caption, lyrics, BPM, duration, key, language, time signature and think mode. The mode switch, instrumental and duration are in the compose card, the rest in a More options fold. The stored entry keeps all of the fields, so Re-run with edits restores the form as it was. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move Transform onto the workspace frame Transform keeps its audio and reference inputs with upload and record, the echo test, the key=value parameters (now in the Advanced fold), the input and output spectra and the three waveform players. The audio that was chosen shows before the run, waiting to be transformed. Re-run with edits puts back the model and parameters and fetches the audio and reference the server kept for that run. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move 3D onto the workspace frame 3D keeps the picture input with paste and webcam, the animation operations a model declares, quality and background, the shape and material steps, guidance and seed, the GLB and animation viewers, the remesh control and the download. A 3D result now has a title from the motion prompt when it has no label, so the strip and the front page name animation results by what was asked. Re-run with edits is shown disabled with the reason, because only a small thumbnail of the picture is kept. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move Diarization onto the workspace frame Diarization keeps its model and recording inputs, the option to prepare speakers to remember, the clean-speech previews, naming and remembering a speaker, and the history entry with only the file name, model and counts. The result now shows a timeline with one lane per speaker, the talk time of each speaker, and the segments with their start time and text. RTTM, SRT (only when the run has text) and JSON are built in the browser from the result. The helpers for talk time, axis ticks and the two text formats have node tests. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): remove the styles and lists the old workspace layout used Nothing renders the two-column workbench, the control column, the old history lists, the generation progress tiles, the TTS voice picker or the result echo any more. The inline-style baseline drops with them. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the workspace frame and one run per type Specs for the type tabs, the compose card and its reason when Generate cannot run, starters, the Advanced fold, the job card with no invented progress, a failed run and its one action, the install note, the strip with its favourites filter, Use in with its disabled steps, Lineage, the parent link, Re-run with edits and its list of changes, deleting and clearing, and the hand-off note. One run through each of Video, TTS, Sound, Transform, 3D and Diarization, the phone layout of all seven, and reduced motion. Existing Studio specs move from the old control column to the compose card with the same intent. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the chat thread with raised user turns, prose replies and one-line activity Your messages are raised blocks on the right at a 760 px measure and the model's replies are plain prose under its name and a warm or not loaded dot. Reasoning, tool calls and their results fold into one quiet line that opens inline into steps. Code blocks carry a Copy button and a Canvas button that opens that block in the canvas, image attachments are thumbnails that open in the lightbox, and files are chips. Per-message actions show on hover, on focus and on the last turn, and a turn takes focus so the arrow keys and C, E, R and B work. A failed reply keeps the text written so far and shows the reason with one Retry action. The Agent chat page keeps the older rules: the new styles are scoped to the chat page and use their own class names. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): use the Home command bar as the Chat composer Chat now ends in the same object as Home: the model chip, the MCP chip, a Canvas chip, the message box with attach buttons, a solid Send and the hint line, with the slash menu on the kit command list. The slash menu lists what Chat can do today (switch model, new chat, conversations, manage mode, canvas, find, settings, export, clear). While a reply is streaming Send becomes Stop, which Esc also presses, and Up in an empty box edits your last message. Attached images show as thumbnails and a line under the bar carries the speed and the token count. HomeComposer takes optional props for this (extra chips, its own slash list, Stop, paste, a stricter Enter); Home passes none of them. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): open conversations from a Ctrl K menu with day groups, undo and a slim header The conversations list opens as a centred menu on Ctrl or Cmd K. It groups chats by day like the Home resume list, shows the model that answered and the time, searches names and message text, and moves with the arrow keys. Enter opens a chat, F2 renames it and Delete removes it. Removing a chat hides the row and shows the kit undo toast; the chat is deleted for good only when the undo time ends. Rename, duplicate, copy and export are on each row, as before. The header is one slim bar: the Chats button, the chat name (click to rename), a context meter when the context size is known, settings and a More menu with rename, duplicate, copy, export, model info, keyboard shortcuts and clear. A dialog lists the shortcuts the page answers to. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): show loaded state, capabilities and fit in the Chat model switcher The model chip in Chat opens the same list as Home, grouped as Loaded now and Installed. Each row says warm or not loaded and marks models that understand images. When the list opens, the page reads the host memory once and asks the server to estimate each listed model at the chat's context size (up to twelve, three at a time), then shows what the model needs and whether it fits: free memory, how much would run on the CPU, or how far over the machine it is. A model with no estimate shows no fit text, and no load time is shown because the API does not report one. A memory bar closes the list. The picker takes the model list from the page when it has one, and useModels can skip its own request. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move chat settings into a sheet and add find in chat, jump to latest and a wider canvas Settings open as a kit sheet: the system prompt, temperature, top P and top K (each says "model default" until it is changed and has a Reset), the context size with quick sizes and a note that it only drives the meter, Manage mode and Focus mode, the model info for admins with its Edit config button, and Clear conversation behind a confirmation. The old slide-out drawer and the model info panel are gone. Ctrl or Cmd Shift F (or the search button, or /find) opens a search bar over the thread. It marks matches in the messages already on the page, shows "n of m" and steps with Enter and Shift+Enter. Nothing is sent to the server. Jump to latest is a pill above the composer. Esc stops a reply, then closes the search, then closes the canvas. The canvas panel gets the kit look: tabs, a Code and Preview switch, Copy and Download, a full-page layout on narrow windows, and translated labels. The Agent chat page shares it and gets the same look. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the empty, no-model, loading and phone states to Chat An empty chat opens with the composer under one line, starters to try, whether the model is loaded, and the Jump back in list: the same rows as Home, read from the chats the page holds. With no chat model installed, an install card offers the starter models for this hardware, the gallery and import, and the composer stays so the text is not lost. While a reply waits for a model, a load card shows what the page knows: the phase the server names, the node, the bytes and the time left when the server reports them, and a progress bar. A model that is just not loaded yet gets a plain note, with no invented phases or estimates. The foot warns when the context is nearly full. On a phone the header drops its labels, the model list and the settings open as sheets from the bottom, per-message actions stay in view and the canvas takes the whole page. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Talk as a calm voice page over what the connection really does Talk is one stage and one transcript. The stage has the pipeline chip, the voice and language chips, an outline orb that follows the real microphone and playback levels, a heading and a sentence for the current state, and the controls. The transcript lists You, Reply, Tool and Result lines and can be copied. Session settings (instructions, voice, language, tools, Manage mode and the pipeline's parts) open in a sheet. The states are the ones the code reaches: no pipeline model, idle, connecting, listening, thinking (also while a tool runs), speaking, an interrupted reply (the server cancelled it; a note marks the cut), a blocked microphone, a link that failed during a session, and any other error with its reason and a link to the traces. Push to talk and hands-free are not on the page, so they are not shown. Diagnostics keep their waveform, spectrum and stats, drawn in theme colours. The page text moves into the talk namespace, and the old Talk and visualizer styles and the inline-style count go down with the rebuild. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): remove the chat styles and strings the rebuilt page replaced The settings drawer, the model info panel, the bubble avatars, the conversation menu popover, the context bar, the recent strip, the staging bar, the file badges and the focus-mode rules have no user now. Their rules, the Chat page's focus class and seven unused empty-state strings are removed. The Agent chat page keeps the shared message, sidebar and input rules it still renders with. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the rebuilt Chat and Talk pages Add a Chat page under features (thread, message actions and keys, the message box and its slash actions, the model list with loaded state and fit, conversations on Ctrl K, settings, find, canvas and the empty, no-model and loading states) and a Talk section to the realtime API page with the states the page shows. Manage mode now turns on from the chat settings or /assistant, and the client MCP steps point at the MCP chip. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): settle the rough edges of the new Chat and Talk pages The undo toast sat under the conversations menu, so the Undo button could not be pressed while the menu was open; the menu, the sheets and the fullscreen canvas now stay below the toast layer. Esc in a rename box saved the text through the blur that follows it; it now cancels. The image viewer closed on Esc only when the page did not re-render on the same key, so its key listener is registered once and reads the latest handlers. Keys on a focused message no longer type their letter into the editor they open, "/" from outside a text field starts a command as it does on Home, and Esc leaves the page's own dialogs alone. Code in the canvas is highlighted for languages that have no preview. The conversations menu drops its key hints on a phone so Clear all stays in view. Talk hides Test tone while connecting and calls a server error "Something went wrong", since the call can still be open. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the rebuilt Chat and Talk pages Specs for the thread layout and the activity fold, code blocks, image thumbnails and the viewer, per-message actions and their keys, a failed reply with its one Retry, Stop and Esc while streaming, the composer and every slash action, the conversations menu (groups, search, resume, rename, delete with undo that ends by itself, one chat left), the model switcher with loaded state, vision and fit text from stubbed estimates, the settings sheet, the canvas panel, find in chat, Jump to latest, the empty, no-model and loading states, the phone layout and reduced motion. Talk is driven over a fake WebRTC link through idle, connecting, listening, thinking, speaking, interrupted, blocked, lost, error and no pipeline, its settings sheet and its phone layout. Node tests cover the message text helpers and the conversation grouping. The existing chat specs move to the new structure with the same intent: the transcript spec now describes the raised turn and the prose reply, and the render smoke accepts Talk's own header. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): keep a bounded run log for agents in the browser The server keeps no run history for an agent, so a run is one task and the events until the agent answers, written to browser storage while the page watches the stream: up to 50 runs per agent, task, step and answer text only. Stored chats from the earlier agent chat page read as runs with stable ids. A run still marked running five minutes after its last event reads as stopped. Helpers read an agent's config into chips, build the list of changed fields against the saved config, hide secret values and offer starting points. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Agents area around runs The Agents page shows what needs a look (work in flight, a run that failed in the last day), then each agent with its model, attached memory and skills, and a strip of its last 14 runs. An agent has its own page: model, tools, memory, skills, instructions, a task box and its runs. A run has an address, shows the thread while it works (steps folded into one line, the tool in use, the answer as it arrives) and settles into a report about a second and a half after the agent answers: task, outcome, follow-ups, evidence and steps, with wide tables opening wider on demand. A failure says in plain words what happened and offers Run again. Create and edit fold into sections with a ready mark and a one-line summary, start from a template or an optional model-written draft, and open a preview sheet with the config as saved and the changes against the saved agent. Status becomes a quiet panel in the same language, and the old chat link opens the agent page. There is no Stop, approval, steer, version or dry-run control, because the agent API has no call behind them. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Agents launcher, agent page, runs and editor Specs for the Now strip and run strip, search and the empty state, the agent page, starting a run, the live thread, settling into the report, the run address across a reload and for a run from another browser, follow-ups with their history, failures, the folding editor with ready marks, templates, the preview sheet with hidden secrets and changes, the status page, and the phone, 1440 and 2560 layouts. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe runs and the new agent create flow Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add helpers for tasks, schedules and job outcomes Reads a cron expression the way the server does (five fields or an @ shortcut), checks it, and puts the common shapes in words. The next run is left out on purpose, because the schedule follows the server clock, which the browser cannot read. Also groups jobs by day, sums the last seven days, and gives each job one outcome line from its result or error. A rerun call starts a new job with the same parameters and media. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Jobs area around runs The Jobs page opens with one sentence about the last seven days, then the tasks (model, schedule in words, last 14 jobs, enabled switch, Run now) and a run history grouped by day. Each row has an outcome sentence and opens to the error or the start of the result with one next action. Deleting a task waits 30 seconds with an undo button. A task opens as a page with its recent runs, its prompt with the gaps marked and its schedule. The task form folds into sections, takes a schedule as a preset or a checked cron expression, warns about prompt gaps the schedule does not fill, and has a preview sheet. A job opens as a document: task, outcome, delivery and the recorded steps; a failed job says what happened and offers Run again. Run now now sends attached media through the job call, which is the only one that takes it. "Clear History" only ever cancelled running jobs, so it is now called Stop running jobs. Webhook headers of a saved task show as JSON instead of [object Object]. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Jobs page, task pages and job pages Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the Jobs page and the task form Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add helpers that say who uses a skill or a collection An agent loads a skill when skills are on and the skill is in its selection (an empty selection means every skill). It reads the one collection that carries its own name, when its knowledge base is on. The helpers derive that from the saved agent configs, build the config that adds or removes a skill or a collection, and estimate tokens as characters divided by four. Removing the last selected skill switches skills off, because an empty selection would mean every skill. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Skills and Memory as one library Skills and collections sit in a list with an open item beside it. Each row says who uses it, read from the saved agent configs, or says it is not used yet. Chat reads neither, so it is never named. An item opens in a pane with a Used by strip (names link to the agent, a small x removes it, with undo) and an Add to menu that shows what the addition costs. A collection can be added only to the agent that carries its name. The Memory pane searches the collection alone and shows ranked passages with scores, lists web sources with their refresh interval and the files, shows the server message when an upload fails, and names the endpoints and where files stay. The Simulate a message sheet runs a collection search and shows an agent's skills with a token estimate. It runs no model. The collection details route now opens the same page. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): show use and cost in the agent form pickers Each skill in the agent form says which other agents use it and what it adds to every message, with a total for the selection. The memory section names the collection the agent reads. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Skills and Memory libraries Specs for the used-by lines (including an agent that uses every skill), the filters, search, add to agent, remove with undo, the last-skill case, an unreadable agent list, the empty states, git repositories, the Memory question box, sources, uploads that fail, the Simulate sheet with the parts the API can run, the agent form hints and the phone layout. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the Skills and Memory libraries Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add helpers for the Operate status page and backend rows Pure functions for the parts that need rules. They work out the memory pool the page measures (GPU memory, system memory, or the workers of a cluster that are answering), which pools are too full, the headline and the four ledger rows, the geometry of the capacity chart, and what removing a backend would leave without a runtime (models name their backend, and a meta backend names the concrete one it points at). A second set says what a backend row states: installing, queued, removing, failed, update available, current or absent. LocalAI keeps no memory history, so the chart reads a bounded buffer of readings the page took itself and says so. A reading with no total is dropped rather than drawn as zero. Two hooks are shared by the pages that need them. One retries a failed operation after moving the failure into the record. The other holds a cancel for an undo window, because the server cannot take a cancel back. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Operate Status, This machine, Backends, Activity and Logs Status opens with one sentence ("2 things need you", or "Everything is running") and four rows: Needs you, Capacity, Running now and Recent failures. A row with a problem opens by itself and holds the button that deals with it: Update a backend, Retry or Dismiss a failed operation, Unload a model. A quiet row stays one line. A new installation gets a first-run screen, a cluster sums the memory of the workers that are answering, and a page still waiting for an answer says so. The chart under the rows is drawn from readings the page took while it was open and is labelled that way, because LocalAI keeps no memory history. This machine leads with GPU memory as one bar, then host memory split by running model, then VRAM, RAM, CPU and disk with a bar each. The running models become a kit table with the same menu and stop dialog. Backends is one list with Installed and Catalog views. A row says what the backend is doing (a progress bar with Cancel, Queued, Failed with Retry, Update 1.2.0, Current), carries the one button that matters, and opens in place. Removing a backend names the models and the meta backends that would stop working. Check for updates, Update all, From URL and a first-run recommendation for llama-cpp are in the header. Activity keeps its three sections as quiet rows. Cancel waits eight seconds with an undo toast, because the server cannot take a cancel back; a cancelled install can be started again from the record. Logs gets a process list, a picker, stream and text filters, Follow and Times switches, and a Clear with an undo window. Not shown, because the API has no data for them: GPU temperature and power, a size per backend, an earlier version to roll back to, a dependency lookup beyond the models and meta backends that name a backend, and models that failed to load. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover Operate Status, This machine, Backends, Activity and Logs New specs for the Status headline and ledger (healthy, needs attention, one thing, a full memory pool alone, loading, first run, cluster), its actions (Update, Retry, Dismiss, Unload with its dialog), the capacity chart built from readings taken while the page is open and bounded, the phone layout, no coloured edge on a row, and reduced motion. The Backends specs cover the two views, install progress with Cancel and its undo window, Retry on a failed install, Update, Update all, Check for updates, removal with the models and meta backends it would break, Install from URL, the first-run recommendation, a cluster, and a phone. Activity gains cancel with undo, Cancel now, a second cancel, leaving the page, progress, and starting a cancelled install again. Logs covers the stream and text filters, Follow, Times, Export, Clear with undo, the process picker and list. This machine covers the GPU strip, several GPUs, no GPU and Add a machine. Existing specs keep their intent and follow the new structure: rows open in place instead of in a pane, Update replaces Upgrade, the notice spec now pins that an update is a row state and not a banner or a rail, and a cancel waits for its undo window. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the rebuilt Operate Status, Backends and Activity pages Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add helpers for the Swarm pages Pure functions for what the pages work out from the cluster API: a node's state in words, which nodes a placement rule may use, what a rule would ask for, what a drain or a lost node would leave without service, the nodes a bulk backend update reaches, and the join commands for a worker, a peer instance and a memory shard. Hooks read the roster, the loaded replicas and the rules. Everything runs in the browser from data the page already holds, and says when it cannot see free memory or disk. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Swarm hub: nodes, node page, placement rules, failover Nodes is a sortable table with comfortable and compact rows, a Needs attention filter by reason, a map of the cluster that is not drawn on a phone, the running models, and a bulk backend update for the nodes that drifted. A node is a page: state, vitals, a drain preview computed from the loaded replicas and the rules, tabs for models, backends, logs and capacity and labels, and Remove that asks for the node's name. Placement rules are written as sentences, show where each model is loaded now, and edit in a side sheet with a preview of the nodes a draft could use. Deleting a rule waits a few seconds so it can be taken back. Failover keeps its chains, adds what the router does when a worker stops answering and a per-node preview of what would stop. Add a node covers a registered worker, a peer instance and a memory shard, with a command to copy and a live line that says when the machine arrived. P2P keeps its page in the same vocabulary, and the node logs page follows the local logs page. Previews are labelled as worked out in the browser. Per-GPU readings and node events are not drawn because the API does not return them. Failover moves to Swarm when distributed mode is on. Legacy fleet components and their styles are removed. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Swarm hub New specs for adding a node (each join method, the command, copy, waiting and found, approve, a single install, P2P, a phone), placement rules (sentences, where models are loaded, the preview matrix, the sheet and its preview, delete with undo) and failover on a cluster. Node detail covers its tabs, the drain preview and its dialog, resume, remove with the typed name, a node that stopped answering, and unload. The nodes specs follow the new structure and keep their intent: the table, filters, grouping, pagination, bulk actions, the map, and running models with stop, logs and the loading, error and empty states. The scheduling, failover, P2P, hub and smoke specs follow the renames. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the rebuilt Swarm pages Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add helpers for the Traffic pages Pure functions for what the pages work out from the usage ledger, the trace summary, the trace buffers and the resources reading: the shared time window, grouping, sorting and filtering of usage rows, chart series and axes that start at zero, the overview figures, per-model statistics, the state of a trace and the words for a failure, the backend operations that ran during a request, CSV export, the Prometheus metric list and scrape config, and a bounded buffer of host readings. A figure whose source cannot say is null, never zero. The trace summary call takes the window in hours, and a helper reads /metrics with its status. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Traffic hub: overview, usage, models, host, traces, middleware Traffic opens on an overview: five figures (requests, failed, p95, tokens in and out) and three charts, each naming its source. A second row of links reaches Usage, Models, GPU and host, Traces, Middleware and Prometheus, and one time window is shared by the first three. Usage groups by model, user or API key, filters, sorts, opens a row on its own chart, exports the rows it holds as CSV or JSON in the browser, and keeps the opt-in cost estimate and the quota forecast. A user who is not an admin sees only their own numbers. Models joins the ledger, the backend-operation buffer and the loaded models. GPU and host shows the current reading and two charts of readings taken since the page opened. Traces gets filters, a settings strip and an explained off state. An API request is a page: the error LocalAI recorded, a timeline with the backend operations that ran meanwhile, and bodies that stay closed until revealed. Middleware draws the pipeline as five steps and shows the rules of the selected step. Prometheus documents /metrics, checks it against the server and gives a scrape config to copy. Alerts is not built: LocalAI has no alert rules. Per-model latency percentiles, GPU utilisation and compare with the previous period are not drawn because the API does not return them. Legacy usage, trace and middleware styles and the usage source components are removed. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Traffic hub New specs for the overview (figures, charts with a data table and arrow key readout, failed and first-run and tracing-off states, the shared window, a phone), usage (group by, filters, sort, export, cost, quotas, a non-admin, empty and loading), models, GPU and host (snapshot, the since-opened labelling, a cluster), the traces list, a trace page (the real error, the timeline, reveal, no headers, a trace that left the buffer), Prometheus and the Middleware pipeline, with shared fixtures. The usage, traces, middleware, hub and smoke specs follow the new structure and keep their intent. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the Traffic hub Add an operations page for the Traffic tab: which record each page reads, what it leaves out and why, the trace page and its reveal, the GPU and host readings kept since the page opened, and the Prometheus endpoint. Link it from the operations index, the tracing page and the middleware page. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): let metric names wrap in the Prometheus table on a phone The long metric names pushed the type and "on this server" columns out of view. Names now wrap inside the table. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Settings with groups by intent, search, a pending bar and history The fifteen sections become eight groups by intent: memory and models, speed and defaults, backends and galleries, access and security, debugging and traces, agents and responses, swarm and sharing, look and feel. Search covers names, descriptions, keys and the old section name, and says where a result used to be. Edits wait in a bar with Discard, Show diff and Apply. The diff lists old and new values and the checks the browser can make: durations parse the way Go parses them, a GPU memory budget is one the server accepts, a gallery box holds JSON, and warnings repeat what the handler and the field text say. Apply sends only the changed keys. Undo saves the previous values again; it is a new save, not a rollback. History lists the changes applied from this browser, since LocalAI keeps no settings log, and Revert stages the old value. A value is marked as changed only where the built-in default is known from the CLI defaults. A row says "Applies now" or "Needs restart" only where the handler or the docs say so. Three things were wrong before and are fixed with the rebuild: the gallery boxes and the shared API keys box were sent under names the server ignores, the "Enable CSRF Protection" switch showed the disable flag the wrong way round, and every save restarted peer-to-peer networking because every field was sent. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Users and keys, Account, sign-in, invite and the 404 page Users and keys is a tabbed page under the Settings tab: people, invites and API keys. The people table filters by state and role, sorts, approves or disables (disabling offers an undo that sets the status back), and opens a side sheet for one person's features, model allow-list and limits. Role, password reset and delete sit in the row menu; delete asks for the name. Invites choose a lifetime of 1, 7 or 30 days and show the link once. API keys can be created with a lifetime, are shown once in full, can be paused, and are revoked after a ten second undo window in which nothing is sent. LocalAI lists keys only to their owner, so the tab shows the signed-in person's own keys and says so. Account has Profile, Security, API keys and Usage. Usage shows the last 30 days, tokens by model and the limits an admin set. The Security tab now shows for a GitHub or SSO account and says the password is not theirs to change. Sign-in asks for one field per step and draws a provider button only for a provider /api/auth/status lists. It has the notice for a sign-up that waits for approval, the first-admin screen, the key-only screen and the invite page. An address outside the app now gets the 404 page too, which names the address and lists the places the sidebar lists, with the same gates. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover Settings, Users and keys, Account, sign-in and the 404 page Settings: groups, search by name, key and old section, the changed marker only where a default is known, apply hints, the pending bar and diff, the checks, apply sending only changed keys, undo as a second save, discard, history, the CSRF inversion and the gallery and API key wire forms, and the phone layout. Users and keys: the table, filters, sort, approve, disable with undo, the row menu, the access sheet, invites, key creation with a one-time reveal, the ten second revoke with undo and with a page leave, and the non-admin redirect. Account, each sign-in variant (error, pending, first admin, key-only, invite, provider buttons) and the 404 page have specs too. Fixtures are shared with the screenshot scripts. Existing specs follow the new structure. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the rebuilt Settings, Users and keys, Account and sign-in pages Runtime settings: the eight groups and where each old section went, search, the pending bar, the diff and its checks, apply, undo, the history, and which settings show a default or an apply note and why. Authentication: the sign-in screen variants, the Account tabs, key lifetimes, the one-time key reveal, the revoke undo window, and the fact that keys are listed only to their owner. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): let the Settings undo toast stand alone and read back a generated P2P token The saved message and the undo toast sat on the same spot at the bottom of the page. The undo toast now carries the saved message. A new P2P token is made by the server when the page sends 0. The page reads it back after the save so the field shows the token and not the placeholder. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): fit the users table, API keys and Account figures on a phone On a phone the users table dropped its Role and Status columns off the screen edge with the row actions. The role and state now sit under the name, so the actions stay in view. API key rows no longer put the key icon on a line of its own, and the three Account figures keep one row. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the phone users table, reduced motion and the empty Account state Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): drop the apply note from three settings the save handler does not mention Size-aware eviction, automatic backend upgrades and development backends said Applies now, but nothing in the handler or the docs says when they take effect. A row now carries a note only where the code or the docs say so. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Voices and Faces as one identity family Voices is one page with three tabs: Speakers (voiceprints for recognising who is speaking), Speech voices (the text-to-speech reference library, kept apart because it is a different store) and From a recording (a link into the diarization workspace). Faces uses the same layout. Who is this and Same person? give the answer in a sentence with the real distance and cut-off, a word for how far inside the cut-off it sits, and a distance scale with the cut-off drawn on it. The cut-off slider re-reads the answer in the browser; the identify call sends the cut-off, and verify uses the threshold the model returns. The old confidence percentage is gone because it is not a probability. The server has no list call, so the people list stays in the browser and the page says so. After a search that asked for more people than it got back, a saved person the server did not return is marked, and people the server returned that the browser does not know are listed. Nothing is claimed from a short or cut-off search. Enrolling is a sheet: sample, name, labels, permission. A copy of the sample in the browser is opt-in, and an administrator can also keep the recording as a speech voice in the same step. Removing a person waits ten seconds behind an Undo toast and sends nothing before then. Errors say what happened (no face found, model missing, call failed), a blocked or missing microphone is explained, and a missing model or a missing permission renders a page that says what turns the feature on instead of a redirect. Analyze, detect and raw embedding move under More tools, with attribute guesses off by default. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Voices and Faces pages Specs for who is this (match, no match, working, failed, missing model), the cut-off slider, same person, a blocked, allowed and insecure microphone, the registry notes and the not-on-the-server marks, the enrol sheet and its opt-in copy, delete with undo on a fake clock, the disabled and no-permission states, the phone layout, reduced motion and Faces. Existing library and diarization specs follow the new structure and keep their intent. Node tests cover the distance words, scale layout, stored list and error mapping. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the Voices and Faces pages Add a WebUI section to the voice and face recognition pages: the two tools, the cut-off, what the people list is and why it can be stale, the undo window, and what is stored where. Point the Voice Library and Fish Audio notes at Build, Voices, Speech voices. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Build landing, Fine-tune, Quantize, Import and Explorer The Build landing says what each tool is for and what it needs from the machine: the installed backend, the GPU memory, RAM and disk the server reports, and a job that is running or the newest one when it failed. A tool that cannot run says why and what enables it. Fine-tune and Quantize share one page: set up, a check list that is redrawn as the form changes, a run view with progress, stages and a log, and a result with real next steps (export, import, chat, Models). The checks state only what the server reports. A job needs no estimate the server cannot make, so none is invented. Stop on a fine-tuning job asks whether to keep a checkpoint, a failed job shows the server's message, and a memory failure offers two changes that are applied to a copy of the setup. Import is a guided flow: source, review, import, done. The server returns no preview before an import starts, so the review reads the spelling of the source, prints the request the form will send and runs the checks that can be made early. The estimate that arrives when the import starts is set against free memory and disk. The ambiguity picker and the Write YAML tab stay. Explorer shows what GET /networks returns and lists a swarm with POST /network/add, with a join sheet that carries the token and commands. Build tools the account may not use say so instead of redirecting. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Build landing, the tool pages, Import and Explorer Specs for the landing (a tool ready, missing a backend, with no GPU, a running or failed job, a feature switched off, a member without admin, phone, reduced motion), the shared tool pattern for Fine-tune and Quantize (set up, live checks, start request, running with progress, chart and log, the stop choice, failure with the server message, finish with next steps, earlier jobs, the account-disabled page, phone), Import (source detection, review, checks, ambiguity, running with the estimate against free memory, done, Write YAML, phone) and Explorer (list, join, list a swarm, empty, not an explorer, retry, phone). Existing specs follow the new structure and keep their intent. Node tests cover the machine facts, tool status, checks, log lines, source detection, the import request and the join commands. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the Build tool pages, the import flow and the Explorer Fine-tuning and quantization now describe the set up, check, run and result steps and what the check list can and cannot say. The import section explains the review step and why the size and memory appear only after the import starts. The distributed page describes the Explorer list, the join sheet and what listing a swarm publishes. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): turn the hardware recommendations into a "Best for this machine" shelf The shelf in the Models inspector put five columns into a 400 px pane, so long model ids wrapped letter by letter underneath the size and the memory figures. Each row now stacks the tag, the id and the size and memory facts beside one Install button, and the id wraps inside its own column. Once a model is installed the shelf narrows to the best fit and keeps the others behind a "N more that fit" toggle. Specs cover the ranking, the layout, the narrowing and the install request against a gallery fixture that carries the 4K estimate the shelf sizes against. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): quiet the Studio tab markers and say their state in words The type tabs drew a saturated green dot for every modality that has a model. The dot now uses a text colour, filled when a model is installed and hollow when none is, and each tab carries "(model installed)" or "(no model installed)" as hidden text so the state is not only a colour. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): stack the model editor empty state and drop its section hues "No fields configured" sat in a flex row, so the icon, the title and the text ran together. It now uses the stacked empty-state layout. The section icons took a different status colour each (amber, red, green); they now share one quiet colour, with the accent on the current section. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): keep the Home memory sentence whole on a phone and drop side rails On a 390 px screen the memory strip clipped "2 models loaded" to make room for the figure. The sentence now takes the first line and the figure and device wrap under it. The sweep also removed coloured left rails from the editor section rail, the skill editor list, the install strip and the audio transform notice (now an outlined note), plus unused chat rules that carried rails and two glow animations that nothing referenced. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): line up hub pages, list the model templates, and stop clipped text Medium-width pages inside Build and Operate were centred while the tab bar above them was flush left, so the title started 60 px right of the first tab. They now start at the bar's edge. Add Model offered nine templates as a grid of identical cards with chip clouds and inline styles. It is now one list of rows, each with the field names it fills in on a single muted line. Two clipped strings are fixed: the Studio voice field cut its placeholder mid-word, and the phone job list ended the schedule line in an ellipsis. The recommendation shelf also separates size and memory with a dot, and the docs describe the shelf. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): keep the hidden Studio tab state inside its tab The hidden state text added to each type tab was absolutely positioned against the page, so on a phone it sat outside the scrolling tab row and widened the page by hundreds of pixels. The tab is now the containing block. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the on-disk sizes on the Installed table, model page and cleanup sheet The fixtures stub GET /api/models/storage. The default report is empty, so existing specs keep the gallery estimates. makeStorage() builds a report from files and the models that use them, the way the server does, and storageSpec() is a models directory with shared and missing files. New specs cover the Size column and its shared line, the fallback when the call fails or the user is not an admin, the files list on the model page, a missing file, the bytes a removal frees with shared files, and the cleanup findings. Node tests cover the storage helpers and the batch arithmetic. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): wait for the page before pressing keys and ticking the clock Two specs failed in loaded full runs and passed alone. The Alt+1 to Alt+7 spec pressed a key before the composer had armed its key handler. The capacity chart spec advanced the fake clock before the poller had mounted, so it counted fewer readings than it expected. Both now wait for the page to mount. The key spec retries a press that lands during a re-render, and the clock spec advances in small steps and polls for the row count. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): hide the Installed footer when the storage report is empty An empty report from the storage call made the footer read "0.0 GB on disk" next to sizes taken from the gallery estimate. An empty report says nothing about the disk, so the footer now shows only the model count. A spec covers it. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): keep Explore pane actions inside the pane The inspector actions sat in a non-wrapping flex row beside the title, so the buttons ran past the pane edge once it got narrow. The row now takes its own line and wraps. The primary action (Install, Retry, Open) comes first. Manage installation becomes a ghost button, and Open details moves to the end of the row, so one action stands out and the others are quiet. No action or test id is removed. Add a spec that checks, in light and dark at several widths and with a pane forced to 320 px, that every action stays inside the pane box and that the pane keeps its inner padding. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * refactor(ui): drop the type chips from the Studio composer The Studio tabs and the composer's type chips listed the same seven modes, so the page said the same thing twice. Keep the tabs as the one place to switch modes. The composer now shows the type it will open as a small label in its header. The type suggestion from the typed words stays as the quiet hint line under the prompt, and Alt+1 to Alt+7 still pick a type. The composer root carries data-type, data-types and data-missing so tests can read the state. Specs pick a type through a shared Alt+digit helper and read a missing model from the tab dot instead of a chip. Remove the unused chip locale strings and CSS. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * refactor(ui): remove the section crumb above page titles Page headers drew a small uppercase crumb with a short rule before it above the title. On the hub pages it repeated the hub name, so Build sat above a heading that also said Build. PageHeader now renders only the title, the supporting line and the actions. Drop the eyebrow prop, the route-derived section name, its CSS and the unused section helper, and remove the explicit eyebrow props from the pages that passed one. Pages stay reachable through the sidebar and the hub tab bar. Add a spec that checks several pages show their title with nothing ahead of it in the header. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * refactor(ui): remove left accent rails from tiles, rows and quotes Several surfaces marked state with a coloured strip on the left edge. Replace each one with a cue that is not a rail: - Stat cards lose the strip; the icon and value still carry the colour. - The highlighted card is a raised surface with a firmer edge. - The selected rail row is an accent wash with a hairline outline. - The status stripe on rail items is a small status dot. - The active failover row is a tinted row. - Quotes in markdown and chat prose are italic instead of barred. - The variant detail panel has a full hairline border. Add a spec that walks the main routes in light and dark and fails on a left border thicker than 1px, a sideways inset shadow, a narrow absolute strip in ::before or ::after, or a narrow tall child pinned to a left edge. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com> |
||
|
|
690f3994b0 |
feat(auth): let users pause and resume their API keys (#12521)
A key can now be paused until the owner resumes it, or until a given
time. A paused key is rejected by validation before last_used is
updated, and a pause time that has passed lifts the pause by itself.
Existing keys stay active.
PATCH /api/auth/api-keys/:id takes {"disabled": bool, "paused_until":
RFC 3339 string or null}. Only the key owner can change it, and a
paused_until in the past is rejected. The key list returns the pause
fields. The Account page gets a Pause and Resume button for each key
and a Paused badge that shows the resume time.
Assisted-by: Claude Code:claude-sonnet-5-5
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
|
||
|
|
99043b442c |
feat(router): route with native decision models (#12449)
* fix(schema): preserve SystemOne image inputs Assisted-by: OpenAI * test(schema): follow Ginkgo conventions for decision inputs Assisted-by: OpenAI * feat(llama-cpp): dispatch native decisions through Score Upgrade the stock dependency and reconcile Score/TTS patches. Reuse native decision parsing, tasks, formatting and response-reader cleanup; preserve ordinary scoring admission and guard older dependencies. Assisted-by: OpenAI * refactor(systemone): share request and model validation Assisted-by: OpenAI:gpt-5 * fix(systemone): preserve HTTP wire-byte validation limit Keep structural validation separate from the serialized internal request bound so HTML escaping cannot reject valid HTTP payloads. Assisted-by: OpenAI:gpt-5 * feat(systemone): bound images and account native decisions Preserve public wire limits independently from router serialization. Reject unsupported NER images, map native request/capability errors, and stamp explicit usage once. Advertise decisions for stock llama-cpp. Assisted-by: OpenAI * fix(systemone): record usage on registered native route Exercise real registration and billing with a mock native backend. Reject empty native responses, malformed image URLs, trailing JSON, and wire overflow including whitespace. Assisted-by: OpenAI * feat(router): add lazy native decision transport Bind named models through internal ModelSystemOne calls with shared validation and bounded abandoned operations. Remove request and echoed-error contents from decision traces. Assisted-by: OpenAI:gpt-5 * feat(router): classify overlapping policies with native decisions Ask independent noul questions, validate probabilities and preserve first-superset routing. Wire the central factory with config-sensitive invalidation and cancellation-safe resolution. Document native framing and bounded operation limits. Assisted-by: OpenAI:gpt-5 * feat(gallery): add pinned Julia-1 native decision model Add a separate text-only llama-cpp Q8 entry with pinned Apache-2.0 source provenance and checksum. Installed using the gallery installer and exercised choice, score and noul on CPU. Assisted-by: OpenAI * test(router): verify native decisions through central factory Add an opt-in real-model Ginkgo integration covering the native Go loader and C++ transport, token usage, independent overlapping labels, and candidate selection. Document owned-server execution and the intentionally non-quality threshold. Assisted-by: Codex:gpt-5 * fix(llama-cpp): align upstream pin and preserve decision signatures Advance to bed0a856 without losing the automated upstream bump. Detect full-request fill_task support at compile time and forward every question for Nimble framing while retaining the earlier native signature. Preserve reconciled SCORE/TTS patches; add standalone compatibility coverage. Assisted-by: Codex:gpt-5 * feat(gallery): add native decision family defaults Pin Laya, Kev-4B, lev, OpenJev and Nimble artifacts. Verify Laya/Kev/lev gallery installs and CPU contracts on both native pins; clearly mark OpenJev/Nimble runtime validation pending and their noncommercial licenses. Assisted-by: OpenAI * docs(decisions): clarify integrated Nimble prerequisite Record the exact combined backend pin while retaining pending OpenJev and Nimble installation/runtime validation status. Assisted-by: Codex:gpt-5 * fix(gallery): indent native decision model sequences Match repository yamllint indentation for Laya, Kev, lev and OpenJev list fields. Parsed gallery data is unchanged; reproduce CI gallery lint failure before the whitespace-only fix and pass the same command afterward. Assisted-by: Codex:gpt-5 * docs(decisions): record OpenJev and Nimble CPU validation Record gallery installation, checksum/metadata verification and multiquestion native smoke results on bed0a856. Retain noncommercial and text-only limitations without accuracy or deterministic-output claims. Assisted-by: OpenAI * fix(ui): expose native Decisions router classifiers Select classifier models using metadata-driven capability routing, retain tuned thresholds, and validate native decision selections before saving. Cover both native backends and create/save/reopen in the real React editor. Assisted-by: Codex:gpt-5 * fix(router): exclude aliases from native decision discovery Check the originally named config before advertising native Decisions eligibility. Retain target capability inheritance for ordinary generation aliases. Exercise the actual capabilities endpoint with native models on both backends, aliases, and disabled models. Assisted-by: Codex:gpt-5 * feat(systemone): share bounded multimodal input validation Preserve text wire limits while admitting bounded PNG/JPEG decision input. Share collection and header validation across internal and public callers and keep the native runner response budget independent. Assisted-by: OpenAI:API-assistant * fix(systemone): bound admission lifetimes and validate complete images Retain shared admission leases through actual work completion, including abandoned internal operations. Decode bounded image pixels, cap public native responses before usage stamping, and preserve oversized malformed text status precedence. Assisted-by: OpenAI:API-assistant * fix(router): classify images before media fetching Preserve ordered structured probes for native decisions. Defer OpenAI media preparation until routing selects the served model, so rejected decision URLs cannot trigger downloads before shared validation. Guard direct image collection with context-aware shared admission. Keep text classifiers and embedding caches from discarding image input. Retain fail-closed classifier configuration and runtime fallback policy. Add middleware, typed-content, admission, cancellation and cache tests. Assisted-by: OpenAI:API-assistant * fix(router): bound extraction before serialization Check probe budgets before copying text or marshaling message state. Count JSON escaping so oversized internal inputs fail before allocation. Preserve typed Anthropic blocks through selected-model conversion and fallback. Keep retry coverage in Ginkgo without global test registration. Assisted-by: OpenAI * fix(router): bound supported probe serialization Arbitrary structs can bypass the probe budget through pointer marshalers, string tags, and promoted fields. Accept concrete chat schema types and plain JSON values instead of emulating arbitrary struct serialization. Budget escaped direct prompts before marshaling so raw length cannot hide serialized expansion. Preserve runtime fallback and reject oversized input before invoking the decision runner. Add Ginkgo allocation, boundary, and marshaler invocation regressions. Six-package tests, three-package race tests, and full-T2 delta lint pass. Assisted-by: OpenAI:GPT-5 golangci-lint * feat(decisions): enable bounded OpenJev images Validate native decision images before permissive media parsing and pixel allocation. Require both decision image support and a vision projector; missing or audio-only projectors cannot silently become text decisions. Pin the OpenJev Q8 projector and document its license and disk footprint. Add native safety tests, canonical limit parity, gallery and load-option checks, and a reproducible CPU direct-RPC contrasting-image smoke. Assisted-by: OpenAI:GPT-5 * fix(decisions): reject incomplete image streams stb accepts corrupt PNG Adler checksums and truncated JPEG scans. Use bounded zlib validation and strict libjpeg decoding before parsing. Keep dimension and aggregate pixel checks ahead of decoder allocations. Wire decoder dependencies into native builds and runtime packaging. Add regressions for appended EOI and embedded marker bypasses. Assisted-by: OpenAI:GPT-5 * fix(ci): gate native decision image validation Run the decoder security tests outside the stdlib-only native suite. Fetch vendor headers at the backend pin and provision decoder dependencies. Gate Go limit parity and production CMake wiring without model downloads. Assisted-by: OpenAI:GPT-5 * test(decisions): cover multimodal public API paths Exercise shared image contracts through the registered HTTP routes and external mock backend. Add opt-in cached gallery installation and real OpenJev image decisions through SystemOne and both routing APIs. Assisted-by: Codex:gpt-5 * test(decisions): assert isolation and cache bypass Observe external RPC calls and compare complete classifier history. Winner-only and cache-miss checks could hide dropped history or cache use. Give real inference its own application and model directory so shared backend mappings and loaded processes cannot affect mixed suite order. Assisted-by: OpenAI:ChatGPT * test(decisions): isolate fixture globals Disable optional global services in the isolated HTTP fixture and register cleanup before setup assertions. Verify meter provider identity survives fixture creation and destruction. Snapshot observed usage before assertions so failures cannot retain the mutex. Require a successful usage stamp before checking error responses. Assisted-by: Codex:gpt-5 golangci-lint * fix(application): honor optional telemetry controls Skip failover gauge registration when metrics are disabled. Register against the application meter rather than looking up the global provider. Allow embedders to retain the bounded routing log without billing stats. Keep the existing default when stats are disabled. The isolated HTTP fixture uses this option without losing its native router assertions. Assisted-by: Codex:gpt-5 golangci-lint --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
9eb5a9e61d |
feat(audio): remember speakers from diarization (#12414)
* feat(schema): validate portable speaker profiles Add the versioned profile schema for explicit speaker enrollment. Validate compatibility against separately supplied loaded-encoder metadata. Reject unusable speakers, invalid vectors, and inconsistent clean spans. This slice does not change HTTP routes, backend integration, or the UI. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(parakeet): export profiles with transcripts Export opt-in speaker profiles and trusted encoder metadata. Replay registrations by ID so duplicate display names keep independent vectors. Use one profile-capable diarization for slots, names, and clean spans. Assign timestamped ASR words to those slots without a second diarization. Preserve legacy opt-out and no-ASR behavior, and propagate failures. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(audio): enroll portable speaker profiles Gate profile exports with voice-recognition permission and validate registration against metadata from the loaded encoder. Preserve audio enrollment and independent registrations with duplicate display names. Exclude diarization and registration exchanges before API trace capture so persisted traces cannot retain profile vectors or JSON audio. Defer candidate dimensions to trusted loaded metadata. Sort candidates by registration ID so incompatible profiles cannot suppress legacy voices through registry iteration order. Keep portable identity checks closed when trusted metadata is unavailable. Test persisted traces, explicit slot zero, and selection through offline and live transport. Document privacy and the ephemeral registry lifecycle. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(ui): remember speakers from diarization Add a Studio page for diarization and opt-in speaker profiles. Preview clean intervals from the original recording before explicit registration. Join profiles by raw speaker labels, preserve duplicate names, and relabel turns only after a successful save. Discard stale results when the model or recording changes. Share registration metadata with voice management without storing vectors or recordings from this flow. Document permissions and the global, ephemeral registry. Cover enrollment, permissions, previews, and asynchronous races with mocked Playwright tests. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: clarify HTTP speaker enrollment support Replace the stale enrollment limitation with the current HTTP workflow. Distinguish native transport from explicit registration and link its docs. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * chore(parakeet): pin merged speaker profile support Use the merged commit from mudler/parakeet.cpp#80. Its tree matches the previously accepted native pin. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: add diarization enrollment setup example Connect the existing gallery modes to the speaker enrollment workflow. Show installation, private profile export, explicit raw-slot registration, and later recognition without another export. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs(blog): explain diarization speaker profiles Put the diarization walkthrough on the LocalAI website in the feature PR. Cover the three gallery modes, explicit enrollment, and privacy limits. Link setup instructions and keep availability conditional on feature support. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs(blog): focus diarization on everyday use Explain what users can do with recordings before the setup steps. Replace the technical walkthrough with a short Studio guide and link readers to the existing reference for model names and developer use. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs(blog): lead with speaker capabilities Present speaker recognition through everyday uses and a short UI flow. Keep technical reference details in the existing documentation. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(diarization): satisfy Go lint checks Avoid copying protobuf message state when extending backend status, check the multipart reader close result, and document the focused testing.T lint exemptions. Assisted-by: nib:gpt-5.6-sol Signed-off-by: Ettore Di Giacinto <mudler@localai.io> --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
f378fe89d0 |
Merge pull request #12373 from mudler/feat/systemone-capability
feat: decisions usecase for decision models, with gallery tagging |
||
|
|
70ce62901f |
refactor: name the capability decisions instead of systemone
The usecase describes what a model can do, and the category is the Decisions API. SystemOne stays as the wire contract: the /v1/systemone routes, the Score RPC question_type and the swagger tag are unchanged. The usecase, flag, auth feature, UI label, gallery tags and docs page are now decisions. Assisted-by: Claude Code:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
bd6863af81 |
fix(react-ui): extract text from PDF attachments in chat and home (#12374)
The React UI read every non-media attachment with file.text(). For a PDF that decodes the binary bytes as UTF-8, so the model received raw "%PDF ... stream ... endobj" noise instead of the document. The legacy Alpine UI ran pdf.js; that step was not ported when the React UI replaced it, but both file pickers still advertise .pdf. Add a shared readAttachmentText helper that routes PDFs through pdfjs-dist and reads other files as before. pdf.js and its worker load on first use, so the main bundle does not grow. A PDF that cannot be parsed or has no text layer (scanned, encrypted, damaged) is rejected with a toast instead of being attached as an empty or garbage file. Cover the chat and home paths with Playwright specs that build a real PDF in the test. Assisted-by: Claude Code:claude-sonnet-5-5 [playwright] [eslint] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
36846466e4 |
feat(ui): show the systemone usecase on installed models
Assisted-by: Claude Code:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
7b11af7f7d |
feat(ui): list failover chains and badge chain models
Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
bdd4a73d05 |
feat(ui): edit failover chain targets with a dedicated field
Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
d65b4d3e6a |
feat(ui): show live failover chain health in the model editor
Editing a failover chain now shows its state, the target serving it, and a per-target health table under the editor header. The strip reads GET /api/failover and follows /api/failover/events, with a 15 s re-list to cover SSE reconnect gaps. Admins can pin a target or unpin the chain after a confirmation. Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
c0b7e64973 |
feat(ui): show running models and host gauges on single-node installs (#12189)
On a single-node install nothing in Operate listed the models loaded on this machine or let an admin stop one. The System page that did was retired in #11548, and its replacements (the Nodes workbench) only work in distributed mode. The Nodes page also mis-detected single-node mode: the cluster routes are not registered there, so /api/nodes answers 404, but only 503 was treated as "distributed off", which sent every single-node install to the empty worker-registration card. The rail hid the entry anyway. Nodes route on a single node becomes "This machine": - the Nodes page's VRAM / RAM / CPU / models-disk gauges, fed from this host by mapping /api/resources onto the worker heartbeat fields - a memory bar splitting host RAM by running model - a running-models table (backend, RSS, CPU share, uptime, PID) with search, sorting, logs and a confirmed Stop - the distributed setup behind an "Add machines" button The Operate overview gains a "Running now" preview (heaviest five, with Stop) on single node and a pointer to Nodes > Running models on a cluster. The rail shows "This machine" in Runtime with a running count. Backend, additive only: - /system: each loaded model carries a `process` block (pid, rss_bytes, memory_percent, cpu_percent, started_at). A sampler keeps one gopsutil handle per PID so CPU is the share since the previous poll rather than the lifetime average; it is omitted on the first reading. - /api/resources: host `cpu` and models-path `disk`, the same readings workers send in their heartbeat. Also fixes the fleet tables widening the page on phones: the headers' absolutely positioned sr-only labels escaped the scroll wrapper. Assisted-by: Claude:claude-opus-5 [Playwright] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
2facfc0d88 |
feat: Add kimodo.cpp and 3D animation API/UI (#12095)
* fix(vulkan): preserve host ICD discovery for packaged backends Add bundled Mesa manifests through VK_ADD_DRIVER_FILES instead of replacing the system driver list. Merge inherited and model-specific additive paths while preserving explicit operator overrides, with regression coverage. Assisted-by: Codex:gpt-5 golangci-lint Assisted-by: Codex:gpt-5.6-sol Signed-off-by: Richard Palethorpe <io@richiejp.com> * feat(3d): add Kimodo CPU and Vulkan animation backend Introduce a distinct animation capability and model-described 3D operations, with a typed /3d/animate API, RPC transport, distributed media staging, permissions, and tracing. Add a persistent kimodo.cpp adapter, skeleton GLB export, CPU/Vulkan packages, model and backend galleries, importer support, CI builds, and documentation. Adapt Studio inputs to each model and provide real-time skeleton playback, seeking, and history. Cover backend validation, packaging, API behavior, importer inventories, distributed staging, and Studio workflows. Validate real-model CPU/Vulkan generation and deploy the integration to the local QA instance. Assisted-by: Codex:gpt-5 golangci-lint Assisted-by: Codex:gpt-5.6-sol Signed-off-by: Richard Palethorpe <io@richiejp.com> * feat(kimodocpp): adopt monolithic encoders and resident inference Update upstream for resident weights, packed execution paths, and cached motion graphs. Default to all 32 text layers while retaining configurable streaming and legacy bundle support. Use monolithic Q8_0 encoders by default and offer all six published quantizations through the gallery and importer. Refresh pinned hashes, tests, and documentation; remove the obsolete thread patch and ensure cached source checkouts follow the upstream pin. Validated CPU and Vulkan generation, lower-bit streaming, gallery/importer suites, packaging, lint, and cold/warm Studio generation on localai-dev. Assisted-by: Codex:gpt-5 golangci-lint Assisted-by: Codex:gpt-5.6-sol Signed-off-by: Richard Palethorpe <io@richiejp.com> --------- Signed-off-by: Richard Palethorpe <io@richiejp.com> |
||
|
|
15de88d361 |
fix(ui): test backend actions through their menu (#12069)
The node detail redesign moved backend operations into an action menu. Four existing specs still search for the removed direct buttons, so the UI E2E workflow fails consistently on master. Open the backend action menu before checking or activating its items. Assisted-by: Codex:gpt-5 Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
d5256a5584 |
fix(ui): restore node operation controls (#12068)
The node restructure hid backend logs and split related controls across inconsistent layouts. Restore contextual log actions and align the detail page with the fleet dashboard. Make multi-node selection clear and accessible. Assisted-by: Codex:gpt-5 Playwright ESLint Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
997d403de4 |
feat(nodes): add fleet operations dashboard (#12046)
* feat(nodes): report CPU telemetry Assisted-by: Codex:gpt-6 * feat(nodes): add fleet view utilities Assisted-by: Codex:gpt-6 * feat(nodes): add fleet operations dashboard Replace the panel roster with aggregate capacity gauges, fleet filtering and selection, bounded bulk actions, and an on-demand node inspector. Extend node details and distributed-mode documentation with CPU and models-disk telemetry. Assisted-by: Codex:gpt-6 * fix(nodes): harden fleet lifecycle actions Assisted-by: Codex:gpt-6 * fix(nodes): restore compact fleet composition Keep fleet health, capacity, and attention in one compact overview at ordinary desktop widths. The inspector now overlays the roster until the workbench can preserve a useful table beside it. Assisted-by: Codex:gpt-6 * feat(nodes): add accessible running models workbench Assisted-by: Codex:gpt-6 * fix(nodes): correct model view ARIA links Keep each tab panel available for its controlling tab while native hidden state removes inactive content from accessibility navigation. Model controls now expose only supported state and valid inspector relationships. Assisted-by: Codex:gpt-6 * fix(nodes): align lifecycle and capacity states Pending nodes now expose approval wherever node actions appear, while other lifecycle controls follow the server transition rules. Capacity totals exclude incomplete readings so missing availability remains unknown. Assisted-by: Codex:gpt-6 * fix(nodes): restore approved dashboard composition Assisted-by: Codex:gpt-6 * fix(nodes): integrate operate navigation Assisted-by: Codex:gpt-6 * fix(nodes): restore low density fleet view Assisted-by: Codex:gpt-6 * fix(nodes): preserve complete operate menu Assisted-by: Codex:gpt-6 * fix(nodes): preserve inspector workspace height Assisted-by: Codex:gpt-6 * fix(nodes): restore standard operate navigation Assisted-by: Codex:gpt-6 * feat(ui): add collapsible console rail Assisted-by: Codex:gpt-6 * feat(nodes): stop models from fleet view Assisted-by: Codex:gpt-6 * fix(nodes): make inspector a full height drawer Assisted-by: Codex:gpt-6 * fix(model): stop mixed local and remote placements Assisted-by: Codex:gpt-6 * fix(ui): announce action menu navigation Assisted-by: Codex:gpt-6 * fix(nodes): keep inspector within viewport Assisted-by: Codex:gpt-6 * fix(ui): preserve focus across model actions Assisted-by: Codex:gpt-6 --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
6983477a71 |
fix(agent-ui): keep chat open for status
Agent Status replaced the chat route and unmounted its EventSource. Any response still in flight could then disappear from the conversation.\n\nOpen status in a separate tab so the chat keeps its live connection until the response completes.\n\nAssisted-by: Codex:gpt-5 [eslint] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
db09452d54 |
feat(tts): support multi-reference personalities
Saved profiles previously resolved to one audio path and transcript, so cloning backends could not use several examples of one personality. Store ordered audio and transcript pairs while preserving the legacy first-reference fields. Fish Speech and audio.cpp receive all pairs, including on distributed workers. Other backends retain their single-reference behavior. Assisted-by: Codex:gpt-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
ddda0436a3 |
fix(ui): exclude navigation from the TTS mock
The API mock also matched navigation to /app/tts and returned a WAV download instead of the React page. Let non-POST requests reach the test server. Assisted-by: Codex:gpt-5 [Playwright] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
1bcbff586b |
fix(ui): target the TTS input in history tests
TTS instructions add a second textarea to the page. Target the speech input by its placeholder so the history test does not depend on the page having one textarea. Assisted-by: Codex:gpt-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
919b5c96fa |
feat(ui): add per-request TTS instructions
Let studio users guide speech delivery for backends that support request instructions. Blank guidance stays out of requests and media history. Assisted-by: Codex:gpt-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
e588d463dc |
fix(ui): preserve percent signs in route parameters (#11883)
fix(ui): preserve decoded route parameters React Router already decodes dynamic path segments before exposing them through useParams. Decoding those values again crashes pages for names containing a literal percent sign and mutates escape-like substrings. Use route parameters as-is, encode the model editor API path at the outbound boundary, and cover all affected pages with Playwright. Fixes #11882 Assisted-by: Codex:gpt-5 eslint playwright Signed-off-by: QiuLG <l237455523@outlook.com> |
||
|
|
de563f17b5 |
fix(ui): omit GPU recommendations that do not fit (#11945)
When no sampled candidate fits GPU memory, ranking falls back to the oversized pool and labels its first model Best fit. Keep GPU picks within the existing 95% budget and hide the section when no candidate qualifies. Remove static GPU starter picks so Home cannot reintroduce the same error. Add browser regressions for both sections and document the empty result. CPU fallback behavior stays unchanged. Assisted-by: Codex:gpt-6 [Codex] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
18d20239df |
fix(ui): keep trace expansion stable during refresh
Squashed merge of #11278. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
718357219b |
fix(ui): send collection intervals as numbers
Squashed merge of #11819. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
80e3240f2d |
feat(distributed): key scheduling rules by a model alias (#11771)
Node placement and replica rules could only name a model, so an operator who pinned "llama3" to the GPU tier had to rewrite the rule whenever a different model took over that job. An alias already gives a stable name for whichever model serves it, and a rule on that name makes it a deployment slot: repoint the alias and the placement follows. A rule keeps the name the operator chose. Reads resolve that name through the config loader to the model the rule governs, so the reconciler counts, schedules and trims replicas of the target, and the router finds an alias-keyed rule from the target it is already routing. An alias that resolves to nothing governs nothing loadable, so the reconciler skips it and the write paths refuse it. A replica is shared by every name that resolves to it, so only one rule can decide where it runs. The REST and MCP write paths reject a rule whose target another rule already governs. A pair that arrives some other way, such as a seed file or an alias repointed onto a model that already has a rule, resolves in favour of the rule named after the model itself and then the oldest, and the rest are listed as shadowed. The eviction guard is the exception: it matches rules to replicas in raw SQL inside a locking transaction and cannot resolve an alias. It reads a stored target that the reconciler refreshes each tick, and falls back to the rule's own name when that target is empty. Assisted-by: Claude:claude-opus-5 golangci-lint eslint Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
29899cd1e0 |
fix(ui): size model fit against the cluster and move node labels into the selector (#11765)
* fix(ui): move node labels into the scheduling selector field The scheduling page kept a node-label browser open above the rules whether or not anyone was writing one, while the field that actually needs labels, the rule's node selector, was two bare text inputs with no hint of what the cluster reports. The browser is gone. The selector's key input now completes against the label keys the cluster uses, and the value input offers only the values that key takes. The roster already loads for the page, so the suggestions cost no request, and a roster that fails to load costs the admin the hints and nothing else. Suggestions stay suggestions: a key no node reports yet still commits as typed, which is how an admin writes a rule before labelling the nodes for it. Assisted-by: Claude:claude-opus-5 golangci-lint eslint playwright Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(distributed): size model fit against the cluster, not the frontend The models page asked the frontend how much memory a model may occupy. In distributed mode the frontend is usually a GPU-less pod while every model runs on a worker, so a fleet of GPU nodes was told it could only run the smallest CPU build. The variant picker's fits flag and its auto-selection came from the same place, as did the hardware recommendations. The registry now reports the largest single healthy backend node. The largest node, not the fleet total: a model loads into one node, so four 16GB workers are not a home for a 40GB model. An operator-set VRAM budget caps a node's contribution, because the scheduler refuses a load above that ceiling anyway, and a GPU node beats a CPU node holding more system RAM. GET /api/resources and GET /api/models carry this as an additional cluster object. Their aggregate and ram fields keep reporting the frontend's own hardware, which is what the resource monitor shows. Variant selection judges backends against the union of the capabilities present in the cluster, the way backend discovery already did. Every path degrades to the local host: no cluster object in single-node mode, and none when the registry cannot be read, so a hiccup narrows the answer back to single-node behaviour rather than marking the whole catalog too large. The verdicts now name the node they belong to, since a model fits somewhere or nowhere. Assisted-by: Claude:claude-opus-5 golangci-lint eslint playwright Signed-off-by: Ettore Di Giacinto <mudler@localai.io> --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
8f56e4e042 |
fix(vram): persist remote probe metadata (#11487)
* fix(vram): persist remote probe metadata The startup warmer repeated remote size and GGUF metadata probes after every restart because both caches lived only in memory. Store successful HTTP probes for 24 hours so frequent restarts reuse the prior results. Bound the cache, reject invalid records, and purge it when gallery data changes. Local model files continue to bypass persistence. Assisted-by: Codex:gpt-5 * fix(vram): check temporary file cleanup The lint gate rejects the unchecked cleanup call in the persistent cache writer. Assisted-by: Codex:gpt-5.6 [golangci-lint] * fix(vram): make persistent cache optional Remote metadata probes can transfer enough data that operators need control over disk reuse and startup warming. Gallery autoload now gates both behaviors, and the runtime setting applies changes immediately. Assisted-by: Codex:gpt-5 * fix(ui): expose gallery startup pre-warm The existing gallery autoload setting also gates the startup metadata warmer. Name both effects in Settings so operators can find the requested boot control. Assisted-by: Codex:gpt-5 --------- Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
9d92139de4 |
feat(ui): edit scheduling rules in place (#11667)
* docs(ui): design scheduling rule editing Document the approved in-place rule editing flow and scalable node-label reference for the scheduling view. Assisted-by: Codex:gpt-5 * feat(ui): improve scheduling rule management Add scalable node-label discovery and editable scheduling rules with responsive, accessible controls. Assisted-by: Codex:gpt-5 * chore(ui): ratchet inline style baseline Record the static inline style removed by the scheduling view enhancement. Assisted-by: Codex:gpt-5 --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
4cad809003 |
fix(ci): test stale chunks in split bundle (#11595)
The V8 coverage build inlines every dynamic import, so the stale chunk tests cannot intercept a page chunk. Run those tests against the normal code-split bundle and exclude them from the inlined coverage pass. Assisted-by: Codex:gpt-5 Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
b806b1fec3 |
fix(ui): reload once when a page chunk 404s
A deploy replaces the whole content-hashed asset set at once. A tab holding an older index.html, or one whose request lands on a replica that the rollout has not swapped yet, asks for a page chunk the server no longer has. The dynamic import rejects and React Router's default error boundary replaces the app with "Unexpected Application Error!" until someone reloads by hand. The router now reloads the page itself when a chunk fails to load. index.html is served no-cache, so the reload lands on a self-consistent asset set. A timestamp in sessionStorage bounds this to one reload per 10 seconds, so a chunk that is genuinely gone reaches the error boundary instead of looping forever. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] |
||
|
|
d10374f849 |
feat(router): make KNN a first-class classifier with a persisted, curated corpus (#10652)
* feat(router): make KNN a first-class classifier with a persisted, curated corpus
Add `classifier: knn` — similarity-weighted voting over labelled
example prompts. Unlike score/colbert it needs no classifier model:
label knowledge lives in a corpus seeded and curated through the
admin API, so routing decisions are deterministic, auditable, and
grounded in graded experience rather than a model's opinion.
Epistemic gate: corpus entries below knn.similarity_threshold cannot
vote; when none clears it the classifier activates no labels and the
router uses the fallback — a prompt unlike all labelled experience is
treated as undecidable, not guessed. Decisions record
nearest_similarity (also on fallback rows) so admins can see how far
the nearest labelled experience was; the Routing tab explains
out-of-corpus fallbacks and shows per-label corpus counts.
Persistence: one JSONL file per router under
<data path>/router-corpus (text, labels, vector, embedder
fingerprint). The file is the source of truth; the local-store index
is rebuilt from it at classifier build time and stays a pure
in-memory index. Entries recorded under a different embedding model
re-embed on load. Also corrects the docs' false claim that
local-store collections persist — the embedding cache never survived
restarts (and still doesn't); the corpus does.
Corpus input is API-only by design (entries may contain example user
content): POST /api/router/{name}/corpus seeds (labels validated
against declared policies, embedded server-side, indexed
immediately), GET .../corpus/stats inspects — label counts only,
entry texts are never returned by any surface — DELETE .../corpus
wipes. Admin-gated like the sibling router endpoints, and exposed as
MCP tools (seed_router_corpus / get_router_corpus_stats /
clear_router_corpus) in both the httpapi and inproc clients with
coverage-test route mappings.
Plumbing: VectorStore gains SearchK (top-K was hardcoded to 1);
local-store gets InsertBatch/Delete as optional fast paths;
RouterConfig gains a knn block (embedding_model, k,
similarity_threshold, vote_threshold, store_name) with meta-registry
fields; the classifier dropdown now offers knn and the
previously-missing colbert; embedding_cache is ignored (with a
warning) for knn — it IS an embedding-KNN lookup; the stale
/api/instructions intelligent-routing entry is rewritten (it
described a classifier that no longer exists); swagger regenerated.
Tests: KNN vote/gate specs with hand-computed vote shares, corpus
manager suite (restart reload without re-embedding, fingerprint
re-embed, dedupe, hostile store names), middleware specs (corpus
routing, gate fallback, config validation, cache-wrap refusal),
corpus endpoint specs pinning the texts-never-returned contract, MCP
catalog + route-mapping gates, and a Playwright spec for corpus
stats and the out-of-corpus decision detail.
Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Richard Palethorpe <io@richiejp.com>
* feat(router): name consulted corpus neighbours in knn decisions
Every knn decision (decision log rows and the /api/router/decide
response) now carries neighbors: the K retrieved corpus entries by
descending similarity - including ones below the epistemic gate, which
is what makes fallback decisions diagnosable - each as {id, similarity,
labels}. The id is the entry's content hash (first 8 bytes of the
SHA-256 of its text, hex): stable across reseeds and re-embeds, and
text-free, so an external platform that seeded the corpus can recompute
text->id on its own copy and bucket decisions by corpus region (per-
region reliability accounting) without corpus text ever leaving the
server. A corrupt index payload surfaces as an id-less neighbour at a
real similarity instead of disappearing.
Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Richard Palethorpe <io@richiejp.com>
* refactor(router): deduplicate knn plumbing and cut corpus hot-path waste
Post-review cleanup of the knn-first-class-router branch; no behaviour
changes on the API surface.
Reuse/altitude:
- RouterKNNConfig.ResolvedStoreName is now the single source of the
router-corpus-<name> default (was hand-derived in four files).
- corpus.ResolveKNNRouter + corpus.Seed carry the shared model
resolution and seed validation; the REST endpoints and the assistant
MCP client are thin transport adapters over them, with sentinel
errors mapped to HTTP statuses at the echo boundary.
- middleware.NewClassifierDeps assembles the classifier dependency set
once for all five entry points (OpenAI, Anthropic, realtime, decide,
corpus) instead of five hand-copied literals.
- router.AllClassifiers feeds both the status endpoint and the
unknown-classifier error, ending the classifier-list drift.
- Per-classifier requirements moved out of validateRouterPolicies into
their buildClassifier arms; the knn arm owns its embedding_cache
opt-out instead of a name-check in the shared wrap tail.
- adminOnly replaces four inline copies of the admin gate in the
middleware routes.
- localVectorStore.Search delegates to SearchK (identical traces).
Efficiency:
- Manager.Add embeds outside the manager mutex and appends to the
JSONL file (O(new) instead of O(corpus) rewrite); a torn tail from a
crash mid-append is tolerated on read and repaired on next write.
- Stats memoises per store keyed on the file's stat fingerprint and no
longer takes the manager mutex, so the 5s status poll stops parsing
vector-laden JSONL and stops blocking behind seeds.
- KNN Classify decodes each neighbour payload once (was twice) and
builds refs and votes in a single pass with one fallback return.
- Corpus file writes fsync before rename/close.
- The corpus manager is built eagerly in newApplication (sync.Once
dropped); test helper dead branch removed.
Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Richard Palethorpe <io@richiejp.com>
* feat(router): bind knn corpus vectors to an embedder fingerprint and fail closed on mismatch
Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Richard Palethorpe <io@richiejp.com>
* chore(mcp): align corpus tool prompts and the mutating-tool safety list
Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Richard Palethorpe <io@richiejp.com>
* feat(proto,backend): report embedding shape from the llama-cpp backend
Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Richard Palethorpe <io@richiejp.com>
* feat(embeddings): Go-side pooling — mean/last/decayed_mean with half-life
Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Richard Palethorpe <io@richiejp.com>
* feat(embeddings): accept chat messages[] and per-request pooling on /v1/embeddings
Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Richard Palethorpe <io@richiejp.com>
* chore(middleware): name the failing fields when post-merge validation 400s
An intermittent post-merge validation failure surfaced as an opaque 400
during integration (pooling scheme mismatch that no client had sent).
Log the model, the request's pooling override, and the merged config's
pooling fields at the failure point so the next occurrence identifies
whether the request or the stored config carried the bad value.
Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Richard Palethorpe <io@richiejp.com>
* fix(embeddings): scheme override must not inherit the config's half-life
A model config defaulting to decayed_mean pooling carries
pooling_half_life_tokens; a request overriding the scheme to mean/last
without its own half-life inherited that value, and post-merge
validation rejected the pair the server itself had assembled. Zero the
inherited half-life when the overridden scheme is not decayed_mean; a
request that explicitly pairs a half-life with a non-decayed scheme
still 400s.
Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Richard Palethorpe <io@richiejp.com>
* fix embedding pooling validation and router bounds
Declare backend embedding layouts and reject incompatible pooling modes. Reset local-store dimensions after a full clear, validate KNN thresholds, and add real backend and store integration coverage.
Assisted-by: Codex:gpt-5
Signed-off-by: Richard Palethorpe <io@richiejp.com>
* ci: run local-store integration tests
Build and install the local-store backend in the Linux test job, then run the existing store integration suite so new specs are discovered automatically.
Assisted-by: Codex:gpt-5
Signed-off-by: Richard Palethorpe <io@richiejp.com>
---------
Signed-off-by: Richard Palethorpe <io@richiejp.com>
|
||
|
|
799cc9f211 |
feat: bound global admission and expose running backend traces (#11560)
feat: bound backend admission and expose running traces Add process-wide backend execution admission without blocking UI or administrative HTTP work. Represent backend operations while they are in flight, surface running traces with immediate log links, and tie streaming admission leases to the gRPC receive lifecycle. Assisted-by: OpenAI Codex: GPT-5 Signed-off-by: Richard Palethorpe <io@richiejp.com> |
||
|
|
0aaff91ebd |
feat(ui): unify model and backend lifecycle (#11548)
* feat(ui): add installed model lifecycle Models now owns catalog exploration and installed runtime controls under one canonical route. URL-owned state keeps lifecycle context recoverable through links and browser history. Assisted-by: Codex:gpt-5 Playwright * feat(ui): add installed backend lifecycle Backends split discovery from backend-binary management. The canonical page now keeps both lifecycle views under one URL-backed shell while it preserves target-node placement. Assisted-by: Codex:gpt-5 Playwright * fix(ui): repair lifecycle state updates Installed models lost distributed refreshes and kept a deleted selection. Backend searches also stopped tracking URL changes, while batch upgrades stopped after their first error. Preserve background refreshes and finish each requested batch action. Drive catalog results from URL-backed state without losing full metadata. Assisted-by: Codex:gpt-5 [Playwright] * feat(ui): make resource pages canonical Replace Host navigation with canonical Models and Backends lifecycle routes, preserve legacy management URLs, and surface shared host capacity on the Operate overview. Assisted-by: Codex:gpt-5 [Playwright] * feat(ui): complete canonical resource lifecycle Finish the responsive list-to-detail behavior, remove the retired Host implementation, and keep Explore focused on discovery while Installed owns destructive actions. Update regression coverage, localization, documentation, and development binding for the canonical resource pages. Assisted-by: Codex:gpt-5 [Playwright] * docs(ui): record the UI design context Record the approved users, brand character, and design principles so future interface work uses the same product direction. Index the context from the repository's agent instructions. Assisted-by: Codex:gpt-5 --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
88edd7fc7f |
fix(distributed): run cold model loads as durable jobs instead of holding the advisory lock (#11514)
* fix(advisorylock): set statement_timeout alongside lock_timeout
WithLockCtx already overrides a deployment-wide lock_timeout on its
dedicated connection so a blocking pg_advisory_lock() waits its turn
instead of failing with 55P03. statement_timeout aborts that exact same
statement independently, with SQLSTATE 57014, and was not overridden.
Production roles commonly carry statement_timeout=60s. Any guarded
section longer than that (a cold model load stages for tens of minutes)
therefore killed every concurrent waiter:
advisorylock: acquiring lock 9003261067483446873: ERROR: canceling
statement due to statement timeout (SQLSTATE 57014)
Derive it from the same context budget as lock_timeout, with a matching
RESET so the pooled connection is returned clean.
Assisted-by: Claude Opus 5 [claude-code]
* feat(distributed): add ModelLoadJob, the durable cold-load record
A cold load in distributed mode is a long-running background job, but it
was modelled as a synchronous side effect of an inference request: the
whole of it (backend install, multi-GB staging, checkpoint load) ran
inside the per-model advisory lock. Loading a 35.7 GB GGUF held that lock
for ~20 minutes, so every concurrent request for the same model blocked
on pg_advisory_lock and died at the role's 60s statement_timeout.
Introduce the row that lets the lock shrink to a decision. Exactly one
ModelLoadJob may be active per tracking key; that uniqueness — not the
lifetime of a lock — is what de-duplicates concurrent loaders across
replicas. ClaimLoadJob does its read-then-write under the advisory lock
and nothing else: no network, file or gRPC I/O inside the guarded
section, so a claim costs milliseconds no matter how long the resulting
load takes.
LastProgress is a heartbeat rather than a byte counter. A checkpoint load
legitimately moves zero bytes for many minutes, so a reaper keyed on byte
movement would reclaim a healthy job mid-load; byte progress stays the
concern of load_deadline.go. A job whose heartbeat stops for longer than
the orphan window is reclaimable, so a replica killed mid-load cannot
wedge a model permanently.
Failed jobs keep their row for a short grace so an immediately-following
request reports the real cause instead of silently starting a fresh load
of a model that just failed.
No caller yet — the router moves onto this in the next commit.
Assisted-by: Claude Opus 5 [claude-code]
* refactor(distributed): run cold loads as jobs, outside the advisory lock
Route wrapped the entire cold load — node selection, backend install,
multi-GB staging and the remote LoadModel — in the per-model advisory
lock. The lock's job is to de-duplicate concurrent loaders, a decision
that takes milliseconds; holding it for the tens of minutes the resulting
work takes is what turned a dedup mechanism into a cluster-wide outage
for that model.
Split it into a claim and a run. The claim is the only thing left inside
the lock. The run is a background job owned by the claiming replica and
bounded by the same progress-extended deadline as before; every other
request for that model — local or on another replica — attaches as a
waiter and is served the moment the model is ready, with no duplicate
load and no lock contention.
Waiters share one broadcast rather than an ordered queue: they all want
the identical outcome, so ordering them would add fairness machinery that
changes no result. The local channel wakes same-replica waiters instantly
and a 2s DB poll is the authority, because a waiter on another replica
has no channel to close. On wake a waiter re-runs the warm path rather
than trusting the signal — the model may have been evicted in between.
A waiter whose client disconnects returns immediately and the job keeps
running; it belongs to the job record, not to the request. A failure is
recorded on the row so every waiter reports the real cause, and the row
survives briefly so the next request does not read "no job" as "not
loading" and start a duplicate load of a model that just failed.
The runner heartbeats the row on a fixed interval whether or not bytes
are moving, which is what keeps a legitimately silent checkpoint load
from being reclaimed as an orphan. Phase (installing/staging/loading) and
placement ride to the heartbeat on the context, the same seam
load_deadline.go already uses, so single-host paths are untouched.
Non-distributed mode (no DB) keeps the inline load exactly as it was.
Assisted-by: Claude Opus 5 [claude-code]
* feat(distributed): bound the wait for a loading model and answer with progress
A request whose model is cold-loading now attaches to the running job and
is served the moment the model is ready. That wait has to be bounded: a
held HTTP request cannot survive real infrastructure, and an ingress or LB
idle timeout kills a twenty-minute request regardless of what LocalAI
does.
New LOCALAI_MODEL_LOAD_WAIT (default 60s) bounds the CALLER, never the
load — the job keeps running either way. On expiry the request gets 503
with Retry-After and a structured body naming the model, the node, the
phase, byte progress and an ETA. The `error` envelope keeps OpenAI
clients working; `loading` is additive so they ignore it.
The ETA comes from the job's own observed rate and is omitted rather than
guessed until enough bytes have moved for that rate to mean anything: a
confidently wrong ETA on a twenty-minute wait is worse than none.
Retry-After is that ETA when known, clamped to [5s, 300s], and the wait
budget otherwise.
LOCALAI_MODEL_LOAD_WAIT=0 waits unbounded, for deployments with no proxy
in front. Zero in the config struct still means "unset, use the default",
so the CLI records the operator's zero as ModelLoadWaitUnbounded rather
than losing the distinction.
The distributed branch of ModelLoader.loadModel wrapped the router's
error with %s, which flattened it to a string. Use %w: the typed error is
what the HTTP layer keys the 503 off.
Assisted-by: Claude Opus 5 [claude-code]
* feat(api): add GET /api/models/{id}/load-status
A client that receives 503 while a model stages onto a worker needs
somewhere to poll. This returns the same `loading` object the 503 carries
— phase, node, byte progress and ETA — or 404 when no load is running.
Read-only and observability-shaped, so it is deliberately neither
admin-gated nor feature-gated: it explains a 503 the caller just
received, and hiding that behind a per-modality feature would make the
explanation for a failed image request depend on chat permissions. It
also gets no MCP tool, since there is nothing here an admin would manage
conversationally.
Registered on the surfaces from .agents/api-endpoints-and-auth.md: the
swagger block (existing `models` tag, so /api/instructions needs no new
area), the endpoint discovery maps in RegisterLocalAIRoutes, regenerated
swagger, and the distributed-mode docs page. No FLAG_* usecase is
involved, so capabilities.js is unchanged.
Assisted-by: Claude Opus 5 [claude-code]
* feat(ui): show cold-load progress in Chat and retry when the model is ready
A chat request for a model that is still staging onto a worker now gets a
503 carrying live progress instead of an error. Render it: the composer
shows the phase (installing / staging / loading), the node, the percent
and the ETA, then polls load-status and re-sends the request the moment
the model is ready.
Reuses the staging progress idiom the page already had rather than
inventing a second one — the two sources are folded into one
loadProgress, with the load job winning because it is authoritative
across frontend replicas and knows the phase, where the staging operation
only knows about a byte transfer this replica happens to be performing.
Waiting is bounded (three send attempts, ~30 min of polling each), so a
load that never finishes still surfaces as an error rather than as a
spinner nobody questions. An aborted generation stops the polling too.
Assisted-by: Claude Opus 5 [claude-code]
* fix(distributed): check warm-path cleanup errors
The router moved legacy cleanup calls onto newly linted lines. Report
cleanup failures while preserving the fallback to a cold load.
Assisted-by: Codex:gpt-5 [golangci-lint]
---------
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
|
||
|
|
95653f221e |
fix(ui): keep agent import action visible (#11488)
* fix(ui): keep agent import action visible The header hid its full import label after the agent list became non-empty. Hide only the nested file input so users can import more agents. Assisted-by: Codex:gpt-5 * test(ui): match the agent import label The Agents page renders the action as Import. The test searched for Import Agent, so it failed before checking visibility. Mock the observables request to remove backend timing from the fixture. Assisted-by: Codex:gpt-5 --------- Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
5c63969760 |
fix: Show MCP connection errors in the UI (#11495)
* fix(mcp): surface configured server failures Keep model-configured MCP servers visible when discovery or connection setup fails, propagate status through distributed discovery, and let the Chat UI show actionable errors while retrying unavailable servers. Add model-editor metadata for remote and stdio configuration and document the expected format, deployment networking boundary, and alternate MCP scopes. Assisted-by: Codex:gpt-5 Ordino golangci-lint Signed-off-by: Richard Palethorpe <io@richiejp.com> * build(compose): match CUDA development image Configure the API image with the cublas, CUDA 13, auth-tagged build settings used by the local development Makefile invocation, including the 24-way Docker build. Assisted-by: Codex:gpt-5 Ordino Signed-off-by: Richard Palethorpe <io@richiejp.com> * revert: keep host build settings out of compose The CUDA development deployment is managed from ~/docker/localai, not the repository example Compose file. Restore the generic example and keep machine-specific build settings in the host deployment. Assisted-by: Codex:gpt-5 Ordino Signed-off-by: Richard Palethorpe <io@richiejp.com> * fix(docker): exclude local agent artifacts Keep Claude worktrees and locally installed verification tools out of the Docker build context. These host-only directories added roughly 1.9 GB to every root image build. Assisted-by: Codex:gpt-5 Ordino Signed-off-by: Richard Palethorpe <io@richiejp.com> --------- Signed-off-by: Richard Palethorpe <io@richiejp.com> |
||
|
|
45cb3983ee |
fix(ui): unmerge the class strings that left buttons in browser chrome (#11462)
Eight header controls across seven pages had two or three elements' classes
collapsed into one string. The wrapper or the icon ended up wearing the
button classes, and the buttons themselves were left with no class at all,
so they rendered in the browser's own chrome. Reported on Agent Jobs; the
grep found the rest.
`fas` does not draw anything by itself: it sets
`font-family: "Font Awesome 6 Free"` and weight 900 on whatever carries it,
and the `fa-*` class supplies the glyph via ::before. So
<button className="btn btn-primary fas fa-plus">
renders its own label "New Task" in the icon font, and
<div className="hstack btn btn-primary btn-sm fas fa-edit btn-secondary fa-arrow-left">
<button>Edit</button>
<button>Back</button>
</div>
styles the flex wrapper as a button that is both primary and secondary,
points two glyphs at one ::before, and leaves both real buttons bare.
Fixed, all of them keeping the correct `<i>` child they already had:
- AgentJobs, AgentTaskDetails (x2), AgentCreate - icon classes off the
button.
- AgentTaskDetails, AgentJobDetails - wrapper back to plain `hstack`, and
the two buttons inside each get the variants the wrapper had been
holding. Back is secondary and leads, Edit/Cancel is the emphatic one
and trails, matching every other detail header.
- VoiceLibrary, VoiceProfileCreate - the title `<i>` had swallowed the
action link's classes, so "Create voice" and "Back to library" were
unstyled anchors. Back was also drawing a "+" because it had inherited
fa-plus while its own fa-arrow-left sat up in the title.
- P2P - a stray fa-circle-info on the title icon.
The ninth instance was ImportModel, where this class of bug was first
found. #11461 rewrote that file and landed first, so nothing is left to fix
there.
Guarded by e2e/class-hygiene.spec.js, which reads the source rather than
walking routes: several of these pages need agent or voice data before they
render a header, so a route walk would skip exactly the pages that had the
bug. It fails on an icon-font class outside an `<i>`/`<span>`, on two glyphs
or two button variants on one element, and on a layout wrapper that is also
a button. Font Awesome modifiers (fa-spin, fa-fw, sizes) are excluded, so
the `fa-spinner fa-spin` idiom stays legal.
e2e: 428 passed.
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] [Playwright]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
|
||
|
|
7a7fb00730 |
feat(ui): rebuild the import form on the restyled design language (#11461)
The import page took the new palette in #11305 but kept its old layout, so it stayed a 760px column with the primary action detached from the form it submits. Two of the problems were outright bugs. The Import button carried no className at all, so the page's single most important control fell through to the user-agent button: system chrome, wrong radius, no design-system focus ring. The YAML button carried `fas fa-save fa-upload`, which sets Font Awesome as the button's own font family (its label text inherits it) and points two glyph classes at one ::before. On the layout: `page--narrow` is documented for "forms / single-record edit views", and in Advanced mode this page held a URI field, a six-section format guide, ten modality chips, nine preference fields, a key-value repeater and a YAML editor at `calc(100vh - 400px)`. The width was the symptom; one column was the disease. - `page--medium` with a work column and a format reference beside it. The reference answers the only question a first-time admin has and used to sit behind a chevron, closed by default. Below 1024px it becomes a disclosure rather than disappearing. - The source field is the hero: monospace, because it holds something you paste, and it carries its own Import button. That removes the hidden aria-hidden submit button that existed only because the real action sat outside the form. - Simple and Advanced are gone. They were ~80% the same surface, and the overlap cost a mode switch, a localStorage key and a three-button Keep/Discard/Cancel dialog whose only job was protecting state that switching modes would hide. One form with a collapsible options panel hides nothing, so none of it is needed. What genuinely differs is the kind of input, which is now the two tabs: a source, or YAML. - The size/VRAM estimate reports under the field that produced it instead of as a banner above the page header, and an import in flight gets the progress, phase and byte counts the poller already returned and the old status card threw away. - ModalityChips resolves its labels through the same `modality.*` keys as the dropdown it filters. It hardcoded English shorthand, so one modality carried two names on one screen ("Speech" on the chip, "Speech recognition" on the group it scrolled to) and seven locales had neither. Its inline styles and its pill radius move onto the design system. - Three inline styles go, including both conditional-padding hacks; the only one left is the progress bar's runtime width. Baseline 538 -> 535. Docs updated in the same change: the WebUI section described a Simple and an Advanced mode and told the reader to "Toggle to Advanced Mode". e2e: 426 passed. The mode-switch suite is replaced by one covering the tabs and the disclosure, and a new layout suite pins the width, the styled primary action, the absence of an icon-font button, the reference column at both widths, and the estimate's position. Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] [Playwright] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
ab52813342 |
feat(modelartifacts): support bounded parallel Hugging Face file downloads (#11162)
* feat(modelartifacts): support bounded parallel Hugging Face file downloads Closes #11114. Snapshot materialization fetched every file through the sequential executor in DownloadFilesWithContext, so a repository split into many shards spent most of its wall clock in per-file request latency rather than moving bytes. Add DownloadFilesWithConcurrency, an errgroup with SetLimit, and keep DownloadFilesWithContext as a wrapper that passes a limit of 1. That leaves the two non-artifact callers (core/gallery and the model config loader) on exactly the path they had: tasks still run in slice order, and the first failure still returns before any later task starts. Only whole files run in parallel. A single file is never split, so the .partial resume machinery and the per-file SHA check in downloadTaskWithRetry are untouched. Two details the parallel path forced: - completedBytes becomes an atomic.Int64. Several AfterDownload hooks add to it while other files' progress callbacks read it; without this the race detector reports three races on the new specs. - The caller's status callback is serialized. The sequential path gave it an implicit guarantee of never being entered twice at once, and it belongs to the caller, so the executor keeps that promise rather than pushing locking onto every caller. AfterDownload is deliberately not serialized -- it does the verify-and-promote work that parallelism exists to overlap. Manifest order needed no work: each hook already writes its own manifest.Files slot by snapshot index, so entries stay in snapshot order whatever the completion order. A spec now pins that. The default is 1, unchanged behaviour. A shared models volume is often the bottleneck rather than the link, so raising it is a deployment decision; --artifact-download-concurrency and LOCALAI_ARTIFACT_DOWNLOAD_CONCURRENCY expose it on both `run` and `models install`. Not done here, per the issue: no chunk-level parallelism within a single file, and no throughput measurements across concurrency 1/2/4/8 -- that needs a representative sharded repo and a real link. Assisted-by: Claude:claude-opus-5 go-test gofmt Signed-off-by: Adira Denis Muhando <dennisadira@gmail.com> * feat(modelartifacts): expose download concurrency in settings Follow-up to review feedback on #11162: - The CLI flag and docs no longer describe the limit as Hugging Face specific. It applies to any artifact source, as @mudler pointed out. - artifact_download_concurrency is now a persisted runtime setting and is editable from the WebUI, so it can be changed without a restart. The manager's limit becomes an atomic.Int64 behind SetDownloadConcurrency, because a live runtime setting can be updated while a materialization is already in flight. Injected materializers stay compatible through an optional setter interface, so a manager that does not implement it is simply left alone. Verified before taking this on: go build, go vet and go test -race all pass for pkg/modelartifacts, pkg/downloader and core/config. The React UI builds with vite, artifact_download_concurrency is present in the built Settings chunk, and eslint reports the same 8 pre-existing warnings on Settings.jsx as it does without the change. Implementation contributed by localai-org-maint-bot on the review thread; reviewed, verified and signed off by me. Assisted-by: Codex:gpt-5 Assisted-by: Claude:claude-opus-5 go-test vite eslint Signed-off-by: Adira Denis Muhando <dennisadira@gmail.com> --------- Signed-off-by: Adira Denis Muhando <dennisadira@gmail.com> Co-authored-by: localai-org-maint-bot <bot-opensource@localaisrl.com> |
||
|
|
5ac445e1d4 |
fix(react-ui): restore 3D Studio results and history (#11393)
* fix(react-ui): restore 3D Studio results and history Keep large conditioning-image payloads out of the rendered request panel so the generated viewer can mount reliably. Accept clipboard images and synchronize 3D history consumers so new results appear in Studio without a reload. Cover clipboard input, bounded request rendering, result display, and cross-view history synchronization with Playwright. Assisted-by: Codex:gpt-5 Playwright * perf(react-ui): idle the 3D viewport when still Limit auto-rotate rendering to 30 FPS and stop scheduling frames when rotation is disabled. Resize, view controls, and pointer input invalidate the still frame on demand. Assisted-by: Codex:gpt-5 Playwright |
||
|
|
147a5ee783 |
fix(react-ui): stop traces page crash when switching trace tabs (#11387)
Switching from Backend Traces back to API Traces crashed the page with "can't access property status, e.response is undefined" (#11376). The API table briefly renders the previous tab's backend rows while the refetch effect is still pending, and those rows carry no `response` envelope. The status column dereferenced it unguarded. Render a neutral placeholder instead of throwing, and cover the tab-switch scenario with a regression spec. Assisted-by: opencode:big-pickle Signed-off-by: Nandana Dileep <110280757+nandanadileep@users.noreply.github.com> |
||
|
|
9f62401fca |
feat(traces): show in-flight API requests (#11368)
Register JSON API exchanges before their handlers run so the traces dashboard can surface active work. Replace the live entry with the completed persisted record under the same ID, and clean it up if a handler panics. Assisted-by: Codex:gpt-5 Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
1741df0bf1 |
fix(ui): scale the chrome audit's timeout to the number of routes it walks (#11319)
chrome-audit.spec.js walks 25 routes in a single test, and has been the
UI E2E suite's failure on 5 of the last 6 master runs. It always dies the
same way, at the 30s per-test default:
Test timeout of 30000ms exceeded.
Error: page.waitForTimeout: Test timeout of 30000ms exceeded.
19 | await page.goto(route)
> 20 | await page.waitForTimeout(400)
The spec is new in 5cb0c1a8; the commit before it was green, and every
run since has been red on this file.
The failure is cumulative rather than one bad route. Across those runs
the clock runs out at line 19, 20 or 21 depending on where the loop
happens to be, and the timeout lands on waitForTimeout rather than on
goto, which is what running out of budget looks like as opposed to a
navigation that hangs. 30s over 25 routes is ~1.2s each, including a
deliberate 400ms settle, so there is very little headroom to begin with.
Give the test a budget proportional to its work: six seconds a route.
That absorbs a slow runner and still fails promptly if a route genuinely
hangs.
Verified: the spec passes on the current UI in 12.2s solo, and the full
suite passes 418 at 8 workers locally. What I could NOT do is reproduce
the CI timeout on this machine, which has 20 cores against the runner's
2 to 4; under synthetic CPU load it still finished in 13.5s. So the fix
is argued from the CI signature and the arithmetic, not from a local
repro, and the proof is this spec going green on the hosted runner.
Note test.setTimeout() has to be called inside the test body. At module
scope Playwright rejects it with "test.setTimeout() can only be called
from a test".
Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
|
||
|
|
fd4ec083b9 |
feat(downloads): add resume-safe pause action (#11222)
Give gallery operations distinct pause and cancel paths. Pause preserves partial download data so reinstalling the same model or backend resumes through HTTP Range, while cancel keeps its destructive semantics. Surface the action in the Activity UI and document the API behavior. Assisted-by: Codex:gpt-5 Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
5cb0c1a872 |
feat(ui): close the gap between the shipped UI and the design mocks (#11307)
* feat(ui): give the Operate overview real numbers and traces a latency shape First two items from a component-by-component comparison against the mocks. The pattern that audit found: everything newly built matched, everything pre-existing got the palette but not the layout, and an "absent rather than empty" rule hid most of the overview exactly when someone was looking at an idle installation. **The headline grid is always rendered**, including at zero, with a fourth cell for host memory. Hiding it removed the page's structure precisely when it was most likely to be read, and "0 failed" is information — an absent panel is not. The quiet case is now said in a line underneath instead of by showing nothing. **The sections state counts** rather than listing their destinations: backends, models, updates and running operations instead of the words "Usage and traces". That needed installed backend and model counts in the summary context, which are two more cheap reads on the poll that was already running. **Traces rows carry latency as a bar as well as a figure**, scaled against the slowest request currently in view and turning amber past two seconds. The table had no latency column at all — the number was buried in the expanded detail, so the shape of the tail was invisible while scanning. Scaling against the view rather than an absolute ceiling is deliberate: what matters when reading a page of traces is which of these are the outliers, and an absolute scale flattens every row on a fast installation into nothing. Full e2e suite: 409 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): name the engine on Home's resident models, and add jump-back-in Third item from the mock comparison. The mock showed each resident model with the engine serving it. /system carried only the id, so the audit recorded this as blocked on a server field — but the config loader is already in scope where that response is built, so it is one lookup. SysInfoModel gains an optional `backend`, resolved from the model's config and omitted rather than guessed when there is none (a loose file, or a config since removed). Home renders the column blank in that case; the test pins both halves of that. Memory per model stays out. It is not one lookup — it would mean asking each backend process — and inventing a number beside a real one is worse than leaving the column off. "Jump back in" is the block the mock had and Home did not. The quick-links row above it is a set of first-run actions; these are the three places someone returns to, each stated with what it currently holds rather than as a bare label. Go: routes suite passes. Full e2e suite: 412 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): rank the recommended models as lanes instead of equal cards The hardware recommendations were a grid of equally-weighted cards. The list is already sorted by fit, and a grid throws that order away: three cards side by side say "pick one", when the page has actually formed an opinion about which one. They are lanes now, read top to bottom in fit order, with the leader carrying the single amber "Best fit" label and the rest marked "Also fits". One opinion per page — the alternatives are alternatives, not runners-up each worth their own colour, which is how a strip of coloured badges ends up meaning nothing. Below 720px the size and VRAM columns drop and the lane keeps the name and the install action, which are the two things a narrow screen needs. The existing panel spec moves off .rec-models-item onto .lane rather than being deleted; dismissal, collapse, keyboard operation and install all still pass unchanged, and there is a new assertion that exactly one row is called out. Full e2e suite: 413 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * fix(ui): drop capsule chips app-wide, and un-break the empty voice library **Pills are gone.** A capsule radius reads as a tag floating on the surface, which fights a system whose structure is hairlines and square corners — and with chips on Discover, Host, Activity and the biometrics pages, "some pages have pills" was the real inconsistency rather than any one page. Sixteen selectors move to the small radius: filter buttons, tab pills, activity and biometrics chips, file and count badges, the jump-to-latest control, the nav badge. Round *buttons* keep their circle — .lightbox__nav and .home-send-btn are circles, not capsules — as do every progress track, status dot and avatar, which are round because they are round, not because they are tags. **The empty voice library was unusable.** `.voice-library-empty` sets min-height: 430px, border: 0 and background: transparent — a description of the empty PANEL — and it had been attached to the action instead. The create button was therefore a 430px transparent box that pushed itself out of the panel and could not be seen. Moved onto the container it describes, which now centres its action rather than letting it fall off the bottom. Same class-mangling shape as the Agents header fixed earlier. Full e2e suite: 416 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): put Host's headline figures on the shared hairline strip Host had shadowed, clickable StatCards above a page that already has a rail, a pane and a tab bar — a second dashboard language on one screen, and a different one again from the figures inside its own detail pane. The Operate overview's figure grid is generalised into a shared `.stat-strip` and Host adopts it, so the two pages read as one system: same cell, same figure scale, same tone vocabulary, and the same hairline grid the split-view StatGrid already uses. The cells stay clickable and still route into the tab and filter they describe, because a count is worth more when it is also the way to the thing counted. Tone is spent only where the number means something — running and updates when non-zero — since a strip where every cell is coloured has no emphasis left. Two bugs made on the way, both now covered: - The first version put `<button>` elements inside a `<dl>` with `<dt>`/`<dd>` inside the buttons. Neither is valid, the browser re-parents both, and the cells collapsed. These cells are a set of controls, so a plain container of buttons is also the honest markup. - Even correct, the strip rendered 2px tall: `.page--app` is a flex column whose split view takes flex:1, so a child with no intrinsic minimum is shrunk away. The old cards survived only because `.stat-card` carried min-height:96px. The strip now declines to shrink, with a test pinning it. The stat-card specs are retargeted rather than deleted: they were written to guard a class collision on a page that no longer uses cards, so they now guard the strip's labels and its height. Full e2e suite: 417 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): make Backends notices an edge rather than a filled card The install and upgrade banners were tinted cards with a full border. A filled panel makes every notice shout at the weight of an error, which is how notices stop being read — and Backends shows one on most visits, so it was shouting routinely. They are now a hairline with a coloured left edge, the same treatment the Operate overview gives rows that want a decision, so "this needs you" looks the same wherever it appears. Counts in the notice take the monospace tabular figures the rest of the console uses. Also drops the last inline style on the page, and refreshes the inline-style baseline, which has read 624 against a real count since #11288 landed. The gate exits 0 either way, so nothing was failing — but a baseline 86 above the truth would have let that many inline styles back in unnoticed. Now at 538, which tightens the ratchet rather than loosening it. The spec creates the upgrade it asserts on rather than skipping when the mock has no notice: a test that skips is a test that proves nothing. Full e2e suite: 418 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): finish the mock parity list, and stop hiding the recommendations The last two items from the audit, plus a correction. **Discover's use-case shelf is lanes.** These are a list of ways in, read in order; a grid of equal cards asks the reader to compare them, which is not the choice on offer. **The request panel reaches every generator.** Video, 3D, Sound and Audio FX join Images and Speech, so each one teaches its own endpoint rather than two of six doing it. Audio FX records the fields that shape the request rather than the bytes, since its payload is multipart. **Recommendations no longer collapse themselves.** They were folded away by default once anything was installed. That is the page's one opinion about this host, and an opinion hidden by default is one the reader never gets. Someone who disagrees can still collapse it and that choice is remembered — the difference is that we no longer make it for them. Three specs asserted the old default and now assert the new one. The use-case heading also sat a line's width from the text it introduces, so the two read as one paragraph. It has air under it now, and the shelf is separated from the recommendations above it. Two tests removed rather than kept: a generator loop whose only real assertion was `expect(endpoint.length).toBeGreaterThan(0)`, and an earlier card-gap guard that could only skip. A test that cannot fail is worse than no test, because it reads as coverage. Full e2e suite: 418 passed, 4 skipped. Inline styles at baseline. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * fix(ui): make the Host figures legible and give the strip its spacing back Three defects introduced by the Host redesign, all found by looking at the running app rather than by the suite. **The figures were invisible.** "Running now" and "Updates available" rendered pure black on the dark ground. Two causes compounding: the `--muted` tone alias never landed, because the source rule has extra spaces before its brace and the exact-match edit missed it silently; and a `<button>` does not inherit colour, so with no tone rule the value fell back to the user agent's `buttontext`. Both fixed, and a test now fails on any figure computing to pure black. **The strip sat flush against the resources panel.** `.stat-strip` declares `margin: 0 0 ...` and is declared later in the file than `.manage-summary`, so the shorthand quietly won and the top margin became zero. Raised to `.stat-strip.manage-summary` so it beats the shorthand on specificity rather than on declaration order, which is the kind of thing that breaks again the next time a rule moves. **Discover's use-case heading had a doubled gap.** `.zero-pane` is a flex column that already separates its children; adding a margin on top of the gap stacked the two. The margin is gone and the heading keeps only its own breathing room. Full e2e suite: 420 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): make Studio's tabs path segments rather than a query parameter `/app/studio?tab=images` reads like a filter applied to a page. It is navigation: a different generator, with its own state and its own deep link. It is now `/app/studio/images`, with the overview at `/app/studio`. Legacy `?tab=` links are redirected once to the path form, replacing the history entry so Back does not bounce between two spellings of the same place. Bookmarks and older links keep working and land on the canonical URL rather than a second version of it, which is the part worth having a test for. The nine `?tab=` references were all in specs, none in docs, so the migration is contained. They move to paths, and a new spec pins the redirect. Full e2e suite: 421 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): make the hardware recommendation a section, not a dismissable card It was a bordered card with a collapse control and a close button, sitting inside a pane that is otherwise hairline sections. Two problems: it read as something bolted onto the page rather than part of it, and treating it as an interruption to be shut is the wrong frame for the one thing the page has to say about the machine it is running on. It is now a plain section with the same heading treatment as the shelves below it. The collapse state, the dismissal, their storage keys and the legacy key read for backwards compatibility all go with it, along with the installedCount prop that existed only to pick a default collapse. Five specs described behaviour that no longer exists and are removed rather than adjusted — collapsing, dismissing, persistence of both, and the toggle's keyboard handling. One new spec asserts the replacement contract: no control with aria-expanded, no dismiss, and no card border. Full e2e suite: 416 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * fix(ui): restore every stripped icon and every default-chrome control You reported two broken icons. They were not two: an earlier automated edit had stripped the `fa-*` class from twenty `<i>` elements across eleven pages, and an `<i>` with no icon class renders nothing at all. Settings' save button, the voice-profile back link, and eighteen others — agent row actions, task and job buttons, import and create actions — were all drawing empty space. Each is restored from its own context rather than a blanket icon: the agent row gets pause/play, pen, comments, file-export and trash; the fine-tune toggle swaps plus for xmark as it opens; the P2P documentation link gets the external-link glyph. The same edit left controls without their classes. Fine-tune's "Import config" was rendering in the browser's own chrome, and `.p2p-cmd__copy` set a border but no background, so it fell back to `buttonface` — a pale grey chip on a dark command block. FineTune's "New job" also had its icon classes folded into the button's className, the same mangling already fixed on the Agents header. Rather than fix the reported two and wait for the next report, this adds a standing audit: twenty-five routes are walked and the test fails on any visible control rendering with user-agent chrome, or any `<i>` without an `fa-*` class. It found the three remaining cases after the first sweep, and it is the reason the next one cannot ship quietly. Full e2e suite: 417 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
58ea2f5d79 |
feat(ui): give Operate and Studio a front door, and fix two layout regressions (#11305)
* feat(ui): give Operate a front door and fold six rail groups into four Opening Operate ran firstVisiblePath() and landed on Backends, because Backends happens to be written first in operateConsole.groups. The section that should answer "is anything wrong" opened on a package manager, and nothing was reported until you visited it. Adds /app/operate. Its one irreplaceable block is "Needs attention", which is empty when nothing is wrong and says so in a line rather than rendering a reassuring green panel. It collects stale backends, failed operations and unhealthy nodes. Everything else on the page is a summary you could already assemble by visiting four others. The rail regroups from six headings to four: Inference and Activity were both "the runtime right now", Access and System were both administration. No destination is removed and no gate changes, so isConsoleItemVisible and consolePaths are untouched. Overview leads the first group, which is what makes firstVisiblePath() return it without knowing it exists. Rail entries now carry a signal beside the label. This does not replace the sidebar badge and is not built as if it does: the badge stays on the always-visible sidebar entry for the reason recorded in Sidebar.jsx, that the rail exists only on Operate routes and can be collapsed. The signals are orientation while inside Operate, so they are aria-hidden and nothing urgent depends on them alone. OperateSummaryContext polls once for the whole console, following OperationsContext, which exists because per-consumer setInterval against one endpoint was the defect it fixed. It is mounted by ConsoleLayout for the Operate console only, so "poll only while in Operate" needs no route check. Built on usePolling, so it pauses on a hidden tab. Operations are read from OperationsContext rather than polled a second time, and each source degrades to no-signal on its own so one dead endpoint cannot blank the rest. It reads the cached GET /api/backends/upgrades and never the POST that forces a real registry check. Traces and Usage get no signal yet: /api/traces returns the list, so a count would mean fetching every trace to render one number. A counts endpoint is the honest fix and is scoped separately. Full e2e suite green (369 passed, 4 skipped), including a render-smoke entry for the new route. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): open Studio on what this machine can actually make Studio was a tab strip over six generators that opened on Images, which was never a decision, only the first entry in BASE_TABS. Nothing said which modalities this installation could run, so the way to learn that video had no model was to pick the tab and find an empty select. Adds an overview tab and makes it the fallback. Explicit tabs still win, so existing deep links keep working; anything unrecognised or gated now lands on the overview rather than Images. Each tab carries a dot: filled when an installed model advertises that modality, hollow when nothing serves it. That is the feature in one detail, turning the strip from navigation into a report of what the machine can do before anything is clicked. The dot is aria-hidden because the overview states the same facts in words and the dots change as models load. Two kinds of unavailable, which had to stop looking alike: - switched off, via a permission: no tab and no lane, unchanged - available with no model: a lane, and a route to installing one Studio now owns one MODALITIES table so the tab strip and the overview cannot disagree about what exists, and calls useModels() once, unfiltered, grouping in the browser. useModels(capability) fetches the whole list and filters locally, so a hook per modality would have been six identical requests to /api/models/capabilities on every mount. There is a test for that. Recent outputs read every localStorage store through a new readAllMediaHistory(), which avoids mounting five hooks that carry save timers the overview has no use for. 3D is read separately through use3DHistory rather than folded in: its entries are GLB blobs in IndexedDB, so they cannot come from the same synchronous read. Typical cost is the median of this machine's own history, not a guess, and renders as a dash when there is nothing to go on. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * fix(ui): stop the stat cards and the console rail breaking on small screens Two unrelated causes behind one report that /app/manage looks wrong when the window is narrow. The stat cards were being laid out by the wrong rule at every width. Two different components both claimed `.stat-grid`: the dashboard card strip that holds .stat-card children, and the detail-pane StatGrid the split views introduced further down App.css. Being later, the second won every shared property, so the cards got its 120px columns and its 1px hairline gap in place of their own 180px columns and spacing-md. Four cards were packed onto a row that fits two, labels wrapped to three lines and clipped, and the icon crowded the value. Renamed the strip to `.stat-cards`, after the children it actually holds, which also removes the mismatch of a `.stat-grid` container full of `.stat-card`s. The split-view component keeps `.stat-grid` and its BEM parts. The expanded console rail had no bounded height. Thirteen destinations stacked in one column is taller than a phone, so opening the menu pushed the page's own heading past the fold: the menu replaced the page rather than annotating it. Capped at 55vh with internal scrolling below 768px, so the content behind stays reachable. Both are asserted on behaviour rather than markup: no stat-card label may be clipped, the card gap must not be the detail pane's hairline, and expanding the rail must leave the page heading on screen. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): retemper the palette to localai.io and add the lane primitive The token half of the style transfer, plus the shared list idiom the two overviews had each grown their own copy of. theme.css moves from Nord to the website's palette, variable names preserved so every consumer moves with it: ground #13171f -> #0d1117, accent frost cyan #88c0d0 -> action blue #4f8cff, success sage -> mint #56d6a4, warning -> the amber #f1b95d the site spends only on the thing asking for a decision. Eyebrows go mint. Dividers become an opaque #29384a hairline rather than alpha over a varying surface, which is what makes stacked surfaces read crisply on the site. Light is derived, not inverted. The site ships one theme and never had to answer this, but the app does: blue darkens to #2f62d8, mint to #0d8b60 and amber to #8a5d0b, all clearing 4.5:1 on a cool paper ground, where the dark-mode values sit near 2:1. Same three roles, different values. Three files restate the palette because CSS variables cannot reach them: cmTheme.js (the whole CodeMirror theme), VoiceVisualizer and WaveformPlayer (canvas). Left alone they would have quietly kept the app half-Nord. The `.lane` primitive replaces the near-identical row CSS that OperateOverview and StudioOverview had each written: a full-bleed row on a hairline that insets on hover, with no card and no shadow. Callers supply only the column template. Both pages now use it, along with `.lane-head` for section rhythm and a `.page-pad` container for top-level pages outside a console shell — without which Studio sat flush against the sidebar with its eyebrow clipped. Studio's tab strip wraps rather than running off the edge at narrow widths. Full e2e suite: 386 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): put Home's resident models on lanes and give the footer one line Home's status line was three chips saying a thing was true. It now reports figures: how many models are resident, how many nodes are healthy, what share of memory is in use, set in tabular monospace so the digits line up. A chip answers whether; a figure answers how much, which is what someone opening the page at a glance is after. Resident models move from status chips to lanes, with the id set in a new `.lane__name--id` because an id is something you might type or paste and the UI face makes it read as a label. /api/system-information carries only the id, so there is deliberately no backend or memory column: inventing one would mean a server change this does not make. The footer was three centred rows and cost the bottom sixth of every page for chrome. It is one line now, version left and links right, wrapping to centred when the viewport is too narrow to hold both. Every link it had, it keeps. Full e2e suite: 392 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * fix(ui): correct three contrast failures and stop a guaranteed-404 poll A contrast audit of the new palette found three values below WCAG AA, one of which the previous commit message claimed was fine: - White on the #4f8cff button is 3.22:1, which is large-text only. The website does exactly this, but a button label in an app is not large text, so the label goes to dark ink at 5.88:1. Light mode keeps white, which is 5.44:1 on its darker blue. - Light-mode success was 4.08:1 on paper, not the 4.5 claimed. Darkened to #0a734f, 5.56:1. - Nord red was already 4.28:1 on raised surfaces, a pre-existing miss carried over unexamined. Lifted to #c96f78, 5.02:1. Lanes gain the two states they were missing: a 44px target on coarse pointers, matching what EntityRail already does so the two list idioms feel the same under a thumb, and a reduced-motion variant that keeps the background feedback while dropping the hover inset, which is a position change. The Operate summary no longer asks for /api/nodes on a single-node install. The cluster API answers 503 when distributed mode is off, so it was a guaranteed miss every fifteen seconds; it is now gated on useDistributedMode, the same condition the rail already uses for the Nodes entry. Full e2e suite: 392 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * fix(ui): restore the gap between overview blocks, and stop claiming zero nodes Two defects a design review surfaced. `.lane-head:first-child { margin-top: 0 }` was meant to stop the first block on a page carrying a top margin. But every <section> makes its lane-head a first child, so the reset applied to all of them and the gap between blocks vanished: "Sections" sat flush against the attention row above it. The header supplies its own bottom margin, so a uniform top margin is correct everywhere. The Cluster summary read "0 nodes" on a single-node install, which looks like a fault when the cluster API is simply switched off. It now says "Single node". Full e2e suite: 392 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): open dark by default, and stop clipping the collapsed sidebar footer Dark is the identity rather than a preference: localai.io ships one theme and it is this one, so an install should look like LocalAI before anyone has chosen anything. The OS setting no longer selects light on first load. The toggle still does, and a stored choice wins forever after, which the tests assert both ways. The collapsed sidebar footer stacked its controls but kept the expanded row's inline padding, so their edges were clipped against the 51px rail. Full e2e suite: 394 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(api): count traces server-side and give the Operate overview real totals The overview's headline block had no source. /api/traces returns the trace list, so "37 errors in 24h" meant fetching every buffered exchange to count it in the browser — waste that grows with the buffer, to produce three integers. Adds GET /api/traces/summary: totals, failures, p95 and a bucketed series for sparklines, over a window that defaults to 24 hours and is capped at a week. Deliberate calls, each with a spec: - A 4xx is the caller getting it wrong, not the installation being unhealthy, so only 5xx and transport errors count as failures. - p95 is a nearest-rank percentile rather than the slowest request, which is what a max would report and what makes latency panels lie. - Buckets are oldest-first so a sparkline reads left to right, and the slice is never nil: nil serialises as null and breaks .map() on the other side, which is a silent runtime error rather than an empty chart. - Exchanges outside the window are not counted at all. The route is registered before /api/traces/:id so "summary" is not captured as a trace ID. On the client, Traces and Usage gain the rail signals they were shipped without, the Observability section summary now states counts instead of listing its destinations, and an installation that has served nothing says so rather than showing three zeroes dressed as telemetry. Sparkline is a bare stroke with an emphasised endpoint and no axes: the figure above it already states the value, so its only job is the shape. Go: 185 middleware specs pass. Full e2e suite: 396 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * fix(ui): stop the memory chart calling a trade-off an error The VRAM-by-context chart rendered any build over the limit in error red, and escalated the verdict to the error tone as soon as two context sizes crossed it. But an over-limit build still installs — #11288 keeps a test on exactly that — so red overstates what is happening. A model that fits at 32k and not 64k is a trade-off, not a fault. Over-limit bars and the limit line now use the warning tone, which is the constraint colour used everywhere else in this branch: know what you are doing, not you may not. The error tone is reserved for "fits nowhere", where the model genuinely cannot run on this host. Full e2e suite: 397 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): give the new surfaces orchestrated motion Uses the reveal system already in the codebase rather than adding a library: pageReveal, .reveal-stagger and staggerStyle() were built for exactly this, and anime.js would be ~17KB duplicating four lines of CSS for list reveals. The overview's headline figures, attention rows and section lanes stagger in, as do Studio's modality lanes and recent outputs, so a page assembles in the order it is read instead of appearing all at once. Two additions beyond stagger. Rail signals transition on opacity when a poll lands, so a number changing reads as an update rather than a jump cut, and it stays on the compositor so it cannot reflow the rail. The attention block animates its left edge in — the one thing on the page that should announce itself, and on the border rather than the text so nothing moves under a reader. Both are dropped entirely under prefers-reduced-motion, alongside the lane hover inset already handled. Full e2e suite: 397 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): put the generators on a hairline field stack and record the request The workbench treatment from the mocks, applied where it costs least: both changes land on shared surfaces, so all six generators get them at once rather than drifting apart page by page. The control column stops being a shadowed card of boxed groups and becomes a hairline field stack — the panel is the page's left half, not an object floating on it — with uppercase micro-labels matching the eyebrow treatment used elsewhere. Because .media-controls is shared, Images, Video, 3D, Speech, Sound and Audio FX all move together. RequestPanel shows the request the form actually built, with a copy-as-curl. LocalAI is API-first and Studio is the best place in the app to teach its own endpoints: the form stops being a black box, and a result worth keeping can be reproduced from a shell without reverse-engineering which fields the page sent. It records what was sent rather than what the form currently holds, and renders nothing until a request has been made — a panel describing a request nobody made is a tutorial, not a record. Wired into Images and Speech. Full e2e suite: 401 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): make Chat a transcript instead of a bubble thread Rounded, filled, asymmetric bubbles fight a system built on hairlines, and they carry the speaker in shape and side rather than in words. The assistant side had already given up its bubble; this finishes the job. Both roles now run full width down one column, separated by a rule, each with a mono role label. The user turn keeps a left edge in the action tone so the two are still told apart at a glance, without a fill or a corner radius. The avatars go: the accent and the label carry the speaker, so the glyph was decoration once neither side had a bubble. Saying who is speaking in words rather than in geometry is also what survives being read aloud, printed, or looked at by someone who cannot pick the sides apart by colour. Full e2e suite: 404 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): dress the API reference in LocalAI's palette The Swagger page was the last surface still shipping in someone else's colours, which is conspicuous now that everything it links from is dark. Swagger UI has no theming hook, so rather than fork it we serve our own index ahead of the library's wildcard and restate the palette over its stylesheet. The library's own bundle and assets are still what load, so a swagger-ui upgrade cannot silently break the page — this is a skin, not a fork. Two things needed real care. Swagger tints the entire operation row per method via .opblock.opblock-post and friends, so the palette had to match that specificity rather than reach for !important; the method now lives on one edge instead of washing across the row, because a page where every row is a status colour has no status colour left. And the filled method chip put white on pale green, which was the least readable thing on the page — it is an outlined mono chip now, carrying the method in its border and text. Palette values are copied from theme.css rather than referenced: this page is served by Go and never sees the app's CSS. The comment says so, and says to keep them in step. Go: routes and middleware suites pass. Full e2e suite: 405 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * fix(ui): make tall split-view pages reachable, repair the Agents header, scale titles Three things found by actually using the app rather than measuring it. **Host was unusable.** The shell above a split view is overflow:hidden so the document cannot grow, which left anything taller than the viewport simply unreachable — and Host stacks a resources card, four stat cards and a tab bar above its split, so the bottom of the pane fell off at every window height with nothing to scroll. Every sweep I ran for this was horizontal, which is why it kept coming back clean. The page now scrolls inside the pinned shell. The pane keeps its own scroller: letting it grow instead pushes the document taller and stretches the rail to match, which is the regression e2e/discover-height.spec.js exists to catch, and which the first version of this fix duly caused. **The Agents header controls were unstyled** — "Create Agent" was rendering with the browser's default chrome. The markup had been mangled at some point: six unrelated classes merged into one string on the link, and the label and button left with none at all and empty icons. Repaired, with the inline flex replaced by a shared .header-actions class. **Page titles take the editorial scale from the site**: larger, tracked at -0.04em, on a line height near 1, so a two-word title reads as a statement rather than a label. The typeface is unchanged — DESIGN.md keeps the existing type system — so the whole difference is scale, tracking and leading, which is where the site gets its voice from. This was the biggest reason the running app still did not look like the mocks. Full e2e suite: 404 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
74b7ea2829 |
feat(ui): replace the gallery and inventory tables with a rail and a detail pane (#11288)
* feat(ui): rename the Install Models nav entry to Discover "Install Models" named the action rather than the destination, and it was the only multi-word entry in a rail of one-word ones (Home, Chat, Studio, Talk, Build, Operate). A bare "Models" was the obvious fix but it collides with the installed-models view under Host, which is a different page for a different job. "Discover" keeps the rhythm and says what the page is for. The icon moves from a download arrow to a compass for the same reason: the page is browsed before it is installed from. Translated in all seven locales rather than left to fall back, so a locale switch does not leave the entry in English. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5 [Claude Code] * feat(ui): replace the gallery table with a rail and a detail pane The eight-column table was not the real problem; the click-to-expand row underneath it was. Variants, files and a VRAM estimate never fitted inside a <tr>, so they were pushed into a drawer that could hold one model at a time, could not be linked to, and had no room to say anything useful. The gallery is now a rail to scan and a pane that answers. The pane has two states and no third: with nothing selected it is the discovery page, and with a model selected it is that model's detail. Selection lives in the URL, so a model is linkable and Back steps out of the detail instead of off the page. The rail groups by capability while browsing and flattens to results the moment a term is typed. That is a rule rather than a toggle: once someone has said what they are looking for, the buckets are between them and the answer, and making the user choose would be handing them our problem. The detail pane plots VRAM against context length with the host's own limit drawn across it. This is new information, not a restyle. A single number invites "so will it run?", and the honest answer is usually "yes, up to a 32k context", which is a shape rather than a number. The estimates were already fetched for every context size, so it costs no new request. Backends that take no context length say so instead of being given a meaningless chart, and a host with no GPU gets no chart at all rather than bars with nothing to compare against. The split-button variant menu goes with the actions column. The pane lists every build with its backend, quantization, size, fit and a details disclosure, each installable, which is what the dropdown was a cramped substitute for. Its tests move onto that list; the three contracts it alone carried (fetch-once caching, the loading state, an unfit build staying installable) are backfilled against the pane. RecommendedModels moves inside the pane, where it has the width to argue for a model instead of listing one, and keeps its own dismissal and collapse. Rail entries carry no description. Two lines is the budget and the second is better spent on whether the thing will run; the stripped-Markdown contract moves to the pane's lede, tooltip included. e2e: 123 passing across models-gallery, navigation, recommended-panel, model-artifact-operation, operations-strip and page-render-smoke. Inline styles in Models.jsx drop from 82 to 41. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5 [Claude Code] * refactor(ui): extract the split view into shared components Discover shipped its rail, pane and detail header as private functions inside Models.jsx. Backends and Host have the same defect and want the same shape, so leaving them there guarantees three rails that drift. SplitView, EntityRail, DetailHeader and StatGrid now live under components/split/. EntityRail is deliberately data-driven: a surface maps its own entity onto { id, name, icon, meta, stripe, groupId } and keeps its vocabulary to itself, which is what stops the rail learning about models, backends and loaded state all at once. The CSS moves with it. What was .discover__rail is .entity-rail, .discover__ pane is .split-view__pane and so on, because a class named after one page is a lie on the next two. Only what is genuinely Discover's stays behind the old prefix: the shelves, the hero and the VRAM-by-context chart. Two additions the shared rail needs and Discover did not: a state stripe, for surfaces read by condition before they are read by name, and an empty label. Discover passes neither. No behaviour change. e2e 100 passing across models-gallery, navigation and models-recommended-panel. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5 [Claude Code] * feat(ui): put the backend gallery on the split view Same defect as the model gallery, so the same shape: a seven-column table over a click-to-expand row that was the only place the repository, licence, tags and links could go. The rail groups backends by the use case they serve, sharing Discover's taxonomy on purpose: a backend is the runtime a use case needs, so "vision" ought to mean the same thing one level down. It flattens on a query for the same reason it does on Discover. The zero state is the one real departure. A backend's fitness is not free memory, it is the accelerator and platform it was built for, so the pane leads with what this host is, then what is not installed yet, then whether anything installed has gone stale. The table listed 37 runtimes and left "which of these can even run here" entirely to the reader. Distribution moves into the pane, which is the one thing a row could never carry: which nodes hold a copy and which do not, with the install-on-more control next to it rather than squeezed against a chip. The distributed and target-node action logic is unchanged, including the guard that keeps a hardware-specific build off the fan-out path. The split-button popover loses its per-row anchoring because there are no rows; one pane, one anchor. Selection lives in ?backend=, preserving the ?target= scope rather than clobbering it. e2e: 139 passing across models-gallery, navigation, backends-management, models-recommended-panel, nodes-per-node-backend-actions, page-render-smoke, operations-strip and model-artifact-operation. The backends spec gains six split-view tests; its three description-cell tests move onto the pane lede. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5 [Claude Code] * feat(ui): put the Host inventory on the split view The last of the three surfaces, and the one that is not a catalog. Both tabs had the same click-to-expand row, so the shell transfers; what does not transfer is the zero state, because there is nothing to discover in your own inventory. With nothing selected the pane reports what is happening: how many models are loaded, what failed, what has an update, and which models are holding VRAM right now. Every number was already on the page. None of them had been assembled into one statement, so "what is going on" was a question the tabs could not answer however long you looked at them. The rail buckets by state rather than capability - Running, Idle, Disabled for models; Update available, Installed for backends - which is the opposite of the galleries and deliberately so: nobody opens Host wondering which of their models does vision. Entries carry a state stripe for the same reason. Load and Stop are promoted out of the kebab, because that is what an operator came for; the rest stays behind the menu rather than diluting it. Adopted, pinned and alias badges follow the model into the pane: they are facts about the thing, not about its state, and the rail line is spent on state. Deliberately NOT done: folding the two tabs into one rail, as the mock had it. It costs five URL parameters, the manage-tab localStorage key and the stat-card shortcuts, all of which are live deep-links today. The tabs stay as the group selector; merging them is a follow-up with its own migration. e2e: full suite 355 passing. New host-split-view spec; alias-template, manage-logs-link, manage-action-menu-position and model-editor-back-nav move off `.table` and the row kebab onto the rail and the pane. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5 [Claude Code] * polish(ui): accessibility and consistency pass over the three split views Findings from a pass over what the previous four commits actually shipped, rather than what they were supposed to. The rail was not a listbox. ARIA lets a listbox contain options and groups, and nothing else, but each group's collapse control is a button that has to sit inside the scroller with the entries it folds. It is now a labelled group of buttons, which is the honest description; selection is announced with aria-current and the arrow keys are unaffected. Every entry was its own tab stop, so tabbing past a forty-entry rail to reach the pane took forty keystrokes. Roving tabindex makes the rail one stop, and arrowing now moves focus with the selection instead of leaving it behind on an entry Tab can no longer reach. The rail rounds its corners with overflow:hidden, which was clipping the focus ring off the first and last entries entirely. Inset outlines fix it. A 30px row is fine under a mouse and too small under a thumb, so coarse pointers get a 44px target without costing density on a desktop. One slot said three different things: "9 models loaded" on Discover, "12 loaded" on Backends, "3 of 9" on Host. All three lists are a page of a larger set, so all three now say so the same way. Also removed: an emptyLabel prop on EntityRail that nothing passed, its dead CSS rule, and MODELS_COLSPAN and ResourceRowDesc, which died with the tables. e2e: full suite 355 passing. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5 [Claude Code] * fix(ui): correct three defects only a real gallery exposed Running the branch against a live instance with 1,595 models and 1,017 backends, rather than against mocked fixtures, surfaced three things the e2e suite could not. Grouping did nothing. The rails matched on the use-case keys the filter chips send (`chat`, `tts`, `transcript`), but those are a server-side vocabulary the handler maps onto entries. What entries actually carry is free-form and inconsistent: models come back tagged `llm`, `gguf`, `vision`, `coding`, and backends `LLM`, `text-to-text`, `audio-transcription`. Nothing matched, so every model landed in "Everything else" and the feature was decorative. Grouping now lives in utils/entityGroups.js, shared by both galleries, matching case-insensitively against the vocabulary the API really uses, with the entry's backend as a fallback signal - a backend named `whisper` is a speech backend whatever its tags say. Order is specific before general and that is load-bearing: a vision model is tagged `llm` too, so testing text first would swallow it. The zero state claimed GPU memory on a machine with no GPU. The resources endpoint reports system RAM in the same field when gpu_count is 0, so the hero read "84.4 GB of GPU memory" next to the recommendations panel correctly saying "No GPU detected". The number was never wrong, only its label; it now says system memory unless a GPU is actually present. The page title still said "Install Models" under a nav entry saying Discover. Also: the keyboard test named the model it expected to arrive at, which made it a hostage of the grouping table and broke the moment the buckets were fixed. It now asserts that the selection moves and returns. e2e: full suite 355 passing. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5 [Claude Code] * fix(ui): the filters and the rail were fighting over the same job Four things you find odd on Discover, and they turn out to be one mistake seen from four sides. The rail grouped the current page. The listing is paginated at nine rows, so those bucket headers described nine entries out of 1,595, and turning a page reshuffled the sections under the reader. The structure was never stable because it was computed over the wrong set. The chips were redundant for the same reason, seen from the other side. They send tag= and filter all 1,595 server-side. The rail grouped nine of them client-side by the same axis. Two controls for one job, and the weaker one was the one this branch added, so it goes. Grouping stays only on Host, where the list is complete, local, and bucketed by state rather than capability. The search bar felt odd because it sat in a full-width band while the thing it narrowed was a 290px rail below and to the left. The whole band now lives in the rail column: search, backend, use cases, refinements, then the list it narrows. One column to say what you want, one to show what you got. Nineteen chips do not fit at that width, so they fold into a disclosure that states the selection. A disclosure and not a popover, deliberately: picking use cases is multi-select and interleaves with the backend select and the toggles below, and a popover dismisses itself the moment you touch either. The header held two counts and two buttons at arm's length from all of it. The counts were the third statement of the same number on one screen, after the rail's "9 of 1,247" and the pane's own headline, so they go. The buttons move into the pane's zero state, which is the surface that answers "what do I do here". Also: the two first-run empty states wore .loading-center, which is display:flex in the default row direction because it exists to centre one spinner. With four children that put the icon, the heading, the sentence and the buttons on a single line with no gap. They are now a proper full-height empty state. e2e: full suite 353 passing. Grouping tests are replaced by ones asserting the rail stays flat; chip tests open the disclosure first; two filter-layout tests that asserted the old three-band arrangement now assert the column. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5 [Claude Code] * polish(ui): make Discover a full-height view, group the chips, name the refinements Four things, all of them the same complaint: the page read as a document with controls scattered on it rather than as one view. The header is fused. A title block with its own padding, a subtitle and two counts made the split view look like an attachment to a document that happened to sit below it. It is now a slim bar carrying the title, the count and the two page-level actions, and the split fills the rest of the window. Rail and pane scroll independently, so the filters and the pane's headline stay put while a long list moves under them. The chips group. Nineteen in a flat row is a lot to scan even behind a disclosure, and they already belong to the four families the rest of the UI speaks, so they are bucketed by those. "All" sits on its own above them without a heading, because it is a reset rather than a use case. The refinements stop looking dumped. When the band became a column they were three controls left where they landed; they now read as a named section with one control per row. The zero state suggests again. It had decayed into a "Browsing / 9 of 1,247 / select a model" line that restated the count for the third time on one screen. It now offers the four use cases as tiles that set the filter, which is the shelf idea from the mock without inventing curation or paying for a second fetch. Two bugs found by looking at it rather than at the tests: the disclosure was clamped to 190px, which cut it off partway through its third section so two of the five never appeared at all; and the creation actions rendered twice, once in the new bar and once in the pane hero a few pixels away. e2e: full suite 353 passing. The chip-row test now holds its contract across the per-family rows rather than a single one, and additionally asserts every family is present and non-empty. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5 [Claude Code] * fix(ui): pin the split view's height so a long detail scrolls the pane Selecting a model with a long description grew the whole page and dragged the rail down with it, which is the opposite of what "full height" was supposed to buy. The flex chain was right and the ceiling was missing. .app-layout and .main-content are min-height:100dvh, which is a floor: flex distributes free space but nothing caps growth, so a pane taller than the viewport expanded the column, the document scrolled, and the rail stretched to match. height:100% on the pane then resolved against an auto-height parent and did nothing. The chat route already solves this by pinning .main-content to 100dvh. The same treatment now applies to any route containing a .page--app, selected with :has() so the shell does not have to learn which pages happen to be split views. Below the stacking breakpoint the pin is lifted, because two stacked halves in two short scrollers is worse than a page that scrolls. Measured on a live instance: document height stays at the viewport across selection (950px either side) and the pane overflows internally instead. Adds discover-height.spec.js, which asserts the page height and the rail height are unchanged by selection and that the pane is the thing that scrolls. The existing specs could not have caught this: they mock short descriptions, and the bug only appears when the pane has more content than the viewport holds. e2e: full suite 355 passing. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5 [Claude Code] * feat(ui): give Backends and Host the full-height view, and fix the Update button Backends now matches Discover: the header fuses into a slim bar carrying the title, the count and the page-level actions, the filters move into the rail column where they narrow the rail and nothing else, and the split fills the window. Its seven chips fit at rail width, so unlike Discover's nineteen they need no disclosure. Host gets the bar and the height; its resource monitor, summary cards and tabs stay above the split, because those are read once while the rail and the pane are worked in. Two things the height change surfaced. The console layout is a flex row with align-items:flex-start, so its body sizes to content. Right for the pages it was built for, wrong for a split view, which needs a ceiling to scroll inside: without it the Backends rail ran past the viewport and over the footer. Pinned with :has() so only split-view routes are affected. The filters vanished when nothing matched. Both galleries swapped the whole shell for an empty state, which took the search box and the chips with it, so the page said "try adjusting your search or filters" while offering neither. The shell now stays and the empty state moves into the pane. Also fixes the Update control on Host, which had no className at all and rendered as bare text, next to a status span that had picked up btn classes and two copies of `fas` and so rendered as a button you cannot press. They have swapped appearances back. e2e: full suite 355 passing. The render-smoke selector learns .view-bar__title, since the pages it checks no longer all use PageHeader. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5 [Claude Code] * fix(ui): keep the view mounted while searching, and bring rail grouping back Searching replaced the whole view with a loader. The search box lives in the rail column, so every debounced refetch unmounted the field being typed into and dropped its focus with it. The list, the filters and the pane went too. The shell now stays and the rail says it is busy: a sweep bar under its header and the stale list dimmed, so the eye knows the answer is being replaced without losing its place. A cold start still gets the skeleton, because there is nothing to keep. The condition for that is "nothing has loaded yet", not "the list is empty". Those differ exactly when someone is editing a query that matched nothing, and getting it wrong there would unmount the view on the keystroke after a no-results search - the worst possible moment. Grouping comes back on both galleries. It was removed because nine rows could not fill five buckets, so a page turn rebuilt the rail's whole structure. That was a symptom of the page size rather than of grouping: the rail now asks for 30 rows instead of 9 (Backends 60 instead of 21), which is enough for the sections to read as structure and turns five times fewer pages. The order of the sections is fixed, so what changes between pages is membership, not arrangement. Grouped while browsing, flat while searching, as before: once a term is typed the buckets stand between the reader and the answer. Also gives GalleryLoader a class and a testid instead of six inline style declarations on a bare div, which is why nothing could select it. e2e: full suite 359 passing, including a new spec asserting the search box keeps its focus and its value across a refetch, and that a cold start still shows the skeleton. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5 [Claude Code] * perf(gallery): stop invalidating the VRAM estimate caches on every request Searching or turning a page felt slow. It was not the search and not the listing: /api/models answers in 3-9ms. It was the VRAM estimate, which the gallery asks for once per row, and which took ~2.3s every single time however often the same model was asked about. pkg/vram already caches what makes that expensive - the remote content-length probes, the GGUF metadata reads and the HF repo sizes. Those caches key on a gallery generation counter, and AvailableGalleryModelsCached triggered a background refresh on every call, with each refresh bumping the counter. One page view is one listing request plus thirty estimate requests, each of which re-read the gallery and started another refresh, so the generation moved constantly and every cache entry was stale before it could ever be read. The caches were dead in production. Three changes, each doing one thing: A refresh interval. The cached list is still served immediately; this only decides how often re-fetching from upstream is worth starting. Five minutes, as a package variable so tests can drive it without waiting. A generation bump only when the gallery actually changed. An unchanged gallery re-fetched on schedule must not throw away work that is still valid, which is the difference between an estimate costing nothing and costing a network round trip. A separate "loaded" flag. The cache engaged on `cached != nil`, so a gallery that legitimately holds nothing read as never-loaded and took the blocking path on every call, bumping the generation each time. Found by the test for the interval, which could not pass while this was true. Measured against a live instance with 1,595 models: one estimate, repeated 2.3s -> 2ms a page of 30, in parallel 10s -> 0.04s A first, genuinely unseen model still costs its remote probe. That is inherent; what changed is that it is now paid once per model per gallery version rather than once per request. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5 [Claude Code] * perf(gallery): warm VRAM estimates at startup, and stop the UI waiting on them Two halves of the same complaint: the gallery stalls on VRAM estimation. Server side, the estimates are now warmed in the background at startup. Estimating an entry nobody has asked about costs a remote probe of its weight files, and the gallery needs one per row, so the first visitor was paying for the whole page. The warm-up walks the gallery in the order the UI lists it, so the first page is ready before anyone reaches it. It is bounded and it never blocks: 300 entries at 4 at a time by default, on its own goroutine, stopping with the server's context. Warming the whole gallery would be thousands of probes on every boot, which is rude to the upstream and slow to finish; warming nothing leaves the first page paying two seconds a row. Anything past the limit still warms itself on first view. LOCALAI_VRAM_WARM_LIMIT=0 turns it off for an air-gapped host, LOCALAI_VRAM_WARM_CONCURRENCY=1 slows it for a metered link. Client side, the page no longer waits on estimates it does not need yet. It fired one request per row at once; a browser allows about six connections per host, so thirty estimates took every slot and the request behind a click - the variant list, an install - queued behind work nobody asked for. That is the freeze: the list was already usable, and the UI was busy fetching sizes. Four at a time leaves room for the interactive request to overtake, and a row whose estimate is still in flight says "sizing…" rather than leaving a blank where a number will appear. buildEstimateInput moves to core/gallery as EstimateInput, since the handler and the warmer both need it. Measured against 1,595 models, from a cold boot: page 1, 30 estimates in parallel 10s -> 0.04s full warm-up (299 of 300 entries) 3m, in the background Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5 [Claude Code] * chore: untrack data/.local_user_id and ignore the runtime data dir `local-ai run` writes its instance state under ./data when started from the repo root, which is exactly what a contributor testing a build does. The identity file ended up committed on this branch by a `git add -A` while verifying the gallery changes against a live instance. Anchored, so it matches the runtime directory at the repo root and not a `data` directory nested inside some package. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5 [Claude Code] --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |