mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-10 15:52:29 -04:00
4046185d554a97bc1dbdb50b6bee6d5cc3edc961
849
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
4046185d55 |
test(traffic): match complete percentile labels
Playwright joins adjacent table-cell text. A llama-cpp backend followed by 99 requests then matches p99, although no percentile is shown. Require word boundaries while keeping the absence assertion. Assisted-by: Codex:gpt-6 |
||
|
|
65623166f0 |
fix(ui): keep model actions visible on phones
Let model descriptions wrap without losing the install action offscreen. Give page table rules precedence over the shared table styles. Separate background facet probes from navigation assertions and wait for asynchronous size sorting and masonry layout. Fix the chat fixture clock so runs near midnight keep the expected conversation groups and order. Assisted-by: Codex:gpt-6 |
||
|
|
8db1d165dd |
feat(ui): add moderated group conversations
Let users give installed models individual turns or ordered rounds in a shared conversation. Attribute completed responses and exclude interrupted output from subsequent prompts. Add streaming and cancellation tests, browser coverage for CI, and a guide for the session-only page. Assisted-by: Codex:GPT-6 |
||
|
|
4d7bdc6ff2 |
feat(ui): redesign the web UI around a shared kit and a calm palette (#12526)
* build(ui): vendor the shared UI kit snapshot at 0.2.0 The restyle needs the kit's tokens, motion layer and component classes. Take a pinned snapshot instead of depending on the kit at build time, and keep a lock file with the version and per-file checksums so a later update shows exactly what changed. The product theme stays outside the vendored directory. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the ink and teal theme and bridge the old variables Define the product colours as the shared UI kit's roles, for light and dark, in theme-localai.css. The kit's contrast check passes on every pair. theme.css keeps the existing --color-* and --shadow-* names but now points each at a role, so App.css and the pages get the new palette without edits. Radii move to the kit scale. index.html now sets data-theme before first paint with the same rule as ThemeContext (stored choice, otherwise dark), because the contract layout of the theme file no longer defaults to dark by itself. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): restyle the shared chrome with the UI kit grammar Adjust the shared classes so every page picks up the same interaction language without per-page edits: - Sidebar sits on the canvas and the current row lifts onto a card. Section labels are tracked uppercase, the badge is a soft pill, and the phone drawer leaves the tab order when closed. - Buttons are flat: hover swaps the surface, press scales to .97, focus is a 2px ring with a 2px offset, danger is a tinted wash. - Inputs use the card surface and the control edge; switches, tabs, filter chips, badges and cards follow the same rules. Cards no longer lift on hover; only linked or button cards react. - Menus and popovers scale in from the trigger corner with 40px items. Dialogs get a veil fade and a spring settle. Toasts become pills at the bottom centre. - The page transition is a 250 ms fade with a 6px rise. It fills backwards so a finished animation no longer leaves a transform that confined dialog veils to the main column. The focus-ring test now checks the outline instead of a box shadow, and new specs cover the theme roles, the first-paint theme and the sidebar lift. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): move leftover hard-coded colours onto the theme roles The YAML editor restated the old blue palette in JavaScript, and a few pages kept literal blues, indigo and violet tints, or fallbacks that only applied because a variable was never defined. Point them at the theme variables so they follow light and dark and the new palette. The status badges in the account pages built their tint by appending "22" to a variable, which is not valid once the variable is defined, so they had no background. Use the wash roles instead. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): raise the type scale and control size toward the kit Body and list text moves to 15px and the rest of the scale follows the kit's 12/13/15/17/21/32/44 steps. Page titles, section headings and stat values are bold with tighter tracking; titles are 32px. Buttons, inputs, selects, tabs and nav rows are 40px high with the 12px radius, compact controls 32px. Tabs become a segmented control. The sidebar widens to 240px (64px collapsed) and nav rows get more room. Identifiers and counts in the split views use the mono face, and the stat grid becomes separate inset tiles. The Geist stack stays: it is bundled, and the thin look came from the size, weight and negative tracking, not the face. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): separate cards, panes and floating surfaces from the canvas Cards, the Models and Installed split panes, the composers and the confirm dialog use a stronger card edge, the rest shadow and the 20px radius, so they read as layers in dark as well as light. Menus and popovers move to a float surface (the hover tone in dark) with the float shadow. The selected rail row gets an accent wash and a 3px accent edge. The send buttons are a clear accent when there is something to send and a quiet inset when not; the Home button carries data-empty for that, since submitting an empty box does nothing. The assistant card becomes an accent wash with a square icon. New surfaces spec checks the pane edge, the selected row, both send buttons and the popover in both themes. The voice library empty-state spec now waits for the layout to settle before comparing two boxes. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): tidy the sidebar header and mark the current row with a dot The header gives the configured horizontal logo a fixed width and centres it in a 72px band, lined up with the nav icons. The collapsed rail shows the configured icon logo centred, and its nav rows become 40px tiles centred in the 64px rail. The current row gets the kit's accent dot, hidden in the rail. The theme, language and account controls stay in the sidebar footer: the app has no global search or command palette to put in a top bar, so a bar would only hold controls that already have a place. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): centre the avatar in the collapsed sidebar rail The collapsed avatar link was set to "flex: 0", which gives it a zero flex basis; with min-width: 0 the link shrank to its padding and the icon overflowed from the link's left edge, about 14px right of the icon column. Use "flex: 0 0 auto" in the collapsed and tablet rail. The footer controls now share the nav icon column in the expanded sidebar too (6px footer padding, 40px control boxes), and the tablet rail gets the same footer padding and hidden language code as the collapsed one. New spec measures the centre x of the nav icons, mark, avatar, language, theme and collapse icons in the collapsed, expanded and tablet states, in both themes, and asserts they agree within 1px. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): size the console and settings rails and stack settings on phones The Operate console rail, the Settings section rail and the account tab bar still used 13px text and the old underline tabs. They now use the 15px nav size, 40px rows and the segmented tab control. Form row labels are 15px with 13px hints. On a phone the Settings section rail sat beside the form and squeezed every row into a few characters. Below 720px the rail stacks above the content as a scrolling row and form rows wrap their control below the label. The save button no longer carries the icon font class, which drew a missing glyph before its label. The language menu is wide enough to keep Bahasa Indonesia on one line. Assisted-by: Claude Code:claude-sonnet-5-5 * build(ui): update the vendored UI kit snapshot to 0.3.0 Take the 0.3.0 snapshot: the sprite now carries the full outline icon set, and the new icons/fa-map.json maps Font Awesome names to icon ids. The map lets the app move off Font Awesome in the following commits. The lock file is regenerated with the new checksums. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add an Icon component backed by the kit sprite Icon draws an inline svg that points into the kit's outline sprite. The sprite is inlined into the page once, so the references resolve under any base path and in the embedded build without a request. Icons size with the font (1em), take currentColor, hide from assistive tech unless given a title, and spin on request. An unknown id draws a neutral circle. FaIcon and iconFromFa resolve Font Awesome names through the kit's map, for names that arrive at run time. iconHtml does the same for markup built as a string. The GitHub and Apple marks are small local glyphs, as the kit ships no brand marks. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): draw shared components and helpers with Icon Replace the Font Awesome elements in the shared components and in the utility modules with the Icon component. Lookup tables now hold kit icon ids instead of class strings. Code-block copy buttons and artifact cards, which build HTML strings, use iconHtml and a sanitizer-safe slot. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): draw model, backend and account pages with Icon Replace the Font Awesome elements on the home, models, backends, import, settings, login, account and users pages with the Icon component. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): draw chat, studio and recognition pages with Icon Replace the Font Awesome elements on the chat, media generation, talk and face and voice pages with the Icon component. The talk status table keeps its spin and pulse states as Icon props. The connected and error states now use a dotted circle and an alert circle, so they differ from the idle ring by shape as well as by colour. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): draw agent, node and operate pages with Icon Replace the Font Awesome elements on the agents, skills, collections, jobs, fine-tune, quantize, nodes, swarm, usage, traces and activity pages with the Icon component. Two class strings on layout elements held leftover button and icon classes from an earlier merge; they are cleaned up so the elements keep only their own classes. Assisted-by: Claude Code:claude-sonnet-5-5 * style(ui): size and align icons for the svg component Icon rules that targeted the font element now target the svg: the descendant "i" selectors in App.css and auth.css become ".lai-icon". The svg is 1.2em with a 2 unit line so it matches the visual size of the old glyphs at the 12 to 16px sizes the app uses, sits on the text baseline, and follows the context font size. Large empty-state marks get a lighter line. Menu icons get a 16px box and the readiness badge icons keep their 20px circle with padding. Add the pulse used by the talk status. Assisted-by: Claude Code:claude-sonnet-5-5 * build(ui): remove Font Awesome No source file references the icon font any more. Drop the package and its stylesheet import. The build no longer ships the solid, regular and brand font files. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): keep focus traps off the svg use references The dialog and drawer focus traps collect focusable elements with a "[href]" selector. An icon's use element carries an href, so it became the first "focusable" element and Tab at the end of the dialog stopped there instead of wrapping to the first button. Match "a[href]" instead. Assisted-by: Claude Code:claude-sonnet-5-5 * style(ui): keep icon sizes overridable and set the line width per svg Give the icon base rule zero specificity so a rule that sizes one icon (nav column, menu box, avatar, language switcher) wins whatever its order in the file. The sprite symbols fix their own line width; the inlined copy drops it so the width set on each svg applies, as the --lai-stroke custom property, and large marks can use a lighter line. Pin the avatar and the language globe to the boxes the sidebar alignment spec expects. Import the map as JSON with an import attribute so Node can load it in the spec. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): select icons by the svg markup and cover the sprite Specs that found icons by their Font Awesome class now select the svg by its data-icon. The dead-icon audit checks that every svg resolves to a sprite symbol and has a size. The class hygiene spec fails on any remaining Font Awesome class. A new spec checks every mapped icon id has a symbol, that the sprite is inlined once, that an icon paints at the root and under a forwarded path prefix, and that Font Awesome names map as documented. Assisted-by: Claude Code:claude-sonnet-5-5 * build(ui): update the vendored UI kit snapshot to 0.4.0 Take the 0.4.0 snapshot: hub tabs with count and attention badges, the six chart series tokens and the grid colour in the theme contract, and sample themes on a calmer palette. The kit headers are renamed and the lock file is regenerated with the new checksums, as for the earlier snapshots. Assisted-by: Claude Code:claude-sonnet-5-5 * style(ui): switch the theme to the calm palette Rewrite the LocalAI theme on the calm palette: a muted teal accent on a near-neutral green-grey canvas, desaturated status colours, no glow and no coloured shadows. The theme fills every role of the shared UI kit's 0.4.0 theme contract for light and dark, including the six chart series and the grid line. The bridge in theme.css keeps the old --color-* names working, adds the dark surface ladder (card, raised, float) and a strong edge, and points the fixed data hues at the chart series. Two values differ from the first sketch. The dark text on the accent fill is #021512 instead of #04201d: it reads 5.58:1 on the fill at rest and 6.4:1 on the hover fill, against 5.08:1 at rest for the lighter value. The light control edge is #6b7d7a. The kit's contrast script passes for all text pairs (4.5:1), control and focus pairs (3:1) and series colours (3:1). Leftovers that no longer fit the palette are fixed: the usage chart takes the six series colours in order, the audio and animation canvases fall back to the new accent, the face box loses its glow, and two gradient fills are now flat. The theme tests expect the new canvas colours. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): replace the console rail with hub tab bars Build and Operate no longer open a second navigation rail beside the page. Each is a hub: one row of the kit's hub tabs above the page, with count and attention badges that scroll sideways on a phone. Every URL and route stays as it was, plus a new /app/build landing page that lists the Build tools with a line each. Build tabs: Overview, Agents, Skills, Memory, Jobs, Fine-Tune, Quantize, Import, Voices (recognition and library) and Faces. Operate tabs: Status, This machine, Swarm (distributed mode only), Runtime (backends, activity, failover), Traffic (usage, traces, middleware) and Settings (settings, users), plus the API link. A tab that holds several pages shows a second row of links, and a sub-page such as a node detail keeps its tab highlighted. The feature and admin gates decide which tabs are drawn, and badges show only values the Operate summary already has. The sidebar lists Build and Operate under a Workspace label next to the Create group. The voice library moves under Build and the model import page gains the Build tab bar. The old rail styles, the rail signals and the console config are removed, and the Operate overview docs describe the tab bar. The specs that drove the rail now drive the tabs, and a new spec covers the tab for each route, gating, badges and the phone layout. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Home as a calm console Home now opens on one command bar: the model chip shows which models are warm, the MCP chip and attach buttons sit beside it, and Send is a solid button with an Enter glyph. Typing "/" opens a grouped, keyboard-driven action list built on the kit command list; every action has a destination in the product. Memory use folds into a one-line strip that opens into the loaded models, with Stop per model and Stop all. It opens by itself while a model is being staged and after a failure, and shows nodes and aggregate memory in a cluster. The list of resident models carries no per-model size because the API reports none. "Jump back in" lists the conversations stored in the browser, one card per day, with j and k to move, Enter to resume and delete with an undo toast. First run keeps the install steps and the recommended models. The assistant prompt is a dismissible line, the library links are one quiet row and the API section is collapsed. Chat accepts an empty new-chat hand-off for /new. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Home console Update the Home specs for the new structure and add specs for the slash menu, the model chip, the memory strip (expand, stop, staging, failure, cluster), the resume list (grouping, j/k, Enter, delete with undo), first run, the send hand-off, a non-admin user and the phone layout. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the fit, disk and cleanup helpers for the models page Pure functions and hooks that the rebuilt Models page reads, with node tests for the rules. modelLedger turns an estimate and the memory budget into one of three verdicts (fits, spills to CPU, over) with the headroom in bytes, and reads the models disk from the resources reading. The disk counts as low under 10 percent or under 20 GB free, and is absent when the server reports none or runs as a cluster controller. cleanupPlan ranks installed models from what the API reports: loaded, pinned, or named by an agent, a task, a failover chain or an alias keeps a model protected; another installed build of the same gallery model is a duplicate; disabled models rank above idle ones. The API records no last use or use count, so none is used. When a lookup fails, nothing is called safe. useModelRemoval holds a removal in the browser for an undo window and sends the existing delete call only when the window ends. Leaving the page drops the batch without deleting anything. The undo toast takes optional labels so other pages can reuse it. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Models as a ledger with a disk strip and cleanup review Explore is one dense table. Each row carries the size, a solid memory bar and the headroom in words ("3.7 free", "+1.5 on CPU", "0.9 over"), worked out from the estimate at the chosen context length. Capability chips show the server's count for each facet, search keeps its meaning and "/" jumps to it, and a density switch (also "d") picks comfortable or compact rows. Selection is a surface step and a check, never a rail. Arrow keys move, Enter installs and Esc closes the inspector, which keeps the fit summary, VRAM by context chart, variants, files, links, tags and licence. A failed install shows its error in the row with a Retry that dismisses the old failure first. A failed or empty listing says which it is, and a host with no GPU is measured against memory and says so. Installed uses the same table with state filters that carry counts, a state per row, Load or Stop on the row, the row menu and the sort by size. Sizes come from the files the gallery lists, so a model it does not know shows a dash. A strip in the header shows the free space on the models disk. It turns amber under 10 percent or under 20 GB free, hides when the server reports no disk or runs as a cluster controller, and opens the cleanup review. Explore says how much an install leaves free. The review ranks installed models as Safe to remove, Probably safe and Your call from real facts only, lists protected models with the reason, and says plainly that usage history is not recorded. A sticky bar shows what a choice frees. Confirming runs a dry run that checks again and lists what will go. Removal waits 30 seconds with an undo; nothing is deleted before that, and leaving the page deletes nothing. The old rail, filter band and popover styles are removed. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Models ledger, Installed table and cleanup review Update the Models, lifecycle, cluster fit, height, search focus and surfaces specs for the table and inspector, keeping what each one checks. New specs, on a shared 41-model gallery stub with three machine profiles: the fit bar and headroom words for a 24 GB card, an 8 GB laptop and a host with no GPU; facet counts, search, "/" and Escape; selection, arrow keys, Enter to install, density; the disk strip when normal, low and hidden; and the states (loading, empty, offline, install failed, phone). Installed covers filters with counts, row actions, the row menu, sizes and sort. The cleanup specs cover grouping, protected models, the honest-data note, the effect bar, the dry run, the undo window, a failed delete, leaving the page, and the phone sheet. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the placement helpers and the estimate hooks The Placement section and the model page need the same few rules, so they sit in plain functions that can be read and tested alone. placement.js holds what gpu_layers, tensor_split and main_gpu mean (unset asks for every layer and the llama.cpp engine trims it, zero is CPU only, 99999999 is the value LocalAI itself writes for all layers), the device list taken from the resources reading, the split by free memory, the part of an estimate that grows with context (read from two lengths, since that term is linear), the fit states with their limit (95 percent of free memory, and the leftover has to fit in system memory too), and a bisection for the largest layer count whose estimate fits. The estimate returns one total and no layer count, so the search runs over 1 to 256 and stops at the first count that no longer changes it. modelWalk.js keeps the order of the list a model page was opened from, in memory and in session storage, for the previous and next buttons. usePlacementEstimate reads /api/models/vram-estimate for a choice, again at twice the context, and with every layer, and keeps readings for the session. useModelPage reads a gallery entry by name, an estimate by context size (from the model's own files when the gallery does not list it), the builds and the loaded models. usePlacementConfig edits the four placement keys of an installed model and saves only what changed. useModelActions is the Load, Stop, disable, pin and remove logic of the Installed table, shared with the model page. MemoryBar is one solid bar with a tick at the capacity of its pool; over capacity it grows past the tick and the tick turns red. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the Placement section to the model editor Run this model on: CPU only (gpu_layers: 0), Auto (the key stays unset) or Custom. Custom takes a number, has an All layers button that writes 99999999, and shows a slider only when the estimate reports the model's layer count, which it does not today. Context size has presets and a number field because the KV cache follows it. With two or more GPUs there is a split (written as percentages, with a button that takes them from the free memory of each card) and a main GPU. A bar per GPU and one for system memory show what other programs use, the model's weights and working memory, and the part that grows with context, with the room left or how far over it is. Under them a verdict in plain words: Fits in GPU, Spills to CPU, Too many layers for the GPU, Runs on CPU only, No GPU found, Not enough memory. It says "slower" and never a multiplier, because the estimate has none. Fit it for me asks the estimate for the largest layer count that fits the free GPU memory and says what it set, with Undo; it is hidden when the estimate is unavailable or the host has no GPU. Loading shows skeletons, an unavailable estimate shows a note with Retry, and a server that schedules onto other machines shows no bars, because its device list is the controller's. The editor shows the section for an installed model, with a link in its section rail. Auto sends null for the key, since a patch only merges, and a null read back opens as Auto. The docs describe the section and what each mode writes. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): open a model on its own page A model has an address, /app/models/<name>, for an installed model and a gallery entry alike. Open it from the arrow at the end of a row, a double click, "o" on the selected row, the inspector's Open details button, or a tap on a phone. The title block holds the main action: Install with a chevron that chooses the build, or Load and Stop with a menu (disable, pin, edit configuration, logs, delete with a confirm). A strip answers whether it fits, what it does and what installing leaves free. Tabs: Overview (about, a memory bar, state, the pages it opens in, and the agents, tasks, chains and aliases that name it); Fit and memory (verdict, context sizes, the bar split into weights and context, and memory by context against the limit, with a data table); Variants and files (builds with size and fit, install any, the files of the chosen build). For an installed model also Usage and history, which says what the API does not record instead of drawing an empty chart, Configuration, which is the Placement section with the file it writes and a link to the full editor, and Logs, the backend log viewer without its page. Keys 1 to 6 switch tabs, [ ] and j k walk the list the page was opened from, Esc or Backspace go back. The list stays mounted behind the page, so Back finds its view, search, filters, selection and scroll as they were, and focus returns to the row's arrow. The page covers loading, an unknown name with the closest matches, the gallery being out of reach, an install in progress with Cancel, and a failed install with Retry. The docs describe the page and its keys. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the model page and the Placement section New specs for the model page: reaching it from Explore, Installed, a double click, "o", a pasted link and a phone tap; the walker and Back with the search, a filter, the selection, the Installed view and the scroll kept, and no second read of the gallery; the title block, the answer strip, tabs by click, keys and arrows; Fit and memory, builds and files with the install call each one makes; an installed model's actions, used-by, the honest usage tab, configuration and logs; loading, an unknown name, offline, an install in flight and a failed one; and the phone. New specs for Placement: every mode and the keys it writes, the slider only when a layer count exists, the context presets, the bars and every verdict, two GPUs, no GPU, a cluster, an unread machine, a loading and an unavailable estimate, Fit it for me and Undo, and the section in the model editor with its save. The phone tap on a row now opens the page, so the two phone specs that expected the inspector as the page check the page and keep the inspector check for a window between a phone and a desk. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): link Studio results and open workspaces from a prompt Each workspace now records the result it was made from (parentId and an edge kind such as take, animate or to-3d) and reads a prompt, model, size, count and source from the query string, so one page can hand work to another. A source result is fetched from the server's own output file and becomes the start image, the picture for 3D, or the audio file. A note on the page says when the source loaded or could not be loaded. Diarization had no history; it now keeps the file name, the model and a speaker count, never the recording. Prompts are cut at 2000 characters when stored. The pure helpers (type suggestion, grouping, lineage layout, favourites, clearing) have node tests. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Studio front page as a composer with your work The front page is a prompt box with a chip per type, a type suggestion from the words, starters, and the options each workspace accepts. Generate opens the workspace with those filled in. A type with no model is a dashed chip that shows a gallery model, its size, memory need and an Install button only when picked; the typed words stay while it installs. Under it, Your work lists results from every workspace as a masonry with filters, counts, favourites and a Clear history action. Results made from each other stack into a project tile and open as a lineage board with a dock for running a new take or branching to the next step; steps the destination cannot start from yet are disabled with the reason. The docs describe the page, what is stored in the browser, and the query parameters a workspace accepts. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Studio composer, your work and the lineage view Specs for the type suggestion, the keys, hand-off to each workspace, the install path for a missing model, the masonry filters, favourites and clearing, stacking, the lineage board, new take and branch, steps that are disabled with a reason, and the phone layout. Existing Studio specs move from lanes to chips with the same intent. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the shared Studio workspace frame and move Images onto it The seven Studio workspaces get one layout: a row of type tabs, a compose card (optional sources as chips, a prompt with starters, a model chip, the essential options as chips, an Advanced fold that names what is inside, the memory the model needs, and one action with the reason when it cannot run), a run area, and a strip of recent results of the type. The run area shows a job card with the time that has passed and an indeterminate bar, because these endpoints report no phase or percentage; a failure with what the server said and one action; or the result with a toolbar: Favourite (the list the front page keeps), Download, Use in (the hand-off targets, disabled with the reason when a destination cannot start from the result), Re-run with edits (the take's values go back in the form, changed fields are outlined and listed) and Lineage. A type with no model shows the install note from the front page. Images is the first workspace on the frame. It keeps its size, count, steps, seed, negative prompt, source image and reference images, and its history writes, including the parent link and edge of a hand-off run. useMediaHistory.addEntry now returns the id of the entry it stored. The docs describe the workspace page. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move Video onto the workspace frame Video keeps its size list, duration, frame rate, steps, seed, CFG scale, frame count, negative prompt, start and end image and avatar audio. The start and end image are source chips, the avatar audio opens the recording and paste input from a chip, and the rest sit in the Advanced fold. A start image from a hand-off shows as a chip with its picture. Results play in the video player with the shared toolbar. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move TTS onto the workspace frame TTS keeps the saved-voice picker for cloning models, the typed voice for the others, the voice library deep link, and the delivery instructions, which now sit in the Advanced fold. The result is the waveform player with the words under it. The stored entry also keeps the voice id so Re-run with edits can select the same saved voice. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move Sound onto the workspace frame Sound keeps its Simple and Advanced modes and every field of both: the description, instrumental, vocal language, caption, lyrics, BPM, duration, key, language, time signature and think mode. The mode switch, instrumental and duration are in the compose card, the rest in a More options fold. The stored entry keeps all of the fields, so Re-run with edits restores the form as it was. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move Transform onto the workspace frame Transform keeps its audio and reference inputs with upload and record, the echo test, the key=value parameters (now in the Advanced fold), the input and output spectra and the three waveform players. The audio that was chosen shows before the run, waiting to be transformed. Re-run with edits puts back the model and parameters and fetches the audio and reference the server kept for that run. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move 3D onto the workspace frame 3D keeps the picture input with paste and webcam, the animation operations a model declares, quality and background, the shape and material steps, guidance and seed, the GLB and animation viewers, the remesh control and the download. A 3D result now has a title from the motion prompt when it has no label, so the strip and the front page name animation results by what was asked. Re-run with edits is shown disabled with the reason, because only a small thumbnail of the picture is kept. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move Diarization onto the workspace frame Diarization keeps its model and recording inputs, the option to prepare speakers to remember, the clean-speech previews, naming and remembering a speaker, and the history entry with only the file name, model and counts. The result now shows a timeline with one lane per speaker, the talk time of each speaker, and the segments with their start time and text. RTTM, SRT (only when the run has text) and JSON are built in the browser from the result. The helpers for talk time, axis ticks and the two text formats have node tests. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): remove the styles and lists the old workspace layout used Nothing renders the two-column workbench, the control column, the old history lists, the generation progress tiles, the TTS voice picker or the result echo any more. The inline-style baseline drops with them. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the workspace frame and one run per type Specs for the type tabs, the compose card and its reason when Generate cannot run, starters, the Advanced fold, the job card with no invented progress, a failed run and its one action, the install note, the strip with its favourites filter, Use in with its disabled steps, Lineage, the parent link, Re-run with edits and its list of changes, deleting and clearing, and the hand-off note. One run through each of Video, TTS, Sound, Transform, 3D and Diarization, the phone layout of all seven, and reduced motion. Existing Studio specs move from the old control column to the compose card with the same intent. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the chat thread with raised user turns, prose replies and one-line activity Your messages are raised blocks on the right at a 760 px measure and the model's replies are plain prose under its name and a warm or not loaded dot. Reasoning, tool calls and their results fold into one quiet line that opens inline into steps. Code blocks carry a Copy button and a Canvas button that opens that block in the canvas, image attachments are thumbnails that open in the lightbox, and files are chips. Per-message actions show on hover, on focus and on the last turn, and a turn takes focus so the arrow keys and C, E, R and B work. A failed reply keeps the text written so far and shows the reason with one Retry action. The Agent chat page keeps the older rules: the new styles are scoped to the chat page and use their own class names. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): use the Home command bar as the Chat composer Chat now ends in the same object as Home: the model chip, the MCP chip, a Canvas chip, the message box with attach buttons, a solid Send and the hint line, with the slash menu on the kit command list. The slash menu lists what Chat can do today (switch model, new chat, conversations, manage mode, canvas, find, settings, export, clear). While a reply is streaming Send becomes Stop, which Esc also presses, and Up in an empty box edits your last message. Attached images show as thumbnails and a line under the bar carries the speed and the token count. HomeComposer takes optional props for this (extra chips, its own slash list, Stop, paste, a stricter Enter); Home passes none of them. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): open conversations from a Ctrl K menu with day groups, undo and a slim header The conversations list opens as a centred menu on Ctrl or Cmd K. It groups chats by day like the Home resume list, shows the model that answered and the time, searches names and message text, and moves with the arrow keys. Enter opens a chat, F2 renames it and Delete removes it. Removing a chat hides the row and shows the kit undo toast; the chat is deleted for good only when the undo time ends. Rename, duplicate, copy and export are on each row, as before. The header is one slim bar: the Chats button, the chat name (click to rename), a context meter when the context size is known, settings and a More menu with rename, duplicate, copy, export, model info, keyboard shortcuts and clear. A dialog lists the shortcuts the page answers to. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): show loaded state, capabilities and fit in the Chat model switcher The model chip in Chat opens the same list as Home, grouped as Loaded now and Installed. Each row says warm or not loaded and marks models that understand images. When the list opens, the page reads the host memory once and asks the server to estimate each listed model at the chat's context size (up to twelve, three at a time), then shows what the model needs and whether it fits: free memory, how much would run on the CPU, or how far over the machine it is. A model with no estimate shows no fit text, and no load time is shown because the API does not report one. A memory bar closes the list. The picker takes the model list from the page when it has one, and useModels can skip its own request. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move chat settings into a sheet and add find in chat, jump to latest and a wider canvas Settings open as a kit sheet: the system prompt, temperature, top P and top K (each says "model default" until it is changed and has a Reset), the context size with quick sizes and a note that it only drives the meter, Manage mode and Focus mode, the model info for admins with its Edit config button, and Clear conversation behind a confirmation. The old slide-out drawer and the model info panel are gone. Ctrl or Cmd Shift F (or the search button, or /find) opens a search bar over the thread. It marks matches in the messages already on the page, shows "n of m" and steps with Enter and Shift+Enter. Nothing is sent to the server. Jump to latest is a pill above the composer. Esc stops a reply, then closes the search, then closes the canvas. The canvas panel gets the kit look: tabs, a Code and Preview switch, Copy and Download, a full-page layout on narrow windows, and translated labels. The Agent chat page shares it and gets the same look. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the empty, no-model, loading and phone states to Chat An empty chat opens with the composer under one line, starters to try, whether the model is loaded, and the Jump back in list: the same rows as Home, read from the chats the page holds. With no chat model installed, an install card offers the starter models for this hardware, the gallery and import, and the composer stays so the text is not lost. While a reply waits for a model, a load card shows what the page knows: the phase the server names, the node, the bytes and the time left when the server reports them, and a progress bar. A model that is just not loaded yet gets a plain note, with no invented phases or estimates. The foot warns when the context is nearly full. On a phone the header drops its labels, the model list and the settings open as sheets from the bottom, per-message actions stay in view and the canvas takes the whole page. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Talk as a calm voice page over what the connection really does Talk is one stage and one transcript. The stage has the pipeline chip, the voice and language chips, an outline orb that follows the real microphone and playback levels, a heading and a sentence for the current state, and the controls. The transcript lists You, Reply, Tool and Result lines and can be copied. Session settings (instructions, voice, language, tools, Manage mode and the pipeline's parts) open in a sheet. The states are the ones the code reaches: no pipeline model, idle, connecting, listening, thinking (also while a tool runs), speaking, an interrupted reply (the server cancelled it; a note marks the cut), a blocked microphone, a link that failed during a session, and any other error with its reason and a link to the traces. Push to talk and hands-free are not on the page, so they are not shown. Diagnostics keep their waveform, spectrum and stats, drawn in theme colours. The page text moves into the talk namespace, and the old Talk and visualizer styles and the inline-style count go down with the rebuild. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): remove the chat styles and strings the rebuilt page replaced The settings drawer, the model info panel, the bubble avatars, the conversation menu popover, the context bar, the recent strip, the staging bar, the file badges and the focus-mode rules have no user now. Their rules, the Chat page's focus class and seven unused empty-state strings are removed. The Agent chat page keeps the shared message, sidebar and input rules it still renders with. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the rebuilt Chat and Talk pages Add a Chat page under features (thread, message actions and keys, the message box and its slash actions, the model list with loaded state and fit, conversations on Ctrl K, settings, find, canvas and the empty, no-model and loading states) and a Talk section to the realtime API page with the states the page shows. Manage mode now turns on from the chat settings or /assistant, and the client MCP steps point at the MCP chip. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): settle the rough edges of the new Chat and Talk pages The undo toast sat under the conversations menu, so the Undo button could not be pressed while the menu was open; the menu, the sheets and the fullscreen canvas now stay below the toast layer. Esc in a rename box saved the text through the blur that follows it; it now cancels. The image viewer closed on Esc only when the page did not re-render on the same key, so its key listener is registered once and reads the latest handlers. Keys on a focused message no longer type their letter into the editor they open, "/" from outside a text field starts a command as it does on Home, and Esc leaves the page's own dialogs alone. Code in the canvas is highlighted for languages that have no preview. The conversations menu drops its key hints on a phone so Clear all stays in view. Talk hides Test tone while connecting and calls a server error "Something went wrong", since the call can still be open. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the rebuilt Chat and Talk pages Specs for the thread layout and the activity fold, code blocks, image thumbnails and the viewer, per-message actions and their keys, a failed reply with its one Retry, Stop and Esc while streaming, the composer and every slash action, the conversations menu (groups, search, resume, rename, delete with undo that ends by itself, one chat left), the model switcher with loaded state, vision and fit text from stubbed estimates, the settings sheet, the canvas panel, find in chat, Jump to latest, the empty, no-model and loading states, the phone layout and reduced motion. Talk is driven over a fake WebRTC link through idle, connecting, listening, thinking, speaking, interrupted, blocked, lost, error and no pipeline, its settings sheet and its phone layout. Node tests cover the message text helpers and the conversation grouping. The existing chat specs move to the new structure with the same intent: the transcript spec now describes the raised turn and the prose reply, and the render smoke accepts Talk's own header. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): keep a bounded run log for agents in the browser The server keeps no run history for an agent, so a run is one task and the events until the agent answers, written to browser storage while the page watches the stream: up to 50 runs per agent, task, step and answer text only. Stored chats from the earlier agent chat page read as runs with stable ids. A run still marked running five minutes after its last event reads as stopped. Helpers read an agent's config into chips, build the list of changed fields against the saved config, hide secret values and offer starting points. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Agents area around runs The Agents page shows what needs a look (work in flight, a run that failed in the last day), then each agent with its model, attached memory and skills, and a strip of its last 14 runs. An agent has its own page: model, tools, memory, skills, instructions, a task box and its runs. A run has an address, shows the thread while it works (steps folded into one line, the tool in use, the answer as it arrives) and settles into a report about a second and a half after the agent answers: task, outcome, follow-ups, evidence and steps, with wide tables opening wider on demand. A failure says in plain words what happened and offers Run again. Create and edit fold into sections with a ready mark and a one-line summary, start from a template or an optional model-written draft, and open a preview sheet with the config as saved and the changes against the saved agent. Status becomes a quiet panel in the same language, and the old chat link opens the agent page. There is no Stop, approval, steer, version or dry-run control, because the agent API has no call behind them. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Agents launcher, agent page, runs and editor Specs for the Now strip and run strip, search and the empty state, the agent page, starting a run, the live thread, settling into the report, the run address across a reload and for a run from another browser, follow-ups with their history, failures, the folding editor with ready marks, templates, the preview sheet with hidden secrets and changes, the status page, and the phone, 1440 and 2560 layouts. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe runs and the new agent create flow Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add helpers for tasks, schedules and job outcomes Reads a cron expression the way the server does (five fields or an @ shortcut), checks it, and puts the common shapes in words. The next run is left out on purpose, because the schedule follows the server clock, which the browser cannot read. Also groups jobs by day, sums the last seven days, and gives each job one outcome line from its result or error. A rerun call starts a new job with the same parameters and media. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Jobs area around runs The Jobs page opens with one sentence about the last seven days, then the tasks (model, schedule in words, last 14 jobs, enabled switch, Run now) and a run history grouped by day. Each row has an outcome sentence and opens to the error or the start of the result with one next action. Deleting a task waits 30 seconds with an undo button. A task opens as a page with its recent runs, its prompt with the gaps marked and its schedule. The task form folds into sections, takes a schedule as a preset or a checked cron expression, warns about prompt gaps the schedule does not fill, and has a preview sheet. A job opens as a document: task, outcome, delivery and the recorded steps; a failed job says what happened and offers Run again. Run now now sends attached media through the job call, which is the only one that takes it. "Clear History" only ever cancelled running jobs, so it is now called Stop running jobs. Webhook headers of a saved task show as JSON instead of [object Object]. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Jobs page, task pages and job pages Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the Jobs page and the task form Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add helpers that say who uses a skill or a collection An agent loads a skill when skills are on and the skill is in its selection (an empty selection means every skill). It reads the one collection that carries its own name, when its knowledge base is on. The helpers derive that from the saved agent configs, build the config that adds or removes a skill or a collection, and estimate tokens as characters divided by four. Removing the last selected skill switches skills off, because an empty selection would mean every skill. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Skills and Memory as one library Skills and collections sit in a list with an open item beside it. Each row says who uses it, read from the saved agent configs, or says it is not used yet. Chat reads neither, so it is never named. An item opens in a pane with a Used by strip (names link to the agent, a small x removes it, with undo) and an Add to menu that shows what the addition costs. A collection can be added only to the agent that carries its name. The Memory pane searches the collection alone and shows ranked passages with scores, lists web sources with their refresh interval and the files, shows the server message when an upload fails, and names the endpoints and where files stay. The Simulate a message sheet runs a collection search and shows an agent's skills with a token estimate. It runs no model. The collection details route now opens the same page. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): show use and cost in the agent form pickers Each skill in the agent form says which other agents use it and what it adds to every message, with a total for the selection. The memory section names the collection the agent reads. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Skills and Memory libraries Specs for the used-by lines (including an agent that uses every skill), the filters, search, add to agent, remove with undo, the last-skill case, an unreadable agent list, the empty states, git repositories, the Memory question box, sources, uploads that fail, the Simulate sheet with the parts the API can run, the agent form hints and the phone layout. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the Skills and Memory libraries Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add helpers for the Operate status page and backend rows Pure functions for the parts that need rules. They work out the memory pool the page measures (GPU memory, system memory, or the workers of a cluster that are answering), which pools are too full, the headline and the four ledger rows, the geometry of the capacity chart, and what removing a backend would leave without a runtime (models name their backend, and a meta backend names the concrete one it points at). A second set says what a backend row states: installing, queued, removing, failed, update available, current or absent. LocalAI keeps no memory history, so the chart reads a bounded buffer of readings the page took itself and says so. A reading with no total is dropped rather than drawn as zero. Two hooks are shared by the pages that need them. One retries a failed operation after moving the failure into the record. The other holds a cancel for an undo window, because the server cannot take a cancel back. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Operate Status, This machine, Backends, Activity and Logs Status opens with one sentence ("2 things need you", or "Everything is running") and four rows: Needs you, Capacity, Running now and Recent failures. A row with a problem opens by itself and holds the button that deals with it: Update a backend, Retry or Dismiss a failed operation, Unload a model. A quiet row stays one line. A new installation gets a first-run screen, a cluster sums the memory of the workers that are answering, and a page still waiting for an answer says so. The chart under the rows is drawn from readings the page took while it was open and is labelled that way, because LocalAI keeps no memory history. This machine leads with GPU memory as one bar, then host memory split by running model, then VRAM, RAM, CPU and disk with a bar each. The running models become a kit table with the same menu and stop dialog. Backends is one list with Installed and Catalog views. A row says what the backend is doing (a progress bar with Cancel, Queued, Failed with Retry, Update 1.2.0, Current), carries the one button that matters, and opens in place. Removing a backend names the models and the meta backends that would stop working. Check for updates, Update all, From URL and a first-run recommendation for llama-cpp are in the header. Activity keeps its three sections as quiet rows. Cancel waits eight seconds with an undo toast, because the server cannot take a cancel back; a cancelled install can be started again from the record. Logs gets a process list, a picker, stream and text filters, Follow and Times switches, and a Clear with an undo window. Not shown, because the API has no data for them: GPU temperature and power, a size per backend, an earlier version to roll back to, a dependency lookup beyond the models and meta backends that name a backend, and models that failed to load. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover Operate Status, This machine, Backends, Activity and Logs New specs for the Status headline and ledger (healthy, needs attention, one thing, a full memory pool alone, loading, first run, cluster), its actions (Update, Retry, Dismiss, Unload with its dialog), the capacity chart built from readings taken while the page is open and bounded, the phone layout, no coloured edge on a row, and reduced motion. The Backends specs cover the two views, install progress with Cancel and its undo window, Retry on a failed install, Update, Update all, Check for updates, removal with the models and meta backends it would break, Install from URL, the first-run recommendation, a cluster, and a phone. Activity gains cancel with undo, Cancel now, a second cancel, leaving the page, progress, and starting a cancelled install again. Logs covers the stream and text filters, Follow, Times, Export, Clear with undo, the process picker and list. This machine covers the GPU strip, several GPUs, no GPU and Add a machine. Existing specs keep their intent and follow the new structure: rows open in place instead of in a pane, Update replaces Upgrade, the notice spec now pins that an update is a row state and not a banner or a rail, and a cancel waits for its undo window. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the rebuilt Operate Status, Backends and Activity pages Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add helpers for the Swarm pages Pure functions for what the pages work out from the cluster API: a node's state in words, which nodes a placement rule may use, what a rule would ask for, what a drain or a lost node would leave without service, the nodes a bulk backend update reaches, and the join commands for a worker, a peer instance and a memory shard. Hooks read the roster, the loaded replicas and the rules. Everything runs in the browser from data the page already holds, and says when it cannot see free memory or disk. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Swarm hub: nodes, node page, placement rules, failover Nodes is a sortable table with comfortable and compact rows, a Needs attention filter by reason, a map of the cluster that is not drawn on a phone, the running models, and a bulk backend update for the nodes that drifted. A node is a page: state, vitals, a drain preview computed from the loaded replicas and the rules, tabs for models, backends, logs and capacity and labels, and Remove that asks for the node's name. Placement rules are written as sentences, show where each model is loaded now, and edit in a side sheet with a preview of the nodes a draft could use. Deleting a rule waits a few seconds so it can be taken back. Failover keeps its chains, adds what the router does when a worker stops answering and a per-node preview of what would stop. Add a node covers a registered worker, a peer instance and a memory shard, with a command to copy and a live line that says when the machine arrived. P2P keeps its page in the same vocabulary, and the node logs page follows the local logs page. Previews are labelled as worked out in the browser. Per-GPU readings and node events are not drawn because the API does not return them. Failover moves to Swarm when distributed mode is on. Legacy fleet components and their styles are removed. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Swarm hub New specs for adding a node (each join method, the command, copy, waiting and found, approve, a single install, P2P, a phone), placement rules (sentences, where models are loaded, the preview matrix, the sheet and its preview, delete with undo) and failover on a cluster. Node detail covers its tabs, the drain preview and its dialog, resume, remove with the typed name, a node that stopped answering, and unload. The nodes specs follow the new structure and keep their intent: the table, filters, grouping, pagination, bulk actions, the map, and running models with stop, logs and the loading, error and empty states. The scheduling, failover, P2P, hub and smoke specs follow the renames. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the rebuilt Swarm pages Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add helpers for the Traffic pages Pure functions for what the pages work out from the usage ledger, the trace summary, the trace buffers and the resources reading: the shared time window, grouping, sorting and filtering of usage rows, chart series and axes that start at zero, the overview figures, per-model statistics, the state of a trace and the words for a failure, the backend operations that ran during a request, CSV export, the Prometheus metric list and scrape config, and a bounded buffer of host readings. A figure whose source cannot say is null, never zero. The trace summary call takes the window in hours, and a helper reads /metrics with its status. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Traffic hub: overview, usage, models, host, traces, middleware Traffic opens on an overview: five figures (requests, failed, p95, tokens in and out) and three charts, each naming its source. A second row of links reaches Usage, Models, GPU and host, Traces, Middleware and Prometheus, and one time window is shared by the first three. Usage groups by model, user or API key, filters, sorts, opens a row on its own chart, exports the rows it holds as CSV or JSON in the browser, and keeps the opt-in cost estimate and the quota forecast. A user who is not an admin sees only their own numbers. Models joins the ledger, the backend-operation buffer and the loaded models. GPU and host shows the current reading and two charts of readings taken since the page opened. Traces gets filters, a settings strip and an explained off state. An API request is a page: the error LocalAI recorded, a timeline with the backend operations that ran meanwhile, and bodies that stay closed until revealed. Middleware draws the pipeline as five steps and shows the rules of the selected step. Prometheus documents /metrics, checks it against the server and gives a scrape config to copy. Alerts is not built: LocalAI has no alert rules. Per-model latency percentiles, GPU utilisation and compare with the previous period are not drawn because the API does not return them. Legacy usage, trace and middleware styles and the usage source components are removed. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Traffic hub New specs for the overview (figures, charts with a data table and arrow key readout, failed and first-run and tracing-off states, the shared window, a phone), usage (group by, filters, sort, export, cost, quotas, a non-admin, empty and loading), models, GPU and host (snapshot, the since-opened labelling, a cluster), the traces list, a trace page (the real error, the timeline, reveal, no headers, a trace that left the buffer), Prometheus and the Middleware pipeline, with shared fixtures. The usage, traces, middleware, hub and smoke specs follow the new structure and keep their intent. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the Traffic hub Add an operations page for the Traffic tab: which record each page reads, what it leaves out and why, the trace page and its reveal, the GPU and host readings kept since the page opened, and the Prometheus endpoint. Link it from the operations index, the tracing page and the middleware page. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): let metric names wrap in the Prometheus table on a phone The long metric names pushed the type and "on this server" columns out of view. Names now wrap inside the table. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Settings with groups by intent, search, a pending bar and history The fifteen sections become eight groups by intent: memory and models, speed and defaults, backends and galleries, access and security, debugging and traces, agents and responses, swarm and sharing, look and feel. Search covers names, descriptions, keys and the old section name, and says where a result used to be. Edits wait in a bar with Discard, Show diff and Apply. The diff lists old and new values and the checks the browser can make: durations parse the way Go parses them, a GPU memory budget is one the server accepts, a gallery box holds JSON, and warnings repeat what the handler and the field text say. Apply sends only the changed keys. Undo saves the previous values again; it is a new save, not a rollback. History lists the changes applied from this browser, since LocalAI keeps no settings log, and Revert stages the old value. A value is marked as changed only where the built-in default is known from the CLI defaults. A row says "Applies now" or "Needs restart" only where the handler or the docs say so. Three things were wrong before and are fixed with the rebuild: the gallery boxes and the shared API keys box were sent under names the server ignores, the "Enable CSRF Protection" switch showed the disable flag the wrong way round, and every save restarted peer-to-peer networking because every field was sent. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Users and keys, Account, sign-in, invite and the 404 page Users and keys is a tabbed page under the Settings tab: people, invites and API keys. The people table filters by state and role, sorts, approves or disables (disabling offers an undo that sets the status back), and opens a side sheet for one person's features, model allow-list and limits. Role, password reset and delete sit in the row menu; delete asks for the name. Invites choose a lifetime of 1, 7 or 30 days and show the link once. API keys can be created with a lifetime, are shown once in full, can be paused, and are revoked after a ten second undo window in which nothing is sent. LocalAI lists keys only to their owner, so the tab shows the signed-in person's own keys and says so. Account has Profile, Security, API keys and Usage. Usage shows the last 30 days, tokens by model and the limits an admin set. The Security tab now shows for a GitHub or SSO account and says the password is not theirs to change. Sign-in asks for one field per step and draws a provider button only for a provider /api/auth/status lists. It has the notice for a sign-up that waits for approval, the first-admin screen, the key-only screen and the invite page. An address outside the app now gets the 404 page too, which names the address and lists the places the sidebar lists, with the same gates. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover Settings, Users and keys, Account, sign-in and the 404 page Settings: groups, search by name, key and old section, the changed marker only where a default is known, apply hints, the pending bar and diff, the checks, apply sending only changed keys, undo as a second save, discard, history, the CSRF inversion and the gallery and API key wire forms, and the phone layout. Users and keys: the table, filters, sort, approve, disable with undo, the row menu, the access sheet, invites, key creation with a one-time reveal, the ten second revoke with undo and with a page leave, and the non-admin redirect. Account, each sign-in variant (error, pending, first admin, key-only, invite, provider buttons) and the 404 page have specs too. Fixtures are shared with the screenshot scripts. Existing specs follow the new structure. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the rebuilt Settings, Users and keys, Account and sign-in pages Runtime settings: the eight groups and where each old section went, search, the pending bar, the diff and its checks, apply, undo, the history, and which settings show a default or an apply note and why. Authentication: the sign-in screen variants, the Account tabs, key lifetimes, the one-time key reveal, the revoke undo window, and the fact that keys are listed only to their owner. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): let the Settings undo toast stand alone and read back a generated P2P token The saved message and the undo toast sat on the same spot at the bottom of the page. The undo toast now carries the saved message. A new P2P token is made by the server when the page sends 0. The page reads it back after the save so the field shows the token and not the placeholder. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): fit the users table, API keys and Account figures on a phone On a phone the users table dropped its Role and Status columns off the screen edge with the row actions. The role and state now sit under the name, so the actions stay in view. API key rows no longer put the key icon on a line of its own, and the three Account figures keep one row. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the phone users table, reduced motion and the empty Account state Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): drop the apply note from three settings the save handler does not mention Size-aware eviction, automatic backend upgrades and development backends said Applies now, but nothing in the handler or the docs says when they take effect. A row now carries a note only where the code or the docs say so. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Voices and Faces as one identity family Voices is one page with three tabs: Speakers (voiceprints for recognising who is speaking), Speech voices (the text-to-speech reference library, kept apart because it is a different store) and From a recording (a link into the diarization workspace). Faces uses the same layout. Who is this and Same person? give the answer in a sentence with the real distance and cut-off, a word for how far inside the cut-off it sits, and a distance scale with the cut-off drawn on it. The cut-off slider re-reads the answer in the browser; the identify call sends the cut-off, and verify uses the threshold the model returns. The old confidence percentage is gone because it is not a probability. The server has no list call, so the people list stays in the browser and the page says so. After a search that asked for more people than it got back, a saved person the server did not return is marked, and people the server returned that the browser does not know are listed. Nothing is claimed from a short or cut-off search. Enrolling is a sheet: sample, name, labels, permission. A copy of the sample in the browser is opt-in, and an administrator can also keep the recording as a speech voice in the same step. Removing a person waits ten seconds behind an Undo toast and sends nothing before then. Errors say what happened (no face found, model missing, call failed), a blocked or missing microphone is explained, and a missing model or a missing permission renders a page that says what turns the feature on instead of a redirect. Analyze, detect and raw embedding move under More tools, with attribute guesses off by default. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Voices and Faces pages Specs for who is this (match, no match, working, failed, missing model), the cut-off slider, same person, a blocked, allowed and insecure microphone, the registry notes and the not-on-the-server marks, the enrol sheet and its opt-in copy, delete with undo on a fake clock, the disabled and no-permission states, the phone layout, reduced motion and Faces. Existing library and diarization specs follow the new structure and keep their intent. Node tests cover the distance words, scale layout, stored list and error mapping. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the Voices and Faces pages Add a WebUI section to the voice and face recognition pages: the two tools, the cut-off, what the people list is and why it can be stale, the undo window, and what is stored where. Point the Voice Library and Fish Audio notes at Build, Voices, Speech voices. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Build landing, Fine-tune, Quantize, Import and Explorer The Build landing says what each tool is for and what it needs from the machine: the installed backend, the GPU memory, RAM and disk the server reports, and a job that is running or the newest one when it failed. A tool that cannot run says why and what enables it. Fine-tune and Quantize share one page: set up, a check list that is redrawn as the form changes, a run view with progress, stages and a log, and a result with real next steps (export, import, chat, Models). The checks state only what the server reports. A job needs no estimate the server cannot make, so none is invented. Stop on a fine-tuning job asks whether to keep a checkpoint, a failed job shows the server's message, and a memory failure offers two changes that are applied to a copy of the setup. Import is a guided flow: source, review, import, done. The server returns no preview before an import starts, so the review reads the spelling of the source, prints the request the form will send and runs the checks that can be made early. The estimate that arrives when the import starts is set against free memory and disk. The ambiguity picker and the Write YAML tab stay. Explorer shows what GET /networks returns and lists a swarm with POST /network/add, with a join sheet that carries the token and commands. Build tools the account may not use say so instead of redirecting. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Build landing, the tool pages, Import and Explorer Specs for the landing (a tool ready, missing a backend, with no GPU, a running or failed job, a feature switched off, a member without admin, phone, reduced motion), the shared tool pattern for Fine-tune and Quantize (set up, live checks, start request, running with progress, chart and log, the stop choice, failure with the server message, finish with next steps, earlier jobs, the account-disabled page, phone), Import (source detection, review, checks, ambiguity, running with the estimate against free memory, done, Write YAML, phone) and Explorer (list, join, list a swarm, empty, not an explorer, retry, phone). Existing specs follow the new structure and keep their intent. Node tests cover the machine facts, tool status, checks, log lines, source detection, the import request and the join commands. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the Build tool pages, the import flow and the Explorer Fine-tuning and quantization now describe the set up, check, run and result steps and what the check list can and cannot say. The import section explains the review step and why the size and memory appear only after the import starts. The distributed page describes the Explorer list, the join sheet and what listing a swarm publishes. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): turn the hardware recommendations into a "Best for this machine" shelf The shelf in the Models inspector put five columns into a 400 px pane, so long model ids wrapped letter by letter underneath the size and the memory figures. Each row now stacks the tag, the id and the size and memory facts beside one Install button, and the id wraps inside its own column. Once a model is installed the shelf narrows to the best fit and keeps the others behind a "N more that fit" toggle. Specs cover the ranking, the layout, the narrowing and the install request against a gallery fixture that carries the 4K estimate the shelf sizes against. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): quiet the Studio tab markers and say their state in words The type tabs drew a saturated green dot for every modality that has a model. The dot now uses a text colour, filled when a model is installed and hollow when none is, and each tab carries "(model installed)" or "(no model installed)" as hidden text so the state is not only a colour. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): stack the model editor empty state and drop its section hues "No fields configured" sat in a flex row, so the icon, the title and the text ran together. It now uses the stacked empty-state layout. The section icons took a different status colour each (amber, red, green); they now share one quiet colour, with the accent on the current section. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): keep the Home memory sentence whole on a phone and drop side rails On a 390 px screen the memory strip clipped "2 models loaded" to make room for the figure. The sentence now takes the first line and the figure and device wrap under it. The sweep also removed coloured left rails from the editor section rail, the skill editor list, the install strip and the audio transform notice (now an outlined note), plus unused chat rules that carried rails and two glow animations that nothing referenced. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): line up hub pages, list the model templates, and stop clipped text Medium-width pages inside Build and Operate were centred while the tab bar above them was flush left, so the title started 60 px right of the first tab. They now start at the bar's edge. Add Model offered nine templates as a grid of identical cards with chip clouds and inline styles. It is now one list of rows, each with the field names it fills in on a single muted line. Two clipped strings are fixed: the Studio voice field cut its placeholder mid-word, and the phone job list ended the schedule line in an ellipsis. The recommendation shelf also separates size and memory with a dot, and the docs describe the shelf. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): keep the hidden Studio tab state inside its tab The hidden state text added to each type tab was absolutely positioned against the page, so on a phone it sat outside the scrolling tab row and widened the page by hundreds of pixels. The tab is now the containing block. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the on-disk sizes on the Installed table, model page and cleanup sheet The fixtures stub GET /api/models/storage. The default report is empty, so existing specs keep the gallery estimates. makeStorage() builds a report from files and the models that use them, the way the server does, and storageSpec() is a models directory with shared and missing files. New specs cover the Size column and its shared line, the fallback when the call fails or the user is not an admin, the files list on the model page, a missing file, the bytes a removal frees with shared files, and the cleanup findings. Node tests cover the storage helpers and the batch arithmetic. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): wait for the page before pressing keys and ticking the clock Two specs failed in loaded full runs and passed alone. The Alt+1 to Alt+7 spec pressed a key before the composer had armed its key handler. The capacity chart spec advanced the fake clock before the poller had mounted, so it counted fewer readings than it expected. Both now wait for the page to mount. The key spec retries a press that lands during a re-render, and the clock spec advances in small steps and polls for the row count. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): hide the Installed footer when the storage report is empty An empty report from the storage call made the footer read "0.0 GB on disk" next to sizes taken from the gallery estimate. An empty report says nothing about the disk, so the footer now shows only the model count. A spec covers it. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): keep Explore pane actions inside the pane The inspector actions sat in a non-wrapping flex row beside the title, so the buttons ran past the pane edge once it got narrow. The row now takes its own line and wraps. The primary action (Install, Retry, Open) comes first. Manage installation becomes a ghost button, and Open details moves to the end of the row, so one action stands out and the others are quiet. No action or test id is removed. Add a spec that checks, in light and dark at several widths and with a pane forced to 320 px, that every action stays inside the pane box and that the pane keeps its inner padding. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * refactor(ui): drop the type chips from the Studio composer The Studio tabs and the composer's type chips listed the same seven modes, so the page said the same thing twice. Keep the tabs as the one place to switch modes. The composer now shows the type it will open as a small label in its header. The type suggestion from the typed words stays as the quiet hint line under the prompt, and Alt+1 to Alt+7 still pick a type. The composer root carries data-type, data-types and data-missing so tests can read the state. Specs pick a type through a shared Alt+digit helper and read a missing model from the tab dot instead of a chip. Remove the unused chip locale strings and CSS. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * refactor(ui): remove the section crumb above page titles Page headers drew a small uppercase crumb with a short rule before it above the title. On the hub pages it repeated the hub name, so Build sat above a heading that also said Build. PageHeader now renders only the title, the supporting line and the actions. Drop the eyebrow prop, the route-derived section name, its CSS and the unused section helper, and remove the explicit eyebrow props from the pages that passed one. Pages stay reachable through the sidebar and the hub tab bar. Add a spec that checks several pages show their title with nothing ahead of it in the header. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * refactor(ui): remove left accent rails from tiles, rows and quotes Several surfaces marked state with a coloured strip on the left edge. Replace each one with a cue that is not a rail: - Stat cards lose the strip; the icon and value still carry the colour. - The highlighted card is a raised surface with a firmer edge. - The selected rail row is an accent wash with a hairline outline. - The status stripe on rail items is a small status dot. - The active failover row is a tinted row. - Quotes in markdown and chat prose are italic instead of barred. - The variant detail panel has a full hairline border. Add a spec that walks the main routes in light and dark and fails on a left border thicker than 1px, a sideways inset shadow, a narrow absolute strip in ::before or ::after, or a narrow tall child pinned to a left edge. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com> |
||
|
|
6343a2dc6f |
feat(voice-detect): report the encoder family so audio-registered voices are fingerprinted (#12499)
* feat(voice-detect): report the encoder family so audio-registered voices are fingerprinted
A voice registered from audio through the voice-detect backend had no
encoder fingerprint, so the parakeet-cpp backend could not tell whether
it was comparable with the loaded speaker model and could only fall back
to the file-name rule.
The voice-detect backend now binds the three new libvoicedetect accessors
with a symbol probe (an older library still loads and reports nothing),
copies the borrowed strings at once and never frees them. It also hashes
the model file once at load. VoiceEmbedResponse gains two optional
fields, encoder_family and encoder_weights ("sha256:<hex>", empty when
the model is not a plain file).
/v1/voice/register stores them as encoder_family and a new
encoder_weights field in the registry entry; model keeps the encoder
name, so the 1:N identify filter by name is unchanged for old entries.
A voice with a family is sent to the backend whatever its file name, and
the backend decides by family. /v1/voice/identify compares the family
when both the stored voice and the probe have one. Old entries load
without the fields and stay unfingerprinted.
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* docs(voice-detect): say how a voice with weights but no family is treated
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(voice-detect): mark the model file open as intended for gosec
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
|
||
|
|
0102751b31 |
fix(mcp): connect to legacy SSE servers (#12304)
* fix(mcp): connect to legacy SSE servers Remote model MCP connections only use Streamable HTTP, so legacy SSE servers fail initialization. Retry with SSE when the initial POST returns 400, 404, or 405. Share one discovery timeout across attempts and cancel failed connections without truncating successful sessions. Preserve HTTP policy and reject foreign SSE message endpoints before attaching credentials. Add SDK integration tests and document automatic transport selection. Assisted-by: Codex:gpt-6 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(mcp): avoid narrowing HTTP status codes Store the initialization status in a 64-bit atomic value to avoid the integer overflow conversion reported by gosec. Assisted-by: Codex:GPT-6 gosec Signed-off-by: Ettore Di Giacinto <mudler@localai.io> --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
c03acb0647 |
feat(models): report per-model disk usage in API and WebUI (#12551)
New admin endpoint GET /api/models/storage reports disk usage of the installed models: each model's files and sizes, files shared between models counted once, and configured files that are missing from disk. The models WebUI page shows a usage summary, each model's file table, and the full file list with per-file status. Assisted-by: Claude Code:claude-fable-5 Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com> |
||
|
|
cd604b5c5b |
fix(distributed): bound, stop and cancel model loads with leases and worker operations (#12524)
Fence model load jobs by generation, lease them on the database clock, bound the work on the worker with operations and a process-group watchdog, and add one stop path with a load-cancel API. See the pull request for the design, the rolling upgrade notes and the test evidence. Assisted-by: Claude Code:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
c525ad16e8 |
fix(audio): record transcription usage and traces (#12546)
* fix(audio): record transcription usage and traces Count successful transcription requests by model, including streams. Keep token counts at zero because transcription exposes no token usage. Capture multipart API trace metadata without reading uploaded audio. Do not record failed transcription or client writes as successful usage. Assisted-by: Codex:gpt-6-astra * fix(audio): preserve aliases in streaming usage Streaming transcription records the resolved target as its usage model. Pass the requested name so JSON and SSE requests share the alias bucket. Assisted-by: Codex:gpt-6-astra * test(http): check multipart reader close errors Assert successful reader cleanup to satisfy the errcheck CI gate. Assisted-by: Codex:gpt-6 --------- Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
6b794651a4 |
feat(diarization): return sound events with include_sounds (#12544)
* feat(diarization): return sound events with include_sounds
A client that wants text, speakers, voice prints and sound events had to
make a diarization call and a separate sound call. Add an include_sounds
request field to /v1/audio/diarization that adds a sounds array of closed
events {start, end, label, confidence}, in seconds.
The parakeet-cpp backend runs a tagger-only scene stream over the clip,
the same stream and thresholds the live path uses, so a clip gives the
same events offline and live. A model with no sound_model companion, or a
backend that does not report sound events, fails with 501 and the stable
code include_sounds_unsupported instead of an empty list. The proto
carries sounds_included so an empty list still means "nothing heard".
The localai-proxy backend forwards the field. Swagger, docs and the
e2e mock backend are updated.
Assisted-by: Claude:claude-sonnet-5-5 [protoc swag go]
* feat(gallery): add parakeet-cpp-multilingual-diarization-speakers-sounds
Same as parakeet-cpp-multilingual-diarization-speakers (TDT 0.6B v3,
Nemotron-3-Diarization, WeSpeaker) plus a CED-Tiny sound_model, so one
model name serves /v1/audio/diarization with include_text,
include_speaker_profiles and include_sounds. It declares the
sound_classification usecase like the realtime scene entries.
Assisted-by: Claude:claude-sonnet-5-5
---------
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
|
||
|
|
690f3994b0 |
feat(auth): let users pause and resume their API keys (#12521)
A key can now be paused until the owner resumes it, or until a given
time. A paused key is rejected by validation before last_used is
updated, and a pause time that has passed lifts the pause by itself.
Existing keys stay active.
PATCH /api/auth/api-keys/:id takes {"disabled": bool, "paused_until":
RFC 3339 string or null}. Only the key owner can change it, and a
paused_until in the past is rejected. The key list returns the pause
fields. The Account page gets a Pause and Resume button for each key
and a Paused badge that shows the resume time.
Assisted-by: Claude Code:claude-sonnet-5-5
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
|
||
|
|
fa25493c9e |
feat(config): let parakeet-cpp declare the speaker_recognition usecase
parakeet-cpp now implements VoiceEmbed and VoiceVerify, so list them in its backend capabilities together with the speaker_recognition usecase. The usecase is not guessed, because most parakeet models have no speaker encoder: a bundle declares it in known_usecases, which is enough for /v1/voice/* and for voice_recognition.model. Add a realtime pipeline test where one bundle model names the vad, transcription, sound_detection and voice_recognition stages and the gate resolves the speaker through the backend's VoiceEmbed. Assisted-by: Claude Code:claude-sonnet-5-5 |
||
|
|
1d117292b1 |
feat(voice): name registered speakers through a bundle speaker component
Naming read only speaker_model:, so a parakeet-cpp bundle that sets speaker_component:voice never received the registered voices, in live transcription or in diarization. Resolve speaker_component: when speaker_model: is not set. A component is not an encoder file, so no file-name tag is derived for it: voices with a hash or family identity and untagged voices reach the backend, and a voice with only a file-name tag matches through the new optional speaker_tag: alias. speaker_model: still wins, including when it points at the bundle file. The bundle gallery entries do not declare speaker_recognition: that usecase means the VoiceEmbed and VoiceVerify RPCs, which the backend does not serve, and it selects the default model for /v1/voice/*. The pinning test now says so. Assisted-by: Claude Code:claude-sonnet-5-5 go golangci-lint |
||
|
|
79a7631cc5 |
feat(parakeet-cpp): encoder fingerprint for speaker naming, VAD trim and word filter options, pin bump (#12491)
* chore(parakeet-cpp): bump parakeet.cpp to 2de154c Brings in the speaker registry encoder fingerprint, the VAD segment trim and the opt-in word filter, a fix for a per-call thread count that stayed set on the process-wide backend after a Silero VAD pass, and bundle components loaded from memory. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(parakeet-cpp): encoder fingerprint for speaker naming, vad_trim and guard_* options Speaker naming. A registered voice now carries the encoder that made it: the embedding family (voicedetect:<arch>:<name>:<dim>) and the sha256 of the weights. The backend reports the family of the loaded speaker model in Status, voice enrollment from speaker_profiles stores it as encoder_family (old entries load without it), and the registry sent to parakeet.cpp is built with parakeet_capi_speaker_registry_add_embedding_fp. The library then refuses a registry of another encoder family and the error names both families; another quantization of the same family only warns. A voice with only a weights hash gets the loaded family when the hashes are equal. Voices without a fingerprint (registered from audio: libvoicedetect cannot report one) keep the file-name rule and are used with a warning. The library cannot mix them with fingerprinted voices in one registry, so a request that has any uses the old registry for all. speaker_strict:true drops them instead. A library without the symbols behaves as before. Transcription. vad_trim (seconds, 0 keeps the whole cuts) goes through the VAD options JSON, so it reaches /v1/vad and the segmenter. The guard_* options guard_min_local_conf, guard_local_radius and guard_drop_punct_only turn on the word filter through parakeet_capi_transcribe_path_json_with, or through the segmenter with vad:true. They are off by default, bad values fail the load, and a library without the symbol fails it with a clear message. The dropped word count is logged at debug level. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
5f5feea15a |
fix(agentpool): keep the conversation in the agent web chat (#12410)
* feat(agents): keep the web chat history between messages The agent chat endpoint ran every message as a fresh job (ag.Ask(WithText(message))), so a follow-up such as "add two days to item 3 and recalculate" never saw the answer it referred to. Agents then rebuilt their reply from scratch instead of changing it. Use the agent's own conversation tracker, as the Telegram and Slack connectors already do: send the earlier turns with the new message and record successful answers. Failed, cancelled or empty runs are not recorded, and the tracker drops a conversation after the agent's last_message_duration of inactivity. The distributed (NATS) chat path is unchanged. Assisted-by: Claude:claude-opus-5-5 ginkgo Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de> * fix(agents): keep each web chat conversation's history separate Review on this PR: the first version kept the history in the agent's conversation tracker under one fixed key per agent. The web UI keeps several conversations per agent (New Chat, switching, Clear) and did not send which one a message belongs to, so the hidden history mixed them: a new chat received the previous chat's turns, and Clear only cleared the screen. The client now sends the earlier turns of the conversation it is showing as `history` with POST /api/agents/:name/chat, and the server keeps no web chat history of its own. Each conversation only ever sees its own turns; New Chat and Clear start without history. The server uses only user and assistant turns with text, bounded to the most recent 40 turns and 64,000 characters. `history` is optional, so existing callers keep the previous behaviour (no history); the distributed (NATS) path does not forward it yet. Tests: two conversations of one agent stay apart, an empty history (New Chat or Clear) starts fresh, system/tool/empty turns are dropped, the bounds keep the most recent turns; the UI helper that builds the history from the visible messages has node --test coverage. Docs: features/agents.md describes `history` and the web UI behaviour. Assisted-by: Claude:claude-opus-5-5 ginkgo Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de> --------- Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de> |
||
|
|
4d681b7f6d |
fix(realtime): detach VAD-commit transcription from barge-in cancellation (#12446)
* fix(realtime): keep VAD-commit transcription alive across barge-in, cancel it at teardown, order commits
Barge-in (new speech onset) cancels the turn's SourceVAD response context
(realtime_turncoord.go: respSink.cancel(SourceVAD)). The VAD commit body
runs under that same context, so an in-flight Whisper STT call was aborted
with 'context canceled' whenever the caller kept talking while the first
chunk was transcribing. The user's turn was lost: no transcript, no
LLM/TTS response.
v1 of this fix ran the transcription with context.WithoutCancel(ctx). The
review correctly pointed out two correctness gaps:
1. Teardown lost its cancellation. WithoutCancel detaches from every
cancellation, so a transcription in flight at session close outlived
the session and blocked respSink.shutdown (which joins the response
goroutines) until the backend finished the job.
2. Out-of-order commits. Consecutive commits run in parallel goroutines,
so a fast second transcription could append its user item before a
slow first one: the conversation became [second, first] and the second
response saw only [second].
Changes (core/http/endpoints/openai/):
- Session gains a session-lifetime context (sessionCtx), cancelled by
conncoord's Teardown BEFORE respSink.shutdown joins the response
goroutines. The transcription (and the voice-gate resolution) run under
it: they survive barge-in (which cancels only the per-response context)
but are cancelled with the session.
- Commit slots order the user-item appends in speech order:
Session.nextCommitSlot() is claimed at commit issue time (VAD CommitTurn
/ client commit), a commit's item append waits on the previous slot's
done (aborts on the session context), and every exit closes the slot so
a failed or torn-down commit never blocks the next. Transcriptions stay
parallel; only the appends are ordered.
- If the turn's response context was cancelled while the (detached)
transcription ran — barge-in, superseded by a newer commit — the user
item still commits (appendUserItem, split out of generateResponse) so
the LLM context keeps the full user input, but no response is generated
for the superseded turn; the newer speech triggers its own response on
the complete history.
- Regression tests (realtime_commit_order_test.go) cover both review
schedules — teardown during an in-flight transcription, and
held-first/finished-second out-of-order completion — plus the
barge-in-during-transcription item survival, driving the real commit
path with a transcription double that honours context cancellation.
- docs/design/realtime-state-machines.md: implementation-status entry for
the committed-turn pipeline (transcription lifetime + commit order).
Fixes #12445
Validated: builds, go vet clean, all openai specs + respcoord/turncoord/
conncoord suites pass under -race (incl. the 3 new regression specs).
Production A/B (call-center voice agent, SIP, silero-vad +
whisper-large-turbo + LLM + TTS, server_vad ~600 ms) on LocalAI v4.11.0:
unpatched — 'transcription_failed: context canceled', first part of the
utterance lost, agent answers only the remainder; patched — full
transcript committed, agent answers the complete utterance, barge-in
still cancels the in-flight assistant TTS response as intended, and
teardown cancels the in-flight transcription instead of waiting for the
backend.
Signed-off-by: nexxtmobile.de <kai@nexxtmobile.de>
* fix(realtime): release commit slots in order on every exit; share slot+issue boundary
Follow-up to the review of
|
||
|
|
99043b442c |
feat(router): route with native decision models (#12449)
* fix(schema): preserve SystemOne image inputs Assisted-by: OpenAI * test(schema): follow Ginkgo conventions for decision inputs Assisted-by: OpenAI * feat(llama-cpp): dispatch native decisions through Score Upgrade the stock dependency and reconcile Score/TTS patches. Reuse native decision parsing, tasks, formatting and response-reader cleanup; preserve ordinary scoring admission and guard older dependencies. Assisted-by: OpenAI * refactor(systemone): share request and model validation Assisted-by: OpenAI:gpt-5 * fix(systemone): preserve HTTP wire-byte validation limit Keep structural validation separate from the serialized internal request bound so HTML escaping cannot reject valid HTTP payloads. Assisted-by: OpenAI:gpt-5 * feat(systemone): bound images and account native decisions Preserve public wire limits independently from router serialization. Reject unsupported NER images, map native request/capability errors, and stamp explicit usage once. Advertise decisions for stock llama-cpp. Assisted-by: OpenAI * fix(systemone): record usage on registered native route Exercise real registration and billing with a mock native backend. Reject empty native responses, malformed image URLs, trailing JSON, and wire overflow including whitespace. Assisted-by: OpenAI * feat(router): add lazy native decision transport Bind named models through internal ModelSystemOne calls with shared validation and bounded abandoned operations. Remove request and echoed-error contents from decision traces. Assisted-by: OpenAI:gpt-5 * feat(router): classify overlapping policies with native decisions Ask independent noul questions, validate probabilities and preserve first-superset routing. Wire the central factory with config-sensitive invalidation and cancellation-safe resolution. Document native framing and bounded operation limits. Assisted-by: OpenAI:gpt-5 * feat(gallery): add pinned Julia-1 native decision model Add a separate text-only llama-cpp Q8 entry with pinned Apache-2.0 source provenance and checksum. Installed using the gallery installer and exercised choice, score and noul on CPU. Assisted-by: OpenAI * test(router): verify native decisions through central factory Add an opt-in real-model Ginkgo integration covering the native Go loader and C++ transport, token usage, independent overlapping labels, and candidate selection. Document owned-server execution and the intentionally non-quality threshold. Assisted-by: Codex:gpt-5 * fix(llama-cpp): align upstream pin and preserve decision signatures Advance to bed0a856 without losing the automated upstream bump. Detect full-request fill_task support at compile time and forward every question for Nimble framing while retaining the earlier native signature. Preserve reconciled SCORE/TTS patches; add standalone compatibility coverage. Assisted-by: Codex:gpt-5 * feat(gallery): add native decision family defaults Pin Laya, Kev-4B, lev, OpenJev and Nimble artifacts. Verify Laya/Kev/lev gallery installs and CPU contracts on both native pins; clearly mark OpenJev/Nimble runtime validation pending and their noncommercial licenses. Assisted-by: OpenAI * docs(decisions): clarify integrated Nimble prerequisite Record the exact combined backend pin while retaining pending OpenJev and Nimble installation/runtime validation status. Assisted-by: Codex:gpt-5 * fix(gallery): indent native decision model sequences Match repository yamllint indentation for Laya, Kev, lev and OpenJev list fields. Parsed gallery data is unchanged; reproduce CI gallery lint failure before the whitespace-only fix and pass the same command afterward. Assisted-by: Codex:gpt-5 * docs(decisions): record OpenJev and Nimble CPU validation Record gallery installation, checksum/metadata verification and multiquestion native smoke results on bed0a856. Retain noncommercial and text-only limitations without accuracy or deterministic-output claims. Assisted-by: OpenAI * fix(ui): expose native Decisions router classifiers Select classifier models using metadata-driven capability routing, retain tuned thresholds, and validate native decision selections before saving. Cover both native backends and create/save/reopen in the real React editor. Assisted-by: Codex:gpt-5 * fix(router): exclude aliases from native decision discovery Check the originally named config before advertising native Decisions eligibility. Retain target capability inheritance for ordinary generation aliases. Exercise the actual capabilities endpoint with native models on both backends, aliases, and disabled models. Assisted-by: Codex:gpt-5 * feat(systemone): share bounded multimodal input validation Preserve text wire limits while admitting bounded PNG/JPEG decision input. Share collection and header validation across internal and public callers and keep the native runner response budget independent. Assisted-by: OpenAI:API-assistant * fix(systemone): bound admission lifetimes and validate complete images Retain shared admission leases through actual work completion, including abandoned internal operations. Decode bounded image pixels, cap public native responses before usage stamping, and preserve oversized malformed text status precedence. Assisted-by: OpenAI:API-assistant * fix(router): classify images before media fetching Preserve ordered structured probes for native decisions. Defer OpenAI media preparation until routing selects the served model, so rejected decision URLs cannot trigger downloads before shared validation. Guard direct image collection with context-aware shared admission. Keep text classifiers and embedding caches from discarding image input. Retain fail-closed classifier configuration and runtime fallback policy. Add middleware, typed-content, admission, cancellation and cache tests. Assisted-by: OpenAI:API-assistant * fix(router): bound extraction before serialization Check probe budgets before copying text or marshaling message state. Count JSON escaping so oversized internal inputs fail before allocation. Preserve typed Anthropic blocks through selected-model conversion and fallback. Keep retry coverage in Ginkgo without global test registration. Assisted-by: OpenAI * fix(router): bound supported probe serialization Arbitrary structs can bypass the probe budget through pointer marshalers, string tags, and promoted fields. Accept concrete chat schema types and plain JSON values instead of emulating arbitrary struct serialization. Budget escaped direct prompts before marshaling so raw length cannot hide serialized expansion. Preserve runtime fallback and reject oversized input before invoking the decision runner. Add Ginkgo allocation, boundary, and marshaler invocation regressions. Six-package tests, three-package race tests, and full-T2 delta lint pass. Assisted-by: OpenAI:GPT-5 golangci-lint * feat(decisions): enable bounded OpenJev images Validate native decision images before permissive media parsing and pixel allocation. Require both decision image support and a vision projector; missing or audio-only projectors cannot silently become text decisions. Pin the OpenJev Q8 projector and document its license and disk footprint. Add native safety tests, canonical limit parity, gallery and load-option checks, and a reproducible CPU direct-RPC contrasting-image smoke. Assisted-by: OpenAI:GPT-5 * fix(decisions): reject incomplete image streams stb accepts corrupt PNG Adler checksums and truncated JPEG scans. Use bounded zlib validation and strict libjpeg decoding before parsing. Keep dimension and aggregate pixel checks ahead of decoder allocations. Wire decoder dependencies into native builds and runtime packaging. Add regressions for appended EOI and embedded marker bypasses. Assisted-by: OpenAI:GPT-5 * fix(ci): gate native decision image validation Run the decoder security tests outside the stdlib-only native suite. Fetch vendor headers at the backend pin and provision decoder dependencies. Gate Go limit parity and production CMake wiring without model downloads. Assisted-by: OpenAI:GPT-5 * test(decisions): cover multimodal public API paths Exercise shared image contracts through the registered HTTP routes and external mock backend. Add opt-in cached gallery installation and real OpenJev image decisions through SystemOne and both routing APIs. Assisted-by: Codex:gpt-5 * test(decisions): assert isolation and cache bypass Observe external RPC calls and compare complete classifier history. Winner-only and cache-miss checks could hide dropped history or cache use. Give real inference its own application and model directory so shared backend mappings and loaded processes cannot affect mixed suite order. Assisted-by: OpenAI:ChatGPT * test(decisions): isolate fixture globals Disable optional global services in the isolated HTTP fixture and register cleanup before setup assertions. Verify meter provider identity survives fixture creation and destruction. Snapshot observed usage before assertions so failures cannot retain the mutex. Require a successful usage stamp before checking error responses. Assisted-by: Codex:gpt-5 golangci-lint * fix(application): honor optional telemetry controls Skip failover gauge registration when metrics are disabled. Register against the application meter rather than looking up the global provider. Allow embedders to retain the bounded routing log without billing stats. Keep the existing default when stats are disabled. The isolated HTTP fixture uses this option without losing its native router assertions. Assisted-by: Codex:gpt-5 golangci-lint --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
e4fa051ee3 |
refactor(distributed): put the NATS-only paths behind interfaces (#12395)
* feat(messaging): add shared subject rules Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * test(messaging): cover BroadcastRoots, ControlRoots and SubjectRoot Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(messaging): add Broadcaster and enforce subject rules in every carrier Broadcaster is the fan-out half of MessagingClient. The NATS client and the in-memory FakeBus now refuse a subject outside the served roots and any wildcard other than a whole single token, and FakeBus shares MatchSubject instead of its own copy. FakeBus Unsubscribe now removes its own subscription instead of the first one with the same subject. A shared conformance suite in messagingtest runs against both carriers. The distributed e2e specs that used invented test.* subjects, and the one that subscribed with a > filter, now use subjects from subjects.go. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor: depend on Broadcaster where only publish and subscribe are used Narrowed to messaging.Broadcaster: nodes/staging_progress.go, nodes/install_progress_publisher.go, galleryop/operation.go, galleryop/service.go, agentpool/user_services.go, agentpool/agent_jobs.go, openresponses/store.go, openresponses/sync.go, syncstate/syncstate.go, finetune/service.go, quantization/service.go and failover/distsync/distsync.go. SubscribeJSON now takes a Broadcaster because it only calls Subscribe, which lets the narrowed consumers use it. Stayed wide: worker/supervisor.go, because its client field also serves the SubscribeReply handlers in worker/lifecycle.go. The request/reply, queue and wiring files (nodes/unloader.go, nodes/file_stager_s3.go, jobs/dispatcher.go, agents/dispatcher.go, agents/events.go, worker/file_staging.go, cli/agent_worker.go, http/app.go) are unchanged by design. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(nodes): name the no-route condition and confine the carrier error Consumers matched nats.ErrNoResponders, which names an absence, to demote a node. They now match ErrNoRoute, the control path maps the carrier's failure onto it, and timeouts and worker refusals are pinned as not being no-route. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs(nodes): state which FileStager implementations return ErrNoRoute Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(nodes): build backend clients through one node-aware seam Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: describe the distributed transport seams Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: correct comments that overclaim after the seams refactor Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(agent-worker): refuse an unserved LOCALAI_AGENT_SUBJECT at startup The messaging client now refuses a subject whose root no carrier serves. An agent worker started with a custom LOCALAI_AGENT_SUBJECT such as tenant-a.agent.execute used to start and then wait on a subject the frontend never publishes to. After the subject rules landed it exited at subscribe time with an error that did not name the setting. Behaviour change: the worker now checks LOCALAI_AGENT_SUBJECT before it registers or connects, and exits with an error that names the variable and says to use a served subject under the agent root, for example agent.execute. The served roots are not widened: a custom root was never delivered by the frontend, and a wider set would reopen the drift the subject rules exist to close. The flag help and the agent worker docs state the constraint. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * test(nodes): pin the reactions to ErrNoRoute Three callers react to ErrNoRoute and had no spec: the reconciler's upgrade drain falls back to the legacy forced install, the reconciler marks the node unhealthy when a pending op has no route, and the backend-op fan-out marks the node unhealthy. Each spec drives the real caller with a scripted no-responders reply and reads the result from the registry or the recorded requests. A fourth spec pins the other side: a pending op that times out leaves the node healthy and only counts the attempt, so mapping timeouts onto ErrNoRoute would fail here. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * test(messaging): pin client subject checks and fail the carrier suite in CI Add specs that call Publish, Request, Subscribe, QueueSubscribe, SubscribeReply and QueueSubscribeReply on a client with no connection. Each call must return ErrUnservedSubject for bogus.thing and ErrUnsupportedWildcard for jobs.>. This proves that the subject check runs before the connection is used, and needs no server. The NATS conformance suite is the only check that runs the subject rules against a real carrier. Before this change it skipped without output when Docker was missing. Now it fails when CI is set, so a Linux runner without Docker cannot hide it. It still skips on local runs and on macOS CI, which has no Docker. Add SubjectNodeBackendInstallProgress to the list of constructors that must build served subjects, and ask contributors to extend the list. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: state what ErrNoRoute may change, and group the distributed guides The seams note said MarkUnhealthy was the only state change allowed on ErrNoRoute. A pending backend op still records the failed attempt, counts toward the reconciler's retry limit and is dead-lettered after the maximum attempts. The note now says that MarkUnhealthy is the only change to the node's own state, and that the per-op accounting is not a verdict about the node. The note also documents that the NATS conformance run fails under CI when Docker is missing. The distributed-seams row moves next to the distributed-state row in the topics table. The liveness ping spec header now says no route is a reason to skip the worker, not proof that the worker is gone. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(nodes): give the backend client factory the node id Mechanical: the method gains a nodeID parameter and the eight test fakes are updated. No behaviour change. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(nodes): drop the optional node-aware factory The node id is now in the main method, so the optional interface and its helper had no behaviour of their own. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(nodes): dial backend probes through the client factory Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(nodes): dial workers' file servers through a per-node dialer Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(http): proxy backend logs through the per-node worker dialer The admin backend-logs proxy (list, lines and the WebSocket stream) now reaches a worker through the same per-node dialer as the HTTP file stager, so every frontend-to-worker dial goes through one seam. The shared direct dialer keeps alive for 15s where the proxy used 30s. Harmless for requests bounded at 15s. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(http): keep the backend-logs proxy independent of the admin connection The proxy request had no context before the dialer change and is bounded only by its 15s timeout. Keep it that way. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor: move the worker control payloads to workerctl Mechanical move of the request and reply structs, the install progress event and the file payloads out of messaging. The verbs no longer belong to one carrier. No alias is left behind. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(worker): serve the lifecycle verbs through a controlServer The worker registers one handler per verb and a NATS server maps each verb to its subject. Registration errors now name the verb. node.stop is served with SubscribeReply, which is identical on the wire because the handler never replies. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(worker): report install progress through the control sink Install and upgrade now emit download progress through the sink the control server hands them. The debounce and the terminal flush stay in the handler path, built over that sink by the new nodes.NewDebouncedInstallProgressSink, which replaces NewDebouncedInstallProgressPublisher. The subject and payload on the wire are unchanged. The supervisor no longer holds the bus, and installFn and upgradeFn let specs drive both verbs without a gallery. The malformed-request log lines are restored for install, upgrade, backend.delete, model.unload, model.stop and model.delete, with the reply bytes unchanged. The signal adapter is renamed noReply, which also lets worker.go import os/signal without an alias again. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(worker): serve the file-staging verbs through a controlServer An empty list-dir answer is now {} rather than {"files":null}, because the typed reply omits an empty Files slice. The frontend decodes both to a nil slice in nodes/file_stager_s3.go. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(messaging): add WorkQueue and the NATS producer This is the producer side of the competing-consumer seam. The work kinds map one to one to today's subjects and queue groups: task to jobs.new and mcp-ci to jobs.mcp-ci.new (both in group workers), agent-run to agent.execute (group agent-workers). Enqueue publishes the payload as Publish does today, with one JSON marshal. FakeBus now records queue groups and keeps reply handlers so later specs can pin and drive them. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(messaging): add the NATS WorkConsumer An in-flight limit of one runs the handler inline on the delivery goroutine, as the MCP CI consumer does today. Any other limit spawns per delivery, as the agent consumer does. Queue groups are unchanged. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor: publish queued work through WorkQueue The job dispatcher, the agent pool and the agent scheduler enqueue through messaging.WorkQueue; the NATS implementation publishes to the same subjects as before. DistributedServices builds the queue next to the NATS client and hands it to the dispatcher and the agent pool, whose distributed mode switch now reads a non-nil WorkQueue. The unused AgentPoolService.SetNATSClient is removed. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor: consume queued work through WorkConsumer The agent dispatcher and the MCP CI consumer register through messaging.WorkConsumer. The NATS implementation keeps the inline one-at-a-time model for MCP CI and the per-delivery model for agent runs. handleMCPCIJob reports on the events publisher the carrier hands it instead of a captured client. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor: delete the consumers nothing in production reached jobs.new has a producer and no production consumer, and the agent dispatcher's Dispatch was only called from tests. Publishing jobs.new is unchanged. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(mcp): send MCP requests to agent workers through AgentControl Timeouts still honour only the deadline, not cancellation, exactly as today. The NATS no-responders error maps to ErrNoRoute and a timeout does not. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(agent-worker): serve MCP requests and backend.stop through agentRPCServer The agent worker's MCP tool and discovery reply subscriptions and its backend stop listener move behind an unexported agentRPCServer interface, served on NATS by nodes.NATSAgentRPCServer. The handlers become typed mcp.ToolHandler and mcp.DiscoveryHandler values that answer every failure with a reply carrying Error. Queue group (agent-workers), inline execution on the delivery goroutine, the background handler context, the unmarshal error reply texts and the reply-less backend stop subscription are unchanged. The backend stop handler takes the decoded backend name, so it can still close that backend's MCP sessions. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(messaging): remove helpers that only tests used BroadcastRoots, ControlRoots and SubjectRoot had no production caller. The roots spec now asserts every served root through ValidateSubject instead. MatchSubject moves back into the test support package, the only place that used it, with its table. NATSAgentRPCServer drops the subscription list it stored and never read, and NewNATSAgentRPCServer gets a doc comment. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * test(mcp): round trip the agent RPC server over a real NATS server One spec sends a tool request and a discovery request through NATSAgentControl to NATSAgentRPCServer and checks that the handlers see the decoded requests and the replies come back. It also puts an undecodable body on the tool subject and checks the server answers with an unmarshal error instead of leaving the requester to time out. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: describe the distributed transport seams The developer note now lists the final seams: fan-out, queues, both halves of the control verbs and of agent RPC, and the dial. It records the open items a second carrier has to handle. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * test: pin the in-flight limit each queue consumer asks for The work queue specs pin what Consume does for a given limit, but nothing pinned which limit each production consumer passes. Changing the agent worker's MCP CI limit from 1 to 0 would have let MCP CI jobs run concurrently on each worker with every test green. Move the MCP CI Consume call into startMCPCIConsumer with the same wiring and pin that it asks for (WorkMCPCI, 1). Pin that NATSDispatcher.Start asks for (WorkAgentRun, maxConcurrent) for several limits. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor: remove helpers the branch left without a caller SubjectJobCancelWildcard lost its last subscriber when the frontend stopped listening on jobs.*.cancel; the NATS permissions and conformance suite spell the subject out, so nothing reads the constant. decodeBackendStopRequest returned a stopAll flag that production dropped and only a test read. decodeBackendStop is now the single decoder with the same semantics: an empty body is stop-all, an empty Backend is stop-all, malformed JSON is an error. stopBackends still derives stop-all from Backend, so no reply changes. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(messaging): keep an explicitly empty agent queue a plain subscription Before the work queue seam the agent worker passed LOCALAI_AGENT_QUEUE straight to QueueSubscribe, so an explicitly empty value made a plain subscription and every agent worker ran every agent run. WithAgentRunRoute replaced an empty queue with agent-workers, which silently changed that. Keep the queue as given once the option is applied. An empty subject still falls back to agent.execute, since it never had a meaning of its own. The flag default stays agent-workers, so only an explicitly empty value reaches this. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: correct comments and record the PR B notes Fix the recordingFactory comment (it also records the parallel flag), document that a negative maxInFlight is unbounded and that Unsubscribe from a handler deadlocks, and say a permanently undecodable payload returns nil. Record controlHandler's undecodable return as a kept exception, and add the second carrier notes to the developer note: the reconciler has no ClientFactory option, the logs proxy honours HTTP_PROXY, verbs one carrier serves need an opt-out, terminal replies come from the result event, and agent runs publish through the NATS-bound EventBridge, which is not an additive change. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
9eb5a9e61d |
feat(audio): remember speakers from diarization (#12414)
* feat(schema): validate portable speaker profiles Add the versioned profile schema for explicit speaker enrollment. Validate compatibility against separately supplied loaded-encoder metadata. Reject unusable speakers, invalid vectors, and inconsistent clean spans. This slice does not change HTTP routes, backend integration, or the UI. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(parakeet): export profiles with transcripts Export opt-in speaker profiles and trusted encoder metadata. Replay registrations by ID so duplicate display names keep independent vectors. Use one profile-capable diarization for slots, names, and clean spans. Assign timestamped ASR words to those slots without a second diarization. Preserve legacy opt-out and no-ASR behavior, and propagate failures. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(audio): enroll portable speaker profiles Gate profile exports with voice-recognition permission and validate registration against metadata from the loaded encoder. Preserve audio enrollment and independent registrations with duplicate display names. Exclude diarization and registration exchanges before API trace capture so persisted traces cannot retain profile vectors or JSON audio. Defer candidate dimensions to trusted loaded metadata. Sort candidates by registration ID so incompatible profiles cannot suppress legacy voices through registry iteration order. Keep portable identity checks closed when trusted metadata is unavailable. Test persisted traces, explicit slot zero, and selection through offline and live transport. Document privacy and the ephemeral registry lifecycle. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(ui): remember speakers from diarization Add a Studio page for diarization and opt-in speaker profiles. Preview clean intervals from the original recording before explicit registration. Join profiles by raw speaker labels, preserve duplicate names, and relabel turns only after a successful save. Discard stale results when the model or recording changes. Share registration metadata with voice management without storing vectors or recordings from this flow. Document permissions and the global, ephemeral registry. Cover enrollment, permissions, previews, and asynchronous races with mocked Playwright tests. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: clarify HTTP speaker enrollment support Replace the stale enrollment limitation with the current HTTP workflow. Distinguish native transport from explicit registration and link its docs. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * chore(parakeet): pin merged speaker profile support Use the merged commit from mudler/parakeet.cpp#80. Its tree matches the previously accepted native pin. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: add diarization enrollment setup example Connect the existing gallery modes to the speaker enrollment workflow. Show installation, private profile export, explicit raw-slot registration, and later recognition without another export. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs(blog): explain diarization speaker profiles Put the diarization walkthrough on the LocalAI website in the feature PR. Cover the three gallery modes, explicit enrollment, and privacy limits. Link setup instructions and keep availability conditional on feature support. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs(blog): focus diarization on everyday use Explain what users can do with recordings before the setup steps. Replace the technical walkthrough with a short Studio guide and link readers to the existing reference for model names and developer use. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs(blog): lead with speaker capabilities Present speaker recognition through everyday uses and a short UI flow. Keep technical reference details in the existing documentation. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(diarization): satisfy Go lint checks Avoid copying protobuf message state when extending backend status, check the multipart reader close result, and document the focused testing.T lint exemptions. Assisted-by: nib:gpt-5.6-sol Signed-off-by: Ettore Di Giacinto <mudler@localai.io> --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
2ae6cae70d |
feat(parakeet-cpp): name speakers from the shared voice registry (#12382)
* feat(voice): list registered voices and record which encoder made them The voice registry could register, identify and forget but not list, and it did not remember which speaker encoder produced an embedding. Add Metadata.Model and Registry.List, answered from the index the store registry already keeps for Forget. Needed so a backend can be given the registered voices that match its own speaker encoder. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(voice): store the encoder model with a registered voice Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(voice): pick the registered voices that match a speaker model Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(proto): carry known voices and speaker names on diarize and live messages Assisted-by: Claude:claude-haiku-4-5 [Claude Code] * feat(diarization): name speakers from the voice registry When a diarization model has a speaker_model option, the endpoint sends the registered voices made by that encoder to the backend. The backend's name and name_score come back as extra fields next to the normalized SPEAKER_NN speaker, and the speakers summary carries the first name seen for each speaker. RTTM output and results without names are unchanged. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(live): pass registered voices to a live session and surface speaker names Live sessions now send the registered voices that match the model's speaker_model to the backend, and each speaker segment carries the name the backend matched. The realtime segment event gains an optional speaker_name field. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(parakeet-cpp): load a speaker model and build per-request voice registries Adds the speaker bindings (ABI v9 and v10, probed separately), the speaker_model, speaker_threshold and speaker_margin options, and a per-request registry builder over the known voices. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(parakeet-cpp): name the speakers in Diarize from the known voices Diarize builds a per-request speaker registry from the known voices when a speaker model is loaded, calls the named C functions, and puts each slot's registered name and score on the segments. The registry is freed on every path. A library without ABI 10 reports Unimplemented instead of dropping the names. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(parakeet-cpp): name speakers in the live scene stream The live scene stream now begins with a known-voice registry when a speaker model is loaded and the live config carries voices, and each closed speaker segment takes its slot's current name from the feed's names map. A segment that closes before its slot is identified has an empty name. The registry is freed after the stream, on every path. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(gallery): speaker naming entries and docs for parakeet-cpp Add three gallery entries that load the WeSpeaker ResNet34 speaker model next to the diarization or realtime scene models, and document speaker names in the voice recognition, diarization, audio to text and realtime pages. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * fix(parakeet-cpp): skip an unusable registered voice instead of failing the request A registered voice with the wrong embedding size, or one the C side refused, failed the whole diarization request, so one legacy voice broke the model for every user. Skip such voices with a warning that does not carry the voice name, and take the plain path when none is left. Also map an exact 0 speaker threshold or margin to a tiny positive value, since the C side reads 0 as "use the default", and fix a stale comment about which contexts Free() walks. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * fix(diarization): warn once per model about voices from another encoder; document the privacy limit The different-encoder warning fired on every request. Log it once per feature and speaker model, then at debug level. Document that the global voice registry lets any caller of a speaker_model model learn matching names, and that skipped wrong-sized voices are logged. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * chore(parakeet-cpp): bump parakeet.cpp to 8c8cec0 (C-API v10) and check speaker naming against the real library The pin moves from 623a968 to 8c8cec0, which brings in everything merged in parakeet.cpp since: the voice identification change (C-API v9, #78) and raw-embedding enroll plus diarize-only speaker naming (C-API v10, #79). New real-library specs (gated on PARAKEET_BACKEND_TEST_SPEAKER_MODEL, _DIAR_MODEL, _WAV and, for the live path, _STREAM_MODEL) name the two speakers of two_speakers.wav from a committed pair of WeSpeaker embeddings, with the voices passed in reversed order. They also check that the float32 threshold reaches C through purego. The shared test loader now registers the v9/v10 and scene symbols as main.go does. The rebase onto origin/master had no conflicts. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
f378fe89d0 |
Merge pull request #12373 from mudler/feat/systemone-capability
feat: decisions usecase for decision models, with gallery tagging |
||
|
|
5613572f38 |
feat(systemone): validate requests and align the docs with Ollama's contract
All three routes now validate the request before it reaches a model: body size (413 over 64 KiB), state, question count, blank ids, option and level counts, and noul criteria keys. Forwarded decision requests skipped this before, so a malformed question surfaced as a backend error. The docs claimed the wire shape matches Ollama's. Field names and question types do; confidence, error shape, keep_alive and state rendering differ, and the docs now say so. Assisted-by: Claude Code:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
70ce62901f |
refactor: name the capability decisions instead of systemone
The usecase describes what a model can do, and the category is the Decisions API. SystemOne stays as the wire contract: the /v1/systemone routes, the Score RPC question_type and the swagger tag are unchanged. The usecase, flag, auth feature, UI label, gallery tags and docs page are now decisions. Assisted-by: Claude Code:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
bd6863af81 |
fix(react-ui): extract text from PDF attachments in chat and home (#12374)
The React UI read every non-media attachment with file.text(). For a PDF that decodes the binary bytes as UTF-8, so the model received raw "%PDF ... stream ... endobj" noise instead of the document. The legacy Alpine UI ran pdf.js; that step was not ported when the React UI replaced it, but both file pickers still advertise .pdf. Add a shared readAttachmentText helper that routes PDFs through pdfjs-dist and reads other files as before. pdf.js and its worker load on first use, so the main bundle does not grow. A PDF that cannot be parsed or has no text layer (scanned, encrypted, damaged) is rejected with a toast instead of being attached as an empty or garbage file. Cover the chat and home paths with Playwright specs that build a real PDF in the test. Assisted-by: Claude Code:claude-sonnet-5-5 [playwright] [eslint] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
b3d65fd538 |
fix(systemone): route NER models to the NER path and refuse decision models on permute and separate
vllm_decide refuses NER architectures and the NER entry point refuses decision architectures, so each model kind 500ed on half of the routes. A token_classify model now goes to the NER path on /v1/systemone, and /permute and /separate return 400 for decision models. Docs and instructions state which kind serves which route. Assisted-by: Claude Code:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
b4852d62d2 |
feat(systemone): register the decisions API on auth and instructions
Adds a default-on systemone route feature for the three /v1/systemone routes and an /api/instructions area for them. No MCP tool is added: the endpoints run inference and are not admin install/edit actions, and the route-map test still passes. Assisted-by: Claude Code:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
36846466e4 |
feat(ui): show the systemone usecase on installed models
Assisted-by: Claude Code:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
e03cf8dac8 |
feat(systemone): refuse models that do not declare the usecase
A chat-only model now gets a 400 naming known_usecases: [systemone] instead of a backend error. Configs declaring no usecases and token_classify models stay allowed so existing laya and GLiNER setups keep working. Assisted-by: Claude Code:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
2fa36e7147 |
feat(parakeet-cpp): speaker diarization, sound detection and live scene events (#12335)
* feat(parakeet-cpp): load diarization and CED models and companions
Repin PARAKEET_VERSION to parakeet.cpp PR #75's head, which adds
parakeet_capi_model_kind (ABI v8). Bind the new diarization, sound
event and combined scene stream C symbols through the same
purego.Dlsym probe pattern already used for the batched JSON entry
point, so the backend still loads against an older libparakeet.so.
Load now classifies the loaded GGUF by role (ASR, diarization or
sound) via parakeet_capi_model_kind and can load up to two companion
models from Options[] (asr_model:, diarization_model:, sound_model:,
paths resolved against opts.ModelPath), verifying each companion's
kind and freeing every context opened so far on any failure. Free
releases the primary and every companion. AudioTranscription now
names the loaded role when it is not ASR instead of a generic model
not loaded error. The dynamic batcher starts only when an ASR context
ends up loaded, primary or companion.
This is groundwork only: the Diarize and SoundDetection RPCs and the
live scene stream that actually use these new roles land in later
commits.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): reset role fields on a failed companion load
loadRoles' freeLoaded only released the C contexts it had opened; it
left ctxPtr/diarCtx/tagCtx and companions pointing at those now-freed
contexts, so a later Free() on the same instance would double-free.
Zero all four alongside the CppFree calls.
Also route AudioTranscriptionStream and AudioTranscriptionLive through
notASRError when ctxPtr is unset but a diarization or sound model is
loaded, matching AudioTranscription: both used to return the generic
model-not-loaded error instead of naming the loaded role.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(parakeet-cpp): add speaker diarization
Implement the Diarize RPC for the parakeet-cpp Go backend, wired to
Nemotron-3-Diarization through libparakeet.so's diarization C-API.
Plain diarization uses parakeet_capi_diarize_pcm; when include_text is
set and an ASR companion is loaded, parakeet_capi_transcribe_and_
diarize_json fills each segment's text instead. Speaker labels are the
decimal index, or "unknown" for -1 (no diarized speaker overlaps).
min_duration_off merges same-speaker segments across a short gap
before min_duration_on drops the segments still too short, then ids
are renumbered. num_speakers/min_speakers/max_speakers/clustering_
threshold have no Sortformer equivalent and are logged at debug
instead of rejected.
Verified against the real Nemotron-3-Diarization + parakeet-tdt_ctc-
110m checkpoints on the two_speakers.wav fixture: correct A-B-A-B
speaker segmentation and matching speaker-attributed transcripts.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(parakeet-cpp): add sound event detection
Wire the SoundDetection RPC to the CED tagger context (p.tagCtx)
loaded by Task 1's role classification. It runs the whole clip
through a one-shot parakeet_capi_sound_stream_* session (window
10s, hop 10s, top_k set to the tagger's class count so every
drained window carries a full score list), averages each class's
score across the drained windows, sorts descending, then applies
the request's threshold and top_k (0 keeps every class).
No tagCtx returns FailedPrecondition; a libparakeet.so missing the
sound_stream symbols returns Unimplemented. Every C call runs under
engineMu, and the stream is always freed, even when a feed or drain
call fails partway through.
Verified against a real ced-tiny-q8_0.gguf on the rooster.wav demo
clip: "Chicken, rooster" tops the list at score 0.91.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): cancel sound detection mid-feed, shrink the lock
SoundDetection now checks ctx before each 10 s feed slice (mirroring
driver.go's feedSlices) and returns Canceled if the caller gave up,
so a long clip can be interrupted instead of feeding to completion
regardless. The stream is still freed on every path, cancellation
included.
Also narrow engineMu to the C calls: the drained JSON document is
now decoded after the lock is released, splitting soundStreamScores
into a locked soundStreamDrain (opts, begin, feed, drain, free) and
an unlocked json.Unmarshal.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(parakeet-cpp): stream speaker and sound events during live transcription
Add two additive proto fields, LiveSpeakerSegment and LiveSoundEvent,
repeated on TranscriptLiveResponse. When a diarization or sound
companion model is loaded, AudioTranscriptionLive now runs a no-ASR
scene stream (parakeet_capi_scene_stream_begin) beside the ASR
streaming session, feeding it the same PCM slices and forwarding any
closed speaker or sound events alongside the matching ASR delta, or
on their own when a slice has no ASR output.
The scene stream is freed and reopened on a mid-stream Config reset,
flushed with is_last before the closing FinalResult, and degrades
gracefully (a warning, not an error) when begin or a later feed call
fails, so live transcription keeps working ASR-only. Existing live
behavior is unchanged when no companion is configured, and no scene
C call is made in that case.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): keep scene events off the ASR critical path in live
Emit each slice's ASR result right after the ASR feed, before the
scene feed for that slice runs, so a companion diarization/sound
model never adds scene compute latency in front of the delta or
<EOU> that drives realtime turn detection. Closed speakers/sounds go
out afterward as their own response, so a slice with both now
produces two responses, ASR first. The live feed log line now
reports ASR and scene wall time separately.
Re-check the diarization/sound contexts a scene stream was begun
with against the live contexts before every feed, under the same
lock: Free() can race between an ASR feed and the matching scene
feed and free the model the stream borrows. A mismatch now returns
without touching the C side. Freeing the stream itself stays
unconditional; the scene stream's destructor only releases its own
buffers and never touches the borrowed contexts.
Also recover a panicking stub inside the live test goroutine instead
of crashing the test binary, and reset the live decode-lag tracker on
a mid-stream config reset, matching what its own comment already
promised.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(realtime): surface live speaker and sound events
Carry the backend's closed speaker segments and sound events
(TranscriptLiveResponse fields 7/8) through LiveTranscriptionEvent
as LiveSpeakerSegment/LiveSoundEvent (nanoseconds mapped to
seconds), and forward them from the semantic_vad live path.
Each speaker segment emits
conversation.item.input_audio_transcription.segment with speaker,
start, end and empty text under the turn's item id. Each sound
event emits conversation.item.sound_detection with one tag
(label, score = peak, index) and the event's new optional
start/end seconds fields, omitted when unset so the existing
unary/windowed sound-detection path is unaffected.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(realtime): keep start/end on a zero-second transcription segment
ConversationItemInputAudioTranscriptionSegmentEvent.Start/End used
omitempty, so a speaker segment starting at 0.0s dropped its
"start" key. Nothing emitted this event before the live scene-event
path, so drop omitempty: the segment always carries real times.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* chore(gallery): add parakeet-cpp diarization, CED and realtime scene models
Add gallery entries for the new parakeet-cpp capabilities: standalone
Nemotron-3-Diarization, the same paired with the Parakeet TDT+CTC
110M ASR model for speaker-attributed text, CED-Tiny and CED-Base
sound classifiers, and a realtime scene bundle combining the
streaming EOU ASR model with diarization and sound companions.
SHA256 taken from the Hub API; licenses from each model card
(openmdw-1.1 for Nemotron-3-Diarization, apache-2.0 for CED,
cc-by-4.0 for the Parakeet ASR models).
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* docs: document parakeet-cpp diarization, sound detection and live scene events
Cover the new parakeet-cpp capabilities across the feature pages:
Nemotron-3-Diarization as a diarization backend (with and without
speaker text, the ignored speaker-count hints, the Sortformer
voice-like-sound quirk), CED as a sound classification backend, the
asr_model/diarization_model/sound_model/diarization_latency companion
options, and the realtime live speaker/sound events (event shapes,
the speech-turn-only limitation, and using this or
pipeline.sound_detection but not both).
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(gallery): correct the realtime-scene license and wording nits
parakeet-cpp-realtime-scene mistakenly copied cc-by-4.0 from the
existing realtime_eou_120m-v1 entry; the model card lists the NVIDIA
open model license instead. Switch to the gallery's usual spelling
for that license and keep the diarization/CED licenses called out in
the description.
Also: audio-diarization.md now says getting per-segment text needs
both an asr_model companion and include_text=true on the request, and
audio-to-text.md's option table reads "Use on" (a pairing the loader
does not enforce) instead of "Allowed on".
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): reject a companion role that duplicates the primary's
loadRoles let a companion option (asr_model:/diarization_model:/
sound_model:) assign into a role field the primary already occupied,
for example asr_model: on an already-ASR primary. The companion's
context silently overwrote ctxPtr/diarCtx/tagCtx, and Free() only
walks those three fields, so the original primary context was never
freed again.
Reject a companion whose role the primary already holds before its
GGUF is even loaded, freeing everything loadRoles opened so far, the
same way a wrong-kind companion is already rejected.
Also warn, rather than silently fall through, when
parakeet_capi_model_kind reports PARAKEET_MODEL_KIND_NONE for a
successfully loaded primary; the primary is still treated as ASR,
matching today's behavior.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): cap live scene sound score retention
sceneBegin started the live diarization/sound companion stream with
the C API's default sound options, whose top_k keeps 5 scores per
window forever until drained. The live scene path never drains sound
scores (only the offline SoundDetection RPC does, with its own fresh
stream), so this window queue on the C side grew for the whole
session's lifetime.
Set opts.Sound.TopK = 0 before starting the scene stream: this
disables score retention while leaving sound event detection (onset/
offset), which the live path actually consumes, unaffected.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): merge diarization segments per speaker, harden Diarize
mergeCloseSegments only compared neighbors in the single start-sorted
segment list, so two same-speaker segments never merged once another
speaker's turn fell between them (A, B, A): the short B segment broke
the adjacency the merge relied on. Group segments by speaker first,
merge within each speaker's own start-ordered run, then re-sort the
result by start so interleaved speakers come back out in timeline
order.
Also harden Diarize's entry points the same way streamFeedDoc/
sceneFeed already are: diarizeCall re-checks p.diarCtx (and, on the
include_text path, p.ctxPtr) under engineMu right before the C call,
so a Free() racing between Diarize's own checks and the lock can no
longer reach the C side with a freed context. When the include_text
call returns NULL, last_error is now read from both contexts and
whichever came back non-empty is reported, since either side of the
pairing can be the one that failed. A WAV decode failure is reported
as InvalidArgument instead of an unwrapped/untyped error.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): harden SoundDetection's engine checks
soundStreamDrain ran every C call under engineMu but never re-checked
p.tagCtx there, so a Free() racing between SoundDetection's own
tagCtx==0 check and this lock could still reach the C side with a
freed context. Re-check p.tagCtx under the lock and return
ModelNotLoaded when it was cleared, mirroring diarizeCall's own
re-check. A WAV decode failure is now reported as InvalidArgument
instead of an unwrapped/untyped error.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* test(parakeet-cpp): cover a mid-session scene feed failure
feedSlicesScene already degrades gracefully when a scene feed call
fails mid-session: it frees the broken stream and carries the ASR-only
session forward. Add a spec covering that path end to end: the scene
stream is freed exactly once, later audio slices still produce ASR
responses, and no speaker/sound events appear before or after the
failure.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* docs: fix the parakeet-cpp companion role table and realtime scene docs
audio-to-text.md's companion option table read "Use on" with a note
that the loader did not enforce the pairing; it now rejects a
companion whose role duplicates the primary's, so restore the
"Allowed on" wording and describe the real enforcement.
openai-realtime.md's live speaker/sound section claimed a mid-stream
session.update resets the companion stream and that it flushes on
session close; neither happens, since the realtime core opens one
live stream (and so one scene stream) per speech turn and closes it
at that turn's commit, with no mid-stream Config in between. Document
that lifecycle instead, state precisely that start/end are seconds
from the start of the turn's own audio, and note that the diarization
model starts a fresh session every turn, so a speaker index is only
meaningful within one turn. The example sound tag ("Rooster", index
17) did not match any real CED label; index 17 in ced-tiny-q8_0.gguf
is "Baby laughter". Replaced with "Chicken, rooster" at its real
index, 99.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): use CED's real index for Chicken, rooster
The scene feed comment and the live test's canned document gave
"Chicken, rooster" index 365. In CED's AudioSet label list it is 99,
which is also what the realtime docs show.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(realtime): call the test event accessor
The scene-event tests range over a method instead of its returned slice.
Call the synchronized accessor so the OpenAI test package compiles.
Assisted-by: Codex:gpt-6
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* chore(parakeet-cpp): pin parakeet.cpp master with sound events
mudler/parakeet.cpp#75 (sound events, scene stream, model kinds) and
#74 (the missing <algorithm> include that broke the image builds) are
on master now. Pin 6dea76a instead of the #75 PR head, and update the
header comment the bump bot reads.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* chore(parakeet-cpp): pin parakeet.cpp with ced.cpp on main
parakeet.cpp #76 moved its ced.cpp submodule from the head of
localai-org/ced.cpp#3 (a branch-only commit) to ced.cpp main, where
#3 landed with an identical tree. Pin 623a968 so the image builds no
longer depend on that branch.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(transcription): carry speaker labels on words and streamed segments
A diarizing backend could label transcript segments, but two paths
dropped the label: TranscriptWord had no speaker field, so live
transcription words and word-level timestamps could not carry one, and
the stream=true transcript.text.done event left the speaker out of
its segments.
TranscriptWord gains an optional speaker (proto field 4, additive).
It flows through the live event and result mapping, the JSON word
output of the endpoint and the CLI, and transcript.text.done now
includes a segment's speaker when there is one. Empty labels are
omitted, so responses without diarization are unchanged.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
(cherry picked from commit
|
||
|
|
d8571a8ee9 |
fix(failover): spill 429 admission rejections to the next target
#12113 changed admission control to reject with 429 instead of 503. failoverWriter only held back responses with status >= 500, so a 429 rejection reached the client and the chain never spilled to its next target. Hold 429 as well. An admission rejection is still flagged and spills without tripping the target. Any other 429 is not retryable, so it is released to the client unchanged. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5-5 [Claude Code] |
||
|
|
9c156656bd |
Merge PR #12285: feat(failover): serve a model name from a chain of local and remote targets
Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
590512d9eb |
feat(distributed): report worker version and show models in node inspector (#12328)
* ci: bump Hugo from 0.146.3 to 0.166.0 The hugo-theme-relearn submodule was bumped to 9.1.x in #12096, which requires Hugo >= 0.165.0. The pinned 0.146.3 broke the docs site build with a template error in alias.html that could not evaluate the Locale field on langs.Language. Bump HUGO_VERSION to 0.166.0 (latest stable) to satisfy the theme minimum and resolve the alias.html template error. Assisted-by: nib:claude-sonnet-4.5 [bash] [read] [edit] * feat(distributed): report worker version and show models in node inspector Workers now send their LocalAI build version and git commit at registration. The controller stores them on BackendNode and exposes them through the existing node list/detail API responses. The node inspector side pane now fetches and renders the list of loaded models (name, state, in-flight) instead of showing only a count, matching what the node detail page already displays. --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
f82efdb43b |
fix(models): fallback to application config default context size in /v1/models/capabilities (#12202) (#12216)
* fix(models): fallback to application config default context size (#12202) Honor appConfig.ContextSize in /v1/models/capabilities when model context_size is unset. * docs(models): explain context size fallback Describe the application default used by capability discovery and preserve the distinction between total context and per-request limits. Assisted-by: Codex:GPT-6 * fix(models): apply the default context size only when context_size is unset The request path applies the application default context size only when a model leaves context_size unset. An explicit 0 or -1 falls through to the backend fallback. The capabilities endpoint now does the same, so it reports the value the backend uses. Assisted-by: Claude:claude-opus-5-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
eec532704c |
fix(functions): honor function_arguments_key when building the tool grammar (#11677)
* fix(functions): honor function_arguments_key when building the tool grammar
All four call sites of `Functions.ToJSONStructure(name, args string)` pass
`FunctionsConfig.FunctionNameKey` as *both* arguments, so
`FunctionArgumentsKey` never reaches the grammar generator.
`ToJSONStructure` writes the two properties into the same map:
property[nameKey] = FunctionName{Const: function.Name}
property[argsKey] = Argument{...}
When `nameKey == argsKey` the second assignment overwrites the first, so a
model configured with `function_name_key` gets a grammar carrying only the
arguments object -- the `{"const": "<function name>"}` constraint is gone and
the grammar can no longer express which function was called.
With `function_name_key: function`, the generated property set collapses from
{"function": {"const": "get_weather"}, "arguments": {...}}
to
{"function": {"type": "object", "properties": {...}}}
Setting only `function_arguments_key` is equally broken in the other
direction: the grammar keeps emitting `arguments` while `ParseFunctionCall`
(pkg/functions/parse.go) looks up the configured key, so the parsed call comes
back with its arguments empty.
The default configuration is unaffected -- with both keys empty
`ToJSONStructure` falls back to `name`/`arguments` for both parameters, which
is why this went unnoticed.
The existing `ToJSONStructure()` unit test already calls the helper with two
distinct keys, so only the call sites were wrong. Extend that test with a case
that keeps both custom keys distinct and asserts the two properties survive.
Signed-off-by: Anai-Guo <antai12232931@outlook.com>
* test(functions): cover configured grammar keys
Route grammar construction through FunctionsConfig so the regression test
covers the key wiring used by every endpoint.
Assisted-by: Codex:gpt-5
* chore: empty commit to trigger workflow approval
Signed-off-by: Tai An <antai12232931@outlook.com>
---------
Signed-off-by: Anai-Guo <antai12232931@outlook.com>
Signed-off-by: Tai An <antai12232931@outlook.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
|
||
|
|
950c271710 |
[router] fix: re-seed the knn corpus index when the vector store comes back empty (#12267)
fix(router): re-seed the knn corpus index when the vector store comes back empty The corpus manager records a store as synced by file fingerprint and embedding fingerprint. The local-store backend behind it is an in-memory gRPC process the model loader may evict (active-backend cap, memory pressure) or the idle watchdog may kill, and relaunch on the next request — empty. The file is unchanged, so EnsureLoaded returned early and the router went blind: every probe fell back with similarity 0 while corpus/stats kept reporting the full count. Measured on a production router (LOCALAI_MAX_ACTIVE_BACKENDS=6, four resident models + two router stores): loading any further backend evicted a store, and the idle watchdog killed both after 15 minutes; /stores/find returned 0 hits against a 100-line corpus file whose stored vectors matched fresh embeddings with cosine 1.000. Two parts, because the knn classifier is built once and cached (GetOrBuildClassifier), so the sync at build time is otherwise the only one for the process lifetime: - corpus.Manager remembers one vector it inserted (probe) and, on the synced path, asks the live index for it. A miss means the index was relaunched — fall through and re-seed from the file (no re-embedding). - The router middleware wraps the knn classifier's store so every lookup runs EnsureLoaded first; the loader gets the raw store, so its probe never re-enters the wrapper. A sync error fails the lookup closed, like the build-time load. Specs: corpus package (relaunched empty store is re-seeded under an unchanged file), middleware (relaunched index behind the cached classifier is re-seeded instead of falling back; the spec is red without the wrapper). The test fake now answers Search for inserted vectors. Folds in the maintainer's follow-up (router-corpus-reseed-after-store-relaunch): reviewed and accepted. Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de> |
||
|
|
52a6d62bbc |
fix: return correct HTTP status codes for saturation and no-nodes-available (#12113)
* feat: return 429 when backends are saturated When backends are at capacity (per-model max_concurrent or the process-wide --max-concurrent-backend-requests ceiling), the response was 503. The OpenAI SDK, litellm, and most agent harnesses key on 429 for rate-limit backoff and treat 503 as a hard error. Both saturation paths now return 429 with the existing Retry-After header and type: "rate_limit_error" in the JSON body. The per-model admission middleware keeps admission_rejected as the code field so existing alerts that match on it still fire. Non-saturation 503s are unchanged: model cold-loading (with progress body), model-load failure cooldown, PII detector fail-closed, and classifier unavailable. These mean "not ready" rather than "busy". Assisted-by: AGENT:regolo/glm5.2 [TOOL] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix: return 503 when scheduler has no available nodes When the scheduler cannot find any healthy node to serve a model — all nodes are full and eviction cannot free a slot, or a node_selector excludes every candidate — the error fell through to 500. A 500 tells clients something is broken when the condition is transient and retryable. The router now wraps these errors with a new ErrNoAvailableNodes sentinel. The HTTP error handler maps it to 503 via applyNoAvailableNodes, following the same pattern as applyBackendAdmission (429). Unrelated scheduler errors (DB timeouts, registry lookups) still return 500. Three return sites are wrapped: - resolveSelectorCandidates: selector matches zero healthy nodes - scheduleNewModel eviction-busy: all models have in-flight requests - scheduleNewModel eviction-failed: eviction itself errored The existing scheduleAndLoad wrapper ("no available nodes: %w") preserves the sentinel through the chain via errors.Is, as does ModelRouterAdapter. Assisted-by: AGENT:regolo/glm5.2 [TOOL] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * test(http): use Ginkgo for admission tests Replace forbidden testing.T calls with Ginkgo and Gomega so the lint check accepts the admission handler tests. Assisted-by: Codex:GPT-6 forbidigo * fix(middleware): show the recorded status for admission rejections The admission audit row now records 429, but the Middleware page still printed a hard-coded 503. Read the status from the event, and update the two package comments that still said 503. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5-5 [Claude Code] --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
afebecc63d |
fix(ollama): report on-disk size for /api/tags and /api/ps (#11989)
Hardcoding size/size_vram as 0 made Ollama clients treat loaded models as free. Prefer ModelFileName+ModelPath Stat when available, and omit size_vram (and size) when the value is unknown instead of emitting literal zeros. Resolve each listed model by its stored ID so tagged variants use their own weights. Fixes #11969 Signed-off-by: lei_lei <imleilei123@gmail.com> |
||
|
|
00bd5ee121 |
fix(failover): skip a target that has been edited into a chain
A chain is checked for nested chains when it is saved, but not when one of its targets is later edited into a chain. Requests then served the inner chain's config as the target, which has no backend and triggers backend auto-detection. Mark such a target missing so no plan picks it, and skip it without a trip in the HTTP path when it is pinned. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5-5 [Claude Code] |
||
|
|
f20cc16033 |
fix(openai): reject an unknown audio response_format as a 400 before the backend runs
Transcription and diarization checked response_format only after the backend had run, and returned a plain error for an unknown value. Failover counts a plain error as a target failure, so one request with a bad response_format ran the backend on every target of a chain and tripped all of them. Check the format before the backend runs and answer 400. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5-5 [Claude Code] |
||
|
|
dae9a431e8 |
Merge remote-tracking branch 'origin/master' into feat/failover-chains
Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5-5 [Claude Code] |
||
|
|
f154bd990a |
feat(system): report per-model DRM VRAM (#12026)
* feat(system): report per-model DRM VRAM Expose optional resident device memory for local backend process trees. Deduplicate DRM clients and omit unsupported or incomplete readings. Document accounting limits and preserve a measured zero in JSON. Closes #11970. Assisted-by: Codex:gpt-6 * fix(system): document trusted procfs reads Scope G304 annotations to paths built from the fixed procfs root, integer process IDs, and kernel directory entries. These reads accept no user-controlled path components. Assisted-by: Codex:GPT-6 gosec --------- Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
5794495a37 |
fix(responses): preserve streamed output items (#12048)
Keep each message and reasoning item at its announced output index. Include the answer in completed responses with reasoning or fallback function calls, and retain reasoning supplied through backend deltas. Add regression coverage for stream indices, final output, plain text, and automatic tool parsing. Assisted-by: Codex:GPT-6 Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
0e52bb657e |
fix(responses): wait for complete JSON tool calls (#12001)
Partial JSON parsing heals a name-only chunk into a tool call. The stream emits that call with empty arguments and skips later chunks. Require complete JSON before emitting terminal tool-call events. Preserve complete calls before an unfinished trailing call, and count only actual tool calls. Add split-chunk regression tests and docs. Refs #11635. The non-streaming report remains unconfirmed. Assisted-by: Codex:GPT-6 Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
27ceda46b0 |
refactor: name the realtime pipeline stages with constants
Review asked for constants instead of the stage literals ("vad",
"transcription", "llm", "tts", "sound_detection") passed to resolveStage,
stageCall and isChainStage, so the uses can be cross-checked. Add
PipelineStage* constants next to the Pipeline type in core/config: the
names match its yaml keys, and core/backend (preload roles) and the openai
realtime endpoint both need them.
Use them in realtime_model.go (stage routing and preload roles),
realtime.go (the voice_recognition preload role) and core/backend
preload.go. model_failover events take their stage from the stageChains
keys, so they now carry the constants too. The failover tests use the
constants for inputs and keep literal wire values in their event
assertions.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
|
||
|
|
8db6c15fe0 |
refactor: name the proxy backends with constants
Review asked for constants instead of the "cloud-proxy" and "localai-proxy" literals. Add CloudProxyBackend and LocalAIProxyBackend next to the other backend-name constants in pkg/model (WhisperBackend, TransformersBackend, ...), which core/config already imports, and use them in every production check: the proxy options builder, IsRemoteProxy, IsCloudProxyBackendPassthrough, the PII defaults, the localai-proxy backend hook and loader warning, and the PII middleware metadata in the routes and the in-process MCP client. The proxy options builder now calls IsRemoteProxy() instead of repeating the two-backend check, so the set of proxy backends is defined once. Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
14c692098e |
fix(openai): answer 400 for a missing or malformed audio upload
TranscriptEndpoint, DiarizationEndpoint and SoundClassificationEndpoint
returned the raw c.FormFile("file") error. Echo turns a non-HTTPError into
a 500, so a request with no multipart boundary or no file field looked like
a server fault. This became visible in tests/e2e once the suite registered
a transcription-capable model (lp-transcription) at runtime: the
default-model middleware then resolves a model, the request reaches the
handler, and "should return mocked transcription" got a 500 in some spec
orders (seed 1790493709).
Read the upload through a small uploadedFile helper that maps any
FormFile failure to 400 with the field name and parser reason. Server-side
failures after that (temp dir, file create, copy) stay 500. The image and
LocalAI upload endpoints already return 400 here.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
|
||
|
|
5ed67be0bf |
fix(failover): skip a target that answers HTTP 501 instead of failing the chain
The non-OpenAI endpoints (depth, detection, face_*, voice_*, images, video, 3d) map a backend's gRPC Unimplemented to an echo 501 without the gRPC status. The retry loop did not see a capability gap, and IsRetryable is false for 501, so the client got 501 and the next target was never tried. Treat a returned or written 501 as a capability gap. Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
7b11af7f7d |
feat(ui): list failover chains and badge chain models
Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
bdd4a73d05 |
feat(ui): edit failover chain targets with a dedicated field
Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
d65b4d3e6a |
feat(ui): show live failover chain health in the model editor
Editing a failover chain now shows its state, the target serving it, and a per-target health table under the editor header. The strip reads GET /api/failover and follows /api/failover/events, with a 15 s re-list to cover SSE reconnect gaps. Admins can pin a target or unpin the chain after a confirmation. Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |