mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-09 22:54:42 -04:00
4046185d554a97bc1dbdb50b6bee6d5cc3edc961
8
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
4d7bdc6ff2 |
feat(ui): redesign the web UI around a shared kit and a calm palette (#12526)
* build(ui): vendor the shared UI kit snapshot at 0.2.0 The restyle needs the kit's tokens, motion layer and component classes. Take a pinned snapshot instead of depending on the kit at build time, and keep a lock file with the version and per-file checksums so a later update shows exactly what changed. The product theme stays outside the vendored directory. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the ink and teal theme and bridge the old variables Define the product colours as the shared UI kit's roles, for light and dark, in theme-localai.css. The kit's contrast check passes on every pair. theme.css keeps the existing --color-* and --shadow-* names but now points each at a role, so App.css and the pages get the new palette without edits. Radii move to the kit scale. index.html now sets data-theme before first paint with the same rule as ThemeContext (stored choice, otherwise dark), because the contract layout of the theme file no longer defaults to dark by itself. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): restyle the shared chrome with the UI kit grammar Adjust the shared classes so every page picks up the same interaction language without per-page edits: - Sidebar sits on the canvas and the current row lifts onto a card. Section labels are tracked uppercase, the badge is a soft pill, and the phone drawer leaves the tab order when closed. - Buttons are flat: hover swaps the surface, press scales to .97, focus is a 2px ring with a 2px offset, danger is a tinted wash. - Inputs use the card surface and the control edge; switches, tabs, filter chips, badges and cards follow the same rules. Cards no longer lift on hover; only linked or button cards react. - Menus and popovers scale in from the trigger corner with 40px items. Dialogs get a veil fade and a spring settle. Toasts become pills at the bottom centre. - The page transition is a 250 ms fade with a 6px rise. It fills backwards so a finished animation no longer leaves a transform that confined dialog veils to the main column. The focus-ring test now checks the outline instead of a box shadow, and new specs cover the theme roles, the first-paint theme and the sidebar lift. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): move leftover hard-coded colours onto the theme roles The YAML editor restated the old blue palette in JavaScript, and a few pages kept literal blues, indigo and violet tints, or fallbacks that only applied because a variable was never defined. Point them at the theme variables so they follow light and dark and the new palette. The status badges in the account pages built their tint by appending "22" to a variable, which is not valid once the variable is defined, so they had no background. Use the wash roles instead. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): raise the type scale and control size toward the kit Body and list text moves to 15px and the rest of the scale follows the kit's 12/13/15/17/21/32/44 steps. Page titles, section headings and stat values are bold with tighter tracking; titles are 32px. Buttons, inputs, selects, tabs and nav rows are 40px high with the 12px radius, compact controls 32px. Tabs become a segmented control. The sidebar widens to 240px (64px collapsed) and nav rows get more room. Identifiers and counts in the split views use the mono face, and the stat grid becomes separate inset tiles. The Geist stack stays: it is bundled, and the thin look came from the size, weight and negative tracking, not the face. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): separate cards, panes and floating surfaces from the canvas Cards, the Models and Installed split panes, the composers and the confirm dialog use a stronger card edge, the rest shadow and the 20px radius, so they read as layers in dark as well as light. Menus and popovers move to a float surface (the hover tone in dark) with the float shadow. The selected rail row gets an accent wash and a 3px accent edge. The send buttons are a clear accent when there is something to send and a quiet inset when not; the Home button carries data-empty for that, since submitting an empty box does nothing. The assistant card becomes an accent wash with a square icon. New surfaces spec checks the pane edge, the selected row, both send buttons and the popover in both themes. The voice library empty-state spec now waits for the layout to settle before comparing two boxes. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): tidy the sidebar header and mark the current row with a dot The header gives the configured horizontal logo a fixed width and centres it in a 72px band, lined up with the nav icons. The collapsed rail shows the configured icon logo centred, and its nav rows become 40px tiles centred in the 64px rail. The current row gets the kit's accent dot, hidden in the rail. The theme, language and account controls stay in the sidebar footer: the app has no global search or command palette to put in a top bar, so a bar would only hold controls that already have a place. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): centre the avatar in the collapsed sidebar rail The collapsed avatar link was set to "flex: 0", which gives it a zero flex basis; with min-width: 0 the link shrank to its padding and the icon overflowed from the link's left edge, about 14px right of the icon column. Use "flex: 0 0 auto" in the collapsed and tablet rail. The footer controls now share the nav icon column in the expanded sidebar too (6px footer padding, 40px control boxes), and the tablet rail gets the same footer padding and hidden language code as the collapsed one. New spec measures the centre x of the nav icons, mark, avatar, language, theme and collapse icons in the collapsed, expanded and tablet states, in both themes, and asserts they agree within 1px. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): size the console and settings rails and stack settings on phones The Operate console rail, the Settings section rail and the account tab bar still used 13px text and the old underline tabs. They now use the 15px nav size, 40px rows and the segmented tab control. Form row labels are 15px with 13px hints. On a phone the Settings section rail sat beside the form and squeezed every row into a few characters. Below 720px the rail stacks above the content as a scrolling row and form rows wrap their control below the label. The save button no longer carries the icon font class, which drew a missing glyph before its label. The language menu is wide enough to keep Bahasa Indonesia on one line. Assisted-by: Claude Code:claude-sonnet-5-5 * build(ui): update the vendored UI kit snapshot to 0.3.0 Take the 0.3.0 snapshot: the sprite now carries the full outline icon set, and the new icons/fa-map.json maps Font Awesome names to icon ids. The map lets the app move off Font Awesome in the following commits. The lock file is regenerated with the new checksums. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add an Icon component backed by the kit sprite Icon draws an inline svg that points into the kit's outline sprite. The sprite is inlined into the page once, so the references resolve under any base path and in the embedded build without a request. Icons size with the font (1em), take currentColor, hide from assistive tech unless given a title, and spin on request. An unknown id draws a neutral circle. FaIcon and iconFromFa resolve Font Awesome names through the kit's map, for names that arrive at run time. iconHtml does the same for markup built as a string. The GitHub and Apple marks are small local glyphs, as the kit ships no brand marks. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): draw shared components and helpers with Icon Replace the Font Awesome elements in the shared components and in the utility modules with the Icon component. Lookup tables now hold kit icon ids instead of class strings. Code-block copy buttons and artifact cards, which build HTML strings, use iconHtml and a sanitizer-safe slot. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): draw model, backend and account pages with Icon Replace the Font Awesome elements on the home, models, backends, import, settings, login, account and users pages with the Icon component. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): draw chat, studio and recognition pages with Icon Replace the Font Awesome elements on the chat, media generation, talk and face and voice pages with the Icon component. The talk status table keeps its spin and pulse states as Icon props. The connected and error states now use a dotted circle and an alert circle, so they differ from the idle ring by shape as well as by colour. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): draw agent, node and operate pages with Icon Replace the Font Awesome elements on the agents, skills, collections, jobs, fine-tune, quantize, nodes, swarm, usage, traces and activity pages with the Icon component. Two class strings on layout elements held leftover button and icon classes from an earlier merge; they are cleaned up so the elements keep only their own classes. Assisted-by: Claude Code:claude-sonnet-5-5 * style(ui): size and align icons for the svg component Icon rules that targeted the font element now target the svg: the descendant "i" selectors in App.css and auth.css become ".lai-icon". The svg is 1.2em with a 2 unit line so it matches the visual size of the old glyphs at the 12 to 16px sizes the app uses, sits on the text baseline, and follows the context font size. Large empty-state marks get a lighter line. Menu icons get a 16px box and the readiness badge icons keep their 20px circle with padding. Add the pulse used by the talk status. Assisted-by: Claude Code:claude-sonnet-5-5 * build(ui): remove Font Awesome No source file references the icon font any more. Drop the package and its stylesheet import. The build no longer ships the solid, regular and brand font files. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): keep focus traps off the svg use references The dialog and drawer focus traps collect focusable elements with a "[href]" selector. An icon's use element carries an href, so it became the first "focusable" element and Tab at the end of the dialog stopped there instead of wrapping to the first button. Match "a[href]" instead. Assisted-by: Claude Code:claude-sonnet-5-5 * style(ui): keep icon sizes overridable and set the line width per svg Give the icon base rule zero specificity so a rule that sizes one icon (nav column, menu box, avatar, language switcher) wins whatever its order in the file. The sprite symbols fix their own line width; the inlined copy drops it so the width set on each svg applies, as the --lai-stroke custom property, and large marks can use a lighter line. Pin the avatar and the language globe to the boxes the sidebar alignment spec expects. Import the map as JSON with an import attribute so Node can load it in the spec. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): select icons by the svg markup and cover the sprite Specs that found icons by their Font Awesome class now select the svg by its data-icon. The dead-icon audit checks that every svg resolves to a sprite symbol and has a size. The class hygiene spec fails on any remaining Font Awesome class. A new spec checks every mapped icon id has a symbol, that the sprite is inlined once, that an icon paints at the root and under a forwarded path prefix, and that Font Awesome names map as documented. Assisted-by: Claude Code:claude-sonnet-5-5 * build(ui): update the vendored UI kit snapshot to 0.4.0 Take the 0.4.0 snapshot: hub tabs with count and attention badges, the six chart series tokens and the grid colour in the theme contract, and sample themes on a calmer palette. The kit headers are renamed and the lock file is regenerated with the new checksums, as for the earlier snapshots. Assisted-by: Claude Code:claude-sonnet-5-5 * style(ui): switch the theme to the calm palette Rewrite the LocalAI theme on the calm palette: a muted teal accent on a near-neutral green-grey canvas, desaturated status colours, no glow and no coloured shadows. The theme fills every role of the shared UI kit's 0.4.0 theme contract for light and dark, including the six chart series and the grid line. The bridge in theme.css keeps the old --color-* names working, adds the dark surface ladder (card, raised, float) and a strong edge, and points the fixed data hues at the chart series. Two values differ from the first sketch. The dark text on the accent fill is #021512 instead of #04201d: it reads 5.58:1 on the fill at rest and 6.4:1 on the hover fill, against 5.08:1 at rest for the lighter value. The light control edge is #6b7d7a. The kit's contrast script passes for all text pairs (4.5:1), control and focus pairs (3:1) and series colours (3:1). Leftovers that no longer fit the palette are fixed: the usage chart takes the six series colours in order, the audio and animation canvases fall back to the new accent, the face box loses its glow, and two gradient fills are now flat. The theme tests expect the new canvas colours. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): replace the console rail with hub tab bars Build and Operate no longer open a second navigation rail beside the page. Each is a hub: one row of the kit's hub tabs above the page, with count and attention badges that scroll sideways on a phone. Every URL and route stays as it was, plus a new /app/build landing page that lists the Build tools with a line each. Build tabs: Overview, Agents, Skills, Memory, Jobs, Fine-Tune, Quantize, Import, Voices (recognition and library) and Faces. Operate tabs: Status, This machine, Swarm (distributed mode only), Runtime (backends, activity, failover), Traffic (usage, traces, middleware) and Settings (settings, users), plus the API link. A tab that holds several pages shows a second row of links, and a sub-page such as a node detail keeps its tab highlighted. The feature and admin gates decide which tabs are drawn, and badges show only values the Operate summary already has. The sidebar lists Build and Operate under a Workspace label next to the Create group. The voice library moves under Build and the model import page gains the Build tab bar. The old rail styles, the rail signals and the console config are removed, and the Operate overview docs describe the tab bar. The specs that drove the rail now drive the tabs, and a new spec covers the tab for each route, gating, badges and the phone layout. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Home as a calm console Home now opens on one command bar: the model chip shows which models are warm, the MCP chip and attach buttons sit beside it, and Send is a solid button with an Enter glyph. Typing "/" opens a grouped, keyboard-driven action list built on the kit command list; every action has a destination in the product. Memory use folds into a one-line strip that opens into the loaded models, with Stop per model and Stop all. It opens by itself while a model is being staged and after a failure, and shows nodes and aggregate memory in a cluster. The list of resident models carries no per-model size because the API reports none. "Jump back in" lists the conversations stored in the browser, one card per day, with j and k to move, Enter to resume and delete with an undo toast. First run keeps the install steps and the recommended models. The assistant prompt is a dismissible line, the library links are one quiet row and the API section is collapsed. Chat accepts an empty new-chat hand-off for /new. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Home console Update the Home specs for the new structure and add specs for the slash menu, the model chip, the memory strip (expand, stop, staging, failure, cluster), the resume list (grouping, j/k, Enter, delete with undo), first run, the send hand-off, a non-admin user and the phone layout. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the fit, disk and cleanup helpers for the models page Pure functions and hooks that the rebuilt Models page reads, with node tests for the rules. modelLedger turns an estimate and the memory budget into one of three verdicts (fits, spills to CPU, over) with the headroom in bytes, and reads the models disk from the resources reading. The disk counts as low under 10 percent or under 20 GB free, and is absent when the server reports none or runs as a cluster controller. cleanupPlan ranks installed models from what the API reports: loaded, pinned, or named by an agent, a task, a failover chain or an alias keeps a model protected; another installed build of the same gallery model is a duplicate; disabled models rank above idle ones. The API records no last use or use count, so none is used. When a lookup fails, nothing is called safe. useModelRemoval holds a removal in the browser for an undo window and sends the existing delete call only when the window ends. Leaving the page drops the batch without deleting anything. The undo toast takes optional labels so other pages can reuse it. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Models as a ledger with a disk strip and cleanup review Explore is one dense table. Each row carries the size, a solid memory bar and the headroom in words ("3.7 free", "+1.5 on CPU", "0.9 over"), worked out from the estimate at the chosen context length. Capability chips show the server's count for each facet, search keeps its meaning and "/" jumps to it, and a density switch (also "d") picks comfortable or compact rows. Selection is a surface step and a check, never a rail. Arrow keys move, Enter installs and Esc closes the inspector, which keeps the fit summary, VRAM by context chart, variants, files, links, tags and licence. A failed install shows its error in the row with a Retry that dismisses the old failure first. A failed or empty listing says which it is, and a host with no GPU is measured against memory and says so. Installed uses the same table with state filters that carry counts, a state per row, Load or Stop on the row, the row menu and the sort by size. Sizes come from the files the gallery lists, so a model it does not know shows a dash. A strip in the header shows the free space on the models disk. It turns amber under 10 percent or under 20 GB free, hides when the server reports no disk or runs as a cluster controller, and opens the cleanup review. Explore says how much an install leaves free. The review ranks installed models as Safe to remove, Probably safe and Your call from real facts only, lists protected models with the reason, and says plainly that usage history is not recorded. A sticky bar shows what a choice frees. Confirming runs a dry run that checks again and lists what will go. Removal waits 30 seconds with an undo; nothing is deleted before that, and leaving the page deletes nothing. The old rail, filter band and popover styles are removed. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Models ledger, Installed table and cleanup review Update the Models, lifecycle, cluster fit, height, search focus and surfaces specs for the table and inspector, keeping what each one checks. New specs, on a shared 41-model gallery stub with three machine profiles: the fit bar and headroom words for a 24 GB card, an 8 GB laptop and a host with no GPU; facet counts, search, "/" and Escape; selection, arrow keys, Enter to install, density; the disk strip when normal, low and hidden; and the states (loading, empty, offline, install failed, phone). Installed covers filters with counts, row actions, the row menu, sizes and sort. The cleanup specs cover grouping, protected models, the honest-data note, the effect bar, the dry run, the undo window, a failed delete, leaving the page, and the phone sheet. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the placement helpers and the estimate hooks The Placement section and the model page need the same few rules, so they sit in plain functions that can be read and tested alone. placement.js holds what gpu_layers, tensor_split and main_gpu mean (unset asks for every layer and the llama.cpp engine trims it, zero is CPU only, 99999999 is the value LocalAI itself writes for all layers), the device list taken from the resources reading, the split by free memory, the part of an estimate that grows with context (read from two lengths, since that term is linear), the fit states with their limit (95 percent of free memory, and the leftover has to fit in system memory too), and a bisection for the largest layer count whose estimate fits. The estimate returns one total and no layer count, so the search runs over 1 to 256 and stops at the first count that no longer changes it. modelWalk.js keeps the order of the list a model page was opened from, in memory and in session storage, for the previous and next buttons. usePlacementEstimate reads /api/models/vram-estimate for a choice, again at twice the context, and with every layer, and keeps readings for the session. useModelPage reads a gallery entry by name, an estimate by context size (from the model's own files when the gallery does not list it), the builds and the loaded models. usePlacementConfig edits the four placement keys of an installed model and saves only what changed. useModelActions is the Load, Stop, disable, pin and remove logic of the Installed table, shared with the model page. MemoryBar is one solid bar with a tick at the capacity of its pool; over capacity it grows past the tick and the tick turns red. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the Placement section to the model editor Run this model on: CPU only (gpu_layers: 0), Auto (the key stays unset) or Custom. Custom takes a number, has an All layers button that writes 99999999, and shows a slider only when the estimate reports the model's layer count, which it does not today. Context size has presets and a number field because the KV cache follows it. With two or more GPUs there is a split (written as percentages, with a button that takes them from the free memory of each card) and a main GPU. A bar per GPU and one for system memory show what other programs use, the model's weights and working memory, and the part that grows with context, with the room left or how far over it is. Under them a verdict in plain words: Fits in GPU, Spills to CPU, Too many layers for the GPU, Runs on CPU only, No GPU found, Not enough memory. It says "slower" and never a multiplier, because the estimate has none. Fit it for me asks the estimate for the largest layer count that fits the free GPU memory and says what it set, with Undo; it is hidden when the estimate is unavailable or the host has no GPU. Loading shows skeletons, an unavailable estimate shows a note with Retry, and a server that schedules onto other machines shows no bars, because its device list is the controller's. The editor shows the section for an installed model, with a link in its section rail. Auto sends null for the key, since a patch only merges, and a null read back opens as Auto. The docs describe the section and what each mode writes. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): open a model on its own page A model has an address, /app/models/<name>, for an installed model and a gallery entry alike. Open it from the arrow at the end of a row, a double click, "o" on the selected row, the inspector's Open details button, or a tap on a phone. The title block holds the main action: Install with a chevron that chooses the build, or Load and Stop with a menu (disable, pin, edit configuration, logs, delete with a confirm). A strip answers whether it fits, what it does and what installing leaves free. Tabs: Overview (about, a memory bar, state, the pages it opens in, and the agents, tasks, chains and aliases that name it); Fit and memory (verdict, context sizes, the bar split into weights and context, and memory by context against the limit, with a data table); Variants and files (builds with size and fit, install any, the files of the chosen build). For an installed model also Usage and history, which says what the API does not record instead of drawing an empty chart, Configuration, which is the Placement section with the file it writes and a link to the full editor, and Logs, the backend log viewer without its page. Keys 1 to 6 switch tabs, [ ] and j k walk the list the page was opened from, Esc or Backspace go back. The list stays mounted behind the page, so Back finds its view, search, filters, selection and scroll as they were, and focus returns to the row's arrow. The page covers loading, an unknown name with the closest matches, the gallery being out of reach, an install in progress with Cancel, and a failed install with Retry. The docs describe the page and its keys. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the model page and the Placement section New specs for the model page: reaching it from Explore, Installed, a double click, "o", a pasted link and a phone tap; the walker and Back with the search, a filter, the selection, the Installed view and the scroll kept, and no second read of the gallery; the title block, the answer strip, tabs by click, keys and arrows; Fit and memory, builds and files with the install call each one makes; an installed model's actions, used-by, the honest usage tab, configuration and logs; loading, an unknown name, offline, an install in flight and a failed one; and the phone. New specs for Placement: every mode and the keys it writes, the slider only when a layer count exists, the context presets, the bars and every verdict, two GPUs, no GPU, a cluster, an unread machine, a loading and an unavailable estimate, Fit it for me and Undo, and the section in the model editor with its save. The phone tap on a row now opens the page, so the two phone specs that expected the inspector as the page check the page and keep the inspector check for a window between a phone and a desk. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): link Studio results and open workspaces from a prompt Each workspace now records the result it was made from (parentId and an edge kind such as take, animate or to-3d) and reads a prompt, model, size, count and source from the query string, so one page can hand work to another. A source result is fetched from the server's own output file and becomes the start image, the picture for 3D, or the audio file. A note on the page says when the source loaded or could not be loaded. Diarization had no history; it now keeps the file name, the model and a speaker count, never the recording. Prompts are cut at 2000 characters when stored. The pure helpers (type suggestion, grouping, lineage layout, favourites, clearing) have node tests. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Studio front page as a composer with your work The front page is a prompt box with a chip per type, a type suggestion from the words, starters, and the options each workspace accepts. Generate opens the workspace with those filled in. A type with no model is a dashed chip that shows a gallery model, its size, memory need and an Install button only when picked; the typed words stay while it installs. Under it, Your work lists results from every workspace as a masonry with filters, counts, favourites and a Clear history action. Results made from each other stack into a project tile and open as a lineage board with a dock for running a new take or branching to the next step; steps the destination cannot start from yet are disabled with the reason. The docs describe the page, what is stored in the browser, and the query parameters a workspace accepts. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Studio composer, your work and the lineage view Specs for the type suggestion, the keys, hand-off to each workspace, the install path for a missing model, the masonry filters, favourites and clearing, stacking, the lineage board, new take and branch, steps that are disabled with a reason, and the phone layout. Existing Studio specs move from lanes to chips with the same intent. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the shared Studio workspace frame and move Images onto it The seven Studio workspaces get one layout: a row of type tabs, a compose card (optional sources as chips, a prompt with starters, a model chip, the essential options as chips, an Advanced fold that names what is inside, the memory the model needs, and one action with the reason when it cannot run), a run area, and a strip of recent results of the type. The run area shows a job card with the time that has passed and an indeterminate bar, because these endpoints report no phase or percentage; a failure with what the server said and one action; or the result with a toolbar: Favourite (the list the front page keeps), Download, Use in (the hand-off targets, disabled with the reason when a destination cannot start from the result), Re-run with edits (the take's values go back in the form, changed fields are outlined and listed) and Lineage. A type with no model shows the install note from the front page. Images is the first workspace on the frame. It keeps its size, count, steps, seed, negative prompt, source image and reference images, and its history writes, including the parent link and edge of a hand-off run. useMediaHistory.addEntry now returns the id of the entry it stored. The docs describe the workspace page. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move Video onto the workspace frame Video keeps its size list, duration, frame rate, steps, seed, CFG scale, frame count, negative prompt, start and end image and avatar audio. The start and end image are source chips, the avatar audio opens the recording and paste input from a chip, and the rest sit in the Advanced fold. A start image from a hand-off shows as a chip with its picture. Results play in the video player with the shared toolbar. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move TTS onto the workspace frame TTS keeps the saved-voice picker for cloning models, the typed voice for the others, the voice library deep link, and the delivery instructions, which now sit in the Advanced fold. The result is the waveform player with the words under it. The stored entry also keeps the voice id so Re-run with edits can select the same saved voice. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move Sound onto the workspace frame Sound keeps its Simple and Advanced modes and every field of both: the description, instrumental, vocal language, caption, lyrics, BPM, duration, key, language, time signature and think mode. The mode switch, instrumental and duration are in the compose card, the rest in a More options fold. The stored entry keeps all of the fields, so Re-run with edits restores the form as it was. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move Transform onto the workspace frame Transform keeps its audio and reference inputs with upload and record, the echo test, the key=value parameters (now in the Advanced fold), the input and output spectra and the three waveform players. The audio that was chosen shows before the run, waiting to be transformed. Re-run with edits puts back the model and parameters and fetches the audio and reference the server kept for that run. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move 3D onto the workspace frame 3D keeps the picture input with paste and webcam, the animation operations a model declares, quality and background, the shape and material steps, guidance and seed, the GLB and animation viewers, the remesh control and the download. A 3D result now has a title from the motion prompt when it has no label, so the strip and the front page name animation results by what was asked. Re-run with edits is shown disabled with the reason, because only a small thumbnail of the picture is kept. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move Diarization onto the workspace frame Diarization keeps its model and recording inputs, the option to prepare speakers to remember, the clean-speech previews, naming and remembering a speaker, and the history entry with only the file name, model and counts. The result now shows a timeline with one lane per speaker, the talk time of each speaker, and the segments with their start time and text. RTTM, SRT (only when the run has text) and JSON are built in the browser from the result. The helpers for talk time, axis ticks and the two text formats have node tests. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): remove the styles and lists the old workspace layout used Nothing renders the two-column workbench, the control column, the old history lists, the generation progress tiles, the TTS voice picker or the result echo any more. The inline-style baseline drops with them. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the workspace frame and one run per type Specs for the type tabs, the compose card and its reason when Generate cannot run, starters, the Advanced fold, the job card with no invented progress, a failed run and its one action, the install note, the strip with its favourites filter, Use in with its disabled steps, Lineage, the parent link, Re-run with edits and its list of changes, deleting and clearing, and the hand-off note. One run through each of Video, TTS, Sound, Transform, 3D and Diarization, the phone layout of all seven, and reduced motion. Existing Studio specs move from the old control column to the compose card with the same intent. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the chat thread with raised user turns, prose replies and one-line activity Your messages are raised blocks on the right at a 760 px measure and the model's replies are plain prose under its name and a warm or not loaded dot. Reasoning, tool calls and their results fold into one quiet line that opens inline into steps. Code blocks carry a Copy button and a Canvas button that opens that block in the canvas, image attachments are thumbnails that open in the lightbox, and files are chips. Per-message actions show on hover, on focus and on the last turn, and a turn takes focus so the arrow keys and C, E, R and B work. A failed reply keeps the text written so far and shows the reason with one Retry action. The Agent chat page keeps the older rules: the new styles are scoped to the chat page and use their own class names. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): use the Home command bar as the Chat composer Chat now ends in the same object as Home: the model chip, the MCP chip, a Canvas chip, the message box with attach buttons, a solid Send and the hint line, with the slash menu on the kit command list. The slash menu lists what Chat can do today (switch model, new chat, conversations, manage mode, canvas, find, settings, export, clear). While a reply is streaming Send becomes Stop, which Esc also presses, and Up in an empty box edits your last message. Attached images show as thumbnails and a line under the bar carries the speed and the token count. HomeComposer takes optional props for this (extra chips, its own slash list, Stop, paste, a stricter Enter); Home passes none of them. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): open conversations from a Ctrl K menu with day groups, undo and a slim header The conversations list opens as a centred menu on Ctrl or Cmd K. It groups chats by day like the Home resume list, shows the model that answered and the time, searches names and message text, and moves with the arrow keys. Enter opens a chat, F2 renames it and Delete removes it. Removing a chat hides the row and shows the kit undo toast; the chat is deleted for good only when the undo time ends. Rename, duplicate, copy and export are on each row, as before. The header is one slim bar: the Chats button, the chat name (click to rename), a context meter when the context size is known, settings and a More menu with rename, duplicate, copy, export, model info, keyboard shortcuts and clear. A dialog lists the shortcuts the page answers to. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): show loaded state, capabilities and fit in the Chat model switcher The model chip in Chat opens the same list as Home, grouped as Loaded now and Installed. Each row says warm or not loaded and marks models that understand images. When the list opens, the page reads the host memory once and asks the server to estimate each listed model at the chat's context size (up to twelve, three at a time), then shows what the model needs and whether it fits: free memory, how much would run on the CPU, or how far over the machine it is. A model with no estimate shows no fit text, and no load time is shown because the API does not report one. A memory bar closes the list. The picker takes the model list from the page when it has one, and useModels can skip its own request. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): move chat settings into a sheet and add find in chat, jump to latest and a wider canvas Settings open as a kit sheet: the system prompt, temperature, top P and top K (each says "model default" until it is changed and has a Reset), the context size with quick sizes and a note that it only drives the meter, Manage mode and Focus mode, the model info for admins with its Edit config button, and Clear conversation behind a confirmation. The old slide-out drawer and the model info panel are gone. Ctrl or Cmd Shift F (or the search button, or /find) opens a search bar over the thread. It marks matches in the messages already on the page, shows "n of m" and steps with Enter and Shift+Enter. Nothing is sent to the server. Jump to latest is a pill above the composer. Esc stops a reply, then closes the search, then closes the canvas. The canvas panel gets the kit look: tabs, a Code and Preview switch, Copy and Download, a full-page layout on narrow windows, and translated labels. The Agent chat page shares it and gets the same look. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add the empty, no-model, loading and phone states to Chat An empty chat opens with the composer under one line, starters to try, whether the model is loaded, and the Jump back in list: the same rows as Home, read from the chats the page holds. With no chat model installed, an install card offers the starter models for this hardware, the gallery and import, and the composer stays so the text is not lost. While a reply waits for a model, a load card shows what the page knows: the phase the server names, the node, the bytes and the time left when the server reports them, and a progress bar. A model that is just not loaded yet gets a plain note, with no invented phases or estimates. The foot warns when the context is nearly full. On a phone the header drops its labels, the model list and the settings open as sheets from the bottom, per-message actions stay in view and the canvas takes the whole page. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Talk as a calm voice page over what the connection really does Talk is one stage and one transcript. The stage has the pipeline chip, the voice and language chips, an outline orb that follows the real microphone and playback levels, a heading and a sentence for the current state, and the controls. The transcript lists You, Reply, Tool and Result lines and can be copied. Session settings (instructions, voice, language, tools, Manage mode and the pipeline's parts) open in a sheet. The states are the ones the code reaches: no pipeline model, idle, connecting, listening, thinking (also while a tool runs), speaking, an interrupted reply (the server cancelled it; a note marks the cut), a blocked microphone, a link that failed during a session, and any other error with its reason and a link to the traces. Push to talk and hands-free are not on the page, so they are not shown. Diagnostics keep their waveform, spectrum and stats, drawn in theme colours. The page text moves into the talk namespace, and the old Talk and visualizer styles and the inline-style count go down with the rebuild. Assisted-by: Claude Code:claude-sonnet-5-5 * refactor(ui): remove the chat styles and strings the rebuilt page replaced The settings drawer, the model info panel, the bubble avatars, the conversation menu popover, the context bar, the recent strip, the staging bar, the file badges and the focus-mode rules have no user now. Their rules, the Chat page's focus class and seven unused empty-state strings are removed. The Agent chat page keeps the shared message, sidebar and input rules it still renders with. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the rebuilt Chat and Talk pages Add a Chat page under features (thread, message actions and keys, the message box and its slash actions, the model list with loaded state and fit, conversations on Ctrl K, settings, find, canvas and the empty, no-model and loading states) and a Talk section to the realtime API page with the states the page shows. Manage mode now turns on from the chat settings or /assistant, and the client MCP steps point at the MCP chip. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): settle the rough edges of the new Chat and Talk pages The undo toast sat under the conversations menu, so the Undo button could not be pressed while the menu was open; the menu, the sheets and the fullscreen canvas now stay below the toast layer. Esc in a rename box saved the text through the blur that follows it; it now cancels. The image viewer closed on Esc only when the page did not re-render on the same key, so its key listener is registered once and reads the latest handlers. Keys on a focused message no longer type their letter into the editor they open, "/" from outside a text field starts a command as it does on Home, and Esc leaves the page's own dialogs alone. Code in the canvas is highlighted for languages that have no preview. The conversations menu drops its key hints on a phone so Clear all stays in view. Talk hides Test tone while connecting and calls a server error "Something went wrong", since the call can still be open. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the rebuilt Chat and Talk pages Specs for the thread layout and the activity fold, code blocks, image thumbnails and the viewer, per-message actions and their keys, a failed reply with its one Retry, Stop and Esc while streaming, the composer and every slash action, the conversations menu (groups, search, resume, rename, delete with undo that ends by itself, one chat left), the model switcher with loaded state, vision and fit text from stubbed estimates, the settings sheet, the canvas panel, find in chat, Jump to latest, the empty, no-model and loading states, the phone layout and reduced motion. Talk is driven over a fake WebRTC link through idle, connecting, listening, thinking, speaking, interrupted, blocked, lost, error and no pipeline, its settings sheet and its phone layout. Node tests cover the message text helpers and the conversation grouping. The existing chat specs move to the new structure with the same intent: the transcript spec now describes the raised turn and the prose reply, and the render smoke accepts Talk's own header. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): keep a bounded run log for agents in the browser The server keeps no run history for an agent, so a run is one task and the events until the agent answers, written to browser storage while the page watches the stream: up to 50 runs per agent, task, step and answer text only. Stored chats from the earlier agent chat page read as runs with stable ids. A run still marked running five minutes after its last event reads as stopped. Helpers read an agent's config into chips, build the list of changed fields against the saved config, hide secret values and offer starting points. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Agents area around runs The Agents page shows what needs a look (work in flight, a run that failed in the last day), then each agent with its model, attached memory and skills, and a strip of its last 14 runs. An agent has its own page: model, tools, memory, skills, instructions, a task box and its runs. A run has an address, shows the thread while it works (steps folded into one line, the tool in use, the answer as it arrives) and settles into a report about a second and a half after the agent answers: task, outcome, follow-ups, evidence and steps, with wide tables opening wider on demand. A failure says in plain words what happened and offers Run again. Create and edit fold into sections with a ready mark and a one-line summary, start from a template or an optional model-written draft, and open a preview sheet with the config as saved and the changes against the saved agent. Status becomes a quiet panel in the same language, and the old chat link opens the agent page. There is no Stop, approval, steer, version or dry-run control, because the agent API has no call behind them. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Agents launcher, agent page, runs and editor Specs for the Now strip and run strip, search and the empty state, the agent page, starting a run, the live thread, settling into the report, the run address across a reload and for a run from another browser, follow-ups with their history, failures, the folding editor with ready marks, templates, the preview sheet with hidden secrets and changes, the status page, and the phone, 1440 and 2560 layouts. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe runs and the new agent create flow Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add helpers for tasks, schedules and job outcomes Reads a cron expression the way the server does (five fields or an @ shortcut), checks it, and puts the common shapes in words. The next run is left out on purpose, because the schedule follows the server clock, which the browser cannot read. Also groups jobs by day, sums the last seven days, and gives each job one outcome line from its result or error. A rerun call starts a new job with the same parameters and media. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Jobs area around runs The Jobs page opens with one sentence about the last seven days, then the tasks (model, schedule in words, last 14 jobs, enabled switch, Run now) and a run history grouped by day. Each row has an outcome sentence and opens to the error or the start of the result with one next action. Deleting a task waits 30 seconds with an undo button. A task opens as a page with its recent runs, its prompt with the gaps marked and its schedule. The task form folds into sections, takes a schedule as a preset or a checked cron expression, warns about prompt gaps the schedule does not fill, and has a preview sheet. A job opens as a document: task, outcome, delivery and the recorded steps; a failed job says what happened and offers Run again. Run now now sends attached media through the job call, which is the only one that takes it. "Clear History" only ever cancelled running jobs, so it is now called Stop running jobs. Webhook headers of a saved task show as JSON instead of [object Object]. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Jobs page, task pages and job pages Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the Jobs page and the task form Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add helpers that say who uses a skill or a collection An agent loads a skill when skills are on and the skill is in its selection (an empty selection means every skill). It reads the one collection that carries its own name, when its knowledge base is on. The helpers derive that from the saved agent configs, build the config that adds or removes a skill or a collection, and estimate tokens as characters divided by four. Removing the last selected skill switches skills off, because an empty selection would mean every skill. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Skills and Memory as one library Skills and collections sit in a list with an open item beside it. Each row says who uses it, read from the saved agent configs, or says it is not used yet. Chat reads neither, so it is never named. An item opens in a pane with a Used by strip (names link to the agent, a small x removes it, with undo) and an Add to menu that shows what the addition costs. A collection can be added only to the agent that carries its name. The Memory pane searches the collection alone and shows ranked passages with scores, lists web sources with their refresh interval and the files, shows the server message when an upload fails, and names the endpoints and where files stay. The Simulate a message sheet runs a collection search and shows an agent's skills with a token estimate. It runs no model. The collection details route now opens the same page. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): show use and cost in the agent form pickers Each skill in the agent form says which other agents use it and what it adds to every message, with a total for the selection. The memory section names the collection the agent reads. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Skills and Memory libraries Specs for the used-by lines (including an agent that uses every skill), the filters, search, add to agent, remove with undo, the last-skill case, an unreadable agent list, the empty states, git repositories, the Memory question box, sources, uploads that fail, the Simulate sheet with the parts the API can run, the agent form hints and the phone layout. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the Skills and Memory libraries Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add helpers for the Operate status page and backend rows Pure functions for the parts that need rules. They work out the memory pool the page measures (GPU memory, system memory, or the workers of a cluster that are answering), which pools are too full, the headline and the four ledger rows, the geometry of the capacity chart, and what removing a backend would leave without a runtime (models name their backend, and a meta backend names the concrete one it points at). A second set says what a backend row states: installing, queued, removing, failed, update available, current or absent. LocalAI keeps no memory history, so the chart reads a bounded buffer of readings the page took itself and says so. A reading with no total is dropped rather than drawn as zero. Two hooks are shared by the pages that need them. One retries a failed operation after moving the failure into the record. The other holds a cancel for an undo window, because the server cannot take a cancel back. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Operate Status, This machine, Backends, Activity and Logs Status opens with one sentence ("2 things need you", or "Everything is running") and four rows: Needs you, Capacity, Running now and Recent failures. A row with a problem opens by itself and holds the button that deals with it: Update a backend, Retry or Dismiss a failed operation, Unload a model. A quiet row stays one line. A new installation gets a first-run screen, a cluster sums the memory of the workers that are answering, and a page still waiting for an answer says so. The chart under the rows is drawn from readings the page took while it was open and is labelled that way, because LocalAI keeps no memory history. This machine leads with GPU memory as one bar, then host memory split by running model, then VRAM, RAM, CPU and disk with a bar each. The running models become a kit table with the same menu and stop dialog. Backends is one list with Installed and Catalog views. A row says what the backend is doing (a progress bar with Cancel, Queued, Failed with Retry, Update 1.2.0, Current), carries the one button that matters, and opens in place. Removing a backend names the models and the meta backends that would stop working. Check for updates, Update all, From URL and a first-run recommendation for llama-cpp are in the header. Activity keeps its three sections as quiet rows. Cancel waits eight seconds with an undo toast, because the server cannot take a cancel back; a cancelled install can be started again from the record. Logs gets a process list, a picker, stream and text filters, Follow and Times switches, and a Clear with an undo window. Not shown, because the API has no data for them: GPU temperature and power, a size per backend, an earlier version to roll back to, a dependency lookup beyond the models and meta backends that name a backend, and models that failed to load. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover Operate Status, This machine, Backends, Activity and Logs New specs for the Status headline and ledger (healthy, needs attention, one thing, a full memory pool alone, loading, first run, cluster), its actions (Update, Retry, Dismiss, Unload with its dialog), the capacity chart built from readings taken while the page is open and bounded, the phone layout, no coloured edge on a row, and reduced motion. The Backends specs cover the two views, install progress with Cancel and its undo window, Retry on a failed install, Update, Update all, Check for updates, removal with the models and meta backends it would break, Install from URL, the first-run recommendation, a cluster, and a phone. Activity gains cancel with undo, Cancel now, a second cancel, leaving the page, progress, and starting a cancelled install again. Logs covers the stream and text filters, Follow, Times, Export, Clear with undo, the process picker and list. This machine covers the GPU strip, several GPUs, no GPU and Add a machine. Existing specs keep their intent and follow the new structure: rows open in place instead of in a pane, Update replaces Upgrade, the notice spec now pins that an update is a row state and not a banner or a rail, and a cancel waits for its undo window. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the rebuilt Operate Status, Backends and Activity pages Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add helpers for the Swarm pages Pure functions for what the pages work out from the cluster API: a node's state in words, which nodes a placement rule may use, what a rule would ask for, what a drain or a lost node would leave without service, the nodes a bulk backend update reaches, and the join commands for a worker, a peer instance and a memory shard. Hooks read the roster, the loaded replicas and the rules. Everything runs in the browser from data the page already holds, and says when it cannot see free memory or disk. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Swarm hub: nodes, node page, placement rules, failover Nodes is a sortable table with comfortable and compact rows, a Needs attention filter by reason, a map of the cluster that is not drawn on a phone, the running models, and a bulk backend update for the nodes that drifted. A node is a page: state, vitals, a drain preview computed from the loaded replicas and the rules, tabs for models, backends, logs and capacity and labels, and Remove that asks for the node's name. Placement rules are written as sentences, show where each model is loaded now, and edit in a side sheet with a preview of the nodes a draft could use. Deleting a rule waits a few seconds so it can be taken back. Failover keeps its chains, adds what the router does when a worker stops answering and a per-node preview of what would stop. Add a node covers a registered worker, a peer instance and a memory shard, with a command to copy and a live line that says when the machine arrived. P2P keeps its page in the same vocabulary, and the node logs page follows the local logs page. Previews are labelled as worked out in the browser. Per-GPU readings and node events are not drawn because the API does not return them. Failover moves to Swarm when distributed mode is on. Legacy fleet components and their styles are removed. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Swarm hub New specs for adding a node (each join method, the command, copy, waiting and found, approve, a single install, P2P, a phone), placement rules (sentences, where models are loaded, the preview matrix, the sheet and its preview, delete with undo) and failover on a cluster. Node detail covers its tabs, the drain preview and its dialog, resume, remove with the typed name, a node that stopped answering, and unload. The nodes specs follow the new structure and keep their intent: the table, filters, grouping, pagination, bulk actions, the map, and running models with stop, logs and the loading, error and empty states. The scheduling, failover, P2P, hub and smoke specs follow the renames. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the rebuilt Swarm pages Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): add helpers for the Traffic pages Pure functions for what the pages work out from the usage ledger, the trace summary, the trace buffers and the resources reading: the shared time window, grouping, sorting and filtering of usage rows, chart series and axes that start at zero, the overview figures, per-model statistics, the state of a trace and the words for a failure, the backend operations that ran during a request, CSV export, the Prometheus metric list and scrape config, and a bounded buffer of host readings. A figure whose source cannot say is null, never zero. The trace summary call takes the window in hours, and a helper reads /metrics with its status. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Traffic hub: overview, usage, models, host, traces, middleware Traffic opens on an overview: five figures (requests, failed, p95, tokens in and out) and three charts, each naming its source. A second row of links reaches Usage, Models, GPU and host, Traces, Middleware and Prometheus, and one time window is shared by the first three. Usage groups by model, user or API key, filters, sorts, opens a row on its own chart, exports the rows it holds as CSV or JSON in the browser, and keeps the opt-in cost estimate and the quota forecast. A user who is not an admin sees only their own numbers. Models joins the ledger, the backend-operation buffer and the loaded models. GPU and host shows the current reading and two charts of readings taken since the page opened. Traces gets filters, a settings strip and an explained off state. An API request is a page: the error LocalAI recorded, a timeline with the backend operations that ran meanwhile, and bodies that stay closed until revealed. Middleware draws the pipeline as five steps and shows the rules of the selected step. Prometheus documents /metrics, checks it against the server and gives a scrape config to copy. Alerts is not built: LocalAI has no alert rules. Per-model latency percentiles, GPU utilisation and compare with the previous period are not drawn because the API does not return them. Legacy usage, trace and middleware styles and the usage source components are removed. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Traffic hub New specs for the overview (figures, charts with a data table and arrow key readout, failed and first-run and tracing-off states, the shared window, a phone), usage (group by, filters, sort, export, cost, quotas, a non-admin, empty and loading), models, GPU and host (snapshot, the since-opened labelling, a cluster), the traces list, a trace page (the real error, the timeline, reveal, no headers, a trace that left the buffer), Prometheus and the Middleware pipeline, with shared fixtures. The usage, traces, middleware, hub and smoke specs follow the new structure and keep their intent. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the Traffic hub Add an operations page for the Traffic tab: which record each page reads, what it leaves out and why, the trace page and its reveal, the GPU and host readings kept since the page opened, and the Prometheus endpoint. Link it from the operations index, the tracing page and the middleware page. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): let metric names wrap in the Prometheus table on a phone The long metric names pushed the type and "on this server" columns out of view. Names now wrap inside the table. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Settings with groups by intent, search, a pending bar and history The fifteen sections become eight groups by intent: memory and models, speed and defaults, backends and galleries, access and security, debugging and traces, agents and responses, swarm and sharing, look and feel. Search covers names, descriptions, keys and the old section name, and says where a result used to be. Edits wait in a bar with Discard, Show diff and Apply. The diff lists old and new values and the checks the browser can make: durations parse the way Go parses them, a GPU memory budget is one the server accepts, a gallery box holds JSON, and warnings repeat what the handler and the field text say. Apply sends only the changed keys. Undo saves the previous values again; it is a new save, not a rollback. History lists the changes applied from this browser, since LocalAI keeps no settings log, and Revert stages the old value. A value is marked as changed only where the built-in default is known from the CLI defaults. A row says "Applies now" or "Needs restart" only where the handler or the docs say so. Three things were wrong before and are fixed with the rebuild: the gallery boxes and the shared API keys box were sent under names the server ignores, the "Enable CSRF Protection" switch showed the disable flag the wrong way round, and every save restarted peer-to-peer networking because every field was sent. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Users and keys, Account, sign-in, invite and the 404 page Users and keys is a tabbed page under the Settings tab: people, invites and API keys. The people table filters by state and role, sorts, approves or disables (disabling offers an undo that sets the status back), and opens a side sheet for one person's features, model allow-list and limits. Role, password reset and delete sit in the row menu; delete asks for the name. Invites choose a lifetime of 1, 7 or 30 days and show the link once. API keys can be created with a lifetime, are shown once in full, can be paused, and are revoked after a ten second undo window in which nothing is sent. LocalAI lists keys only to their owner, so the tab shows the signed-in person's own keys and says so. Account has Profile, Security, API keys and Usage. Usage shows the last 30 days, tokens by model and the limits an admin set. The Security tab now shows for a GitHub or SSO account and says the password is not theirs to change. Sign-in asks for one field per step and draws a provider button only for a provider /api/auth/status lists. It has the notice for a sign-up that waits for approval, the first-admin screen, the key-only screen and the invite page. An address outside the app now gets the 404 page too, which names the address and lists the places the sidebar lists, with the same gates. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover Settings, Users and keys, Account, sign-in and the 404 page Settings: groups, search by name, key and old section, the changed marker only where a default is known, apply hints, the pending bar and diff, the checks, apply sending only changed keys, undo as a second save, discard, history, the CSRF inversion and the gallery and API key wire forms, and the phone layout. Users and keys: the table, filters, sort, approve, disable with undo, the row menu, the access sheet, invites, key creation with a one-time reveal, the ten second revoke with undo and with a page leave, and the non-admin redirect. Account, each sign-in variant (error, pending, first admin, key-only, invite, provider buttons) and the 404 page have specs too. Fixtures are shared with the screenshot scripts. Existing specs follow the new structure. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the rebuilt Settings, Users and keys, Account and sign-in pages Runtime settings: the eight groups and where each old section went, search, the pending bar, the diff and its checks, apply, undo, the history, and which settings show a default or an apply note and why. Authentication: the sign-in screen variants, the Account tabs, key lifetimes, the one-time key reveal, the revoke undo window, and the fact that keys are listed only to their owner. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): let the Settings undo toast stand alone and read back a generated P2P token The saved message and the undo toast sat on the same spot at the bottom of the page. The undo toast now carries the saved message. A new P2P token is made by the server when the page sends 0. The page reads it back after the save so the field shows the token and not the placeholder. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): fit the users table, API keys and Account figures on a phone On a phone the users table dropped its Role and Status columns off the screen edge with the row actions. The role and state now sit under the name, so the actions stay in view. API key rows no longer put the key icon on a line of its own, and the three Account figures keep one row. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the phone users table, reduced motion and the empty Account state Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): drop the apply note from three settings the save handler does not mention Size-aware eviction, automatic backend upgrades and development backends said Applies now, but nothing in the handler or the docs says when they take effect. A row now carries a note only where the code or the docs say so. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild Voices and Faces as one identity family Voices is one page with three tabs: Speakers (voiceprints for recognising who is speaking), Speech voices (the text-to-speech reference library, kept apart because it is a different store) and From a recording (a link into the diarization workspace). Faces uses the same layout. Who is this and Same person? give the answer in a sentence with the real distance and cut-off, a word for how far inside the cut-off it sits, and a distance scale with the cut-off drawn on it. The cut-off slider re-reads the answer in the browser; the identify call sends the cut-off, and verify uses the threshold the model returns. The old confidence percentage is gone because it is not a probability. The server has no list call, so the people list stays in the browser and the page says so. After a search that asked for more people than it got back, a saved person the server did not return is marked, and people the server returned that the browser does not know are listed. Nothing is claimed from a short or cut-off search. Enrolling is a sheet: sample, name, labels, permission. A copy of the sample in the browser is opt-in, and an administrator can also keep the recording as a speech voice in the same step. Removing a person waits ten seconds behind an Undo toast and sends nothing before then. Errors say what happened (no face found, model missing, call failed), a blocked or missing microphone is explained, and a missing model or a missing permission renders a page that says what turns the feature on instead of a redirect. Analyze, detect and raw embedding move under More tools, with attribute guesses off by default. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Voices and Faces pages Specs for who is this (match, no match, working, failed, missing model), the cut-off slider, same person, a blocked, allowed and insecure microphone, the registry notes and the not-on-the-server marks, the enrol sheet and its opt-in copy, delete with undo on a fake clock, the disabled and no-permission states, the phone layout, reduced motion and Faces. Existing library and diarization specs follow the new structure and keep their intent. Node tests cover the distance words, scale layout, stored list and error mapping. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the Voices and Faces pages Add a WebUI section to the voice and face recognition pages: the two tools, the cut-off, what the people list is and why it can be stale, the undo window, and what is stored where. Point the Voice Library and Fish Audio notes at Build, Voices, Speech voices. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): rebuild the Build landing, Fine-tune, Quantize, Import and Explorer The Build landing says what each tool is for and what it needs from the machine: the installed backend, the GPU memory, RAM and disk the server reports, and a job that is running or the newest one when it failed. A tool that cannot run says why and what enables it. Fine-tune and Quantize share one page: set up, a check list that is redrawn as the form changes, a run view with progress, stages and a log, and a result with real next steps (export, import, chat, Models). The checks state only what the server reports. A job needs no estimate the server cannot make, so none is invented. Stop on a fine-tuning job asks whether to keep a checkpoint, a failed job shows the server's message, and a memory failure offers two changes that are applied to a copy of the setup. Import is a guided flow: source, review, import, done. The server returns no preview before an import starts, so the review reads the spelling of the source, prints the request the form will send and runs the checks that can be made early. The estimate that arrives when the import starts is set against free memory and disk. The ambiguity picker and the Write YAML tab stay. Explorer shows what GET /networks returns and lists a swarm with POST /network/add, with a join sheet that carries the token and commands. Build tools the account may not use say so instead of redirecting. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the Build landing, the tool pages, Import and Explorer Specs for the landing (a tool ready, missing a backend, with no GPU, a running or failed job, a feature switched off, a member without admin, phone, reduced motion), the shared tool pattern for Fine-tune and Quantize (set up, live checks, start request, running with progress, chart and log, the stop choice, failure with the server message, finish with next steps, earlier jobs, the account-disabled page, phone), Import (source detection, review, checks, ambiguity, running with the estimate against free memory, done, Write YAML, phone) and Explorer (list, join, list a swarm, empty, not an explorer, retry, phone). Existing specs follow the new structure and keep their intent. Node tests cover the machine facts, tool status, checks, log lines, source detection, the import request and the join commands. Assisted-by: Claude Code:claude-sonnet-5-5 * docs: describe the Build tool pages, the import flow and the Explorer Fine-tuning and quantization now describe the set up, check, run and result steps and what the check list can and cannot say. The import section explains the review step and why the size and memory appear only after the import starts. The distributed page describes the Explorer list, the join sheet and what listing a swarm publishes. Assisted-by: Claude Code:claude-sonnet-5-5 * feat(ui): turn the hardware recommendations into a "Best for this machine" shelf The shelf in the Models inspector put five columns into a 400 px pane, so long model ids wrapped letter by letter underneath the size and the memory figures. Each row now stacks the tag, the id and the size and memory facts beside one Install button, and the id wraps inside its own column. Once a model is installed the shelf narrows to the best fit and keeps the others behind a "N more that fit" toggle. Specs cover the ranking, the layout, the narrowing and the install request against a gallery fixture that carries the 4K estimate the shelf sizes against. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): quiet the Studio tab markers and say their state in words The type tabs drew a saturated green dot for every modality that has a model. The dot now uses a text colour, filled when a model is installed and hollow when none is, and each tab carries "(model installed)" or "(no model installed)" as hidden text so the state is not only a colour. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): stack the model editor empty state and drop its section hues "No fields configured" sat in a flex row, so the icon, the title and the text ran together. It now uses the stacked empty-state layout. The section icons took a different status colour each (amber, red, green); they now share one quiet colour, with the accent on the current section. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): keep the Home memory sentence whole on a phone and drop side rails On a 390 px screen the memory strip clipped "2 models loaded" to make room for the figure. The sentence now takes the first line and the figure and device wrap under it. The sweep also removed coloured left rails from the editor section rail, the skill editor list, the install strip and the audio transform notice (now an outlined note), plus unused chat rules that carried rails and two glow animations that nothing referenced. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): line up hub pages, list the model templates, and stop clipped text Medium-width pages inside Build and Operate were centred while the tab bar above them was flush left, so the title started 60 px right of the first tab. They now start at the bar's edge. Add Model offered nine templates as a grid of identical cards with chip clouds and inline styles. It is now one list of rows, each with the field names it fills in on a single muted line. Two clipped strings are fixed: the Studio voice field cut its placeholder mid-word, and the phone job list ended the schedule line in an ellipsis. The recommendation shelf also separates size and memory with a dot, and the docs describe the shelf. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): keep the hidden Studio tab state inside its tab The hidden state text added to each type tab was absolutely positioned against the page, so on a phone it sat outside the scrolling tab row and widened the page by hundreds of pixels. The tab is now the containing block. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): cover the on-disk sizes on the Installed table, model page and cleanup sheet The fixtures stub GET /api/models/storage. The default report is empty, so existing specs keep the gallery estimates. makeStorage() builds a report from files and the models that use them, the way the server does, and storageSpec() is a models directory with shared and missing files. New specs cover the Size column and its shared line, the fallback when the call fails or the user is not an admin, the files list on the model page, a missing file, the bytes a removal frees with shared files, and the cleanup findings. Node tests cover the storage helpers and the batch arithmetic. Assisted-by: Claude Code:claude-sonnet-5-5 * test(ui): wait for the page before pressing keys and ticking the clock Two specs failed in loaded full runs and passed alone. The Alt+1 to Alt+7 spec pressed a key before the composer had armed its key handler. The capacity chart spec advanced the fake clock before the poller had mounted, so it counted fewer readings than it expected. Both now wait for the page to mount. The key spec retries a press that lands during a re-render, and the clock spec advances in small steps and polls for the row count. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): hide the Installed footer when the storage report is empty An empty report from the storage call made the footer read "0.0 GB on disk" next to sizes taken from the gallery estimate. An empty report says nothing about the disk, so the footer now shows only the model count. A spec covers it. Assisted-by: Claude Code:claude-sonnet-5-5 * fix(ui): keep Explore pane actions inside the pane The inspector actions sat in a non-wrapping flex row beside the title, so the buttons ran past the pane edge once it got narrow. The row now takes its own line and wraps. The primary action (Install, Retry, Open) comes first. Manage installation becomes a ghost button, and Open details moves to the end of the row, so one action stands out and the others are quiet. No action or test id is removed. Add a spec that checks, in light and dark at several widths and with a pane forced to 320 px, that every action stays inside the pane box and that the pane keeps its inner padding. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * refactor(ui): drop the type chips from the Studio composer The Studio tabs and the composer's type chips listed the same seven modes, so the page said the same thing twice. Keep the tabs as the one place to switch modes. The composer now shows the type it will open as a small label in its header. The type suggestion from the typed words stays as the quiet hint line under the prompt, and Alt+1 to Alt+7 still pick a type. The composer root carries data-type, data-types and data-missing so tests can read the state. Specs pick a type through a shared Alt+digit helper and read a missing model from the tab dot instead of a chip. Remove the unused chip locale strings and CSS. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * refactor(ui): remove the section crumb above page titles Page headers drew a small uppercase crumb with a short rule before it above the title. On the hub pages it repeated the hub name, so Build sat above a heading that also said Build. PageHeader now renders only the title, the supporting line and the actions. Drop the eyebrow prop, the route-derived section name, its CSS and the unused section helper, and remove the explicit eyebrow props from the pages that passed one. Pages stay reachable through the sidebar and the hub tab bar. Add a spec that checks several pages show their title with nothing ahead of it in the header. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * refactor(ui): remove left accent rails from tiles, rows and quotes Several surfaces marked state with a coloured strip on the left edge. Replace each one with a cue that is not a rail: - Stat cards lose the strip; the icon and value still carry the colour. - The highlighted card is a raised surface with a firmer edge. - The selected rail row is an accent wash with a hairline outline. - The status stripe on rail items is a small status dot. - The active failover row is a tinted row. - Quotes in markdown and chat prose are italic instead of barred. - The variant detail panel has a full hairline border. Add a spec that walks the main routes in light and dark and fails on a left border thicker than 1px, a sideways inset shadow, a narrow absolute strip in ::before or ::after, or a narrow tall child pinned to a left edge. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com> |
||
|
|
4a1089be35 |
chore: ⬆️ Update 0xShug0/audio.cpp to cf124a67cc55d8f65a9a15eec69edbff0fb212c8 (#12421)
* ⬆️ Update 0xShug0/audio.cpp Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> * fix(audio-cpp): mirror the turn detection task The upstream enum adds TurnDetection, so the exhaustive conversion fails to compile. Extend both conversions and preserve the RPC admission rules. Test that turn detection cannot route through VAD or another existing RPC. Assisted-by: Codex:gpt-6 --------- Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
3d586cc3c5 |
fix(audio-cpp): forward voice reference transcripts (#11997)
Saved voices send ref_text, but Fish Audio requires reference_text. Derive the canonical parameter while preserving explicit overrides. Both TTS modes use the shared builder. Add regression cases and document the parameter alias. Assisted-by: Codex:gpt-6 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
e7306a087a |
feat(audio-cpp): AUDIOCPP_DEFAULT_BACKEND fallback for models without a backend option (#12133)
Models whose options carry no explicit backend: open their session on the CPU backend even in accelerator images. The gallery entries carry backend:best since #11892; this covers hand-written model configurations the same way, per deployment: the environment variable supplies the fallback, an explicit backend: option always wins (merged beside the existing threads and maingpu fallbacks), and validation reuses the option parser. Assisted-by: Claude:claude-fable-5 Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com> |
||
|
|
d1ceeaa99a |
feat(gallery): consolidate 25 pending gallery PRs (#12124)
Consolidates all 25 pending gallery bot PRs into a single merge to resolve the conflict cascade — every PR branched from a different point in master and they all touch gallery/index.yaml, so merging them individually was blocked by constant conflicts. Changes: - gallery/index.yaml: +739 lines (new model entries and fixes) - docs/content/features/model-gallery.md: +116 lines (new model docs) - docs/content/features/audio-cpp.md: +12 lines (Sortformer checksum fix) Entry count: 1597 -> 1890 (293 new entries, no duplicates, YAML validated). Supersedes: #11986 #11992 #11994 #11996 #11999 #12002 #12017 #12019 #12021 #12025 #12027 #12029 #12032 #12036 #12037 #12038 #12041 #12042 #12043 #12047 #12050 #12064 #12065 #12066 #12118 Assisted-by: MAKI:regolo/glm5.2 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
878b99384f |
feat(audio-cpp): add ROCm backend image
The pinned audio.cpp revision supports HIP, but LocalAI neither builds a ROCm image nor accepts its backend option. AMD hosts therefore fall back to the CPU image. Build and publish the HIP variant, connect it to AMD capability selection, and accept both upstream HIP names. Assisted-by: Codex:gpt-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
e8546965c7 |
fix(gallery): default audio-cpp models to backend:best (#11892)
* fix(gallery): default audio-cpp models to backend:best The audio-cpp engine creates its session on the CPU backend when no backend option is given, so every gallery model ran CPU-only even on machines where a CUDA/Vulkan/Metal device was registered. backend:best selects the best available backend and falls back to CPU. Assisted-by: Claude:claude-fable-5 Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com> * docs(audio-cpp): explain gallery device selection Document automatic compute backend selection and the CPU override. Assisted-by: Codex:gpt-6 Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com> --------- Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com> Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
9c85cacfe3 |
feat(audio-cpp): add the audio.cpp native backend (#11141)
* backend(audio-cpp): add the native build scaffold Links 0xShug0/audio.cpp engine_runtime through its public framework headers and serves Health/Status. Model loading and the audio RPCs follow. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): keep the build-tree rpath at $ORIGIN Upstream sets CMAKE_BUILD_WITH_INSTALL_RPATH in its own directory scope, so CMake was appending its build-tree library dir to our target and baking an absolute build-host path into the shipped binary. Set BUILD_WITH_INSTALL_RPATH on the target so a package that forgets to bundle libggml*.so fails on the build machine too, instead of only on a user's box. Also document why EXCLUDE_FROM_ALL must stay on the add_subdirectory call, correct the claim that Ubuntu ships no gRPC CMake config, stop the pin comment from repeating the assignment token that bump_deps.sh rewrites, and make test-engine fail rather than pass when no test is registered. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): parse namespaced model options Splits option entries on the first colon so path values survive, and routes load./session. prefixes to the upstream load and session option maps. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): reject out-of-range numeric model options std::atoi is undefined once the digits exceed long and in practice wraps, so device:2147483648 was accepted and handed the ggml backend selector a device index of -2147483648 from a function whose error text promises a non-negative integer. Parse with strtol and reject on ERANGE, on a value above INT_MAX, and on any unconsumed trailing input. The error strings are unchanged. Name the whole entry in the unknown-key error too: an entry like ':value' has an empty key and left the user nothing to grep for in their YAML. Tests look keys up through a helper instead of map::at, so a prefix off-by-one fails one named check rather than aborting the binary and skipping the rest of the suite, and cover the overflow, negative and non-numeric paths. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): route LocalAI RPCs onto audio.cpp tasks Task-major resolution over the family's advertised capability set, with the voice-reference and instructions signals selecting cloning and voice design, and a streaming-to-offline fallback for server-streaming transcription only. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): use upstream's 'spk' task name and pin the preference order The SpeakerRecognition short name was 'spkrec', which audio.cpp neither prints nor parses; a name copied out of audio.cpp was rejected and a pinned 'spkrec' would not survive the engine boundary. Emit 'spk', keep 'spkrec' as an input-only alias, and correct the known-tasks lists. Three assertions were vacuous because their fixtures advertised a single task, so reversing a preference order or dropping the RPC name and the attempted pairs from the capability error all passed. Give them fixtures that can tell the orderings apart. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): convert sample, time and PCM units Integer nanosecond conversion so 44.1 kHz stays exact, float seconds for the VAD and diarization messages, and saturating s16le encode so an overshooting sample cannot wrap to the opposite sign. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): harden seconds_to_samples against NaN and overflow seconds_to_samples is the one entry point fed by untrusted-shaped input: a float-seconds timestamp off the wire, or a boundary from a model that diverged. Its guard covered only the low side, so NaN and out-of-range values fell through to an undefined double-to-int64 cast and came back as INT64_MIN. A hugely negative sample index used later as an offset or a length is a wild pointer rather than merely a wrong timestamp. Reject NaN with the !(x > 0) form and saturate before the cast. Also round instead of truncating there. These functions exist to cross the float seconds boundary the VAD and diarize messages use, and truncation lost a sample about half the time on the samples-to-seconds-and-back round trip, starting at n=1. Pin the decode scale at INT16_MIN, pin nanosecond truncation on a nonzero fraction, and record why the clamp argument order in f32_to_s16le is load-bearing for NaN. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): map NaN PCM samples to silence explicitly f32_to_s16le relied on std::min argument order to keep a NaN sample away from std::lround, whose result is unspecified for NaN. That was too subtle to rest on a comment, and the comment was itself wrong: it warned against a spelling that the outer std::max already catches, while three real spellings leak, including std::clamp, which is the idiomatic C++17 way to write the same clamp and so the likeliest future edit. Divert NaN before the clamp and encode it as 0. A NaN sample rendered as a full-scale click is worse audio than a dropped one, and this unit converts audio that may have originated off the wire. Pin it with an exact-value check rather than a range check, since all three outcomes the plausible spellings produce are finite and inside full scale, plus an invalid-operation check that fails unless the NaN is diverted before any ordered comparison. That second check is what catches modernizing the clamp and dropping the guard together. Also bound the seconds round-trip comment, which claimed unconditionally what holds only below roughly 2^23 samples, and document NaN, saturation and that bound in the header. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): assemble transcripts from runtime spans The top-level transcript text is TaskResult.text_output verbatim. audio.cpp carries text nowhere else: speech_segments, speaker_turns and word_timestamps hold spans and labels only, so deriving the text from them empties the transcript for any producer that omits word timing, VibeVoice diarized ASR included. Fixtures cover every observed producer shape. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): keep a nested speaker turn's own label A segment sourced from speaker_turns re-derived its speaker by greatest overlap. A turn's overlap with its own span is the largest possible, so a turn nested inside another speaker's turn could only tie with the container, and the tie went to whichever came first. sortformer_diar binarizes each speaker independently and sorts by start sample, so the container always comes first and the interjecting speaker was silently erased from DiarizeSegment.speaker. choose_segment_spans now carries the label out with the span. Also pins the nearest-segment fallback against measuring from either endpoint or from segment position, which a trailing-only stray word could not do, and exercises the empty-word guard in join_words. Two fixtures that pin a rule but do not mirror any pinned family are relabelled defensive. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): serialize runs with a wedge-aware guard audio.cpp sessions are not reentrant and a wedged CUDA call cannot be cancelled, so a plain mutex would pile every worker thread behind a stuck GPU. Callers waiting past the configured bound, or arriving while the holder has already overrun it, fail fast instead. A caller that queues behind a healthy run deliberately does not stamp the clock: only the thread that takes the lock does. Stamping on arrival would restart the wedge clock on every request and hide a stuck run from everyone behind it, which is the pile-up this guard exists to prevent. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): serialize inference through an InferenceLane One audio.cpp model is loaded per backend process and its sessions are not reentrant, so concurrent gRPC handlers have to take turns. Serialization alone is not enough: a wedged GPU call cannot be cancelled from userspace, so an unbounded queue behind one stuck run would swallow every gRPC worker thread until the process is useless. InferenceLane gives handlers a lane with room for one runner. LaneEntry occupies it for a scope and gives it back on every exit, including an exception, and is the only way to take the lane at all: occupy/vacate are private with LaneEntry as the sole friend, so a caller cannot acquire without holding something that releases. LaneEntry is immovable on purpose, because a moved-from entry would have to stop releasing while the lane still recorded it as occupied. A caller either waits indefinitely or brings a millisecond budget. A bounded caller that cannot get in fails instead of waiting on, and a bounded caller whose budget is already shorter than the age of the run in the lane fails immediately, which is what stops a queue forming behind a wedged run. The two failures carry different text: one names the wait it exhausted, the other states the measured age of the run without claiming to know why it is long, since a short budget meeting a legitimately long run lands there too. The run's age is stamped only after acquisition. A waiter that published itself as holder would restart the measurement and hide a genuinely stuck holder from every caller behind it. Budget negotiation and the overrun decision are pure functions taking their inputs explicitly, so both are covered without threads or sleeping. The per-model ceiling arrives as an int of milliseconds; a request may tighten it and may never loosen it. Replaces the previous run_guard unit, which was a derivative of an Apache-2.0 file upstream and could not stay in an MIT tree. Written from a behaviour contract with no reference to the removed code. Tests: 65 checks, standard library only, single translation unit, clean under -Wall -Wextra. Mutation tested at 23/23 killed; two of those mutants exposed missing coverage and the tests were extended until they died. ThreadSanitizer clean. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): make the B10 test able to fail, and document LaneEntry Review of the previous commit found the B10 test could not fail for the reason it was named. It aged the in-flight run to about 120 ms and then tried two budgets, 30 ms and 60 ms, both under that age, so both callers took the fail-fast path. "The two failure modes do not share one message" was comparing two fail-fast messages that differ only in the budget they print, and the timeout path was never reached. The second budget is now 400 ms, well over the run's age, so that caller queues and times out, and a new check asserts which path each caller took instead of inferring it from inequality. A mutant that makes the fail-fast path emit the timeout message previously died only on B4 and B8 checks; it now also dies on B10. Comment-only changes elsewhere. LaneEntry now says it is not reentrant and does not detect reentrancy: a second entry on a thread that already holds the lane surfaces as LaneUnavailable with a positive budget, but parks silently in unbounded mode, which matters because a handler may hold one across a whole stream. The immovability note now names the shapes that work, an optional emplaced in place or a unique_ptr, rather than saying to hold the entry indirectly without saying how; all three documented forms were compiled before being written down, which is how the note came to say that an optional of an immovable type cannot itself be returned. The header's explanation of why fail-fast exists is reworded. Two clauses traced back to a specification written after reading the Apache-2.0 upstream header, and while that was judged de minimis, this unit was rewritten precisely to carry no upstream expression at all. The margin table in the report was also wrong about which wall-clock margins are load-sensitive: there are four, not one, and the tightest is the B3 arrival check, which is now flagged at the call site. No margin value changed and none moved across 65 runs. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): gate model loading on the audio.cpp family Refuses any GGUF without an audiocpp.model_spec.family key and any non-GGUF path without an explicit family option, so the model loader's greedy backend probe cannot bind an unrelated llama.cpp GGUF to this backend (#9287). Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): load models and cache sessions per task Loads one ILoadedVoiceModel and creates an IVoiceTaskSession lazily per (task, mode), so the same model serves both the unary and streaming RPCs. LoadModel derives the family from GGUF metadata or an explicit option and fails with INVALID_ARGUMENT otherwise, so a failed load is a gRPC error the backend probe can see. audiocpp_backend::Task mirrors engine::runtime::VoiceTaskKind positionally, and drift there is silent: every unit still compiles and every test still passes while the backend runs a different task. Two mechanisms pin it. The static_asserts in loaded_model.cpp catch an insertion or a reorder, and -Werror=switch on that one file turns an appended upstream enumerator into a build failure rather than a warning in a 600 file log. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): stop aborting the process on SIGTERM The signal handler called grpc::Server::Shutdown directly. Shutdown takes an absl::Mutex, which is not async-signal-safe: the handler can interrupt a thread already holding that mutex, and abseil's deadlock detector responds by aborting. Every SIGTERM therefore ended in exit 134 and a 'dying due to potential deadlock' stack rather than a drained shutdown. The handler now sets a lock-free atomic and returns. Server::Wait moves to a helper thread so the main thread can poll that flag and call Shutdown itself, outside any signal context. A condition variable would not have helped, because notifying one from a handler is not async-signal-safe either. SIGTERM and SIGINT both exit 0 with no stack trace, where both previously exited 134. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): correct the status, lifetime and state contracts of LoadedModel An environment fault during session creation was reported as UNIMPLEMENTED. A missing libggml-cpu-*.so surfaced to the client as 'family silero_vad advertises vad/offline but refused to create the session: Failed to initialize CPU backend', which tells LocalAI the model cannot do this and must never be retried, and sends an operator hunting a capability bug instead of a packaging one. A throw from create_task_session is now a plain runtime_error, so it maps to INTERNAL. Only a null return, where the family genuinely declined, stays a CapabilityError. The model.'s task: option was parsed and then dropped: it lived in a local that died at the end of LoadModel and had no route to RequestShape::pinned_task. LoadedModel now keeps it and exposes pinned_task(). The global model becomes a shared_ptr reached through snapshot(). An audio RPC runs for seconds and cannot hold g_model_mu for its duration, so under a unique_ptr a Free arriving mid-request would destroy the model underneath it. Handlers now take a counted reference and whichever finishes last does the teardown, outside the lock. session_for documents the streaming state contract rather than resetting the session itself. Resetting on a cache hit was tried first and is not possible: silero_vad throws 'session prepare() must be called before Silero VAD reset()', so it would turn an ordinary second fetch into a hard error. start_stream's base implementation is already a reset, so a caller that runs prepare then start_stream per stream gets a clean session; a probe against the bundled silero_vad confirms an identical replay when it does and a carried-over stream when it does not. Also: an unknown backend: name is rejected before the model loads rather than after; MainGPU is parsed instead of passed through std::atoi, which turned 'gpu1' into device 0 silently; and device carries a device_set flag, because 0 is both the default and a real device index, so MainGPU was overriding an explicit device:0 that the neighbouring threads: handling promises will win. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): serve the VAD and Diarize RPCs Both emit float seconds, converted from the runtime's sample-index spans, and both take a counted reference to the loaded model through snapshot() and hold it for the whole call: a Free arriving mid-request drops only the global's reference, so whichever request finishes last destroys the model instead of one of them running on freed weights. An AddressSanitizer build reproduces exactly that heap-use-after-free inside ggml_vec_dot_f32 when the handler keeps a raw pointer instead, which is why the shape is what it is. The inference lane is taken before session_for, not after. session_for reads and writes an unsynchronised session cache and the offline run calls prepare(), which mutates the session, so both belong inside the lane. Diarize routes before it reads the input file, so a family that cannot diarize at all says so rather than complaining about the audio first. Its per-segment text stays empty because audio.cpp's SpeakerTurn carries a span and a speaker label only, and nested or overlapping turns are passed through untouched: a sortformer turn inside another speaker's turn is correct output for overlapped speech, and LocalAI is overlap-tolerant downstream. Duration counts frames rather than floats, so a stereo input does not report twice its length. Verified end to end against upstream's bundled silero_vad, which needs no download, using the bundled 16 kHz speech asset: a synthetic tone returns nothing, correctly, because silero detects speech and a sine is not speech. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): enforce ModelIdentity on VAD and Diarize audio-cpp was the only C++ backend without the model-identity guard, and no later task in the plan added it. pkg/grpc/server.go enforces checkModelIdentity on exactly these two RPCs, for the reason #10952 records: in distributed mode a worker can recycle a stopped backend's gRPC port for another model's backend, and the controller's liveness-only probe cannot tell a stale cached route from a live one. Without this guard a stale route gets a different model's VAD or diarization answer back with a 200. The loaded identity lives on LoadedModel rather than in a separate global, which is where this differs from llama-cpp. A handler holding the model through snapshot() then necessarily judges against the identity that model was loaded with, and a concurrent reload cannot swap one without the other. The refusal is NOT_FOUND carrying the verbatim grpcerrors.ModelMismatchSentinel substring. session_for and run_offline now take a const LaneEntry & proof-of-holding parameter. The rule that both must run under the inference lane was prose, which is exactly how the plan came to specify the inverted order; it is now a compile error. Restoring the inverted order fails to build rather than racing on an unsynchronised session map with a mutating prepare(). Diarize's speaker-hint comment claimed the dropped hints were "not a silent failure". From the caller's side that is what they are, and backend.proto documents num_speakers as forcing, so the comment now says plainly that the forwarding is dead for sortformer and that the family which lands must either honour num_speakers or refuse it. read_audio_file inspects the error_code from exists(), so an unsearchable parent directory no longer reports as a missing file. The VAD handler records the stimulus that actually works, since silero correctly ignores synthetic tones and the next task would otherwise rediscover that. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): make the lane and identity guards structural Two hardenings ahead of the eleven handlers still to be written, both of which get harder to retrofit later. The lane proof-of-holding parameter was a const reference, which binds to a temporary, so session_for(rpc, shape, model->acquire(0)) compiled. Each such temporary dies at the end of its own full-expression, releasing the lane between two calls that must share one: precisely the split the parameter exists to prevent, and the form a future author is most likely to reach for because it reads as tidy. A non-const reference requires an lvalue, so the temporary form now fails to compile while the named-local handlers build unchanged. The header comment no longer implies the check is total either: it proves a lane was taken, not that it is this model's lane. The identity check was two lines each handler had to remember, with nothing failing if a new one forgot them and no C++ equivalent of model_identity_modalities_test.go to notice. snapshot() becomes snapshot_unchecked(), whose only legitimate caller is Status, since HealthMessage carries no ModelIdentity. Handlers go through snapshot_for(), which takes the counted reference, refuses when nothing is loaded, and runs the identity check before anything can route. Every handler already has to call something to obtain the model, so the guarded call is now the shortest path and skipping it means deliberately typing snapshot_unchecked. A convention that has to be remembered can rot; this cannot. Verified: the temporary-argument and inverted-order forms each fail to compile with the expected diagnostic, the real handlers build, and bypassing the guard in Diarize alone turns the identity test red on that RPC while VAD stays green. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): serve the AudioTranscription RPC Adds result_map, the engine-to-proto boundary, and wires the offline transcription RPC. The handler branches on the ROUTED task: for Asr the request's prompt is whisper-style decoding context and becomes a request option, for Alignment the same field IS the transcript to align and becomes the text input. Routing has already decided which. The result text is TaskResult.text_output verbatim and is never derived from the segments. audio.cpp carries transcript text in text_output and nowhere else, so deriving it returns an empty transcript for every producer that reports segments without word timing. transcript_assembly already enforces that; this commit's job is not to undo it at the proto boundary, and result_map_ctest pins it there. read_audio_file now takes the sample rate the caller needs. Both file-fed speech handlers ask for 16 kHz mono, for two reasons: silero_vad and sortformer_diar refuse anything else outright, which turned an ordinary 44.1 kHz upload into INTERNAL, and nemotron_asr emits word timestamps in its own 16 kHz feature domain whatever the input was, so only a 16 kHz buffer makes the emitted nanoseconds right. Zero keeps the file's native rate and channels, which is what source separation will need. LoadedModel::check_can_serve answers a capability refusal before the lane is taken and before the input file is read. Routing is a pure read of the immutable capabilities, so a model that cannot serve an RPC no longer waits out somebody else's run to say so. VAD and Diarize use it too. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): stop linking sentencepiece's vendored protobuf engine_runtime links sentencepiece, whose default SPM_PROTOBUF_PROVIDER builds the protobuf-lite 3.14.0 sources it vendors. The generated backend.pb.cc is built against the toolchain's protobuf 3.21.12. Both ended up in the binary: 476 google::protobuf:: symbols came from the archive, 278 of them also defined by libprotobuf.so, and the archive won, because once ld pulls a member in for sentencepiece's own code every reference binds to the definitions that member carries. The visible symptom is one function. ParseContext::ParseMessage(MessageLite*, const char*) is what a generated _InternalParse calls for a submessage field and for nothing else, so flat messages parsed and nested ones did not: a TranscriptResult carrying segments serialized to correct bytes that the same process could not read back, and TranscriptLiveRequest, a oneof of submessages, could not have been parsed at all. Underneath that, 3.21 generated code was running 3.14 arena, ArenaStringPtr and ExtensionSet code. -Wl,--exclude-libs does not fix it. It makes those symbols LOCAL in .dynsym and the parse still fails, because the binding was decided at static link time and no visibility flag revisits it. Setting SPM_PROTOBUF_PROVIDER to "package" before add_subdirectory points sentencepiece at the protobuf the generated code was already built against. Zero google::protobuf:: definitions remain in the executable afterwards, every nested message round trips, and citrinet_asr, which parses a SentencePiece ModelProto at load time and would break first if this were wrong, still tokenizes and transcribes correctly. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): fix the segment text a transcription response is built from Segment text is not decoration. core/http/endpoints/openai/transcription.go routes response_format text, srt, vtt and lrc through schema.TranscriptionResponse, which builds the entire body out of Segments[].Text and never reads the top-level text. So for those four formats the segment text IS the response. nemotron_asr emits one word_timestamp per SentencePiece token, and the word boundary is carried as a LEADING SPACE on the piece ("So", "me", " call"). join_words inserted a space unconditionally, so response_format=text returned "So me call me na ture ," while the correct sentence sat unread in the top-level field. The separator is now chosen from the words themselves: whole words are space-joined, subword pieces are concatenated, and one leading space anywhere selects the latter. Concatenating the real nemotron pieces reproduces text_output exactly, verified end to end. This does not touch the top-level text, which is still text_output verbatim. The rule that forbids deriving the transcript from the segments is about the direction segments -> text; segment text has no source other than its words. Two smaller corrections in the same area: timestamp_granularities ["word"] set only "word_timestamps", a key no family in the pinned upstream reads. It now sets "return_timestamps", which qwen3_asr does read and which both runs its forced aligner and shortens its chunk window, so asking for word granularity no longer silently returns nothing. The request-option comment claimed more than it delivered. prompt, translate and temperature are read by no ASR family, and are forwarded only so a family adopting them works unchanged; the comment now says so per key, and gives TranscriptRequest.diarize the same explicit treatment threads already had. Also: the shipping target now carries -Wall -Wextra -Wpedantic, which it never did, so "the build is clean" starts meaning something; and fill_transcript_result no longer swallows a null response pointer, since answering OK with an empty transcript is the one failure mode this unit exists to prevent. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): serve the AudioTransform RPC Covers voice conversion, singing voice conversion, speech to speech and source separation, the four tasks LocalAI's AudioTransform can represent. AudioTransformResult carries one dst while htdemucs and mel_band_roformer produce several named stems from a single run, so inference runs ONCE, every stem is written as a sibling file <dst-stem>.<name>.<ext>, and params[stem] selects which one dst receives, defaulting to vocals and falling back to the first output. An unknown stem name is INVALID_ARGUMENT listing the real stem names rather than a silent substitution, and the selection happens before the first write so a refused request leaves no files behind. params[stem] is consumed here and is not forwarded into the engine's request options. The stem decision lives in stem_selection, which is stdlib only and therefore tested by backend/cpp/run-unit-tests.sh. It also validates the names, because they come from the model (htdemucs reads them from the GGUF's config.sources) and each becomes a component of a path this backend writes: a name carrying a path separator would escape the caller's output directory, and two stems sharing a name would silently overwrite one another. Both files are read at their native rate and channel count. Separation forces it, since demucs and roformer refuse any rate but 44.1 kHz and lose the stereo image that separates a centred vocal from a wide mix. The conversion families all resample internally (seed_vc, vevo2, miocodec, chatterbox were each checked), so passing the file through unchanged is also strictly better than band limiting it to 16 kHz first. Verified end to end against htdemucs f16 on a 44.1 kHz stereo mix: four stems plus dst, dst byte identical to the selected stem, params[stem] selecting a different one, an unknown stem refused with no files written, and mono input preserved as mono output. Also against miocodec for the single output path, where params[stem] is refused rather than ignored. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): refuse an impossible stem early, and stop blaming the caller for a failed write Four fixes from the first review of the AudioTransform RPC. check_can_serve now returns the resolved route, so params[stem] on a route that is not source separation is refused from the route instead of after a full inference: 11 ms rather than the 4.5 s a miocodec conversion costs, and far worse on seed_vc or vevo2. The post-run refusal stays as the backstop for a separation-routed family that returns no stems anyway. The typo'd-stem-name case still needs the run, since no framework header publishes the stem names before one. Stem names carrying control bytes are refused. GGUF strings are length prefixed and demucs reads its sources from JSON, so an embedded NUL survives to here: two names differing only after the NUL are distinct std::strings, so the duplicate check passes them, and then path::c_str() truncates both and they open the same file. That is exactly the silent overwrite the duplicate check exists to prevent, with the .wav lost as well. A failed write is now INTERNAL rather than INVALID_ARGUMENT. The destination is LocalAI's own generated-content directory, not anything the caller named, so a full disk or a permission fault there is a server fault and is worth retrying, which is the opposite of what INVALID_ARGUMENT tells a client. An empty output path stays INVALID_ARGUMENT. Two comment corrections and one clarification: the separators' required rate is their checkpoint's declared samplerate rather than a hardcoded 44100, seed_vc resamples with soxr and falls back to sinc-hann, and the "no files left behind" guarantee covers a refused request, not a write that fails partway through the loop. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(audio-transform): stop folding every upload to 16 kHz mono, and name the separation stems Two defects that made source separation unusable through LocalAI's own API, even though the backend served it correctly over gRPC. /audio/transform normalized every upload to 16 kHz mono s16 through utils.AudioToWav, with no way past it. htdemucs and mel_band_roformer refuse any rate but their checkpoint's own and separate a centred vocal from a wide mix using the stereo image, so every separation request through the HTTP API died with "HTDemucs prepare() sample rate mismatch: expected 44100, got 16000" while the same call over gRPC worked. The fold is not wrong, it is backend-specific: LocalVQE's echo cancellation genuinely wants 16 kHz mono and needs the reference in the same shape. So it becomes a declaration, BackendCapability.AudioTransformInputMono16k, set for localvqe and for nothing else. A backend that declares nothing gets its upload unchanged, which means no backend has to opt in to work. utils.AudioToWavPreservingShape is the non-folding conversion: a 16-bit PCM WAV passes through byte for byte at any rate and channel count, anything else is transcoded to WAV with its rate and channel layout kept. The other defect is that the run-once stem design bought nothing. A separation backend writes every stem beside dst from one inference, but AudioTransformResult carried only dst, so the other three were files no caller could find and a caller wanting all four had to run four separations. AudioTransformResult grows a repeated AudioTransformStem, the backend fills it, core/backend validates that each path really is inside the generated-content directory it handed over, and the endpoint publishes them as an X-Audio-Stems JSON header beside the existing X-Audio-Input-Url. JSON because a stem name is the model's own string and could contain any separator a hand-rolled format would use. Verified end to end through the HTTP endpoint with htdemucs f16 on a 44.1 kHz stereo file: 200 with a 44.1 kHz stereo body, all four stems named and fetchable through /generated-audio/, body byte identical to the selected stem, and params[stem]=drums returning a different one. The same upload sent to a model whose backend is localvqe still reaches the backend as 16 kHz mono, confirmed both by the engine's own rate refusal and by the persisted input file. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(audio-transform): reject extensible WAV from the passthrough, escape stem URLs, convert stems with dst Four fixes from the second review, plus one bug they made visible. isPCM16Wav tested only the bit depth, and go-audio's IsValidFile never looks at the format tag, so a 16-bit WAVE_FORMAT_EXTENSIBLE (0xFFFE) upload was passed through untouched where the old fold would have transcoded it. audio.cpp's WAV reader accepts 16-bit only when the tag is 1, so such a file died with "unsupported WAV encoding". Extensible is what many DAWs and Windows tools write and music files are this endpoint's new headline input, so it is a first-contact failure rather than a corner. The check now requires tag 1, with a spec that fails against the old implementation. Stem URLs are percent-escaped. A stem name is the model's own string and legally contains a space, a '#', a '?' or a '%'; an unescaped '#' truncates the URL before the request is even sent. The name field keeps the raw name. sample_rate and response_format are applied to the stems as well as to dst. Applying beat documenting: dst IS one of those stems, so leaving them alone broke the "dst duplicates the selected stem" invariant the whole design rests on, and both conversions are no-ops when unset. A stem whose conversion fails is dropped from the header rather than advertised in the wrong shape. Verifying that turned up why it had never been noticed: the two fields were never bound at all. The request arrives as multipart/form-data and echo's binder falls back to the FIELD NAME without a form tag, matching only case-insensitively, so "SampleRate" never matched "sample_rate" and "Format" never matched "response_format". Both were documented in the endpoint table and silently ignored. Two form tags fix it, and with them the conversion is observable end to end. Docs: audio-transform.md now documents what LocalAI does to an upload before the backend sees it, which backend gets the 16 kHz mono fold and why, params[stem], and the X-Audio-Stems header with a worked example. Also records the known limitation that the fold lookup is on the bare backend name, so pinned variants (vulkan-localvqe) do not match, and points at IsLlamaCppBackend as the suffix-tolerant precedent. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): serve the TTS and SoundGeneration RPCs TTSRequest.voice is treated as a speaker reference clip when it names an existing regular file, which makes routing prefer VoiceCloning, and as a named preset otherwise, in which case it travels as VoiceReference::cached_voice_id. Both the clip and SoundGenerationRequest.src are read at the file's own rate and channel count: upstream's own CLI and server do exactly that, every consuming family resamples internally and mostly with a better resampler than ours, and ace_step and stable_audio resample their input per channel, so a downmix here would delete the stereo image they are built to consume. The request builders live in their own unit rather than in grpc-server.cpp's anonymous namespace so they can be tested; grpc-server.cpp has a main() and cannot be linked into a test binary. The option keys are the whole point of these functions, so each one was grepped against the pinned upstream and the accounting is written down beside it. instructions maps to "instruct", which is what upstream's own server maps the OpenAI field to and what qwen3_tts and omnivoice read, and to "caption" for irodori_tts; the style tag is spelled "instruct" too, because "instructions" is looked up nowhere. duration maps to "duration_seconds", read by all three generation families, with the proto's own name kept only as a forward-tolerant alias. Keys that no family reads say so. Both handlers answer a capability refusal before taking the lane and before any file read, so a model that cannot synthesise does not queue behind somebody else's run to be told no. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): stop emitting an empty style language, and name the missing clip StyleCondition::language was set whenever has_language() was true, with no !empty() guard, while the language option twelve lines below had one. core/backend/tts.go sets Language unconditionally, so has_language() is true on every request LocalAI sends and carries "" when the caller named none. An engaged-but-empty style language is worse than an absent one: supertonic reads text_input->language behind its own !empty() guard and then overrides it from style->language with no guard at all, so "" replaced its "en" default and its tokenizer threw "invalid Supertonic language: ". Every /v1/audio/speech request that set instructions and no language would have been an INTERNAL against a supertonic model. A plain request never saw it, because the style condition only exists when instructions are non-empty, which is why the chatterbox end to end run did not catch it. TTS also stops discarding the Route that check_can_serve already returns. A family routed to voice cloning without a reference clip used to be refused from inside its own prepare(), which meant an INTERNAL naming neither the RPC nor the field to set; chatterbox advertises clon and no tts, so that was every preset-only request to it. It is now an INVALID_ARGUMENT naming TTSRequest.voice, answered in about 4 ms, and it cannot misfire because has_voice_reference is what selected cloning in the first place. Reading CapabilitySet::supports_speaker_reference to generalise this stays a follow-up. The src read carries a written caveat rather than a family blocklist, because ace_step's editing routes legitimately need src: setting src on a stable_audio model corrupts the heap and aborts the process in the pinned upstream, and the only thing keeping that off the network is that schema.ElevenLabsSoundGenerationRequest has no field for it. Nobody reading that Go schema would know why, so the reason is recorded where the field is read. build_tts_shape is extracted so TTSStream cannot describe the same request differently, and it arrived untested: two mutations of it survived until a test_tts_shape case was added. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): serve the TTSStream and AudioTranscriptionStream RPCs TTSStream leads with a streaming WAV header carrying 0xFFFFFFFF sizes, matching the convention backend/go/vibevoice-cpp established, so an HTTP client can start playback before the full PCM exists. Its chunks are read from StreamEvent::named_audio_outputs and not audio_output: supertonic, omnivoice and voxcpm2 all put their streamed audio there and leave audio_output empty until the very end, so reading the obvious field yields a stream with no audio in it. The finish_stream result is the family's own merged whole rather than a tail, so it is emitted only when nothing was streamed. Streaming transcription sends incremental deltas and degrades to a single delta plus the final result on families that offer no streaming ASR, which is the same message sequence with fewer deltas. The four streaming ASR families disagree on what partial_text means: nemotron_asr, vibevoice_asr and higgs_audio_stt report incremental fragments while voxtral_realtime reports the whole hypothesis and reports it twice, so the reconciliation lives in one tested unit rather than in the handler. nemotron_asr reports only through the stream event sink, and only from inside finalize, so the audio driver installs one and clears it again before returning: the session is cached and a sink left holding the caller's frame is a use after free waiting for the next stream. begin_stream is now the only implementation of the streaming state obligation, prepare then start_stream. Streaming sessions are cached, and what clears the previous stream is start_stream's reset; a family override that dropped it would break every call site with no compile error, so there is one call site. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): keep streaming deltas on UTF-8 boundaries, refuse dtypes that abort TranscriptStreamResponse.delta is a proto3 string, whose wire format requires valid UTF-8. voxtral_realtime reports its hypothesis as a concatenation of raw token BYTES (tokenizer_text.cpp:171-183), so the cumulative difference between two consecutive reports is eventually a lone continuation byte, and the C++ runtime serializes that with only a logged warning while the Go runtime refuses to unmarshal it: the client loses the remaining deltas AND the final_result. Measured on a trace of a non-ASCII sentence, 11 of 31 messages failed to unmarshal and every accented character was lost. TranscriptDeltaTracker now holds back an incomplete trailing sequence and merges it into the next fragment; reconcile flushes it, which it always can because the final text is complete. The same trace now unmarshals in full with zero failures. A streaming buffer whose float count is not a whole number of frames is refused rather than truncated. The integer division dropped the tail floats from the fed audio and therefore from the transcript, with no diagnostic; vibevoice_asr refuses the same thing from the other side of the call. A supertonic GGUF whose weights are not f32 is refused at load. It reaches ggml_concat with mismatched operand types and ggml_abort takes the whole backend process down on the first request, so nothing downstream can report it: the model loads, then every request kills the process. Attributed rather than assumed, the unary TTS path aborts identically, and upstream records that package as untested. The refusal names the orig package and says what to run before deleting the guard. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): stop a repeated lead byte from orphaning the next delta The first UTF-8 fix closed the cumulative half only. Rule 2 discards a fragment the known text already starts with, and when that fragment is the LEAD BYTE of a new character it looks exactly like a repeat of an older character beginning with the same byte. It was discarded rather than held, its continuation bytes then arrived alone and began the next delta, and utf8_complete_prefix_length only ever inspected the trailing sequence, so a delta invalid at the FRONT went out whole. Through a real Go proto.Unmarshal the review's four-character repro gave 3 deltas, 2 unmarshal failures and a lost transcript. Reachable from the incremental families, not only from voxtral: nemotron_asr's decoder cuts at a byte offset and vibevoice_asr's common_prefix_size compares bytes, so both split characters. Measured over 30,000 randomized incremental traces, 53.28% of Japanese traces and 9.52% of French ones carried at least one delta the Go runtime refuses. Two changes. Rule 2 no longer judges a fragment that ends mid-character, so the lead byte is held instead of swallowed and the character survives intact; the cost is a few duplicated bytes in a shrinking cumulative report, which no pinned family produces. release() additionally drops leading orphan continuation bytes, so no delta can begin mid-character whatever the rules above it decide. Losing a byte keeps the stream alive; emitting one ends the RPC and takes the final_result with it. Post-fix all 60,000 traces produce zero unmarshal failures, and the cumulative streams plus both pure-ASCII incremental streams are byte-identical to the previous commit, so nothing changed for the families already working. The weight-dtype allow list moves to family_gate, where it is stdlib-only and pinned by a test rather than only by a comment. Two comment citations corrected. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): read only an exact repeat as a repeat, not any prefix Rule 2 discarded any partial the known text merely started with. For a cumulative family that is a duplicate; for an incremental family it is an ordinary short fragment that happens to coincide with the start of the transcript, and it was dropped, silently corrupting the text. Pure ASCII, no multi-byte character anywhere: the fragments "pure ", "ascii ", "trans", "c", "ri", "p", "t" left the client holding "pure ascii transcrit". Over 5,000 randomized traces per transcript, 9.50% of pure-ASCII and 29.12% of French traces ended with the client holding something other than final_result.text, with a 200 and no diagnostic. Both incremental families emit fragments that small routinely, since nemotron_asr cuts at a byte offset and vibevoice_asr at a common prefix. Narrowing rule 2 to an exact repeat drives that to zero on all six transcripts and changes no cumulative stream at all: 30,000 randomized cumulative traces are byte-identical to the previous commit. What rule 2 guarded was established from upstream rather than from its own comment. The only duplicate any pinned family produces is voxtral_realtime's, where process_available_stream_chunks feeds each event to the sink from inside its loop and returns the last of the batch, so that event arrives twice with byte-equal text. A duplicate is an exact repeat, so equality still covers it. The case given up is a cumulative report that SHRINKS, which no pinned family can produce: voxtral decodes a token vector that is only push_back'ed and cleared by reset(), so within a stream it can only grow. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): serve the AudioTranscriptionLive RPC The one bidirectional stream this backend serves. The client sends a TranscriptLiveConfig, then TranscriptLiveAudio frames; the server acknowledges with ready, emits deltas as the audio arrives, and sends final_result once the read side closes. There is no offline fallback: live transcription has to consume audio incrementally, so a family with no streaming ASR is refused rather than served a batch run, which is what this RPC's Streaming-only mode_candidates list already says. The driver is a new sibling of run_streaming_audio, run_streaming_live, because the audio does not exist yet: instead of slicing a buffer it pulls frames from the caller until the read side closes. It installs the same ScopedStreamSink in the same order, which is not optional, since nemotron_asr returns a bare event from process_audio_chunk and reports every partial through the sink from inside finalize(). It buffers the wire's frames up to the family's own preferred window rather than feeding whatever size the client's audio callback produced, and it does not call finish_stream at all when no audio arrived, because nemotron_asr throws "finalize requires streamed audio" and an empty transcript is the truthful answer to transcribing nothing. Three things the handler had to get right and one it cannot: - The audio contract. A live request carries no samples, but nemotron_asr's streaming prepare() throws without an audio contract, and build_preparation_request derives it from TaskRequest::audio_input, so that field is an EMPTY buffer holding only the rate and the channel count. - 16 kHz or a refusal. The families express their spans in their own 16 kHz feature domain whatever the input was, and live frames cannot be resampled on the way in the way a file can, so an 8 kHz session would return timestamps 2x off with a 200. core/backend hardcodes 16000 anyway. - A mid-stream Config is refused. backend.proto calls it a decoder reset, but deltas already on the wire cannot be retracted, so a reset would leave the final text contradicting the transcript the client assembled. Ignoring the message would hand a client that believes it reset the decoder a transcript that silently continues the audio it thought it discarded. - The stale-route identity check cannot run here: TranscriptLiveRequest carries no ModelIdentity in either arm of its oneof, so snapshot_for does not instantiate for it. snapshot_unchecked's comment now names that as a second legitimate class of caller and says the fix is a proto change. eou and eob stay false. They exist for cache-aware models that emit end-of-utterance and end-of-backchannel tokens; audio.cpp's StreamEvent has no equivalent signal, and a client uses eou to decide the speaker yielded the turn, so a guess inferred from silence cuts people off mid-sentence. The lane is held for the whole stream, which is as long as the user keeps talking: the streaming session is stateful and cached, so a concurrent run would interleave two callers' audio and corrupt both transcripts. Verified against nemotron_asr over a real connection with a 14 s WAV in 512-sample frames: ready first, 59 incremental deltas with no repeated prefix, concat(deltas) equal to final_result.text, word timestamps in nanoseconds, eou and eob false. citrinet_asr answers UNIMPLEMENTED naming the family and listing asr/offline. A config followed by a close returns an empty final_result rather than hanging, and a first message that is not a config is INVALID_ARGUMENT. Two concurrent streams both return the complete transcript. Two cleanups on lines Task 12 touched, folded in. The DtypeAllowList terminator is now asserted at compile time: the reported out-of-bounds read did not exist, the single entry does terminate, but the loops have no other bound and any edit that widened an entry would walk off the end. And the dtype guard now short-circuits on "is there a table entry" through a new predicate rather than on the emptiness of the description string, which would have skipped the check on an entry with an empty allow list, i.e. on precisely the entry that refuses every dtype. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): bound the lane a live stream can hold AudioTranscriptionLive holds the model's inference lane for the whole stream, which is correct (the streaming session is stateful and a concurrent run would interleave two callers' audio) and newly dangerous. Every other RPC holds the lane across compute, or across a write to a slow reader, and both of those terminate on their own. A live stream instead blocks in a client-driven read, and a peer that goes silent WITHOUT closing the stream never terminates anything: the lane stays taken and every other request against that model queues behind a client that stopped speaking. live_watchdog is a one-shot idle timer that ends the stream when no frame has arrived inside a window. It is standard library only, so it is unit tested without an engine. gRPC's synchronous Read has no timeout and cannot be given one, so the only way to unblock it is ServerContext::TryCancel, which decides the wire status itself: the client sees CANCELLED rather than the DEADLINE_EXCEEDED the handler returns, the reason is logged, and the lane coming back is the point. When it fires the read loop throws rather than reporting end-of-input, so the driver does not go on to finalize a decode nobody is waiting for. It is armed only after the lane is taken and disarmed as soon as the read side closes, and both ends matter. Arming earlier would cover acquire(), which legitimately blocks while another live stream runs, so a queued caller would be cancelled for waiting its turn. Disarming later would cover our own decode, where a window overrun is not a peer going quiet and cancelling would throw away the transcript the client is waiting for. The window is the new live_idle_timeout_ms option, 30 s by default, 0 meaning no limit. core/http/endpoints/openai/realtime.go drives a 300 ms ticker and feeds every tick that produced new audio while a turn is open, so 30 s of silence is a hundred ticks that delivered nothing. It is also longer than any pause a speaker takes mid-utterance, which is the case that must never be cut off, and backend.proto lets one stream span many utterances, so a client that pauses longer between them raises the option rather than discovering it. Two smaller corrections in the same handler: - check_can_serve now runs BEFORE the sample rate check. pkg/grpc/grpcerrors/errors.go degrades to the file path on UNIMPLEMENTED and on nothing else, so a live-incapable model asked at a wrong rate was answering INVALID_ARGUMENT and costing the caller its fallback. - a negative sample rate is refused instead of silently becoming 16000. Zero still means 16000, which is what the proto documents; -1 is malformed rather than absent and gets the same refusal every other bad rate gets. And one thing recorded rather than changed, at the handler: "live" here means incremental INPUT, not low latency, and with the pinned families it does not yet mean incremental OUTPUT either. nemotron_asr's process_audio_chunk only appends to its buffer, so its whole decode and every delta happen inside finalize(), after the client closes its send side. The policy-window buffering is inert for that family and matters only for vibevoice_asr and higgs_audio_stt. Verified on the wire with live_idle_timeout_ms:3000. A silent client acked at 371 ms and was cancelled at 3.371 s; a second live stream opened one second later received its ack 2.37 s in, i.e. at the instant the first was cancelled, and then transcribed successfully on the same cached session. Without the watchdog it would still be waiting. Re-ran the live transcription (ready first, 59 incremental deltas, concat equal to the final text, word timestamps in nanoseconds, eou and eob false), the citrinet refusal at both a right and a wrong rate (UNIMPLEMENTED either way now), and Task 12's AudioTranscriptionStream on nemotron_asr, which is unchanged. Mutation testing the watchdog found a weakness in its own test: the destructor test slept past the window inside the watched scope, so a destructor that DETACHED the thread instead of joining it passed unnoticed. The test now uses a window longer than the scope, which kills that mutant, and says why. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): refuse the unsupported RPCs with a reason AudioEncode, AudioDecode, AudioTransformStream, AudioToAudioStream and VoiceEmbed have no counterpart in audio.cpp's VoiceTaskKind. Each now returns UNIMPLEMENTED naming the loaded family, what that family does support, and the upstream limitation, instead of the generated base class's bare status. The reasons live in a table in capability_routing.cpp so they are data rather than literals copied into five handlers, and so a test can assert every one of them. The five claims this was planned against were re-read at the pinned upstream e800d435d130dc776baf6f3e6129bb62b1495c89, and one did not hold. "audio.cpp streams tts and asr only" is false: silero_vad advertises vad with RunMode::Streaming. The refusal stands on the narrower claim that survives, that no family advertises streaming for any task AudioTransform routes to, and a test asserts the refuted wording does not come back. VoiceEmbed is the one refusal whose request carries a ModelIdentity, so it runs the #10952 check before answering: a stale route must get NOT_FOUND and the router's sentinel, not "audio.cpp cannot embed speakers" about a model that is not loaded here. It cannot use snapshot_for, whose no-model branch would tell the caller to load a model when no model can help, so it takes the reference through snapshot_unchecked and checks identity itself. That function's comment now names three classes of caller instead of two. The two bidirectional surfaces refuse without reading their stream, verified with a client that writes a config and eight frames first and gets the status rather than hanging. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): correct the vevo2 clause, and assert the absences Review found a false clause in the AudioToAudioStream refusal. It said s2s is "offline voice conversion ... which converts one clip into another speaker's voice", which is true of miocodec and false of vevo2: vevo2's s2s route is `editing` and only `editing` (default_route_for_task and route_matches_task in src/models/vevo2/session.cpp), documented as "Edit source speech into new target text while using the target voice" and requiring --target-text, so it rewrites what was said. vevo2's voice conversion is its separate vc task. It now reads "offline clip-to-clip processing against a target voice, declared only by miocodec (voice conversion) and vevo2 (speech editing)", and a test asserts the miscast cannot come back. The conclusion is unchanged: neither family converses. That defect was undetectable on the wire, since vevo2 does not load here, which is the argument for upstream_absence_ctest.cpp. It links engine_runtime purely to interrogate make_default_registry() and asserts the five premises the refusal reasons rest on: no codec task kind, no family advertising spk, no streaming for sep/vc/svc/s2s, miocodec advertising exactly vc and s2s, and s2s advertised by exactly miocodec and vevo2. The last two are exact sets, so an addition fails here rather than leaving a message stale. A positive control proves the registry is populated and the query works before any absence is believed, and every assertion has a reproduced negative control. This turns an AUDIO_CPP_VERSION bump from "remember to re-read five prose paragraphs" into a test failure. unsupported_surface now switches over UnsupportedRpc with no default label, so -Wswitch reports a sixth enumerator added without a row at build time; the runtime bounds guard it replaces is deleted. The AudioTransformStream reason had a true premise and an overreaching conclusion: an offline sep family could be buffered into a stream, as other LocalAI backends do. It now says this backend declines to offer a buffered offline call in disguise, rather than implying impossibility. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): make the missing-switch-case diagnostic fatal unsupported_surface() switches UnsupportedRpc onto the table row that explains it, with no default label, so -Wswitch reports an enumerator nobody handled. As a warning that is not enough: adding a sixth enumerator and building the shipping target gives exit 0, a binary and one warning, and the trailing `return surfaces[0];` then answers the new RPC with AudioEncode's codec reason. That is a confident, specific and false statement about audio.cpp on the wire, on the one code path whose entire job is to be truthful about what this backend cannot do, and it is worse than the runtime fallback it replaced, which at least named itself as a bug in this file. capability_routing.cpp therefore joins loaded_model.cpp on the existing -Werror=switch pin, whose comment already made this argument for the engine enum. The comment now covers both files. The pin stays per-file rather than project-wide because upstream's own ace_step/vae_decoder.cpp has unhandled -Wswitch cases of its own. Verified: a sixth enumerator now fails `make grpc-server` with exit 2 and no binary; appending a 14th VoiceTaskKind upstream still fails loaded_model.cpp, so the two pins fire independently; both reverted clean. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): package the backend image Bundles the dependency closure for the from-scratch image, the dlopened ggml CPU-variant shared objects that ldd cannot see, and upstream's bundled silero_vad and marblenet_vad assets so VAD works with no download. The bundled loader sits in the package ROOT rather than at lib/ld.so. run.sh execs it, which makes /proc/self/exe name the loader, and this backend has two consumers of that path: ggml discovers the libggml-cpu-*.so by listing dirname(/proc/self/exe), and resolve_model_path expands bundled:<name> under the same directory. Rooting the loader makes the binary, the ggml objects and assets/ share the one directory all three resolution mechanisms agree on. llama-cpp's lib/ld.so layout would need assets/ moved into lib/ as well. The image builds against apt gRPC and protobuf, like Dockerfile.ds4 and unlike Dockerfile.privacy-filter. The from-source gRPC that install-base-deps.sh and the base-grpc-* images supply vendors protobuf 26, which pulls abseil into message_lite.h; with SPM_PROTOBUF_PROVIDER=package that collides with sentencepiece's vendored mini-abseil and every absl::internal reference becomes ambiguous. Noble's protobuf 3.21.12 predates the abseil dependency and is the pair every earlier verification of this backend ran against. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): exempt the driver libraries from the packaging gate package.sh already left libcuda.so* and libnvidia-* to the host when copying, because the driver has to match the kernel module on whatever host runs the image, but the validation gate had no matching exemption. With BUILD_TYPE=cublas ggml is static and links CUDA::cuda_driver, so grpc-server carries DT_NEEDED libcuda.so.1 and the gate would have rejected the very absence the copy loop created, failing every cublas build in CI. One regex now feeds both. Building a control for that found a second defect: ld.so --list refuses to trace an object with an unresolvable dependency at all, exiting 127 without emitting a per-library line, so the "=> not found" rule was dead code and no exemption could have applied to it. The gate now traces with LD_TRACE_LOADED_OBJECTS and LD_LIBRARY_PATH, which reports the missing name and exits 0, and which is also what run.sh does at run time. Adds a layout assertion so a future move of the loader into lib/ fails the build instead of shipping a package that resolves bundled: models into lib/assets and finds no ggml CPU backend, and records for Task 16 that the Darwin script must not be a straight copy of privacy-filter-darwin.sh, which never calls package.sh and would silently drop assets/. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): register the backend with CI and the gallery Adds the five Linux matrix entries (cpu amd64/arm64 sharing a tag-suffix so the manifest merge fires, cuda 12, cuda 13, vulkan), the path-filter case that keeps later PRs touching backend/cpp/audio-cpp/ from getting zero CI jobs, the bump-bot entry pointing at the AUDIO_CPP_VERSION pin in the backend Makefile, the gallery meta plus its -development variant and the image entries for every variant, and the Makefile docker-build wiring. The matrix entries carry base-image only, with no builder-base-image, unlike the llama-cpp and privacy-filter blocks they sit next to. The prebuilt quay.io/go-skynet/ci-cache:base-grpc-* images ship a from-source gRPC whose protobuf v26 depends on abseil, and this backend's sentencepiece is built with SPM_PROTOBUF_PROVIDER=package, so it sees real abseil's absl::lts_20240116:: internal alongside its own vendored plain absl::internal and every absl::internal:: reference becomes ambiguous. Building against base-grpc-amd64 fails at sentencepiece-static.dir/error.cc.o with "reference to 'internal' is ambiguous". Dockerfile.audio-cpp installs apt's gRPC/protobuf 3.21.12 itself, which is also the pair every unit and end-to-end run of this backend has been verified against, and the CUDA toolkit therefore has to come from base-image. No Darwin matrix entry and no metal gallery entries: the Metal build needs scripts/build/audio-cpp-darwin.sh, a backends/audio-cpp-darwin make target and a routing step in backend_build_darwin.yml, none of which exist yet, so an entry added now would be routed to build-darwin-go-backend and look for backend/go/audio-cpp/. The inferBackendPathDarwin case and the DARWIN_BESPOKE_BUILDERS membership are in place, inert, so that adding the entry later is a one-line change that cannot be claimed by the generic Go path. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): pin the CUDA architectures, drop the vulkan variant Upstream sets CUDA_ARCHITECTURES to `native` on the engine_runtime target whenever CMAKE_CUDA_ARCHITECTURES is unset at root scope, and docs/build/ linux.md says so outright. ggml's own default does not rescue it: it list(APPEND)s in the ggml subdirectory scope, which never reaches the root scope where the engine_runtime property is decided. No CI runner has a GPU for `native` to enumerate, so both cublas entries would have gone red on the very commit that first turns a CUDA build on. Pin the list in backend/cpp/audio-cpp/Makefile, selected by CUDA_MAJOR_VERSION, which Dockerfile.audio-cpp now forwards from the CI build-arg it was previously discarding. The values are copied from ggml's own version guards rather than invented, so engine_runtime and ggml compile for the same set: CUDA 12 keeps the Maxwell/Pascal/Volta virtual archs and stops at 120a-real, CUDA 13 drops them and adds 121a-real. The `a` suffix is used rather than `f` because the latter needs CMake 3.31.8 and Ubuntu Noble ships 3.28.3. Verified by driving CMake 3.28.3's own CUDA architecture validator over both lists, with 120f-virtual as the rejected control. Drop the vulkan matrix entry, its two gallery entries, the vulkan capability key on both metas and the Vulkan tag. Every other vulkan backend gets its Mesa ICD drivers from .docker/install-base-deps.sh, which package-gpu-libs.sh then bundles; Dockerfile.audio-cpp calls neither and installs only libvulkan-dev and glslc, so the image would ship a Vulkan loader that finds no GPU. No CI job runs a vulkan image against real hardware, so that would have passed green and failed in users' hands. BUILD_TYPE=vulkan stays supported for local builds. Also note on the cublas entries that cuda-major-version now selects the architecture list and that cuda-minor-version and the base-image tag encode the same toolkit, and correct the stale entry counts on matrixEntryKey. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): build for Darwin Metal Bespoke C++ Darwin path like ds4 and privacy-filter: an includeDarwin matrix entry, a backends/audio-cpp-darwin make target, a gated workflow step, and the metal image entries plus metal/metal-darwin-arm64 capability keys in the backend gallery. The build script deliberately does NOT reassemble the package the way privacy-filter-darwin.sh does. It runs the backend's own `make package` and copies the result, so the Darwin package keeps the root-level layout the Linux one has: grpc-server, run.sh, the ggml objects and assets/ in one directory, with lib/ for the dylib closure. Hand-assembling would drop assets/, and assets/ is what makes the bundled: model paths resolve with nothing downloaded. The dylib walk is a full transitive closure rather than the single level ds4 and llama-cpp do, because Homebrew's grpc++ pulls libgrpc, abseil, upb, cares and OpenSSL that grpc-server does not link itself, and a level-1 walk ships a package that only works on a machine that already has Homebrew grpc. Two fixes folded in, both in the backend Makefile: - an EMPTY CUDA_MAJOR_VERSION fell through to the CUDA 12 architecture list, which contains 120a-real and so needs nvcc >= 12.8. A local BUILD_TYPE=cublas build on a 12.0-12.7 host failed to compile where upstream's documented default (native) worked. EMPTY now maps to native, 12 and 13 keep their lists, and any other non-empty value is an error on cublas builds. CI always passes a major, so CI is unaffected. - the Darwin branch now points CMake at Homebrew's keg-only libomp. AppleClang ships no OpenMP runtime and nothing is symlinked into /opt/homebrew, so FindOpenMP finds neither the library nor the header, and audio.cpp calls find_package(OpenMP REQUIRED) whenever ENGINE_ENABLE_OPENMP is on. Without the hint the macOS build would have died at configure time. If the keg is absent the build disables OpenMP instead of failing. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): make the Darwin fallbacks loud and the rpath walk complete Review follow-up on the Darwin Metal build. The OpenMP fallback was silent. If brew --prefix libomp ever comes back empty, CI produced a green Metal package with 108 #pragma omp directives across ~30 files compiled out, and clang says nothing about an ignored omp pragma without -Wsource-uses-openmp, so the only trace was one absent flag inside a set -x cmake line. That regression would have been blamed on Metal. It now warns. The @rpath arm of the dylib walk had no live candidate when it was written, on the reasoning that a Metal build links ggml statically. The OpenMP fix in the same commit made libomp.dylib one, and whether Homebrew records it as an absolute opt path or as @rpath/libomp.dylib is not observable from Linux. The walk now expands @rpath, @loader_path and @executable_path against the object's own LC_RPATH entries, and only fails when nothing on disk answers, printing the rpath list with the error so a failure on a machine nobody can attach to explains itself. Also: ADDITIONAL_LIBS now go through the closure rather than a bare cp, so they are deduplicated and their own dependencies bundled; build/darwin/lib is created explicitly instead of relying on package.sh pre-creating it; the libomp probe uses nested ifneq rather than $(and ...), which needs GNU make 3.81 and would otherwise expand empty and take the OFF branch on an older make; and -DOpenMP_ROOT is quoted like its CUDA sibling. Verified with a Linux harness that runs the script verbatim against a stubbed otool: a level-2 transitive dep, an @rpath dep reachable only through LC_RPATH, and an ADDITIONAL_LIBS dep are all bundled, a dependency cycle terminates, system libraries are skipped, the packaged tree has assets/ at the root beside grpc-server with the dylibs in lib/, and both failure paths exit non-zero. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): make bundled: reachable from a model YAML resolve_model_path() tested the bundled: prefix on `candidate`, which prefers ModelFile and falls back to Model. LocalAI fills ModelFile by joining ModelPath onto the configured model string (pkg/model/loader.go, LoadModelWithFile), and only sets it from a managed artifact otherwise, so a model YAML saying `model: bundled:silero_vad` arrives as ModelFile "/models/bundled:silero_vad" and Model "bundled:silero_vad". The prefix therefore never matched through the normal load path: it matched only for a hand-written LoadModel call that left ModelFile empty, which is exactly how task 15 verified it, and every model YAML using the form failed with "model path does not exist: /models/bundled:silero_vad". Both fields are now checked, Model first, so the zero-download VAD path the package ships assets for is reachable the way it is documented. A caller that puts the form in ModelFile still works, so task 15's verification stands. Compiled clean; the runtime check could not run on this host, whose system libprotobuf/libre2 have gone missing (the pre-existing grpc-server binary no longer resolves its libraries either), so it wants a container run. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): advertise the backend and document its options Registers audio-cpp as preference-only in /backends/known: the family lives in GGUF metadata that an importer cannot read from a remote repo, and one repo hosts thirty families, so there is no honest auto-detect signal. Modality is a single string and the import form chips on a fixed key set, so it registers as tts with the other modalities named in the description rather than under an invented key the UI would bucket as "other". Adds a features page covering the option namespacing, the routing table per endpoint, the RPCs this backend declines and why, the bundled VAD path, the separation stem behaviour, and the family gotchas (supertonic needs the orig package; chatterbox advertises cloning and no plain tts; nemotron_asr defers its whole decode to finalize so live transcription emits nothing until the client half-closes, unlike higgs_audio_stt and voxtral_realtime). Every option name and family capability in it was read off the pinned upstream checkout. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * backend(audio-cpp): test resolve_model_path, and correct the family names The bundled: fix in 842443cd7 shipped without a test, which is how the bug got there: task 15 verified the form with a hand-written LoadModel that left ModelFile empty, and that is the one shape the server never produces. Four cases in streaming_driver_ctest, which already links loaded_model.cpp, pin the PRODUCTION shapes instead. The first fails against the pre-fix source (returns the joined /models/bundled:silero_vad); the other three are the branches the bundled: lookup now runs in front of and must fall through for. Three family names in the docs were the source directory rather than the registered family, on pages whose whole argument is that these names cannot be guessed: demucs is htdemucs (demucs/loader.cpp:22), roformer is mel_band_roformer (roformer/assets.h:15), and moss is TWO families, moss_tts_local and moss_tts_nano. The hyphenated ASR names are underscored to match, here and in the compatibility table. The supertonic dtype note claimed more than the evidence carries. The f16 abort is a local observation, identical through TTS and TTSStream; upstream's docs/gguf.md leaves the 16-bit column untested and records q8_0 as "No (unsupported weight dtype)", which says unusable rather than fatal. Both are still refused, because the allow list is what the family can run. Corrected in family_gate.h, family_gate.cpp and the docs together, since the docs inherited the wording from the code. The importers tripwire says in the file that it is a tripwire: it exercises no audio-cpp behaviour, and the registration assertion lives in backend_test.go. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * gallery: add audio.cpp models covering every served RPC One representative model per RPC group of the audio-cpp backend, plus the two bundled VAD models, which need no download at all because the assets ship inside the backend package. Every hash was computed with sha256sum on the downloaded file. Quantizations come from upstream's tested-status table in docs/gguf.md rather than a default of q8_0: supertonic ships the orig package (its q8_0 is recorded as an unsupported weight dtype and its f16 aborts in ggml_concat), and nemotron_asr and htdemucs ship f16 because their q8_0 builds are recorded with drift while 16-bit is a clean pass. Diarization and separation use the diarization and audio_transform usecases, not transcript: /v1/audio/diarization and /audio/transform filter the default model on FLAG_DIARIZATION and FLAG_AUDIO_TRANSFORM respectively, so a transcript flag would have hidden both models from their own endpoints. The forced aligner sets parameters.language, which the transcription endpoint uses as the fallback when no language form field is sent, because the family requires both a transcript and a language. All ten entries were run twice: once against the raw gRPC server, and once installed with local-ai models install and called through the HTTP endpoint. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * gallery: correct the audio.cpp entries' licenses Swept all ten entries against the real upstream named in audio.cpp's tools/model_manager.py rather than against the audio.cpp repo's own license. Three were wrong: supertonic apache-2.0 -> openrail weights come from mlx-community/supertonic-3-mlx, and both it and Supertone/supertonic are openrail citrinet apache-2.0 -> other pulled from NGC nvidia/nemo/stt_en_citrinet_256, governed by the NGC Terms of Use sortformer other -> cc-by-nc-4.0 nvidia/diar_sortformer_4spk-v1 is CC BY-NC 4.0, and the gallery already uses that exact string, so there is no reason to obscure a non-commercial bar The license field is one word, so citrinet and sortformer also gained a sentence saying why they are restricted. The other seven were confirmed correct against their sources. Also drops an unverified claim from the nemotron description. It said the model drives the realtime transcription session; that endpoint actually calls TranscribeStream, and the live RPC reaches LocalAI only through realtime_semantic_vad.go. Neither path was exercised here, so the description now states only the two calls that were. MarbleNet gains the NeMo upstream under urls: for parity with silero. No sha256, quantization, usecase or model choice changed. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(audio-transform): bound sample_rate, keep same-named uploads apart Four defects the whole-branch review found on the Go side, plus two comment corrections. sample_rate is a disk-exhaustion hazard. The branch added the `form:` tag that makes the field bind for the first time, so the resample path went from dead to live, and utils.AudioResample interpolates the int straight into ffmpeg's -ar with no bound. Measured with ffmpeg 7: -ar 999999999 on a 0.01 s clip writes 20 MB and exits 0, which scales linearly to the reported 3.9 GB for one second, into a GeneratedContentDir nothing sweeps, and convertStems repeats it once per separation stem. Clamped to 8000..192000 in the handler, before the temp dir and before the model is touched, and rejected with a 400 outside it. The low end was reported as "a 0-byte file". It is not: -ar 1 writes a 78-byte header with no audio behind it, whose declared data size still claims 70 bytes, so go-audio parses it as a 35 SECOND file and a size check does not see it. The guard therefore compares the declared data chunk against the bytes actually on disk, and AudioResample now fails rather than returning a WAV carrying nothing. Both parts of a transform request land in one temp dir, and the raw copy was named only after the client's basename, so `-F audio=@mic/clip.wav -F reference=@loopback/clip.wav` wrote "raw-clip.wav" twice. Since AudioToWavPreservingShape hardlinks an already-PCM16 WAV rather than copying it, the reference part's os.Create truncated the inode audio.wav pointed at: mic and reference came out identical, which makes an echo canceller null everything and return near-silence with a 200. The raw copy now carries the form field name. audio-cpp had no BackendCapabilities entry, so VoiceCloningForModel returned nil before it ever consulted the model's tts.voice_cloning override and every `voice: "profile:<id>"` request was refused with a 400, on a backend that ships audio-cpp-chatterbox whose family serves cloning and not plain TTS. Registered with its RPCs, usecases and the reference-audio contract, and deliberately without the 16 kHz mono fold, which its separation families cannot survive. GetBackendCapability was exact-match only, so every pinned gallery variant read as an unknown backend: vulkan-localvqe lost the 16 kHz mono fold that used to be unconditional and started failing inside LocalVQE, and the usecase gate does not stand in for it because BuildFilteredFirstAvailableDefaultModel returns early once the client names a model. Lookup now falls back to the meta name by stripping the gallery's hardware prefix and release-channel suffix, exact match first so nothing can be shadowed. Same class as #10945. Also corrected: the AudioTransformRequest comment claimed echo's binder falls back to the field name, which it does not in either direction (bindData binds ONLY tagged fields and `continue`s otherwise; `model` arrives from setModelNameFromRequest's c.FormValue). And the stable_audio `src` heap corruption caveat now lives on ElevenLabsSoundGenerationRequest, where the Go developer who would add the field can see it, instead of only in C++. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(audio-cpp): refuse a task pin the RPC cannot serve, and stop empty frames holding the lane The model's `task:` option is copied into the request shape by all nine handlers, which is correct, but resolve_route then replaced the RPC's candidate list with the pin WHOLESALE and never asked whether the pin was something that RPC routes to. One pin therefore bled across all nine surfaces, and because the family still supported the pinned task the result was a wrong 200 rather than an error. Reproduced live: nemotron with task:asr made Vad return 200 with zero segments after a full ASR decode, so 14 seconds of speech was reported as silence, and Diarize did the same; silero_vad with task:vad made AudioTranscription return 200 with empty text and four segments whose spans were VAD segments, which combined with response_format in {text,srt,vtt,lrc} building the body solely from Segments[].Text yields a well formed SRT of four timed EMPTY cues. It also contradicted the documented contract, that a family which cannot serve a request is refused rather than rerouted. A pin is now checked against the RPC's admissible task set before it is adopted, and the refusal names both the pin and the RPC. The set is derived from task_candidates with every shape flag set rather than restated, so a task added to an RPC's candidates cannot become inadmissible by omission. Every legitimate pin survives, and the test asserts all fifteen of them alongside the eight crossings that must not. The live watchdog was defeated by empty frames. idle.touch() ran on ANY message, before the has_audio and pcm.empty() filters, so a peer writing unset-oneof or zero-length frames faster than the window held the lane indefinitely while feeding the decoder nothing. There is one lane per model and one model per process, so that is a single client denying the whole backend, which is what the watchdog exists to prevent, and the thrown text already said "no audio frame arrived". The touch moved below the filters, which are now a named predicate so the distinction is testable rather than a call order nobody can see. Three comments corrected against measurement rather than reasoning: - CMakeLists claimed zero google::protobuf:: definitions remain in the executable. nm -C --defined-only reports 2515, and that is expected: they are generated code, sentencepiece::ModelProto's own _InternalParse among them. The claim that holds, and the one the ABI fix is actually about, is that no vendored protobuf RUNTIME is linked and ParseContext::ParseMessage is UNDEFINED in the executable, resolving to libprotobuf.so. - refuse_cloning_without_a_clip's "cannot misfire" paragraph had its reasoning backwards. Routing picks VoiceCloning as the FALLBACK when there is no clip, which is the case being caught; chatterbox, which ships in the gallery, advertises clon and no tts at all, so every voice-less request lands there. - audio_units read "2.1 min at 96 kHz" for index 11289602, which is 1.96 min. 2.1 min is 96 kHz's OWN first failure at 12288002. Both were remeasured and the note is now a per-rate table. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * build(audio-cpp): exclude the upstream checkout from the C++ gate, harden the darwin walk run-unit-tests.sh pruned */llama.cpp/* but not */audio.cpp/*. It is safe today only by luck: upstream's 44 tests all put "test" at the FRONT of the filename (17 test-*.cpp, 27 test_*.cpp, zero *_test.cpp), so the glob misses every one of them, and nothing enforces that. This gate runs on every PR for every backend and compiles each match as a standalone translation unit with nothing but nlohmann/json on the include path, so the day upstream adds or renames one test the gate goes red repo-wide on an Apache-2.0 file nobody here wrote. audio-cpp-darwin.sh now logs the raw otool -L output and the parsed LC_RPATH list unconditionally, before the walk. Both awk filters in that script assume a column layout nobody working on this can observe, since it runs only on the CI Mac, and a green first Darwin run proves nothing about the assumption: an awk that silently matched nothing yields an empty dependency list, which reads exactly like "no non-system dependencies" and packages happily. Both filters otherwise feed process substitutions, so their input never reached the log. It also lists every symlink in the package and fails on one that cannot resolve inside the image. A dangling link does not fail anything else here, because every assertion tests with -e, which follows links; it fails at dlopen on a user's Mac. Links are NOT banned outright, which the review suggested but which would break the libggml.dylib -> libggml.0.dylib chain the `cp -a` above exists to preserve. What is banned is a link that resolves on the build host and will not resolve in the image: a broken one, or an absolute one pointing outside the package. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * style(audio-cpp): drop em dashes from the audio-cpp capability entry Follow-up to a84b3c4b9, no behaviour change. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(config): key the voice-cloning model rule on the resolved backend Making GetBackendCapability strip the gallery hardware prefix and release channel fixed pinned variants of /audio/transform, but VoiceCloningForModel kept keying its per-backend switch on the caller's spelling. A pinned name therefore resolved the capability by stripping and then missed every case in the switch, falling through to the permissive default: cuda12-vibevoice-cpp advertised voice cloning for the realtime 0.5B model, metal-coqui for tacotron2, cuda12-crispasr for a pure ASR model, cpu-qwen3-tts-cpp for CustomVoice. Each of those is a model that cannot clone, so /v1/audio/speech accepted a profile: voice it had to fail on inside the backend rather than rejecting it with a 400, and the UI advertised the capability too. resolveBackendCapability now returns the key the entry was found under, and callers that branch on backend identity use that key instead of the name they were handed. The exact-match-first order is unchanged, so a backend genuinely registered under a variant-looking name still keys on its own name. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5 [Claude Code] * gallery(audio-cpp): declare audio_transform on the chatterbox entry Chatterbox advertises VoiceCloning AND VoiceConversion (src/models/chatterbox), and the entry's own description already said so, but known_usecases listed only tts. /audio/transform selects its default model by FLAG_AUDIO_TRANSFORM, so voice conversion was reachable only by naming the model explicitly and was invisible to every usecase-driven surface. It is the one audio.cpp task with a shipped gallery model and no way to find it. Verified against the real model rather than inferred from the capability list: AudioTransform with chatterbox-q8_0, speech as audio_path and a speaker clip as reference_path, returns a 5.08 s 24 kHz mono WAV at -25.5 dB mean and zero stems, which is the single-output shape voice conversion should have. The description now says which endpoint reaches that half and warns that installing this next to a source-separation model gives /audio/transform two candidates, so the model should be named rather than defaulted. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5 [Claude Code] * gallery(audio-cpp): add voice-design and singing-voice-conversion entries Two of the three audio.cpp task kinds that had no gallery model now have one. Both were driven end to end against the real weights through the backend before being written, not inferred from the capability tables. audio-cpp-irodori-voicedesign covers vdes. TTS carrying `instructions` routes to the vdes task, so the voice is described in words rather than supplied as a clip. Verified: "a calm elderly woman speaking slowly with a warm, gentle tone" over an 8.76 s 48 kHz mono render at -16.8 dB mean, and a closed-loop citrinet pass recovers the sentence with the accent drift expected from a Japanese-first model read by an English recogniser. audio-cpp-seedvc-singing covers svc, and pins task:svc because nothing else can reach it. seed_vc advertises svc and ordinary voice conversion, no request signal means "this input is singing", and auto-routing resolves the tie to voice conversion every time. Verified with the pin: 5.04 s 44.1 kHz output whose closed-loop citrinet transcription is exact. s2s deliberately has no entry, and the reason is not effort. miocodec is the only upstream family whose speech-to-speech route needs no text, and it returned audio with correct duration and level but no recoverable speech in four independent attempts: the stale build, v2 q8_0, v2 orig (the variant upstream records as a clean Pass), both tasks, and matched 44.1 kHz inputs on both sides. vevo2's route refuses with "Vevo2 text/prosody route requires text_input or target_text", and session.cpp:897 fills target_text only from request.text_input, which AudioTransform has no field to carry. The same vevo2 weights convert voice correctly through the default route with an exact ASR round trip, so the model and the plumbing are both healthy; it is the s2s route specifically that this RPC cannot express. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5 [Claude Code] * backend(audio-cpp): carry transform text through params, add the s2s entry AudioTransform is audio-in / audio-out and its proto message has no text field, but not every task it routes to is audio-only. vevo2's speech-to-speech route is a text and prosody route: session.cpp:897 fills refs.target_text from request.text_input and nowhere else, and the run refuses without one with "Vevo2 text/prosody route requires text_input or target_text". The params map is the only channel this RPC has that reaches the engine, so the text travels through it and apply_transform_text_input unpacks it after the params have been copied into task.options. Before this, s2s was not awkward to reach through /audio/transform, it was unreachable, and it was the last audio.cpp task kind with a real model and no way to get to it. target_text is canonical and text is its alias, the order vevo2's own option table declares them in, so a request setting both gets the canonical one rather than whichever the map happened to store first. An empty value falls through to the next candidate instead of ending the search. language rides along only when a text was found: on its own it conditions nothing, and manufacturing a text_input for it would route a plain separation request carrying a language hint through the text path. The keys are left in task.options rather than erased, because vevo2's loader advertises target_text as a request option and a family reading it there keeps working. Nine tests, all confirmed failing on behaviour against a stub that returned false before the implementation was written. Verified end to end afterwards: vevo2-q8_0 with task:s2s and params[text] returns a 5.12 s 24 kHz output whose closed-loop citrinet transcription is exact, and htdemucs separation with no text param still returns its four stems, with and without params[stem]. audio-cpp-vevo2-speech-to-speech ships that route. Every audio.cpp task kind with a loadable family now has a gallery entry; spk remains the only gap and has no family upstream at all. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5 [Claude Code] * docs(audio-cpp): document params[text] and the pinned transform tasks The text channel and the two task pins are both invisible from the endpoint contract alone: nothing in the AudioTransform form tells a reader that a speech-to-speech model needs the line it is resynthesising, and nothing says that asking for singing voice conversion without task:svc silently gets plain voice conversion instead. Both are the kind of thing a user only discovers from a refusal or, worse, from output that looks right and is not. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5 [Claude Code] * fix(utils): annotate the two G304 sites this branch introduced gosec flags os.Open on a variable path, and both new call sites in ffmpeg.go are its alerts on this PR. Neither is reachable by an outside caller: isPCM16Wav opens the exact path it is about to hand ffmpeg as input, which in the upload path is a server-created temp file named from path.Base of the client name so no traversal survives, and wavAudioBytes opens AudioResample's own dst, a name this package derives from src and has just had ffmpeg write. Annotated in the repo's existing style rather than restructured, with the reason spelled out, because a bare suppression is worth nothing to the next reader. The three other G304 sites in this file, in passthroughWAV and isTargetWav, predate the branch and are left untouched. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5 [Claude Code] * fix(audio-cpp): build arm64 with gcc-14 for the armv9.2 SME variants The arm64 CPU image failed to build: cc1: error: invalid feature modifier 'sme' in '-march=armv9.2-a+dotprod+fp16+sve+i8mm+sve2+sme' ggml's CPU_ALL_VARIANTS table includes armv9.2 variants compiled with +sme, and Ubuntu Noble's default gcc-13 rejects that feature modifier. Every entry in the table has to compile even though a host only ever dlopens the one its own CPU supports, so a single unbuildable variant fails the whole image. gcc-14 accepts it, which is exactly the fix llama-cpp already carries in .docker/llama-cpp-compile.sh; this is the same problem reached by a different Dockerfile. Applied to every arm64 BUILD_TYPE rather than to the CPU one alone, and that differs from llama-cpp on purpose. llama-cpp needs it only for its pure-CPU image because its GPU builds run llama-cpp-fallback, which builds no variant table. This backend's Makefile turns ENGINE_ENABLE_CPU_ALL_VARIANTS on for every non-Darwin build, GPU included, so an arm64 GPU image would hit the identical error. The matrix has no arm64 GPU entry today, which is precisely why gating on an empty BUILD_TYPE would leave the trap armed for whoever adds the first one. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5 [Claude Code] --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |