Commit Graph
900 Commits
Author SHA1 Message Date
4d7bdc6ff2 feat(ui): redesign the web UI around a shared kit and a calm palette (#12526)
* build(ui): vendor the shared UI kit snapshot at 0.2.0

The restyle needs the kit's tokens, motion layer and component classes.
Take a pinned snapshot instead of depending on the kit at build time,
and keep a lock file with the version and per-file checksums so a later
update shows exactly what changed. The product theme stays outside the
vendored directory.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add the ink and teal theme and bridge the old variables

Define the product colours as the shared UI kit's roles, for light and
dark, in theme-localai.css. The kit's contrast check passes on every
pair. theme.css keeps the existing --color-* and --shadow-* names but
now points each at a role, so App.css and the pages get the new palette
without edits. Radii move to the kit scale.

index.html now sets data-theme before first paint with the same rule as
ThemeContext (stored choice, otherwise dark), because the contract
layout of the theme file no longer defaults to dark by itself.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): restyle the shared chrome with the UI kit grammar

Adjust the shared classes so every page picks up the same interaction
language without per-page edits:

- Sidebar sits on the canvas and the current row lifts onto a card.
  Section labels are tracked uppercase, the badge is a soft pill, and
  the phone drawer leaves the tab order when closed.
- Buttons are flat: hover swaps the surface, press scales to .97, focus
  is a 2px ring with a 2px offset, danger is a tinted wash.
- Inputs use the card surface and the control edge; switches, tabs,
  filter chips, badges and cards follow the same rules. Cards no longer
  lift on hover; only linked or button cards react.
- Menus and popovers scale in from the trigger corner with 40px items.
  Dialogs get a veil fade and a spring settle. Toasts become pills at
  the bottom centre.
- The page transition is a 250 ms fade with a 6px rise. It fills
  backwards so a finished animation no longer leaves a transform that
  confined dialog veils to the main column.

The focus-ring test now checks the outline instead of a box shadow, and
new specs cover the theme roles, the first-paint theme and the sidebar
lift.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): move leftover hard-coded colours onto the theme roles

The YAML editor restated the old blue palette in JavaScript, and a few
pages kept literal blues, indigo and violet tints, or fallbacks that
only applied because a variable was never defined. Point them at the
theme variables so they follow light and dark and the new palette.

The status badges in the account pages built their tint by appending
"22" to a variable, which is not valid once the variable is defined, so
they had no background. Use the wash roles instead.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): raise the type scale and control size toward the kit

Body and list text moves to 15px and the rest of the scale follows the
kit's 12/13/15/17/21/32/44 steps. Page titles, section headings and
stat values are bold with tighter tracking; titles are 32px.

Buttons, inputs, selects, tabs and nav rows are 40px high with the 12px
radius, compact controls 32px. Tabs become a segmented control. The
sidebar widens to 240px (64px collapsed) and nav rows get more room.
Identifiers and counts in the split views use the mono face, and the
stat grid becomes separate inset tiles.

The Geist stack stays: it is bundled, and the thin look came from the
size, weight and negative tracking, not the face.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): separate cards, panes and floating surfaces from the canvas

Cards, the Models and Installed split panes, the composers and the
confirm dialog use a stronger card edge, the rest shadow and the 20px
radius, so they read as layers in dark as well as light. Menus and
popovers move to a float surface (the hover tone in dark) with the
float shadow.

The selected rail row gets an accent wash and a 3px accent edge. The
send buttons are a clear accent when there is something to send and a
quiet inset when not; the Home button carries data-empty for that, since
submitting an empty box does nothing. The assistant card becomes an
accent wash with a square icon.

New surfaces spec checks the pane edge, the selected row, both send
buttons and the popover in both themes. The voice library empty-state
spec now waits for the layout to settle before comparing two boxes.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): tidy the sidebar header and mark the current row with a dot

The header gives the configured horizontal logo a fixed width and
centres it in a 72px band, lined up with the nav icons. The collapsed
rail shows the configured icon logo centred, and its nav rows become
40px tiles centred in the 64px rail. The current row gets the kit's
accent dot, hidden in the rail.

The theme, language and account controls stay in the sidebar footer:
the app has no global search or command palette to put in a top bar, so
a bar would only hold controls that already have a place.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): centre the avatar in the collapsed sidebar rail

The collapsed avatar link was set to "flex: 0", which gives it a zero
flex basis; with min-width: 0 the link shrank to its padding and the
icon overflowed from the link's left edge, about 14px right of the
icon column. Use "flex: 0 0 auto" in the collapsed and tablet rail.

The footer controls now share the nav icon column in the expanded
sidebar too (6px footer padding, 40px control boxes), and the tablet
rail gets the same footer padding and hidden language code as the
collapsed one.

New spec measures the centre x of the nav icons, mark, avatar,
language, theme and collapse icons in the collapsed, expanded and
tablet states, in both themes, and asserts they agree within 1px.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): size the console and settings rails and stack settings on phones

The Operate console rail, the Settings section rail and the account tab
bar still used 13px text and the old underline tabs. They now use the
15px nav size, 40px rows and the segmented tab control. Form row labels
are 15px with 13px hints.

On a phone the Settings section rail sat beside the form and squeezed
every row into a few characters. Below 720px the rail stacks above the
content as a scrolling row and form rows wrap their control below the
label. The save button no longer carries the icon font class, which
drew a missing glyph before its label. The language menu is wide enough
to keep Bahasa Indonesia on one line.

Assisted-by: Claude Code:claude-sonnet-5-5

* build(ui): update the vendored UI kit snapshot to 0.3.0

Take the 0.3.0 snapshot: the sprite now carries the full outline icon set,
and the new icons/fa-map.json maps Font Awesome names to icon ids. The map
lets the app move off Font Awesome in the following commits. The lock file
is regenerated with the new checksums.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add an Icon component backed by the kit sprite

Icon draws an inline svg that points into the kit's outline sprite. The
sprite is inlined into the page once, so the references resolve under any
base path and in the embedded build without a request. Icons size with the
font (1em), take currentColor, hide from assistive tech unless given a
title, and spin on request. An unknown id draws a neutral circle.

FaIcon and iconFromFa resolve Font Awesome names through the kit's map,
for names that arrive at run time. iconHtml does the same for markup built
as a string. The GitHub and Apple marks are small local glyphs, as the kit
ships no brand marks.

Assisted-by: Claude Code:claude-sonnet-5-5

* refactor(ui): draw shared components and helpers with Icon

Replace the Font Awesome elements in the shared components and in the
utility modules with the Icon component. Lookup tables now hold kit icon
ids instead of class strings. Code-block copy buttons and artifact cards,
which build HTML strings, use iconHtml and a sanitizer-safe slot.

Assisted-by: Claude Code:claude-sonnet-5-5

* refactor(ui): draw model, backend and account pages with Icon

Replace the Font Awesome elements on the home, models, backends, import,
settings, login, account and users pages with the Icon component.

Assisted-by: Claude Code:claude-sonnet-5-5

* refactor(ui): draw chat, studio and recognition pages with Icon

Replace the Font Awesome elements on the chat, media generation, talk and
face and voice pages with the Icon component. The talk status table keeps
its spin and pulse states as Icon props. The connected and error states
now use a dotted circle and an alert circle, so they differ from the idle
ring by shape as well as by colour.

Assisted-by: Claude Code:claude-sonnet-5-5

* refactor(ui): draw agent, node and operate pages with Icon

Replace the Font Awesome elements on the agents, skills, collections,
jobs, fine-tune, quantize, nodes, swarm, usage, traces and activity
pages with the Icon component. Two class strings on layout elements held
leftover button and icon classes from an earlier merge; they are cleaned
up so the elements keep only their own classes.

Assisted-by: Claude Code:claude-sonnet-5-5

* style(ui): size and align icons for the svg component

Icon rules that targeted the font element now target the svg: the
descendant "i" selectors in App.css and auth.css become ".lai-icon". The
svg is 1.2em with a 2 unit line so it matches the visual size of the old
glyphs at the 12 to 16px sizes the app uses, sits on the text baseline,
and follows the context font size. Large empty-state marks get a lighter
line. Menu icons get a 16px box and the readiness badge icons keep their
20px circle with padding. Add the pulse used by the talk status.

Assisted-by: Claude Code:claude-sonnet-5-5

* build(ui): remove Font Awesome

No source file references the icon font any more. Drop the package and its
stylesheet import. The build no longer ships the solid, regular and brand
font files.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): keep focus traps off the svg use references

The dialog and drawer focus traps collect focusable elements with a
"[href]" selector. An icon's use element carries an href, so it became the
first "focusable" element and Tab at the end of the dialog stopped there
instead of wrapping to the first button. Match "a[href]" instead.

Assisted-by: Claude Code:claude-sonnet-5-5

* style(ui): keep icon sizes overridable and set the line width per svg

Give the icon base rule zero specificity so a rule that sizes one icon
(nav column, menu box, avatar, language switcher) wins whatever its order
in the file. The sprite symbols fix their own line width; the inlined copy
drops it so the width set on each svg applies, as the --lai-stroke custom
property, and large marks can use a lighter line. Pin the avatar and the
language globe to the boxes the sidebar alignment spec expects. Import the
map as JSON with an import attribute so Node can load it in the spec.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): select icons by the svg markup and cover the sprite

Specs that found icons by their Font Awesome class now select the svg by
its data-icon. The dead-icon audit checks that every svg resolves to a
sprite symbol and has a size. The class hygiene spec fails on any
remaining Font Awesome class. A new spec checks every mapped icon id has
a symbol, that the sprite is inlined once, that an icon paints at the root
and under a forwarded path prefix, and that Font Awesome names map as
documented.

Assisted-by: Claude Code:claude-sonnet-5-5

* build(ui): update the vendored UI kit snapshot to 0.4.0

Take the 0.4.0 snapshot: hub tabs with count and attention badges, the six
chart series tokens and the grid colour in the theme contract, and sample
themes on a calmer palette. The kit headers are renamed and the lock file
is regenerated with the new checksums, as for the earlier snapshots.

Assisted-by: Claude Code:claude-sonnet-5-5

* style(ui): switch the theme to the calm palette

Rewrite the LocalAI theme on the calm palette: a muted teal accent on a
near-neutral green-grey canvas, desaturated status colours, no glow and no
coloured shadows. The theme fills every role of the shared UI kit's 0.4.0
theme contract for light and dark, including the six chart series and the
grid line. The bridge in theme.css keeps the old --color-* names working,
adds the dark surface ladder (card, raised, float) and a strong edge, and
points the fixed data hues at the chart series.

Two values differ from the first sketch. The dark text on the accent fill is
#021512 instead of #04201d: it reads 5.58:1 on the fill at rest and 6.4:1 on
the hover fill, against 5.08:1 at rest for the lighter value. The light
control edge is #6b7d7a. The kit's contrast script passes for all text pairs
(4.5:1), control and focus pairs (3:1) and series colours (3:1).

Leftovers that no longer fit the palette are fixed: the usage chart takes
the six series colours in order, the audio and animation canvases fall back
to the new accent, the face box loses its glow, and two gradient fills are
now flat. The theme tests expect the new canvas colours.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): replace the console rail with hub tab bars

Build and Operate no longer open a second navigation rail beside the page.
Each is a hub: one row of the kit's hub tabs above the page, with count and
attention badges that scroll sideways on a phone. Every URL and route stays
as it was, plus a new /app/build landing page that lists the Build tools
with a line each.

Build tabs: Overview, Agents, Skills, Memory, Jobs, Fine-Tune, Quantize,
Import, Voices (recognition and library) and Faces. Operate tabs: Status,
This machine, Swarm (distributed mode only), Runtime (backends, activity,
failover), Traffic (usage, traces, middleware) and Settings (settings,
users), plus the API link. A tab that holds several pages shows a second row
of links, and a sub-page such as a node detail keeps its tab highlighted.
The feature and admin gates decide which tabs are drawn, and badges show only
values the Operate summary already has.

The sidebar lists Build and Operate under a Workspace label next to the
Create group. The voice library moves under Build and the model import page
gains the Build tab bar. The old rail styles, the rail signals and the
console config are removed, and the Operate overview docs describe the tab
bar. The specs that drove the rail now drive the tabs, and a new spec covers
the tab for each route, gating, badges and the phone layout.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild Home as a calm console

Home now opens on one command bar: the model chip shows which models are
warm, the MCP chip and attach buttons sit beside it, and Send is a solid
button with an Enter glyph. Typing "/" opens a grouped, keyboard-driven
action list built on the kit command list; every action has a destination
in the product.

Memory use folds into a one-line strip that opens into the loaded models,
with Stop per model and Stop all. It opens by itself while a model is being
staged and after a failure, and shows nodes and aggregate memory in a
cluster. The list of resident models carries no per-model size because the
API reports none.

"Jump back in" lists the conversations stored in the browser, one card per
day, with j and k to move, Enter to resume and delete with an undo toast.
First run keeps the install steps and the recommended models. The assistant
prompt is a dismissible line, the library links are one quiet row and the
API section is collapsed. Chat accepts an empty new-chat hand-off for /new.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Home console

Update the Home specs for the new structure and add specs for the slash
menu, the model chip, the memory strip (expand, stop, staging, failure,
cluster), the resume list (grouping, j/k, Enter, delete with undo), first
run, the send hand-off, a non-admin user and the phone layout.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add the fit, disk and cleanup helpers for the models page

Pure functions and hooks that the rebuilt Models page reads, with node
tests for the rules.

modelLedger turns an estimate and the memory budget into one of three
verdicts (fits, spills to CPU, over) with the headroom in bytes, and
reads the models disk from the resources reading. The disk counts as low
under 10 percent or under 20 GB free, and is absent when the server
reports none or runs as a cluster controller.

cleanupPlan ranks installed models from what the API reports: loaded,
pinned, or named by an agent, a task, a failover chain or an alias keeps
a model protected; another installed build of the same gallery model is a
duplicate; disabled models rank above idle ones. The API records no last
use or use count, so none is used. When a lookup fails, nothing is called
safe.

useModelRemoval holds a removal in the browser for an undo window and
sends the existing delete call only when the window ends. Leaving the page
drops the batch without deleting anything. The undo toast takes optional
labels so other pages can reuse it.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild Models as a ledger with a disk strip and cleanup review

Explore is one dense table. Each row carries the size, a solid memory bar
and the headroom in words ("3.7 free", "+1.5 on CPU", "0.9 over"), worked
out from the estimate at the chosen context length. Capability chips show
the server's count for each facet, search keeps its meaning and "/" jumps
to it, and a density switch (also "d") picks comfortable or compact rows.
Selection is a surface step and a check, never a rail. Arrow keys move,
Enter installs and Esc closes the inspector, which keeps the fit summary,
VRAM by context chart, variants, files, links, tags and licence. A failed
install shows its error in the row with a Retry that dismisses the old
failure first. A failed or empty listing says which it is, and a host with
no GPU is measured against memory and says so.

Installed uses the same table with state filters that carry counts, a
state per row, Load or Stop on the row, the row menu and the sort by size.
Sizes come from the files the gallery lists, so a model it does not know
shows a dash.

A strip in the header shows the free space on the models disk. It turns
amber under 10 percent or under 20 GB free, hides when the server reports
no disk or runs as a cluster controller, and opens the cleanup review.
Explore says how much an install leaves free.

The review ranks installed models as Safe to remove, Probably safe and
Your call from real facts only, lists protected models with the reason,
and says plainly that usage history is not recorded. A sticky bar shows
what a choice frees. Confirming runs a dry run that checks again and lists
what will go. Removal waits 30 seconds with an undo; nothing is deleted
before that, and leaving the page deletes nothing.

The old rail, filter band and popover styles are removed.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Models ledger, Installed table and cleanup review

Update the Models, lifecycle, cluster fit, height, search focus and
surfaces specs for the table and inspector, keeping what each one checks.

New specs, on a shared 41-model gallery stub with three machine profiles:
the fit bar and headroom words for a 24 GB card, an 8 GB laptop and a host
with no GPU; facet counts, search, "/" and Escape; selection, arrow keys,
Enter to install, density; the disk strip when normal, low and hidden; and
the states (loading, empty, offline, install failed, phone). Installed
covers filters with counts, row actions, the row menu, sizes and sort.
The cleanup specs cover grouping, protected models, the honest-data note,
the effect bar, the dry run, the undo window, a failed delete, leaving the
page, and the phone sheet.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add the placement helpers and the estimate hooks

The Placement section and the model page need the same few rules, so they
sit in plain functions that can be read and tested alone.

placement.js holds what gpu_layers, tensor_split and main_gpu mean (unset
asks for every layer and the llama.cpp engine trims it, zero is CPU only,
99999999 is the value LocalAI itself writes for all layers), the device
list taken from the resources reading, the split by free memory, the part
of an estimate that grows with context (read from two lengths, since that
term is linear), the fit states with their limit (95 percent of free
memory, and the leftover has to fit in system memory too), and a bisection
for the largest layer count whose estimate fits. The estimate returns one
total and no layer count, so the search runs over 1 to 256 and stops at
the first count that no longer changes it.

modelWalk.js keeps the order of the list a model page was opened from, in
memory and in session storage, for the previous and next buttons.

usePlacementEstimate reads /api/models/vram-estimate for a choice, again at
twice the context, and with every layer, and keeps readings for the
session. useModelPage reads a gallery entry by name, an estimate by
context size (from the model's own files when the gallery does not list
it), the builds and the loaded models. usePlacementConfig edits the four
placement keys of an installed model and saves only what changed.
useModelActions is the Load, Stop, disable, pin and remove logic of the
Installed table, shared with the model page. MemoryBar is one solid bar
with a tick at the capacity of its pool; over capacity it grows past the
tick and the tick turns red.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add the Placement section to the model editor

Run this model on: CPU only (gpu_layers: 0), Auto (the key stays unset) or
Custom. Custom takes a number, has an All layers button that writes
99999999, and shows a slider only when the estimate reports the model's
layer count, which it does not today. Context size has presets and a
number field because the KV cache follows it. With two or more GPUs there
is a split (written as percentages, with a button that takes them from the
free memory of each card) and a main GPU.

A bar per GPU and one for system memory show what other programs use, the
model's weights and working memory, and the part that grows with context,
with the room left or how far over it is. Under them a verdict in plain
words: Fits in GPU, Spills to CPU, Too many layers for the GPU, Runs on CPU
only, No GPU found, Not enough memory. It says "slower" and never a
multiplier, because the estimate has none. Fit it for me asks the estimate
for the largest layer count that fits the free GPU memory and says what it
set, with Undo; it is hidden when the estimate is unavailable or the host
has no GPU. Loading shows skeletons, an unavailable estimate shows a note
with Retry, and a server that schedules onto other machines shows no bars,
because its device list is the controller's.

The editor shows the section for an installed model, with a link in its
section rail. Auto sends null for the key, since a patch only merges, and
a null read back opens as Auto. The docs describe the section and what each
mode writes.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): open a model on its own page

A model has an address, /app/models/<name>, for an installed model and a
gallery entry alike. Open it from the arrow at the end of a row, a double
click, "o" on the selected row, the inspector's Open details button, or a
tap on a phone. The title block holds the main action: Install with a
chevron that chooses the build, or Load and Stop with a menu (disable, pin,
edit configuration, logs, delete with a confirm). A strip answers whether
it fits, what it does and what installing leaves free.

Tabs: Overview (about, a memory bar, state, the pages it opens in, and the
agents, tasks, chains and aliases that name it); Fit and memory (verdict,
context sizes, the bar split into weights and context, and memory by
context against the limit, with a data table); Variants and files (builds
with size and fit, install any, the files of the chosen build). For an
installed model also Usage and history, which says what the API does not
record instead of drawing an empty chart, Configuration, which is the
Placement section with the file it writes and a link to the full editor,
and Logs, the backend log viewer without its page. Keys 1 to 6 switch
tabs, [ ] and j k walk the list the page was opened from, Esc or Backspace
go back.

The list stays mounted behind the page, so Back finds its view, search,
filters, selection and scroll as they were, and focus returns to the row's
arrow. The page covers loading, an unknown name with the closest matches,
the gallery being out of reach, an install in progress with Cancel, and a
failed install with Retry.

The docs describe the page and its keys.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the model page and the Placement section

New specs for the model page: reaching it from Explore, Installed, a
double click, "o", a pasted link and a phone tap; the walker and Back with
the search, a filter, the selection, the Installed view and the scroll
kept, and no second read of the gallery; the title block, the answer strip,
tabs by click, keys and arrows; Fit and memory, builds and files with the
install call each one makes; an installed model's actions, used-by,
the honest usage tab, configuration and logs; loading, an unknown name,
offline, an install in flight and a failed one; and the phone.

New specs for Placement: every mode and the keys it writes, the slider
only when a layer count exists, the context presets, the bars and every
verdict, two GPUs, no GPU, a cluster, an unread machine, a loading and an
unavailable estimate, Fit it for me and Undo, and the section in the model
editor with its save.

The phone tap on a row now opens the page, so the two phone specs that
expected the inspector as the page check the page and keep the inspector
check for a window between a phone and a desk.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): link Studio results and open workspaces from a prompt

Each workspace now records the result it was made from (parentId and an
edge kind such as take, animate or to-3d) and reads a prompt, model, size,
count and source from the query string, so one page can hand work to
another. A source result is fetched from the server's own output file and
becomes the start image, the picture for 3D, or the audio file. A note on
the page says when the source loaded or could not be loaded.

Diarization had no history; it now keeps the file name, the model and a
speaker count, never the recording. Prompts are cut at 2000 characters
when stored. The pure helpers (type suggestion, grouping, lineage layout,
favourites, clearing) have node tests.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild the Studio front page as a composer with your work

The front page is a prompt box with a chip per type, a type suggestion
from the words, starters, and the options each workspace accepts. Generate
opens the workspace with those filled in. A type with no model is a dashed
chip that shows a gallery model, its size, memory need and an Install
button only when picked; the typed words stay while it installs.

Under it, Your work lists results from every workspace as a masonry with
filters, counts, favourites and a Clear history action. Results made from
each other stack into a project tile and open as a lineage board with a
dock for running a new take or branching to the next step; steps the
destination cannot start from yet are disabled with the reason.

The docs describe the page, what is stored in the browser, and the query
parameters a workspace accepts.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Studio composer, your work and the lineage view

Specs for the type suggestion, the keys, hand-off to each workspace, the
install path for a missing model, the masonry filters, favourites and
clearing, stacking, the lineage board, new take and branch, steps that
are disabled with a reason, and the phone layout. Existing Studio specs
move from lanes to chips with the same intent.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add the shared Studio workspace frame and move Images onto it

The seven Studio workspaces get one layout: a row of type tabs, a compose
card (optional sources as chips, a prompt with starters, a model chip,
the essential options as chips, an Advanced fold that names what is
inside, the memory the model needs, and one action with the reason when
it cannot run), a run area, and a strip of recent results of the type.

The run area shows a job card with the time that has passed and an
indeterminate bar, because these endpoints report no phase or percentage;
a failure with what the server said and one action; or the result with a
toolbar: Favourite (the list the front page keeps), Download, Use in (the
hand-off targets, disabled with the reason when a destination cannot
start from the result), Re-run with edits (the take's values go back in
the form, changed fields are outlined and listed) and Lineage. A type
with no model shows the install note from the front page.

Images is the first workspace on the frame. It keeps its size, count,
steps, seed, negative prompt, source image and reference images, and its
history writes, including the parent link and edge of a hand-off run.
useMediaHistory.addEntry now returns the id of the entry it stored. The
docs describe the workspace page.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): move Video onto the workspace frame

Video keeps its size list, duration, frame rate, steps, seed, CFG scale,
frame count, negative prompt, start and end image and avatar audio.
The start and end image are source chips, the avatar audio opens the
recording and paste input from a chip, and the rest sit in the Advanced
fold. A start image from a hand-off shows as a chip with its picture.
Results play in the video player with the shared toolbar.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): move TTS onto the workspace frame

TTS keeps the saved-voice picker for cloning models, the typed voice for
the others, the voice library deep link, and the delivery instructions,
which now sit in the Advanced fold. The result is the waveform player
with the words under it. The stored entry also keeps the voice id so
Re-run with edits can select the same saved voice.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): move Sound onto the workspace frame

Sound keeps its Simple and Advanced modes and every field of both: the
description, instrumental, vocal language, caption, lyrics, BPM,
duration, key, language, time signature and think mode. The mode switch,
instrumental and duration are in the compose card, the rest in a More
options fold. The stored entry keeps all of the fields, so Re-run with
edits restores the form as it was.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): move Transform onto the workspace frame

Transform keeps its audio and reference inputs with upload and record,
the echo test, the key=value parameters (now in the Advanced fold), the
input and output spectra and the three waveform players. The audio that
was chosen shows before the run, waiting to be transformed. Re-run with
edits puts back the model and parameters and fetches the audio and
reference the server kept for that run.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): move 3D onto the workspace frame

3D keeps the picture input with paste and webcam, the animation
operations a model declares, quality and background, the shape and
material steps, guidance and seed, the GLB and animation viewers, the
remesh control and the download. A 3D result now has a title from the
motion prompt when it has no label, so the strip and the front page name
animation results by what was asked. Re-run with edits is shown disabled
with the reason, because only a small thumbnail of the picture is kept.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): move Diarization onto the workspace frame

Diarization keeps its model and recording inputs, the option to prepare
speakers to remember, the clean-speech previews, naming and remembering a
speaker, and the history entry with only the file name, model and
counts. The result now shows a timeline with one lane per speaker, the
talk time of each speaker, and the segments with their start time and
text. RTTM, SRT (only when the run has text) and JSON are built in the
browser from the result. The helpers for talk time, axis ticks and the
two text formats have node tests.

Assisted-by: Claude Code:claude-sonnet-5-5

* refactor(ui): remove the styles and lists the old workspace layout used

Nothing renders the two-column workbench, the control column, the old
history lists, the generation progress tiles, the TTS voice picker or the
result echo any more. The inline-style baseline drops with them.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the workspace frame and one run per type

Specs for the type tabs, the compose card and its reason when Generate
cannot run, starters, the Advanced fold, the job card with no invented
progress, a failed run and its one action, the install note, the strip
with its favourites filter, Use in with its disabled steps, Lineage, the
parent link, Re-run with edits and its list of changes, deleting and
clearing, and the hand-off note. One run through each of Video, TTS,
Sound, Transform, 3D and Diarization, the phone layout of all seven, and
reduced motion. Existing Studio specs move from the old control column to
the compose card with the same intent.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild the chat thread with raised user turns, prose replies and one-line activity

Your messages are raised blocks on the right at a 760 px measure and the
model's replies are plain prose under its name and a warm or not loaded
dot. Reasoning, tool calls and their results fold into one quiet line
that opens inline into steps. Code blocks carry a Copy button and a
Canvas button that opens that block in the canvas, image attachments are
thumbnails that open in the lightbox, and files are chips. Per-message
actions show on hover, on focus and on the last turn, and a turn takes
focus so the arrow keys and C, E, R and B work. A failed reply keeps the
text written so far and shows the reason with one Retry action.

The Agent chat page keeps the older rules: the new styles are scoped to
the chat page and use their own class names.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): use the Home command bar as the Chat composer

Chat now ends in the same object as Home: the model chip, the MCP chip,
a Canvas chip, the message box with attach buttons, a solid Send and the
hint line, with the slash menu on the kit command list. The slash menu
lists what Chat can do today (switch model, new chat, conversations,
manage mode, canvas, find, settings, export, clear). While a reply is
streaming Send becomes Stop, which Esc also presses, and Up in an empty
box edits your last message. Attached images show as thumbnails and a
line under the bar carries the speed and the token count.

HomeComposer takes optional props for this (extra chips, its own slash
list, Stop, paste, a stricter Enter); Home passes none of them.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): open conversations from a Ctrl K menu with day groups, undo and a slim header

The conversations list opens as a centred menu on Ctrl or Cmd K. It
groups chats by day like the Home resume list, shows the model that
answered and the time, searches names and message text, and moves with
the arrow keys. Enter opens a chat, F2 renames it and Delete removes it.
Removing a chat hides the row and shows the kit undo toast; the chat is
deleted for good only when the undo time ends. Rename, duplicate, copy
and export are on each row, as before.

The header is one slim bar: the Chats button, the chat name (click to
rename), a context meter when the context size is known, settings and a
More menu with rename, duplicate, copy, export, model info, keyboard
shortcuts and clear. A dialog lists the shortcuts the page answers to.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): show loaded state, capabilities and fit in the Chat model switcher

The model chip in Chat opens the same list as Home, grouped as Loaded
now and Installed. Each row says warm or not loaded and marks models that
understand images. When the list opens, the page reads the host memory
once and asks the server to estimate each listed model at the chat's
context size (up to twelve, three at a time), then shows what the model
needs and whether it fits: free memory, how much would run on the CPU, or
how far over the machine it is. A model with no estimate shows no fit
text, and no load time is shown because the API does not report one. A
memory bar closes the list.

The picker takes the model list from the page when it has one, and
useModels can skip its own request.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): move chat settings into a sheet and add find in chat, jump to latest and a wider canvas

Settings open as a kit sheet: the system prompt, temperature, top P and
top K (each says "model default" until it is changed and has a Reset),
the context size with quick sizes and a note that it only drives the
meter, Manage mode and Focus mode, the model info for admins with its
Edit config button, and Clear conversation behind a confirmation. The old
slide-out drawer and the model info panel are gone.

Ctrl or Cmd Shift F (or the search button, or /find) opens a search bar
over the thread. It marks matches in the messages already on the page,
shows "n of m" and steps with Enter and Shift+Enter. Nothing is sent to
the server. Jump to latest is a pill above the composer. Esc stops a
reply, then closes the search, then closes the canvas.

The canvas panel gets the kit look: tabs, a Code and Preview switch, Copy
and Download, a full-page layout on narrow windows, and translated
labels. The Agent chat page shares it and gets the same look.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add the empty, no-model, loading and phone states to Chat

An empty chat opens with the composer under one line, starters to try,
whether the model is loaded, and the Jump back in list: the same rows as
Home, read from the chats the page holds. With no chat model installed,
an install card offers the starter models for this hardware, the gallery
and import, and the composer stays so the text is not lost.

While a reply waits for a model, a load card shows what the page knows:
the phase the server names, the node, the bytes and the time left when the
server reports them, and a progress bar. A model that is just not loaded
yet gets a plain note, with no invented phases or estimates. The foot
warns when the context is nearly full.

On a phone the header drops its labels, the model list and the settings
open as sheets from the bottom, per-message actions stay in view and the
canvas takes the whole page.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild Talk as a calm voice page over what the connection really does

Talk is one stage and one transcript. The stage has the pipeline chip,
the voice and language chips, an outline orb that follows the real
microphone and playback levels, a heading and a sentence for the current
state, and the controls. The transcript lists You, Reply, Tool and Result
lines and can be copied. Session settings (instructions, voice, language,
tools, Manage mode and the pipeline's parts) open in a sheet.

The states are the ones the code reaches: no pipeline model, idle,
connecting, listening, thinking (also while a tool runs), speaking, an
interrupted reply (the server cancelled it; a note marks the cut), a
blocked microphone, a link that failed during a session, and any other
error with its reason and a link to the traces. Push to talk and
hands-free are not on the page, so they are not shown. Diagnostics keep
their waveform, spectrum and stats, drawn in theme colours.

The page text moves into the talk namespace, and the old Talk and
visualizer styles and the inline-style count go down with the rebuild.

Assisted-by: Claude Code:claude-sonnet-5-5

* refactor(ui): remove the chat styles and strings the rebuilt page replaced

The settings drawer, the model info panel, the bubble avatars, the
conversation menu popover, the context bar, the recent strip, the
staging bar, the file badges and the focus-mode rules have no user now.
Their rules, the Chat page's focus class and seven unused empty-state
strings are removed. The Agent chat page keeps the shared message,
sidebar and input rules it still renders with.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the rebuilt Chat and Talk pages

Add a Chat page under features (thread, message actions and keys, the
message box and its slash actions, the model list with loaded state and
fit, conversations on Ctrl K, settings, find, canvas and the empty,
no-model and loading states) and a Talk section to the realtime API page
with the states the page shows. Manage mode now turns on from the chat
settings or /assistant, and the client MCP steps point at the MCP chip.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): settle the rough edges of the new Chat and Talk pages

The undo toast sat under the conversations menu, so the Undo button could
not be pressed while the menu was open; the menu, the sheets and the
fullscreen canvas now stay below the toast layer. Esc in a rename box
saved the text through the blur that follows it; it now cancels. The
image viewer closed on Esc only when the page did not re-render on the
same key, so its key listener is registered once and reads the latest
handlers. Keys on a focused message no longer type their letter into the
editor they open, "/" from outside a text field starts a command as it
does on Home, and Esc leaves the page's own dialogs alone.

Code in the canvas is highlighted for languages that have no preview.
The conversations menu drops its key hints on a phone so Clear all
stays in view. Talk hides Test tone while connecting and calls a server
error "Something went wrong", since the call can still be open.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the rebuilt Chat and Talk pages

Specs for the thread layout and the activity fold, code blocks, image
thumbnails and the viewer, per-message actions and their keys, a failed
reply with its one Retry, Stop and Esc while streaming, the composer and
every slash action, the conversations menu (groups, search, resume,
rename, delete with undo that ends by itself, one chat left), the model
switcher with loaded state, vision and fit text from stubbed estimates,
the settings sheet, the canvas panel, find in chat, Jump to latest, the
empty, no-model and loading states, the phone layout and reduced motion.
Talk is driven over a fake WebRTC link through idle, connecting,
listening, thinking, speaking, interrupted, blocked, lost, error and no
pipeline, its settings sheet and its phone layout. Node tests cover the
message text helpers and the conversation grouping.

The existing chat specs move to the new structure with the same intent:
the transcript spec now describes the raised turn and the prose reply, and
the render smoke accepts Talk's own header.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): keep a bounded run log for agents in the browser

The server keeps no run history for an agent, so a run is one task and
the events until the agent answers, written to browser storage while the
page watches the stream: up to 50 runs per agent, task, step and answer
text only. Stored chats from the earlier agent chat page read as runs
with stable ids. A run still marked running five minutes after its last
event reads as stopped. Helpers read an agent's config into chips, build
the list of changed fields against the saved config, hide secret values
and offer starting points.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild the Agents area around runs

The Agents page shows what needs a look (work in flight, a run that
failed in the last day), then each agent with its model, attached memory
and skills, and a strip of its last 14 runs. An agent has its own page: model,
tools, memory, skills, instructions, a task box and its runs. A run has
an address, shows the thread while it works (steps folded into one line,
the tool in use, the answer as it arrives) and settles into a report
about a second and a half after the agent answers: task, outcome,
follow-ups, evidence and steps, with wide tables opening wider on demand.
A failure says in plain words what happened and offers Run again.

Create and edit fold into sections with a ready mark and a one-line
summary, start from a template or an optional model-written draft, and
open a preview sheet with the config as saved and the changes against
the saved agent. Status becomes a quiet panel in the same language, and
the old chat link opens the agent page.

There is no Stop, approval, steer, version or dry-run control, because
the agent API has no call behind them.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Agents launcher, agent page, runs and editor

Specs for the Now strip and run strip, search and the empty state, the
agent page, starting a run, the live thread, settling into the report,
the run address across a reload and for a run from another browser,
follow-ups with their history, failures, the folding editor with ready
marks, templates, the preview sheet with hidden secrets and changes, the
status page, and the phone, 1440 and 2560 layouts.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe runs and the new agent create flow

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add helpers for tasks, schedules and job outcomes

Reads a cron expression the way the server does (five fields or an @
shortcut), checks it, and puts the common shapes in words. The next run
is left out on purpose, because the schedule follows the server clock,
which the browser cannot read. Also groups jobs by day, sums the last
seven days, and gives each job one outcome line from its result or
error. A rerun call starts a new job with the same parameters and media.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild the Jobs area around runs

The Jobs page opens with one sentence about the last seven days, then
the tasks (model, schedule in words, last 14 jobs, enabled switch, Run
now) and a run history grouped by day. Each row has an outcome sentence
and opens to the error or the start of the result with one next action.
Deleting a task waits 30 seconds with an undo button.

A task opens as a page with its recent runs, its prompt with the gaps
marked and its schedule. The task form folds into sections, takes a
schedule as a preset or a checked cron expression, warns about prompt
gaps the schedule does not fill, and has a preview sheet. A job opens
as a document: task, outcome, delivery and the recorded steps; a failed
job says what happened and offers Run again.

Run now now sends attached media through the job call, which is the
only one that takes it. "Clear History" only ever cancelled running
jobs, so it is now called Stop running jobs. Webhook headers of a saved
task show as JSON instead of [object Object].

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Jobs page, task pages and job pages

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the Jobs page and the task form

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add helpers that say who uses a skill or a collection

An agent loads a skill when skills are on and the skill is in its selection (an empty selection means every skill). It reads the one collection that carries its own name, when its knowledge base is on. The helpers derive that from the saved agent configs, build the config that adds or removes a skill or a collection, and estimate tokens as characters divided by four. Removing the last selected skill switches skills off, because an empty selection would mean every skill.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild Skills and Memory as one library

Skills and collections sit in a list with an open item beside it. Each row says who uses it, read from the saved agent configs, or says it is not used yet. Chat reads neither, so it is never named. An item opens in a pane with a Used by strip (names link to the agent, a small x removes it, with undo) and an Add to menu that shows what the addition costs. A collection can be added only to the agent that carries its name.

The Memory pane searches the collection alone and shows ranked passages with scores, lists web sources with their refresh interval and the files, shows the server message when an upload fails, and names the endpoints and where files stay. The Simulate a message sheet runs a collection search and shows an agent's skills with a token estimate. It runs no model. The collection details route now opens the same page.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): show use and cost in the agent form pickers

Each skill in the agent form says which other agents use it and what it adds to every message, with a total for the selection. The memory section names the collection the agent reads.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Skills and Memory libraries

Specs for the used-by lines (including an agent that uses every skill), the filters, search, add to agent, remove with undo, the last-skill case, an unreadable agent list, the empty states, git repositories, the Memory question box, sources, uploads that fail, the Simulate sheet with the parts the API can run, the agent form hints and the phone layout.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the Skills and Memory libraries

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add helpers for the Operate status page and backend rows

Pure functions for the parts that need rules. They work out the memory
pool the page measures (GPU memory, system memory, or the workers of a
cluster that are answering), which pools are too full, the headline and
the four ledger rows, the geometry of the capacity chart, and what
removing a backend would leave without a runtime (models name their
backend, and a meta backend names the concrete one it points at). A
second set says what a backend row states: installing, queued, removing,
failed, update available, current or absent.

LocalAI keeps no memory history, so the chart reads a bounded buffer of
readings the page took itself and says so. A reading with no total is
dropped rather than drawn as zero.

Two hooks are shared by the pages that need them. One retries a failed
operation after moving the failure into the record. The other holds a
cancel for an undo window, because the server cannot take a cancel back.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild Operate Status, This machine, Backends, Activity and Logs

Status opens with one sentence ("2 things need you", or "Everything is
running") and four rows: Needs you, Capacity, Running now and Recent
failures. A row with a problem opens by itself and holds the button that
deals with it: Update a backend, Retry or Dismiss a failed operation,
Unload a model. A quiet row stays one line. A new installation gets a
first-run screen, a cluster sums the memory of the workers that are
answering, and a page still waiting for an answer says so. The chart
under the rows is drawn from readings the page took while it was open
and is labelled that way, because LocalAI keeps no memory history.

This machine leads with GPU memory as one bar, then host memory split by
running model, then VRAM, RAM, CPU and disk with a bar each. The running
models become a kit table with the same menu and stop dialog.

Backends is one list with Installed and Catalog views. A row says what
the backend is doing (a progress bar with Cancel, Queued, Failed with
Retry, Update 1.2.0, Current), carries the one button that matters, and
opens in place. Removing a backend names the models and the meta
backends that would stop working. Check for updates, Update all, From
URL and a first-run recommendation for llama-cpp are in the header.

Activity keeps its three sections as quiet rows. Cancel waits eight
seconds with an undo toast, because the server cannot take a cancel
back; a cancelled install can be started again from the record. Logs
gets a process list, a picker, stream and text filters, Follow and
Times switches, and a Clear with an undo window.

Not shown, because the API has no data for them: GPU temperature and
power, a size per backend, an earlier version to roll back to, a
dependency lookup beyond the models and meta backends that name a
backend, and models that failed to load.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover Operate Status, This machine, Backends, Activity and Logs

New specs for the Status headline and ledger (healthy, needs attention,
one thing, a full memory pool alone, loading, first run, cluster), its
actions (Update, Retry, Dismiss, Unload with its dialog), the capacity
chart built from readings taken while the page is open and bounded, the
phone layout, no coloured edge on a row, and reduced motion.

The Backends specs cover the two views, install progress with Cancel and
its undo window, Retry on a failed install, Update, Update all, Check for
updates, removal with the models and meta backends it would break,
Install from URL, the first-run recommendation, a cluster, and a phone.
Activity gains cancel with undo, Cancel now, a second cancel, leaving the
page, progress, and starting a cancelled install again. Logs covers the
stream and text filters, Follow, Times, Export, Clear with undo, the
process picker and list. This machine covers the GPU strip, several GPUs,
no GPU and Add a machine.

Existing specs keep their intent and follow the new structure: rows open
in place instead of in a pane, Update replaces Upgrade, the notice spec
now pins that an update is a row state and not a banner or a rail, and a
cancel waits for its undo window.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the rebuilt Operate Status, Backends and Activity pages

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add helpers for the Swarm pages

Pure functions for what the pages work out from the cluster API: a node's
state in words, which nodes a placement rule may use, what a rule would
ask for, what a drain or a lost node would leave without service, the
nodes a bulk backend update reaches, and the join commands for a worker,
a peer instance and a memory shard. Hooks read the roster, the loaded
replicas and the rules.

Everything runs in the browser from data the page already holds, and
says when it cannot see free memory or disk.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild the Swarm hub: nodes, node page, placement rules, failover

Nodes is a sortable table with comfortable and compact rows, a Needs
attention filter by reason, a map of the cluster that is not drawn on a
phone, the running models, and a bulk backend update for the nodes that
drifted. A node is a page: state, vitals, a drain preview computed from
the loaded replicas and the rules, tabs for models, backends, logs and
capacity and labels, and Remove that asks for the node's name.

Placement rules are written as sentences, show where each model is
loaded now, and edit in a side sheet with a preview of the nodes a draft
could use. Deleting a rule waits a few seconds so it can be taken back.
Failover keeps its chains, adds what the router does when a worker stops
answering and a per-node preview of what would stop. Add a node covers a
registered worker, a peer instance and a memory shard, with a command to
copy and a live line that says when the machine arrived. P2P keeps its
page in the same vocabulary, and the node logs page follows the local
logs page.

Previews are labelled as worked out in the browser. Per-GPU readings and
node events are not drawn because the API does not return them. Failover
moves to Swarm when distributed mode is on. Legacy fleet components and
their styles are removed.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Swarm hub

New specs for adding a node (each join method, the command, copy,
waiting and found, approve, a single install, P2P, a phone), placement
rules (sentences, where models are loaded, the preview matrix, the sheet
and its preview, delete with undo) and failover on a cluster. Node
detail covers its tabs, the drain preview and its dialog, resume, remove
with the typed name, a node that stopped answering, and unload.

The nodes specs follow the new structure and keep their intent: the
table, filters, grouping, pagination, bulk actions, the map, and running
models with stop, logs and the loading, error and empty states. The
scheduling, failover, P2P, hub and smoke specs follow the renames.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the rebuilt Swarm pages

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): add helpers for the Traffic pages

Pure functions for what the pages work out from the usage ledger, the
trace summary, the trace buffers and the resources reading: the shared
time window, grouping, sorting and filtering of usage rows, chart series
and axes that start at zero, the overview figures, per-model statistics,
the state of a trace and the words for a failure, the backend operations
that ran during a request, CSV export, the Prometheus metric list and
scrape config, and a bounded buffer of host readings.

A figure whose source cannot say is null, never zero. The trace summary
call takes the window in hours, and a helper reads /metrics with its
status.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild the Traffic hub: overview, usage, models, host, traces, middleware

Traffic opens on an overview: five figures (requests, failed, p95, tokens
in and out) and three charts, each naming its source. A second row of
links reaches Usage, Models, GPU and host, Traces, Middleware and
Prometheus, and one time window is shared by the first three.

Usage groups by model, user or API key, filters, sorts, opens a row on its
own chart, exports the rows it holds as CSV or JSON in the browser, and
keeps the opt-in cost estimate and the quota forecast. A user who is not
an admin sees only their own numbers. Models joins the ledger, the
backend-operation buffer and the loaded models. GPU and host shows the
current reading and two charts of readings taken since the page opened.

Traces gets filters, a settings strip and an explained off state. An API
request is a page: the error LocalAI recorded, a timeline with the backend
operations that ran meanwhile, and bodies that stay closed until revealed.
Middleware draws the pipeline as five steps and shows the rules of the
selected step. Prometheus documents /metrics, checks it against the
server and gives a scrape config to copy.

Alerts is not built: LocalAI has no alert rules. Per-model latency
percentiles, GPU utilisation and compare with the previous period are not
drawn because the API does not return them. Legacy usage, trace and
middleware styles and the usage source components are removed.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Traffic hub

New specs for the overview (figures, charts with a data table and arrow
key readout, failed and first-run and tracing-off states, the shared
window, a phone), usage (group by, filters, sort, export, cost, quotas, a
non-admin, empty and loading), models, GPU and host (snapshot, the
since-opened labelling, a cluster), the traces list, a trace page (the
real error, the timeline, reveal, no headers, a trace that left the
buffer), Prometheus and the Middleware pipeline, with shared fixtures.

The usage, traces, middleware, hub and smoke specs follow the new
structure and keep their intent.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the Traffic hub

Add an operations page for the Traffic tab: which record each page reads,
what it leaves out and why, the trace page and its reveal, the GPU and
host readings kept since the page opened, and the Prometheus endpoint.
Link it from the operations index, the tracing page and the middleware
page.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): let metric names wrap in the Prometheus table on a phone

The long metric names pushed the type and "on this server" columns out of
view. Names now wrap inside the table.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild Settings with groups by intent, search, a pending bar and history

The fifteen sections become eight groups by intent: memory and models,
speed and defaults, backends and galleries, access and security,
debugging and traces, agents and responses, swarm and sharing, look and
feel. Search covers names, descriptions, keys and the old section name, and
says where a result used to be.

Edits wait in a bar with Discard, Show diff and Apply. The diff lists old and
new values and the checks the browser can make: durations parse the way Go
parses them, a GPU memory budget is one the server accepts, a gallery box
holds JSON, and warnings repeat what the handler and the field text say.
Apply sends only the changed keys. Undo saves the previous values again; it is
a new save, not a rollback. History lists the changes applied from this
browser, since LocalAI keeps no settings log, and Revert stages the old value.

A value is marked as changed only where the built-in default is known from
the CLI defaults. A row says "Applies now" or "Needs restart" only where the
handler or the docs say so.

Three things were wrong before and are fixed with the rebuild: the gallery
boxes and the shared API keys box were sent under names the server ignores,
the "Enable CSRF Protection" switch showed the disable flag the wrong way
round, and every save restarted peer-to-peer networking because every field
was sent.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild Users and keys, Account, sign-in, invite and the 404 page

Users and keys is a tabbed page under the Settings tab: people, invites and
API keys. The people table filters by state and role, sorts, approves or
disables (disabling offers an undo that sets the status back), and opens a
side sheet for one person's features, model allow-list and limits. Role,
password reset and delete sit in the row menu; delete asks for the name.
Invites choose a lifetime of 1, 7 or 30 days and show the link once. API
keys can be created with a lifetime, are shown once in full, can be paused,
and are revoked after a ten second undo window in which nothing is sent.
LocalAI lists keys only to their owner, so the tab shows the signed-in
person's own keys and says so.

Account has Profile, Security, API keys and Usage. Usage shows the last 30
days, tokens by model and the limits an admin set. The Security tab now
shows for a GitHub or SSO account and says the password is not theirs to
change.

Sign-in asks for one field per step and draws a provider button only for a
provider /api/auth/status lists. It has the notice for a sign-up that waits
for approval, the first-admin screen, the key-only screen and the invite
page. An address outside the app now gets the 404 page too, which names the
address and lists the places the sidebar lists, with the same gates.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover Settings, Users and keys, Account, sign-in and the 404 page

Settings: groups, search by name, key and old section, the changed marker
only where a default is known, apply hints, the pending bar and diff, the
checks, apply sending only changed keys, undo as a second save, discard,
history, the CSRF inversion and the gallery and API key wire forms, and the
phone layout.

Users and keys: the table, filters, sort, approve, disable with undo, the row
menu, the access sheet, invites, key creation with a one-time reveal, the
ten second revoke with undo and with a page leave, and the non-admin redirect.
Account, each sign-in variant (error, pending, first admin, key-only, invite,
provider buttons) and the 404 page have specs too. Fixtures are shared with
the screenshot scripts. Existing specs follow the new structure.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the rebuilt Settings, Users and keys, Account and sign-in pages

Runtime settings: the eight groups and where each old section went, search,
the pending bar, the diff and its checks, apply, undo, the history, and which
settings show a default or an apply note and why. Authentication: the
sign-in screen variants, the Account tabs, key lifetimes, the one-time key
reveal, the revoke undo window, and the fact that keys are listed only to
their owner.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): let the Settings undo toast stand alone and read back a generated P2P token

The saved message and the undo toast sat on the same spot at the bottom of the
page. The undo toast now carries the saved message.

A new P2P token is made by the server when the page sends 0. The page reads
it back after the save so the field shows the token and not the placeholder.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): fit the users table, API keys and Account figures on a phone

On a phone the users table dropped its Role and Status columns off the screen
edge with the row actions. The role and state now sit under the name, so the
actions stay in view. API key rows no longer put the key icon on a line of its
own, and the three Account figures keep one row.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the phone users table, reduced motion and the empty Account state

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): drop the apply note from three settings the save handler does not mention

Size-aware eviction, automatic backend upgrades and development backends said
Applies now, but nothing in the handler or the docs says when they take effect.
A row now carries a note only where the code or the docs say so.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild Voices and Faces as one identity family

Voices is one page with three tabs: Speakers (voiceprints for recognising
who is speaking), Speech voices (the text-to-speech reference library,
kept apart because it is a different store) and From a recording (a link
into the diarization workspace). Faces uses the same layout.

Who is this and Same person? give the answer in a sentence with the real
distance and cut-off, a word for how far inside the cut-off it sits, and
a distance scale with the cut-off drawn on it. The cut-off slider re-reads
the answer in the browser; the identify call sends the cut-off, and verify
uses the threshold the model returns. The old confidence percentage is
gone because it is not a probability.

The server has no list call, so the people list stays in the browser and
the page says so. After a search that asked for more people than it got
back, a saved person the server did not return is marked, and people the
server returned that the browser does not know are listed. Nothing is
claimed from a short or cut-off search.

Enrolling is a sheet: sample, name, labels, permission. A copy of the
sample in the browser is opt-in, and an administrator can also keep the
recording as a speech voice in the same step. Removing a person waits ten
seconds behind an Undo toast and sends nothing before then.

Errors say what happened (no face found, model missing, call failed), a
blocked or missing microphone is explained, and a missing model or a
missing permission renders a page that says what turns the feature on
instead of a redirect. Analyze, detect and raw embedding move under
More tools, with attribute guesses off by default.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Voices and Faces pages

Specs for who is this (match, no match, working, failed, missing model),
the cut-off slider, same person, a blocked, allowed and insecure
microphone, the registry notes and the not-on-the-server marks, the
enrol sheet and its opt-in copy, delete with undo on a fake clock, the
disabled and no-permission states, the phone layout, reduced motion and
Faces. Existing library and diarization specs follow the new structure
and keep their intent. Node tests cover the distance words, scale
layout, stored list and error mapping.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the Voices and Faces pages

Add a WebUI section to the voice and face recognition pages: the two
tools, the cut-off, what the people list is and why it can be stale, the
undo window, and what is stored where. Point the Voice Library and
Fish Audio notes at Build, Voices, Speech voices.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): rebuild the Build landing, Fine-tune, Quantize, Import and Explorer

The Build landing says what each tool is for and what it needs from the
machine: the installed backend, the GPU memory, RAM and disk the server
reports, and a job that is running or the newest one when it failed. A
tool that cannot run says why and what enables it.

Fine-tune and Quantize share one page: set up, a check list that is
redrawn as the form changes, a run view with progress, stages and a log,
and a result with real next steps (export, import, chat, Models). The
checks state only what the server reports. A job needs no estimate the
server cannot make, so none is invented. Stop on a fine-tuning job asks
whether to keep a checkpoint, a failed job shows the server's message,
and a memory failure offers two changes that are applied to a copy of
the setup.

Import is a guided flow: source, review, import, done. The server
returns no preview before an import starts, so the review reads the
spelling of the source, prints the request the form will send and runs
the checks that can be made early. The estimate that arrives when the
import starts is set against free memory and disk. The ambiguity picker
and the Write YAML tab stay.

Explorer shows what GET /networks returns and lists a swarm with POST
/network/add, with a join sheet that carries the token and commands.
Build tools the account may not use say so instead of redirecting.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the Build landing, the tool pages, Import and Explorer

Specs for the landing (a tool ready, missing a backend, with no GPU, a
running or failed job, a feature switched off, a member without admin,
phone, reduced motion), the shared tool pattern for Fine-tune and
Quantize (set up, live checks, start request, running with progress,
chart and log, the stop choice, failure with the server message, finish
with next steps, earlier jobs, the account-disabled page, phone), Import
(source detection, review, checks, ambiguity, running with the estimate
against free memory, done, Write YAML, phone) and Explorer (list, join,
list a swarm, empty, not an explorer, retry, phone).

Existing specs follow the new structure and keep their intent. Node
tests cover the machine facts, tool status, checks, log lines, source
detection, the import request and the join commands.

Assisted-by: Claude Code:claude-sonnet-5-5

* docs: describe the Build tool pages, the import flow and the Explorer

Fine-tuning and quantization now describe the set up, check, run and
result steps and what the check list can and cannot say. The import
section explains the review step and why the size and memory appear only
after the import starts. The distributed page describes the Explorer
list, the join sheet and what listing a swarm publishes.

Assisted-by: Claude Code:claude-sonnet-5-5

* feat(ui): turn the hardware recommendations into a "Best for this machine" shelf

The shelf in the Models inspector put five columns into a 400 px pane,
so long model ids wrapped letter by letter underneath the size and the
memory figures. Each row now stacks the tag, the id and the size and
memory facts beside one Install button, and the id wraps inside its own
column.

Once a model is installed the shelf narrows to the best fit and keeps
the others behind a "N more that fit" toggle. Specs cover the ranking,
the layout, the narrowing and the install request against a gallery
fixture that carries the 4K estimate the shelf sizes against.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): quiet the Studio tab markers and say their state in words

The type tabs drew a saturated green dot for every modality that has a
model. The dot now uses a text colour, filled when a model is installed
and hollow when none is, and each tab carries "(model installed)" or
"(no model installed)" as hidden text so the state is not only a colour.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): stack the model editor empty state and drop its section hues

"No fields configured" sat in a flex row, so the icon, the title and
the text ran together. It now uses the stacked empty-state layout. The
section icons took a different status colour each (amber, red, green);
they now share one quiet colour, with the accent on the current section.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): keep the Home memory sentence whole on a phone and drop side rails

On a 390 px screen the memory strip clipped "2 models loaded" to make
room for the figure. The sentence now takes the first line and the
figure and device wrap under it.

The sweep also removed coloured left rails from the editor section
rail, the skill editor list, the install strip and the audio transform
notice (now an outlined note), plus unused chat rules that carried
rails and two glow animations that nothing referenced.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): line up hub pages, list the model templates, and stop clipped text

Medium-width pages inside Build and Operate were centred while the tab
bar above them was flush left, so the title started 60 px right of the
first tab. They now start at the bar's edge.

Add Model offered nine templates as a grid of identical cards with chip
clouds and inline styles. It is now one list of rows, each with the
field names it fills in on a single muted line.

Two clipped strings are fixed: the Studio voice field cut its
placeholder mid-word, and the phone job list ended the schedule line in
an ellipsis. The recommendation shelf also separates size and memory
with a dot, and the docs describe the shelf.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): keep the hidden Studio tab state inside its tab

The hidden state text added to each type tab was absolutely positioned
against the page, so on a phone it sat outside the scrolling tab row and
widened the page by hundreds of pixels. The tab is now the containing
block.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): cover the on-disk sizes on the Installed table, model page and cleanup sheet

The fixtures stub GET /api/models/storage. The default report is empty,
so existing specs keep the gallery estimates. makeStorage() builds a
report from files and the models that use them, the way the server
does, and storageSpec() is a models directory with shared and missing
files.

New specs cover the Size column and its shared line, the fallback when
the call fails or the user is not an admin, the files list on the model
page, a missing file, the bytes a removal frees with shared files, and
the cleanup findings. Node tests cover the storage helpers and the
batch arithmetic.

Assisted-by: Claude Code:claude-sonnet-5-5

* test(ui): wait for the page before pressing keys and ticking the clock

Two specs failed in loaded full runs and passed alone. The Alt+1 to
Alt+7 spec pressed a key before the composer had armed its key
handler. The capacity chart spec advanced the fake clock before the
poller had mounted, so it counted fewer readings than it expected.

Both now wait for the page to mount. The key spec retries a press that
lands during a re-render, and the clock spec advances in small steps and
polls for the row count.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): hide the Installed footer when the storage report is empty

An empty report from the storage call made the footer read "0.0 GB on
disk" next to sizes taken from the gallery estimate. An empty report
says nothing about the disk, so the footer now shows only the model
count. A spec covers it.

Assisted-by: Claude Code:claude-sonnet-5-5

* fix(ui): keep Explore pane actions inside the pane

The inspector actions sat in a non-wrapping flex row beside the title,
so the buttons ran past the pane edge once it got narrow. The row now
takes its own line and wraps.

The primary action (Install, Retry, Open) comes first. Manage
installation becomes a ghost button, and Open details moves to the end
of the row, so one action stands out and the others are quiet. No
action or test id is removed.

Add a spec that checks, in light and dark at several widths and with a
pane forced to 320 px, that every action stays inside the pane box and
that the pane keeps its inner padding.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* refactor(ui): drop the type chips from the Studio composer

The Studio tabs and the composer's type chips listed the same seven
modes, so the page said the same thing twice. Keep the tabs as the one
place to switch modes.

The composer now shows the type it will open as a small label in its
header. The type suggestion from the typed words stays as the quiet
hint line under the prompt, and Alt+1 to Alt+7 still pick a type. The
composer root carries data-type, data-types and data-missing so tests
can read the state.

Specs pick a type through a shared Alt+digit helper and read a
missing model from the tab dot instead of a chip. Remove the unused
chip locale strings and CSS.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* refactor(ui): remove the section crumb above page titles

Page headers drew a small uppercase crumb with a short rule before it
above the title. On the hub pages it repeated the hub name, so Build
sat above a heading that also said Build.

PageHeader now renders only the title, the supporting line and the
actions. Drop the eyebrow prop, the route-derived section name, its CSS
and the unused section helper, and remove the explicit eyebrow props
from the pages that passed one. Pages stay reachable through the
sidebar and the hub tab bar.

Add a spec that checks several pages show their title with nothing
ahead of it in the header.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* refactor(ui): remove left accent rails from tiles, rows and quotes

Several surfaces marked state with a coloured strip on the left edge.
Replace each one with a cue that is not a rail:

- Stat cards lose the strip; the icon and value still carry the colour.
- The highlighted card is a raised surface with a firmer edge.
- The selected rail row is an accent wash with a hairline outline.
- The status stripe on rail items is a small status dot.
- The active failover row is a tinted row.
- Quotes in markdown and chat prose are italic instead of barred.
- The variant detail panel has a full hairline border.

Add a spec that walks the main routes in light and dark and fails on a
left border thicker than 1px, a sideways inset shadow, a narrow
absolute strip in ::before or ::after, or a narrow tall child pinned to
a left edge.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-08 22:03:07 +02:00
mudler-agentandEttore Di Giacinto 6343a2dc6f feat(voice-detect): report the encoder family so audio-registered voices are fingerprinted (#12499)
* feat(voice-detect): report the encoder family so audio-registered voices are fingerprinted

A voice registered from audio through the voice-detect backend had no
encoder fingerprint, so the parakeet-cpp backend could not tell whether
it was comparable with the loaded speaker model and could only fall back
to the file-name rule.

The voice-detect backend now binds the three new libvoicedetect accessors
with a symbol probe (an older library still loads and reports nothing),
copies the borrowed strings at once and never frees them. It also hashes
the model file once at load. VoiceEmbedResponse gains two optional
fields, encoder_family and encoder_weights ("sha256:<hex>", empty when
the model is not a plain file).

/v1/voice/register stores them as encoder_family and a new
encoder_weights field in the registry entry; model keeps the encoder
name, so the 1:N identify filter by name is unchanged for old entries.
A voice with a family is sent to the backend whatever its file name, and
the backend decides by family. /v1/voice/identify compares the family
when both the stored voice and the probe have one. Old entries load
without the fields and stay unfingerprinted.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs(voice-detect): say how a voice with weights but no family is treated

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(voice-detect): mark the model file open as intended for gosec

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-08 19:26:11 +02:00
localai-org-maint-botandlocalai-org-maint-bot 0102751b31 fix(mcp): connect to legacy SSE servers (#12304)
* fix(mcp): connect to legacy SSE servers

Remote model MCP connections only use Streamable HTTP, so legacy SSE
servers fail initialization. Retry with SSE when the initial POST returns
400, 404, or 405.

Share one discovery timeout across attempts and cancel failed connections
without truncating successful sessions. Preserve HTTP policy and reject
foreign SSE message endpoints before attaching credentials.

Add SDK integration tests and document automatic transport selection.

Assisted-by: Codex:gpt-6
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(mcp): avoid narrowing HTTP status codes

Store the initialization status in a 64-bit atomic value to avoid the
integer overflow conversion reported by gosec.

Assisted-by: Codex:GPT-6 gosec
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-10-08 16:41:33 +02:00
4a1089be35 chore: ⬆️ Update 0xShug0/audio.cpp to cf124a67cc55d8f65a9a15eec69edbff0fb212c8 (#12421)
* ⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* fix(audio-cpp): mirror the turn detection task

The upstream enum adds TurnDetection, so the exhaustive conversion fails
to compile. Extend both conversions and preserve the RPC admission rules.
Test that turn detection cannot route through VAD or another existing RPC.

Assisted-by: Codex:gpt-6

---------

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-10-08 16:41:02 +02:00
localai-org-maint-botandlocalai-org-maint-bot f4a09b4b77 fix(vram): reject oversized GGUF metadata before allocation (#12560)
Upgrade gguf-parser-go to v0.26.3 so string lengths are checked against the remaining file size before allocation. Master AIO CI crashed in the background gallery warmer when v0.25.0 tried to allocate several terabytes; panic recovery cannot catch a fatal runtime OOM.

Check both overflow-sized and file-exceeding strings through the remote reader, and document the size-only estimate fallback.

Assisted-by: Codex:gpt-6

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-10-08 09:56:59 +02:00
mudler-agentandEttore Di Giacinto 895d50385f fix(distributed): converge model configs across frontends (#12558)
* fix(galleryop): announce model changes before the preload

After a gallery install or delete, the replica that ran it replaced its
config loader, then preloaded every installed model, and only then
published the models invalidation. The preload does remote lookups and
checksums for each model, so on a large models directory peers learned
about the change minutes after the originator listed it. When the
preload failed or the operation was cancelled, the event was never sent.

Publish the invalidation, and apply the delete lifecycle, as soon as
the loader holds the new set. The preload still runs afterwards with
its own error handling. Its failure is reported on the operation, but
it no longer rolls back a deletion that peers have already applied.

Assisted-by: Claude Code:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(distributed): resync model configs from the models directory

Frontends refresh their model configs only when a models invalidation
arrives on NATS. NATS keeps no history, so a frontend that is
disconnected when the message is published never applies the change.
It keeps serving the old config, for example an alias that points at
the previous model, until some later change happens to touch it.

Each frontend now reruns the peer reconcile against the shared models
directory after every NATS reconnect, and every
--model-config-resync-interval (default 30s) when a config file
changed. The pass names no model, so only models whose file changed
get a revision transition, and an unchanged directory costs one read
of the config files.

The reconcile replaced the whole loader with a parse of the models
directory, which dropped models loaded with --config-file and
published a deletion revision for them. Configs defined outside the
directory are now kept, both there and after a gallery install.

Assisted-by: Claude Code:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(nodes): stop stale frontends from retargeting alias rules

A scheduling rule keyed by an alias derives its target from the alias
mapping of the frontend that reads it. Every frontend keeps its own
copy of the model configs, so after an alias is repointed a frontend
that has not reloaded it still resolves the old target. Two frontends
then rewrote the rule's stored target_model against each other on
alternate reconciler ticks, and the outdated one scaled up the model
the alias used to point at.

The registry already records the accepted config revision of each
model. A frontend now derives a rule's target from its own alias
mapping only when its config revision for the rule's name matches
that record. Otherwise it keeps the stored target_model: it neither
writes the column nor reconciles replicas of the old target. The check
reads the database only for a rule whose stored and derived targets
differ. With no accepted revision on record, the old behaviour stays.

Assisted-by: Claude Code:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-08 09:56:12 +02:00
Plamen K. Kosseff c03acb0647 feat(models): report per-model disk usage in API and WebUI (#12551)
New admin endpoint GET /api/models/storage reports disk usage of the
installed models: each model's files and sizes, files shared between
models counted once, and configured files that are missing from disk.
The models WebUI page shows a usage summary, each model's file table,
and the full file list with per-file status.

Assisted-by: Claude Code:claude-fable-5

Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>
2026-10-07 22:03:48 +02:00
mudler-agent cd604b5c5b fix(distributed): bound, stop and cancel model loads with leases and worker operations (#12524)
Fence model load jobs by generation, lease them on the database clock, bound the work on the worker with operations and a process-group watchdog, and add one stop path with a load-cancel API. See the pull request for the design, the rolling upgrade notes and the test evidence.

Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-07 16:59:56 +02:00
localai-org-maint-botandlocalai-org-maint-bot c525ad16e8 fix(audio): record transcription usage and traces (#12546)
* fix(audio): record transcription usage and traces

Count successful transcription requests by model, including streams.
Keep token counts at zero because transcription exposes no token usage.

Capture multipart API trace metadata without reading uploaded audio.
Do not record failed transcription or client writes as successful usage.

Assisted-by: Codex:gpt-6-astra

* fix(audio): preserve aliases in streaming usage

Streaming transcription records the resolved target as its usage model.
Pass the requested name so JSON and SSE requests share the alias bucket.

Assisted-by: Codex:gpt-6-astra

* test(http): check multipart reader close errors

Assert successful reader cleanup to satisfy the errcheck CI gate.

Assisted-by: Codex:gpt-6

---------

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-10-07 16:57:23 +02:00
localai-org-maint-botandEttore Di Giacinto 6b794651a4 feat(diarization): return sound events with include_sounds (#12544)
* feat(diarization): return sound events with include_sounds

A client that wants text, speakers, voice prints and sound events had to
make a diarization call and a separate sound call. Add an include_sounds
request field to /v1/audio/diarization that adds a sounds array of closed
events {start, end, label, confidence}, in seconds.

The parakeet-cpp backend runs a tagger-only scene stream over the clip,
the same stream and thresholds the live path uses, so a clip gives the
same events offline and live. A model with no sound_model companion, or a
backend that does not report sound events, fails with 501 and the stable
code include_sounds_unsupported instead of an empty list. The proto
carries sounds_included so an empty list still means "nothing heard".

The localai-proxy backend forwards the field. Swagger, docs and the
e2e mock backend are updated.

Assisted-by: Claude:claude-sonnet-5-5 [protoc swag go]

* feat(gallery): add parakeet-cpp-multilingual-diarization-speakers-sounds

Same as parakeet-cpp-multilingual-diarization-speakers (TDT 0.6B v3,
Nemotron-3-Diarization, WeSpeaker) plus a CED-Tiny sound_model, so one
model name serves /v1/audio/diarization with include_text,
include_speaker_profiles and include_sounds. It declares the
sound_classification usecase like the realtime scene entries.

Assisted-by: Claude:claude-sonnet-5-5

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-07 14:10:45 +02:00
localai-org-maint-botandEttore Di Giacinto 217038d4fe fix(agents): share the collection list across frontends (#12536)
With the PostgreSQL vector engine, each frontend kept the collections it
had opened in an in-memory map and answered "collection not found" for
any other. A collection created through one frontend was unknown to the
others until they restarted, and each frontend listed a different set.

Wrap the in-process collections backend in distributed mode so that the
database is the source of truth:

- lists come from the registry,
- a lookup miss checks the registry before it returns 404, and opens a
  collection that exists there (once per name, even under concurrency),
- a cached collection that left the registry is dropped and closed, with
  a re-check at most every 5 seconds,
- create and reset publish an event on the existing collection
  invalidation subject, so other frontends re-check at once.

Without the postgres engine, or outside distributed mode, nothing changes.

Assisted-by: Claude:claude-sonnet-5-5 go

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-06 22:45:23 +02:00
mudler-agentandEttore Di Giacinto 690f3994b0 feat(auth): let users pause and resume their API keys (#12521)
A key can now be paused until the owner resumes it, or until a given
time. A paused key is rejected by validation before last_used is
updated, and a pause time that has passed lifts the pause by itself.
Existing keys stay active.

PATCH /api/auth/api-keys/:id takes {"disabled": bool, "paused_until":
RFC 3339 string or null}. Only the key owner can change it, and a
paused_until in the past is rejected. The key list returns the pause
fields. The Account page gets a Pause and Resume button for each key
and a Paused badge that shows the resume time.

Assisted-by: Claude Code:claude-sonnet-5-5

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-06 18:25:21 +02:00
Ettore Di Giacinto 7467b51281 test(gallery): expect speaker_recognition on the bundles with a voice component
The parakeet-cpp backend now answers VoiceEmbed and VoiceVerify, so the
pinning test asserts the usecase on the entries that declare it instead of
its absence. Document voice_verify_threshold.

Assisted-by: Claude Code:claude-sonnet-5-5
2026-10-05 23:48:20 +00:00
Ettore Di Giacinto e9af44b90b docs(voice): document a parakeet-cpp bundle as the embedding model
Describe the bundle as an embedding model for /v1/voice/*, the realtime
voice_recognition stage, and the limits: identify and plain verify only,
a 256-dimension space shared with voice-detect-wespeaker-resnet34, and the
libparakeet.so symbol it needs.

Assisted-by: Claude Code:claude-sonnet-5-5
2026-10-05 23:46:14 +00:00
Ettore Di Giacinto 04df10a082 docs(voice): document speaker naming through a bundle component
Cover speaker_component: and speaker_tag: in the audio-to-text option
tables and bundle section, add a section to the voice recognition page
that explains which registered voices a bundle component receives, and
update the speaker_name condition in the realtime and diarization pages.

Assisted-by: Claude Code:claude-sonnet-5-5
2026-10-05 23:45:14 +00:00
Ettore Di Giacinto a2c498359b fix(voice): keep one voice store per embedding dimension
The in-memory local-store rejects vectors of another size than the ones
it holds. With 192-value voices registered, adding a 256-value voice
failed, and identifying with a 256-value probe returned an error.

The registry now keeps one store per embedding dimension. The first
dimension seen uses the configured store name, so single-encoder
instances are unchanged. Later dimensions use "<name>-<dim>". Identify
searches only the store of the probe size and returns no match when no
voice of that size exists. Forget finds the store from the stored
embedding. A name may hold one voice per encoder.

Assisted-by: Claude Code:claude-sonnet-5-5 golangci-lint
2026-10-05 23:42:49 +00:00
localai-org-maint-botandEttore Di Giacinto 68c980f3cd chore(deps): bump localrecall to v0.6.6 (#12508)
* chore(deps): bump localrecall to v0.6.6

LocalRecall v0.6.6 closes the PostgreSQL connection pool when a
collection fails to open. Before, each failed collection create left
its pool open. With a failing embedding model, every retry leaked one
more pool until PostgreSQL refused new clients.

LocalAI gets LocalRecall through LocalAGI, so this raises the indirect
requirement directly instead of waiting for a LocalAGI bump.

The release also limits each collection pool to 4 connections by
default. POSTGRES_POOL_MAX_CONNS changes the limit.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]

* docs(agents): document the PostgreSQL pool size limit

LocalRecall v0.6.6 limits the connection pool of each collection to 4
connections and reads POSTGRES_POOL_MAX_CONNS to change it. Document
the variable next to the other PostgreSQL settings of the embedded
store.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-06 00:51:53 +02:00
mudler-agentandEttore Di Giacinto 79a7631cc5 feat(parakeet-cpp): encoder fingerprint for speaker naming, VAD trim and word filter options, pin bump (#12491)
* chore(parakeet-cpp): bump parakeet.cpp to 2de154c

Brings in the speaker registry encoder fingerprint, the VAD segment trim
and the opt-in word filter, a fix for a per-call thread count that stayed
set on the process-wide backend after a Silero VAD pass, and bundle
components loaded from memory.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(parakeet-cpp): encoder fingerprint for speaker naming, vad_trim and guard_* options

Speaker naming. A registered voice now carries the encoder that made it:
the embedding family (voicedetect:<arch>:<name>:<dim>) and the sha256 of
the weights. The backend reports the family of the loaded speaker model in
Status, voice enrollment from speaker_profiles stores it as encoder_family
(old entries load without it), and the registry sent to parakeet.cpp is
built with parakeet_capi_speaker_registry_add_embedding_fp. The library
then refuses a registry of another encoder family and the error names both
families; another quantization of the same family only warns. A voice with
only a weights hash gets the loaded family when the hashes are equal.

Voices without a fingerprint (registered from audio: libvoicedetect cannot
report one) keep the file-name rule and are used with a warning. The
library cannot mix them with fingerprinted voices in one registry, so a
request that has any uses the old registry for all. speaker_strict:true
drops them instead. A library without the symbols behaves as before.

Transcription. vad_trim (seconds, 0 keeps the whole cuts) goes through the
VAD options JSON, so it reaches /v1/vad and the segmenter. The guard_*
options guard_min_local_conf, guard_local_radius and guard_drop_punct_only
turn on the word filter through parakeet_capi_transcribe_path_json_with,
or through the segmenter with vad:true. They are off by default, bad
values fail the load, and a library without the symbol fails it with a
clear message. The dropped word count is logged at debug level.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-05 15:57:16 +02:00
localai-org-maint-botandlocalai-org-maint-bot 3d586cc3c5 fix(audio-cpp): forward voice reference transcripts (#11997)
Saved voices send ref_text, but Fish Audio requires reference_text.
Derive the canonical parameter while preserving explicit overrides.
Both TTS modes use the shared builder.

Add regression cases and document the parameter alias.

Assisted-by: Codex:gpt-6

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-10-05 11:05:43 +02:00
Stefan Walcz 5f5feea15a fix(agentpool): keep the conversation in the agent web chat (#12410)
* feat(agents): keep the web chat history between messages

The agent chat endpoint ran every message as a fresh job
(ag.Ask(WithText(message))), so a follow-up such as "add two days to
item 3 and recalculate" never saw the answer it referred to. Agents
then rebuilt their reply from scratch instead of changing it.

Use the agent's own conversation tracker, as the Telegram and Slack
connectors already do: send the earlier turns with the new message and
record successful answers. Failed, cancelled or empty runs are not
recorded, and the tracker drops a conversation after the agent's
last_message_duration of inactivity. The distributed (NATS) chat path is
unchanged.

Assisted-by: Claude:claude-opus-5-5 ginkgo
Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>

* fix(agents): keep each web chat conversation's history separate

Review on this PR: the first version kept the history in the agent's
conversation tracker under one fixed key per agent. The web UI keeps
several conversations per agent (New Chat, switching, Clear) and did not
send which one a message belongs to, so the hidden history mixed them:
a new chat received the previous chat's turns, and Clear only cleared
the screen.

The client now sends the earlier turns of the conversation it is
showing as `history` with POST /api/agents/:name/chat, and the server
keeps no web chat history of its own. Each conversation only ever sees
its own turns; New Chat and Clear start without history. The server
uses only user and assistant turns with text, bounded to the most
recent 40 turns and 64,000 characters. `history` is optional, so
existing callers keep the previous behaviour (no history); the
distributed (NATS) path does not forward it yet.

Tests: two conversations of one agent stay apart, an empty history
(New Chat or Clear) starts fresh, system/tool/empty turns are dropped,
the bounds keep the most recent turns; the UI helper that builds the
history from the visible messages has node --test coverage. Docs:
features/agents.md describes `history` and the web UI behaviour.

Assisted-by: Claude:claude-opus-5-5 ginkgo
Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>

---------

Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
2026-10-05 01:38:29 +02:00
leilei3167andlocalai-org-maint-bot d90f703bd9 fix(xsysinfo): read process VRAM without proc children (#12482)
* fix(xsysinfo): read process VRAM without proc children

Kernels without CONFIG_PROC_CHILDREN have no task children file, so
ProcessVRAM dropped every DRM reading. When that file is missing, walk
child processes from /proc/<pid>/stat ppid links instead. Other read
errors still drop the reading.

Fixes #12481

Signed-off-by: leilei3167 <imleilei123@gmail.com>

* docs(system): describe the proc children fallback

Document VRAM reporting on kernels without CONFIG_PROC_CHILDREN.

Assisted-by: Codex:GPT-6

---------

Signed-off-by: leilei3167 <imleilei123@gmail.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-10-05 01:37:12 +02:00
Plamen K. Kosseff 104c2fb440 fix(oci): stage image downloads in a configured directory, not in /tmp (#12280)
* fix(oci): stage image downloads in a configured directory, not in /tmp

Problem:
- the image tar and its compressed layers staged in os.TempDir()
- /tmp is commonly a RAM-backed tmpfs far smaller than a backend image
- big installs failed with a full /tmp or silently ate RAM

Change:
- staging path resolved like the other storage paths
  (LOCALAI_DOWNLOAD_STAGING_PATH, default ${basepath}/downloading),
  threaded to the extractor as a download option
- every download works in its own subdirectory holding its tar and
  layers: nothing shared between concurrent downloads, one removal
  cleans a download up, a crash leaves one self-contained orphan
- callers that pass no staging directory keep the OS temp behavior

Assisted-by: Claude:claude-fable-5
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>

* fix(oci): handle staging cleanup errors

Log staging cleanup failures to satisfy errcheck. Test extraction with an
unavailable OS temp directory and verify that staging is removed.

Assisted-by: Codex:GPT-6
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>

---------

Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>
2026-10-05 01:27:45 +02:00
nexxtmobile.deandnexxtmobile.de 4d681b7f6d fix(realtime): detach VAD-commit transcription from barge-in cancellation (#12446)
* fix(realtime): keep VAD-commit transcription alive across barge-in, cancel it at teardown, order commits

Barge-in (new speech onset) cancels the turn's SourceVAD response context
(realtime_turncoord.go: respSink.cancel(SourceVAD)). The VAD commit body
runs under that same context, so an in-flight Whisper STT call was aborted
with 'context canceled' whenever the caller kept talking while the first
chunk was transcribing. The user's turn was lost: no transcript, no
LLM/TTS response.

v1 of this fix ran the transcription with context.WithoutCancel(ctx). The
review correctly pointed out two correctness gaps:

1. Teardown lost its cancellation. WithoutCancel detaches from every
   cancellation, so a transcription in flight at session close outlived
   the session and blocked respSink.shutdown (which joins the response
   goroutines) until the backend finished the job.
2. Out-of-order commits. Consecutive commits run in parallel goroutines,
   so a fast second transcription could append its user item before a
   slow first one: the conversation became [second, first] and the second
   response saw only [second].

Changes (core/http/endpoints/openai/):
- Session gains a session-lifetime context (sessionCtx), cancelled by
  conncoord's Teardown BEFORE respSink.shutdown joins the response
  goroutines. The transcription (and the voice-gate resolution) run under
  it: they survive barge-in (which cancels only the per-response context)
  but are cancelled with the session.
- Commit slots order the user-item appends in speech order:
  Session.nextCommitSlot() is claimed at commit issue time (VAD CommitTurn
  / client commit), a commit's item append waits on the previous slot's
  done (aborts on the session context), and every exit closes the slot so
  a failed or torn-down commit never blocks the next. Transcriptions stay
  parallel; only the appends are ordered.
- If the turn's response context was cancelled while the (detached)
  transcription ran — barge-in, superseded by a newer commit — the user
  item still commits (appendUserItem, split out of generateResponse) so
  the LLM context keeps the full user input, but no response is generated
  for the superseded turn; the newer speech triggers its own response on
  the complete history.
- Regression tests (realtime_commit_order_test.go) cover both review
  schedules — teardown during an in-flight transcription, and
  held-first/finished-second out-of-order completion — plus the
  barge-in-during-transcription item survival, driving the real commit
  path with a transcription double that honours context cancellation.
- docs/design/realtime-state-machines.md: implementation-status entry for
  the committed-turn pipeline (transcription lifetime + commit order).

Fixes #12445

Validated: builds, go vet clean, all openai specs + respcoord/turncoord/
conncoord suites pass under -race (incl. the 3 new regression specs).
Production A/B (call-center voice agent, SIP, silero-vad +
whisper-large-turbo + LLM + TTS, server_vad ~600 ms) on LocalAI v4.11.0:
unpatched — 'transcription_failed: context canceled', first part of the
utterance lost, agent answers only the remainder; patched — full
transcript committed, agent answers the complete utterance, barge-in
still cancels the in-flight assistant TTS response as intended, and
teardown cancels the in-flight transcription instead of waiting for the
backend.

Signed-off-by: nexxtmobile.de <kai@nexxtmobile.de>

* fix(realtime): release commit slots in order on every exit; share slot+issue boundary

Follow-up to the review of c8f991be (issue #12445): two schedules still
broke commit ordering.

1. A failed middle commit released later turns before earlier turns
   finished. slot.done closed on every return, but the wait on
   slot.prevDone happened only on the successful nonempty-transcript
   path. Hold transcription A, let B fail (or return an empty
   transcript / be rejected by the voice gate), then complete C: B
   closed its channel without waiting for A, so C appended and started
   its response without A (response history [third], final history
   [third first]). Fix: the slot now releases (done closes) only AFTER
   the predecessor has finished — on EVERY exit path, including errors,
   empty transcripts, gate rejections and teardown (the session context
   can still stop the wait, so teardown never blocks on a
   never-finishing predecessor). The success path keeps its append gate
   (wait before appending the user item); the deferred release gate
   enforces the same order on every other exit.

2. Slot order and response issue order could disagree between the two
   producers. The VAD CommitTurn and the client
   input_audio_buffer.commit reserved the slot and called
   respSink.issue separately; a pause between the two let the other
   producer reserve AND issue first, so the later issue superseded the
   EARLIER turn's response (response history [first], final history
   [first second], second turn un-answered). Fix: both producers now go
   through Session.issueCommit, which claims the slot and issues the
   body under one lock (commitOrderMu) — slot order == issue order.
   respSink.issue is non-blocking, so the lock never stalls
   VAD/barge-in handling.

Regression tests (realtime_commit_order_test.go) now drive the REAL
issue path — Session.issueCommit into the real responseSink/respcoord,
so coordinator supersession and the spawned response goroutines are
exercised — and cover: teardown during an in-flight transcription;
held-first/finished-second out-of-order completion; barge-in
(respSink.cancel) item survival; a FAILED middle commit; an EMPTY
middle commit; interleaved VAD/client producers in both directions.
The failed/empty middle specs fail deterministically without the
release gate (verified against the pre-fix code).

docs/design/realtime-state-machines.md: implementation-status entry
updated (append gate + release gate + shared issue boundary).

Validated: builds, go vet clean, all openai specs + respcoord/turncoord/
conncoord suites pass under -race (476 specs, incl. the 4 new ones).

Fixes #12445

Signed-off-by: nexxtmobile.de <kai@nexxtmobile.de>

---------

Signed-off-by: nexxtmobile.de <kai@nexxtmobile.de>
Co-authored-by: nexxtmobile.de <kai@nexxtmobile.de>
2026-10-05 01:27:40 +02:00
mudler-agentandEttore Di Giacinto 3185b6fcf8 feat(parakeet-cpp): load bundle GGUF files, bundle gallery entries, pin bump (#12479)
* feat(parakeet-cpp): load bundle GGUF files and use their components by role

A bundle GGUF holds several models (ASR, VAD, diarization, sound events,
speaker encoder) in one file, each with its own licence. Detect a bundle
at load through parakeet_capi_bundle_components_json and open components
with parakeet_capi_load_component. The three symbols are probed together,
so an older libparakeet.so still loads plain files as before.

The only ASR component is the primary model; bundle_asr:<name> picks one
when there are several. A Silero VAD component of the primary bundle is
loaded without an option and serves /v1/vad and vad:true. The diar, ced
and voice components load on request: diar_component, sound_component and
speaker_component, or a companion option (diarization_model, sound_model,
speaker_model, vad_model) that names a bundle, even the model file itself.
vad_component picks a VAD component and implies vad:true.

A role the bundle cannot fill fails the load with the component list, and
a diarization or sound request on a model without that role names the
bundle components. Every existing option and single-file model behaves as
before.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* chore(parakeet-cpp): bump parakeet.cpp to 781a973

Brings in the bundle GGUF format and its C-API (parakeet_capi_load_component,
parakeet_capi_bundle_components_json, parakeet_capi_load_error).

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(parakeet-cpp): gallery entries for the bundle GGUF files, docs

Add parakeet-cpp-bundle-small (338 MB: Parakeet TDT+CTC 110M, Nemotron-3-
Diarization, CED-Small, WeSpeaker ResNet34-LM, Silero VAD), -standard
(1.1 GB, Parakeet TDT 0.6B v3 instead of the 110M model) and
-moondream-redux (215 MB: packed Redux and Silero VAD, CPU only). One
install serves transcription, VAD, diarization, sound events and speaker
naming through the component options. The existing single-purpose entries
stay.

A bundle has no single licence, so the entries use license: other and
state the licence and credit of each component in the description, with
the upstream inconsistency of the CED licence. The docs get a section on
bundles in audio-to-text with the entries, the roles, the options and the
licence notice, and pointers from the VAD, diarization and sound
classification pages. A gallery test checks the file names, checksums,
usecases and options.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-04 20:33:13 +02:00
mudler-agentandEttore Di Giacinto d3724eef50 fix(nodes): stage dedicated diarization audio (#12474)
Dedicated diarization forwards frontend audio paths to remote workers,
which cannot read those temporary files. Stage the input before the RPC
and release it afterward, following the transcription lifecycle.

Clone the request so staging does not change caller-owned data. Cover
input bytes, request fields, cleanup, and error propagation in tests.
Document distributed diarization staging on the existing feature page.

Assisted-by: nib:gpt-6-astra

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-04 13:14:00 +02:00
mudler-agentandEttore Di Giacinto ed4a3975be feat(parakeet-cpp): gallery entries for the VAD-only Moondream slices, pin bump (#12469)
* feat(parakeet-cpp): add gallery entries for the VAD-only Moondream slices

Add parakeet-cpp-vad-moondream-redux and parakeet-cpp-vad-moondream-ultra.
They install the VAD head of Moondream Redux and Ultra (Q8_0) as small
files of 10 MB and 6 MB, cut out of the full models without retraining,
for the VAD endpoint. The files cannot transcribe, and a transcription
request fails with a clear error.

The files load only with a parakeet.cpp build that has VAD-only GGUF
support (parakeet.cpp pull request 87). The backend pin must move to a
commit that includes it before these entries work in a released image.
The parakeet-cpp-vad entry keeps installing Silero.

The docs list the files with the size, load time and memory compared
with loading a whole model. A gallery test checks the usecase, the file
name and the checksum of each entry.

Assisted-by: Claude Code:claude-sonnet-5-5 [golangci-lint]

* chore(parakeet-cpp): bump parakeet.cpp to e53a253

Brings in the VAD-only GGUF loader.

Assisted-by: Claude Code:claude-sonnet-5-5 [git] [gh]

* docs(gallery): link the parakeet.cpp VAD docs instead of the merged PR

Assisted-by: Claude Code:claude-sonnet-5-5 [git]

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-04 09:34:40 +02:00
mudler-agentandEttore Di Giacinto 99043b442c feat(router): route with native decision models (#12449)
* fix(schema): preserve SystemOne image inputs

Assisted-by: OpenAI

* test(schema): follow Ginkgo conventions for decision inputs

Assisted-by: OpenAI

* feat(llama-cpp): dispatch native decisions through Score

Upgrade the stock dependency and reconcile Score/TTS patches. Reuse native decision parsing, tasks, formatting and response-reader cleanup; preserve ordinary scoring admission and guard older dependencies.

Assisted-by: OpenAI

* refactor(systemone): share request and model validation

Assisted-by: OpenAI:gpt-5

* fix(systemone): preserve HTTP wire-byte validation limit

Keep structural validation separate from the serialized internal request bound so HTML escaping cannot reject valid HTTP payloads.

Assisted-by: OpenAI:gpt-5

* feat(systemone): bound images and account native decisions

Preserve public wire limits independently from router serialization. Reject unsupported NER images, map native request/capability errors, and stamp explicit usage once. Advertise decisions for stock llama-cpp.

Assisted-by: OpenAI

* fix(systemone): record usage on registered native route

Exercise real registration and billing with a mock native backend. Reject empty native responses, malformed image URLs, trailing JSON, and wire overflow including whitespace.

Assisted-by: OpenAI

* feat(router): add lazy native decision transport

Bind named models through internal ModelSystemOne calls with shared validation and bounded abandoned operations. Remove request and echoed-error contents from decision traces.

Assisted-by: OpenAI:gpt-5

* feat(router): classify overlapping policies with native decisions

Ask independent noul questions, validate probabilities and preserve first-superset routing. Wire the central factory with config-sensitive invalidation and cancellation-safe resolution. Document native framing and bounded operation limits.

Assisted-by: OpenAI:gpt-5

* feat(gallery): add pinned Julia-1 native decision model

Add a separate text-only llama-cpp Q8 entry with pinned Apache-2.0 source provenance and checksum. Installed using the gallery installer and exercised choice, score and noul on CPU.

Assisted-by: OpenAI

* test(router): verify native decisions through central factory

Add an opt-in real-model Ginkgo integration covering the native Go loader and C++ transport, token usage, independent overlapping labels, and candidate selection. Document owned-server execution and the intentionally non-quality threshold.

Assisted-by: Codex:gpt-5

* fix(llama-cpp): align upstream pin and preserve decision signatures

Advance to bed0a856 without losing the automated upstream bump. Detect full-request fill_task support at compile time and forward every question for Nimble framing while retaining the earlier native signature. Preserve reconciled SCORE/TTS patches; add standalone compatibility coverage.

Assisted-by: Codex:gpt-5

* feat(gallery): add native decision family defaults

Pin Laya, Kev-4B, lev, OpenJev and Nimble artifacts. Verify Laya/Kev/lev gallery installs and CPU contracts on both native pins; clearly mark OpenJev/Nimble runtime validation pending and their noncommercial licenses.

Assisted-by: OpenAI

* docs(decisions): clarify integrated Nimble prerequisite

Record the exact combined backend pin while retaining pending OpenJev and Nimble installation/runtime validation status.

Assisted-by: Codex:gpt-5

* fix(gallery): indent native decision model sequences

Match repository yamllint indentation for Laya, Kev, lev and OpenJev list fields. Parsed gallery data is unchanged; reproduce CI gallery lint failure before the whitespace-only fix and pass the same command afterward.

Assisted-by: Codex:gpt-5

* docs(decisions): record OpenJev and Nimble CPU validation

Record gallery installation, checksum/metadata verification and multiquestion native smoke results on bed0a856. Retain noncommercial and text-only limitations without accuracy or deterministic-output claims.

Assisted-by: OpenAI

* fix(ui): expose native Decisions router classifiers

Select classifier models using metadata-driven capability routing, retain tuned thresholds, and validate native decision selections before saving. Cover both native backends and create/save/reopen in the real React editor.

Assisted-by: Codex:gpt-5

* fix(router): exclude aliases from native decision discovery

Check the originally named config before advertising native Decisions eligibility. Retain target capability inheritance for ordinary generation aliases. Exercise the actual capabilities endpoint with native models on both backends, aliases, and disabled models.

Assisted-by: Codex:gpt-5

* feat(systemone): share bounded multimodal input validation

Preserve text wire limits while admitting bounded PNG/JPEG decision input. Share collection and header validation across internal and public callers and keep the native runner response budget independent.

Assisted-by: OpenAI:API-assistant

* fix(systemone): bound admission lifetimes and validate complete images

Retain shared admission leases through actual work completion, including abandoned internal operations. Decode bounded image pixels, cap public native responses before usage stamping, and preserve oversized malformed text status precedence.

Assisted-by: OpenAI:API-assistant

* fix(router): classify images before media fetching

Preserve ordered structured probes for native decisions. Defer OpenAI
media preparation until routing selects the served model, so rejected
decision URLs cannot trigger downloads before shared validation.

Guard direct image collection with context-aware shared admission.
Keep text classifiers and embedding caches from discarding image input.
Retain fail-closed classifier configuration and runtime fallback policy.

Add middleware, typed-content, admission, cancellation and cache tests.

Assisted-by: OpenAI:API-assistant

* fix(router): bound extraction before serialization

Check probe budgets before copying text or marshaling message state.
Count JSON escaping so oversized internal inputs fail before allocation.

Preserve typed Anthropic blocks through selected-model conversion and
fallback. Keep retry coverage in Ginkgo without global test registration.

Assisted-by: OpenAI

* fix(router): bound supported probe serialization

Arbitrary structs can bypass the probe budget through pointer marshalers,
string tags, and promoted fields. Accept concrete chat schema types and
plain JSON values instead of emulating arbitrary struct serialization.

Budget escaped direct prompts before marshaling so raw length cannot hide
serialized expansion. Preserve runtime fallback and reject oversized
input before invoking the decision runner.

Add Ginkgo allocation, boundary, and marshaler invocation regressions.
Six-package tests, three-package race tests, and full-T2 delta lint pass.

Assisted-by: OpenAI:GPT-5 golangci-lint

* feat(decisions): enable bounded OpenJev images

Validate native decision images before permissive media parsing and pixel
allocation. Require both decision image support and a vision projector;
missing or audio-only projectors cannot silently become text decisions.

Pin the OpenJev Q8 projector and document its license and disk footprint.
Add native safety tests, canonical limit parity, gallery and load-option
checks, and a reproducible CPU direct-RPC contrasting-image smoke.

Assisted-by: OpenAI:GPT-5

* fix(decisions): reject incomplete image streams

stb accepts corrupt PNG Adler checksums and truncated JPEG scans.
Use bounded zlib validation and strict libjpeg decoding before parsing.
Keep dimension and aggregate pixel checks ahead of decoder allocations.

Wire decoder dependencies into native builds and runtime packaging.
Add regressions for appended EOI and embedded marker bypasses.

Assisted-by: OpenAI:GPT-5

* fix(ci): gate native decision image validation

Run the decoder security tests outside the stdlib-only native suite.
Fetch vendor headers at the backend pin and provision decoder dependencies.
Gate Go limit parity and production CMake wiring without model downloads.

Assisted-by: OpenAI:GPT-5

* test(decisions): cover multimodal public API paths

Exercise shared image contracts through the registered HTTP routes and
external mock backend. Add opt-in cached gallery installation and real
OpenJev image decisions through SystemOne and both routing APIs.

Assisted-by: Codex:gpt-5

* test(decisions): assert isolation and cache bypass

Observe external RPC calls and compare complete classifier history.
Winner-only and cache-miss checks could hide dropped history or cache use.

Give real inference its own application and model directory so shared
backend mappings and loaded processes cannot affect mixed suite order.

Assisted-by: OpenAI:ChatGPT

* test(decisions): isolate fixture globals

Disable optional global services in the isolated HTTP fixture and register
cleanup before setup assertions. Verify meter provider identity survives
fixture creation and destruction.

Snapshot observed usage before assertions so failures cannot retain the
mutex. Require a successful usage stamp before checking error responses.

Assisted-by: Codex:gpt-5 golangci-lint

* fix(application): honor optional telemetry controls

Skip failover gauge registration when metrics are disabled. Register
against the application meter rather than looking up the global provider.

Allow embedders to retain the bounded routing log without billing stats.
Keep the existing default when stats are disabled. The isolated HTTP
fixture uses this option without losing its native router assertions.

Assisted-by: Codex:gpt-5 golangci-lint

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-04 09:34:21 +02:00
localai-org-maint-botandmudler f035746db9 docs: ⬆️ update docs version mudler/LocalAI (#12458)
⬆️ Update docs version mudler/LocalAI

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-04 08:48:14 +02:00
mudler-agentandEttore Di Giacinto 6c718fdda7 feat(parakeet-cpp): VAD (Silero and the Moondream head), vad_model option and gallery entries (#12463)
* feat(parakeet-cpp): implement the VAD call and add the vad_model option

The backend now serves the VAD gRPC call (POST /vad and /v1/vad) with the
standalone VAD of libparakeet. It accepts a Silero VAD GGUF as the model
file, or an ASR model with a VAD head (Moondream Ultra and Redux). The
request audio is float32 PCM at 16 kHz; the response lists the speech
segments in seconds, like the silero-vad backend. A model with neither
fails the request with the library message.

The vad_threshold, vad_min_pause, vad_min_speech, vad_speech_pad and
vad_max_segment options tune the segmenter. Unset values keep the
defaults of the detector in use, and a bad value fails the load.

The vad_model option names a Silero GGUF, resolved against the models
directory like the other companion files. It lets any ASR model cut long
audio at pauses through parakeet_capi_transcribe_path_json_vad_with, and
it implies vad. vad:true alone still uses the model's own head.

The new symbols are probed like the existing optional ones. A library
without them still loads; the feature that needs one fails with a clear
message only when it is used.

Assisted-by: Claude:claude-sonnet-5-5 [go test]

* feat(gallery): add parakeet-cpp VAD entries and a v3 plus Silero example

Add VAD-only entries for the parakeet-cpp backend: the VAD heads of
Moondream Redux (packed, CPU) and Ultra (Q8_0), which share their files
with the existing ASR entries, and Silero VAD v6.2.3 as a GGUF (MIT,
Silero Team). The parakeet-cpp-vad entry installs Silero; it has no variants,
because variant ranking prefers the larger build that fits and these are
different detectors.

Add parakeet-cpp-tdt-0.6b-v3-silero-vad, a v3 entry that sets vad_model
so long audio is cut at pauses by Silero.

The Silero GGUF entries point at the intended Hugging Face URL of the
file; the existing silero-vad entries are unchanged. A test checks the
usecases, the shared files and the default entry and the vad_model reference.

Assisted-by: Claude:claude-sonnet-5-5 [go test]

* docs: describe parakeet-cpp VAD and the vad_model option

Document the VAD endpoint on the parakeet-cpp backend (Silero GGUF and
the VAD heads of Moondream Ultra and Redux), the vad_* tuning options,
and the vad_model option that lets an ASR model without a VAD head cut
long audio with Silero.

Assisted-by: Claude:claude-sonnet-5-5

* chore(parakeet-cpp): bump parakeet.cpp to 6165e3d

Pin the release that adds the standalone VAD (Ultra/Redux head and
Silero) and the C API calls the backend now uses.

Assisted-by: Claude:claude-sonnet-5-5

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-04 01:02:09 +02:00
mudler-agentandEttore Di Giacinto 00faf4ec26 feat(parakeet-cpp): Moondream Ultra and Redux gallery entries, VAD option, pin bump (#12455)
* chore(parakeet-cpp): bump parakeet.cpp to 11c1a0f

Picks up support for the Moondream Ultra and Redux models and the
VAD-segmented transcription entry point (C-API ABI is still 10).

Assisted-by: Claude:claude-sonnet-5-5

* feat(parakeet-cpp): add vad option for long-audio transcription

Models with a VAD head (Moondream Ultra and Redux) can cut long audio at
pauses. Setting vad:true in the model options routes offline
transcription through parakeet_capi_transcribe_path_json_vad. The symbol
is probed at startup like the other optional entry points, and vad:true
fails the load with a clear message when the library lacks it. A model
without a VAD head fails the request with the library's own message.
The option is off by default and does not affect streaming.

Also say in the load error that a packed ternary Redux model is CPU only,
because the library reports its refusal on a GPU backend through its log,
not through the C API.

Assisted-by: Claude:claude-sonnet-5-5

* feat(gallery): add Moondream Ultra and Redux for parakeet-cpp

Add five entries from the public parakeet-cpp GGUF repository: Ultra in
F16 and Q8_0, and Redux as packed ternary (CPU only, offline only) and
as dequantized F16 and Q8_0 (any backend). The entries enable vad:true so
long audio is cut at pauses. Checksums come from the repository's LFS
metadata. The weights are CC-BY-4.0.

Assisted-by: Claude:claude-sonnet-5-5

* docs(parakeet-cpp): document Moondream Ultra, Redux and the vad option

Assisted-by: Claude:claude-sonnet-5-5

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-03 22:08:12 +02:00
372b1f8983 feat(silero-vad): allow threshold/silence/pad via model options (#12430)
* feat(silero-vad): allow threshold/silence/pad via model options

Signed-off-by: anton ziderer <Antonziderer@mail.ru>
Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(silero-vad): ignore NaN thresholds

NaN passes the detector validation but prevents speech comparisons from succeeding.
Ignore it like malformed input and document the option validation.
Convert the option tests to Ginkgo and cover invalid overrides.

Assisted-by: Codex:gpt-6
Signed-off-by: anton ziderer <Antonziderer@mail.ru>

---------

Signed-off-by: anton ziderer <Antonziderer@mail.ru>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-10-02 23:59:49 +02:00
mudler-agentandEttore Di Giacinto e4fa051ee3 refactor(distributed): put the NATS-only paths behind interfaces (#12395)
* feat(messaging): add shared subject rules

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* test(messaging): cover BroadcastRoots, ControlRoots and SubjectRoot

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(messaging): add Broadcaster and enforce subject rules in every carrier

Broadcaster is the fan-out half of MessagingClient. The NATS client and
the in-memory FakeBus now refuse a subject outside the served roots and any
wildcard other than a whole single token, and FakeBus shares MatchSubject
instead of its own copy. FakeBus Unsubscribe now removes its own
subscription instead of the first one with the same subject.

A shared conformance suite in messagingtest runs against both carriers.
The distributed e2e specs that used invented test.* subjects, and the one
that subscribed with a > filter, now use subjects from subjects.go.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* refactor: depend on Broadcaster where only publish and subscribe are used

Narrowed to messaging.Broadcaster: nodes/staging_progress.go,
nodes/install_progress_publisher.go, galleryop/operation.go,
galleryop/service.go, agentpool/user_services.go, agentpool/agent_jobs.go,
openresponses/store.go, openresponses/sync.go, syncstate/syncstate.go,
finetune/service.go, quantization/service.go and
failover/distsync/distsync.go. SubscribeJSON now takes a Broadcaster
because it only calls Subscribe, which lets the narrowed consumers use it.

Stayed wide: worker/supervisor.go, because its client field also serves
the SubscribeReply handlers in worker/lifecycle.go. The request/reply,
queue and wiring files (nodes/unloader.go, nodes/file_stager_s3.go,
jobs/dispatcher.go, agents/dispatcher.go, agents/events.go,
worker/file_staging.go, cli/agent_worker.go, http/app.go) are unchanged by
design.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* refactor(nodes): name the no-route condition and confine the carrier error

Consumers matched nats.ErrNoResponders, which names an absence, to demote a node. They now match ErrNoRoute, the control path maps the carrier's failure onto it, and timeouts and worker refusals are pinned as not being no-route.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs(nodes): state which FileStager implementations return ErrNoRoute

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* refactor(nodes): build backend clients through one node-aware seam

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: describe the distributed transport seams

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: correct comments that overclaim after the seams refactor

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(agent-worker): refuse an unserved LOCALAI_AGENT_SUBJECT at startup

The messaging client now refuses a subject whose root no carrier serves.
An agent worker started with a custom LOCALAI_AGENT_SUBJECT such as
tenant-a.agent.execute used to start and then wait on a subject the
frontend never publishes to. After the subject rules landed it exited
at subscribe time with an error that did not name the setting.

Behaviour change: the worker now checks LOCALAI_AGENT_SUBJECT before it
registers or connects, and exits with an error that names the variable
and says to use a served subject under the agent root, for example
agent.execute. The served roots are not widened: a custom root was
never delivered by the frontend, and a wider set would reopen the
drift the subject rules exist to close. The flag help and the agent
worker docs state the constraint.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* test(nodes): pin the reactions to ErrNoRoute

Three callers react to ErrNoRoute and had no spec: the reconciler's
upgrade drain falls back to the legacy forced install, the reconciler
marks the node unhealthy when a pending op has no route, and the
backend-op fan-out marks the node unhealthy. Each spec drives the real
caller with a scripted no-responders reply and reads the result from
the registry or the recorded requests.

A fourth spec pins the other side: a pending op that times out leaves
the node healthy and only counts the attempt, so mapping timeouts onto
ErrNoRoute would fail here.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* test(messaging): pin client subject checks and fail the carrier suite in CI

Add specs that call Publish, Request, Subscribe, QueueSubscribe,
SubscribeReply and QueueSubscribeReply on a client with no connection.
Each call must return ErrUnservedSubject for bogus.thing and
ErrUnsupportedWildcard for jobs.>. This proves that the subject check
runs before the connection is used, and needs no server.

The NATS conformance suite is the only check that runs the subject rules
against a real carrier. Before this change it skipped without output
when Docker was missing. Now it fails when CI is set, so a Linux runner
without Docker cannot hide it. It still skips on local runs and on macOS
CI, which has no Docker.

Add SubjectNodeBackendInstallProgress to the list of constructors that
must build served subjects, and ask contributors to extend the list.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: state what ErrNoRoute may change, and group the distributed guides

The seams note said MarkUnhealthy was the only state change allowed
on ErrNoRoute. A pending backend op still records the failed attempt,
counts toward the reconciler's retry limit and is dead-lettered after
the maximum attempts. The note now says that MarkUnhealthy is the only
change to the node's own state, and that the per-op accounting is not
a verdict about the node.

The note also documents that the NATS conformance run fails under CI
when Docker is missing. The distributed-seams row moves next to the
distributed-state row in the topics table. The liveness ping spec
header now says no route is a reason to skip the worker, not proof
that the worker is gone.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* refactor(nodes): give the backend client factory the node id

Mechanical: the method gains a nodeID parameter and the eight test fakes are updated. No behaviour change.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* refactor(nodes): drop the optional node-aware factory

The node id is now in the main method, so the optional interface and its helper had no behaviour of their own.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* refactor(nodes): dial backend probes through the client factory

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* refactor(nodes): dial workers' file servers through a per-node dialer

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* refactor(http): proxy backend logs through the per-node worker dialer

The admin backend-logs proxy (list, lines and the WebSocket stream) now reaches a worker through the same per-node dialer as the HTTP file stager, so every frontend-to-worker dial goes through one seam. The shared direct dialer keeps alive for 15s where the proxy used 30s. Harmless for requests bounded at 15s.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(http): keep the backend-logs proxy independent of the admin connection

The proxy request had no context before the dialer change and is bounded only by its 15s timeout. Keep it that way.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* refactor: move the worker control payloads to workerctl

Mechanical move of the request and reply structs, the install progress event and the file payloads out of messaging. The verbs no longer belong to one carrier. No alias is left behind.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* refactor(worker): serve the lifecycle verbs through a controlServer

The worker registers one handler per verb and a NATS server maps each verb to its subject. Registration errors now name the verb. node.stop is served with SubscribeReply, which is identical on the wire because the handler never replies.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* refactor(worker): report install progress through the control sink

Install and upgrade now emit download progress through the sink the control server hands them. The debounce and the terminal flush stay in the handler path, built over that sink by the new nodes.NewDebouncedInstallProgressSink, which replaces NewDebouncedInstallProgressPublisher. The subject and payload on the wire are unchanged. The supervisor no longer holds the bus, and installFn and upgradeFn let specs drive both verbs without a gallery.

The malformed-request log lines are restored for install, upgrade, backend.delete, model.unload, model.stop and model.delete, with the reply bytes unchanged. The signal adapter is renamed noReply, which also lets worker.go import os/signal without an alias again.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* refactor(worker): serve the file-staging verbs through a controlServer

An empty list-dir answer is now {} rather than {"files":null}, because the typed reply omits an empty Files slice. The frontend decodes both to a nil slice in nodes/file_stager_s3.go.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(messaging): add WorkQueue and the NATS producer

This is the producer side of the competing-consumer seam. The work kinds map one to one to today's subjects and queue groups: task to jobs.new and mcp-ci to jobs.mcp-ci.new (both in group workers), agent-run to agent.execute (group agent-workers). Enqueue publishes the payload as Publish does today, with one JSON marshal. FakeBus now records queue groups and keeps reply handlers so later specs can pin and drive them.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(messaging): add the NATS WorkConsumer

An in-flight limit of one runs the handler inline on the delivery goroutine, as the MCP CI consumer does today. Any other limit spawns per delivery, as the agent consumer does. Queue groups are unchanged.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* refactor: publish queued work through WorkQueue

The job dispatcher, the agent pool and the agent scheduler enqueue through messaging.WorkQueue; the NATS implementation publishes to the same subjects as before. DistributedServices builds the queue next to the NATS client and hands it to the dispatcher and the agent pool, whose distributed mode switch now reads a non-nil WorkQueue. The unused AgentPoolService.SetNATSClient is removed.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* refactor: consume queued work through WorkConsumer

The agent dispatcher and the MCP CI consumer register through messaging.WorkConsumer. The NATS implementation keeps the inline one-at-a-time model for MCP CI and the per-delivery model for agent runs. handleMCPCIJob reports on the events publisher the carrier hands it instead of a captured client.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* refactor: delete the consumers nothing in production reached

jobs.new has a producer and no production consumer, and the agent dispatcher's Dispatch was only called from tests. Publishing jobs.new is unchanged.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* refactor(mcp): send MCP requests to agent workers through AgentControl

Timeouts still honour only the deadline, not cancellation, exactly as today. The NATS no-responders error maps to ErrNoRoute and a timeout does not.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* refactor(agent-worker): serve MCP requests and backend.stop through agentRPCServer

The agent worker's MCP tool and discovery reply subscriptions and its backend stop listener move behind an unexported agentRPCServer interface, served on NATS by nodes.NATSAgentRPCServer. The handlers become typed mcp.ToolHandler and mcp.DiscoveryHandler values that answer every failure with a reply carrying Error.

Queue group (agent-workers), inline execution on the delivery goroutine, the background handler context, the unmarshal error reply texts and the reply-less backend stop subscription are unchanged. The backend stop handler takes the decoded backend name, so it can still close that backend's MCP sessions.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* refactor(messaging): remove helpers that only tests used

BroadcastRoots, ControlRoots and SubjectRoot had no production caller. The roots spec now asserts every served root through ValidateSubject instead. MatchSubject moves back into the test support package, the only place that used it, with its table. NATSAgentRPCServer drops the subscription list it stored and never read, and NewNATSAgentRPCServer gets a doc comment.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* test(mcp): round trip the agent RPC server over a real NATS server

One spec sends a tool request and a discovery request through NATSAgentControl to NATSAgentRPCServer and checks that the handlers see the decoded requests and the replies come back. It also puts an undecodable body on the tool subject and checks the server answers with an unmarshal error instead of leaving the requester to time out.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: describe the distributed transport seams

The developer note now lists the final seams: fan-out, queues, both halves of the control verbs and of agent RPC, and the dial. It records the open items a second carrier has to handle.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* test: pin the in-flight limit each queue consumer asks for

The work queue specs pin what Consume does for a given limit, but nothing
pinned which limit each production consumer passes. Changing the agent
worker's MCP CI limit from 1 to 0 would have let MCP CI jobs run
concurrently on each worker with every test green.

Move the MCP CI Consume call into startMCPCIConsumer with the same wiring
and pin that it asks for (WorkMCPCI, 1). Pin that NATSDispatcher.Start asks
for (WorkAgentRun, maxConcurrent) for several limits.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* refactor: remove helpers the branch left without a caller

SubjectJobCancelWildcard lost its last subscriber when the frontend
stopped listening on jobs.*.cancel; the NATS permissions and conformance
suite spell the subject out, so nothing reads the constant.

decodeBackendStopRequest returned a stopAll flag that production dropped
and only a test read. decodeBackendStop is now the single decoder with the
same semantics: an empty body is stop-all, an empty Backend is stop-all,
malformed JSON is an error. stopBackends still derives stop-all from
Backend, so no reply changes.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(messaging): keep an explicitly empty agent queue a plain subscription

Before the work queue seam the agent worker passed LOCALAI_AGENT_QUEUE
straight to QueueSubscribe, so an explicitly empty value made a plain
subscription and every agent worker ran every agent run.
WithAgentRunRoute replaced an empty queue with agent-workers, which
silently changed that.

Keep the queue as given once the option is applied. An empty subject still
falls back to agent.execute, since it never had a meaning of its own. The
flag default stays agent-workers, so only an explicitly empty value
reaches this.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: correct comments and record the PR B notes

Fix the recordingFactory comment (it also records the parallel flag),
document that a negative maxInFlight is unbounded and that Unsubscribe from
a handler deadlocks, and say a permanently undecodable payload returns nil.

Record controlHandler's undecodable return as a kept exception, and add the
second carrier notes to the developer note: the reconciler has no
ClientFactory option, the logs proxy honours HTTP_PROXY, verbs one carrier
serves need an opt-out, terminal replies come from the result event, and
agent runs publish through the NATS-bound EventBridge, which is not an
additive change.

Assisted-by: Claude:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-02 23:58:41 +02:00
mudler-agentandEttore Di Giacinto 38a2aa45fe feat(gallery): add Nimble 9B and CLM decision models, bump vllm.cpp to a19294a9 (#12397)
* feat(gallery): add nimble-9b-vllm-cpp decision model

Add Bespoke Nimble 9B, converted for vllm.cpp and pinned to the weights
commit 52eead25 of mudler/Bespoke-Nimble-9B-vllm-cpp (HEAD only adds the
model card). It is a redistribution of bespokelabs/Bespoke-Nimble-9B
with the LoRA merged into Qwen3.5-9B; config.json names NimbleModel, so
no hf_overrides are needed.

The artifact sits under overrides, where the installer reads it. The
entry sets an 8192-token context, Nimble's own prompt limit, and a KV
pool of 1024 blocks of 32 tokens for 4 sequences (about 1 GiB at 32 KiB
per token for the 8 full-attention layers).

Installed with local-ai models install and served on CPU through the
vllm-cpp backend: the model card's billing request gives billing
(0.986), refund 0.998 and urgency 0.33. Peak resident memory was
18.4 GB, so the description asks for about 20 GB of free RAM.

List the entry in the decisions gallery table. CLM stays out of the
gallery: the pinned engine cannot load the published head layout.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5

* feat(gallery): add clm-v0.1-8b-vllm-cpp and bump vllm.cpp for CLM

The CLM checkpoint on mudler/CLM-v0.1-8B-vllm-cpp stores the heads with
the reference's own tensor names (state_head.inp, hidden.N, norms.N,
out). The pinned vllm.cpp 96788348 still expects the old .0/.2/.4/.6
layout and refuses the load with "head.safetensors incomplete for
state_head". vllm.cpp a19294a9 matches the reference layout and adds the
converter that produced the upload, so move the pin there. The ABI stays
at v30.

Add the CLM entry, pinned to the weights commit 0d1903b1 (HEAD only adds
the model card), with a 4096-token context and a KV pool for 4 sequences
(about 2.25 GiB at 144 KiB per token for Qwen3-8B).

Installed with local-ai models install and served on CPU against a
libvllm built at a19294a9: the model card example (john works at google,
entity type) gives person 0.950, the same as the card. Peak resident
memory was 17.9 GB.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-02 09:41:31 +02:00
Stefan Walcz c3bea567fe fix(llama-cpp): do not stream the error text as content on pre-stream failures (#12425)
When the first result of a streamed request is an error (for example a
prompt that exceeds the context), PredictStream wrote the error message
as a Reply and only then returned the error status. LocalAI treated that
Reply as the first token: it sent the assistant role chunk and the error
text as `content` on an HTTP 200 stream. Because a chunk had already been
written, the pre-stream HTTP error path from #12204 never triggered, so
streaming clients still got a 200 with the error as model output, while
the same request without streaming correctly returns a 400.

Return the error only as the gRPC status. The e2e backend suite gets a
`context_overflow` capability (enabled for llama-cpp) that streams an
over-long prompt and asserts an error status with no content.

Assisted-by: Claude:claude-opus-5-5

Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
2026-10-02 09:23:57 +02:00
Stefan Walcz 1056c62f4c fix(cloud-proxy): surface Anthropic refusals instead of empty replies (#12424)
When an Anthropic model declines to answer, the Messages API returns
HTTP 200 with `stop_reason: "refusal"` and empty `content`. The
translate mode mapped that to a normal reply with no content, so the
OpenAI-compatible response looked like a successful completion
(`finish_reason: "stop"`, empty message). Routers, agents and UIs could
not tell "the model declined" from "the model had nothing to say", and
no fallback was triggered.

Return an explicit error for `stop_reason: "refusal"` in both the
non-streaming path and the streaming path (`message_delta`). Regular
replies, including empty `end_turn` replies, are unchanged.

Assisted-by: Claude:claude-opus-5-5

Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
2026-10-02 09:22:49 +02:00
Stefan Walcz 985df46d24 fix(whisperx): keep the transcript when diarization fails, report real errors (#12427)
AudioTranscription caught every exception and returned an empty
TranscriptResult. A failed diarization step therefore discarded a
transcript that was already finished: with an HF token that has not
accepted the terms of the gated pyannote pipeline, the download fails
with 403 and every transcription came back as an empty text with
HTTP 200.

Diarization now degrades: if it fails, the transcript is returned
without speaker labels and the reason is logged. Any other failure
aborts the call with INTERNAL instead of pretending success with an
empty text.

Assisted-by: Claude:claude-opus-5-5

Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
2026-10-02 09:18:12 +02:00
Stefan Walcz 6aa7b9b871 fix(llama-cpp): let parallel:1 in the model options win over LLAMACPP_PARALLEL (#12426)
The environment fallback was applied whenever n_parallel was still 1
after option parsing. An explicit `parallel: 1` in the model YAML is
indistinguishable from the default that way, so it was replaced by
LLAMACPP_PARALLEL. The docs say options in the YAML take precedence
over environment variables; a single model could not be forced to one
slot while the global variable was set.

Track whether the options set the slot count and resolve it in a small
helper (parallel_params.h): option first, then LLAMACPP_PARALLEL, then
1. The helper gets a standalone unit test picked up by
`make test-backend-cpp`.

Assisted-by: Claude:claude-opus-5-5

Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
2026-10-02 09:17:44 +02:00
mudler-agentandEttore Di Giacinto 9eb5a9e61d feat(audio): remember speakers from diarization (#12414)
* feat(schema): validate portable speaker profiles

Add the versioned profile schema for explicit speaker enrollment.
Validate compatibility against separately supplied loaded-encoder metadata.
Reject unusable speakers, invalid vectors, and inconsistent clean spans.

This slice does not change HTTP routes, backend integration, or the UI.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(parakeet): export profiles with transcripts

Export opt-in speaker profiles and trusted encoder metadata.
Replay registrations by ID so duplicate display names keep independent
vectors.

Use one profile-capable diarization for slots, names, and clean spans.
Assign timestamped ASR words to those slots without a second diarization.
Preserve legacy opt-out and no-ASR behavior, and propagate failures.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(audio): enroll portable speaker profiles

Gate profile exports with voice-recognition permission and validate
registration against metadata from the loaded encoder. Preserve audio
enrollment and independent registrations with duplicate display names.

Exclude diarization and registration exchanges before API trace capture
so persisted traces cannot retain profile vectors or JSON audio.

Defer candidate dimensions to trusted loaded metadata. Sort candidates
by registration ID so incompatible profiles cannot suppress legacy voices
through registry iteration order. Keep portable identity checks closed
when trusted metadata is unavailable.

Test persisted traces, explicit slot zero, and selection through offline
and live transport. Document privacy and the ephemeral registry lifecycle.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(ui): remember speakers from diarization

Add a Studio page for diarization and opt-in speaker profiles. Preview
clean intervals from the original recording before explicit registration.

Join profiles by raw speaker labels, preserve duplicate names, and relabel
turns only after a successful save. Discard stale results when the model
or recording changes. Share registration metadata with voice management
without storing vectors or recordings from this flow.

Document permissions and the global, ephemeral registry. Cover enrollment,
permissions, previews, and asynchronous races with mocked Playwright tests.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: clarify HTTP speaker enrollment support

Replace the stale enrollment limitation with the current HTTP workflow.
Distinguish native transport from explicit registration and link its docs.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* chore(parakeet): pin merged speaker profile support

Use the merged commit from mudler/parakeet.cpp#80.
Its tree matches the previously accepted native pin.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: add diarization enrollment setup example

Connect the existing gallery modes to the speaker enrollment workflow.
Show installation, private profile export, explicit raw-slot registration,
and later recognition without another export.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs(blog): explain diarization speaker profiles

Put the diarization walkthrough on the LocalAI website in the feature PR.
Cover the three gallery modes, explicit enrollment, and privacy limits.
Link setup instructions and keep availability conditional on feature support.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs(blog): focus diarization on everyday use

Explain what users can do with recordings before the setup steps.
Replace the technical walkthrough with a short Studio guide and link
readers to the existing reference for model names and developer use.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs(blog): lead with speaker capabilities

Present speaker recognition through everyday uses and a short UI flow.
Keep technical reference details in the existing documentation.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(diarization): satisfy Go lint checks

Avoid copying protobuf message state when extending backend status, check the multipart reader close result, and document the focused testing.T lint exemptions.

Assisted-by: nib:gpt-5.6-sol

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-02 08:14:00 +02:00
localai-org-maint-botandlocalai-org-maint-bot 668802d6ca fix(watchdog): ignore stale backend evictions (#12333)
Remove watchdog tracking when backends stop, crash, or fail to start.
Validate eviction addresses under the model lifecycle lock so delayed
shutdowns cannot stop a replacement backend. Preserve replacement size
estimates when removing an old address.

Add lifecycle regression tests and document shutdown behavior.

Fixes #12331

Assisted-by: Codex:gpt-6

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-10-01 14:46:41 +02:00
Pratik Gandhi 3d0311e640 docs: fix Discord, Slack and Telegram example links (#12392)
The chat bot examples moved to mudler/LocalAI-examples; the old paths in
LocalAGI and LocalAI no longer exist.

Assisted-by: Claude:claude-opus-5-5

Signed-off-by: Pratik Gandhi <travpreneur@gmail.com>
2026-10-01 08:40:26 +02:00
localai-org-maint-botandEttore Di Giacinto 2ae6cae70d feat(parakeet-cpp): name speakers from the shared voice registry (#12382)
* feat(voice): list registered voices and record which encoder made them

The voice registry could register, identify and forget but not list, and
it did not remember which speaker encoder produced an embedding. Add
Metadata.Model and Registry.List, answered from the index the store
registry already keeps for Forget. Needed so a backend can be given the
registered voices that match its own speaker encoder.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(voice): store the encoder model with a registered voice

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(voice): pick the registered voices that match a speaker model

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(proto): carry known voices and speaker names on diarize and live messages

Assisted-by: Claude:claude-haiku-4-5 [Claude Code]

* feat(diarization): name speakers from the voice registry

When a diarization model has a speaker_model option, the endpoint sends
the registered voices made by that encoder to the backend. The backend's
name and name_score come back as extra fields next to the normalized
SPEAKER_NN speaker, and the speakers summary carries the first name seen
for each speaker. RTTM output and results without names are unchanged.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(live): pass registered voices to a live session and surface speaker names

Live sessions now send the registered voices that match the model's
speaker_model to the backend, and each speaker segment carries the name
the backend matched. The realtime segment event gains an optional
speaker_name field.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(parakeet-cpp): load a speaker model and build per-request voice registries

Adds the speaker bindings (ABI v9 and v10, probed separately), the
speaker_model, speaker_threshold and speaker_margin options, and a
per-request registry builder over the known voices.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(parakeet-cpp): name the speakers in Diarize from the known voices

Diarize builds a per-request speaker registry from the known voices when a
speaker model is loaded, calls the named C functions, and puts each slot's
registered name and score on the segments. The registry is freed on every
path. A library without ABI 10 reports Unimplemented instead of dropping
the names.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(parakeet-cpp): name speakers in the live scene stream

The live scene stream now begins with a known-voice registry when a
speaker model is loaded and the live config carries voices, and each
closed speaker segment takes its slot's current name from the feed's
names map. A segment that closes before its slot is identified has an
empty name. The registry is freed after the stream, on every path.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(gallery): speaker naming entries and docs for parakeet-cpp

Add three gallery entries that load the WeSpeaker ResNet34 speaker model
next to the diarization or realtime scene models, and document speaker
names in the voice recognition, diarization, audio to text and realtime
pages.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* fix(parakeet-cpp): skip an unusable registered voice instead of failing the request

A registered voice with the wrong embedding size, or one the C side
refused, failed the whole diarization request, so one legacy voice broke
the model for every user. Skip such voices with a warning that does not
carry the voice name, and take the plain path when none is left.

Also map an exact 0 speaker threshold or margin to a tiny positive value,
since the C side reads 0 as "use the default", and fix a stale comment
about which contexts Free() walks.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* fix(diarization): warn once per model about voices from another encoder; document the privacy limit

The different-encoder warning fired on every request. Log it once per
feature and speaker model, then at debug level. Document that the global
voice registry lets any caller of a speaker_model model learn matching
names, and that skipped wrong-sized voices are logged.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* chore(parakeet-cpp): bump parakeet.cpp to 8c8cec0 (C-API v10) and check speaker naming against the real library

The pin moves from 623a968 to 8c8cec0, which brings in everything merged
in parakeet.cpp since: the voice identification change (C-API v9, #78) and
raw-embedding enroll plus diarize-only speaker naming (C-API v10, #79).

New real-library specs (gated on PARAKEET_BACKEND_TEST_SPEAKER_MODEL,
_DIAR_MODEL, _WAV and, for the live path, _STREAM_MODEL) name the two
speakers of two_speakers.wav from a committed pair of WeSpeaker embeddings,
with the voices passed in reversed order. They also check that the float32
threshold reaches C through purego. The shared test loader now registers
the v9/v10 and scene symbols as main.go does.

The rebase onto origin/master had no conflicts.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-01 08:25:52 +02:00
localai-org-maint-botandlocalai-org-maint-bot 51be7b48e5 chore(gallery): add Cyber-Ornith 1.5 variants (#12383)
Add pinned Q4_K_M and Q6_K builds for text chat with llama.cpp.
Document installation and explicit variant selection.

Assisted-by: Codex:gpt-6

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-10-01 08:24:00 +02:00
Ettore Di Giacinto 975ff5ca42 feat(gallery): add kev-0.8b-vllm-cpp decision model
Add the kev 0.8B decision model, converted for vllm.cpp and pinned to
revision c17e7366 of mudler/kev-0.8b-vllm-cpp. It is a redistribution of
jaredpalmer/kev-0.8b with the LoRA merged and the PointerHead stored as
head.safetensors, so only the vllm-cpp backend can load it.

The artifact sits under overrides, where the installer reads it. The
entry sets a 2048-token context and an explicit KV pool: with the
default 4096-token context the CPU KV pool holds only 4064 tokens and the
load fails.

List the entry in the decisions gallery table and drop kev from the
list of decision models that are not gallery entries yet.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5
2026-09-30 22:28:13 +00:00
mudler-agentandEttore Di Giacinto b540c3e1fd chore(vllm-cpp): bump to 967883486 (ABI v30), add hf_overrides and Tev1 entries, fix vllm-cpp gallery installs (#12379)
* chore(vllm-cpp): bump vllm.cpp to 967883486 (ABI v30)

Moves the pin from c3bebc357 to 967883486. On top of the Nimble decision
adapter and the Qwen3.5 vision-loader fix, this brings Tev1 on
/v1/systemone and vllm_decide (opt-in through a "Tev1Model" architecture
in config.json), a tokenizer/ subdirectory fallback so the Laya HF
snapshot loads as downloaded, a stop-token fix, a logprobs fix under async
scheduling and a pinned parakeet.cpp fetch for the diarization build.

ABI v30 only adds the diarization and speaker-attributed ASR entry
points; no existing struct or signature changed, so the purego mirrors
keep their layout and only abiVersion moves to 30. Between 4479dc99f and
967883486 vllm.h changed only in a comment.

v30 turns VLLM_CPP_WITH_DIARIZATION on by default. The fetch is pinned
now, but ON still downloads parakeet.cpp at configure time and links a
second ggml into libvllm for calls this backend never makes, so build
with the option off: the symbols stay present as refusing stubs.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5

* feat(vllm-cpp): add the hf_overrides engine arg

vLLM parity: engine_args.hf_overrides is a JSON object of top-level
config.json keys merged over the model directory's own config.json. The
main use is opting a published checkpoint into an engine adapter its
config does not name, such as {"architectures": ["Tev1Model"]} on the
Tev1 snapshots, which declare Qwen3_5ForConditionalGeneration.

The C ABI has no override input and the engine reads config.json from
the directory it is given, so Load builds a private overlay directory:
the merged config.json plus a symlink to every other entry of the model
directory, and passes that to the engine. The download is never written.
Free, a failed load and the next Load remove the overlay.
validModelPath and the DFlash draft resolution still see the real
directory.

A value that is not a JSON object, a .gguf model or a directory without
config.json fails the load instead of being skipped like an unknown
engine_args key, because loading the unmodified config would serve a
different architecture than the one configured.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5

* fix(gallery): nest vllm-cpp artifacts under overrides

artifacts: is a model-config key, and the installer reads model-config
keys only from overrides:. Five vllm-cpp entries (laya, gliner25-decide,
qwen3-vl-4b, cua-s1-forms and gliner2.5) declared it at the entry top
level, where it is silently dropped: the install reports success, writes
a config whose model is the bare HF repo id and downloads nothing, and
vllm-cpp (which does not infer artifacts) then fails the first load with
"model path not found".

Move each block under overrides:, and add a guard test that refuses a
top-level artifacts: key in the index.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5

* feat(gallery): add Tev1 4B and 0.8B on vllm-cpp

Two decisions entries for Together AI's Tev1 checkpoints, pinned to the
current HF revisions. Tev1 is autoregressive: vllm.cpp answers
/v1/systemone by scoring the option letters, and the same engine still
serves chat completions. The published config.json names
Qwen3_5ForConditionalGeneration, so each entry sets
hf_overrides: {architectures: [Tev1Model]} to enable the decision route
without editing the download. known_usecases is [decisions] only, since
a declared decisions list is authoritative for reservation.

The descriptions state what was checked: agreement with transformers on
CPU over seven questions (4B 7/7, max probability difference 0.0004;
0.8B 6/7 with one near tie), CPU-only for the decision route, and a
fine-tune license the model card says is still being finalized, so no
license key is set.

The Decisions API page lists both entries, drops the note that Tev1
does not serve /v1/systemone and documents the 24-option limit (Ollama
allows 26).

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-30 20:22:49 +02:00
Ettore Di Giacinto 5613572f38 feat(systemone): validate requests and align the docs with Ollama's contract
All three routes now validate the request before it reaches a model: body
size (413 over 64 KiB), state, question count, blank ids, option and level
counts, and noul criteria keys. Forwarded decision requests skipped this
before, so a malformed question surfaced as a backend error.

The docs claimed the wire shape matches Ollama's. Field names and question
types do; confidence, error shape, keep_alive and state rendering differ, and
the docs now say so.

Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-30 14:31:50 +00:00
Ettore Di Giacinto 70ce62901f refactor: name the capability decisions instead of systemone
The usecase describes what a model can do, and the category is the Decisions
API. SystemOne stays as the wire contract: the /v1/systemone routes, the
Score RPC question_type and the swagger tag are unchanged. The usecase,
flag, auth feature, UI label, gallery tags and docs page are now decisions.

Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-30 14:14:09 +00:00
Ettore Di Giacinto b3d65fd538 fix(systemone): route NER models to the NER path and refuse decision models on permute and separate
vllm_decide refuses NER architectures and the NER entry point refuses decision
architectures, so each model kind 500ed on half of the routes. A token_classify
model now goes to the NER path on /v1/systemone, and /permute and /separate
return 400 for decision models. Docs and instructions state which kind serves
which route.

Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-30 11:55:26 +00:00
Ettore Di Giacinto c7f278dd0d docs: document the systemone usecase and decisions API
Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-30 11:37:48 +00:00
2fa36e7147 feat(parakeet-cpp): speaker diarization, sound detection and live scene events (#12335)
* feat(parakeet-cpp): load diarization and CED models and companions

Repin PARAKEET_VERSION to parakeet.cpp PR #75's head, which adds
parakeet_capi_model_kind (ABI v8). Bind the new diarization, sound
event and combined scene stream C symbols through the same
purego.Dlsym probe pattern already used for the batched JSON entry
point, so the backend still loads against an older libparakeet.so.

Load now classifies the loaded GGUF by role (ASR, diarization or
sound) via parakeet_capi_model_kind and can load up to two companion
models from Options[] (asr_model:, diarization_model:, sound_model:,
paths resolved against opts.ModelPath), verifying each companion's
kind and freeing every context opened so far on any failure. Free
releases the primary and every companion. AudioTranscription now
names the loaded role when it is not ASR instead of a generic model
not loaded error. The dynamic batcher starts only when an ASR context
ends up loaded, primary or companion.

This is groundwork only: the Diarize and SoundDetection RPCs and the
live scene stream that actually use these new roles land in later
commits.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): reset role fields on a failed companion load

loadRoles' freeLoaded only released the C contexts it had opened; it
left ctxPtr/diarCtx/tagCtx and companions pointing at those now-freed
contexts, so a later Free() on the same instance would double-free.
Zero all four alongside the CppFree calls.

Also route AudioTranscriptionStream and AudioTranscriptionLive through
notASRError when ctxPtr is unset but a diarization or sound model is
loaded, matching AudioTranscription: both used to return the generic
model-not-loaded error instead of naming the loaded role.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(parakeet-cpp): add speaker diarization

Implement the Diarize RPC for the parakeet-cpp Go backend, wired to
Nemotron-3-Diarization through libparakeet.so's diarization C-API.

Plain diarization uses parakeet_capi_diarize_pcm; when include_text is
set and an ASR companion is loaded, parakeet_capi_transcribe_and_
diarize_json fills each segment's text instead. Speaker labels are the
decimal index, or "unknown" for -1 (no diarized speaker overlaps).
min_duration_off merges same-speaker segments across a short gap
before min_duration_on drops the segments still too short, then ids
are renumbered. num_speakers/min_speakers/max_speakers/clustering_
threshold have no Sortformer equivalent and are logged at debug
instead of rejected.

Verified against the real Nemotron-3-Diarization + parakeet-tdt_ctc-
110m checkpoints on the two_speakers.wav fixture: correct A-B-A-B
speaker segmentation and matching speaker-attributed transcripts.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(parakeet-cpp): add sound event detection

Wire the SoundDetection RPC to the CED tagger context (p.tagCtx)
loaded by Task 1's role classification. It runs the whole clip
through a one-shot parakeet_capi_sound_stream_* session (window
10s, hop 10s, top_k set to the tagger's class count so every
drained window carries a full score list), averages each class's
score across the drained windows, sorts descending, then applies
the request's threshold and top_k (0 keeps every class).

No tagCtx returns FailedPrecondition; a libparakeet.so missing the
sound_stream symbols returns Unimplemented. Every C call runs under
engineMu, and the stream is always freed, even when a feed or drain
call fails partway through.

Verified against a real ced-tiny-q8_0.gguf on the rooster.wav demo
clip: "Chicken, rooster" tops the list at score 0.91.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): cancel sound detection mid-feed, shrink the lock

SoundDetection now checks ctx before each 10 s feed slice (mirroring
driver.go's feedSlices) and returns Canceled if the caller gave up,
so a long clip can be interrupted instead of feeding to completion
regardless. The stream is still freed on every path, cancellation
included.

Also narrow engineMu to the C calls: the drained JSON document is
now decoded after the lock is released, splitting soundStreamScores
into a locked soundStreamDrain (opts, begin, feed, drain, free) and
an unlocked json.Unmarshal.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(parakeet-cpp): stream speaker and sound events during live transcription

Add two additive proto fields, LiveSpeakerSegment and LiveSoundEvent,
repeated on TranscriptLiveResponse. When a diarization or sound
companion model is loaded, AudioTranscriptionLive now runs a no-ASR
scene stream (parakeet_capi_scene_stream_begin) beside the ASR
streaming session, feeding it the same PCM slices and forwarding any
closed speaker or sound events alongside the matching ASR delta, or
on their own when a slice has no ASR output.

The scene stream is freed and reopened on a mid-stream Config reset,
flushed with is_last before the closing FinalResult, and degrades
gracefully (a warning, not an error) when begin or a later feed call
fails, so live transcription keeps working ASR-only. Existing live
behavior is unchanged when no companion is configured, and no scene
C call is made in that case.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): keep scene events off the ASR critical path in live

Emit each slice's ASR result right after the ASR feed, before the
scene feed for that slice runs, so a companion diarization/sound
model never adds scene compute latency in front of the delta or
<EOU> that drives realtime turn detection. Closed speakers/sounds go
out afterward as their own response, so a slice with both now
produces two responses, ASR first. The live feed log line now
reports ASR and scene wall time separately.

Re-check the diarization/sound contexts a scene stream was begun
with against the live contexts before every feed, under the same
lock: Free() can race between an ASR feed and the matching scene
feed and free the model the stream borrows. A mismatch now returns
without touching the C side. Freeing the stream itself stays
unconditional; the scene stream's destructor only releases its own
buffers and never touches the borrowed contexts.

Also recover a panicking stub inside the live test goroutine instead
of crashing the test binary, and reset the live decode-lag tracker on
a mid-stream config reset, matching what its own comment already
promised.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(realtime): surface live speaker and sound events

Carry the backend's closed speaker segments and sound events
(TranscriptLiveResponse fields 7/8) through LiveTranscriptionEvent
as LiveSpeakerSegment/LiveSoundEvent (nanoseconds mapped to
seconds), and forward them from the semantic_vad live path.

Each speaker segment emits
conversation.item.input_audio_transcription.segment with speaker,
start, end and empty text under the turn's item id. Each sound
event emits conversation.item.sound_detection with one tag
(label, score = peak, index) and the event's new optional
start/end seconds fields, omitted when unset so the existing
unary/windowed sound-detection path is unaffected.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(realtime): keep start/end on a zero-second transcription segment

ConversationItemInputAudioTranscriptionSegmentEvent.Start/End used
omitempty, so a speaker segment starting at 0.0s dropped its
"start" key. Nothing emitted this event before the live scene-event
path, so drop omitempty: the segment always carries real times.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* chore(gallery): add parakeet-cpp diarization, CED and realtime scene models

Add gallery entries for the new parakeet-cpp capabilities: standalone
Nemotron-3-Diarization, the same paired with the Parakeet TDT+CTC
110M ASR model for speaker-attributed text, CED-Tiny and CED-Base
sound classifiers, and a realtime scene bundle combining the
streaming EOU ASR model with diarization and sound companions.

SHA256 taken from the Hub API; licenses from each model card
(openmdw-1.1 for Nemotron-3-Diarization, apache-2.0 for CED,
cc-by-4.0 for the Parakeet ASR models).

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: document parakeet-cpp diarization, sound detection and live scene events

Cover the new parakeet-cpp capabilities across the feature pages:
Nemotron-3-Diarization as a diarization backend (with and without
speaker text, the ignored speaker-count hints, the Sortformer
voice-like-sound quirk), CED as a sound classification backend, the
asr_model/diarization_model/sound_model/diarization_latency companion
options, and the realtime live speaker/sound events (event shapes,
the speech-turn-only limitation, and using this or
pipeline.sound_detection but not both).

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(gallery): correct the realtime-scene license and wording nits

parakeet-cpp-realtime-scene mistakenly copied cc-by-4.0 from the
existing realtime_eou_120m-v1 entry; the model card lists the NVIDIA
open model license instead. Switch to the gallery's usual spelling
for that license and keep the diarization/CED licenses called out in
the description.

Also: audio-diarization.md now says getting per-segment text needs
both an asr_model companion and include_text=true on the request, and
audio-to-text.md's option table reads "Use on" (a pairing the loader
does not enforce) instead of "Allowed on".

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): reject a companion role that duplicates the primary's

loadRoles let a companion option (asr_model:/diarization_model:/
sound_model:) assign into a role field the primary already occupied,
for example asr_model: on an already-ASR primary. The companion's
context silently overwrote ctxPtr/diarCtx/tagCtx, and Free() only
walks those three fields, so the original primary context was never
freed again.

Reject a companion whose role the primary already holds before its
GGUF is even loaded, freeing everything loadRoles opened so far, the
same way a wrong-kind companion is already rejected.

Also warn, rather than silently fall through, when
parakeet_capi_model_kind reports PARAKEET_MODEL_KIND_NONE for a
successfully loaded primary; the primary is still treated as ASR,
matching today's behavior.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): cap live scene sound score retention

sceneBegin started the live diarization/sound companion stream with
the C API's default sound options, whose top_k keeps 5 scores per
window forever until drained. The live scene path never drains sound
scores (only the offline SoundDetection RPC does, with its own fresh
stream), so this window queue on the C side grew for the whole
session's lifetime.

Set opts.Sound.TopK = 0 before starting the scene stream: this
disables score retention while leaving sound event detection (onset/
offset), which the live path actually consumes, unaffected.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): merge diarization segments per speaker, harden Diarize

mergeCloseSegments only compared neighbors in the single start-sorted
segment list, so two same-speaker segments never merged once another
speaker's turn fell between them (A, B, A): the short B segment broke
the adjacency the merge relied on. Group segments by speaker first,
merge within each speaker's own start-ordered run, then re-sort the
result by start so interleaved speakers come back out in timeline
order.

Also harden Diarize's entry points the same way streamFeedDoc/
sceneFeed already are: diarizeCall re-checks p.diarCtx (and, on the
include_text path, p.ctxPtr) under engineMu right before the C call,
so a Free() racing between Diarize's own checks and the lock can no
longer reach the C side with a freed context. When the include_text
call returns NULL, last_error is now read from both contexts and
whichever came back non-empty is reported, since either side of the
pairing can be the one that failed. A WAV decode failure is reported
as InvalidArgument instead of an unwrapped/untyped error.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): harden SoundDetection's engine checks

soundStreamDrain ran every C call under engineMu but never re-checked
p.tagCtx there, so a Free() racing between SoundDetection's own
tagCtx==0 check and this lock could still reach the C side with a
freed context. Re-check p.tagCtx under the lock and return
ModelNotLoaded when it was cleared, mirroring diarizeCall's own
re-check. A WAV decode failure is now reported as InvalidArgument
instead of an unwrapped/untyped error.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* test(parakeet-cpp): cover a mid-session scene feed failure

feedSlicesScene already degrades gracefully when a scene feed call
fails mid-session: it frees the broken stream and carries the ASR-only
session forward. Add a spec covering that path end to end: the scene
stream is freed exactly once, later audio slices still produce ASR
responses, and no speaker/sound events appear before or after the
failure.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: fix the parakeet-cpp companion role table and realtime scene docs

audio-to-text.md's companion option table read "Use on" with a note
that the loader did not enforce the pairing; it now rejects a
companion whose role duplicates the primary's, so restore the
"Allowed on" wording and describe the real enforcement.

openai-realtime.md's live speaker/sound section claimed a mid-stream
session.update resets the companion stream and that it flushes on
session close; neither happens, since the realtime core opens one
live stream (and so one scene stream) per speech turn and closes it
at that turn's commit, with no mid-stream Config in between. Document
that lifecycle instead, state precisely that start/end are seconds
from the start of the turn's own audio, and note that the diarization
model starts a fresh session every turn, so a speaker index is only
meaningful within one turn. The example sound tag ("Rooster", index
17) did not match any real CED label; index 17 in ced-tiny-q8_0.gguf
is "Baby laughter". Replaced with "Chicken, rooster" at its real
index, 99.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): use CED's real index for Chicken, rooster

The scene feed comment and the live test's canned document gave
"Chicken, rooster" index 365. In CED's AudioSet label list it is 99,
which is also what the realtime docs show.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(realtime): call the test event accessor

The scene-event tests range over a method instead of its returned slice.
Call the synchronized accessor so the OpenAI test package compiles.

Assisted-by: Codex:gpt-6
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* chore(parakeet-cpp): pin parakeet.cpp master with sound events

mudler/parakeet.cpp#75 (sound events, scene stream, model kinds) and
#74 (the missing <algorithm> include that broke the image builds) are
on master now. Pin 6dea76a instead of the #75 PR head, and update the
header comment the bump bot reads.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* chore(parakeet-cpp): pin parakeet.cpp with ced.cpp on main

parakeet.cpp #76 moved its ced.cpp submodule from the head of
localai-org/ced.cpp#3 (a branch-only commit) to ced.cpp main, where
#3 landed with an identical tree. Pin 623a968 so the image builds no
longer depend on that branch.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(transcription): carry speaker labels on words and streamed segments

A diarizing backend could label transcript segments, but two paths
dropped the label: TranscriptWord had no speaker field, so live
transcription words and word-level timestamps could not carry one, and
the stream=true transcript.text.done event left the speaker out of
its segments.

TranscriptWord gains an optional speaker (proto field 4, additive).
It flows through the live event and result mapping, the JSON word
output of the endpoint and the CLI, and transcript.text.done now
includes a segment's speaker when there is one. Empty labels are
omitted, so responses without diarization are unchanged.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
(cherry picked from commit 2f0049f979)
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(importers): detect the parakeet.cpp diarization GGUF

The Nemotron-3-Diarization GGUFs are published in
mudler/parakeet-cpp-gguf as nemotron-3-diarization-<quant>.gguf. The
parakeet-cpp importer did not recognise that name, so a direct
`local-ai models import` of the file fell through to another importer.

A direct URL to the file now imports with the diarization usecase. A
repo import still picks ASR weights when the repo also ships the
diarization model, and falls back to the diarization weights only
when there are no others.

Ported from #12323.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(config): advertise diarization and sound detection for parakeet-cpp

The capability table listed parakeet-cpp as transcription only, though
the backend now answers Diarize (Nemotron-3-Diarization) and
SoundDetection (CED) depending on the model kind it loads.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(parakeet-cpp): label transcript segments with the diarization companion

A diarization_model companion only fed live speaker events and
Diarize; /v1/audio/transcriptions ignored it.

With the companion attached and diarize=true (the OpenAI endpoint's
default), unary transcription now labels each segment with its
speaker and splits segments at speaker turns; with word timestamps
each word carries its speaker. The stream=true final result labels
each utterance with the speaker who said most of it. Both use the
checkpoint's own diarization over the whole clip, as NeMo's diarize()
does. Words take the speaker whose segments overlap them most, or the
nearest segment within 0.5 s, the same rule as parakeet.cpp's
speaker-attributed ASR.

Docs: the diarization_model row and a paragraph on transcript
speakers; Nemotron-3-Diarization handles up to 8 speakers.

Ported from #12323.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(realtime): speaker segments from committed-turn transcription

Speaker events reached a realtime session only from the live
semantic_vad path, which needs a cache-aware streaming transcription
model. Committed-turn transcription (server_vad, or any offline
model) always asked the backend for diarize=false and dropped the
segments' speakers.

pipeline.diarization (off by default) asks the transcription model for
speaker labels on each committed turn and emits every labelled segment
as a conversation.item.input_audio_transcription.segment event, with
its text, before the turn's completed event. It is opt-in because some
backends fail a diarization request they cannot serve.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(gallery): add parakeet-cpp-realtime-scene-tdt

parakeet-cpp-realtime-scene pairs the streaming EOU model with the
diarization and CED companions; its speaker and sound events need a
cache-aware streaming model. This entry does the same with Parakeet
TDT 0.6B v3 (multilingual, offline) for realtime under server_vad:
set it as both transcription and sound_detection and turn on
pipeline.diarization, and each committed turn gets speaker segments
and sound tags from one parakeet-cpp backend.

Files and sha256 match the Hub and are shared with the existing TDT v3,
diarization and CED-Tiny entries. A real-model spec checks the
combination on a clip with two speakers and a rooster: A-B-A-B speaker
turns, and "Chicken, rooster" among the sound tags. The test loader
now binds the sound entry points like main.go.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(gallery): add CED-Base variants of the parakeet-cpp scene models

parakeet-cpp-realtime-scene and parakeet-cpp-realtime-scene-tdt ship
with CED-Tiny. The -base variants use CED-Base (86M), which tags sounds
more confidently (on the rooster clip "Crowing" 0.65 against 0.49 for
Tiny).

Measured on CPU over a 37 s clip: the live diarization + sound stream
runs at 0.125 of real time with CED-Base against 0.103 with CED-Tiny,
because diarization dominates; sound detection per committed turn costs
0.031 against 0.005. The realtime docs list both and note that any CED
size works as sound_model.

Files and sha256 match the Hub and are shared with the existing
parakeet-cpp-ced-base entry. The TDT variant passes the real-model
scene spec with CED-Base (A-B-A-B speakers, "Chicken, rooster" found).

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(config): register pipeline.diarization in the config metadata

TestAllFieldsHaveRegistryEntries fails on the branch because the new
pipeline.diarization field has no registry entry. Add one so the model
editor shows it as a toggle next to the sound detection options.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-29 23:58:17 +02:00
Ettore Di Giacinto 9c156656bd Merge PR #12285: feat(failover): serve a model name from a chain of local and remote targets
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 15:32:30 +00:00