Files
LocalAI/docs/content/operations/overview.md
T
localai-org-maint-botandEttore Di Giacinto c0b7e64973 feat(ui): show running models and host gauges on single-node installs (#12189)
On a single-node install nothing in Operate listed the models loaded on
this machine or let an admin stop one. The System page that did was
retired in #11548, and its replacements (the Nodes workbench) only work
in distributed mode. The Nodes page also mis-detected single-node mode:
the cluster routes are not registered there, so /api/nodes answers 404,
but only 503 was treated as "distributed off", which sent every
single-node install to the empty worker-registration card. The rail hid
the entry anyway.

Nodes route on a single node becomes "This machine":
- the Nodes page's VRAM / RAM / CPU / models-disk gauges, fed from this
  host by mapping /api/resources onto the worker heartbeat fields
- a memory bar splitting host RAM by running model
- a running-models table (backend, RSS, CPU share, uptime, PID) with
  search, sorting, logs and a confirmed Stop
- the distributed setup behind an "Add machines" button

The Operate overview gains a "Running now" preview (heaviest five, with
Stop) on single node and a pointer to Nodes > Running models on a
cluster. The rail shows "This machine" in Runtime with a running count.

Backend, additive only:
- /system: each loaded model carries a `process` block (pid, rss_bytes,
  memory_percent, cpu_percent, started_at). A sampler keeps one gopsutil
  handle per PID so CPU is the share since the previous poll rather than
  the lifetime average; it is omitted on the first reading.
- /api/resources: host `cpu` and models-path `disk`, the same readings
  workers send in their heartbeat.

Also fixes the fleet tables widening the page on phones: the headers'
absolutely positioned sr-only labels escaped the scroll wrapper.

Assisted-by: Claude:claude-opus-5 [Playwright]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-21 19:12:35 +02:00

4.4 KiB

+++ title = "Operate overview" weight = 1 +++

/app/operate is the front door to the Operate console. It answers one question — is anything wrong — without you having to open four other pages.

Needs attention

The block the page exists for. It lists only things that want a decision:

  • a backend with an update available
  • an operation that failed
  • a node reporting unhealthy

When nothing needs attention it says so in one line and renders nothing else. There is no green panel: a status page that shouts when everything is fine teaches you to stop reading it.

Headline totals

Requests, failed requests and p95 latency over the last 24 hours, each with a sparkline of the trend. These come from GET /api/traces/summary, which counts the trace buffer server-side:

curl http://localhost:8080/api/traces/summary?hours=24 \
  -H "Authorization: Bearer <admin-key>"
{
  "total": 18402,
  "errors": 37,
  "p95_ms": 842,
  "window_hours": 24,
  "buckets": [{ "start": "2026-08-02T09:00:00Z", "count": 1520, "errors": 3 }]
}

hours defaults to 24 and is capped at 168. Only 5xx responses and transport errors count as failures — a 4xx is the caller getting it wrong, not the installation being unhealthy. p95_ms is a nearest-rank percentile, not the slowest request.

The endpoint exists so a dashboard wanting three numbers does not fetch the whole trace list to count it. An installation that has served nothing yet says so rather than showing three zeroes dressed as telemetry.

Host capacity

The overview also shows the host's current RAM or GPU capacity, utilization, and model storage. Loading, unavailable, and empty states are explicit. This uses the same 15-second Operate summary poll as the rail and attention data, so opening the overview does not start a second resource poller.

Running now

On a single-node install the overview lists the models loaded on this machine, heaviest first, up to five. Each row shows the backend, resident memory, CPU share and uptime, with View logs and Stop model… in the row menu. Open this machine leads to the full list.

With distributed mode on, models run on workers rather than on the controller, so this section links to Operate → Nodes → Running models instead.

This machine

On a single-node install, Operate → This machine (/app/nodes) shows the host and everything loaded on it:

  • Capacity gauges for VRAM, RAM, CPU and the models disk, the same gauges the Nodes page draws for a cluster. A host without a GPU says so rather than showing an empty VRAM gauge.
  • A memory bar splitting host RAM by running model, so you can see which model is holding memory.
  • Running models: search, sort by memory, CPU or uptime, open a model's logs, or stop it. Stopping asks for confirmation; the model loads again on its next request.

The page polls GET /system and GET /api/resources every five seconds. The per-model readings come from the process block of [GET /system]({{% relref "reference/system-info" %}}); the host CPU and disk readings come from the cpu and disk fields of GET /api/resources.

Add machines reveals the command to start LocalAI in distributed mode. Once distributed mode is on, the same route becomes the Nodes page and the rail entry moves to the Cluster group.

Models and backends no longer live under a nested Host page. Use Models → Installed for model runtime and configuration actions, and Operate → Backends → Installed for installed backend actions. The overview links into the canonical Operate sections rather than duplicating those inventories.

Old /app/manage bookmarks remain supported. They redirect with replace semantics to the matching Installed Models or Installed Backends view while preserving legacy search, filter, selection, variant, and development flags.

The rail

The Operate rail groups its destinations under four headings — Runtime, Cluster, Observability and Administration — and shows a live value beside several of them: pending backend updates, running operations, healthy node count, request volume and error count. Host capacity lives on the overview instead of appearing as a separate destination.

Those values are orientation, not an alarm. The rail only exists on Operate routes and can be collapsed, so anything urgent also appears in Needs attention and on the operations badge attached to the sidebar entry, which is always visible.