Files
LocalAI/docs/content/reference/system-info.md
T
localai-org-maint-botandEttore Di Giacinto c0b7e64973 feat(ui): show running models and host gauges on single-node installs (#12189)
On a single-node install nothing in Operate listed the models loaded on
this machine or let an admin stop one. The System page that did was
retired in #11548, and its replacements (the Nodes workbench) only work
in distributed mode. The Nodes page also mis-detected single-node mode:
the cluster routes are not registered there, so /api/nodes answers 404,
but only 503 was treated as "distributed off", which sent every
single-node install to the empty worker-registration card. The rail hid
the entry anyway.

Nodes route on a single node becomes "This machine":
- the Nodes page's VRAM / RAM / CPU / models-disk gauges, fed from this
  host by mapping /api/resources onto the worker heartbeat fields
- a memory bar splitting host RAM by running model
- a running-models table (backend, RSS, CPU share, uptime, PID) with
  search, sorting, logs and a confirmed Stop
- the distributed setup behind an "Add machines" button

The Operate overview gains a "Running now" preview (heaviest five, with
Stop) on single node and a pointer to Nodes > Running models on a
cluster. The rail shows "This machine" in Runtime with a running count.

Backend, additive only:
- /system: each loaded model carries a `process` block (pid, rss_bytes,
  memory_percent, cpu_percent, started_at). A sampler keeps one gopsutil
  handle per PID so CPU is the share since the previous poll rather than
  the lifetime average; it is omitted on the first reading.
- /api/resources: host `cpu` and models-path `disk`, the same readings
  workers send in their heartbeat.

Also fixes the fleet tables widening the page on phones: the headers'
absolutely positioned sr-only labels escaped the scroll wrapper.

Assisted-by: Claude:claude-opus-5 [Playwright]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-21 19:12:35 +02:00

3.0 KiB

+++ disableToc = false title = "System Info and Version" weight = 23 url = "/reference/system-info/" +++

LocalAI provides endpoints to inspect the running instance, including available backends, loaded models, and version information.

System Information

  • Method: GET
  • Endpoint: /system

Returns available backends and currently loaded models.

Response

Field Type Description
backends array List of available backend names (strings)
loaded_models array List of currently loaded models
loaded_models[].id string Model identifier
loaded_models[].backend string Backend serving the model, from its config. Omitted when the model was loaded without one
loaded_models[].process object The backend process serving the model on this host. Omitted when there is no local process (a distributed worker holds the model) or it could not be read
loaded_models[].process.pid integer Process ID
loaded_models[].process.rss_bytes integer Resident host memory, in bytes. Weights offloaded to a GPU are not included
loaded_models[].process.memory_percent number rss_bytes as a percentage of host RAM
loaded_models[].process.cpu_percent number Share of the whole host's CPU used since the previous call, 0-100. Omitted on the first call that sees the process, because there is no earlier reading to compare against
loaded_models[].process.started_at string When the process started (RFC 3339)

Usage

curl http://localhost:8080/system

Example response

{
  "backends": [
    "llama-cpp",
    "huggingface",
    "diffusers",
    "whisper"
  ],
  "loaded_models": [
    {
      "id": "my-llama-model",
      "backend": "llama-cpp",
      "process": {
        "pid": 48213,
        "rss_bytes": 5368709120,
        "memory_percent": 7.8,
        "cpu_percent": 42.5,
        "started_at": "2026-09-21T09:12:44Z"
      }
    },
    {
      "id": "whisper-1",
      "backend": "whisper"
    }
  ]
}

cpu_percent covers the time since the previous call to this endpoint, so a dashboard polling every few seconds gets a current reading. The WebUI's Operate → This machine page polls it every five seconds.


Version

  • Method: GET
  • Endpoint: /version

Returns the LocalAI version and build commit.

Response

Field Type Description
version string Version string in the format version (commit)

Usage

curl http://localhost:8080/version

Example response

{
  "version": "2.26.0 (a1b2c3d4)"
}

Error Responses

Status Code Description
500 Internal server error