mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-01 10:34:36 -04:00
The three NATS queue groups jobs.new, jobs.mcp-ci.new and agent.execute are gone. Dispatching work is now a row in a work_claims table, taken by one frontend replica with SELECT ... FOR UPDATE SKIP LOCKED and driven on an agent worker as a streaming control RPC over that worker's tunnel. Exactly-one delivery among competing consumers is a database problem, not a broker feature. An agent worker has no database, so it never claims; it executes what the claiming replica hands it. A claim must not outlive the replica that took it. The reap releases a claim whose owner is no longer a live replica in the instances table, on the database clock, and never asks how long the claim has been held. A job that legitimately runs for an hour on a heartbeating replica is left alone, while a claim whose owner stopped heartbeating becomes claimable again within one liveness window. A replica with no advertised address has no instances row at all, so it refuses to claim rather than have its work reaped out from under it mid-run. The settle rule is stated once, in settleClaim, and every exit path calls it. A transport failure releases the claim and never completes or discards it; only a decoded reply line completes it. That line is deliberately not cluster.IsWorkerAnswer, which accepts the stream refusals a worker's tunnel writes before any request body reaches its control server: completing on those would discard work that never ran. The terminal line is persisted before the claim is completed, so a store that refuses leaves the claim standing rather than leaving the job running for ever. That is the dropped-result defect fixed structurally rather than by retry. This also surfaces a pre-existing gap rather than causing one: no worker has ever served plain task jobs, and publishing them into an empty queue group left them running with no trace. Such a claim is now failed with a reason. Removes QueueWorkers, --agent-subject and --agent-queue, and narrows an agent worker's minted JWT by agent.execute and jobs.mcp-ci.new. Assisted-by: Claude Opus 5 [claude-code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
68 lines
2.7 KiB
Go
68 lines
2.7 KiB
Go
package natsauth
|
|
|
|
import "strings"
|
|
|
|
// workerSubjectToken mirrors messaging.sanitizeSubjectToken without importing unexported logic.
|
|
func workerSubjectToken(nodeID string) string {
|
|
r := strings.NewReplacer(".", "-", "*", "-", ">", "-", " ", "-", "\t", "-", "\n", "-")
|
|
return r.Replace(nodeID)
|
|
}
|
|
|
|
// WorkerPermissions returns NATS pub/sub allow lists for a registered node.
|
|
//
|
|
// It serves AGENT nodes. They are the only workers left that connect to the
|
|
// bus: an agent worker subscribes to the queue subjects listed below, while a
|
|
// backend worker connects to no bus at all, because every verb a frontend gives
|
|
// it is an HTTP route on its own server reached through its outbound tunnel
|
|
// (core/services/workerctl).
|
|
//
|
|
// The non-agent branch is therefore a grant of nothing, and it has to be
|
|
// spelled that way rather than deleted. NATS reads an EMPTY allow list as no
|
|
// restriction, so a function that returned nil here would upgrade every JWT the
|
|
// frontend still mints for a backend node from "its own inbox" to "the entire
|
|
// account". The inbox is self-scoped and reaches no cluster subject.
|
|
//
|
|
// nodeID no longer narrows anything: no allow list below is per-node, because
|
|
// the only per-node subject an agent worker ever subscribed to was its
|
|
// backend.stop, which is a control RPC on its tunnel now. It stays in the
|
|
// signature so a future per-node grant has somewhere to come from, and because
|
|
// mint.go names the JWT user after it.
|
|
func WorkerPermissions(nodeID, nodeType string) (pubAllow, subAllow []string) {
|
|
switch nodeType {
|
|
case "agent":
|
|
// Keep this list in sync with the subscriptions in core/cli/agent_worker.go.
|
|
//
|
|
// MCP tool execution, discovery, agent execution, MCP CI runs and the
|
|
// per-node backend.stop are all absent: every one of them is a control
|
|
// RPC on the tunnel the worker holds, addressed by the frontend rather
|
|
// than by a subject. The last two left when the queue groups became
|
|
// claim rows on the job store.
|
|
//
|
|
// Removing them narrowed this list; it must never be narrowed to
|
|
// nothing, because NATS reads an EMPTY allow list as no restriction at
|
|
// all, which would widen an agent JWT to the whole account.
|
|
subAllow = []string{
|
|
"agent.*.cancel",
|
|
"gallery.*.cancel",
|
|
"gallery.*.progress",
|
|
"jobs.*.cancel",
|
|
"jobs.*.progress",
|
|
"jobs.*.result",
|
|
"staging.*.progress",
|
|
"_INBOX.>",
|
|
}
|
|
pubAllow = []string{
|
|
"agent.>",
|
|
"jobs.>",
|
|
"_INBOX.>",
|
|
}
|
|
default:
|
|
// Backend worker: nothing, held open at its own inbox for the reason in
|
|
// the doc comment. The node subtree it used to subscribe on went with
|
|
// the connection itself, which this worker no longer opens.
|
|
subAllow = []string{"_INBOX.>"}
|
|
pubAllow = []string{"_INBOX.>"}
|
|
}
|
|
return pubAllow, subAllow
|
|
}
|