mirror of
https://github.com/mudler/LocalAI.git
synced 2026-09-12 22:33:54 -04:00
The claim queue's attempts counter grew without bound and nothing read it. At the default two-second poll a permanently undispatchable row cost about 43000 UPDATEs a day, and it cost more than writes: rows are claimed oldest first, so the oldest stuck row was re-claimed ahead of every newer one on every tick and held a dispatch slot while it failed. One poison row starved the queue behind it. No dead letter, and that is the decision rather than the omission. Read settleClaim: the only outcome that releases a claim is one where NOTHING was learned about the work. No agent worker was connected, the tunnel broke, a peer could not be reached, the stream was refused before the request body left this replica. Not one of those is a worker saying it ran the job and it failed, and an attempt ceiling would turn "the fleet was away long enough" into a job failure nobody reported, which is the collapse this whole design exists to prevent pointed at work instead of at nodes. The one verdict available here, that no build of any worker serves this kind, is already settled as an answer. So the retry stays unbounded and the RATE does not. Each release stamps the row with the earliest it may be claimed again, doubling from two seconds to a cap of sixty, computed in the release statement from the row's own attempts count and stamped on the DATABASE clock, because that is the clock competing replicas order the queue on. Queued work becomes claimable again within one cap of the fleet returning, and a stuck row no longer holds the head of the queue. A claim released by the reap carries no delay at all: that work was never handed to anyone, so there is nothing to back off from. Assisted-by: Claude Opus 5 [claude-code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>