mirror of
https://github.com/mudler/LocalAI.git
synced 2026-09-29 09:35:02 -04:00
fix(failover): free the leader lock soon after the leader's host dies
A leader whose host died without closing its connection kept the advisory lock for about two hours of OS keepalive defaults, and no other frontend could probe. The lock session now sets short TCP keepalives and tcp_user_timeout, so the server drops it within about 30 seconds. Shutdown now closes the lock for good, so a tick that runs after it cannot take the lock back. Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
This commit is contained in:
1 parent
b5792e4d17
commit
e59fb854ee
4 files changed
+175
-31
No files matched your search
@@ -187,8 +187,10 @@ share one failover state:
|
||||
all frontends converge on the same target for a chain.
|
||||
- One frontend, the probe leader, runs the health checks, decides fail-over
|
||||
and fail-back, and loads warm targets. The leader holds a PostgreSQL
|
||||
advisory lock and keeps it until it shuts down or its database connection
|
||||
fails. Then another frontend takes the lock and becomes the leader.
|
||||
advisory lock and keeps it until it stops or its database connection fails.
|
||||
Then another frontend takes the lock and becomes the leader: immediately
|
||||
when the leader shuts down or its process exits, and within about 30 seconds
|
||||
when the leader's host or network fails.
|
||||
- Warm targets stay loaded on the workers. The router and the replica
|
||||
reconciler treat them like pinned models and do not evict them.
|
||||
- A frontend that starts late gets the current state within 10 seconds,
|
||||
|
||||
Reference in new issue
Block a user