fix(failover): free the leader lock soon after the leader's host dies

A leader whose host died without closing its connection kept the
advisory lock for about two hours of OS keepalive defaults, and no other
frontend could probe. The lock session now sets short TCP keepalives and
tcp_user_timeout, so the server drops it within about 30 seconds.

Shutdown now closes the lock for good, so a tick that runs after it
cannot take the lock back.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
This commit is contained in:
Ettore Di Giacinto committed 2026-09-27 07:42:20 +00:00
1 parent b5792e4d17
commit e59fb854ee
4 files changed
+175 -31

No files matched your search

+4 -2
View File
@@ -187,8 +187,10 @@ share one failover state:
all frontends converge on the same target for a chain.
- One frontend, the probe leader, runs the health checks, decides fail-over
and fail-back, and loads warm targets. The leader holds a PostgreSQL
advisory lock and keeps it until it shuts down or its database connection
fails. Then another frontend takes the lock and becomes the leader.
advisory lock and keeps it until it stops or its database connection fails.
Then another frontend takes the lock and becomes the leader: immediately
when the leader shuts down or its process exits, and within about 30 seconds
when the leader's host or network fails.
- Warm targets stay loaded on the workers. The router and the replica
reconciler treat them like pinned models and do not evict them.
- A frontend that starts late gets the current state within 10 seconds,