The unit test only covers the backoff schedule. The retry itself cannot be
reached in-process: the errors come from InnoDB choosing a deadlock victim
or timing out a lock wait, and nothing can fake that. These drive it for
real, tagged [notCI] alongside the existing socket tests since they need a
reachable database and credentials in /etc/zm/zm.conf.
Three cases, each against a scratch table the fixture creates and drops. No
ZoneMinder table is read or written.
- Riding out a lock another session holds. Session lock wait is set to 1s
and the lock is held for 2.5s, so the update only lands if it is retried.
This one does NOT discriminate: the old code retried lock waits too, just
immediately and forever. It is here to catch a regression, not to
demonstrate the fix.
- Giving up on a lock nobody releases. Verified to hang indefinitely
against the old semantics, killed at 45s; it now returns inside the
budget. This is the case the bound exists for.
- A real deadlock. Both sessions reach for a row the other holds. InnoDB
weights victim selection by rows changed, so the other session dirties a
hundred rows first to make the ZoneMinder connection the cheaper one and
therefore the victim; without that the victim varied per run and the test
passed without ever reaching the retry, which is how the first version of
it was wrong. Verified 12/12 pass with the retry and 12/12 fail without.
The assertion that carries this last one is the row value, not the return
code: when the other session wins the toss its transaction is discarded and
only one increment survives, so checking that both landed is what pins the
retry actually happening.
Full non-live suite unchanged at 12171 assertions in 133 test cases. Live
suite 36 assertions in 3 cases, run five times consecutively. Scratch table
confirmed dropped afterwards.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Y6FieTwEXuLhhR4e2yiax