select_zms() has three ways out, and two of them send a stream command before
the tail of the function sets started. streamCommand() drops anything sent while
that is false, so such a branch reports success having sent nothing and the
picture sits on its last keepalive frame.
The resume branch was fixed on this branch already. The branch below it, for a
page rendered with mode=paused, has the same shape and was missed. With auth on
it needed a still-valid hash to reach, so it was intermittent; with auth off,
where there is no hash that can go stale, the srcAuthCurrent change on this
branch makes it the path every initial load takes. Set started there too.
Add tests/js/monitorstream-resume.test.js, which asserts what reaches the wire
rather than which branch ran: the resume and initial-paused paths each send
exactly one CMD_PLAY on the connkey they are supposed to address, resuming
leaves src and connkey alone, and a stale auth hash still rebuilds and quits the
process the old connkey addressed. Removing the one-line fix fails the
initial-paused case and leaves the other three passing.
Full JS suite passes, ESLint clean on both files.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkQwahn9pi1y4wJe9BTxjM
select_zms() has a branch that resumes an existing zms with CMD_PLAY rather
than rebuilding the stream, but three things stopped it ever working. Every one
of them showed up only when the page was driven for real.
stop() ends by clearing activePlayer, which is the only thing that branch
tests, so after any stop it was unreachable and start() fell through to "new
src, new connkey" - a second zms, the first left running and no longer
addressable. That is the montage page's accumulating processes: hide the tab,
the handler stops each stream, and returning replaces rather than resumes. The
same happens when a monitor is scrolled out of view and back. stop() now
remembers what it shut down, and select_zms() resumes on that as well as on
activePlayer, provided we still hold the connkey to address it.
srcAuthCurrent required a non-empty zmAuth.hash, which is '' whenever auth is
off or the relay carries no hash. There is nothing that can go stale in that
case, so the src is as current as it will ever be; requiring a hash sent every
such install down the rebuild path for no reason.
streamCommand() drops anything sent while !started, and started is not set
until the end of select_zms(), so the resume issued its CMD_PLAY into nothing
and the stream stayed stopped. Resuming after a pause worked only because
pause() leaves started set. It is now set before the command goes out.
Finally, the rebuild path clears the "Loading..." info block from img_onload,
which cannot fire on a resume because src never changes. Without clearing it,
a stream that had in fact resumed sat behind that block and its still image and
looked frozen - the fault that made this look unfixable at first.
restart() is excluded from all of it: it is the error path, whatever failed may
be that very zms, and a broken img is not repaired by CMD_PLAY.
Verified on a live montage page with two monitors: across two hide/show cycles
the same two zms processes are kept - no orphans, no respawn - and both
pictures are live afterwards, with the zms status reporting stopped=0 and
~15 fps.
refs #4706
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
select_zms() mints a fresh connkey whenever it has to rebuild the stream src,
without telling the process the old connkey belonged to that it is finished.
Once the key is replaced nothing can reach that process again: no CMD_QUIT can
be delivered, and a stopped one will not notice on its own.
getStreamCmdResponse() already learned this - its reload path calls
quitConnKey() first, with a comment saying why - but the path every ordinary
start() takes did not. quitConnKey had exactly two references in the file: its
own definition and that one use.
refs #4706
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
-- Check isActive first.
-- Run error registration
- When checking for errors in nextCycleView(), check monitorStream.player instead of (monitorStream.selectedPlayer || monitorStream.activePlayer)
- When executing MonitorStream.restart(), do not clear the player's error count. This will allow us to analyze playback issues. We do this clearing elsewhere.
- Instead of Math.random(), we'll use Date.now().
This will make the code cleaner and more understandable.
- We also now use refreshStreamSrc() when stopping stream playback.
- When closing a socket, check the session before clearing the WebSocket, not after.
- Execute resetCountStreamErrors() when playback starts successfully.
- Recheck the session inside videoEl.onplay after the 500ms timeout.
- Exit this.img_onload if the monitor is already stopped.
- Assign the 'zm:tracksReceived' listener only after stream playback has started, not during start().
- Additional processing of MonitorStream.playbackSessionId
- Calling stream.load() inside stop() for the rtsp2web type MSE is redundant and can lead to unnecessary warnings. We do everything necessary in stopMse().
- When restarting a stream, call updatePlayerControls() instead of streamCmdStop() to avoid duplicating the stop() command (relevant for the Watch page).
- Now, to switch to the next player, we won't call MonitorStream.selectNextPlayer() in different places in the code. Instead of selectNextPlayer(), we will always call restart(), which will select the next player. This will allow us to more accurately clear all unnecessary objects. To achieve this, we have slightly modified restart(). - For go2rtc, execute playbackSessionId = generateUUID() when changing the SRC, not when selecting a player, since go2rtc itself can change src.
- Avoid looping when executing selectNextPlayer() during errors in the ZMS player.
- Added the "fatal" argument to streamErrorRegistration(), which prevents unnecessary relaunches of the same player that was previously playing in the event of a fatal playback error.
- Added a setter for src to the VideoStream class for go2rtc.
- Added support for monitorStream.isActive for the Watch page.
- Added the updatePlayerControls() function for the Watch page to manage the state of the player buttons. In the future, we need to create a single function for managing the state of the player buttons. It's a bit of a mess right now. We'll do this in the next PR to avoid breaking anything.
- To optimize performance, part of the code from getTracksFromStream() has been moved directly to the listener in MonitorStream.
MonitorStream.kill() unconditionally did `stream.onerror = null` and
`stream.onload = null` on whatever element the monitor was using. That was
written for the zms <img>, whose onerror/onload are inherited event-handler
accessors.
With go2rtc the element is <video-stream>, where onerror is a method on
VideoRTC.prototype (video-rtc.js) overridden by VideoStream (video-stream.js).
Assigning null there finds a writable data property on the prototype chain and
so creates an *own* property on the instance, shadowing the method for as long
as the element lives. replaceDOMElement() returns the same node when the tag
already matches, so select_go2rtc() handed the poisoned element back on the
next start, and the listener VideoRTC.onconnect() registers,
this.ws.addEventListener('error', (ev) => this.onerror(ev));
threw "TypeError: this.onerror is not a function" on the next websocket
failure. Any kill()-then-start() path reached it: switching monitors on watch,
the stop/play buttons, montage viewport handling. It also meant the restart
that VideoStream.onerror performs was silently dead after a kill().
Guard the assignments on the element actually being an IMG.
Add tests/js/monitorstream-kill.test.js, which evaluates the real
MonitorStream.js in a vm context and checks that kill() leaves a prototype
onerror callable on a <video-stream>, adds no own onerror/onload to it, and
still clears both on an <img>.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MXtwyhzssj24Xwjmx8EA2z
The page kept two copies of the same secret: auth_relay, the query fragment
every AJAX call is authenticated with, and auth_hash, the bare hash stamped
into stream <img> URLs. Different responses updated different copies, so
they could drift, and a drifted auth_hash produced stream URLs that zms
rejects.
ZMAuth stores only the relay and derives the hash from it, so the two cannot
disagree. Its helpers cover the four shapes the call sites used:
zmAuth.hash derived, '' under the plain/none relay forms
zmAuth.update(data) absorb the auth fields of any response
zmAuth.appendTo(url) authenticate a url, no-op when auth is off
zmAuth.applyTo(src) point a stream url at the current credential
appendTo also removes the `x ? '&'+x : ''` guard repeated at every call
site, some of which had omitted it and emitted a dangling '?'.
Migrates all call sites across web/js and the classic skin, and drops both
globals from skin.js.php.
Tests: tests/js/auth-helpers.test.js, 44 passing.
getStreamCmdResponse() responded to every ajax/stream.php failure the same way:
mint a fresh connkey and reload the img src. ajaxError() returns HTTP 200 with
result=Error, so these arrive in jQuery's done() rather than fail(), and all
twelve error paths in stream.php took that branch.
Only one of them means zms is gone. For the rest the process is still running
and streaming, and replacing the connkey makes it unaddressable: CMD_STOP,
CMD_QUIT and mode=single all then go to the new key, so nothing can reach the
old process and only SIGPIPE can stop it, which we know is unreliable. That is
why the reports of lingering zms after switching monitors were unaffected by
changes to what the stop path sends.
The timeout path made this routine rather than rare. On select() expiry
ajaxError is commented out, so the script carries on to socket_recvfrom() on a
now non-blocking socket. That returns false, and false == 0 under switch's loose
comparison, so a merely slow zms was reported as 'No data to read from socket'
and torn down.
stream.php now classifies each failure as no_socket, timeout, transient or
invalid, and sends it as 'reason'. The client restarts the stream only for
no_socket. A missing reason is still treated as fatal, so a php that predates
this keeps the old behaviour.
Before replacing the connkey the client now sends CMD_QUIT to the old one, so
the process we are about to lose track of is asked to exit. That is deliberately
not routed through streamCommand(): it must name its target explicitly, since
this.connKey is about to change, and its response must not feed back into
getStreamCmdResponse(), or a QUIT that also failed would re-enter the error path
and loop.
ajaxError() takes the classification as a third argument, named $reason because
$code is already the HTTP status, and only includes it when set, so the other
131 callers are unaffected.
Tests: tests/js covers the fatal/non-fatal decision including the no-reason
fallback, tests/php pins the classification mapping and the switch(false)
semantics the timeout branch depends on. Both verified to fail when the
behaviour is reverted.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- If a video track is missing when it is required, we generate the event:
dispatchTracksReceived(videoFeedStream, {
status: 'aborted',
reason: 'playback-videoTrack-missing'
});
kill() cleared this.started before calling this.stop(), but stop() returns
early when !started. For the zms path that meant clearInterval() on
statusCmdTimer and streamCmdTimer never ran, activePlayer was never reset and
mediaStream/audioTrack/videoTrack were never released. Every kill() leaked a
pair of intervals, which adds up over a montage or watch page that cycles
monitors every few seconds.
Keep started set until stop() has done its work, and pass skipStreamCommand so
stop() doesn't follow CMD_QUIT with a CMD_STOP against a socket zms is already
tearing down. Clear connkey afterwards.
stop() already sets started=false and activePlayer='' at the end, so kill()
doesn't need to.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
$.ajaxQueue defers the $.ajax() call until earlier queued requests
finish, and jQuery serialises the data object at that later point.
streamCmdReq() passed its data object by reference, so callers that
hand it this.streamCmdParms directly had their command overwritten by
the streamCmdQuery timer setting command=CMD_QUERY before the request
was actually sent.
show_analyse_frames() is the only such caller, which is why the Show
Analysis button in the live view appeared to do nothing: zms received
CMD_QUERY instead of CMD_ANALYZE_ON and never switched to
FRAME_ANALYSIS.
Copy the params inside streamCmdReq() so every caller is covered.
streamCommand() and alarmCommand() already copied by hand.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>