Commit Graph
217 Commits
Author SHA1 Message Date
Isaac ConnorandClaude Opus 5.5 82901ac1fb fix: schedule ONVIF subscription renewal for short and unreported lifetimes (#5181)
* fix: schedule ONVIF subscription renewal for short and unreported lifetimes refs #5179

The Beward camera from #5165 grants a 60 second subscription and its
RenewResponse carries no TerminationTime. Three things combined so that
ZM renewed after every PullMessages that returned messages (61 renewals
in 74 seconds of the reporter's log):

- The renewal was scheduled a fixed ONVIF_RENEWAL_ADVANCE_SECONDS (60)
  before termination, which for a 60 second subscription is its creation
  time. ONVIFNextRenewalTime() now uses the smaller of that advance and
  half the remaining lifetime.
- Renew() only logged a missing TerminationTime and left
  next_renewal_time in the past. assume_renewal_times() now schedules
  from the lifetime we asked for, capped at what the camera last granted
  (ONVIFAssumedLifetime()), so this camera renews every 30 seconds.
- IsRenewalNeeded() was only checked after a response carrying messages,
  so a camera quiet for longer than its subscription was never renewed.
  It is now also checked after an empty (SOAP_EOF) poll.

Since renewal is now checked after every poll, a camera that answers
Renew with ActionNotSupported has renewal disabled rather than being
asked again on each poll, matching the existing handling of unusable
TerminationTimes.

Tests cover the renewal time for long, short and very short
subscriptions, and the assumed lifetime with no, smaller and larger
previously granted lifetimes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: only disable ONVIF renewal for ActionNotSupported, keep sub-second grants refs #5179

Review follow-ups on the renewal scheduling change:

- Renew() treated soap->error == 12 as ActionNotSupported, but 12 is
  gSOAP's generic SOAP_FAULT. With renewal now disabled on that path, a
  NotAuthorized or InvalidArgVal fault from Renew would have stopped
  renewal for good while marking the subscription healthy.
  ONVIFIsActionNotSupported() checks the fault subcode and string for
  ActionNotSupported (wsa: or ter:); any other fault now takes the
  existing cleanup and re-subscribe path.
- granted_lifetime truncated the remaining time to whole seconds. A
  termination less than a second away stored 0, which reads as "never
  reported", so the next Renew without a TerminationTime assumed the full
  requested 300 seconds. ONVIFGrantedLifetime() rounds up instead.

Tests cover ActionNotSupported subcodes and strings against other faults
and non-fault results, and the granted lifetime for whole, fractional and
sub-second remainders.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: count an assumed ONVIF renewal lifetime from the request, not the response refs #5179

When a RenewResponse has no TerminationTime, assume_renewal_times()
started the assumed lifetime when the response arrived, but the camera
counts it from when it received the request. A slow response pushed the
deadline and the next renewal later by the response time: with
subscription_timeout=10 and a 6 second response, the camera's deadline
was request+10 but the renewal was scheduled for request+11. In absolute
renewal mode it also ignored the exact deadline that was sent.

Renew() now records the time just before building the request and the
deadline it asks for: request time plus subscription_timeout, or the
absolute whole-second time it sends. ONVIFAssumedTermination() replaces
ONVIFAssumedLifetime() and returns that deadline, capped at request time
plus the lifetime the camera last granted. The renewal is scheduled from
the request time; if the response arrived after it, it is already due and
the next poll renews.

Tests cover no, smaller and larger previous grants, an absolute deadline
kept exactly, and the slow-response case from review.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-04 12:33:00 -04:00
Isaac ConnorandClaude Opus 5.5 d7238f745a fix: stop busy-polling ONVIF cameras that don't hold PullMessages refs #5180 (#5182)
ZM asks the camera to hold each PullMessages for up to pull_timeout
seconds (default 1). The Beward camera from #5165 answers in 10-30ms with
no messages, which gSOAP reports as SOAP_EOF, and Run() issued the next
request at once: ~2,500 PullMessages in 74 seconds of the reporter's log
(~34/s), each with a fresh UsernameToken digest.

Time each PullMessages. When it comes back with no messages (SOAP_EOF, or
a SOAP_OK response with no NotificationMessage) before the Timeout,
WaitForMessage() now waits out the remainder in 100ms steps, staying
responsive to terminate_ and zm_terminate. A camera that ignores the
long-poll is then polled at most once per pull_timeout; one that holds it
is polled again immediately as before. Events raised during the wait stay
queued in the camera's pull point and arrive with the next request, so
the added latency is at most pull_timeout.

ONVIFEarlyPollWait() computes the remainder and is tested for an
immediate answer, a full hold, a partial hold and sub-millisecond
elapsed times.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-04 11:03:24 -04:00
Isaac ConnorandClaude Opus 5.5 98c5d535e1 fix: ignore an ONVIF PullMessages TerminationTime that is not in the future refs #5165 (#5178)
Per-topic alarm expiry takes each alarm's deadline from the
PullMessagesResponse TerminationTime. Beward cameras send their
CurrentTime as the TerminationTime in every response, so an alarm was
stored with a deadline that had already passed and the sweep at the end
of the same WaitForMessage() pass removed it about 25us later. The
analysis thread checks onvif->isAlarmed() once per frame, never saw the
monitor alarmed, and no event was recorded even though the log showed
"ONVIF Triggered Start Event".

Move the TerminationTime handling into a free function,
ONVIFAlarmTermination(), which applies the camera clock offset as before
and returns false when the adjusted time is not after now. The alarm then
gets no expiry and is cleared by the camera's explicit State=false
message, as it was before per-topic expiry was added. The Debug line now
also prints CurrentTime so this camera behaviour is visible in logs.

Tests cover TerminationTime equal to and before CurrentTime, a future
TerminationTime, a skewed camera clock, a missing CurrentTime keeping the
previous offset, and a missing TerminationTime.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-04 09:52:58 -04:00
Isaac ConnorandClaude Opus 5 8694032b49 fix: validate Range requests in view_video instead of trusting them refs #5174
view_video.php took the end of the range straight from the request and never
checked it against the file:

  if (!empty($matches[2])) $end = intval($matches[2]);
  $length = $end - $begin + 1;

so on a 1000 byte file:

  bytes=0-99999    206, Content-Length 100000, body 1000 bytes
  bytes=5000-6000  206, Content-Range naming bytes that do not exist
  bytes=500-100    206, Content-Length -399
  bytes=1000-      206, Content-Length 0
  bytes=-100       200 with the whole file, not the last 100 bytes

A client that is told to expect 100000 bytes and gets 1000 does not see a bad
request, it sees a truncated file, and reports the video as broken. The suffix
form was not recognised at all because the pattern required a digit before the
dash.

This is reached once per fragment by the byte-range HLS manifest VideoStore
writes, every fragment being a Range against the one mp4, so a player that
asks for anything the file cannot supply gets a body that does not match its
own Content-Length rather than an answer it can act on.

Parse the header properly: clamp a range that runs past the end, because a
client may ask for more than is there and is entitled to what is there;
answer 416 with "Content-Range: bytes */size" when the range cannot be
satisfied at all, so the client learns the real length; and read "-N" as the
last N bytes. Length is now derived from the range being served rather than
the one requested, and the send loop counts down by the bytes it actually
read, so Content-Length and the body cannot disagree.

Only the first range of a multi-range request is served, as before. A
multipart/byteranges body is not worth building for this, and falling back to
sending the whole representation is not an option when these are event videos
of hundreds of megabytes; Content-Range names exactly what was sent.

The parsing is its own dependency-free include so it can be tested without a
database, and tests/php/test_http_range.php covers each case above plus a
sweep asserting that every range it ever returns lies inside the file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-10-04 08:26:42 -05:00
Isaac ConnorandClaude Opus 5 0bee8990be test: pin that manifest byte ranges cover whole fragments refs #5174
#5174 reports an event's m3u8 as invalid and proposes moving every fragment's
byte range 8 bytes further into the file, from @17203 to @17211.

That reading comes from ffprobe's trace, whose second column is the offset of
the box *body* -- the start plus the 8 byte box header -- not the start. In
the manifest quoted there the first fragment is 1434876@17203, and the trace
has the moof body at 17211 and the mdat ending at 1452079, so the range spans
exactly moof(264) + mdat(1434612) = 1434876 from the start of the moof. The
proposed change would cut the moof header off every segment.

The manifest is not contiguous, which is what draws the eye: the init range
is ftyp+moov and the fragments start 16KB later, because reserve_region puts
the leading sidx in between and the sidx has to END where the fragments
BEGIN. Nothing fetches those bytes and HLS does not require byte ranges to
abut.

So assert it rather than argue it. Against sidx-moof.mp4, using the box
walker the sidx tests already keep for the purpose, check that a fragment
starts at its moof box rather than 8 bytes in, that fragments run back to
back, that the init range is the header boxes and nothing else, and that the
gap between them is exactly the reserved index.

This says nothing about the browser error that prompted the report, which is
not quoted in the issue; it only rules the byte ranges in or out.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-10-04 08:26:42 -05:00
Isaac ConnorandClaude Opus 5 a97cc6ab51 fix: append to the HLS manifest instead of rewriting it per fragment
writeM3U8 truncated and rewrote the whole playlist every time a fragment
completed. That is O(fragments) of writing per fragment, so an event pays
O(fragments squared) overall. For the ten minute events a section length
produces the manifest is about 55KB and the cost is invisible. It stops being
invisible when an event does not close.

On a box here, four monitors stopped closing their events after a reboot and
ran for 51 hours. Their manifests reached 17MB and 470,000 lines, and strace
showed where the writes were going: in one window the mp4 took 3 writes while
index.m3u8 took 195, opened O_WRONLY|O_CREAT|O_TRUNC. The cameras produced
0.4MB/s of video between them and the disk was absorbing 24MB/s at 100%
utilisation, 106ms average write latency.

That is a loop rather than just waste. One rewrite took 3.48s of wall clock
against a fragment arriving every 1.2s, so the event thread could never catch
up, the packet queue stayed full ("Analysis is not keeping up" every three
seconds), and the analysis thread is where the section length check that would
have closed the event lives. The growth starved the only thing that could stop
it.

An EVENT playlist is append only: a fragment's three lines never change once
written, and only the header depends on anything global. So write the new
fragments to the end, and fall back to a full rewrite when something above
them would differ -- a changed target duration or init segment end, the
different url the close path passes, a different path, or the closing
ENDLIST. The remembered byte count is checked against the file before
appending, so a manifest that something else has truncated, replaced or
removed is rebuilt rather than appended to; any failure zeroes the state,
which makes a rewrite the answer to anything unexpected.

The header and fragment text are split into m3u8Header, m3u8Fragment and
m3u8TargetDuration so the property that matters can be tested without
standing up a VideoStore: appending one fragment at a time produces a byte
identical manifest to writing the whole thing at once.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-10-04 07:49:12 -05:00
Isaac ConnorandClaude Opus 5.5 c4ef3fdd62 feat: key Frames by (EventId, FrameId) and drop the surrogate Id refs #5164
Every query on Frames filters on EventId and orders or ranges by FrameId,
and nothing needs a frame by Id. The Id primary key was a clustered index
nothing read, and EventId_FrameId_idx was a secondary index every query
used. On a copy of a local table that secondary index was half the
table's size. With (EventId, FrameId) as the primary key the index is
gone, and an event's rows are stored together, so deleting an event is a
range delete.

zm_update-1.39.36.sql:
- converts AI_Detections.FrameId from Frames.Id to the per-event frame
  number, drops its foreign key to Frames and indexes (EventId, FrameId).
  A composite foreign key cannot replace it: ON DELETE SET NULL would
  have to null the NOT NULL EventId, and Frames rows are written in
  batches, so a detection can be recorded before its frame row.
- removes duplicate (EventId, FrameId) rows, keeping the earliest.
- rebuilds Frames with the new primary key.
- removes ON UPDATE CURRENT_TIMESTAMP from Frames.TimeStamp. Any UPDATE
  of a frame row was overwriting its capture time.
Each step checks the current schema first, so the migration can be
re-run.

REST API: view, edit and delete take /frames/<action>/<EventId>/<FrameId>.json.
The old single-Id URLs return 404. CakePHP 2 has no composite keys, so
the model's primaryKey is EventId. That keeps Event's dependent cascade
delete limited to the event's own frames. The controller writes with
explicit (EventId, FrameId) conditions instead of save(), which would
match rows on EventId alone. Edit no longer changes EventId or FrameId.

view=image with fid but no eid used to look up Frames.Id. It now returns
404.

The Perl Frame class is identified by (EventId, FrameId).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 11:48:26 -04:00
Isaac ConnorandClaude Opus 5.5 b79239940f fix: scope the frames list search to its event and match on FrameId refs #5164
ajax/frames.php collected Frames.Id values from the event's rows and then
filtered with a FrameId IN (...) term, which FilterTerm translated to
Id IN (...). The search query had no EventId condition of its own. Once
the term matches on FrameId, which is only unique within an event, it
returns rows from every event, so the query is now anchored on the event
with the same WHERE EventId clause as the unfiltered list.

FilterTerm now maps FrameId to the FrameId column. The frames list
thumbnail links address the image by eid and fid instead of Frames.Id.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 11:48:26 -04:00
Isaac ConnorandClaude Opus 5.5 3d858e2ccb refactor: find the mfra in finalize() with zm_mp4::media_end
finalize() located the mfra by reading the trailing mfro with fopen and
fseeko and decoding it by hand, the same job zm_mp4::media_end() does for
the sidx scan. Use media_end() for both, so the last HLS fragment and the
index agree on where the media ends.

media_end() is also stricter: it requires an mfra box of the stated size
at the offset the mfro points to, where the old code accepted any size up
to the file length. A trailer that does not check out now leaves the final
fragment running to EOF, as a missing trailer already did.

A new test covers media_end() against the fixture's real mfra and three
damaged trailers: an mfro size off by four, one larger than the file, and
no mfro at all.

refs #5144

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 07:46:15 -04:00
Isaac ConnorandClaude Opus 5.5 26e35beed9 fix: mark sidx references as SAPs only when their fragment starts on one
build_sidx_region() wrote starts_with_SAP=1, SAP type 1 into every
reference. That holds for frag_keyframe recordings, where every fragment
opens on a keyframe, but a monitor whose encoder options add frag_duration
or frag_size gets fragments cut mid-GOP, and the index then claimed a
decodable start where there was none.

scan_fragments() now reads the sync flag of each fragment's first video
sample: trun first_sample_flags, else the first entry's sample_flags, else
the tfhd default_sample_flags, else the trex default (read_video_track()
now keeps it). A reference with no SAP gets 0 for starts_with_SAP, type and
delta; a merged reference takes the flag of its first fragment. parse_traf()
also skips tfhd default_sample_size, which it never needed before.

Tests: remuxing the fixture with frag_keyframe and a 250 ms frag_duration
gives about a dozen references, of which exactly the three that begin at a
keyframe are SAPs; the frag_keyframe fixture scans as all SAPs, and the
byte-for-byte comparison with the reference tool is unchanged; merging takes
the first fragment's flag. The remux sequence the muxer test used is now a
helper shared by both.

refs #5144

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 07:27:37 -04:00
Isaac ConnorandClaude Opus 5.5 05eadfc946 perf: size the sidx region for the event instead of a flat 64 KiB
Every fragmented event reserved 64 KiB for its leading sidx, whatever its
length. On a short motion clip of 1-2 MB that is 3-6% of the file.

zm_mp4::reserve_size() sizes the region for twice the fragments expected
in one section: section length divided by the GOP, which is the encoder's
gop_size when encoding and the packet queue's longest keyframe interval
otherwise, over the capture fps. It rounds up to whole 4 KiB blocks and
stays between 4 KiB (337 references) and the previous 64 KiB, which is
still taken when the section length, GOP or fps is unknown. The default
600 s section at a 1 s GOP now reserves 16 KiB.

A low guess does not lose the index: an event with more fragments than
the region holds has neighbouring fragments merged into one reference,
as before, which only coarsens seeking. VideoStore keeps the size it
reserved so finalize() fills that exact region, and Monitor gains a
GetSectionLength() accessor.

refs #5144

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 22:15:25 -04:00
dm32768 b912fc7848 fix: index event MP4s with a leading sidx so players start at once (#5145)
VideoStore::open() reserves a 64 KiB free box between moov and the first
fragment, but only when the MOV muxer wrote the moov in the header and
moof+mdat fragments follow it. finalize() scans the fragments and fills the
region with free padding followed by a sidx that ends at the first moof, so
FFmpeg-based players (Chromium, Android WebView, Electron) treat the index as
complete and start playback without visiting every fragment. On any parse or
write failure the region stays a free box.

fixes #5144
2026-09-30 20:58:38 -04:00
Isaac ConnorandClaude Opus 5.5 b27d5707ba fix: size the buffer for the file's dimensions in Image::ReadJpeg refs GHSA-rpp4-xmqm-84ff
ReadJpeg() stored the JPEG header's width and height in the Image before
calling WriteBuffer(). WriteBuffer() reallocates only when the requested
size differs from the current one, so it saw no change and kept the
existing buffer and linesize, and the scanline loop then decoded every
row of a larger file past its end. A File monitor re-reads its source
into a monitor-sized image on every capture, so whoever can write that
file could overflow zmc's heap. Stored event JPEGs read by zms take the
same path.

Leave width and height to WriteBuffer(), as DecodeJpeg() already does.
FileCamera::Capture() also now refuses a file whose dimensions do not
match the monitor, since everything downstream is sized for the monitor.

Add a Catch2 case that reads a 256x192 JPEG into a 64x48 image. It
segfaulted before this change and passes after.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 19:44:57 -04:00
Isaac Connor 9f3f6c6770 fix: only trust X-Forwarded-For from configured proxies for auth hash IPs
With ZM_AUTH_HASH_IPS on, the client address bound into the auth hash was
taken from the left-most X-Forwarded-For value whenever the header was
present, both when PHP generated the hash (getRemoteAddr()) and when PHP or
zms validated it (getAuthUser(), zmLoadAuthUser()). The header is client
controlled, so anyone holding a leaked hash could replay it from anywhere by
sending the address it was bound to.

Add ZM_AUTH_TRUSTED_PROXIES, a list of exact reverse proxy addresses.
X-Forwarded-For is now used only when REMOTE_ADDR is one of them, and is read
from the right, skipping hops that are themselves listed proxies, so values a
client prepends are never chosen. With the option empty, the default, the
header is ignored and REMOTE_ADDR is used.

PHP (web/includes/Network.php getRemoteAddr()) and C++ (ClientAddress() in
zm_utils, used by zmLoadAuthUser()) implement the same rule so generation and
validation continue to agree. Every PHP caller already routes through
getRemoteAddr(), so session.php and auth.php need no change.

Reverse proxy users who enable ZM_AUTH_HASH_IPS must list their proxy in the
new option; until they do, hashes bind to the proxy address, which still
validates but no longer distinguishes clients. This is the behaviour change
that the #4921 work avoided by trusting the header.

The option is added through ConfigData only, like other recent options;
zmupdate.pl --freshen inserts it, zms falls back to the compiled-in default and
PHP treats an undefined constant as empty, so no schema migration is needed.

Tests: ClientAddress Catch2 case; tests/php/test_remote_addr.php updated for
the trusted-proxy rule.

refs GHSA-72rf-54rm-798c

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit acea5889dec0826f595cb736147a5fdc9c94a2a9)
2026-09-24 19:51:27 -04:00
Isaac Connor e924bfb999 fix: treat a request parameter that is not a string as absent in auth
Saving a user from the web ui took the whole auth path down:

  PHP Fatal error: Uncaught TypeError: strcasecmp(): Argument #1 ($string1)
  must be of type string, array given in includes/auth.php:197
  #0 auth.php(197): strcasecmp()
  #1 auth.php(528): getAuthUser()
  #2 auth.php(687): userFromSession()

reached from ?view=user&uid=2. That page's form posts user[Username],
user[Password], user[Name] and the rest, so $_REQUEST['user'] is an array on
every save from it, and getAuthUser() read that parameter as the username to
filter on and handed it to strcasecmp(). Under PHP 8 a string function given an
array is a TypeError rather than a warning, so the request died with a 500.

The same shape arrives from anyone who cares to send it, and not only on a page
that needs a session. userFromSession() reads user, pass, username, password
and auth straight out of the request, and the credential branches run before
anyone is logged in, so ?username[]=x&password[]=y reaches validateUser() with
arrays on an install that has never seen the caller before.

requestString() returns a parameter only when it is a string and null
otherwise, which is what the callers already do with a parameter that was not
sent. An array is not a username, a password or an auth hash.

master no longer has the strcasecmp line the report names, so it does not fatal
in that exact spot, but it reads the same unvalidated values: $filterUser is
bound as a query parameter and the credentials still reach validateUser(). This
fixes the class rather than the one line, and backports to 1.38 where the
reported line lives.

The test lifts requestString() out of auth.php and evaluates it alone, because
including auth.php needs a database; test_auth_no_include_side_effects.php
sidesteps the same dependency the same way. 8 cases, covering the form's array,
the login parameters, a nested array and the strings that must still pass
through. 5 of them fail without this change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
(cherry picked from commit 1f4aa1a16c1b63fdf34a6f2c6aef6590dbd45acd)
2026-09-24 19:44:16 -04:00
Isaac Connor 8f5a0af862 Merge pull request #5143 from SteveGilvarry/feature/stream-socket-events
Replace the per-monitor media FIFOs with a unix stream socket
2026-09-23 18:17:50 -04:00
Steve GilvarryandClaude Fable 5.1 f023032488 test: keep the stream socket benchmark's pacing deadline signed
kPacketInterval * i with a size_t i gives the deadline an unsigned
duration; once the producer falls behind, libc++'s sleep_until turns the
negative remaining time into a near-infinite sleep and the benchmark
hangs. Multiply by a signed index.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-23 15:07:15 +10:00
Steve GilvarryandClaude Fable 5.1 0ac5a2cc83 fix: refuse over-long stream socket paths and media for unknown stream ids
zm::UnixSocket copies the path with a truncating strncpy, so a
PATH_SOCKS long enough to overflow sun_path made the server bind one
file while chmod, chown and unlink acted on another, and the client
connect to a truncated, different path. Both now refuse such a path with
an error. Start() also closes the listener on its own failure paths.

SendMedia indexed the two-entry sequence array with the stream id; a
caller passing StreamId::Monitor would have written past it. Ignore
anything that is not video or audio.

Tests: server and client refuse a 200 character path; a Monitor stream
id is ignored and does not disturb the video sequence.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-23 15:07:15 +10:00
Steve GilvarryandClaude Fable 5.1 e917315f12 fix: gate rtsp media on a session confirmed against the latest HELLO set
on_media admitted a packet whenever its generation matched the one the
session was built for. A restarted zmc starts again at generation 0, so
if the camera came back with different parameters its replayed keyframe
passed that check and went into the old codec's packer before the main
thread rebuilt the session.

Move the HELLO and generation bookkeeping into RtspSessionTracker, a
small class with no locking or RTSP types. Every HELLO and every
disconnect clears the confirmed flag; Update() asks the tracker for a
plan (none, adopt, rebuild), acts on it and confirms; on_media feeds the
packers only for a confirmed session at the confirmed generation. A set
that failed to build (unsupported codec) is not retried until a new
HELLO arrives, so a bad camera does not rebuild on every pass.

Tests: a new suite for the tracker covers the complete-set rule, the
HELLO-to-rebuild window, generation reuse after a producer restart with
changed and with unchanged parameters, audio appearing, disappearing and
changing, a generation without video, a failed build, and HELLOs for
streams the server does not serve.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-23 15:07:15 +10:00
Steve GilvarryandClaude Fable 5.1 3d98b216ca fix: frame the stream socket snapshot when a consumer connects
The cached snapshot was framed when the status last changed, so after a
generation bump or further events a new consumer received it stamped
with a stale generation and an old event-sequence baseline. Keep only
the body and frame it in AcceptClient, so the header carries the
generation and sequence in effect at the moment of connection.

Tests: a snapshot cached at generation 0 with no events is delivered to
a later consumer with the current generation and sequence.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-23 15:07:15 +10:00
Steve GilvarryandClaude Fable 5.1 7d9d97f387 fix: apply both stream parameter sets to the stream socket in one generation
Two problems with announcing streams one at a time. SetAudioParams only
bumped the generation when audio had been announced before, so audio
joining a video-only stream was appended to the current generation after
its video HELLO, which the protocol says is the last HELLO of a
generation. And a re-prime that changed both streams bumped twice:
consumers saw, and a connecting consumer was handed, an intermediate
generation pairing the new audio with the old video, which
zm_rtsp_server built a session for and tore down again.

Add StreamSocket::SetStreams(video, audio), which applies the whole set
under one lock: unchanged is a no-op, the first announcement stays in
generation 0, and any change once a video HELLO has gone out - new
parameters, a stream appearing, a stream disappearing - is exactly one
bump with every remaining stream re-announced, audio first. The single
stream setters and the new ClearVideoParams are wrappers over the same
logic, and a null or codec-less parameter set means "no such stream".
PrimeCapture passes both streams in one call.

Tests: audio joining an announced video stream bumps the generation;
SetStreams keeps the initial announcement in generation 0, changes both
streams in one generation with consistent pairing, and drops a video
stream the source no longer has.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-23 15:07:14 +10:00
Isaac ConnorandClaude Opus 5 bdf688eecb fix: carry the zone's alarm colour on the alarmed and filtered pixel methods
The alarmed/filtered pixel check methods handed Overlay() the raw GRAY8
scoring mask. Overlay keys on a non-zero source pixel and copies that byte, so
the highlight took whatever shape one byte has in the destination format: the
red channel on an RGB32 monitor, which is the whole reason alarms have always
come out red; all three channels on RGB24, giving white; luma alone on YUV420.
The zone's configured Alarm Colour was honoured only on the blob path, which
goes through HighlightEdges.

Generalise HighlightEdges into BuildHighlight, which takes an edges_only flag
and otherwise fills every marked pixel, and build the highlight for the pixel
methods the same way the blob path already builds its outline: in the
capture's own pixel format, carrying alarm_rgb, once scoring is finished with
the GRAY8 mask. HighlightEdges stays as a thin wrapper so the blob path and
its callers are unchanged.

This is a behaviour change: monitors left on the default red see no
difference, but a zone configured with any other Alarm Colour now paints that
colour instead of red or white.

Reverting the zone hunk fails the new test in all three of its sections.
Full suite: 148 cases, 12535 assertions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WBHBB95RBX7D9p8ge2WDZb
2026-09-22 09:57:19 -04:00
Isaac ConnorandClaude Opus 5 5b2dd2abbe fix: do not fault in the shm time accessors before connect() has mapped
Monitor::connect() returns false with shared_data still null on every one of
its failure paths: the mmap file cannot be opened (wrong ownership, e.g.
after a package upgrade), fstat fails, ftruncate cannot grow it (/dev/shm out
of space -- a container with the default 64MB tmpfs hits this quickly, since
one 720x480 monitor with 10 buffers already asks for ~20MB and a 1080x720 one
asks for ~62MB), or mmap itself fails.

zmc's startup loop reacts by retrying:

    while (!monitor->connect() and !zm_terminate) {
      Warning("Couldn't connect to monitor %d", monitor->Id());
      monitor->SetHeartbeatTime(std::chrono::system_clock::now());
      sleep(1);
    }

so the first thing it does after a failed connect is write through the null
pointer. zmc dies with SIGSEGV at address 0x80 instead of retrying, which
presents as a monitor that will not start and a capture daemon that keeps
crashing. The accessors on either side of these already guard with
`if (shared_data && shared_data->valid)`.

Reverting the guard fails the new test with SIGSEGV at zm_monitor_shm.cpp:66,
and restoring it passes. Same fix as release-1.38's a550545e4.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WBHBB95RBX7D9p8ge2WDZb
2026-09-22 09:57:19 -04:00
Isaac ConnorandClaude Opus 5 428c048bac feat: measure the audio level on demand, with a meter in the editor
Reverts the previous commit's always-on measurement. Decoding audio for every
monitor that has it spends CPU on a number almost nothing reads, which is the
wrong trade even though it did solve the chicken and egg of picking a
threshold without ever seeing a level.

Measure when something is actually going to use the reading instead:

 - AudioDetection is on, as before, so nothing changes for a monitor that
   scores on audio; or
 - somebody asked. SharedData gains audio_level_until, a wall clock second
   the capture thread keeps measuring up to. The monitor editor's new level
   meter pushes it forward while it is on screen and the measurement lapses a
   few seconds after the page is left, so nothing has to send a stop and a
   crashed browser cannot leave a monitor decoding forever.

When the reading stops being wanted the decoder is released and the published
level and peak are cleared, so a stale number is not left looking current and
an old peak does not land on the next frame row written.

audio_level_until is carved out of analysis_pad rather than appended, so
SharedData stays 888 bytes and no existing offset moves; the static_asserts,
Memory.pm and Monitor.php are updated together and all three now agree the
field is at +880.

The meter itself is on the audio settings, shown whether or not
AudioDetection is checked, because the level is what you need in order to
choose a threshold. It draws the threshold currently in the input as a mark on
the bar so a reading can be judged against it before saving, and a monitor
whose zmc is not running reads "no reading" rather than a confident 0, which
would be indistinguishable from silence.

Frames.AudioLevel is therefore 0 again on monitors that do not score on audio.
That is what the graph already treats as "no audio data", so it draws no line
rather than a flat one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JpiSWBmtQkR5bcgpHWY4ME
2026-09-20 18:17:44 -05:00
Isaac ConnorandClaude Opus 5 4619258293 feat: measure the audio level whatever AudioDetection is set to
Gating the measurement on AudioDetection made the graph useless for the job
it is most wanted for. AudioThreshold is a per-device number -- the floor on
one camera's mic is nothing like another's -- so it has to be measured before
it can be set, but nothing was measured until it was already set. Enabling
detection with a guessed threshold to find out what the real one should be is
backwards.

The level is now read for every monitor with decodable audio.  AudioDetection
governs only whether crossing the threshold contributes a score, which is
what the setting is named for. shared_data->audio_alarm stays 0 when it is
off, so nothing downstream changes for a monitor that does not want audio
alarms.

Nothing here depends on Analysing either. The measurement is in
Monitor::Capture, which runs on whatever Analysing is set to, and frame rows
come from Event::AddFrame, which a continuously recording monitor reaches
through the RECORDING_ALWAYS path with motion detection off. So a monitor
that only records continuously still gets levels on its rows.

Since Monitor::Capture retries Open on every audio packet until it succeeds,
and that now happens for every monitor with audio rather than the handful with
detection on, AudioDetector remembers a codec it has already failed to find a
decoder for. Without it a stream ZoneMinder cannot decode logs a warning at
the audio packet rate for as long as the monitor runs. A reconnect bringing a
different codec is still tried.

The cost is one audio decode per monitor with audio, where before it was one
per monitor with detection enabled. That is small next to the video path, but
it is not nothing on a box with many cameras.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JpiSWBmtQkR5bcgpHWY4ME
2026-09-20 18:17:44 -05:00
Isaac ConnorandClaude Opus 5 4ea307a1cf feat: graph motion score and audio level under the event video
The cue strip under the video showed one flat red band per alarm period. It
tried to encode the score as a bar height, 'height: '+frame.Score+'px', but
never could: the frames ajax did not return Score, so every bar got
'height: undefinedpx' and fell back to the stylesheet's height: 100%. So the
motion level it looks like it is drawing has never actually been drawn.

Replace it with a line graph of both series over the length of the event. The
alarm periods stay, as a pale wash behind the lines, so nothing that was
readable before is lost.

The two series do not share a vertical scale. Audio level is 0-100 by
construction, but a motion score is a sum over zones with no upper bound, so
pinning both to 0-100 would flatten the audio line against the floor on any
event scoring above 100. Each is scaled to its own maximum, and hovering reads
out the exact values, which is what the numbers are wanted for. An event whose
rows are all zero -- recorded before the column existed, or by a monitor with
AudioDetection off -- draws no audio line at all, rather than a flat line
claiming silence was measured.

The readout goes into the existing #indicator, which already tracks the mouse
across the whole progress bar, instead of a second tooltip competing for the
same pixels.

The geometry lives in web/js/LevelGraph.js so it can be tested without a DOM,
and so the polyline and the hover readout cannot disagree about where a given
second sits. The bar grows from 1.25em to fit the graph, which needed two
consequential CSS changes: #indicator now spans the taller bar, and
.progressBox drops from 0.66 opacity to 0.25, because at full strength the
played part of the graph is unreadable behind it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JpiSWBmtQkR5bcgpHWY4ME
2026-09-20 18:17:44 -05:00
Isaac ConnorandClaude Opus 5 33661907d1 feat: persist the peak audio level on each frame row
zm_update-1.39.31.sql gave monitors audio detection, but the level only ever
existed in shared memory, so it was gone the moment the frame passed and there
was nothing for the event view to plot. Add Frames.AudioLevel next to Score, on
the same 0-100 dBFS-derived scale the threshold uses.

What is stored is the peak since the previous row, not the level at the instant
the row was written. Frames rows are written well below the capture rate --
only alarm, bulk and score-increasing frames get one -- so sampling at write
time would drop exactly the short loud noises worth seeing on a timeline.
AudioDetector accumulates the peak as it decodes and Event::AddFrame takes it
where the row is built, which clears it so each row covers its own interval.
The Event constructor takes and discards it once, otherwise an event's first
row reports the loudest moment since the previous event ended.

This needs no shared memory change: zma is now an offline re-analysis tool and
the live analysis runs in a thread of zmc, alongside the capture thread that
runs the decoder, so the peak can stay in the AudioDetector. SharedData keeps
its documented 888-byte layout and its fixed offsets.

The frames ajax returns the column, and Score with it. elements in
web/ajax/status.php is a whitelist that never listed Score, which is why the
event view's cue strip has been reading an undefined Score off every frame.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JpiSWBmtQkR5bcgpHWY4ME
2026-09-20 18:17:43 -05:00
Isaac ConnorandClaude Opus 5 40c4cd3c5d fix: run reverse playback at the rate the user picked
changeRate's reverse branch did not use the selected rate as the reverse
speed. It looked it up as rates[rates.indexOf(-rate)-1]/100, stepping one
entry further down the shared rate list, so every reverse rate ran a notch
too slow: -1/2x played at 1/4x and -16x at 10x. -1/4x was worse than slow,
because one step below 25 in that list is 0, so revSpeed came out 0 and the
video sat still while the ui claimed it was rewinding.

The rate the user picked is the speed, so use it.

Leaving reverse through the dropdown also leaked the rewind interval, which
only pauseClicked and vjsPlay ever stopped. Picking a forward rate after a
reverse one left it running, so it went on dragging currentTime backwards and
resetting playbackRate to 0 on every tick while the player was supposedly
running forwards. The teardown is now stopRewind(), split out of
stopFastRev() because stopFastRev rewrites the rate select to 1x, which would
undo the choice changeRate is in the middle of applying.

stopFastRev no longer reads the rate back out of the player to decide what to
put in the select and the cookie, for the same reason streamFastFwd stopped
doing it: videojs can defer the set until the tech is ready, so the getter
still answers with the rate we just left.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JpiSWBmtQkR5bcgpHWY4ME
2026-09-20 18:17:43 -05:00
Isaac ConnorandClaude Opus 5 ef361dc3ab fix: stop the event playback rate buttons stepping off the end of the rate list
Clicking fast forward once too many at 16x threw out of video.js:

  TypeError: HTMLMediaElement.playbackRate setter: Value being assigned is
  not a finite floating-point value.
    playbackRate@video.min.js
    streamFastFwd@..._views_js_event-....js

streamFastFwd stepped the shared rate list by indexing it directly,
rates[rates.indexOf(current)+1]. At the top of the list that is rates[15],
undefined, and undefined/100 is NaN, which Firefox refuses outright.

The guard meant to prevent this ran after the assignment rather than before
it, and read the rate back from the player to decide, so it could only
disable the button once the bad value had already been sent. It was reachable
in normal use because streamPlay() re-enables the button whatever rate we are
at, so play-then-fast-forward at 16x throws every time; picking 16x from the
rate dropdown gets there too, since changeRate does not touch button state.

indexOf also answers -1 for a rate that is not in the list, and -1+1 indexes
rates[0], which is -1600: stepping forwards from an unlisted rate asked for
16x reverse.

streamFastRev had the same fault at the other end. rates[0-1] is undefined, so
revSpeed became NaN and every tick of the rewind interval then handed
currentTime a NaN.

Both now go through stepRate, which snaps an unlisted rate to the nearest
listed one and returns null rather than walking off either end, so the caller
disables the button and leaves the player alone instead of assigning
something it cannot use. streamFastFwd also sets the dropdown and cookie from
the rate it just asked for rather than reading it back, because the stack in
the report shows videojs deferring the set until the tech is ready, at which
point the getter still answers with the old rate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JpiSWBmtQkR5bcgpHWY4ME
2026-09-20 18:17:43 -05:00
Isaac ConnorandClaude Opus 5 f36645dfd9 fix: point the audioMotion-analyzer install instructions at the ES module
The library is AGPL-3.0-or-later so it is not shipped, and the admin installs
it at skins/<skin>/assets/audioMotion-analyzer/src/audioMotion-analyzer.js.
The instructions for doing that had drifted from what the code loads:

- help.txt and the OPTIONS_WHATTODISPLAY help gave the bare
  https://cdn.jsdelivr.net/npm/audiomotion-analyzer@X.X.X URL, which resolves
  to the package's "main" entry, the minified UMD bundle dist/index.js. The
  install path is the package's src/ ES module, so the URL needs the explicit
  /src/audioMotion-analyzer.js suffix. Same for the download links in the
  AudioMotionVersionNotInstalled and AudioMotionVersionWrongVersion messages.
- The install path was written as /skins/MySkin/..., a placeholder that does
  not correspond to any skin.
- RequiresAudioMotionEnabled named only the file, not where it goes.
- assets/version documented every other asset in that directory but not this
  one, leaving no explanation for the otherwise empty directory.

help.txt no longer restates the required version, so 4.5.4 stays declared only
by SUPPORTED_AUDIO_MOTION_ANALYZER_VERSION as intended, and assets/version
points at that constant rather than duplicating it.

Add tests/js/audiomotion-paths.test.js to hold the PHP feature probe, the
dynamic import, help.txt and both lang catalogues to the same path, the same
download URL and the same version.

Also gitignore the installed library so a local install is not committed back
into a GPL-2.0 tree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JpiSWBmtQkR5bcgpHWY4ME
2026-09-20 18:17:43 -05:00
Claude 21042ba2f7 fix: send audio HELLO first and keep rtsp sessions across a zmc restart
Correct three problems in the stream socket generation tracking added by
the previous review fix-up:

- ClearAudioParams bumped the generation without restarting the video
  sequence or dropping the cached keyframe, so a late joiner received a
  HELLO at generation N+1 followed by a KEYFRAME still stamped N, and
  sequences did not restart as documented. It now does the same full bump
  as SetVideoParams/SetAudioParams.
- zm_rtsp_server rebuilt the xop session whenever the generation changed,
  even with identical codec parameters. Generations restart at 0 when zmc
  restarts, so every zmc restart dropped the RTSP clients; the original
  code kept the session in that case. Unchanged parameters now just adopt
  the new generation without a teardown.
- Deciding whether audio belongs to the current generation by comparing
  per-stream generation numbers raced the two HELLOs of a generation
  (double rebuild when Update() ran between them) and cannot tell a
  producer restart apart. The producer now sends the audio HELLO before
  the video HELLO within a generation (on connect and on every bump, and
  PrimeCapture announces audio before video), so the video HELLO always
  completes a generation's parameter set. The consumer forgets the
  previous generation's HELLOs when a new generation starts and on
  disconnect, and builds only once the video HELLO of the latest
  generation is in. It also records whether audio was announced at build
  time rather than whether a packer was created, so an unsupported audio
  codec no longer triggers a rebuild on every pass.

The ordering guarantee is documented in the protocol header and the
stream socket docs. Tests pin the audio-first order on connect and on a
video reconfigure, and the keyframe drop and sequence restart on audio
removal.

refs #5143

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4UcdJLt1bxwdpcigGxZRD
2026-09-20 00:36:35 +00:00
Claude 281a7d967e fix: harden stream socket transport and protocol from PR review
Address several stream socket review findings on the transport, its
consumer client and the wire protocol:

- ParseAllowedUids rejects negative, out-of-range and non-round-tripping
  uids instead of wrapping or truncating them (e.g. 2^32 no longer
  becomes uid 0).
- StreamSocketClient backs off after a connection the producer closes
  before any message, so a rejected consumer (uid allow-list, client
  limit) no longer busy-loops; a rejection is not reported as a
  disconnect.
- SendMedia drops packets for a stream that has no announced HELLO, and
  ClearAudioParams forgets a previously announced audio stream (bumping
  the generation and re-issuing the surviving video HELLO), so a stale
  audio HELLO is never replayed and media never precedes its HELLO.
- Header pts_us is encoded as signed (two's-complement) microseconds so
  negative and AV_NOPTS_VALUE timestamps survive the wire; the dump tool
  decodes it as signed and tracks sequence gaps per generation so a
  generation reset is not mistaken for packet loss.

Tests cover the uid rejections, the audio HELLO clearing and media
guard, the connection-rejection backoff, and signed pts round-trips.

refs #5143

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4UcdJLt1bxwdpcigGxZRD
2026-09-19 23:00:03 +00:00
Steve GilvarryandClaude Fable 5.1 0e88b390c3 test: add a stream socket cost benchmark
A hidden Catch2 case ([.benchmark], run with ./tests "[benchmark]") that
pushes 1000 synthetic H.264-like packets (150 KB keyframe every 50, 15 KB
deltas) through a StreamSocket at 250 packets a second with 0, 1 and 8
consumers, and prints per scenario: SendMedia wall time on the producer
thread (p50/p99/max), CPU per packet for the producer, the listener
thread and an average consumer, and how many packets each consumer
received. Producer CPU is sampled around each SendMedia with the clock
read overhead calibrated out; the listener figure is the process CPU
left after subtracting the producer thread and the consumers.

Run on a 4-core VM after the unlocked drain change: about 1.3 us per
SendMedia with no consumer, 16 us with one, 18 to 24 us with eight,
and every consumer receives every packet.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-16 05:51:32 +10:00
Steve GilvarryandClaude Fable 5.1 c11e484eca test: use a per-process socket path in the stream socket tests
Both suites bound a fixed path under /tmp, so two test binaries running
at once (parallel ctest, a developer and CI on one box) would unlink
each other's listener. Include the pid in the path.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-16 05:12:43 +10:00
Steve GilvarryandClaude Fable 5.1 65fdfaecae fix: omit oversized extradata from the stream socket HELLO instead of truncating
BuildHello cast extradata_size to the u16 TLV length, so extradata over
64 KiB was sent truncated with a length that happened to match, and a
consumer would try to use a cut-off parameter set. Parameter sets are a
few hundred bytes, so anything that large is unusable anyway: warn and
leave the tag out, which keeps the HELLO valid. Name the TLV limit and
use it for the string clamp too.

Tests: extradata one byte over the limit is omitted and the HELLO still
parses; exactly the limit travels intact.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-16 05:12:42 +10:00
Steve GilvarryandClaude Fable 5.1 84b1c06430 fix: drop the cached stream socket keyframe when the capture source closes
The keyframe cached for late joiners survived Monitor::Close(). A
consumer connecting while the camera reconnected was primed with a
keyframe from the previous capture session, whose pts can be ahead of
what the new session produces, and with identical stream parameters
there is no generation bump to warn it. Add StreamSocket::InvalidateKeyframe()
and call it from Close(); the next keyframe from the new session fills
the cache again.

Tests: after InvalidateKeyframe() a new consumer gets HELLO and then the
next live packet, with no KEYFRAME replay in between.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-16 05:12:42 +10:00
Steve GilvarryandClaude Fable 5.1 d5d997da41 feat: deliver EVENT frames through StreamSocketClient
The consumer class dispatched HELLO, MEDIA, KEYFRAME, STATS and BYE but
had no case for the EVENT type added for the monitor lifecycle channel,
so every event fell into the unknown-type branch and was dropped at
debug level 2. Add an on_event callback that receives the parsed
MonitorEvent together with the header (event sequence and media
generation), and reject malformed payloads with a warning like HELLO.

Tests: a client connected to a StreamSocket receives the cached snapshot
on connect and a broadcast state_changed, with the codes, state names,
health code, message and wall clock intact.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-16 05:12:42 +10:00
SteveGilvarryandClaude Fable 5.1 f69700eb6a feat: serve monitor lifecycle events on the stream socket
Add SendMonitorEvent() to broadcast an EVENT frame to every connected
consumer, framed with a per-monitor event sequence (independent of the
media sequences, so it is not reset by a media generation bump) and the
current media generation for correlation. Events are control messages and
are never dropped from a client queue; the sequence still advances when no
consumer is connected, so a late joiner sees the loss as a gap.

Add SetSnapshotEvent() to cache the current-status snapshot replayed to
each new consumer on connect, the events analogue of the cached keyframe.
AcceptClient now enqueues HELLO(s), the snapshot, then the keyframe.

Tests: a broadcast EVENT round-trips with the right type/stream/sequence;
the event sequence advances across a clientless gap; the snapshot is
replayed after HELLO on connect. Ran ./tests/tests '[stream_socket]':
130 assertions in 11 cases pass.

refs #2875

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-16 05:02:24 +10:00
SteveGilvarryandClaude Fable 5.1 2964b8ed59 feat: add EVENT frame to stream socket wire protocol
Add MessageType::Event (0x06) and StreamId::Monitor (0x02) to the v1
stream socket protocol, with an event-code + TLV payload (BuildEvent/
ParseEvent) for the monitor lifecycle channel. Event codes cover the
capture-fault edges (connection/prime/capture failed and restored), a
state_changed transition, and a snapshot sent on consumer connect. TLV
tags carry wall-clock microseconds, a human-readable message, current and
previous state id, an errno/ffmpeg detail code, a state name, and a health
code that a snapshot uses to report the active fault.

The framing is unchanged and length-prefixed, so the new type is
forward-compatible: a pre-event consumer skips it by length, a pre-event
zmc never emits one. No protocol version bump.

Adds round-trip, unknown-tag-skip, and truncation tests. Ran
./tests/tests 'stream_socket::*': 93 assertions in 14 cases pass.

refs #2875

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-16 05:02:24 +10:00
SteveGilvarryandClaude Fable 5 d0c0a64754 feat: add StreamSocketClient consumer class
Reader side of the stream socket protocol, for zm_rtsp_server and any
in-tree consumer: connects to PATH_SOCKS/stream_{id}.sock with 1s retry,
reads exact-size length-prefixed messages into one reused buffer (no
per-read allocation churn, no resync scanning - contrast the FIFO
reader's 4KB chunking and memmem hunting), validates headers and
dispatches HELLO/MEDIA/KEYFRAME/STATS/BYE through callbacks on the
reader thread. Reconnects automatically on EOF or protocol error; a
fresh HELLO arrives after every reconnect. An adopt-fd constructor
supports tests and single-shot uses.

Tests: 5 new Catch2 cases - end-to-end against a real StreamSocket
(HELLO fields, media payload integrity), BYE + reconnect across a
server restart, fragmented byte-stream delivery via socketpair,
malformed-header disconnect, unknown-message-type skip. Full suite:
103/103 pass via ctest.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-09-16 05:00:44 +10:00
SteveGilvarryandClaude Fable 5 b46be59622 feat: add StreamSocket unix-socket media server class
Per-monitor media stream server for the wire protocol added previously:
a poll()-driven listener thread serving multiple consumers at once over
PATH_SOCKS/stream_{monitor_id}.sock.

Memory and blocking behaviour:
- each message is serialized once and shared across all client queues;
  media payloads are reference-counted via av_packet_clone, no copies
- header and payload are written with writev, never concatenated
- the producer (capture thread) never blocks: per-client queues are
  bounded by bytes and message count; on overflow the oldest non-control
  messages are dropped for that client only, observable as sequence gaps
  and in STATS; clients making no progress are disconnected
- the latest keyframe access unit is cached (refcount only) and replayed
  to late joiners after HELLO for immediate first-frame rendering

Parameter changes (SetVideoParams/SetAudioParams) bump the generation,
reset sequences and rebroadcast HELLO. Peers are checked via SO_PEERCRED
against ZM_STREAM_SOCKET_ALLOWED_UIDS; sockets are chmod 0660 with group
ZM_STREAM_SOCKET_GROUP. Stop() sends BYE so consumers can distinguish
shutdown from failure.

Tests: 8 new Catch2 test cases (lifecycle/permissions, HELLO-first
ordering, late-joiner keyframe replay, queue overflow with sequence-gap
and STATS accounting, stalled-vs-live client isolation, generation bump,
BYE on stop, allowed-uids parsing). Full suite: 98/98 pass via ctest.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-09-16 05:00:44 +10:00
SteveGilvarryandClaude Fable 5 c4a8d4ddf6 feat: add stream socket wire protocol encoder/decoder
First part of the monitor stream socket series: a length-prefixed binary
protocol (version 1) intended to replace the per-monitor media FIFO text
framing. Pure encode/decode functions with no I/O.

- 24-byte fixed header (little-endian): length, version, type, stream,
  flags, sequence, generation, pts_us
- Message types HELLO/MEDIA/KEYFRAME/STATS/BYE
- HELLO payload is a TLV list built from AVCodecParameters, carrying
  codec id, extradata (SPS/PPS/VPS, AAC AudioSpecificConfig, AV1
  sequence header), dimensions, frame rate, sample rate, channels,
  profile and level; unknown tags are skipped by parsers
- STATS payload carries per-client sent/dropped counters

Tests: 8 new Catch2 test cases (header roundtrip and boundary/reject
cases, HELLO video/audio/no-extradata roundtrips, unknown-tag skip,
malformed-TLV rejection, STATS roundtrip). Full suite: 90/90 pass
locally via ctest with BUILD_TEST_SUITE=ON.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-09-16 05:00:44 +10:00
Isaac ConnorandClaude Opus 5 599627a940 fix: stop the cycle timer leaking an interval on every restore fixes #5135
cycleStart() took a fresh setInterval id straight into cycleIntervalId without
clearing what was already there. Once overwritten the old id is unrecoverable,
so the orphan keeps calling nextCycleView every second and cyclePause() can only
ever stop the last one armed.

Several callers reach cycleStart() with no cyclePause() in between: the play
button, the are-you-still-watching modal closing, and startPage(). That last one
is the routine path - it runs from visibilitychange, from resume and from
pageshow, and a restore fires more than one of those, so a tab coming back while
cycling was active armed two. The monitor restart in the same function is
protected, since it nulls prevStateStarted on the way through, but the cycle
branch below it never cleared prevStateCycle.

Clear the interval at the top of cycleStart(), which makes every caller safe
whatever order they arrive in, and clear prevStateCycle in startPage() the way
prevStateStarted already is. cycle.js has the same shape in its own cycleStart()
behind a play button, so it gets the same guard.

Tests in tests/js/watch-cycle-interval.test.js drive watch.js under stubbed
timers and count what is left running: two starts leave one interval, a pause
after three starts leaves none, and a second startPage() does not re-arm. Three
of the four fail against the unfixed file.

Full JS suite passes, ESLint clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkQwahn9pi1y4wJe9BTxjM
2026-09-13 11:51:26 -04:00
Isaac ConnorandClaude Opus 5 da99928235 fix: revalidate on every resume rather than trusting a time window
The staleness window was unsound. calculateAuthHash() keys the hash to the
clock hour it was minted in and getAuthUser() accepts the last
ZM_AUTH_HASH_TTL hourly buckets, so a hash dies at the top of an hour rather
than at some age. generateAuthHash() then serves the cached one until it is
half a TTL old, so what arrives can already be nearly spent: on the defaults a
hash minted at 10:59 is still handed out at 11:58 and is refused at 12:00. A
client stamping that arrival as fresh for an hour skips the probe until 12:58
and restarts its streams on a dead hash - the exact failure this was written to
prevent. No fixed window is safe, because the remaining life of a hash we hold
can be anything down to zero, and AUTH_STALE_MS also ignored the configured
ZM_AUTH_HASH_TTL.

So drop AUTH_STALE_MS, authIsStale() and authFreshAt, and have whenAuthFresh()
revalidate. The one case that can still skip the probe is having no hash at all
- authentication off, or a relay form that does not use one - where there is
nothing that can expire and nothing a probe would report. revalidateAuth()
already shares one request between concurrent callers, so a resume that wakes
several of these still costs a single probe, and that is what the montage code
did unconditionally before any of this.

refreshTablesPendingVisibility() now returns as soon as it finds nothing was
deferred. It is bound on every classic page including the unauthenticated ones,
and the version before this ran the whole auth path on an empty queue, so
merely becoming visible could fire a probe with no work behind it.

Tests: two authIsStale cases removed with the function, two whenAuthFresh cases
added - a probe is sent and the callback held until it answers, and no probe is
sent when there is no hash. Reintroducing a fast path fails the first. Full JS
suite green, ESLint clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkQwahn9pi1y4wJe9BTxjM
2026-09-13 10:37:02 -04:00
Isaac ConnorandClaude Opus 5 4c878eeebe fix: send the auth revalidation probe without a credential
The probe went out as zmAuth.appendTo(...), carrying the very hash it existed to
replace. zm_authenticate_request() resolves a request against exactly one source:
a non-empty auth= in the URL enters the ZM_AUTH_HASH_LOGINS branch (on by
default), and when getAuthUser() rejects it the chain has already been taken, so
the userFromSession() arm below it never runs. A live session cookie then
authenticates as nobody.

Past ZM_AUTH_HASH_TTL - a tab hidden longer than two hours on the defaults, which
is exactly the case this change is for - the probe was therefore the one request
guaranteed to fail, and its failure is read as 'login', so the user was bounced
to the login page with a perfectly good session. That is worse than the 403s in
the log this set out to remove.

Send the probe bare. The session cookie is what answers, which is the question
being asked: who am I, and what is my current hash?

That also makes the failure handling mean what its comment claimed. A rejection
now really is a dead session rather than a dead hash, so redirecting to login on
it is right - and both 401 and 403 reach it, which the comment now says.

Also correct the AUTH_STALE_MS comment: authIsStale() is a strict comparison, so
a credential confirmed exactly AUTH_STALE_MS ago is still fresh, as the test
asserts.

Tests: tests/js/auth-helpers.test.js, 51 passed (4 new for revalidateAuth,
covering the bare probe, shared in-flight request, callbacks surviving a
transient failure, and login on a rejected session). Reverting the probe to the
credentialed form fails two of them.

refs #5093

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nr76CednxtDt2nPuq6WrbL
2026-09-13 00:03:13 -04:00
Isaac ConnorandClaude Opus 5 1d01c966d4 refactor: track credential age instead of hidden time, fold montage in
whenAuthFresh() gated on how long the tab had been hidden, which only covers
one of the ways the page stops hearing about the credential. montage's idle
timeout is another: hitting ZM_WEB_VIEWING_TIMEOUT stops every monitor, and
their status polls with them, while the tab stays visible the whole time. The
Are You Still Watching modal can then sit there for hours, so the auth hash
baked into the monitor srcs is just as dead as after an overnight sleep, and
a hidden-time gate would wave it straight through.

Track the age of the credential itself instead. ZMAuth.update() stamps it
whenever a reply carries auth_relay or auth - even when the value is
unchanged, since the server has still just confirmed it - and revalidateAuth()
stamps it too, so authentication being off doesn't leave every caller
revalidating once an hour forever. authHiddenTooLong(hiddenAt, now) becomes
authIsStale(freshAt, now); onAuthVisible() no longer keeps a timestamp.

That subsumes montage's refreshAuthAndStartMonitors(), which duplicated
revalidateAuth()'s navBar probe and fired a second one on every refocus.
Both call sites are now whenAuthFresh(startVisibleMonitors), which shares
the in-flight probe rather than racing it.

Tested: node tests/js/auth-helpers.test.js (47 passed), node
tests/js/table-helpers.test.js (8 passed), npx eslint on the changed files.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JkE45gkMjnkTiJiUySbe7V
2026-09-13 00:03:13 -04:00
Isaac ConnorandClaude Opus 5 c83a15e048 fix: refresh auth hash before restarting streams after a long tab hide
Hidden tabs have their timers throttled, and frozen/slept ones stopped
outright, so nothing refreshes the credential while the tab is in the
background. The server rotates the auth hash at half of AUTH_HASH_TTL
(generateAuthHash()), so the hash baked into a stream <img> src or a
deferred table url is usually dead by the time the tab is refocused. Every
restarted stream and table poll then fires a request that 403s and logs an
auth error before anything gets around to revalidating.

Record when the tab goes hidden and treat the credential as stale once more
than an hour has passed. Add whenAuthFresh(cb), which runs cb straight away
after a short alt-tab but queues it behind a revalidation after a long
sleep, so nothing makes an authenticated request on the expired hash.

revalidateAuth() now queues callbacks rather than dropping them when a probe
is already in flight, clears the hidden timestamp only once a reply actually
arrives, and still runs the queued callbacks on a transient failure - a
network blip is no reason to leave the page's streams stopped. A 401 still
redirects to login and drops the queue.

Gate the two visibility-driven callers on it: watch.js startPage() and
table-helpers.js refreshTablesPendingVisibility(). montage.js already
refreshes auth before restarting its monitors.

Tested: node tests/js/auth-helpers.test.js (48 passed), node
tests/js/table-helpers.test.js (8 passed), npx eslint on the changed files.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JkE45gkMjnkTiJiUySbe7V
2026-09-13 00:03:13 -04:00
Isaac ConnorandClaude Opus 5 1361f93804 fix: send the initial mode=paused CMD_PLAY that streamCommand was dropping
select_zms() has three ways out, and two of them send a stream command before
the tail of the function sets started. streamCommand() drops anything sent while
that is false, so such a branch reports success having sent nothing and the
picture sits on its last keepalive frame.

The resume branch was fixed on this branch already. The branch below it, for a
page rendered with mode=paused, has the same shape and was missed. With auth on
it needed a still-valid hash to reach, so it was intermittent; with auth off,
where there is no hash that can go stale, the srcAuthCurrent change on this
branch makes it the path every initial load takes. Set started there too.

Add tests/js/monitorstream-resume.test.js, which asserts what reaches the wire
rather than which branch ran: the resume and initial-paused paths each send
exactly one CMD_PLAY on the connkey they are supposed to address, resuming
leaves src and connkey alone, and a stale auth hash still rebuilds and quits the
process the old connkey addressed. Removing the one-line fix fails the
initial-paused case and leaves the other three passing.

Full JS suite passes, ESLint clean on both files.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkQwahn9pi1y4wJe9BTxjM
2026-09-12 15:51:14 -04:00
Isaac Connor 76002987f9 Merge pull request #5128 from connortechnology/5127-packetqueue-iterator-leak
fix: free the event start iterator when openEvent cannot lock its packet
2026-09-12 10:24:25 -04:00
Steve GilvarryandClaude Opus 5 74453bf783 test: guard the WS-Security test on WITH_GSOAP refs #4998
src/CMakeLists.txt wraps the gsoap sources in if(GSOAP_FOUND), so the
daemons build fine without gsoap. tests/ had no such guard, and
zm_onvif_wsse.cpp includes soapH.h and plugin/wsseapi.h unconditionally,
so BUILD_TEST_SUITE=ON plus no gsoap failed to compile:

  tests/zm_onvif_wsse.cpp:23:10: fatal error: 'soapH.h' file not found

gsoap was therefore optional for the binaries and mandatory for the test
suite. Nothing noticed because ci-cpp-tests.yml installs libgsoap-dev, so
the no-gsoap path is never exercised.

Guarded in the file rather than in CMake, matching what
zm_onvif_auth_error.cpp already does. zm_onvif_renewal.cpp deliberately
stays unguarded - it pulls in no gsoap headers and its four cases are
plain timing and formatting logic that run either way.

Verified on macOS: 144 cases with gsoap, 142 without, both clean builds.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B5KL9Xbi7K5aGsauLtd8tG
2026-09-12 15:52:55 +10:00