Commit Graph
12739 Commits
Author SHA1 Message Date
b1470cd719 fix(heltec): sleep the T1 and T096 panels on screen-off to stop image retention (#11894)
* fix(heltec-t1): sleep the panel on screen-off to stop image retention

The T1 is excluded from the LovyanGFX sleep()/wakeup() calls because it uses
TFT_eSPI, so DISPLAYOFF only dropped the backlight and the ST7735 kept driving
the last frame unlit for the whole screen-off timeout. That constant static
image is what burns ghost pixels into the panel.

Add opt-in TFT_SLEEP_WHEN_OFF for the TFT_eSPI path: DISPOFF + SLPIN on
screen-off, SLPOUT (120 ms) + DISPON on wake, issued as raw MIPI DCS since
TFT_eSPI exposes no sleep API. Frame memory survives sleep-in, so the previous
frame reappears and the dirty-window diff continues unchanged. The sleep flag
keeps the double displayOn() in Screen::handleSetOn() from paying the delay
twice.

* fix(heltec-t096): sleep the panel on screen-off to stop image retention

The T096 sits behind the same TFT_eSPI exclusion as the T1, so DISPLAYOFF only
dropped its backlight while the ST7735S kept driving the last frame unlit for
the whole screen-off timeout. Same panel, same bus, same burn-in.

Opt in to TFT_SLEEP_WHEN_OFF; the TFTDisplay side of the fix is already generic.

* fix(tft): harden the TFT_SLEEP_WHEN_OFF wake/sleep sequence

Drive VTFT_CTRL LOW before SLPOUT, so the rail is up before the panel is
addressed. Wait out the remainder of the 120 ms the controller needs after
SLPIN before sending SLPOUT, so a wake landing as the screen timeout fires is
not dropped. Guard DISPLAYOFF on panelAsleep to match DISPLAYON.

Correct the T1 VTFT_CTRL comment: LOW enables the rail, not HIGH.

* fix(tft): wait out only what is left of the sleep-in window before SLPOUT

Throttle::isWithinTimespanMs() is true for the whole 120 ms after SLPIN, so the
wake path paid a fresh 120 ms on top of however much had already elapsed. A wake
119 ms after the SLPIN waited ~120 ms rather than ~1 ms, up to 119 ms of
avoidable latency on every quick off/on.

Add Throttle::remainingMs(), which returns what is left of the interval and
saturates at 0 instead of underflowing to a ~49 day wait. It reads the clock
once, so a caller that tests and then waits cannot be preempted between the two
and land on that underflow - which a separate isWithinTimespanMs() plus
subtraction at the call site could.

heltec-mesh-node-t1 and -t096 both build; neither is board_level = pr, so CI
does not compile this path. Docker native suite green, 1515/1515.

* trunk

---------

Co-authored-by: Thomas Göttgens <tgoettgens@gmail.com>
Co-authored-by: Ben Meadors <benmmeadors@gmail.com>
Co-authored-by: Jonathan Bennett <jbennett@incomsystems.biz>
2026-09-23 09:10:38 +00:00
Tom 08cd97ea2d test(harness): make run-tests.sh drivable by a caller that cannot see the terminal (#11862)
* test(harness): make run-tests.sh drivable by a caller that cannot see the terminal

A tool call or a fresh session reads a captured output file, gets interrupted
mid-run, and starts cold. Two things in the wrapper tripped that caller: an
interrupt left pio's build tree running with no record of it (the next
invocation started a second build into the same .pio/build/, or pgrep'd and
matched itself); and pio's own "[PASSED]" / "N succeeded" lines made a
half-finished output file read as green.

run-tests.sh:
- Run record in .pio/runtests/current.tsv for the life of a run. A second
  invocation prints RESULT: BUSY and exits 4 without touching the build
  directory. Valid while the holder pid OR the recorded process group is
  alive, so a SIGKILLed wrapper with live scons children reads ORPHANED
  rather than clear.
- --status / --wait / --abort. --abort kills the whole tree by pgid.
- pio runs under setsid with its pgid recorded; INT/TERM/HUP kill the tree
  and record RESULT: ABORTED (exit 5), log kept.
- One result() for every verdict: prints to the stdout the script started
  with (a signal can arrive inside a redirected pio call) and writes
  .pio/runtests/last-result.tsv with head, args, env, finish time, kept log
  and a tree fingerprint (HEAD + working-tree diff + untracked files).
  --status marks the last verdict STALE when the tree has changed since.
- Banner naming the final RESULT: line as the only verdict.
- Non-Linux host: RESULT: UNSUPPORTED, exit 6, instead of the AMBER code.
- A failed build removes .pio/build/<env>/meshtasticd, which run bare would
  reprint the last good run.
- FILTERED lists the not-run count, not 77 suite names.

bin/run-tests.cmd forwards into WSL with the exit code passed through, so the
same command line works from cmd.exe and PowerShell; no logic is duplicated.

test/README.md, copilot-instructions.md and the mirrors document the new
codes and the rule.

* test(harness): one-suite warm-up, and name the build phase instead of freezing the counter

Measured on a full native run: the warm-up, `pio test --without-testing`
with no filter, builds AND links every suite - 78 links, 2320 s, 29.7 s
each, 39 minutes before the first test ran - and prints a "[PASSED]" line
for each program it merely linked. CI never did this; its warm-up is one
`platformio run`. The shared src objects are the same whichever suite links
them, so the warm-up now links one: the filtered suite when there is one,
else test_utf8. The run itself still builds every suite, as it must.

The heartbeat counted objects newer than its marker, which sits still through
PlatformIO's single-threaded scons dependency scan and through each link -
twelve minutes at "430/754 objs, ETA 17m" on that run, which reads as a hung
build to a caller who cannot run ps. It now names the phase from the
processes in the recorded group: [scons] / [compile] (with the ETA) / [link]
/ [test], and --status prints the same phase word.

* test(harness): define phase_of_run before --status can call it

* test(harness): bind --wait to the run it observed; serialize the run-record publish

Review findings on #11862, all three valid:

- --wait stored a state and then waited for any record to clear, so a run
  that finished and a second that started between polls would be followed to
  the second's verdict, and a run that turned ORPHANED mid-wait could print a
  stale last-result. Each run now has an id (pid-start) in current.tsv and
  last-result.tsv; --wait captures it and reports only a matching verdict,
  else ABORTED-without-verdict.
- run_state() then the current.tsv write was a check-then-act pair: two
  invocations in the same instant could both see IDLE. The pair is now one
  critical section under a short-lived flock, and both records are published
  by rename so no reader can see a partial file. The record stays the
  ownership token; the lock only serializes the handoff (a SIGKILLed holder
  releases flock but not the record, which is why flock alone was rejected).
- The FILTERED and AMBER examples in copilot-instructions.md carried a
  literal suite count, which the same document says never to do.

Verified on a live run: --status RUNNING, a concurrent start refused BUSY, a
--wait started before --abort reported the aborted run's own verdict with the
matching id, exit 5; no build process survived.

* fix(waypoints): do not create an empty store file on clear

clearAllWaypoints() wrote a two-byte empty store unconditionally, so test_waypoint_expiry left
Waypoints_default.wpts behind in every run and the suite has read AMBER (undeclared shared state)
since it landed. An existing file - stale or unreadable included - is still rewritten as empty, so
a reset after a failed load clears flash as before; a file that is not there is left not there.
2026-09-22 11:10:53 +00:00
Austin 43479aa4e1 Actions: exclude mutable action tag rule from Semgrep scans (#11944)
We are not pedantic enough to pin Actions versions to a SHA.
Re-enable semgrep scans on previously-ignored .github/workflows and disable
yaml.github-actions.security.github-actions-mutable-action-tag.github-actions-mutable-action-tag
2026-09-21 19:09:28 +00:00
Austin bf2ef8f2d6 Actions: Fix setup-python v7 version pinning (#11942)
We pin official actions to their major-version floating tag, we don't care about minor-revs here.
2026-09-21 14:56:30 +00:00
github-actions[bot]andjp-bennett 9fe0cd8c19 Update protobufs (#11936)
Co-authored-by: jp-bennett <5630967+jp-bennett@users.noreply.github.com>
2026-09-21 06:41:35 -05:00
Tom f49cc46476 docs(agents): list nRF54 as a platform, drop the phantom nRF52833/nRF52832 (#11865)
* docs(agents): list nRF54L15 as a platform, drop the phantom nRF52833/nRF52832

The supported-platform list named nRF52833 and the warm-tier section
excluded "bare nRF52832"; neither part exists anywhere in the tree. The
nRF54L15 (ARCH_NRF54L on the nRF52 platform layer, out-of-tree
meshtastic/platform-nordicnrf54 with the s145 SoftDevice core) does,
with xiao_nrf54l15_lr2021 as the per-PR canary, and was missing.

The warm tier's real exclusion is MEM_CLASS_TINY (STM32WL only); the
nRF54L15 falls into MEM_CLASS_SMALL and keeps its warm records in
/prefs/warm.dat, not the nRF52840 raw-flash ring.

* docs(copilot-instructions): nRF54 has its own platform layer; warm.dat only where WARM_NODE_COUNT > 0

Matches #11867: src/platform/nrf54 (architecture.h, main-nrf54.cpp), env base nrf54_base in
variants/nrf54l15/nrf54.ini, sharing only NRF52Bluetooth.cpp, Nrf52SaadcLock.cpp, alloc.cpp and
hardfault.cpp with src/platform/nrf52. The persistence line named the condition instead of
"everywhere else", which read as including STM32WL, which has no warm tier.

* docs(copilot-instructions): the platform is nRF54; no chip qualifier

* docs(copilot-instructions): state the warm-tier gate before the exception

The section opened with the arch condition folded into a parenthetical
beside STM32WL - "On every arch except STM32WL (`WARM_NODE_COUNT > 0`;
the only `MEM_CLASS_TINY` part)" - which reads as though STM32WL is the
part with the tier, contradicting the bullets below it.

Lead with the rule instead. The gate is the memory class, not the part:
mesh-pb-constants.h zeroes WARM_NODE_COUNT under MESHTASTIC_MEM_CLASS <=
MEM_CLASS_TINY, and STM32WL is merely the only part in that class today,
so naming the part as the rule would go stale the moment another tiny
board lands. Saying it once up front also lets the Persistence bullet
drop its own copy of the caveat.

* ci: run the PR matrix on docs-only PRs so ci-gate reports

ci-gate is the required status check for develop, and it lives in this
workflow, whose pull_request trigger ignored "**.md". A PR that touches only
Markdown never runs the workflow, never reports ci-gate, and stays BLOCKED
with every other check green - which is what this PR has been for a week.

Drop the paths-ignore on pull_request only. push keeps its filter: nothing
gates on the post-merge run. A docs-only PR now costs one --level pr matrix,
which is rare enough to be cheaper than an unmergeable class of PRs.
2026-09-21 10:05:25 +00:00
Austin d49cf21c3a Actions: Add a basic retry for PPA uploads (#11923) 2026-09-20 18:18:14 +00:00
Thomas Göttgens 46e009d66d fix(portduino): detach the CH341 poll thread when it detaches its own interrupt on Windows (#11882)
* fix(portduino): detach the CH341 poll thread when it detaches its own interrupt on Windows

* fix(portduino): stop a superseded CH341 poll thread on Windows

* fix(portduino): wait for CH341 poll threads before deinit closes the device

* fix(portduino): stop a superseded CH341 poll thread before it rewrites pin state

* fix(portduino): keep a same-thread re-arm's CH341 pin state sentinel intact

* fix(portduino): close the CH341 attach/deinit race and recognize a superseded poll thread

* style(portduino): trim the new CH341 comments to the two-line house limit
2026-09-20 17:20:03 +00:00
Ben MeadorsandWayenWeng d4bb6eea91 fix(gps): wake AG3335 from software RTC sleep (#11889)
* fix(gps): wake AG3335 from software RTC sleep

On the Airoha trackers, GPS_HARDSLEEP simply dropped PIN_GPS_EN. Cutting VCC
with no prior command leaves the receiver in hardware RTC mode, from which
nothing in the firmware ever brings it back: the only Airoha wake sequence
that existed, wakeAirohaForActiveProbe(), is reachable from probe() alone and
never from the normal GPS_HARDSLEEP -> GPS_ACTIVE transition. The tracker
stops producing fixes until it is rebooted.

Park the receiver in software RTC mode with $PAIR650,0 before the power cut,
and pulse GPS_RTC_INT after VCC comes back to bring it out again. The receiver
may already have auto-slept and missed the first command, so it is resent
until it acks, bounded at 400 ms rather than repeated a fixed number of times:
setPowerState() runs on the GPS thread and from the notifyDeepSleep observer,
and every mesh thread shares loopTask on nRF52, so an unconditional wait here
stalls LoRa servicing, the screen and buttons for its full duration on every
sleep cycle, on top of holding the receiver powered that much longer.

The RTC_INT pulse becomes a shared helper so the probe path and the power
state machine no longer carry separate copies, and the wake is guarded on
GPS_RTC_INT as well as GNSS_AIROHA, since the family flag is not a promise
that the board routed that line.

Also drops the raw digitalWrite(PIN_GPS_EN, LOW) calls that followed
writePinEN(false) in GPS_HARDSLEEP and GPS_OFF, and the one in
toggleGpsMode(). writePinEN() already drives the pin through the GpioVirtPin
chain built in createGps(); the raw writes duplicated it while bypassing both
that abstraction and the RAK4631/WISMESH_TAP guard inside writePinEN().

Affects tracker-t1000-e, seeed_mesh_tracker_X1 and wio-t1000-s.

Co-Authored-By: WayenWeng <jinyuan.weng@seeed.cc>

* fix(gps): keep the $PAIR650 retry inside its stated budget

The while-condition was evaluated after getACK, so a final attempt starting
just under the budget could add another ack window on top of it. Stop starting
attempts once a whole window no longer fits, which makes 400 ms a real ceiling
rather than a soft one.

* fix(gps): gate the AG3335 park on a wake path, the probed model, and a miss count

Review found three holes in the soft-RTC park, all from review by Thomas
Goettgens.

The sleep was guarded on GNSS_AIROHA while the wake was guarded on
GNSS_AIROHA && GPS_RTC_INT, so a board that did not route RTC_INT would have
been parked with no way back out - worse than the bare power cut this is
meant to fix. Both now derive from a single HAS_AIROHA_SOFT_RTC, so they
cannot be guarded separately again.

The retry stamped 'start' from raw millis() but tested it with
Throttle::isWithinTimespanMs, which reads Time::getMillis(). Under
Time::setTestMillis() the two diverge and the loop either falls through or
never exits; getACK just below already uses the injectable clock.

The park was gated only at compile time, so a probe that fell back to
GENERIC_NMEA would still be sent PAIR650 and block for the full budget every
cycle. It now checks the probed model the way the constellation setup at
L975 does, and gives up after three consecutive unacked sleeps so a receiver
that is present but wedged cannot stall loopTask indefinitely.

* fix(gps): drive the Airoha probe wake from the capability, not the board

Requested by Manuel Verch. wakeAirohaForActiveProbe() asserted EN and pulsed
RTC_INT only under TRACKER_T1000_E, so seeed_mesh_tracker_X1 and wio-t1000-s
got a bare $PAIR382 during probe even though both route the same pins. Keying
it on HAS_AIROHA_SOFT_RTC gives every board that routed RTC_INT the physical
wake, which is what the probe's own hardware reset needs undone, and lets a
future variant opt in by declaring the pins rather than by name.

The 1000 ms $PAIR382 repeat loop goes with it: it was compensating for the
pulse being absent at this point, so a single command after the pulse is
enough, and T1000-E boots a second sooner.

A board that declares GNSS_AIROHA without routing RTC_INT keeps the bare
command. Waking and parking must match, per the previous commit, but probing
only reads, so it cannot strand the receiver the way the park can. The EN
write is guarded on PIN_GPS_EN the way createGps() already guards its own.

---------

Co-authored-by: WayenWeng <jinyuan.weng@seeed.cc>
2026-09-20 17:18:01 +00:00
Manuel d971d507c4 Thinknode M9 V2: keyboard and GPS (#11905)
* thinknode v2 keyboard and GPS

* update device-ui commit reference
2026-09-20 15:29:51 +00:00
Tom 332c4d7c6f Narrow the ad-hoc NodeInfo greeting (#11897)
* feat(nodedb): greet only while the node store is under half full

The ad-hoc greeting in MeshService::handleFromRadio() was gated on
!isFull(), so a node kept sending unsolicited NodeInfo right up to the
last free slot - on a dense mesh that is the regime where the store is
already churning and the greeting is least likely to buy a lasting
entry.

Add NodeDB::isHalfEmpty(), true only when strictly more than half the
slots are free, and gate the greeting on it instead. The comparison is
written as 2 * numMeshNodes < cap so a half-full store reads false with
no integer rounding, and MAX_NUM_NODES is read into a local because
portduino resolves it through a runtime call.

The helper keeps the MINIMUM_SAFE_FREE_HEAP term that !isFull() used to
contribute: low heap disqualifies the store regardless of occupancy, so
a sparse database on a memory-starved device still does not transmit.

Admission is untouched - updateFrom() and getOrCreateMeshNode() still
fill to capacity. Only greeting stops early.

* fix(nodeinfo): raise the minimum greeting window to 30 minutes

The !shorterTimeout branch of NodeInfoModule::allocReply() used a
10-minute base, so a node that had just greeted one neighbour could
greet the next ten minutes later. Raise the base to 30 minutes.

This is the floor, not the window: getConfiguredOrDefaultMsScaled()
still multiplies by the congestion coefficient for the roles that scale,
so a busy mesh stretches it further. ROUTER/ROUTER_LATE and the
tracker/sensor roles bypass the scaling and get a flat 30 minutes.

The interactive paths are unaffected - they pass shorterTimeout and keep
their own 60-second gate. The periodic broadcast is unaffected too:
default_node_info_broadcast_secs is 3 hours with a 1-hour minimum, both
clear of the new floor, so the timer is not swallowed by the throttle.

* fix(nodeinfo): a send restarts the routine broadcast countdown

sendOurNodeInfo() left the OSThread schedule alone, so an ad-hoc send
had no effect on the periodic broadcast: run() anchors the next run at
runned() + interval, and nothing re-anchored it when the send came from
a greeting, a PKI decrypt failure or a completed key verification. The
routine copy could follow minutes behind an ad-hoc one, putting two
NodeInfos on the air for no gain.

Call setIntervalFromNow() with the configured broadcast interval once
the packet is queued, so the next periodic copy is a full interval from
the send rather than from the last tick.

It sits on the return-true path only: a send vetoed by allocReply() -
throttle, airtime ceiling, reply suppression - must not be able to
silence the routine broadcast. Calling it from inside runOnce() is
harmless, since run() then applies the same interval from a last_run of
effectively now.

* test(nodeinfo): cover the send window, the countdown reset and the greeting gate

Three behaviours from this branch had no coverage: isHalfEmpty()'s exclusive
boundary, the 30-minute send floor, and the countdown reset on a send.

isHalfEmpty() goes to test_nodedb_blocked, which already owns the full-store
cases and clears the hot store per test. Three tests sweep the cap over the
sizes real deployments have - portduino resolves MAX_NUM_NODES from
General.MaxNodes on every read, so a predicate that cached it would greet at the
wrong occupancy - and pin the band where admission outlives greeting. That suite
had no tearDown; it has one now, restoring the cap so an assertion firing
mid-sweep cannot leak a 2-node cap into the tests after it.

test_nodeinfo_send_window is new because nothing in the tree stands up
NodeInfoModule's send path. Six tests: the floor at 30 minutes with 10 refused,
the interactive 60-second gate staying separate, the countdown re-armed by a
broadcast and by an ad-hoc unicast, left alone by a refused send, and a preset
change consumed only by a send that goes out.

The scaling above 40 online nodes is deliberately not retested here -
getConfiguredOrDefaultMsScaled() is test_default's contract, per preset and per
role. These tests pin the base and leave the multiplier alone.

NodeInfoModule gains two PIO_UNIT_TESTING accessors for the countdown:
concurrency::OSThread is a private base, so a test shim cannot reach it and only
the class itself can. They compile out of a shipping build.

The heap term in isHalfEmpty()/isFull() stays uncovered: memGet.getFreeHeap()
returns UINT32_MAX on portduino, so a native test could only pin a stub.

* chore(trunk): exempt test_nodedb_blocked from the trufflehog Lob detector

test_removeNodeByNum_presentNodeOnFullDb is exactly 35 characters after the
test_ prefix, which is the length of a Lob API key, and trufflehog's detector
matches the bare identifier. The name is years old; it surfaces now only because
this branch touches the file, and the pre-push gate reports a finding in a
changed file as new.

Added to the ignore block that already carries the same detector's hex-literal
false positives, with the reason stated alongside them. Nothing in that file is
a credential.

* fix(nodeinfo): exempt a licensed station from the floor, delay only on a real send

Two review findings on the 30-minute window.

Ham mode sets node_info_broadcast_secs to 600 s for the FCC minimum call-sign
announcement (AdminModule.cpp). The new floor refused every one of those sends
until 30 minutes had passed, so a licensed station's call sign went out three
times less often than the regulation asks - a regression the old 10-minute base
did not have. A licensed station now keeps its own interval whenever that is
shorter than the floor. The exemption is exactly the licensed case because
nothing else can get under the floor: a set-config clamps the field to an hour,
and the userprefs path clamps identically.

sendOurNodeInfo() ignored what sendToMesh() returned, so a packet the router
declined - no interface, queue full - still re-armed the routine broadcast and
still reported success, which let runOnce() consume a pending channel change
for a send that never reached the air. Only ERRNO_OK and ERRNO_SHOULD_RELEASE
now count; sendToMesh() has already released the packet in both cases.

Both are pinned by tests that fail without them, measured: the licensed case
fails at "11 min is past it, and the floor must not override it", the declined
send at "a declined send is not a send". The licensed test carries an unlicensed
control on the same configuration, so deleting the floor outright would not
satisfy it.

* test(nodeinfo): assert the deadline the scheduler reads, from an aged last_run

The countdown cases asserted Thread::interval, which is not what schedules the
next run: shouldRun() keys off _cached_next_run, and the two ways of writing it
differ. setIntervalFromNow() recomputes it from now; Thread::setInterval()
recomputes it from last_run. Swap the call in sendOurNodeInfo() for the latter
and the period still reads three hours while the deadline lands wherever the
last tick was - firing the routine copy right behind an ad-hoc send, the exact
thing the reset exists to prevent. Every test passed.

Assert the deadline instead, from a fixture where the two answers are
distinguishable: ageLastRunForTests() calls Thread::runned() with an hour-old
timestamp, the state a periodic thread is genuinely in between runs, so a
deadline off last_run lands an hour early against a five second tolerance.

Measured: with setInterval() in place of setIntervalFromNow(), the new case
fails by 3600004 ms and the eight others pass, including the one asserting the
period - which is what says the old assertion could not see this.

runned() and _cached_next_run are protected in Thread and OSThread is a private
base, so the hooks live on NodeInfoModule, with the two already there.

Raised by Copilot on #11897.

* fix(nodeinfo): a declined send must not start the throttle window either

allocReply() stamped TransmitHistory when it built the packet, before anything
had been sent. The previous commit made sendOurNodeInfo() report a router
rejection instead of swallowing it, but the stamp was already written by then,
so a packet that never reached the air still started the window - and with the
floor now at 30 minutes, that silences the node for half an hour over a send
that failed.

allocReply() has two callers and only one of them can see the outcome: the
module framework sends its own reply through currentReply, with no post-send
hook a module can reach (MeshModule::sendResponse is not virtual). So the stamp
stays there for that path, and sendOurNodeInfo() defers it across its own
allocReply() call and stamps once the router has accepted the packet.
deferHistoryStamp mirrors the shorterTimeout member alongside it - same
call-scoped signal, same lifetime.

test_sendWindow_aRejectedSendDoesNotStartTheWindow asserts both halves: no stamp
after the rejection, and the retry immediately after goes out. The existing
rejected-send case checked the first failure and the countdown only, which is
how this survived it.

248/248 across every suite that touches NodeInfoModule (admin_session_repro,
admin_radio, nodeinfo_send_window, traffic_management, fuzz_packets) plus
transmit_history, whose subject this is.

Raised by CodeRabbit on #11897.
2026-09-20 10:32:41 +00:00
TomandBen Meadors 20f7ab1be9 Relay a PKI unicast with a known party in LOCAL_ONLY and KNOWN_ONLY (#11898)
* fix(router): relay a PKI unicast with a known party in LOCAL_ONLY and KNOWN_ONLY

A frame the relay cannot decrypt takes the OPAQUE_RELAY_ONLY path in
Router::perhapsHandleReceived() and returns before handleReceived(), so it never
reaches a module. The rebroadcast_mode rule for opaque traffic is therefore the
IS_ONE_OF list in relayOpaquePacket(), and LOCAL_ONLY and KNOWN_ONLY were not on
it: a node in either mode dropped every opaque frame, including a PKI unicast
with a party it knows. The gate in RoutingModule::handleReceivedProtobuf() that
used to allow exactly that is unreachable for these packets and no longer
decides anything.

Add both modes to the list, with the identity rule the unreachable gate carried:
a PKI-shaped unicast (channel 0, not broadcast) with `from` or `to` known to us.
An unreadable broadcast and a unicast between two strangers stay dropped in
these modes, which is what the proto documents - they ignore foreign meshes.

This is not only direct messages. Remote administration and key verification are
PKI unicasts too, and a KNOWN_ONLY relay was black-holing those between two
other nodes just the same.

CORE_PORTNUMS_ONLY reached this list the same way in #11843; this is the
remaining pair.

* test(rebroadcast_mode): pin the relay decision per mode, in both directions

What this node carries for other nodes, per DeviceConfig.rebroadcast_mode, for
packets it can read and packets it cannot. The harness pushes a real PKI frame
through Router::perhapsHandleReceived() and counts what reaches the radio.

Both directions are load-bearing, and each is guarded by cases the other leaves
green. Measured by rebuilding the firmware three ways:

  - as shipped: 9/9 pass.
  - with the two modes taken back out of relayOpaquePacket(): the known-party
    and remote-admin cases fail, the stranger and foreign-mesh cases still pass.
  - with the modes listed but the identity qualifier deleted: the stranger and
    foreign-mesh cases fail, the known-party cases still pass.

So a revert and an over-broadening each fail their own tests, and neither can be
satisfied by breaking the other.

The suite and its harness come from the opaque-packet-handling branch and
compile against develop unmodified. Two assertions were dropped because they
pin behaviour this branch does not add: delivery of unreadable frames to the
phone, and a queued copy of our own suppressing the originator's repeat, which
needs relayOpaquePacket() to consult the TX queue.

* fix(test): guard the harness include, drop a comment its test outlived

Two review findings on the imported suite.

The harness builds real PKI frames through CryptoEngine entry points a
MESHTASTIC_EXCLUDE_PKI build does not declare, but it was included above the
guard, so the empty-suite branch that exists for those builds could not compile.
Move the include inside the guard and pull the base includes the stub branch
needs above it.

The phone-delivery case was dropped when the suite came across - this branch
does not deliver unreadable frames to the phone - but its comment stayed behind
and now described the test below it, which is about the signature policy.

* test(rebroadcast_mode): make the from-known and channel-0 operands load-bearing

The qualifier has three operands, and the suite only exercised one of them.

Every existing case that a known party carries has a known DESTINATION: the
remote-admin case marks both parties, the known-destination case marks the
target. Rewriting the identity test to consult p->to alone passed all nine. And
the only non-PKI-shaped frame in the suite was a broadcast, which
!isBroadcast(p->to) rejects before p->channel is ever read, so deleting the
channel gate passed all nine too.

Add the two cases that close it: a known SOURCE with a destination we have never
heard of, which must relay in both modes, and a known party on a channel hash we
do not hold, which must not - that is someone else's channel traffic, addressed,
not PKI.

Measured both ways. With the from operand and the channel gate removed from
relayOpaquePacket(), the two new cases fail and the other nine pass; with the
shipping code, 11/11.

Raised by CodeRabbit on #11898.

---------

Co-authored-by: Ben Meadors <benmmeadors@gmail.com>
2026-09-20 10:29:48 +00:00
Jonathan BennettandClaude Opus 5 d3b4b343e7 Send an ack over PKC when no channel can carry it (#11891)
* Send an ack over PKC when no channel can carry it

PKI needs only the two keys, so a DM can reach us over a channel we do not
carry. Its ack is a ROUTING packet, which wouldEncryptWithPKC() excludes, so
today it is channel-encoded, fails at setActiveByIndex() with NO_CHANNEL, and is
never sent. The sender sees nothing and retransmits to exhaustion for a message
that was in fact delivered.

Fall back to PKC for exactly that case. This is the one place an ack is
deliberately made opaque to relays; normally that costs next-hop learning and
intermediate retransmission cancel, which is why ROUTING is PKC-excluded in
general, but here there is no readable alternative to lose, because without this
the ack does not exist.

The predicate is scoped as tightly as that argument reaches: a unicast ROUTING
packet we originate, carrying a request_id, to a destination whose key we hold,
under the same ham/sim/private-key preconditions PKC always has, and only when
the channel index does not resolve. It tests channels.getHash() rather than
setActiveByIndex() so it has no side effect; generateHash already returns -1 for
an invalid key, so the two agree on which indexes are unusable.

Four cases in test_packet_signing pin the corners: the fallback fires, it does
not paper over an ack with no destination key, it does not catch a non-ack on
the same unusable channel, and an ack on a channel that does resolve still goes
out readable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012hLcVif8GDEmA2k77hmFG8

* Short-circuit the fallback so an out-of-range channel logs no error

wouldEncryptWithPKC() reaches channels.getName(chIndex) before its portnum
exclusion, and getByIndex() logs "Invalid channel index" on the way past. With
the general predicate tested first, an ack on an out-of-range index printed that
error and then went on to encode successfully. Test ackFallback first so the
case that is about to succeed never asks.

Also record why the range check leads inside the predicate: getHash() is a bare
hashes[i] with no bounds test of its own.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012hLcVif8GDEmA2k77hmFG8

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-20 05:21:06 +00:00
Jonathan BennettandClaude Opus 5 a0c230e091 fix(baseui): stop the connection footer erasing the last body row (#11918)
drawCommonFooter() black-fills the bottom (connection_icon_height + 2) rows
before drawing the API-connected link icon. On colour builds that fill spans the
whole screen width; the monochrome path fills only the icon's own width.

The body grid does not shrink with the panel. textSixthLine is 58 rows down
whatever the display height, while the footer band starts at
SCREEN_HEIGHT - 1 - connection_icon_height. On a 64-row panel that is row 58 --
exactly the sixth body line -- so the bar erases the last thing the frame drew.
On BaseUI that is the sixth body row, the LoRa frame's ChUtil bar (rows 52-59)
and the bottom of the clock. Taller panels put the band clear of the grid and
are unaffected: at 80 rows it starts at 74, and on high-res it is far below.

It only shows once a client is connected, since the function early-returns on
!isAPIConnected(), which is why it reads as the blue link icon eating the
screen rather than as a layout bug.

Only the icon's own rect is registered for colour tinting, so the wide fill buys
the tint nothing. Keep it where it clears the body, and fall back to the
icon-width fill -- what the monochrome path already does -- where it would not.
Panels with room below the body are byte-identical to before.

Seen on a 128x64 HUB75 running meshtasticd; esp32s3/visualizer-hub75 is the
same geometry.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 04:12:35 +00:00
Jonathan BennettandClaude Opus 5 ee02cc3426 Games joystick input (#11917)
* fix(games): correct the high-score announcement argument order

GAMES_HIGH_SCORE_STRING is "New %s high score %lu by %s!" but the arguments
were passed as (name, initials, score): the initials string was formatted
through %lu and the score integer through %s. That is a format/argument
mismatch, so the announcement printed garbage at best and dereferenced the
score as a pointer at worst.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(input): report which physical gamepad button produced an event

A joystick event only carried the action it was mapped to, so a consumer could
not tell two buttons apart once they shared one action, and games were limited
to the handful of actions the broker defines.

Carry the originating evdev button code in InputEvent::kbchar, encoded into a
reserved 0xC0..0xDF range that misses printable ASCII and every
INPUT_BROKER_MSG_ value (SystemCommands switches on kbchar without looking at
inputEvent, so a collision there would reboot the node rather than move a
paddle). D-pad events are axes, not buttons, and keep leaving kbchar at 0 --
which is exactly what lets a consumer tell stick from button.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(portduino): let one joystick action bind several buttons

Input.JoystickButtons took a single evdev code per action, so a pad's A and Y
could not both select, and the shoulder buttons could not sit alongside the
D-pad. Accept a list of codes as well as a bare scalar; the config writer
inverts its code->action map back out, emitting a list only where an action
has more than one button.

ConfigCheck gains a real checker for the section (it was previously waved
through as free-form) covering the three ways a mapping silently does nothing:
an action name the driver does not know, an evdev name where the numeric code
belongs, and one code claimed by two actions. Two fixtures and shell-test
cases cover the clean list form and those three faults.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(games): use the gamepad's extra buttons, and return home when idle

Games now receive the physical button alongside the action, so a pad with more
than two usable buttons controls more than two things:

- Snake: a shoulder button mapped to left/right turns relative to the snake's
  heading (L counter-clockwise, R clockwise) while the D-pad keeps steering
  absolutely. The two are told apart by kbchar, not by hardcoding one pad's
  codes.
- Breakout: the ball now rides the paddle after each serve until the player
  fires it with B or A, so a life is not lost to a ball already in flight when
  the player looks up. The paddle also keeps its position between lives. A game
  can claim BACK for the duration (Game::wantsBackButton) so B serves instead
  of pausing, and releases it once the ball is live.
- Start (BTN_BASE4 / BTN_START) is mapped to select like any other button, so
  it launches games and drives the menus; inside a running game GamesModule
  picks it out of kbchar and pauses instead.

Separately, the games frame no longer holds a walked-away device hostage: after
15 s with no input it returns to the home frame, so the device still reads as a
Meshtastic node. The timer is suspended while a picker or banner is up (e.g.
high-score initials entry, which the input handler never sees) so it cannot
yank the user out mid-entry.

Screen::isInteractionBusy() generalises the old module-intercept check --
modal module, intercepting module, game, or an open interactive overlay --
and MessageRenderer uses it before popping an incoming-message banner. A
transient banner REPLACES an active overlay, so an arriving message could
otherwise discard a half-entered high score. The message is still stored, its
thread still selected, and the unread indicator still set.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ui): compose freetext on the on-screen keyboard from a gamepad

A gamepad can drive the on-screen keyboard but cannot type, so on a host with
a joystick and no configured keyboard device the OSK is the only way to compose
freetext. Set osk_found there, and gate the "Freetext" menu entries on whether
the device can enter text at all (physical keyboard, OSK, or touchscreen
virtual keyboard) rather than on kb_found alone -- those entries were hidden on
exactly the devices that needed them.

The OSK prompt that CannedMessageModule already had inline in the message
selector becomes showOnScreenKeyboard(), so the menu path can reach it too.
Menus call in from a banner callback and the banner is torn down as soon as
that callback returns, which would take the keyboard down with it, so the menu
path defers the launch to runOnce().

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(games): address review on frame fallback and joystick input gating

Breakout: the paddle suppression was far too broad. aLinuxJoystick is
constructed on every Linux host whether or not a gamepad is configured
(InputBroker.cpp), so `aLinuxJoystick && kbchar == 0` was true everywhere and
swallowed LEFT/RIGHT from the keyboard, trackball and ExpressLRS -- on a host
with no joystick attached at all. Gate on the stick actually driving the paddle
instead: LinuxJoystick assigns heldX before it emits and only auto-repeats while
heldX is set, so every axis LEFT/RIGHT arrives with a zone held and nothing else
does. kbchar == 0 still distinguishes an axis from a shoulder button mapped to
left/right, which must keep nudging the paddle.

Screen: showHomeFrame() did nothing when the home frame was hidden, since
setFrames() only assigns positions.home for !hiddenFrames.home. That stranded
the games inactivity bounce on the frame it was trying to leave. Fall back to
the messages frame, which setFrames() always adds.

Test: rename test_ballWaitsOnPaddleUntilLaunched to
test_ball_waitsOnPaddleUntilLaunched, matching the repo convention and its
neighbours in the file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(games): give games the whole InputEvent so Breakout can identify the source

Follow-up to review on #11917. The previous narrowing still could not tell
sources apart: kbchar == 0 is shared by the joystick's D-pad axis and by every
other driver that sends a bare LEFT/RIGHT, so while the D-pad was held a
keyboard or touchscreen press was still discarded. heldXZone() proves the axis
is driving, not that this particular event came from it.

Pass the event itself to Game::handleInput() rather than (ev, kbchar). Games
that only care about the action read event->inputEvent; Snake keeps using
kbchar for shoulder steering; Breakout now also checks event->source against
LinuxJoystick's origin name, so only that driver's own axis repeats are
suppressed.

Chose the event over a third positional parameter so the signature does not
have to grow again the next time a game needs something the event already
carries.

All three conditions in Breakout are load-bearing: source says it came from
this gamepad, kbchar == 0 says it is the axis rather than a shoulder button
mapped to left/right, and heldXZone() != 0 says the axis is what is driving
right now so tick() already has it covered.

LinuxJoystick::originName() exposes the name the driver stamps into
InputEvent::source, alongside the existing heldXZone()/heldYZone() accessors.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 04:11:18 +00:00
renovate[bot] fa87e7d29c chore(deps): update libch341-spi-userspace digest to d85aceb (#11907)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
2026-09-19 08:44:00 +00:00
github-actions[bot]andvidplace7 24dc0e9d62 Upgrade trunk (#11909)
Co-authored-by: vidplace7 <1779290+vidplace7@users.noreply.github.com>
2026-09-19 06:21:30 -05:00
github-actions[bot]andvidplace7 28f5610922 Upgrade trunk (#11369)
Co-authored-by: vidplace7 <1779290+vidplace7@users.noreply.github.com>
2026-09-18 22:44:25 +00:00
James Rich 2a398fb2fd Cancel superseded runs of Check PR Labels and Semgrep Differential Scan (#11906)
* Cancel superseded runs of Check PR Labels and Semgrep Differential Scan

Both workflows trigger on pull_request without a concurrency group, so
every push, label or edit on a PR queues a fresh run while the earlier
ones keep their place in the org's shared runner queue. Check PR Labels
listens to six event types, so opening, labelling and editing one PR
queued three identical runs within a minute today.

Give each the same head_ref-keyed group with cancel-in-progress that
CI, Tests and the trunk check already use. Develop pushes are
unaffected: they fall through to run_id, as before.

* Key the new concurrency groups by PR number, not head_ref

Fork PRs often come from a branch named develop, so two of them share
github.head_ref and one would cancel the other's run. The PR number is
unique per PR.
2026-09-18 20:34:59 -04:00
rcarteraz 60f82b4488 Fix Mesh Node T1 device image filename (#11893)
custom_meshtastic_images pointed at heltec-mesh-node-t1.svg, which does
not exist in web-flasher. The file there is heltec-meshnode-t1.svg. The
wrong path returns HTTP 200 with a 7 KB HTML error page rather than the
203 KB SVG, so it does not surface as a 404.
2026-09-18 14:58:11 +00:00
Thomas Göttgens f13faa6aba fix(esp32s3): bound the SerialConsole idle sleep on hardware USB CDC (#11901)
HWCDC::isPlugged() is a SOF watchdog that reads false transiently while USB
is connected and working. runOnce() answered that with a 20 s sleep, and
nothing wakes the thread on RX, so host traffic sat in the CDC RX ring and
reached the API as a burst. Cap the sleep at 250 ms, the rate readStream()
already idles at.

IS_USB_SERIAL only tested ARDUINO_USB_CDC_ON_BOOT, so ARDUINO_USB_MODE=0
boards ran the same check against a USB-Serial/JTAG peripheral that is not
attached to the PHY and never sees a SOF. Gate the check on IS_USB_HWCDC.

Measured on tlora-t3s3-v1, 900 s of 1 Hz ToRadio/FromRadio round trips:
before 12 stalls, rtt_max 19.96 s, console asleep 26.4% of wall time.
After 0 stalls, rtt_max 0.147 s, p50 unchanged at 0.028 s.

Fixes #11864
2026-09-18 14:46:43 +00:00
Thomas Göttgens d96c690a91 fix(detect): check LPS22HB before SFA30 at 0x5D and CRC-validate SFA30 probe (#11881)
The SFA30 probe only compared the requestFrom() length, which equals the
requested length for any device that ACKs, so an LPS33HW/LPS35HW at 0x5D
was reported as SFA30. Probe WHO_AM_I first so ST sensors never receive
the SFA30 command, and require valid Sensirion CRC-8 on every word of the
device marking response.

Fixes #11880
2026-09-18 12:19:37 +00:00
Thomas Göttgens c29bd00971 Rewrite the MQTT region root topic only on the default broker (#11899)
* fix(mqtt): rewrite the region root topic only on the default broker

A region change rewrote any root starting with "msh", on any broker. Custom roots such as msh/home were clobbered, and private brokers had their topics moved even though a regional broker is regional already. An empty root, which MQTT treats as the default, was never updated.

Region changes now go through MQTT::applyRegionRootTopic(), which rewrites the root only on the default broker and only when the root is empty, the default, or a msh/<region> the firmware wrote itself.

* fix(mqtt): parse the broker address before the default-server check

Persist module config only when the root actually changed. Replace a stray NUL byte in the test with the \0 escape.

* fix(mqtt): count the regional roots as the default root topic
2026-09-18 10:34:53 +00:00
Thomas GöttgensandClaude a5dce941dd fix(telemetry): follow AS3935Config rename to AS3935State (#11903)
meshtastic/protobufs#1045: admin.proto's AS3935_config and
telemetry.proto's AS3935Config collide after name mangling. The
flash-persisted message is renamed to AS3935State upstream.

Field numbers are unchanged, so existing /prefs/as3935.dat files still
decode.

Depends on the protobufs rename and the regenerated sources landing
first.

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-18 17:32:26 +02:00
github-actions[bot]andcaveman99 11550fa3bd Update protobufs (#11904)
Co-authored-by: caveman99 <25002+caveman99@users.noreply.github.com>
2026-09-18 16:41:30 +02:00
Benjamin Faershtein 3aac179397 fix(phoneapi): wake clients after config sync (#11818) 2026-09-18 08:50:20 +00:00
Thomas GöttgensandBen Meadors 0dafcc90fe fix(api): retain the unwritten tail on a short TCP API write (#11890)
* fix(api): retain the unwritten tail on a short TCP API write

ServerAPI closed the session whenever stream->write() returned fewer bytes than
requested. A short write is transmit-buffer backpressure, not a dead socket, and
it is most likely during the back-to-back frames of the initial NodeDB dump, so
a node at its node cap dropped clients on effectively every connect.

Route TCP frames through StreamFrameWriter, the retained-tail path the USB CDC
console already uses: the remainder is re-offered on the next pass and the
session is closed only when the link itself is gone. Poll at 25ms while output
is still undelivered, since nothing wakes the thread when the socket frees
transmit space.

Fixes #11822

* fix(api): block log re-encoding while a TCP frame is retained

emitLogRecord() writes into txBufLog and StreamFrameWriter can now hold that
buffer as a retained tail, so a second log record would overwrite bytes the
transport has not sent yet. Gate encoding on the retained-frame state, matching
SerialConsole.

No caller reaches this today (emitLogRecord() is only used by SerialConsole),
but retaining the buffer at all is new here.

---------

Co-authored-by: Ben Meadors <benmmeadors@gmail.com>
2026-09-18 06:54:50 +00:00
renovate[bot] d4166d0426 chore(deps): update libch341-spi-userspace digest to eaaef01 (#11895)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
2026-09-18 06:45:22 +00:00
Ben Meadors 585ce17f59 fix(nrf52): stop concurrent flash writers corrupting LittleFS, and stop a failed save formatting it (#11872)
* fix(nrf52): serialise the warm-node ring against LittleFS on the shared flash cache

On nRF52840 the warm-node store writes its 3-page record ring straight through
flash_nrf5x_write/erase/flush, holding only spiLock. Every LittleFS writer instead
holds Adafruit_LittleFS's own mutex, and two of them run on other tasks entirely:
Bluefruit's bond saves on the callback task, and - since phone config writes moved
into BLE context - a whole saveToDisk on the BLE task. Neither takes spiLock.

Both writers share one 4 KB page cache, one SoftDevice flash semaphore and one
result word. flash_cache_write repoints that cache when the requested page differs
from the cached one, so a second writer arriving mid-write flushes the first
writer's page and re-points the buffer; the first writer's remaining memcpy then
lands in the wrong page's image. Ring records end up inside LittleFS metadata, or
the reverse. The collision also exhausts the flash layer's 20 x 1 ms busy-retry
budget against an 85 ms page erase, and flash_cache_flush discards the failure, so
32 LittleFS blocks vanish with no error reaching the filesystem. What the user sees
is a torn directory pair on the next mount, a format, and critical error 13.

Take the filesystem mutex in the five ring entry points that reach flash, after
spiLock and never before - the order every existing path already uses. The ring
touches no LittleFS call itself, so the non-recursive mutex is never re-entered.

Longest new hold is a page rotation at roughly half a second, against a 2 s
supervision timeout and a 90 s watchdog. Non-nRF52840 backends are untouched.

* fix(nodedb): make saveProto report a failed readback or rename

SafeFile::close() already verifies the .tmp by hash and renames it over the
live file, and saveProto captured that result, logged it, and then returned
the pb_encode status alone. A torn or half-programmed page therefore counted
as a successful save: for the fullAtomic files the old contents silently
survived, for nodes.proto (written in place) the file was simply gone, and
saveToDisk's recovery path never fired for the one failure it exists for.

* fix(nodedb): retry a failed save before formatting, and never format on a low rail

saveToDisk answered any failed write with an immediate fsFormat(), which is
where most "critical error 12/13" reports and the total config wipe behind
them come from. A write that fails once is far more often a busy SoftDevice
or a VDD dip mid-save than a corrupt filesystem, so:

- retry twice, 150 ms apart, re-checking powerHAL_isPowerLevelSafe() before
  each attempt and before the format; on a low rail return false and leave
  the filesystem alone (the next save lands once the rail recovers, and boot
  already waits for a safe level)
- check fsFormat()'s result instead of assuming it worked
- after a successful format rewrite every segment, not only the ones this
  call asked for: the format took config.proto and the node identity with it,
  so a nodes-only save that ended in a format used to come back up as a new
  node
- with encrypted storage a format also destroys the DEK; skip the resave
  rather than land the private key and PSKs on flash in plaintext

RP2040 feeds its watchdog across the delays, as the neighbouring code does.

* fix(nrf52): quiesce flash before every software reset and power-off

The Adafruit flash layer keeps one 4 KB page image and one SoftDevice flash
semaphore for the whole chip. Every reset path we own - Power::reboot(),
enterDfuMode() (admin enter_dfu_mode_request, which arrives on the BLE task
since #10967), cpuDeepSleep()'s reset and system-off arms, and the
wio-t1000-s secure DFU handler - went straight to NVIC_SystemReset or
sd_power_system_off while another task could be half-way through a page
program or erase. A reset in that window leaves the page erased or partly
programmed; LittleFS finds the torn metadata on the next mount and the
corruption handler formats the filesystem.

nrf52FlashQuiesce() takes spiLock and the LittleFS mutex, waits out whatever
write is in flight, flushes the page cache, and keeps both locks because the
caller resets next. The corruption-reboot handler and __assert_func are left
alone: they run inside the filesystem call stack or a fault, where taking the
mutex would deadlock.

nRF54L is a second copy of these paths since #11867 and still defines
ARCH_NRF52, so it gets the same function on the same core flash layer.

* fix(nrf52): quiesce flash before the library BLE DFU handler jumps to the bootloader

On every board except wio-t1000-s the Nordic DFU service is the framework's
BLEDfu, whose START_DFU handler runs on the callback task and jumps to the
bootloader with no regard for a flash write in progress on the loop task.
That is the OTA path the Apple app and nRF Connect use (Android sends
enter_dfu_mode_request instead, which the previous commit covers).

QuiescingBLEDfu re-installs the control-point write callback after
BLEDfu::begin() and wraps the library's: flush under both locks, then drop
the LittleFS mutex before handing over, because the library reloads the bond
keys through LittleFS on its way to the jump and the mutex is not recursive.
spiLock stays held across the handler: every LittleFS writer on the BLE task
takes it first, the loop task cannot preempt the callback task, and the
handler never blocks after the flush, so nothing can dirty flash before
bootloader_util_app_start(). If the handler returns, nothing jumped, and the
lock is released.

The library callback is a file-static, so it is read back out of the
characteristic through a pointer-to-member obtained via a using-declaration;
that is well-formed C++ and compiles under the pinned GCC 9.3 with LTO.

* test(nodedb): pin the save-failure contract of saveProto and saveToDisk

A failed rename must come back as false from saveProto, a one-off unsafe
rail reading during a write must be retried and land, and a rail still
unsafe at the retry gate must make saveToDisk return false with the
filesystem untouched. The rail is scripted through a strong
powerHAL_isPowerLevelSafe() over the weak native default; on Windows the
default is strong, so only the rename case runs there. The format branch
itself is unreachable natively (a FLASH_CORRUPTION critical error exits
the portduino process), which is what the survival assertions pin.

* fix(nodedb): only format when the filesystem itself is unreadable

Making saveProto honest about write failures gave the recovery path a new way in:
any persistent write failure now reached fsFormat(), which takes every file with
it. A busy or lock-protected nRF52 flash fails every write for as long as it lasts,
so two retries are not enough to tell that apart from a corrupt filesystem, and
guessing wrong costs the node its config, keys and bonds.

Reads settle it. They never touch the SoftDevice write path that a busy flash
fails on, so if /prefs still walks and a stored proto still opens and reads, the
metadata chain is intact and the write failure was transient - return false and
let the caller try again later. Genuine corruption is not silently tolerated: lfs
asserts on it, and the nRF52 handler reboots and formats on the way back up.

Covered by a test that fails without this: a save whose rename cannot succeed,
against an otherwise healthy filesystem, must leave devicestate untouched.

* trunk: exempt Unity test entry points from trufflehog

trufflehog's Lob detector matches "test_" followed by alphanumerics, which
describes every Unity test function name. It fired on a new test in
test_nodedb_save_retry and will fire again on the next suite added. Scoped to
test/**/test_main.cpp, alongside the existing gitleaks exemption for the
synthetic node-DB fixtures.

* fix(nodedb): feed the RP2040 watchdog around the format and the resave

saveToDisk() only feeds the watchdog at the top of each retry. The last
retry, the readable probe, fsFormat() and the five-segment resave then
share one 8 s budget (watchdog_enable in main-rp2xx0.cpp) with no loop
left to feed it. A timeout during the resave leaves the filesystem empty
and the node boots on defaults with a new identity - the exact outcome
this PR exists to prevent, reached by a different road.

Feed once before the probe and again before the resave. Both feeds sit
outside any lock: filesystemStillReadable() takes spiLock itself, and
the format has already released it. ARCH_RP2040 covers rp2040 and rp2350
alike, and the blocks compile out everywhere else, so no other platform
and no native test changes.

Raised by @caveman99 in review.

* fix(nodedb): narrow the save-probe comment and name the full-filesystem case

The comment on filesystemStillReadable() claimed "real corruption asserts
in lfs and formats on reboot". That does hold on nRF52 - nrf52.ini builds
with -DLFS_NO_ASSERT and force-includes cpp_overrides/lfs_util.h, whose
LFS_NO_ASSERT arm routes LFS_ASSERT to the lfs_assert() in main-nrf52.cpp,
which stamps NRF52_MAGIC_LFS_IS_CORRUPT and resets into the format - but
NodeDB.cpp compiles for ESP32, RP2040 and portduino too, where nothing of
the sort is wired up. It is also not true on nRF52 under POFWARN, where
lfs_assert() deliberately skips the stamp. Drop the claim rather than
qualify it three ways.

The log line now names what a field log actually needs to tell apart: a
filesystem that still reads but cannot be written is either busy or full.

Raised by @caveman99 in review.

* fix(sx128x): quiesce flash before the 2.4GHz region reset

reinitChip() saves the region, waits 2 s and resets. On nRF52 that was
the last software reset still going straight to NVIC_SystemReset with a
page program possibly in flight, so "every software reset" in the earlier
commit did not quite hold.

The quiesce stays inside the ARCH_NRF52 arm on purpose. The #else arm
logs and falls through to lora.setCRC() further down, which re-enters
spiBeginTransaction(); a quiesce hoisted above the #if would take spiLock
and never give it back, self-deadlocking portduino and stm32wl. Routing
this through Power::reboot() is wrong for the same class of reason:
setupModules() runs before initLoRa, so its notifyReboot observers and
waypointStore.saveToFlash() are live and would add a flash write to an
aborted radio init.

Raised by @caveman99 in review.

* fix(nrf52): only quiesce on the DFU control write that actually resets

QuiescingBLEDfu wrapped every control-point write, so a write that was
never going to reset still blocked the Bluefruit callback task on spiLock,
forced an early page-cache commit and held back the GATT authorize reply.
Only START_DFU resets; gate on that.

Deliberately no "request->len &&" term. The library's own test is
`request->data[0] == START_DFU` with no length check (BLEDfu.cpp:110 in
both the nRF52 and nRF54 cores), and Bluefruit hands the callback a copy
of a reused event buffer, so a zero-length write carrying a stale 0x01
still resets inside the library. A len term here would let exactly that
reset run unquiesced, which is the case this wrapper exists for. Reading
data[0] is always in bounds: ble_gatts_evt_write_t declares uint8_t
data[1] and the copy covers it.

Raised by @caveman99 in review.
2026-09-17 22:31:32 +00:00
Jonathan BennettandClaude Opus 5 6f3f0bd7c2 Prove explicit acks with Routing.ack_proof (#11877)
* Prove explicit acks with Routing.ack_proof

Explicit acks are ROUTING_APP packets, and ROUTING_APP is excluded from PKC, so
an ack travels under channel encryption alone - and the default channel key is
public. Anyone in range can forge one, and the client grants its strongest
delivery claim on the strength of the ack's unauthenticated `from`.

Where the acknowledged packet was PKI encrypted the endpoints already share a
Curve25519 secret, so the recipient can prove receipt in ~10 encoded bytes:

    ack_proof = HMAC-SHA256(shared_key,
                            "ack" | LE32(from) | LE32(to) | LE32(request_id)
                                  | routing)[0..8)

where `routing` is the encoded Routing message without the ack_proof field,
taken as received with that byte range removed rather than re-encoded. Excising
keeps the value a function of the received bytes alone, so it does not depend on
two implementations' encoders agreeing and does not drop fields this build has
never heard of. For the same reason the sender appends the field rather than
setting it on a decoded struct and re-encoding.

What this does and does not buy. It buys an authenticated delivery receipt from
the actual recipient, which is the property a forged ack costs a user and which
matters where people act on a delivery confirmation. It does NOT protect the
retransmission loop, and must not be described as if it does:
perhapsGenerateImplicitAckForOwnOverheard clears a pending retransmission on any
overheard rebroadcast of our own (from, id), header-only and keyless, so
replaying the originator's own ciphertext stops their retries more cheaply than
forging an ack. No ack authentication of any kind closes that path.

Advisory only, and deliberately not a step toward enforcement. A rule requiring
a proof once a peer has sent one would make a missing proof destroy the only
delivery signal we have, on state the user cannot see: the proof needs the peer
to hold our key, and peer-side eviction, a downgrade or a factory reset are all
invisible to us. A valid proof marks the ack verified; anything else behaves
exactly as today. The client renders the difference.

Verification costs one X25519 and nothing caches the shared secret, so
ackProofPermitsAction gates on findPendingPacket first - otherwise a forged ack
naming any packet id, which is visible in the cleartext header, would force a DH.

Depends on meshtastic/protobufs#1094 for the generated field.

* test: correct a comment that predates ack_proof being generated

The wire-roundtrip test described ack_proof as an unknown field, which was true
while the prototype hand-encoded it. The field is generated now, so this build
understands it - but older firmware does not, which is the case the assertion
actually covers.

* Fix the EXCLUDE_PKI test build, and format

Two copies of test_proof_binds_error_reason and test_proof_binds_direction were
sitting in the #else branch of test_ack_proof, where Identity, makeIdentity,
makeAck, becomeNode, crypto and ACK_PROOF_SIZE do not exist. They were also
unregistered, so they were dead code that only served to break the build. The
native suite never compiles that branch, so a local run could not see it.

Also apply clang-format to a declaration that had been wrapped by hand.

* Exempt test_ack_proof from trufflehog's Lob false positive

Same detector and same shape as the three suites already listed here: it
stitches nearby hex literals into one candidate string, and this suite's node
numbers and request ids (0x0A0A0A0A, 0x0B0B0B0B, 0xABCD1234) happen to match a
Lob API key. The suite holds no literal key material - every key it uses comes
from crypto->generateKeyPair at runtime.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012hLcVif8GDEmA2k77hmFG8

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-17 22:24:29 +00:00
54334ff936 Initial firmware support for Axiometa Genesis Mini (#11852)
* Initial firmware support for Axiometa Genesis Mini

* Report the real AXIOMETA_GENESIS_MINI hardware model

The protobufs now carry AXIOMETA_GENESIS_MINI = 148, so the board no
longer has to masquerade as private hardware.

On ESP32 the -D PRIVATE_HW flag never selected the model on its own -
architecture.h has no arm for it, so the board fell through to the
PRIVATE_HW default at the end of the chain. Give it its own arm and drop
the flag.

* fix(input): sample the encoder button after light-sleep wake

The edge that wakes the device lands while beforeLightSleep() has the
interrupts detached, and a button that is still held produces no further edge
until it is released. The thread stayed parked at INT32_MAX, so the press was
never sampled - no event was emitted, PowerFSM's GPIO-wake branch reads
BUTTON_PIN rather than the encoder pin, and the node dropped straight back into
light sleep with the press swallowed entirely. The second press worked, the
first did not.

Sample once on wake, and only when the button is asserted, so a timer or radio
wake leaves the thread alone. The press then follows the ordinary path and
InputBroker drops the event because the screen was off, so it wakes the screen
and does nothing more - the same behaviour every other input device has.

Rotation stays deliberately non-waking: only the button pin is armed in
doLightSleep(), and the abState re-seed discards a shaft moved during sleep
rather than replaying it as detents.

Reported by CodeRabbit on #11852.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: rcarteraz <robert.l.carter2@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 22:22:05 +00:00
Austin 683b3533fd setup-base: Move python dependencies into a requirements.txt, pin versions for caching (#11876)
Moves python dependencies declared in `setup-base` to a separate requirements.txt file (which renovate can keep updated) and aligns it with gh-action-firmware.

Also removes `pio upgrade`, this is pointless (we just installed the latest platformio in the previous step)
2026-09-17 21:12:59 +00:00
Simplycissmus ff3cc66827 fix(rp2xx0): log and reset on a failed assert instead of hanging (#11853)
RP2xx0 had no __assert_func, so newlib's ran: it prints to stdio and abort()
reaches arduino-pico's _exit, a breakpoint loop. Before rp2040Loop() arms the
watchdog that hangs the node until power is removed; afterwards it costs a
silent stall of up to 8 s.

Install one along the lines of the nRF52 handler: log the failed expression
and reboot through watchdog_reboot().

Refs #11795
2026-09-17 20:52:32 +00:00
github-actions[bot]andjp-bennett 7c3c730a50 Update protobufs (#11887)
Co-authored-by: jp-bennett <5630967+jp-bennett@users.noreply.github.com>
2026-09-17 18:34:04 +00:00
Jason P a4e8b9444f baseui_fixfavoritesonsmalllcds (#11883) 2026-09-17 15:46:42 +00:00
rcarteraz 7bfa062f27 fix(esp32): raise T-Watch S3 to support level 1 (#11884) 2026-09-17 15:46:38 +00:00
renovate[bot] dbbd42687b chore(deps): update meshtastic/crypto digest to 1c817c2 (#11879)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
2026-09-17 12:03:28 +00:00
renovate[bot] d1cefc8382 chore(deps): update fusion digest to 2051197 (#11878)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
2026-09-17 12:03:19 +00:00
Thomas Göttgens f36a1ea821 nRF52: reclaim flash to bring rak4631 back under its size budget (#11873)
* build(nrf52): drop unused TinyUSB classes and assert function-name strings

Only the CDC class is used on nRF52. Disable the MSC, HID, MIDI, vendor and
video class drivers in the Adafruit TinyUSB config, and pass an empty
__ASSERT_FUNC so assert() no longer embeds __PRETTY_FUNCTION__ strings.
File and line are still reported.

rak4631 estimate: ~7.4 KB flash, ~2.3 KB RAM.

* fix(nrf52): link only the secp256r1 cc310 curve domain

CRYS_ECPKI_GetEcDomain indexes ecDomainsFuncP, which references the
parameter tables of all eleven cc310 curves. Bluefruit LESC pairing only
requests secp256r1, so override the lookup to return that domain alone.

rak4631 estimate: ~7.4 KB flash.

* fix(airtime): replace powf in the channel-utilization EMA fold

foldChannelUtil was the only powf caller on nRF52. The exponent is an
integer step count, so raise the EMA factor by squaring instead; a
multi-day sleep still folds in at most 32 multiplications.

rak4631 estimate: ~1.9 KB flash.

* fix(graphics): use double sin/cos in the compass renderers

The compass renderers were the only sinf/cosf callers on nRF52 screen
builds, pulling in the float trig kernels next to the double ones GeoCoord
already links. Call the double variants instead.

rak4631 estimate: ~3.2 KB flash.

* fix(motion): use double atan2 for magnetometer heading fallbacks

MMC5983MA, QMC6309 and the InkHUD map centre were the remaining
application atan2f callers. The double atan2 is already linked, so the
float variant only added atan2f, __ieee754_atan2f and atanf. The saving
lands once meshtastic/Fusion#1 removes the library's atan2f as well.

rak4631 estimate: ~0.8 KB flash with Fusion#1.

* fix(hopscale): trim diagnostic logging to state changes and anomalies

Drop the save/restore confirmations, the hourly histogram and trend dumps,
the denominator step logs and the per-packet hop_limit log (printPacket
already reports HopLim). Keep the save-failure and histogram-full warnings,
the congestion on/off transition and a single periodic status line, and
remove lastScaledPerHop, which only fed the logs.

* fix(hopscale): silence cppcheck uselessAssignmentArg on restored count

* perf(crypto): use full-schedule AES128/AES256 for AES-CCM

aesSetKey used AESSmall128/AESSmall256, which re-derive round keys for
every block. AES128/AES256 precompute the schedule, encrypt faster and are
already linked by encryptAESCtr, so the AESSmall*/AESTiny* code drops out.
No change on ESP32, where AESSmall* already aliases AES128/AES256.

rak4631 estimate: ~3.9 KB flash; cipher object up to 184 bytes larger.

* perf(nrf52): use the shared software CTR for AES-256 and remove tiny-aes

CryptoCell only accelerates AES-128, which stays on hardware. AES-256 CTR
now calls CryptoEngine::encryptAESCtr (rweather CTR<AES256>, already
linked) instead of the in-tree tiny-aes copy, whose sources were removed
in the previous commit. Output is identical.

rak4631 estimate: ~0.8 KB flash.

* perf(mesh): use std::map for pending retransmissions and API port timestamps

NextHopRouter::pending and PhoneAPI::lastPortNumToRadio were the only
unordered_map instances linked on nRF52. Switching them to std::map, which
is already linked, drops the libstdc++ hashtable, rehash policy and prime
table. GlobalPacketId gains operator<; the unused hash functor is removed.

rak4631 estimate: ~2.1 KB flash.

* perf: parse sensor decimals without strtod

The WS85 serial parser (strtof) and DFRobotLarkSensor (String::toFloat)
were the only callers of newlib's strtod. Add parseDecimalFloat to
meshUtils for plain [+-]digits[.digits] fields and use it at both sites.
Covered by test_type_conversions against strtof.

rak4631 estimate: ~4.5 KB flash.

* perf(gps): compute tan from sin/cos in UTM and OSGR conversion

latLongToUTM and latLongToOSGR were the only tan callers. sin and cos are
already linked, so deriving tan from them drops tan and __kernel_tan.

rak4631 estimate: ~1.1 KB flash.

* fix(graphics): only dispatch the theme menu when TFT coloring is enabled

The Theme option is only offered with GRAPHICS_TFT_COLORING_ENABLED, but
handleMenuSwitch dispatched ThemeMenu unconditionally, linking kThemes and
the theme accessors into monochrome builds where the menu is unreachable.

rak4631 estimate: ~1 KB flash.

* fix(senxx): trim diagnostic logging to errors and user-visible actions

Keep all errors and warnings and a single version line; shorten the admin
action messages; drop progress chatter, state save/restore confirmations and
the per-reading and VOC-state debug dumps. The nested VOC restore branch
collapses to one condition with the same behaviour.

rak4631 estimate: ~2 KB flash.

* build(nrf52): define CRYPTO_AES_NO_DECRYPT

CTR and CCM only encrypt, so the AES inverse tables and round helpers are
dead code on nRF52. Takes effect once the Crypto dependency includes
meshtastic/Crypto#5.

rak4631 estimate: ~1.0 KB flash.
2026-09-17 09:36:41 +00:00
Thomas Göttgens 67e8aafef7 fix(http): hold spiLock only for filesystem calls in the HTTP file handlers (#11870)
* fix(http): hold spiLock only for filesystem calls in the static and upload handlers

* fix(http): hold spiLock only for filesystem calls in the browse and delete handlers

* fix(http): abort an upload when a write comes up short
2026-09-16 22:01:55 +00:00
ae8dee9582 Fix: MQTT topic not updated when LoRa region changes (#10565)
* Initial plan

* Fix: update MQTT topics when LoRa region changes

When the LoRa region is changed via AdminModule::handleSetConfig,
moduleConfig.mqtt.root is updated (e.g. from msh/US to msh/EU_868)
but the running MQTT instance kept using the stale topic strings
(cryptTopic / jsonTopic / mapTopic) that were set at construction time.

Introduce MQTT::reinitTopics() which:
- resets the topic strings to their base values and prepends the
  current moduleConfig.mqtt.root, and
- disconnects from the broker so the next reconnect re-subscribes
  under the new topic prefix.

Call reinitTopics() from MQTT's constructor (replacing the inline
block) so the logic lives in one place, and call it from
AdminModule::handleSetConfig right after moduleConfig.mqtt.root is
rewritten on a region change.

Add a unit test (test_reinitTopicsUpdatesOnRegionChange) that verifies
both the updated subscriptions and the updated publish topic after a
simulated region change.

* Fix: call mqtt->reinitTopics() on region change via menuhandler

* Format MQTT test with trunk style

* fix(mqtt): drop undeclared jsonTopic refs in reinitTopics()

reinitTopics() assigned to a jsonTopic member that does not exist on this
branch (the MQTT class only has cryptTopic and mapTopic), so MQTT.cpp failed
to compile ("'jsonTopic' was not declared in this scope") and broke every
build that compiles it. Remove the jsonTopic lines so reinitTopics() rebuilds
exactly the topics the original constructor did.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* refactor(mqtt): simplify reinitTopics and its call sites

* fix(mqtt): rebuild topics in runOnce when the root changes

Replaces the per-call-site reinitTopics() calls. Also keep the device state and node database segments when the EU clamp swaps the region.

* fix(mqtt): refresh topics in onSend when the root changed

Rename the region change test to underscore-separated segments.

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Ben Meadors <benmmeadors@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Thomas Göttgens <tgoettgens@gmail.com>
2026-09-16 21:47:27 +00:00
Thomas Göttgens ce7e6e448d Separate the nRF54 platform code from src/platform/nrf52 (#11867)
* Move the nRF54L platform code into src/platform/nrf54l15 and drop the ARCH_NRF54L branches from src/platform/nrf52

* Name the platform directory nrf54 so future nRF54 variants can share it

* Rename the nRF54 platform base to nrf54_base in variants/nrf54l15/nrf54.ini

* Leave the SoftDevice random seed to the Bluefruit core, which seeds in begin() and answers NRF_EVT_RAND_SEED_REQUEST

* Seed the SoftDevice from checkSDEvents() when it pops NRF_EVT_RAND_SEED_REQUEST
2026-09-16 21:21:41 +00:00
github-actions[bot]andjp-bennett 7469d52900 Update protobufs (#11874)
Co-authored-by: jp-bennett <5630967+jp-bennett@users.noreply.github.com>
2026-09-16 23:19:47 +02:00
Jonathan Bennett 19dfa1385b Use new store-and-forward original_id field (#11849) 2026-09-16 17:28:24 +00:00
Ben MeadorsandThomas Göttgens ff9a2b1b86 fix(http): give a TLS session the contiguous heap it actually needs (#6960) (#11838)
* fix(http): give a TLS session the contiguous heap it actually needs (#6960)

An HTTPS connection to a no-PSRAM ESP32-S3 fails its handshake with
MBEDTLS_ERR_SSL_ALLOC_FAILED and the client gets a TCP reset. Four things
were in the way, and the first two are the bug.

mbedtls_ssl_setup() makes two independent calloc() calls, one inbound
record buffer and one outbound, each needing its own contiguous block of
16717 bytes. The admission gate compared ESP.getFreeHeap(), which is a sum
over every free block and says nothing about the largest one, so on a
fragmented heap it waved the connection through and let the handshake be
the thing that discovered there was no room.

Size the two record buffers separately. Inbound stays at the RFC maximum
because browsers and our own MQTT and OTA clients receive full-size
records; outbound only ever carries what this device sends, 4 kB at a time
at most. That is 12288 bytes of contiguous heap back per session, which is
the difference between one usable HTTPS connection and two. These are the
ESP-IDF defaults - Espressif's arduino-lib-builder is what turns the
asymmetric split off. It replaces CONFIG_MBEDTLS_SSL_MAX_CONTENT_LEN,
which no longer applies: that symbol is `depends on
!MBEDTLS_ASYMMETRIC_CONTENT_LEN`, and mbedTLS dropped it in 3.0.

Then ask the allocator the question mbedTLS is about to ask it, rather
than a different one. The probe holds both record buffers at once, as
setup() does, through the same heap that CONFIG_MBEDTLS_*_MEM_ALLOC pins
mbedTLS to - a plain malloc would be served from PSRAM that mbedTLS never
touches. Not heap_caps_get_largest_free_block(), which walks every block
under the allocator lock and tripped the interrupt watchdog in #11666.
Closed connections are reaped before the measurement, so the library
cannot free a slot behind it and accept into it unmeasured, and so the
probe sees the memory a finished session just returned. The answer is
advisory: the heap can still turn between the probe and setup(), which
makes the failure rare rather than routine.

A smaller outgoing record makes an existing bug much easier to hit.
mbedtls_ssl_write() accepts at most one record per call and returns a
short count for the rest, and neither the server library nor Print::print()
loops on that, so any response body larger than one record was quietly cut
off - /json/nodes has been truncating at 16 kB already. Loop at the bulk
write sites. Looping on the returned count also covers a client that
negotiates a smaller maximum fragment length, which a fixed-size chunker
would not.

/json/nodes also built its whole body in one std::string. A full node DB
is 200 entries of ~240 bytes, and std::string grows by doubling, so it
asked for a 64 kB contiguous block while still holding a 32 kB one, on a
board with ~50 kB of heap, through operator new, which aborts rather than
throws here. Stream it a node at a time instead.

Finally, stop every node in the fleet paying for geofencing. The crossing
table reserved all 256 slots at construction, 4 kB, though a node with no
geofenced waypoints never tracks a crossing at all. Allocate on the first
one instead, stepping the capacity explicitly so there is a size to probe
for, and hand the block back once the last waypoint is gone. The probe
matters because this allocation has moved off the pristine boot heap onto
a live one, where a reserve() that cannot find the block would abort;
a refusal now degrades into the bounded-drop path that already exists.

heltec-v3 builds clean from scratch with the IDF recompiled, and comes out
20,656 bytes smaller in flash and 172 bytes smaller in static RAM. Full
native suite is green at 1414 cases.

* style(http,geofence): cut the added comments to the repo's one-or-two-line rule

The rationale belongs in the commit message and the PR, not in block comments
above every new symbol. No behaviour change; the binary is byte-identical.

* fix(geofence): index the capacity probe write to clear cppcheck uninitdata

* fix(http): stop the /json/nodes body at the first failed write

* fix(geofence): grow crossing state with realloc so a refused allocation drops instead of aborting

* refactor(http): probe through esp_mbedtls_mem_calloc and fold the free-slot scan into reaping

* refactor(geofence): allocate the crossing table once on first use and free it on store changes

* fix(http): probe through mbedtls_calloc so the calls keep C linkage

* fix(http): name the check that refused a TLS session, and sum the right heap

Both failure modes logged the same line, so a refusal on the free-heap
threshold read as fragmentation. The probe returns a verdict now and the
warning names the check that failed.

That threshold also read ESP.getFreeHeap(), the internal heap, while the probe
below it goes through mbedtls_calloc() - PSRAM wherever EXTERNAL_MEM_ALLOC
sends mbedTLS, which [device-ui_base] sets. There is no free-size query on the
allocator, and the handshake's remaining small blocks want a sum rather than a
block to probe for, so the sum asks heap_caps for the same heap.

* fix(http): probe the TLS heap only when a client is waiting for a slot

The probe holds ~21 kB across two callocs and ran whenever a slot was free,
which on a board with nobody connecting is every 50 ms while the web server is
active and every second when it is idle - on exactly the low-heap boards this
branch is for.

It now runs only when a zero-timeout select() on the listen socket says a
client is already in the backlog, which is the question loop() asks a few
microseconds later. With nothing knocking the server drives its open
connections and accepts nothing, so a connection is never accepted past an
unmeasured gate; one that lands between the two select()s waits a tick.

* perf(http): flush /json/nodes a couple of kB at a time

One writeAll() per node is one mbedtls_ssl_write() per node, about 200 of them
for a full node DB. Filling a 2 kB buffer first cuts that by roughly 9x and
still bounds the allocation, which is why the whole-body string went away in
the first place.

* fix(geofence): keep the table-full warning for a full table

A refused malloc shared the one-shot table-full warning, so a momentary heap
dip spent the message that means the 256-entry cap was reached, and the real
thing then never printed. They are separate now, and the allocation failure
throttles rather than latching, because the next position may well find the
4 kB the table needs.

* fix(http): close a waiting client the TLS heap check refuses

* fix(http): hold spiLock only for filesystem calls in the browse and delete handlers

* refactor(http): leave the browse and delete spiLock scoping to #11870

---------

Co-authored-by: Thomas Göttgens <tgoettgens@gmail.com>
2026-09-16 16:52:00 +00:00
Manuel 4b3f2a2beb fix drop_stale_sdkconfig_defaults() (#11861) 2026-09-16 16:15:22 +00:00
Austin a70fe9ea10 Actions: Build MacOS 27, Drop 15 (#11851)
MacOS 15 will still be built in homebrew, but not tested as part of CI here.
2026-09-16 14:04:51 +00:00
Tom ee76117835 fix(router): relay opaque packets in CORE_PORTNUMS_ONLY (#11844)
* fix(router): relay opaque packets per rebroadcast_mode, not only in ALL

d6b12ea3f (#10967) moved undecryptable packets onto relayOpaquePacket(), which
relays only in ALL and ALL_SKIP_DECODING. A PKI unicast between two other
nodes - remote admin, a DM, key verification - is opaque to a relay, so a
router in CORE_PORTNUMS_ONLY (the ROUTER role default) stopped carrying any
of it, and KNOWN_ONLY / LOCAL_ONLY lost the rule 6eabbaf43 added in 2024 that
relays a PKI-shaped unicast with one known party. The old rule was still in
RoutingModule::handleReceivedProtobuf, unreachable: encrypted packets no
longer reach modules.

opaqueRelayAllowedByMode() gives each mode an explicit opaque rule:
ALL / ALL_SKIP_DECODING / CORE_PORTNUMS_ONLY relay (the port list cannot apply
to a packet with no readable port); KNOWN_ONLY / LOCAL_ONLY relay a channel-0
unicast whose sender or destination has a User in NodeDB; NONE relays nothing.
The dead RoutingModule block is removed; its licensed-party check stays for
decoded packets. Nothing here reads packet_signature_policy.

* fix(router): NAK, phone delivery and MQTT uplink for packets we cannot read

Before #10967 an undecryptable packet addressed to us reached RoutingModule,
which handed it to the phone and, via ReliableRouter::sniffReceived, answered
a want_ack unicast from an unknown sender with PKI_UNKNOWN_PUBKEY - the NAK
that makes the sender transmit its NodeInfo so its retry decrypts. A PKI DM
between two other nodes was uplinked to MQTT as ciphertext when encrypted
uplink was on. All three stopped: the gate REJECTs a to-us decode failure
before any module runs, and the pki_encrypted marking in dispatchReceived is
unreachable for opaque ingress.

passesRoutingAuthGate() now treats every DECODE_FAILURE not from us as opaque
(isFromUs stays REJECT, #11544). The opaque branch calls handleOpaqueForUs(),
which NAKs a want_ack unicast to us (PKI_UNKNOWN_PUBKEY when we hold no key
for the sender, NO_CHANNEL otherwise) and queues a frame we had no way to
read - PKI without the sender's key, or a channel hash matching nothing we
hold - straight to the phone via sendToPhone(), bypassing handleFromRadio()
so an unverified sender never touches NodeDB. A matched-and-failed frame
(bad key, tampering, junk) is NAKed but not delivered. uplinkOpaqueUnicast()
restores the MQTT path for channel-0 unicasts not to or from us, gated on
mqtt.enabled and mqtt.encryption_enabled; the dead marking is removed.

* test(packet_signing): pin what a relay does with traffic it cannot read

Group R builds genuine PKI-encrypted packets between two generated
identities and runs them through ingress under every rebroadcast_mode:

  R1  remote admin between two known nodes relays in every mode but NONE
  R2  between strangers: KNOWN_ONLY / LOCAL_ONLY decline, the rest relay
  R3  one known party satisfies KNOWN_ONLY / LOCAL_ONLY
  R4  an unknown-channel broadcast relays in ALL / ALL_SKIP / CORE only
  R5  undecryptable DM to us: one PKI_UNKNOWN_PUBKEY NAK, phone gets the
      frame, nothing relayed, sender not added to NodeDB - in every mode
  R6  the same frame claiming to be from us gets no reaction
  R7  sender key held but wrong: NAK NO_CHANNEL, no phone delivery
  R8  opaque PKI unicast is uplinked only with encrypted MQTT uplink
  R9  unknown-channel broadcast reaches the phone without touching NodeDB
  R10 the relay decision is identical under all three signature policies

C6 no longer lists CORE_PORTNUMS_ONLY as a mode that suppresses opaque relay
and expects the phone to see an unreadable frame; the RoutingModule mock
records the NAK reason.

* test(rebroadcast_mode): give the relay policy its own suite, sharing the ingress harness

test/support/AuthPipelineHarness.h now holds the mock NodeDB, the counting radio /
router / routing-module / module / MQTT, the packet builders (decoded, channel-
encrypted, PKI between two generated identities) and the per-process / per-test
lifecycle that test_packet_signing kept locally.

test_rebroadcast_mode pins what this node carries for others, per
DeviceConfig.rebroadcast_mode: remote admin between two known nodes relays in
every mode but NONE; strangers are declined by KNOWN_ONLY / LOCAL_ONLY and
carried elsewhere; one known party suffices; an unknown-channel broadcast relays
in ALL / ALL_SKIP / CORE only; a licensed node never relays ciphertext and
relays plaintext unless a party is known unlicensed; hop_limit 0, id 0, a
foreign next_hop and CLIENT_MUTE each stop a relay; the signature policy
changes none of it. Registered in state-manifest.tsv and the routing shard.

test_packet_signing keeps what follows from the auth gate's verdict: C9-C11 now
carry want_ack and hops so their "nothing happens" assertions are no longer
vacuous (junk on a held channel relays but never reaches the phone; a legacy
DM and a malformed PKI plaintext to us are NAKed NO_CHANNEL once and nothing
else), and C18-C22 cover the to-us NAK / phone / MQTT outcomes. Comments on the
src side trimmed to the two-line rule.

* fix(router): classify an opaque frame from the decode attempt, not the header

handleOpaqueForUs() re-derived "unreadable" from the wire header and NodeDB,
which disagreed with what perhapsDecode() had just found: a hash-0 broadcast
from a sender whose key we hold read as readable although PKI never applies to
a broadcast, and a pending-key decrypt rejected as malformed read as unreadable
because no stored key existed. Both changed the NAK reason and whether
ciphertext reached the phone.

passesRoutingAuthGate() now hands the attempt's DecodeState out and the handler
takes unreadable = (state == DECODE_OPAQUE). For that to be precise,
perhapsDecode() sets pkiAttempted only when a sender, pending or admin key was
actually tried, and the KNOWN_ONLY short-circuit - which declines before any
attempt - reports OPAQUE for a PKI-shaped unicast to us or an unheld hash and
FAILURE for a held channel. isUnreadableToUs() is gone; Channels::hasHash()
replaces its loop.

* test(support): free the harness AirTime and NodeStatus before restoring the originals

pipelineHarnessDestroy() restored the saved pointers and orphaned the two
objects it had installed. LeakSanitizer reported the NodeStatus as a 160-byte
direct leak and errored test_packet_signing and test_rebroadcast_mode at exit
in the coverage shards while every case passed.

* fix(router): a failed admin-key fallback does not count as a decrypt attempt

Every configured admin key set pkiAttempted, so on a node with any admin key
an unknown sender's DM read as DECODE_FAILURE: NO_CHANNEL instead of
PKI_UNKNOWN_PUBKEY, and withheld from the phone, so the sender never learned to
send its NodeInfo. An admin key that fails says nothing about the sender; only
the sender's own (or pending) key counts. A successful-but-malformed admin
decrypt already returns DECODE_FAILURE directly.

* test(packet_signing): pin decode provenance for admin keys, forged from-us frames, hash-0 broadcasts and KNOWN_ONLY strangers

C18 configures an unrelated admin key so the fallback runs and fails. C19
installs our identity so the forged frame is a real decrypt attempt and asserts
the gate's REJECT before the side effects. C23: a hash-0 broadcast from a keyed
sender on a channel we do not hold is unreadable and reaches the phone. C24:
KNOWN_ONLY declines a stranger on a held channel as matched, so the phone never
sees it, while the same stranger on an unheld channel is unreadable.

* test(support): restore the caller's DH key, model the TX queue and ACK/NAK log, clear per-process state between tests

The ingress harness builds its router, radio, routing module and crypto engine
once per process, so anything they carry decides the next test's outcome. Three
of those carried surfaces were already deciding one.

makePkiUnicastBetween() ended by installing a fresh random DH key, so a caller
that set its own key before building a frame silently lost it. C19 did exactly
that: its REJECT assertion passed because the decrypt failed on a key mismatch,
not because the from-us arm rejected the forgery, and would have stayed green
with that arm deleted. CryptoEngine::private_key is public under
PIO_UNIT_TESTING, so the helper now saves the engine's key and puts it back;
the trap is closed for every caller rather than worked around in one. C19 also
builds the frame before installing our identity and gains a control assertion:
without our key in NodeDB the same frame is OPAQUE_RELAY_ONLY, which is what
makes the REJECT attributable.

The opaque dedup ring survived a whole suite unreset. That was tolerable while
it only gated relay; it is about to gate the NAK, phone delivery and MQTT
uplink too, where a stale (from,id) would silently zero a later test's
expectations instead of failing it. Cleared per test, along with the DH key,
any pending handshake key, and the admin-key fallback budget that C18 drains
six tokens from. resetAdminKeyFallbackBudget() is defined under
!MESHTASTIC_EXCLUDE_PKI but declared unguarded, so the call site is guarded.

installOurIdentity() now marks HAS_USER on our own node, as NodeDB does on a
device. The rebroadcast_mode predicates read that bit, and markOurselvesLicensed()
keeps owner.is_licensed and our NodeDB record in agreement for the same reason:
getLicenseStatus(us) must say Licensed, not NotLicensed.

The radio and routing-module mocks become models rather than counters, since
every suite including this header gets them. The radio holds a real TX queue
that findInTxQueue() consults and that cancelSending()/removePendingTXPacket()
take entries out of, and it records each frame it was handed so a test can
assert a hop limit or relay_node instead of a call count. The routing module
keeps every ACK/NAK with its destination, channel and hop limit, so a second
NAK can be pinned without losing the first; its reset() replaces the six sites
that zeroed ackCalls by hand, which would otherwise desync the log from the
counter. Phone-queue draining moves into the harness for the suites that both
need it.

* refactor(router): the opaque path lives in Router, not NextHopRouter

A pure move. Nothing about handling a frame we cannot read is next-hop
specific: relayOpaquePacket() reads iface, isToUs/isFromUs, the device role,
owner.is_licensed, the last byte of our node number and the packet's own
header, then calls Router::send(). The one thing that held it in the subclass
was the (from,id) dedup ring, and that landed in NextHopRouter next to the
pending and route-health tables by proximity rather than dependency - its
whole purpose is to stay isolated from routing state, which argues for sitting
beside PacketHistory instead.

So the ring, opaqueWasSeenRecently(), relayOpaquePacket() and the
rebroadcast_mode predicate (now Router::opaqueAllowedByMode) move down, and
relayOpaquePacket() stops being virtual: there is one router chain, nothing
else overrode it, and the base implementation returned false to no one. The
alternative was a second virtual to reach the same array from the same caller,
which is what the dedup work that follows would otherwise have needed.

isRebroadcaster() comes along because relayOpaquePacket() needs it and it reads
only config.device - no FloodingRouter state - so it was already misplaced.
capEventRelayHops() moves too, and is now declared for NextHopRouter's own
rebroadcast path rather than being file-static.

* fix(router): apply the packet's own rules to every consumer of an opaque frame

Five rules that pre-#10967 applied to a packet we could not read were left
applying to the relay alone. This puts them back on all four consumers - relay,
NAK, phone, MQTT - which is one change of shape, so it lands as one commit
rather than five: the branch head now computes what is true of the frame once
and every consumer below reads the same answer.

Duplicate suppression. The only dedup sat inside relayOpaquePacket(), behind
its isToUs() early return, so it never saw a frame addressed to us. Three
neighbours rebroadcasting a stranger's want_ack DM to us cost three
PKI_UNKNOWN_PUBKEY NAKs on the air, three encrypted frames queued for the
phone, and at a gateway three publishes of every opaque PKI DM between other
nodes. Before these packets stopped going through PacketHistory,
shouldFilterReceived() ran first and made each of those once per (from,id).
The originator's own retransmission keeps its exemption, and gains the two
rules the decoded path already applies to a repeat: do not queue a second copy
while the first is still in the TX queue, and answer again at hop 0, since only
a direct neighbour ever sees hop_start == hop_limit.

rebroadcast_mode. LOCAL_ONLY and KNOWN_ONLY say the node ignores what it cannot
decrypt; the deleted RoutingModule branch gated phone delivery as well as
relay, and only the relay half was carried over, so a stranger's unknown-channel
ciphertext reached the phone in every mode. The phone now follows the same
predicate. NONE still delivers: it means do not relay, not do not listen.

Licence. A licensed station transmits in the clear and may not answer, or hand
on, traffic to or from a node it knows to be unlicensed - the rule RoutingModule
applies to decoded packets, which the opaque path never got.

NAK reason. PKI_UNKNOWN_PUBKEY claimed a missing key even for a channel-0 frame
too short to have carried PKI overhead. Such a frame was never a candidate, so
the reason is NO_CHANNEL, matching the size test perhapsDecode uses.

Uplink. uplinkOpaqueUnicast() read the header only, so a channel-0 unicast that
matched a held hash-0 channel and failed its AEAD was published to the PKI
topic as ciphertext we never tried to read. It now takes the gate's verdict.

One consequence worth stating: the dedup ring records frames the relay gate
used to reject before reaching it - to us, from us, hop-exhausted, mode-blocked,
and everything on a licensed node - so its 32 slots serve four consumers on
nodes that previously never touched it.

Two clean-ups ride along because the same rules move: ReliableRouter's
sniffReceived() loses its own undecryptable-NAK arms, which radio ingress has
not been able to reach since the auth gate started answering those frames
before handleReceived() (deliverLocal, the only other caller, is always
decoded), and test_rebroadcast_mode stops declaring a warm.dat write it never
makes - no case in it reads a signer back from the warm store.

* test(rebroadcast_mode): declare the warm-store write again

The suite stopped writing warm.dat only until this branch gave it a test that
installs an identity of our own, which puts a key through the warm store. The
declaration was removed on the evidence of a run that predated that test, in
the same commit that added it; the harness caught the undeclared write.

* fix(router): put back the undecodable-NAK arms in ReliableRouter::sniffReceived

Deleted as dead code, and they are not. Radio ingress genuinely cannot reach
them any more - the auth gate answers an unreadable frame in
handleOpaqueForUs() and returns before handleReceived(), and deliverLocal(),
the only other caller, is always decoded - but sniffReceived() has a contract
of its own that test_reliable_ack_matrix drives directly, and a local or
SimRadio caller still arrives with an encrypted packet. CI caught it in the
misc-4 shard.

Restored verbatim, with the reachability noted where the next reader will look
rather than in a commit message nobody greps.

* fix(router): classify a PKI-shaped unicast from key material alone

Addresses the review on #11844.

A channel we hold whose hash is 0 matched every PKI DM on the mesh and
failed every one, and that failure was read as "we tried". One unlucky
1-in-256 channel hash therefore withheld every PKI DM from the phone and
the broker, answered NO_CHANNEL where the sender needs PKI_UNKNOWN_PUBKEY
to recover, and classified our own overheard DMs as a forgery so the
implicit "Delivered to mesh" ACK never fired. Hash 0 on a unicast is the
PKI sentinel, so isPkiShapedUnicast() now decides it in one place, used by
the KNOWN_ONLY short-circuit, the decode provenance and the NAK reason.
The isToUs asymmetry in the short-circuit is gone with it.

Also from the review:

- id 0 cannot be deduped by the (from,id) ring, so an undecryptable
  want_ack DM carrying it drew a NAK and a phone frame on every copy
  heard. Every consumer now declines it, as relay already did; a NAK for
  request_id 0 is unmatchable at the sender anyway.
- The ring records only frames some consumer can act on. Our own
  overheard rebroadcasts and unicasts that can neither be relayed nor
  uplinked were evicting live entries, spending the anti-amplification
  bound #11522 added it for.
- The MQTT uplink applies the licensed-station rule, and deliberately not
  rebroadcast_mode: that setting governs what goes back on the air, and
  MQTT has its own switches for what leaves over IP, which is where the
  pre-#10967 uplink sat.
- gateState is initialised rather than relying on the gate's first
  statement.

Comments on the opaque path trimmed to the two-line rule; the reasoning
lives in the tests, which are exempt.

* ci(size-budget): raise the rak4631 flash budget to 748000

The opaque-relay restore lands at 746,080 bytes on rak4631, 80 bytes over
the previous 746,000 limit. Image ends at 0xDC260, 55 KB clear of the warm
region.

* fix(router): keep #10967's removal of the undecodable-frame NAK

The PKI_UNKNOWN_PUBKEY / NO_CHANNEL NAK for a frame we cannot read is not
restored. Every input to that decision - to, from, id, want_ack, hop_start,
hop_limit - is unauthenticated cleartext, so the NAK is a reflector: one
frame in from a node with no key material, one flooded reply out to
whichever `from` it names, with a hop budget the sender chooses. #10967
removed it as an ACK side effect on purpose; the security review of this PR
shows why, and the relay regression it fixes does not need it.

handleOpaqueForUs() now only delivers to the phone. The originator-retx
re-NAK and its `repeat` plumbing go with the NAK. A to-us DECODE_FAILURE is
REJECT again at the gate, as #10967 had it: nothing on the opaque path acts
on a frame we matched and failed on. ReliableRouter::sniffReceived() is back
to develop byte for byte; its undecodable arms are reachable by local and
SimRadio callers only and stay as they are.

A PKI DM to a node that does not hold the sender's key fails silently, as on
develop; the receiving phone still sees the frame. The sender-side recovery
this NAK used to trigger is the follow-on's problem to solve without a
header-driven reply.

Tests: C10, C11, C18, C20, C26, C27, C29, C31 assert no NAK; C28 (the NAK's
hop budget) is deleted; the matched-failure helper drops its want_ack arm.

* fix(router): do not re-decode a frame the gate already classified; scope capEventRelayHops

handleOpaqueForUs() hands the phone a copy the auth gate has already run
perhapsDecode() on. MeshService::sendToPhone() ran it again, and for a
PKI-shaped DM from a sender whose key we lack that second pass re-enters the
admin-key fallback and spends a second token from a budget the code documents
as global and attacker-facing: the sustained rate halved from 4/s to 2/s on
any node with an admin key configured. sendToPhone() takes an
alreadyClassified flag and skips the decode; nothing else calls it that way.
Nodes with no admin key configured were never affected, since
adminKeyFallbackAllowed() returns before touching the bucket.

test_C34 freezes the clock, configures an unrelated admin key, and pins one
token spent per unreadable frame delivered to the phone.
adminKeyFallbackTokensRemaining() is a PIO_UNIT_TESTING accessor beside
resetAdminKeyFallbackBudget().

ReliableRouter::sniffReceived()'s PKI_UNKNOWN_PUBKEY arm now requires the
frame to be long enough to have carried the PKI overhead, the same shape test
Router's opaque classification applies; a shorter channel-0 frame was never a
PKI candidate, so no key was missing and it gets NO_CHANNEL. Reachable by
local and SimRadio callers; pinned in test_reliable_ack_matrix.

capEventRelayHops() moved out of NextHopRouter.cpp as a file-static and
became a free function at global scope. Both callers are Router subclasses,
so it is now a static member of Router next to isRebroadcaster().

* fix(router): size the opaque dedup ring like PACKETHISTORY_MAX

Every ROUTER relays opaque frames now, not just nodes set to ALL, so the
churn through this ring is what bounds a duplicate storm on the backbone.
32 untimed slots shared by three consumers is thin for that; 128 on every
target but STM32WL, which keeps 32 alongside its 20-entry PacketHistory.
8 B per slot in .bss: +768 B, no flash.

* revert(router): return the branch to develop

* fix(router): relay opaque packets in CORE_PORTNUMS_ONLY

A packet a relay cannot decrypt - a PKI unicast between two other nodes, an
unknown-channel broadcast - is relayed only from its header, and only in the
rebroadcast modes relayOpaquePacket() lists. CORE_PORTNUMS_ONLY was not among
them, so a node in that mode, which is the ROUTER role default, dropped every
such packet. The portnum filter that mode exists for cannot be applied to a
payload the relay cannot read; add the mode to the list.

Fixes #11843.

* test(packet_signing): CORE_PORTNUMS_ONLY carries an opaque frame

C6 listed CORE_PORTNUMS_ONLY among the modes that must suppress an opaque
relay, pinning the behaviour #11843 reports. It now asserts the frame is
relayed in that mode, with the same no-side-effect checks as the ALL case, and
keeps LOCAL_ONLY and NONE as the suppressing modes.
2026-09-16 13:14:47 +00:00
Tom 0caf3a09f8 ci(size-budget): raise the rak4631 flash budget 746 000 -> 752 000 (#11866)
develop has been over the 746 000 line since #11826 (748 088 on the
CI runner at 3468af94a), so the size-budget-gate fails on every PR that
does not itself shed 2 KB, and the queue behind it cannot land. 6 KB
covers the current overage plus the pending NodeDB/NodeInfo work
(+584) and one or two of the mid-sized branches, while staying 47 KB
inside the 0xEA000 warm-region guard that is the real wall (image ends
at 0xDCA38 today). Raised deliberately, per the file's own rule.
2026-09-16 11:20:14 +00:00
Tom 2d6dad9ee9 Portduino: Fix LR2021 switch tables, power ceilings and IRQ handling (#11382)
* fix(portduino): recognise the LR2021 power ceilings in --check

loadConfig() has read Lora.LR2021_MAX_POWER and Lora.LR2021_MAX_POWER_HF
since LR2021 support landed, but neither was listed in the config checker's
schema. --check therefore reported both as "unknown key ... ignored by
meshtasticd" -- false, and actively misleading: it tells the user to delete a
key that is doing exactly what they wanted.

This breaks the contract stated above schema(), that a key taught to
loadConfig() is added there too. CI enforces that by running --check over
bin/config.d/**, but no shipped config sets either key -- or mentions lr2021
at all -- so nothing ever tripped over the omission. It could only surface
for someone hand-writing an LR2021 config.

LR20x0 is the only module with two power ceilings, one per band, selected at
runtime by region; every other family expresses the split as separate module
names and needs a single key. That is the likely reason the pair was missed
while every other *_MAX_POWER key was added.

Also adds both to valueSpecs(), so a wrong-typed value is reported rather
than silently replaced by the default.

* feat(portduino): configurable IRQ DIO and a chip-neutral RF switch table for LR20x0

Two gaps found bringing an LR2021 up under meshtasticd on a Luckfox Lyra
Zero W. Both sit in the LR20x0 support added in #11252, and they interact:
the switch table has to be written slightly wrong to pass validation, and
the interrupt lands on a pin that table is driving. Every symptom is silent,
because begin() only exercises SPI and BUSY -- the radio reports init
success and then receives nothing.

IRQ DIO could not be set on Portduino
-------------------------------------
LR20x0Interface picked the IRQ DIO purely at compile time, and neither
LR2021_IRQ_DIO_NUM nor IRQ_DIO_NUM exists for a Portduino target, so
meshtasticd always fell through to RadioLib's default of DIO5 -- which is
also the first RF switch line on carriers using the DIO5-DIO8 table. A
variant says this with a #define (the pro-micro DIY board uses DIO9); a
carrier has only the YAML, and had no way to say it.

Adds Lora.IRQ_DIO_NUM, and an ARCH_PORTDUINO branch after the two existing
#define branches, so a variant that already sets one still wins.

The switch table was parsed as LR11xx-only
------------------------------------------
Pin names resolved to RADIOLIB_LR11X0_DIOn whatever the radio, and the mode
set was the LR11xx's, so MODE_RX_HF -- a mode the LR20x0 really has -- was
rejected as an unknown key and had to be omitted. The two families are not
interchangeable: an LR11xx has no DIO9, so its fifth switch slot is DIO10,
while an LR20x0's fifth slot is DIO9 and DIO10 is its sixth. A table naming
DIO10 was therefore driving the wrong pin on an LR20x0.

The YAML layer now stores what was written -- a DIO number and a neutral
mode id -- and each interface supplies its own DIO constants and OpMode_t
map to a shared builder. Neither family's constants are assumed to coincide
with the other's.

This also fixes a round trip in the config writer, which decoded pins by
comparing against RADIOLIB_LR11X0_* and always emitted five values per mode
row: for a four-pin table it produced YAML that --check would reject for
mismatched row lengths.

--check
-------
Findings are now judged against the resolved module rather than a fixed
list, so a mode or pin the part does have can no longer be rejected, and one
it does not have is named instead of silently accepted. The claim that the
table "is only applied to LR11xx radios" was stale and is corrected, and the
missing-table warning now covers both families. "auto" is excluded
throughout: the module has not been probed yet, so absence cannot be judged.

The IRQ/switch-pin collision is reported in both directions, including the
harder case where no key is set and the radio default collides -- nothing in
the file looks wrong. Note that listing a pin is what breaks it, not driving
it: setRfSwitchTable() reassigns the DIO function for every pin in the list
whatever the levels say, so an all-LOW column is still a collision.

Seven fixtures cover these, including a false-positive guard: DIO5 as the
interrupt is normal, and must stay silent when the table is elsewhere.

* feat(portduino): let the YAML ask for a TCXO probe, across every family that has one

A variant declares "a TCXO may or may not be fitted" at compile time with
TCXO_OPTIONAL, because the board is known when the image is built. A
Portduino carrier cannot: the same meshtasticd binary runs on hardware
populated either way, so the statement has to arrive as YAML and be answered
at runtime.

Adds Lora.TCXO_OPTIONAL, and TCXO_OPTIONAL_ENABLED in RadioLibInterface.h to
unify the two, so each driver asks the question once rather than growing a
second, Portduino-shaped code path. On an embedded target it stays a
compile-time constant, so `if (TCXO_OPTIONAL_ENABLED)` folds away exactly as
the old `#if` did: the nrf52_promicro_diy_tcxo image, which defines
TCXO_OPTIONAL and so exercises the converted branches, still ends at
0xDF1D0 -- the same address as before this change.

It is defined there rather than in a header of its own because
InterfacesTemplates.cpp includes all three interface .cpp files into one
translation unit, where a per-file definition would collide.

Covers every family that has a TCXO reference to probe for: SX126x
(sx1262/sx1268/LLCC68), LR11xx and LR20x0. With no DIO3_TCXO_VOLTAGE given,
the TCXO attempt uses RadioLib's own 1.6 V default rather than being skipped
-- otherwise there is nothing to fall back FROM and the flag would silently
do nothing. This is also the FIXME that sat on the Portduino branch in
LR20x0Interface: an unset voltage now means "no TCXO" explicitly.

Two things are deliberately left alone:

Each family keeps its own probe order. LR11xx tries XTAL first, because a
TCXO-first attempt hangs RadioLib's unbounded calibration wait on a module
with no TCXO fitted, whereas XTAL fails fast and cleanly on a module that
has one; LR20x0 and SX126x try the TCXO first. A carrier therefore behaves
the same way in a Portduino build as in an embedded one, and changing an
order stays a hardware-behaviour decision rather than a tidying-up one.

The SX126x retry is Portduino-only. An embedded TCXO_OPTIONAL board already
gets this from initLoRa(), which constructs a second SX126x interface with
no Vref when the first fails; retrying inside init() as well would leave
that ladder step unreachable and change how every existing t-echo-class
board reports its oscillator. A Portduino build has no ladder to fall
through, because the module is named in YAML rather than probed.

--check learns the key, reports which Vref will actually be tried, and warns
when it is set on a radio with no TCXO reference, where it is read, stored
and inert.

* docs(portduino): condense the comments on this branch

The repo asks for one or two lines and no multi-paragraph blocks, on the
grounds that the diff and the commit message carry the rationale while the
code carries the behaviour. What landed here was well past that: 163 added
comment lines, including a 26-line block above a single macro.

Removes the rhetoric, the issue numbers and the before-and-after asides, and
the notes on where a thing used to live. No added block is longer than three
lines now.

Two facts needed stating and are stated once each rather than repeated at
every use: the slot/DIO divergence between the families, in PortduinoGlue.h,
and the per-family TCXO probe order, in RadioLibInterface.h. The longest
surviving explanation is why an all-LOW switch column still collides with the
interrupt, which sits in the fixtures README because without it that pair of
fixtures reads as contradictory.

Comments only; no functional change.

* address CodeRabbit review on #11382

- SX126xInterface: distinguish an explicit DIO3_TCXO_VOLTAGE from the
  TCXO_OPTIONAL probing default in the debug log instead of always
  claiming the config field was set.
- ConfigCheck: modesFor() now reports an unresolved use_autoconf against
  the union of both radio families' modes, not the LR11xx subset - fixes
  a false "not a mode this part has" warning for valid LR2021-only modes
  (e.g. RFSW_RX_HF) before autodetection resolves the module.
- PortduinoGlue loadConfig: build rfswitch_mode_high[m] as a fresh
  per-row bitmask instead of OR-accumulating onto a stale value, so a
  config re-parse can clear a slot back to LOW.
- PortduinoGlue YAML serialization: gate rfswitch_table emission on
  has_rfswitch_table rather than rfswitch_dio_num[0] >= 0 (missed sparse
  pin lists), and track each emitted pin's original slot so row values
  line up correctly instead of shifting when a low slot is absent.
- config-dist.yaml: document the per-family TCXO/XTAL probe order
  (SX126x/LR20x0 TCXO-first, LR11x0 XTAL-first).
- Trim three overlong comments per the coding-guideline nitpicks.

Left the SX126x XTAL-retry-on-oscillator-failure nitpick alone -
RadioLib's begin() already does its own XOSC_START_ERR recovery
internally, and narrowing our wrapper's retry condition on top of that
needs hardware to verify it doesn't regress a real failure path.

Verified: bin/test-config-check.sh GREEN 69/69 against an isolated
native build; pio test -e native -f test_rtc PASSED.

* fix CI: cppcheck duplicateValueTernary, harden kRfSwitchModes init

LR11x0Interface::init(): work around cppcheck's duplicateValueTernary
on `TCXO_OPTIONAL_ENABLED ? 0 : tcxoVoltage` (both branches fold to 0
on a board with no ARCH_PORTDUINO, no TCXO_OPTIONAL, and no explicit
Vref, since tcxoVoltage already reduces to 0 via the same macro chain)
by splitting it into a plain assignment + if, rather than suppressing
the warning. The other TCXO_OPTIONAL_ENABLED ternaries in this PR
(LR11x0Interface.cpp:75, SX126xInterface.cpp:76, LR20x0Interface.cpp:
86,240) pick between TCXO_OPTIONAL_DEFAULT_VOLTAGE (1.6f) and 0, which
can never coincide, so they're unaffected and left as-is.

ConfigCheck.cpp: kRfSwitchModes was a namespace-scope global with
dynamic initialization (a lambda IIFE) reading kRfSwitchModeNames,
which is defined in a different translation unit (PortduinoGlue.cpp).
Currently safe only because kRfSwitchModeNames's initializer is
constant-expression-only (string literals + enum constants), which
the standard guarantees completes before any TU's dynamic
initializers - but that safety is silent and would break if
PortduinoGlue.cpp's array initializer ever stopped being a constant
expression, with nothing to warn a future editor. Converted to a
function-local static (Meyers' singleton), which is correct by
construction regardless of the other TU's initializer, updating all
4 call sites (definition + 3 uses) from kRfSwitchModes to
kRfSwitchModes().

Verified: pio test -e native -f test_radio PASSED; bin/test-config-
check.sh GREEN 69/69 against an isolated native build.

* refactor: simplify TCXO voltage handling across interfaces and improve comments

* fix rfswitch_table cross-file merge; drop now-stale checker warning

Three CodeRabbit findings on 09e0d390c, addressed together since #2
and #3 are the same root cause:

1. PortduinoGlue.cpp: require an exact "DIO<n>" match when parsing
   rfswitch_table.pins. sscanf's %d stops at the first non-digit, so
   "DIO5invalid" silently parsed as DIO5 at runtime even though
   ConfigCheck.cpp's static validator (exact match against
   kRfSwitchPins) already rejected it - checker and loader disagreed.

2. PortduinoGlue.cpp: reset all 5 pin slots and all 8 mode rows before
   applying a table, rather than only overwriting what the new table
   mentions. A later config.d file that omitted a mode a prior file
   had set (e.g. only redefining MODE_TX) let the earlier file's
   MODE_RX leak through, contradicting "last file wins" - the rule
   every other Lora: key already follows.

3. ConfigCheck.cpp: with #2 fixed, rfswitch_table behaves like any
   other cross-file key, so removed the special-cased ERROR in
   checkCrossFileOverlap ("These do NOT override each other... OR of
   every table") - it described the pre-fix OR-accumulation bug and
   is no longer accurate. Falls through to the generic "last file
   wins" INFO now. Renamed/repurposed the rfswitch-sticky fixture to
   rfswitch-last-wins and updated its assertion (was rc=1 asserting
   the old error text, now rc=0 asserting the generic info) and the
   fixtures README.

Verified: bin/test-config-check.sh GREEN 69/69 against an isolated
native build, including the renamed assertion.

* Assert the effective rfswitch table, not just the overlap diagnostic

The "last one wins" case checked that the cross-file info fires and that the
result is clean. Neither observes the table the loader actually ended up with,
so the merge bug ee9b5b81e fixed - a later table leaving an earlier file's pins
and mode rows as carryover - would still have passed it. Raised by CodeRabbit.

Asserting the value needs two things the existing case cannot supply.

The winner has to be deterministic. Both files in rfswitch-last-wins/ sit in
config.d/, which is walked with a bare directory_iterator and no sort, so which
one lands last is up to the filesystem - the point configd-conflict/ exists to
make, and the reason the checker warns rather than assuming alphabetical order.
rfswitch-replace/ puts the losing table in config.yaml instead, which is always
loaded before config.d/. The loser is the wider of the two, four pins and three
all-HIGH mode rows against the winner's two pins and one all-LOW row, so
carryover shows up as a surviving pin, a surviving mode row, or a HIGH that
should be LOW.

The table has to be observable. The check report says no more than "RF switch
table   : set", and check-yaml cannot help: it is --check --output-yaml, and
--check wins and exits before the dump - which the case just below it asserts.
emit_yaml() does serialise the effective table, so the assert helper grows a
yaml mode that passes --output-yaml alone.

Confirmed non-vacuous: with the reset loop in loadConfig() removed, the new
assertion fails and the old one still passes.

Config-check suite GREEN 70/70. Native suite GREEN 44/44, 968 cases.

* Say nothing about rows the radio will never read

Two of the RF-switch diagnostics judged a table against a family's mode
list without first asking whether the module reads a table at all.

modesFor() treated every module that was not an LR20x0 as LR11xx-like, so
an sx1262 carrying a table was told which of its rows were "not a mode
sx1262 has" and which modes it had omitted - alongside the correct warning
that the whole table is inert. pinsFor() already returned an empty set for
these parts and its caller already guarded on that; the mode path now
matches.

The missing-mode advice is dropped under "auto" as well. The module has
not been probed, so the union of both families is all there is to compare
against, and naming its absent modes would advise adding MODE_TX_HP,
MODE_GNSS and MODE_WIFI rows to what may turn out to be an LR20x0.

Fixtures for both silences, and an assert() needle prefixed with '!' to
hold them: a line that is merely absent today is otherwise nobody's
regression.

Also corrects the kLr11x0SwitchDios/kLr20x0SwitchDios comments. They
describe a slot mapping, but buildRfSwitchTable() searches them by value
to find the parallel pin constant - and with 7 DIOs against 5 YAML pin
slots, the LR20x0 array could not be positional.

* Warn when the two TCXO keys ask for opposite things

DIO3_TCXO_VOLTAGE written out as false or 0 asks for DIO3 to be left
alone, and stores identically to the key being absent - so TCXO_OPTIONAL
then probes DIO3 at the radio default anyway. Both keys behave exactly as
documented; only together are they wrong, which is what makes the outcome
surprising. loadConfig() now keeps the distinction that the store loses,
and --check reports the contradiction and which key to drop.

The flag is diagnostic only and is not serialized: an explicit false and
an absent key both round-trip as absent, as they did before.

Also fixes the Portduino TCXO log lines, which named the variant define
SX126X_DIO3_TCXO_VOLTAGE on a path where the knob is the YAML key, and
adds a TODO over the SX126x XTAL retry. RadioLib has autocorrected that
case itself since 7.5.0 - SX126x::modSetup() retries config() on the XTAL
when begin() fails with SPI_CMD_FAILED and XOSC_START_ERR - so the
ordinary case never reaches our retry and what does is mostly invalid
settings, logged as a TCXO fault.

* feat(portduino): bound Lora.IRQ_DIO_NUM, accept the older spelling, document both

An LR20x0 raises its interrupt on DIO5 through DIO11. Anything else was read
straight out of the YAML and programmed into RadioLib, where it routes the IRQ
nowhere: begin() touches only SPI and BUSY, so the radio reports init success
and then never receives a packet - the same failure the switch-pin collision
check already covers, reached by a typo instead.

Refuse it in loadConfig() and again in the driver before it reaches RadioLib,
warning both times. Because loadConfig() discards the value, the merged config
cannot tell a rejected number from an absent key, so --check judges the range in
its per-file pass where the offending line is still known.

LR2021_IRQ_DIO_NUM, the spelling carried by the two earlier LR2021 branches, is
read when IRQ_DIO_NUM is absent and reported as shadowed when it is not. Both
keys, the DIO range and the collision that makes the setting matter are now
described in config-dist.yaml.

* fix(portduino): a rejected IRQ_DIO_NUM returns to the radio default

loadConfig() runs once per file - the main config, then each file in
config.d/ - and they all write the same portduino_config. An out-of-range
Lora.IRQ_DIO_NUM warned that it was falling back to the radio default but
left any valid value an earlier file had set, so LR20x0Interface went on
programming that stale DIO.

Reset it to -1, the unset sentinel every other reader already tests for.

Pinned by a new fixture: DIO9 in the main config, out of range in
config.d/. The main config is always read first, so the ordering is
deterministic, unlike two files in config.d/ (see rfswitch-last-wins).
The assertion requires the summary to name the radio default and not DIO9;
with the reset removed and rebuilt, it is the only assertion that fails.
2026-09-16 10:44:06 +00:00