Commit Graph
7314 Commits
Author SHA1 Message Date
Georg von StackelbergandClaude Sonnet 5.5 e0c670a65d Meshnology W12: report MESHNOLOGY_W12 hardware model (#12037)
The MESHNOLOGY_W12 enum (145) was added in meshtastic/protobufs#1042, but
the W12 variant predates it and still reports PRIVATE_HW (255). Add the
HW_VENDOR mapping in architecture.h, set custom_meshtastic_hw_model to 145,
and drop the outdated "no hardware model yet" comment.

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-01 22:46:39 +00:00
Jonathan BennettandClaude Opus 5 e691bd3790 Claude/dmshell lock (#12024)
* Lock: give Portduino a real mutex instead of the empty fallback

Lock.cpp has a FreeRTOS implementation and an empty one, and Portduino takes
the empty one: every lock() and unlock() on a Linux build is a no-op, so
concurrency::Lock protects nothing there. TrafficManagementModule's cacheLock
and SPILock are both built on it, and native meshtasticd runs the radio and the
API on separate threads.

Add a pthread implementation under ARCH_PORTDUINO. The timed lock(uint32_t)
blocks rather than returning early, because there is no portable timed
pthread_mutex_lock across Linux and macOS and returning true without acquiring
would leave a caller such as SPILock unlocking a mutex it never took. Targets
that have neither FreeRTOS nor pthreads, such as STM32WL, keep the existing
empty implementation byte for byte.

Co-Authored-By: Jonathan Bennett <jbennett@incomsystems.biz>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01REkPVFh6kvG4AZJ5A8AtM9

* test: create spiLock in the shared setup, which NodeDB needs and no test had

Making Portduino's Lock real turns a latent null dereference into a crash. spiLock
is a bare pointer that initSPI() fills in, and only main.cpp calls that, so in a
test binary it stays null. NodeDB's constructor reaches it through loadFromDisk(),
and while Lock::lock() was an empty function the call never touched `this`, so
23 suites have been calling a method on a null pointer and getting away with it.
With a pthread mutex behind it the same call reads through the null pointer and
takes SIGSEGV at offset 0x10, which is what test_phone_api_config_dump,
test_muted_source, test_nodeinfo_send_window and test_module_config hit.

Create it once in initializeTestEnvironment(), which every affected suite calls as
the first statement of setup(), before any of them constructs a NodeDB. The guard
is the idiom test_xmodem and test_nodedb_identity_hygiene already use; theirs stay
correct and become no-ops. test_safefile called initSPI() bare right after the
harness, which would now trip its assert, so that call goes away.

No firmware behaviour changes: main.cpp still calls initSPI() exactly once, and
nothing outside the test harness is touched.

Co-Authored-By: Jonathan Bennett <jbennett@incomsystems.biz>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01REkPVFh6kvG4AZJ5A8AtM9

* test: create cryptLock where a suite reaches it without building a Router

Second instance of the same latent null dereference the previous commit fixed
for spiLock. cryptLock is a bare pointer that Router's constructor creates
(Router.cpp:246); AdminModule::setPassKey takes a LockGuard on it, and a suite
that exercises an admin path without standing up a Router leaves it null. While
Lock::lock() was empty on Portduino the guard never touched `this`; with a
pthread mutex it reads through null, which is test_tak_config's SIGSEGV in
handleGetModuleConfig.

It cannot go in initializeTestEnvironment() the way spiLock did, because Router
asserts cryptLock is unset before allocating its own, so creating it for every
suite would break the ones that do build a Router. It is a named helper instead,
testEnsureCryptLock(), called by the six suites that reach a cryptLock path with
no Router: test_ack_proof, test_admin_session_repro, test_fuzz_packets,
test_hop_scaling, test_module_config and test_tak_config. The three that define
setup() twice behind a PKI #if get the call only in the branch that compiles the
tests in.

No firmware behaviour changes; nothing outside test/ is touched.

Co-Authored-By: Jonathan Bennett <jbennett@incomsystems.biz>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01REkPVFh6kvG4AZJ5A8AtM9

* test: create cryptLock in the harness instead of chasing suites

Router's constructor asserted cryptLock was unset, then allocated it. That
assert is why ten suites carry a mock-router destructor whose only job is to
delete the global and null it so the next router can be built. While
Lock::lock() was an empty function on Portduino a null cryptLock cost nothing,
so those null windows were invisible; with a real mutex, anything reaching
perhapsDecode() or the ack-proof paths after one of those destructors runs
dereferences null.

The fix is the idiom already on the next line of the same constructor, which
routingAuthCacheLock has used all along: reuse the lock if one exists. Nothing
in src/ ever deleted cryptLock, so a Router that finds one is finding the
process's only one. initializeTestEnvironment() can then create it for every
suite, the way it now does for spiLock, and the ten teardowns and the
per-suite helper from the previous commit all go away.

Replaces the six testEnsureCryptLock() call sites with one creation point, and
removes the null windows in test_admin_radio, test_mesh_beacon,
test_mesh_module, test_mqtt, test_nexthop_routing, test_nodeinfo_send_window,
test_traffic_management, test_event_channel_phone_api,
test_event_channel_router and test_phone_api_config_dump.

Co-Authored-By: Jonathan Bennett <jbennett@incomsystems.biz>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01REkPVFh6kvG4AZJ5A8AtM9

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-01 21:33:14 +00:00
Jonathan BennettandClaude Opus 5 8c0abbd522 feat(portduino): BLE peripheral support via BlueZ for meshtasticd (Raspberry Pi) (#11396)
* feat(portduino): BLE peripheral support via BlueZ for meshtasticd on Linux

Adds the standard Meshtastic BLE service (toRadio/fromRadio/fromNum/logRadio)
to the Linux native target, so a Raspberry Pi running meshtasticd can be
paired and used over BLE like any other Meshtastic device.

Implementation: a new LinuxBluetooth backend registers a GATT application,
LE advertisement and pairing agent with bluetoothd over the org.bluez D-Bus
APIs, using sdbus-c++ (both the 1.x and 2.x major versions, via a small
compat shim - Debian bookworm/Ubuntu 24.04 ship 1.x, trixie/Fedora ship 2.x).
When the sdbus-c++ dev package is absent the whole backend compiles out via
__has_include, the same optional-dependency idiom as the ulfius webserver.

Threading follows the NimbleBluetooth model, simplified: the sdbus event
loop runs its own thread, and all PhoneAPI calls happen on the main thread.
Writes queue to the main loop; reads park the D-Bus reply and are completed
from the main thread after queued writes, so write-then-read clients see
their answer without any busy-waiting.

Enablement is a double opt-in: a new `Bluetooth:` config.yaml section
(Enabled, default false; AdapterId, default hci0) must turn BLE on for the
host, and the regular device config bluetooth.enabled must be on. The
config-check schema and fixtures cover the new section.

Pairing honors config.bluetooth.mode: NO_PIN maps to a NoInputNoOutput
just-works agent; RANDOM_PIN to DisplayOnly with the kernel-generated
passkey shown on screen/log via the existing BluetoothStatus plumbing.
FIXED_PIN falls back to random-passkey semantics with a warning - BlueZ
does not support forcing a passkey. PIN modes enforce
encrypt-authenticated-read/write on all mesh characteristics.

Packaging: install a D-Bus system policy so the meshtasticd user may talk
to org.bluez, add it to the bluetooth group, order the unit after
bluetooth.service, and add libsdbus-c++-dev to debian/rpm/docker/CI deps.

Verified in-container against a mock bluetoothd: registration flow, GATT
tree enumeration, advertisement properties, and a full config download
(ToRadio wantConfig -> 47 FromRadio packets) through the D-Bus bridge.
Real-hardware pairing/notify testing on a Pi still pending.

Known limitations (v1): meshtasticd must be restarted if bluetoothd
restarts; FIXED_PIN degrades to a random passkey; getRssi() returns 0
(same as nRF52).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0184v7MyCLuJHW2ebZ9r8NmQ

* fix(portduino): BLE fixes from first real-hardware pass (Pi CM5 + RAK6421)

Findings from testing PR #11396 on a Raspberry Pi CM5 (Pi OS trixie,
BlueZ/sdbus-c++ 2.1 - the v2 compat path) with a RAK6421 HAT and an
Android phone:

- getMacAddr() leaked its HCI socket on every call and never closed it,
  and on failure returned without touching the caller's buffer - which
  getDeviceName() passed in uninitialized. Close the socket on all paths
  and cache the MAC after the first successful read; it cannot change at
  runtime and this now runs on every bluetoothd property read.
- getDeviceName() zero-initializes its MAC buffer, and LinuxBluetooth
  snapshots the name once at setup() on the main thread: the
  advertisement's LocalName getter runs on the D-Bus event-loop thread
  and getDeviceName()'s static buffer is not thread-safe.
- Restore NimBLE-style config-phase packet prefetch (depth 3). The
  initial port answered every FromRadio read with a D-Bus -> main-loop
  round trip, which made the config download noticeably slow; ReadValue
  now answers straight from the prefetch queue on the event-loop thread,
  with NimBLE's safety rules (never in STATE_SEND_PACKETS, writes always
  observed before reads, queue cleared on disconnect).
- Set advertising MinInterval/MaxInterval to 20-100ms (BlueZ >= 5.71;
  older versions ignore the properties). btmon showed the kernel default
  of 1.28s otherwise, and Android's background-connect scan windows are
  sparse enough that tap-to-connect took 8-14s; 20ms is the same floor
  NimBLE uses on ESP32.

Verified on hardware: scan, passkey pairing, connect, config download,
reconnect after bond wipe. Also diagnosed (no code change): the node
identity MAC comes from the RAK HAT EEPROM by design, so the BLE name
suffix follows the HAT rather than the BT adapter.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0184v7MyCLuJHW2ebZ9r8NmQ

* fix(portduino): address CodeRabbit review on BLE support

- setBluetoothEnable: handle disable before the config gate, so a running
  BLE stack is always stoppable even after the device config turns
  Bluetooth off underneath it
- getMacAddr: read the adapter configured as Bluetooth.AdapterId instead
  of hardcoding hci0, falling back to hci0 for unparseable names
- systemd unit: Wants=bluetooth.service so bluetoothd is pulled up when
  present (After= only orders, it does not start it)
- debian/rpm: Recommends: bluez as the runtime contract for BLE
- dbus policy: document why the org.bluez rule is destination-wide
  rather than a per-interface allowlist

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0184v7MyCLuJHW2ebZ9r8NmQ

* fix(portduino): show the BLE pairing code on BaseUI screens

onDisplayPasskey published the passkey to bluetoothStatus and triggered
PowerFSM, but never called screen->startAlert(), so on BaseUI the code only
ever reached the log. BluetoothStatus has no BaseUI consumer -- only InkHUD's
PairingApplet and StatusLEDModule read it -- so a Pi driving a HUB75/OLED
panel showed nothing while BlueZ sat waiting for the user to type a code they
could not see. NimBLE and nRF52 draw it via startAlert(); this adds the
missing half for Linux.

The agent callbacks run on the sdbus event-loop thread while the screen is
owned by the main thread, so the passkey is handed over as a pending flag and
drawn from runOnce(), matching the existing disconnectCleanupPending pattern
rather than reaching into the screen from the event loop.

Dismissed on all four exits, so a stale code cannot stick on an always-on
panel: Paired -> true (newly watched in PropertiesChanged, which previously
only looked at Connected), agent Cancel, peer disconnect (moved out of the
lastGone branch so a peer leaving mid-pairing clears the code even when
another device is still connected), and doDeinit() -- applied inline there
because runOnce() may never be scheduled again after teardown.

Verified on a Pi 5 + BlueZ 5.66 in RANDOM_PIN mode: the code renders on a
HUB75 panel and clears once the phone completes pairing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* ci: install libsdbus-c++-dev for the native test build

setup-native-test landed on develop while this branch was adding
libsdbus-c++-dev to setup-native, so the new action's "full setup-native
list" of C libraries is missing it. Without the package the test job
builds with HAS_BLUETOOTH 0 and never compiles LinuxBluetooth.cpp.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0184v7MyCLuJHW2ebZ9r8NmQ

* fix(portduino): warn when a factory reset cannot clear BLE bonds

factoryReset(eraseBleBonds) silently did nothing on Linux when the BLE
backend was not running, so the reset reported success while the host's
pairings stayed. Removing them needs a live connection to bluetoothd that
a disabled backend never opened, so say so rather than imply they went.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0184v7MyCLuJHW2ebZ9r8NmQ

* fix(portduino): gate the factory-reset bond clear on an enabled backend

setup() leaves linuxBluetooth allocated with its bus torn down when it
throws, so a pointer check alone let factoryReset log "Clear bluetooth
bonds" for a clear that clearBonds() then declined to perform. isEnabled()
is only true after setup() completes, which routes that case to the
warning that says the bonds were left alone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0184v7MyCLuJHW2ebZ9r8NmQ

* style: reformat under clang-format 20

#11909 moved trunk from clang-format 16 to 20, which spaces C-style casts
differently and reindents the comment above the HAS_WIFI block. Both files
are ones this branch already touches, and trunk's fmt linter grades whole
files, so its check fails until they are reformatted. No behaviour change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0184v7MyCLuJHW2ebZ9r8NmQ

* fix(portduino): three BLE config and lifecycle fixes from review

Bluetooth config keys are now assigned individually rather than per section.
loadConfig() runs once for every file in config.d, so reading an absent key as
its default let a later file that named only one of them silently reset the
other: `AdapterId: hci1` alone turned Bluetooth off, and `Enabled: true` alone
dragged the adapter back to hci0. Only what a file actually states should
override what an earlier one set.

A backend that failed to come up is now retried. setup() can leave
linuxBluetooth non-null but disabled - bluetoothd not ready, adapter missing,
policy refusing - and every later enable then called resumeAdvertising(), which
returns immediately while disabled. A transient failure at boot kept BLE off
until the process restarted. doSetup() already opens with `if (enabled) return`
and tears the bus down on every failure path, so calling it again is safe.

Bluetooth.AdapterId is now checked for the hci<digits> form. LinuxBluetooth uses
the value verbatim as the BlueZ object path while the MAC fallback reads only
the leading hciN, so "hci1junk" looks plausible, yields a MAC, then finds no
adapter and BLE never comes up. Covered by a new fixture and suite case, which
is the kind of silent no-op that directory exists to catalogue.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-01 21:20:09 +00:00
e0c76fd41e Improve performance of encrypted packets with a shared secret cache (#11979)
* Improve performance of encrypted packets with a shared secret cache

Every PKI encrypt, decrypt and ack-proof ran Curve25519::dh2 plus a SHA256
to derive the pairwise key, so a node in a conversation paid a full X25519
per packet, on the main loop, under cryptLock. On a RAK4631 that is ~96 ms
of the ~210 ms it takes to handle a DM.

setCryptoSharedSecret() derives the key only when it is not already held
for that peer, keeping the last MAX_CACHED_SHARED_SECRETS derivations
(8 on nRF52, 2 on STM32WL, 10 elsewhere, under 400 bytes) and evicting the
least recently used. encryptCurve25519, decryptCurve25519 and
ackProofCompute all go through it, so the ack proof gets the same cache
without a second DH path. The cache is emptied whenever our own private
key changes, since every secret in it is then stale.

The lookup key is the first 4 bytes of the peer's public key. A collision
makes us derive against the wrong cached secret, which costs a failed
decrypt for that pair; it cannot disclose either secret.

Co-Authored-By: Jonathan Bennett <jbennett@incomsystems.biz>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JLHxcWJuvSoSSMpz3LWdt3

* Do not let an empty cache slot answer for a zero-prefixed peer key

The cache used lookup_key == 0 to mean "slot unused", so a peer key whose
first 4 bytes are zero matched every unused slot and was handed that slot's
zeroed secret as a hit. The all-zero key is exactly such a key, which is how
test_proof_rejects_weak_peer_key caught it: dh2's weak-point check never ran.

Entries carry an explicit valid flag instead, which the struct's existing
padding absorbs. A weak peer key is also now rejected before the cache is
consulted, since with 4 bytes of lookup key it could otherwise collide with a
cached peer and be served that peer's secret rather than being refused.

Co-Authored-By: Jonathan Bennett <jbennett@incomsystems.biz>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JLHxcWJuvSoSSMpz3LWdt3

* Key the shared secret cache by the whole peer public key

A four-byte lookup key is grindable: anyone can generate a keypair whose
public key shares those bytes with a peer they want to shadow, get their own
entry cached, and then be handed the secret this node uses to talk to that
peer - readable by them, since they hold the matching private key. That is
disclosure of traffic meant for the peer, and a forgeable ack proof in its
name, not the failed exchange a chance collision would cause.

Entries hold the peer key itself and are matched on all 32 bytes, so the
weak-key check dh2 does on a miss can no longer be skipped either and the
explicit isWeakPoint guard goes away with it. The cache costs 66 bytes per
entry: 528 on nRF52, 660 where the default 10 entries apply.

Co-Authored-By: Jonathan Bennett <jbennett@incomsystems.biz>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JLHxcWJuvSoSSMpz3LWdt3

* Allowlist the cache's uptime shift for the millis deadline guard

The guard's regex reads `millis() >> 22` as a comparison against the uptime
clock. It is a right shift, coarsening uptime into the ~1.165 hour units the
cache stamps entries with, and the eviction arithmetic handles that stamp's
8-bit wrap itself.

Co-Authored-By: Jonathan Bennett <jbennett@incomsystems.biz>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JLHxcWJuvSoSSMpz3LWdt3

---------

Co-authored-by: Jason B. Cox <contact@jasonbcox.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-10-01 17:47:25 +00:00
f5940ab5dd Meshtasticd notifications (#10216)
* First attempt at sending message notifications on Linux

* Add files via upload

* Notifications: Use libnotify

Use libnotify for meshtasticd desktop notifications.

Add "HAS_LIBNOTIFY" macro to guard for builds where libnotify is expected.
This is not included in buildroot/openwrt builds. (no libnotify except when extra repos are added).

Install desktop icon in the correct location on Debian and Fedora packages.
Update dependencies in packaging and dockerfiles.

* Add libnotify to setup-native (GitHub Actions)

* Address review feedback on the meshtasticd notification path

- platformio.ini: only define HAS_LIBNOTIFY when pkg-config actually finds
  libnotify. The probe previously ran unguarded and the macro was defined
  unconditionally, so a native build without libnotify-dev both hard-failed at
  config time and claimed the feature was available.
- Fix the sender lookup for develop's flattened NodeInfoLite: has_user/user.*
  are gone, replaced by nodeInfoLiteHasUser() and direct long_name/short_name.
  This is a silent semantic conflict - it merges cleanly but does not compile.
- Move the desktop notification off the packet path. notify_notification_show()
  is a synchronous DBus round trip and meshtasticd's packet handling is
  single-threaded, so a slow or wedged notification daemon could stall the
  radio. handleReceived() now resolves the strings and queues them (bounded at
  16); a dedicated worker owns every libnotify call.
- Stop retrying forever: a failed notify_init(), or three consecutive failed
  shows, latches desktop notifications off instead of re-logging per message.
- NodeDB: the ARCH_PORTDUINO default-enable block was nested inside
  #ifdef HAS_I2S, which Portduino never defines, so it never ran. Hoist it out.

* fix(portduino): close input-broker guard before the libnotify block

The libnotify implementation was inserted ahead of the #endif that closed
#if !MESHTASTIC_EXCLUDE_INPUTBROKER, so the file's trailing #endif closed
#if HAS_LIBNOTIFY instead and the input-broker conditional was never closed.
Every build target failed to preprocess:

  src/modules/ExternalNotificationModule.cpp:630: error: unterminated #if

Close the guard immediately after handleInputEvent(), as develop does, so the
two conditionals stay independent and HAS_LIBNOTIFY still resolves when the
input broker is excluded.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013EZstXBJtzRyBmUD1h2FLs

* fix(portduino): back off desktop notifications instead of latching them off

notifyDisabled was doing two jobs: "libnotify looks unusable" and "the
destructor wants the worker to exit". Because the worker's only exit path was
also its only failure path, three failed shows turned notifications off for the
process lifetime with no way back. The packaged daemon is exactly that case:
bin/meshtasticd.service runs as User=meshtasticd and the rpm spec creates that
account with /sbin/nologin, so there is no session bus and every show() fails.
A stock install logged three warnings and then went silent forever, while
NodeDB now enables the module by default on Portduino.

Split the flag. notifyShutdown is destructor-only and remains the worker's one
exit; a retry window replaces the latch. After maxNotifyFailures the worker
arms a backoff (30s, doubling to a 15min cap), drops the queue rather than
holding stale popups, and keeps looping. The producer refuses to queue while
the window is open, so the retry is driven by the next message after it expires
rather than by a timer - no idle wakeups, no probe notifications. Any success
resets the backoff. notify_init() moved inside the loop so a retry can pick up
a session bus that was absent at startup.

The worker also no longer calls the LOG_ macros. RedirectablePrint formats into
a shared static buffer that nothing guards and every other writer to it is on
the main thread, so the worker now records the state change under the mutex it
already holds and reportNotifyStatus() emits it from portduinoNotify() and
runOnce(). Per-message failure spam becomes one line per transition: the reason
and retry interval when it goes down, and a line when it comes back.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013EZstXBJtzRyBmUD1h2FLs

* fix(portduino): sanitize mesh text before it reaches libnotify

The notification body was the raw decoded payload and the summary was a node
name, both attacker-controlled, and neither was checked before being handed to
libnotify. g_variant_new_string() rejects invalid UTF-8: a GLib CRITICAL and an
"[Invalid UTF-8]" body by default, and a hard abort under G_DEBUG=fatal-criticals,
so an unauthenticated mesh packet could terminate meshtasticd on any install
running with that flag. An embedded NUL separately truncated the body at the
first one, hiding the rest of the message.

Route both strings through sanitizedMeshText(), which replaces embedded NULs and
then applies the existing sanitizeUtf8() helper. TypeConversions already
sanitizes names on the way into NodeDB, but an abort is too sharp an edge to
leave resting on an invariant owned by another file.

Escape the body for Pango markup as well. Servers advertising "body-markup"
parse a markup subset there, so a message can inject formatting to dress itself
up as trusted UI, and where the server also advertises body-images or
body-hyperlinks it can inject tags that make the notification daemon fetch a
remote URL. The escaping is unconditional rather than gated on
notify_get_server_caps(): a caps query that fails, or goes stale across a daemon
restart, fails open, while on a server without body-markup the only cost is
entities rendering literally, and mesh text carries no markup worth preserving.
The summary is not escaped - the spec gives it no markup, so escaping it would
only ever show entities.

Also call notify_uninit() in the destructor, after the join: the worker owns
every libnotify call, so tearing down while it is still running would be a
use-after-uninit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013EZstXBJtzRyBmUD1h2FLs

* fix(portduino): gate the notification default-on behind HAS_LIBNOTIFY

The default-enable block was gated on ARCH_PORTDUINO, so a native build where
pkg-config could not find libnotify still shipped external_notification enabled.
portduinoNotify() is not compiled into that build, and native defines
EXT_NOTIFICATION_MODULE_OUTPUT as 0 with setup() guarding the pin writes behind
output > 0, so nothing was driven - it only exposed a config surface that can
do nothing. Gate it on the same macro that governs the code it exists to feed.

HAS_LIBNOTIFY is a build flag for the whole env and is simply absent when the
probe fails, so this reads as 0 on every other target, matching how
ExternalNotificationModule.cpp already tests it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013EZstXBJtzRyBmUD1h2FLs

* fix(portduino): fail closed when markup escaping returns NULL

escapedNotificationBody() returned its input unescaped if g_markup_escape_text()
gave back NULL, which would hand libnotify the exact attacker-controlled string
the function exists to neutralize. The branch is unreachable for the input we
pass - already sanitized to valid UTF-8, and GLib documents no NULL return for
it - but an error path whose fallback is the unsafe action is the wrong shape
for a helper on this boundary. Drop the body instead.

Also note in the doc comment why running this on the caller's thread does not
weaken the rule that the worker owns every libnotify call: it is a pure GLib
string function that touches no libnotify or DBus state.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013EZstXBJtzRyBmUD1h2FLs

* style: format ExternalNotificationModule.h per the pinned clang-format

Trunk Check Runner failed on c0d5ebff with "1 unformatted file" against this
header. The deviation is a brace-spacing fix that clang-format wants on the
rtttl stub's begin(): the formatter emitted it while reformatting this file for
an earlier commit, and I reverted it then as unrelated churn. It was not -
trunk's fmt check reports a modified file as a whole, so touching this header at
all surfaces that line, and dropping the formatter's output is what turned the
check red.

Verified with the version trunk.yaml pins, clang-format 20.1.0, rather than the
older one available locally: it is the only remaining deviation across the three
files this branch touches.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013EZstXBJtzRyBmUD1h2FLs

* fix(portduino): stop the libnotify probe swallowing native-tft's link flags

The docker arm64 builds failed to link native-tft:

  libmeshtastic-device-ui.a(CURLService.cpp.o): undefined reference to symbol
  'curl_easy_reset@@CURL_GNUTLS_3'
  /lib/aarch64-linux-gnu/libcurl-gnutls.so.4: DSO missing from command line

-lcurl was never on the link line. The libnotify probe was the last line of
[native_base].build_flags, and an env extending it writes

  build_flags = ${native_base.build_flags} -Os -lcurl -lX11 ...

so the interpolation appended those flags to the probe's own line. The result is
a single shell command ending in `... && echo -D HAS_LIBNOTIFY=1 || : -Os -lcurl
-lX11 -linput -lxkbcommon -ffunction-sections -fdata-sections -Wl,--gc-sections`,
where the trailing flags are arguments to echo when the probe succeeds and to `:`
when it fails. They never reach the compiler, and nothing reports it.

Three envs lost flags this way: native-tft and native-tft-debug (-lcurl -lX11
-linput -lxkbcommon), and native-fb (-lcurl, --gc-sections). env:native appends
nothing on that line, which is why the native test suite stayed green and this
stayed hidden until the module compiled and something actually tried to link.

Move the probe above `-I /usr/include` so a plain flag line ends the value, and
record the constraint so the next flag added here does not re-break it.

Verified with `pio project config --json-output`: before, the three envs carried
the flags inside the probe's command string; after, all three carry them as
build flags and no env still swallows any.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013EZstXBJtzRyBmUD1h2FLs

---------

Co-authored-by: Austin Lane <vidplace7@gmail.com>
Co-authored-by: Tom Fifield <tom@tomfifield.net>
Co-authored-by: Claude <noreply@anthropic.com>
2026-10-01 17:38:14 +00:00
Jason P c50b294e6f BaseUI: use function pointers for banner callbacks (#12026) 2026-10-01 15:28:38 +00:00
Jonathan BennettandClaude Opus 5.5 790944a75e tftSetup: don't claim the SPI bus after a timed-out take (#12025)
ReentrantSpiLock::lock(uint32_t) recorded the calling thread as owner
with depth 1 whether or not spiLock->lock(timeout) succeeded. After a
timeout the thread's next lock() then took the reentrant path and skipped
the real acquire, and the matching unlock() released a semaphore the
thread never held, so device-ui and the radio could both reach the bus.

Record ownership only when the take succeeds, and return false without
touching owner or depth when it times out.

Found while reviewing CodeRabbit's note on #12024.



Claude-Session: https://claude.ai/code/session_01REkPVFh6kvG4AZJ5A8AtM9

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 20:53:42 +00:00
Jonathan BennettandClaude Opus 5.5 778184c7a1 Log: report LOG_TRACE lines to the API as TRACE rather than UNSET (#11977)
RedirectablePrint::getLogLevel switches on the first character of the level
string and has cases for D, I, W, E and C, but none for the "TRACE" that
LOG_TRACE passes. Every trace line therefore reached a client as
meshtastic_LogRecord_Level_UNSET, so a capture taken over the API could not be
filtered or sorted by level, and a reader had no way to tell a trace line from
one the firmware never labelled.

meshtastic_LogRecord_Level_TRACE already exists in the protobuf as 5. Add the
missing case, ordered with the others by ascending severity.



Claude-Session: https://claude.ai/code/session_01REkPVFh6kvG4AZJ5A8AtM9

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 20:21:44 +00:00
Manuel 17cc649a64 fix(gps): stop the AIROHA engine before power-down on boards with GPS_SLEEP_INT (ThinkNode M9) (#12004)
* fix AIROHA GPS and M9 sleep/wake

* fix nrf compile error

* make coderabbit happy as well
2026-09-29 16:11:52 +00:00
Thomas Göttgens 642076baed Show Ethernet status on the WiFi screen (#11991)
* Show Ethernet status on the WiFi screen

On W5500 and CH390 boards the WiFi frame shows Ethernet link state, speed,
duplex and IP when Ethernet is enabled and has link, or WiFi is not
configured. If WiFi is also configured, one line shows its IP or
connection state.

* Use the CH390 driver for UDP multicast link checks

On USE_CH390D builds ETH resolved to the unused Arduino ETHClass, so
onSend() always saw Ethernet as disconnected and never sent.

* Require HAS_ETHERNET for USE_WS5500 and USE_CH390D
2026-09-29 11:13:51 +00:00
danandClaude Sonnet 5 d482dc78a1 Fix fixed-position broadcasts going out as lat/lon 0,0 (#11985)
handleReceivedProtobuf() returns early for a fixed-position self-update, to
protect the pinned coordinates from being overwritten by a phone/GPS update -
but that early return also skipped the line that refreshes the `precision`
member for the packet's channel, leaving it at its 0 "safe starting value".

alterReceivedProtobuf() then runs anyway on the same from-us packet (whether
it's our own broadcast looped back locally, or a phone-submitted one) and
truncates it to whatever `precision` currently holds. applyPositionPrecision(_,
0) doesn't clamp - it wipes the whole Position back to defaults, intentionally,
for the channel-privacy case this function was written for. Router::deliverLocal()
runs this same dispatch on the same packet object before it's actually
transmitted, so the wipe landed in the real outgoing radio packet: every
fixed-position broadcast went out as lat=0/lon=0 instead of the fixed
coordinates.

Fix: refresh `precision` on this early-return path too, instead of skipping
it. For our own already-correctly-truncated broadcast this makes
alterReceivedProtobuf()'s re-truncation a no-op (same precision, idempotent).
For a phone-submitted position while fixed_position is on, it's what
correctly applies the channel's privacy truncation - preserving that behavior
was the point of not touching alterReceivedProtobuf() itself. (An earlier
version of this fix skipped alterReceivedProtobuf() entirely whenever
fixed_position was set, which would have let a phone-submitted position
broadcast at full precision regardless of channel setting - caught in review.)

Regression range: v2.7.24 through current develop (introduced in #10383,
which removed the `precision > 0` guard that used to make this a no-op).
v2.7.23 and earlier are unaffected.

Tested on a Portduino/meshtasticd build with position.fixed_position set:
confirmed periodic position broadcasts carried the configured coordinates
instead of 0,0 after this change. Not tested against ESP32/nRF52 hardware -
the affected code path is platform-generic (no ARCH_* gating), so the fix
should apply equally, but I don't have that hardware on hand to verify.

Fixes #11547


Claude-Session: https://claude.ai/code/session_01N7q3Vu58yzcdcP5hCW2au6

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-27 12:58:55 +00:00
GeorgeandThomas Göttgens cb4efd8f41 Classify 0x55 as BQ27220 on the T-Lora Pager (#11961)
* Classify 0x55 as BQ27220 on the T-Lora Pager

The scanner tells the T-Deck keyboard and the BQ27220 fuel gauge apart
at 0x55 by reading register 0x04: nonzero means BQ27220, zero means
TDECKKB. A Pager gauge that needs a configuration reset reads zero, so
it was detected as a T-Deck keyboard. firstKeyboard() prefers TDECKKB
over TCA8418KB, so the cardKB thread polled the gauge every 300 ms and
fed its bytes to the UI as keystrokes, producing phantom input and
unsent-by-user messages.

The Pager always has the BQ27220 at 0x55 and the TCA8418 keyboard at
0x34, so skip the heuristic there and classify 0x55 as BQ27220.

Fixes #11959

* Classify 0x55 as BQ27220 on every HAS_BQ27220 board

---------

Co-authored-by: Thomas Göttgens <tgoettgens@gmail.com>
2026-09-25 20:15:53 +00:00
Thomas Göttgens 09e514602e feat(variants): add L1 Pro 1W battery curve and brightness levels (#11972)
- Add OCV_ARRAY for the L1 Pro 1W battery.
- Let variants override the brightness menu levels via
  SCREEN_BRIGHTNESS_LEVEL_{MEDIUM,HIGH,VERY_HIGH}; defaults unchanged.
- Set L1 Pro 1W brightness levels to 100/160/255.
- Drop USE_SSD1306 from the L1 Pro 1W variant; USE_SH1106 is set in
  configuration.h.
2026-09-25 20:10:14 +02:00
James Rich 20284840a8 Report the ack proof verdict to the client (#11965)
* Report the ack proof verdict to the client

* Test that a proven ack reaches the phone as VALID

* Say why the verdict clear and the pending gate sit where they do
2026-09-25 13:33:58 +00:00
57937c5ccc Fix NTP time detection on OpenWrt Portduino (#11919)
* Fix NTP time detection on OpenWrt Portduino

* chore: trunk fmt

---------

Co-authored-by: stm32repo <orangepimaster@gmail.com>
Co-authored-by: Ben Meadors <benmmeadors@gmail.com>
2026-09-25 12:04:40 +00:00
DoctorRFer c57da9a09d fix(logging): don't route USE_SEGGER LOG_* through SEGGER_RTT_printf (#11970)
SEGGER_RTT_printf() only supports %c %d %u %x %X %s %p. A %f is
skipped without consuming its double from the va_list, so a later %s
reads part of that double as a pointer and HardFaults.

On CanaryOne (USE_SEGGER defined in variant.h) this resets the device
on every packet retransmission, via PacketHistory's
"Reusing slot aged %.3fs TRACE %s" log line. It also means CanaryOne
emits no logs over USB or the API, only over RTT.

RedirectablePrint::write() already mirrors every character to RTT
when USE_SEGGER is set, so drop the SEGGER-specific LOG_* macros and
use the normal logging path: full printf support, the existing log
semaphore, and RTT output preserved.

Tested on CanaryOne (v2.7.26.54e0d8d + this change, USE_SEGGER still
enabled): unacknowledged want_ack sends now retransmit twice and NAK
(err=5) without a reset, and logs reach USB/API again.
2026-09-25 11:09:19 +00:00
Tom f90b48ea6c chore: trunk fmt --all (#11938)
Whitespace and comment-alignment only: the output of `trunk fmt --all` on
develop @ d49cf21c3 with the pinned clang-format@20.1.0 and prettier. 48
files had drifted; trunk-action only checks the files a PR touches, so the
drift never fails CI and instead lands as noise in the next PR to edit any of
them. No code change.
2026-09-24 22:41:37 +00:00
95906609db Sign the whole Data envelope, in one unambiguous layout (#11422)
* Bind request_id and reply_id into the XEdDSA signing buffer

A signed reply can be re-pointed at a different message today. The client sets
reply_id on an outgoing text to make a tapback, firmware signs the resulting
broadcast, but reply_id lives in the Data envelope rather than the payload the
signature covers - and channel crypto is AES-CTR with no MAC, so anyone holding
the PSK can rewrite it in flight and the signature still verifies. request_id
has the same shape and is bound with it.

Packets carrying neither field keep the existing [from|id|portnum|payload]
layout, byte-identical to what v2.8.0 alphas are signing today, so the bulk of
signed traffic - broadcasts - stays verifiable in both directions across the
upgrade. Only packets that actually carry one of the two fields use the extended
layout. Both sides pick the layout from the packet's own decoded fields, so
nothing is transmitted to select it.

Open question for review, deliberately not decided here: a format/version byte
in the buffer would be cleaner than a conditional layout, because the safety
argument for the conditional form has to be re-derived whenever a portnum is
added. It costs nothing on the wire since the buffer is never transmitted, but
it changes every signature and so breaks verification against the alphas that
are already signing. If we want it, better done once and before 2.8.0 leaves
alpha.

Tests: request_id/reply_id flips alongside the existing from/id/portnum negative
cases; a hand-built alpha-format signature that must still verify, and must not
be reinterpretable as the extended layout or vice versa; and receive-path cases
for a retargeted tapback, a retargeted response, and an ordinary signed
broadcast that must be unaffected.

* Sign the whole Data envelope, in one unambiguous layout

Replaces the conditional two-layout signing buffer with a single fixed one:

  version(1) | from(4) | id(4) | to(4) | portnum(4) | request_id(4)
            | reply_id(4) | emoji(4) | bitfield(4) | flags(1) | payload(N)

The conditional scheme was ambiguous. Base was header || arbitrary payload, so
any byte string the extended layout emitted was also a legal base payload: an
attacker could move eight payload bytes into request_id/reply_id and truncate a
signed message while its signature still verified. No marker placed only in the
extended layout fixes that, because base can always reproduce it. A fixed-length
header does - the payload boundary is total - XEDDSA_SIGNED_HEADER_LEN and never
depends on content.

Binding reply_id alone was also not enough, because the fields around it are just
as malleable:

  emoji         a reaction is a text packet with the emoji in the payload,
                reply_id naming the parent, and this flag telling the client to
                render it as a reaction. Flipping it turns a signed reply into a
                signed reaction, so it has to travel with reply_id.
  bitfield      bit 0 is OK_TO_MQTT, the sender's consent to upload to a public
                broker, and the exploitable direction is the one that leaks. The
                whole uint32 is signed so bits 2..31 are covered in advance, and
                presence is signed separately so stripping it is not the same as
                sending it zero.
  want_response bit 1 of bitfield mirrors it and Router merges the two with |=,
                so signing either alone protects neither.
  to            without it a signed broadcast can be re-addressed as a direct
                message and still verify, delivering a public statement as an
                apparent private one. Relays rewrite hop_limit, next_hop and
                relay_node, never `to`.

Left out: dest and source (one write in the tree, no readers), channel (the wire
carries a hash where the decoded packet carries an index), and the hop fields,
which relays rewrite by design. Signing the encoded Data wholesale is not an
option either - a relay that holds the channel key decodes and re-encodes it, so
byte fidelity is lost and unknown fields are stripped. The ack_proof excision
trick does not transfer for the same reason: Routing survives because it rides
inside the opaque payload, which Data itself does not.

sign/verify now take the Data rather than a field list, so adding to the covered
set cannot silently miss a call site. Integers are explicitly little-endian, as
in ackProofCompute. The buffer is sized from the schema's own maximum payload
rather than from what the fits-on-air gate currently admits, because an overflow
makes buildSigningBuffer return 0 and signing fail with no error; MAX_BLOCKSIZE
is left alone so the AES-CTR guard and scratch buffers sharing it are unaffected.

This changes every signature, so it is not compatible with the v2.8.0 alphas. It
is deliberately being done while 2.8.0 is still prerelease: a verify failure
against an authoritative key is an unconditional drop, so the same change after
2.8.0 goes stable would split signed traffic across versions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012hLcVif8GDEmA2k77hmFG8

---------

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Ben Meadors <benmmeadors@gmail.com>
2026-09-24 07:26:37 +00:00
Jonathan BennettandClaude Opus 5 57bdedf324 Bind the ack proof to the node we addressed, not the ack's sender (#11932)
* Bind the ack proof to the node we addressed, not the ack's sender

The MAC proves only that its author holds a pairwise key with us, and every
keyed peer holds one. ackProofVerify looks the key up by getFrom(p) - the ack's
claimed sender - and nothing compared that against the node we actually sent to,
so C, whose authoritative key we hold, could read our packet id out of the
cleartext header and mint a receipt for a packet that went to B. It verified
VALID. "An authenticated delivery receipt from the actual recipient" was not
what the code delivered.

Guard on orig->packet->to before verifying. Since that establishes
getFrom(p) == orig->packet->to, the existing key lookup is then correct and
AckProof.cpp is untouched.

Bailing out rather than verifying against the recipient key and reporting
INVALID is deliberate twice over: it skips the X25519 an attacker would
otherwise choose when we pay, and a third-party ack is "not a receipt" rather
than "a forged receipt" - naks from intermediates (NO_CHANNEL,
PKI_UNKNOWN_PUBKEY, MAX_RETRANSMIT) legitimately come from a node that is not
the destination and must not be logged as proof mismatches. A broadcast original
has no single recipient, so there is nothing to bind to.

Also correct the header doc. It claimed channel (non-PKI) traffic gets nothing
from this, but isProvableAck tests only the ack's shape: a DM that travelled
under channel encryption still gets a proven ack when we hold the peer's key,
because the secret comes from X25519 rather than from the channel. That is more
coverage than the receipt needs and it costs one X25519 per ack generated - both
worth stating rather than implying the opposite.

test_proof_from_third_peer_fails_under_recipient_key pins the property the guard
relies on: Carol's proof for a packet Alice sent to Bob verifies under Carol's
key and fails under Bob's. The guard itself is not directly assertable while
every branch of the verdict switch returns true.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012hLcVif8GDEmA2k77hmFG8

* Make the sender check honor ACK_PROOF_ENFORCE

The mismatch branch returned true unconditionally. Today every branch of this
function returns true, so that reads as equivalent - but if enforcement is ever
switched on it is a hole rather than a no-op. The sender field is not
authenticated, so an attacker would simply address the ack from anyone other
than the node we sent to, take the early return, and skip the proof requirement
entirely.

Hold only success acks to it. A nak legitimately arrives from an intermediate
rather than from the destination - NO_CHANNEL, PKI_UNKNOWN_PUBKEY and
MAX_RETRANSMIT all do - so naks keep today's behavior in either mode.

Broadcast gets its own unconditional return rather than sharing the condition.
orig->packet->to is NODENUM_BROADCAST there and can never equal any sender, so
folding it into the sender check would, under enforcement, stop every reliable
broadcast from ever being acked.

This does not make enforcement sound on its own and the comment says so: ABSENT
still permits, so spoofing the sender and omitting the proof gets through
regardless, and ABSENT cannot be made to block for the reasons recorded at
ACK_PROOF_ENFORCE. The narrower point is that a flag named "enforce" should not
have a branch that silently ignores it.

Raised by CodeRabbit on #11932. Its reading - that this is a live authorization
bypass a third party can use to suppress retransmissions - does not hold: the
sender field is unauthenticated either way, and perhapsGenerateImplicitAckForOwn
Overheard already clears a pending retransmission on a replayed copy of our own
ciphertext, with no ack and no key involved. The structural point stands on its
own merits.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012hLcVif8GDEmA2k77hmFG8

* Avoid cppcheck's duplicateValueTernary in the enforce check

`return isAck ? !ACK_PROOF_ENFORCE : true` has the same value in both arms
while the flag is off, which is the whole point of the line - and is exactly
what cppcheck reports:

  style: Same value in both branches of ternary operator. [duplicateValueTernary]

That failed the `check` matrix on every platform. Write it as an if, which
expresses the same thing and does not trip the rule, and say so in a comment so
it does not get folded back into a ternary.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012hLcVif8GDEmA2k77hmFG8

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-24 07:15:54 +00:00
Ben Meadors 8ef996f5b4 fix(radio): never read a RadioLib error code as a packet's time on air (#11940)
* fix(radio): never read a RadioLib error code as a packet's time on air

getTimeOnAir() reads the packet type back over SPI on every chip and answers RADIOLIB_ERR_WRONG_MODEM when the chip is no longer in the mode we left it in (a reset or brownout leaves it in FSK). RadioLib returns times as an unsigned microsecond count and its negative int16_t error codes through that same value, so computePacketTime() divided -20 by 1000 and handed 4294967ms up as the time one packet spent on the air.

That single sample is enough to take the node off the air until it is rebooted:

  - channelUtilizationPercent() reads 7158% and utilizationTXPercent() 119%,
    matching the 7000%+ ChUtil spike and the 120% AirUtil in #11935, so both TX
    gates stay shut for the hour it takes that bucket to age out;
  - receiveDetected() uses the same figure as its false-header timeout, so
    activeReceiveStart is never cleared and canSendImmediately() sees a busy
    channel for 72 minutes.

Classify the returned value instead of trusting it. A failed getTimeOnAir() now falls back to calculateTimeOnAir() on the modem parameters we configured - arithmetic with no readback that can fail - and the RX path, which already used calculateTimeOnAir(), is guarded the same way.

Fixes #11935

* style: trunk fmt RadioLibInterface.h
2026-09-23 10:43:34 +00:00
b1470cd719 fix(heltec): sleep the T1 and T096 panels on screen-off to stop image retention (#11894)
* fix(heltec-t1): sleep the panel on screen-off to stop image retention

The T1 is excluded from the LovyanGFX sleep()/wakeup() calls because it uses
TFT_eSPI, so DISPLAYOFF only dropped the backlight and the ST7735 kept driving
the last frame unlit for the whole screen-off timeout. That constant static
image is what burns ghost pixels into the panel.

Add opt-in TFT_SLEEP_WHEN_OFF for the TFT_eSPI path: DISPOFF + SLPIN on
screen-off, SLPOUT (120 ms) + DISPON on wake, issued as raw MIPI DCS since
TFT_eSPI exposes no sleep API. Frame memory survives sleep-in, so the previous
frame reappears and the dirty-window diff continues unchanged. The sleep flag
keeps the double displayOn() in Screen::handleSetOn() from paying the delay
twice.

* fix(heltec-t096): sleep the panel on screen-off to stop image retention

The T096 sits behind the same TFT_eSPI exclusion as the T1, so DISPLAYOFF only
dropped its backlight while the ST7735S kept driving the last frame unlit for
the whole screen-off timeout. Same panel, same bus, same burn-in.

Opt in to TFT_SLEEP_WHEN_OFF; the TFTDisplay side of the fix is already generic.

* fix(tft): harden the TFT_SLEEP_WHEN_OFF wake/sleep sequence

Drive VTFT_CTRL LOW before SLPOUT, so the rail is up before the panel is
addressed. Wait out the remainder of the 120 ms the controller needs after
SLPIN before sending SLPOUT, so a wake landing as the screen timeout fires is
not dropped. Guard DISPLAYOFF on panelAsleep to match DISPLAYON.

Correct the T1 VTFT_CTRL comment: LOW enables the rail, not HIGH.

* fix(tft): wait out only what is left of the sleep-in window before SLPOUT

Throttle::isWithinTimespanMs() is true for the whole 120 ms after SLPIN, so the
wake path paid a fresh 120 ms on top of however much had already elapsed. A wake
119 ms after the SLPIN waited ~120 ms rather than ~1 ms, up to 119 ms of
avoidable latency on every quick off/on.

Add Throttle::remainingMs(), which returns what is left of the interval and
saturates at 0 instead of underflowing to a ~49 day wait. It reads the clock
once, so a caller that tests and then waits cannot be preempted between the two
and land on that underflow - which a separate isWithinTimespanMs() plus
subtraction at the call site could.

heltec-mesh-node-t1 and -t096 both build; neither is board_level = pr, so CI
does not compile this path. Docker native suite green, 1515/1515.

* trunk

---------

Co-authored-by: Thomas Göttgens <tgoettgens@gmail.com>
Co-authored-by: Ben Meadors <benmmeadors@gmail.com>
Co-authored-by: Jonathan Bennett <jbennett@incomsystems.biz>
2026-09-23 09:10:38 +00:00
Tom 08cd97ea2d test(harness): make run-tests.sh drivable by a caller that cannot see the terminal (#11862)
* test(harness): make run-tests.sh drivable by a caller that cannot see the terminal

A tool call or a fresh session reads a captured output file, gets interrupted
mid-run, and starts cold. Two things in the wrapper tripped that caller: an
interrupt left pio's build tree running with no record of it (the next
invocation started a second build into the same .pio/build/, or pgrep'd and
matched itself); and pio's own "[PASSED]" / "N succeeded" lines made a
half-finished output file read as green.

run-tests.sh:
- Run record in .pio/runtests/current.tsv for the life of a run. A second
  invocation prints RESULT: BUSY and exits 4 without touching the build
  directory. Valid while the holder pid OR the recorded process group is
  alive, so a SIGKILLed wrapper with live scons children reads ORPHANED
  rather than clear.
- --status / --wait / --abort. --abort kills the whole tree by pgid.
- pio runs under setsid with its pgid recorded; INT/TERM/HUP kill the tree
  and record RESULT: ABORTED (exit 5), log kept.
- One result() for every verdict: prints to the stdout the script started
  with (a signal can arrive inside a redirected pio call) and writes
  .pio/runtests/last-result.tsv with head, args, env, finish time, kept log
  and a tree fingerprint (HEAD + working-tree diff + untracked files).
  --status marks the last verdict STALE when the tree has changed since.
- Banner naming the final RESULT: line as the only verdict.
- Non-Linux host: RESULT: UNSUPPORTED, exit 6, instead of the AMBER code.
- A failed build removes .pio/build/<env>/meshtasticd, which run bare would
  reprint the last good run.
- FILTERED lists the not-run count, not 77 suite names.

bin/run-tests.cmd forwards into WSL with the exit code passed through, so the
same command line works from cmd.exe and PowerShell; no logic is duplicated.

test/README.md, copilot-instructions.md and the mirrors document the new
codes and the rule.

* test(harness): one-suite warm-up, and name the build phase instead of freezing the counter

Measured on a full native run: the warm-up, `pio test --without-testing`
with no filter, builds AND links every suite - 78 links, 2320 s, 29.7 s
each, 39 minutes before the first test ran - and prints a "[PASSED]" line
for each program it merely linked. CI never did this; its warm-up is one
`platformio run`. The shared src objects are the same whichever suite links
them, so the warm-up now links one: the filtered suite when there is one,
else test_utf8. The run itself still builds every suite, as it must.

The heartbeat counted objects newer than its marker, which sits still through
PlatformIO's single-threaded scons dependency scan and through each link -
twelve minutes at "430/754 objs, ETA 17m" on that run, which reads as a hung
build to a caller who cannot run ps. It now names the phase from the
processes in the recorded group: [scons] / [compile] (with the ETA) / [link]
/ [test], and --status prints the same phase word.

* test(harness): define phase_of_run before --status can call it

* test(harness): bind --wait to the run it observed; serialize the run-record publish

Review findings on #11862, all three valid:

- --wait stored a state and then waited for any record to clear, so a run
  that finished and a second that started between polls would be followed to
  the second's verdict, and a run that turned ORPHANED mid-wait could print a
  stale last-result. Each run now has an id (pid-start) in current.tsv and
  last-result.tsv; --wait captures it and reports only a matching verdict,
  else ABORTED-without-verdict.
- run_state() then the current.tsv write was a check-then-act pair: two
  invocations in the same instant could both see IDLE. The pair is now one
  critical section under a short-lived flock, and both records are published
  by rename so no reader can see a partial file. The record stays the
  ownership token; the lock only serializes the handoff (a SIGKILLed holder
  releases flock but not the record, which is why flock alone was rejected).
- The FILTERED and AMBER examples in copilot-instructions.md carried a
  literal suite count, which the same document says never to do.

Verified on a live run: --status RUNNING, a concurrent start refused BUSY, a
--wait started before --abort reported the aborted run's own verdict with the
matching id, exit 5; no build process survived.

* fix(waypoints): do not create an empty store file on clear

clearAllWaypoints() wrote a two-byte empty store unconditionally, so test_waypoint_expiry left
Waypoints_default.wpts behind in every run and the suite has read AMBER (undeclared shared state)
since it landed. An existing file - stale or unreadable included - is still rewritten as empty, so
a reset after a failed load clears flash as before; a file that is not there is left not there.
2026-09-22 11:10:53 +00:00
github-actions[bot]andjp-bennett 9fe0cd8c19 Update protobufs (#11936)
Co-authored-by: jp-bennett <5630967+jp-bennett@users.noreply.github.com>
2026-09-21 06:41:35 -05:00
Thomas Göttgens 46e009d66d fix(portduino): detach the CH341 poll thread when it detaches its own interrupt on Windows (#11882)
* fix(portduino): detach the CH341 poll thread when it detaches its own interrupt on Windows

* fix(portduino): stop a superseded CH341 poll thread on Windows

* fix(portduino): wait for CH341 poll threads before deinit closes the device

* fix(portduino): stop a superseded CH341 poll thread before it rewrites pin state

* fix(portduino): keep a same-thread re-arm's CH341 pin state sentinel intact

* fix(portduino): close the CH341 attach/deinit race and recognize a superseded poll thread

* style(portduino): trim the new CH341 comments to the two-line house limit
2026-09-20 17:20:03 +00:00
Ben MeadorsandWayenWeng d4bb6eea91 fix(gps): wake AG3335 from software RTC sleep (#11889)
* fix(gps): wake AG3335 from software RTC sleep

On the Airoha trackers, GPS_HARDSLEEP simply dropped PIN_GPS_EN. Cutting VCC
with no prior command leaves the receiver in hardware RTC mode, from which
nothing in the firmware ever brings it back: the only Airoha wake sequence
that existed, wakeAirohaForActiveProbe(), is reachable from probe() alone and
never from the normal GPS_HARDSLEEP -> GPS_ACTIVE transition. The tracker
stops producing fixes until it is rebooted.

Park the receiver in software RTC mode with $PAIR650,0 before the power cut,
and pulse GPS_RTC_INT after VCC comes back to bring it out again. The receiver
may already have auto-slept and missed the first command, so it is resent
until it acks, bounded at 400 ms rather than repeated a fixed number of times:
setPowerState() runs on the GPS thread and from the notifyDeepSleep observer,
and every mesh thread shares loopTask on nRF52, so an unconditional wait here
stalls LoRa servicing, the screen and buttons for its full duration on every
sleep cycle, on top of holding the receiver powered that much longer.

The RTC_INT pulse becomes a shared helper so the probe path and the power
state machine no longer carry separate copies, and the wake is guarded on
GPS_RTC_INT as well as GNSS_AIROHA, since the family flag is not a promise
that the board routed that line.

Also drops the raw digitalWrite(PIN_GPS_EN, LOW) calls that followed
writePinEN(false) in GPS_HARDSLEEP and GPS_OFF, and the one in
toggleGpsMode(). writePinEN() already drives the pin through the GpioVirtPin
chain built in createGps(); the raw writes duplicated it while bypassing both
that abstraction and the RAK4631/WISMESH_TAP guard inside writePinEN().

Affects tracker-t1000-e, seeed_mesh_tracker_X1 and wio-t1000-s.

Co-Authored-By: WayenWeng <jinyuan.weng@seeed.cc>

* fix(gps): keep the $PAIR650 retry inside its stated budget

The while-condition was evaluated after getACK, so a final attempt starting
just under the budget could add another ack window on top of it. Stop starting
attempts once a whole window no longer fits, which makes 400 ms a real ceiling
rather than a soft one.

* fix(gps): gate the AG3335 park on a wake path, the probed model, and a miss count

Review found three holes in the soft-RTC park, all from review by Thomas
Goettgens.

The sleep was guarded on GNSS_AIROHA while the wake was guarded on
GNSS_AIROHA && GPS_RTC_INT, so a board that did not route RTC_INT would have
been parked with no way back out - worse than the bare power cut this is
meant to fix. Both now derive from a single HAS_AIROHA_SOFT_RTC, so they
cannot be guarded separately again.

The retry stamped 'start' from raw millis() but tested it with
Throttle::isWithinTimespanMs, which reads Time::getMillis(). Under
Time::setTestMillis() the two diverge and the loop either falls through or
never exits; getACK just below already uses the injectable clock.

The park was gated only at compile time, so a probe that fell back to
GENERIC_NMEA would still be sent PAIR650 and block for the full budget every
cycle. It now checks the probed model the way the constellation setup at
L975 does, and gives up after three consecutive unacked sleeps so a receiver
that is present but wedged cannot stall loopTask indefinitely.

* fix(gps): drive the Airoha probe wake from the capability, not the board

Requested by Manuel Verch. wakeAirohaForActiveProbe() asserted EN and pulsed
RTC_INT only under TRACKER_T1000_E, so seeed_mesh_tracker_X1 and wio-t1000-s
got a bare $PAIR382 during probe even though both route the same pins. Keying
it on HAS_AIROHA_SOFT_RTC gives every board that routed RTC_INT the physical
wake, which is what the probe's own hardware reset needs undone, and lets a
future variant opt in by declaring the pins rather than by name.

The 1000 ms $PAIR382 repeat loop goes with it: it was compensating for the
pulse being absent at this point, so a single command after the pulse is
enough, and T1000-E boots a second sooner.

A board that declares GNSS_AIROHA without routing RTC_INT keeps the bare
command. Waking and parking must match, per the previous commit, but probing
only reads, so it cannot strand the receiver the way the park can. The EN
write is guarded on PIN_GPS_EN the way createGps() already guards its own.

---------

Co-authored-by: WayenWeng <jinyuan.weng@seeed.cc>
2026-09-20 17:18:01 +00:00
Manuel d971d507c4 Thinknode M9 V2: keyboard and GPS (#11905)
* thinknode v2 keyboard and GPS

* update device-ui commit reference
2026-09-20 15:29:51 +00:00
Tom 332c4d7c6f Narrow the ad-hoc NodeInfo greeting (#11897)
* feat(nodedb): greet only while the node store is under half full

The ad-hoc greeting in MeshService::handleFromRadio() was gated on
!isFull(), so a node kept sending unsolicited NodeInfo right up to the
last free slot - on a dense mesh that is the regime where the store is
already churning and the greeting is least likely to buy a lasting
entry.

Add NodeDB::isHalfEmpty(), true only when strictly more than half the
slots are free, and gate the greeting on it instead. The comparison is
written as 2 * numMeshNodes < cap so a half-full store reads false with
no integer rounding, and MAX_NUM_NODES is read into a local because
portduino resolves it through a runtime call.

The helper keeps the MINIMUM_SAFE_FREE_HEAP term that !isFull() used to
contribute: low heap disqualifies the store regardless of occupancy, so
a sparse database on a memory-starved device still does not transmit.

Admission is untouched - updateFrom() and getOrCreateMeshNode() still
fill to capacity. Only greeting stops early.

* fix(nodeinfo): raise the minimum greeting window to 30 minutes

The !shorterTimeout branch of NodeInfoModule::allocReply() used a
10-minute base, so a node that had just greeted one neighbour could
greet the next ten minutes later. Raise the base to 30 minutes.

This is the floor, not the window: getConfiguredOrDefaultMsScaled()
still multiplies by the congestion coefficient for the roles that scale,
so a busy mesh stretches it further. ROUTER/ROUTER_LATE and the
tracker/sensor roles bypass the scaling and get a flat 30 minutes.

The interactive paths are unaffected - they pass shorterTimeout and keep
their own 60-second gate. The periodic broadcast is unaffected too:
default_node_info_broadcast_secs is 3 hours with a 1-hour minimum, both
clear of the new floor, so the timer is not swallowed by the throttle.

* fix(nodeinfo): a send restarts the routine broadcast countdown

sendOurNodeInfo() left the OSThread schedule alone, so an ad-hoc send
had no effect on the periodic broadcast: run() anchors the next run at
runned() + interval, and nothing re-anchored it when the send came from
a greeting, a PKI decrypt failure or a completed key verification. The
routine copy could follow minutes behind an ad-hoc one, putting two
NodeInfos on the air for no gain.

Call setIntervalFromNow() with the configured broadcast interval once
the packet is queued, so the next periodic copy is a full interval from
the send rather than from the last tick.

It sits on the return-true path only: a send vetoed by allocReply() -
throttle, airtime ceiling, reply suppression - must not be able to
silence the routine broadcast. Calling it from inside runOnce() is
harmless, since run() then applies the same interval from a last_run of
effectively now.

* test(nodeinfo): cover the send window, the countdown reset and the greeting gate

Three behaviours from this branch had no coverage: isHalfEmpty()'s exclusive
boundary, the 30-minute send floor, and the countdown reset on a send.

isHalfEmpty() goes to test_nodedb_blocked, which already owns the full-store
cases and clears the hot store per test. Three tests sweep the cap over the
sizes real deployments have - portduino resolves MAX_NUM_NODES from
General.MaxNodes on every read, so a predicate that cached it would greet at the
wrong occupancy - and pin the band where admission outlives greeting. That suite
had no tearDown; it has one now, restoring the cap so an assertion firing
mid-sweep cannot leak a 2-node cap into the tests after it.

test_nodeinfo_send_window is new because nothing in the tree stands up
NodeInfoModule's send path. Six tests: the floor at 30 minutes with 10 refused,
the interactive 60-second gate staying separate, the countdown re-armed by a
broadcast and by an ad-hoc unicast, left alone by a refused send, and a preset
change consumed only by a send that goes out.

The scaling above 40 online nodes is deliberately not retested here -
getConfiguredOrDefaultMsScaled() is test_default's contract, per preset and per
role. These tests pin the base and leave the multiplier alone.

NodeInfoModule gains two PIO_UNIT_TESTING accessors for the countdown:
concurrency::OSThread is a private base, so a test shim cannot reach it and only
the class itself can. They compile out of a shipping build.

The heap term in isHalfEmpty()/isFull() stays uncovered: memGet.getFreeHeap()
returns UINT32_MAX on portduino, so a native test could only pin a stub.

* chore(trunk): exempt test_nodedb_blocked from the trufflehog Lob detector

test_removeNodeByNum_presentNodeOnFullDb is exactly 35 characters after the
test_ prefix, which is the length of a Lob API key, and trufflehog's detector
matches the bare identifier. The name is years old; it surfaces now only because
this branch touches the file, and the pre-push gate reports a finding in a
changed file as new.

Added to the ignore block that already carries the same detector's hex-literal
false positives, with the reason stated alongside them. Nothing in that file is
a credential.

* fix(nodeinfo): exempt a licensed station from the floor, delay only on a real send

Two review findings on the 30-minute window.

Ham mode sets node_info_broadcast_secs to 600 s for the FCC minimum call-sign
announcement (AdminModule.cpp). The new floor refused every one of those sends
until 30 minutes had passed, so a licensed station's call sign went out three
times less often than the regulation asks - a regression the old 10-minute base
did not have. A licensed station now keeps its own interval whenever that is
shorter than the floor. The exemption is exactly the licensed case because
nothing else can get under the floor: a set-config clamps the field to an hour,
and the userprefs path clamps identically.

sendOurNodeInfo() ignored what sendToMesh() returned, so a packet the router
declined - no interface, queue full - still re-armed the routine broadcast and
still reported success, which let runOnce() consume a pending channel change
for a send that never reached the air. Only ERRNO_OK and ERRNO_SHOULD_RELEASE
now count; sendToMesh() has already released the packet in both cases.

Both are pinned by tests that fail without them, measured: the licensed case
fails at "11 min is past it, and the floor must not override it", the declined
send at "a declined send is not a send". The licensed test carries an unlicensed
control on the same configuration, so deleting the floor outright would not
satisfy it.

* test(nodeinfo): assert the deadline the scheduler reads, from an aged last_run

The countdown cases asserted Thread::interval, which is not what schedules the
next run: shouldRun() keys off _cached_next_run, and the two ways of writing it
differ. setIntervalFromNow() recomputes it from now; Thread::setInterval()
recomputes it from last_run. Swap the call in sendOurNodeInfo() for the latter
and the period still reads three hours while the deadline lands wherever the
last tick was - firing the routine copy right behind an ad-hoc send, the exact
thing the reset exists to prevent. Every test passed.

Assert the deadline instead, from a fixture where the two answers are
distinguishable: ageLastRunForTests() calls Thread::runned() with an hour-old
timestamp, the state a periodic thread is genuinely in between runs, so a
deadline off last_run lands an hour early against a five second tolerance.

Measured: with setInterval() in place of setIntervalFromNow(), the new case
fails by 3600004 ms and the eight others pass, including the one asserting the
period - which is what says the old assertion could not see this.

runned() and _cached_next_run are protected in Thread and OSThread is a private
base, so the hooks live on NodeInfoModule, with the two already there.

Raised by Copilot on #11897.

* fix(nodeinfo): a declined send must not start the throttle window either

allocReply() stamped TransmitHistory when it built the packet, before anything
had been sent. The previous commit made sendOurNodeInfo() report a router
rejection instead of swallowing it, but the stamp was already written by then,
so a packet that never reached the air still started the window - and with the
floor now at 30 minutes, that silences the node for half an hour over a send
that failed.

allocReply() has two callers and only one of them can see the outcome: the
module framework sends its own reply through currentReply, with no post-send
hook a module can reach (MeshModule::sendResponse is not virtual). So the stamp
stays there for that path, and sendOurNodeInfo() defers it across its own
allocReply() call and stamps once the router has accepted the packet.
deferHistoryStamp mirrors the shorterTimeout member alongside it - same
call-scoped signal, same lifetime.

test_sendWindow_aRejectedSendDoesNotStartTheWindow asserts both halves: no stamp
after the rejection, and the retry immediately after goes out. The existing
rejected-send case checked the first failure and the countdown only, which is
how this survived it.

248/248 across every suite that touches NodeInfoModule (admin_session_repro,
admin_radio, nodeinfo_send_window, traffic_management, fuzz_packets) plus
transmit_history, whose subject this is.

Raised by CodeRabbit on #11897.
2026-09-20 10:32:41 +00:00
TomandBen Meadors 20f7ab1be9 Relay a PKI unicast with a known party in LOCAL_ONLY and KNOWN_ONLY (#11898)
* fix(router): relay a PKI unicast with a known party in LOCAL_ONLY and KNOWN_ONLY

A frame the relay cannot decrypt takes the OPAQUE_RELAY_ONLY path in
Router::perhapsHandleReceived() and returns before handleReceived(), so it never
reaches a module. The rebroadcast_mode rule for opaque traffic is therefore the
IS_ONE_OF list in relayOpaquePacket(), and LOCAL_ONLY and KNOWN_ONLY were not on
it: a node in either mode dropped every opaque frame, including a PKI unicast
with a party it knows. The gate in RoutingModule::handleReceivedProtobuf() that
used to allow exactly that is unreachable for these packets and no longer
decides anything.

Add both modes to the list, with the identity rule the unreachable gate carried:
a PKI-shaped unicast (channel 0, not broadcast) with `from` or `to` known to us.
An unreadable broadcast and a unicast between two strangers stay dropped in
these modes, which is what the proto documents - they ignore foreign meshes.

This is not only direct messages. Remote administration and key verification are
PKI unicasts too, and a KNOWN_ONLY relay was black-holing those between two
other nodes just the same.

CORE_PORTNUMS_ONLY reached this list the same way in #11843; this is the
remaining pair.

* test(rebroadcast_mode): pin the relay decision per mode, in both directions

What this node carries for other nodes, per DeviceConfig.rebroadcast_mode, for
packets it can read and packets it cannot. The harness pushes a real PKI frame
through Router::perhapsHandleReceived() and counts what reaches the radio.

Both directions are load-bearing, and each is guarded by cases the other leaves
green. Measured by rebuilding the firmware three ways:

  - as shipped: 9/9 pass.
  - with the two modes taken back out of relayOpaquePacket(): the known-party
    and remote-admin cases fail, the stranger and foreign-mesh cases still pass.
  - with the modes listed but the identity qualifier deleted: the stranger and
    foreign-mesh cases fail, the known-party cases still pass.

So a revert and an over-broadening each fail their own tests, and neither can be
satisfied by breaking the other.

The suite and its harness come from the opaque-packet-handling branch and
compile against develop unmodified. Two assertions were dropped because they
pin behaviour this branch does not add: delivery of unreadable frames to the
phone, and a queued copy of our own suppressing the originator's repeat, which
needs relayOpaquePacket() to consult the TX queue.

* fix(test): guard the harness include, drop a comment its test outlived

Two review findings on the imported suite.

The harness builds real PKI frames through CryptoEngine entry points a
MESHTASTIC_EXCLUDE_PKI build does not declare, but it was included above the
guard, so the empty-suite branch that exists for those builds could not compile.
Move the include inside the guard and pull the base includes the stub branch
needs above it.

The phone-delivery case was dropped when the suite came across - this branch
does not deliver unreadable frames to the phone - but its comment stayed behind
and now described the test below it, which is about the signature policy.

* test(rebroadcast_mode): make the from-known and channel-0 operands load-bearing

The qualifier has three operands, and the suite only exercised one of them.

Every existing case that a known party carries has a known DESTINATION: the
remote-admin case marks both parties, the known-destination case marks the
target. Rewriting the identity test to consult p->to alone passed all nine. And
the only non-PKI-shaped frame in the suite was a broadcast, which
!isBroadcast(p->to) rejects before p->channel is ever read, so deleting the
channel gate passed all nine too.

Add the two cases that close it: a known SOURCE with a destination we have never
heard of, which must relay in both modes, and a known party on a channel hash we
do not hold, which must not - that is someone else's channel traffic, addressed,
not PKI.

Measured both ways. With the from operand and the channel gate removed from
relayOpaquePacket(), the two new cases fail and the other nine pass; with the
shipping code, 11/11.

Raised by CodeRabbit on #11898.

---------

Co-authored-by: Ben Meadors <benmmeadors@gmail.com>
2026-09-20 10:29:48 +00:00
Jonathan BennettandClaude Opus 5 d3b4b343e7 Send an ack over PKC when no channel can carry it (#11891)
* Send an ack over PKC when no channel can carry it

PKI needs only the two keys, so a DM can reach us over a channel we do not
carry. Its ack is a ROUTING packet, which wouldEncryptWithPKC() excludes, so
today it is channel-encoded, fails at setActiveByIndex() with NO_CHANNEL, and is
never sent. The sender sees nothing and retransmits to exhaustion for a message
that was in fact delivered.

Fall back to PKC for exactly that case. This is the one place an ack is
deliberately made opaque to relays; normally that costs next-hop learning and
intermediate retransmission cancel, which is why ROUTING is PKC-excluded in
general, but here there is no readable alternative to lose, because without this
the ack does not exist.

The predicate is scoped as tightly as that argument reaches: a unicast ROUTING
packet we originate, carrying a request_id, to a destination whose key we hold,
under the same ham/sim/private-key preconditions PKC always has, and only when
the channel index does not resolve. It tests channels.getHash() rather than
setActiveByIndex() so it has no side effect; generateHash already returns -1 for
an invalid key, so the two agree on which indexes are unusable.

Four cases in test_packet_signing pin the corners: the fallback fires, it does
not paper over an ack with no destination key, it does not catch a non-ack on
the same unusable channel, and an ack on a channel that does resolve still goes
out readable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012hLcVif8GDEmA2k77hmFG8

* Short-circuit the fallback so an out-of-range channel logs no error

wouldEncryptWithPKC() reaches channels.getName(chIndex) before its portnum
exclusion, and getByIndex() logs "Invalid channel index" on the way past. With
the general predicate tested first, an ack on an out-of-range index printed that
error and then went on to encode successfully. Test ackFallback first so the
case that is about to succeed never asks.

Also record why the range check leads inside the predicate: getHash() is a bare
hashes[i] with no bounds test of its own.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012hLcVif8GDEmA2k77hmFG8

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-20 05:21:06 +00:00
Jonathan BennettandClaude Opus 5 a0c230e091 fix(baseui): stop the connection footer erasing the last body row (#11918)
drawCommonFooter() black-fills the bottom (connection_icon_height + 2) rows
before drawing the API-connected link icon. On colour builds that fill spans the
whole screen width; the monochrome path fills only the icon's own width.

The body grid does not shrink with the panel. textSixthLine is 58 rows down
whatever the display height, while the footer band starts at
SCREEN_HEIGHT - 1 - connection_icon_height. On a 64-row panel that is row 58 --
exactly the sixth body line -- so the bar erases the last thing the frame drew.
On BaseUI that is the sixth body row, the LoRa frame's ChUtil bar (rows 52-59)
and the bottom of the clock. Taller panels put the band clear of the grid and
are unaffected: at 80 rows it starts at 74, and on high-res it is far below.

It only shows once a client is connected, since the function early-returns on
!isAPIConnected(), which is why it reads as the blue link icon eating the
screen rather than as a layout bug.

Only the icon's own rect is registered for colour tinting, so the wide fill buys
the tint nothing. Keep it where it clears the body, and fall back to the
icon-width fill -- what the monochrome path already does -- where it would not.
Panels with room below the body are byte-identical to before.

Seen on a 128x64 HUB75 running meshtasticd; esp32s3/visualizer-hub75 is the
same geometry.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 04:12:35 +00:00
Jonathan BennettandClaude Opus 5 ee02cc3426 Games joystick input (#11917)
* fix(games): correct the high-score announcement argument order

GAMES_HIGH_SCORE_STRING is "New %s high score %lu by %s!" but the arguments
were passed as (name, initials, score): the initials string was formatted
through %lu and the score integer through %s. That is a format/argument
mismatch, so the announcement printed garbage at best and dereferenced the
score as a pointer at worst.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(input): report which physical gamepad button produced an event

A joystick event only carried the action it was mapped to, so a consumer could
not tell two buttons apart once they shared one action, and games were limited
to the handful of actions the broker defines.

Carry the originating evdev button code in InputEvent::kbchar, encoded into a
reserved 0xC0..0xDF range that misses printable ASCII and every
INPUT_BROKER_MSG_ value (SystemCommands switches on kbchar without looking at
inputEvent, so a collision there would reboot the node rather than move a
paddle). D-pad events are axes, not buttons, and keep leaving kbchar at 0 --
which is exactly what lets a consumer tell stick from button.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(portduino): let one joystick action bind several buttons

Input.JoystickButtons took a single evdev code per action, so a pad's A and Y
could not both select, and the shoulder buttons could not sit alongside the
D-pad. Accept a list of codes as well as a bare scalar; the config writer
inverts its code->action map back out, emitting a list only where an action
has more than one button.

ConfigCheck gains a real checker for the section (it was previously waved
through as free-form) covering the three ways a mapping silently does nothing:
an action name the driver does not know, an evdev name where the numeric code
belongs, and one code claimed by two actions. Two fixtures and shell-test
cases cover the clean list form and those three faults.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(games): use the gamepad's extra buttons, and return home when idle

Games now receive the physical button alongside the action, so a pad with more
than two usable buttons controls more than two things:

- Snake: a shoulder button mapped to left/right turns relative to the snake's
  heading (L counter-clockwise, R clockwise) while the D-pad keeps steering
  absolutely. The two are told apart by kbchar, not by hardcoding one pad's
  codes.
- Breakout: the ball now rides the paddle after each serve until the player
  fires it with B or A, so a life is not lost to a ball already in flight when
  the player looks up. The paddle also keeps its position between lives. A game
  can claim BACK for the duration (Game::wantsBackButton) so B serves instead
  of pausing, and releases it once the ball is live.
- Start (BTN_BASE4 / BTN_START) is mapped to select like any other button, so
  it launches games and drives the menus; inside a running game GamesModule
  picks it out of kbchar and pauses instead.

Separately, the games frame no longer holds a walked-away device hostage: after
15 s with no input it returns to the home frame, so the device still reads as a
Meshtastic node. The timer is suspended while a picker or banner is up (e.g.
high-score initials entry, which the input handler never sees) so it cannot
yank the user out mid-entry.

Screen::isInteractionBusy() generalises the old module-intercept check --
modal module, intercepting module, game, or an open interactive overlay --
and MessageRenderer uses it before popping an incoming-message banner. A
transient banner REPLACES an active overlay, so an arriving message could
otherwise discard a half-entered high score. The message is still stored, its
thread still selected, and the unread indicator still set.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ui): compose freetext on the on-screen keyboard from a gamepad

A gamepad can drive the on-screen keyboard but cannot type, so on a host with
a joystick and no configured keyboard device the OSK is the only way to compose
freetext. Set osk_found there, and gate the "Freetext" menu entries on whether
the device can enter text at all (physical keyboard, OSK, or touchscreen
virtual keyboard) rather than on kb_found alone -- those entries were hidden on
exactly the devices that needed them.

The OSK prompt that CannedMessageModule already had inline in the message
selector becomes showOnScreenKeyboard(), so the menu path can reach it too.
Menus call in from a banner callback and the banner is torn down as soon as
that callback returns, which would take the keyboard down with it, so the menu
path defers the launch to runOnce().

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(games): address review on frame fallback and joystick input gating

Breakout: the paddle suppression was far too broad. aLinuxJoystick is
constructed on every Linux host whether or not a gamepad is configured
(InputBroker.cpp), so `aLinuxJoystick && kbchar == 0` was true everywhere and
swallowed LEFT/RIGHT from the keyboard, trackball and ExpressLRS -- on a host
with no joystick attached at all. Gate on the stick actually driving the paddle
instead: LinuxJoystick assigns heldX before it emits and only auto-repeats while
heldX is set, so every axis LEFT/RIGHT arrives with a zone held and nothing else
does. kbchar == 0 still distinguishes an axis from a shoulder button mapped to
left/right, which must keep nudging the paddle.

Screen: showHomeFrame() did nothing when the home frame was hidden, since
setFrames() only assigns positions.home for !hiddenFrames.home. That stranded
the games inactivity bounce on the frame it was trying to leave. Fall back to
the messages frame, which setFrames() always adds.

Test: rename test_ballWaitsOnPaddleUntilLaunched to
test_ball_waitsOnPaddleUntilLaunched, matching the repo convention and its
neighbours in the file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(games): give games the whole InputEvent so Breakout can identify the source

Follow-up to review on #11917. The previous narrowing still could not tell
sources apart: kbchar == 0 is shared by the joystick's D-pad axis and by every
other driver that sends a bare LEFT/RIGHT, so while the D-pad was held a
keyboard or touchscreen press was still discarded. heldXZone() proves the axis
is driving, not that this particular event came from it.

Pass the event itself to Game::handleInput() rather than (ev, kbchar). Games
that only care about the action read event->inputEvent; Snake keeps using
kbchar for shoulder steering; Breakout now also checks event->source against
LinuxJoystick's origin name, so only that driver's own axis repeats are
suppressed.

Chose the event over a third positional parameter so the signature does not
have to grow again the next time a game needs something the event already
carries.

All three conditions in Breakout are load-bearing: source says it came from
this gamepad, kbchar == 0 says it is the axis rather than a shoulder button
mapped to left/right, and heldXZone() != 0 says the axis is what is driving
right now so tick() already has it covered.

LinuxJoystick::originName() exposes the name the driver stamps into
InputEvent::source, alongside the existing heldXZone()/heldYZone() accessors.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 04:11:18 +00:00
Thomas Göttgens f13faa6aba fix(esp32s3): bound the SerialConsole idle sleep on hardware USB CDC (#11901)
HWCDC::isPlugged() is a SOF watchdog that reads false transiently while USB
is connected and working. runOnce() answered that with a 20 s sleep, and
nothing wakes the thread on RX, so host traffic sat in the CDC RX ring and
reached the API as a burst. Cap the sleep at 250 ms, the rate readStream()
already idles at.

IS_USB_SERIAL only tested ARDUINO_USB_CDC_ON_BOOT, so ARDUINO_USB_MODE=0
boards ran the same check against a USB-Serial/JTAG peripheral that is not
attached to the PHY and never sees a SOF. Gate the check on IS_USB_HWCDC.

Measured on tlora-t3s3-v1, 900 s of 1 Hz ToRadio/FromRadio round trips:
before 12 stalls, rtt_max 19.96 s, console asleep 26.4% of wall time.
After 0 stalls, rtt_max 0.147 s, p50 unchanged at 0.028 s.

Fixes #11864
2026-09-18 14:46:43 +00:00
Thomas Göttgens d96c690a91 fix(detect): check LPS22HB before SFA30 at 0x5D and CRC-validate SFA30 probe (#11881)
The SFA30 probe only compared the requestFrom() length, which equals the
requested length for any device that ACKs, so an LPS33HW/LPS35HW at 0x5D
was reported as SFA30. Probe WHO_AM_I first so ST sensors never receive
the SFA30 command, and require valid Sensirion CRC-8 on every word of the
device marking response.

Fixes #11880
2026-09-18 12:19:37 +00:00
Thomas Göttgens c29bd00971 Rewrite the MQTT region root topic only on the default broker (#11899)
* fix(mqtt): rewrite the region root topic only on the default broker

A region change rewrote any root starting with "msh", on any broker. Custom roots such as msh/home were clobbered, and private brokers had their topics moved even though a regional broker is regional already. An empty root, which MQTT treats as the default, was never updated.

Region changes now go through MQTT::applyRegionRootTopic(), which rewrites the root only on the default broker and only when the root is empty, the default, or a msh/<region> the firmware wrote itself.

* fix(mqtt): parse the broker address before the default-server check

Persist module config only when the root actually changed. Replace a stray NUL byte in the test with the \0 escape.

* fix(mqtt): count the regional roots as the default root topic
2026-09-18 10:34:53 +00:00
Thomas GöttgensandClaude a5dce941dd fix(telemetry): follow AS3935Config rename to AS3935State (#11903)
meshtastic/protobufs#1045: admin.proto's AS3935_config and
telemetry.proto's AS3935Config collide after name mangling. The
flash-persisted message is renamed to AS3935State upstream.

Field numbers are unchanged, so existing /prefs/as3935.dat files still
decode.

Depends on the protobufs rename and the regenerated sources landing
first.

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-18 17:32:26 +02:00
github-actions[bot]andcaveman99 11550fa3bd Update protobufs (#11904)
Co-authored-by: caveman99 <25002+caveman99@users.noreply.github.com>
2026-09-18 16:41:30 +02:00
Benjamin Faershtein 3aac179397 fix(phoneapi): wake clients after config sync (#11818) 2026-09-18 08:50:20 +00:00
Thomas GöttgensandBen Meadors 0dafcc90fe fix(api): retain the unwritten tail on a short TCP API write (#11890)
* fix(api): retain the unwritten tail on a short TCP API write

ServerAPI closed the session whenever stream->write() returned fewer bytes than
requested. A short write is transmit-buffer backpressure, not a dead socket, and
it is most likely during the back-to-back frames of the initial NodeDB dump, so
a node at its node cap dropped clients on effectively every connect.

Route TCP frames through StreamFrameWriter, the retained-tail path the USB CDC
console already uses: the remainder is re-offered on the next pass and the
session is closed only when the link itself is gone. Poll at 25ms while output
is still undelivered, since nothing wakes the thread when the socket frees
transmit space.

Fixes #11822

* fix(api): block log re-encoding while a TCP frame is retained

emitLogRecord() writes into txBufLog and StreamFrameWriter can now hold that
buffer as a retained tail, so a second log record would overwrite bytes the
transport has not sent yet. Gate encoding on the retained-frame state, matching
SerialConsole.

No caller reaches this today (emitLogRecord() is only used by SerialConsole),
but retaining the buffer at all is new here.

---------

Co-authored-by: Ben Meadors <benmmeadors@gmail.com>
2026-09-18 06:54:50 +00:00
Ben Meadors 585ce17f59 fix(nrf52): stop concurrent flash writers corrupting LittleFS, and stop a failed save formatting it (#11872)
* fix(nrf52): serialise the warm-node ring against LittleFS on the shared flash cache

On nRF52840 the warm-node store writes its 3-page record ring straight through
flash_nrf5x_write/erase/flush, holding only spiLock. Every LittleFS writer instead
holds Adafruit_LittleFS's own mutex, and two of them run on other tasks entirely:
Bluefruit's bond saves on the callback task, and - since phone config writes moved
into BLE context - a whole saveToDisk on the BLE task. Neither takes spiLock.

Both writers share one 4 KB page cache, one SoftDevice flash semaphore and one
result word. flash_cache_write repoints that cache when the requested page differs
from the cached one, so a second writer arriving mid-write flushes the first
writer's page and re-points the buffer; the first writer's remaining memcpy then
lands in the wrong page's image. Ring records end up inside LittleFS metadata, or
the reverse. The collision also exhausts the flash layer's 20 x 1 ms busy-retry
budget against an 85 ms page erase, and flash_cache_flush discards the failure, so
32 LittleFS blocks vanish with no error reaching the filesystem. What the user sees
is a torn directory pair on the next mount, a format, and critical error 13.

Take the filesystem mutex in the five ring entry points that reach flash, after
spiLock and never before - the order every existing path already uses. The ring
touches no LittleFS call itself, so the non-recursive mutex is never re-entered.

Longest new hold is a page rotation at roughly half a second, against a 2 s
supervision timeout and a 90 s watchdog. Non-nRF52840 backends are untouched.

* fix(nodedb): make saveProto report a failed readback or rename

SafeFile::close() already verifies the .tmp by hash and renames it over the
live file, and saveProto captured that result, logged it, and then returned
the pb_encode status alone. A torn or half-programmed page therefore counted
as a successful save: for the fullAtomic files the old contents silently
survived, for nodes.proto (written in place) the file was simply gone, and
saveToDisk's recovery path never fired for the one failure it exists for.

* fix(nodedb): retry a failed save before formatting, and never format on a low rail

saveToDisk answered any failed write with an immediate fsFormat(), which is
where most "critical error 12/13" reports and the total config wipe behind
them come from. A write that fails once is far more often a busy SoftDevice
or a VDD dip mid-save than a corrupt filesystem, so:

- retry twice, 150 ms apart, re-checking powerHAL_isPowerLevelSafe() before
  each attempt and before the format; on a low rail return false and leave
  the filesystem alone (the next save lands once the rail recovers, and boot
  already waits for a safe level)
- check fsFormat()'s result instead of assuming it worked
- after a successful format rewrite every segment, not only the ones this
  call asked for: the format took config.proto and the node identity with it,
  so a nodes-only save that ended in a format used to come back up as a new
  node
- with encrypted storage a format also destroys the DEK; skip the resave
  rather than land the private key and PSKs on flash in plaintext

RP2040 feeds its watchdog across the delays, as the neighbouring code does.

* fix(nrf52): quiesce flash before every software reset and power-off

The Adafruit flash layer keeps one 4 KB page image and one SoftDevice flash
semaphore for the whole chip. Every reset path we own - Power::reboot(),
enterDfuMode() (admin enter_dfu_mode_request, which arrives on the BLE task
since #10967), cpuDeepSleep()'s reset and system-off arms, and the
wio-t1000-s secure DFU handler - went straight to NVIC_SystemReset or
sd_power_system_off while another task could be half-way through a page
program or erase. A reset in that window leaves the page erased or partly
programmed; LittleFS finds the torn metadata on the next mount and the
corruption handler formats the filesystem.

nrf52FlashQuiesce() takes spiLock and the LittleFS mutex, waits out whatever
write is in flight, flushes the page cache, and keeps both locks because the
caller resets next. The corruption-reboot handler and __assert_func are left
alone: they run inside the filesystem call stack or a fault, where taking the
mutex would deadlock.

nRF54L is a second copy of these paths since #11867 and still defines
ARCH_NRF52, so it gets the same function on the same core flash layer.

* fix(nrf52): quiesce flash before the library BLE DFU handler jumps to the bootloader

On every board except wio-t1000-s the Nordic DFU service is the framework's
BLEDfu, whose START_DFU handler runs on the callback task and jumps to the
bootloader with no regard for a flash write in progress on the loop task.
That is the OTA path the Apple app and nRF Connect use (Android sends
enter_dfu_mode_request instead, which the previous commit covers).

QuiescingBLEDfu re-installs the control-point write callback after
BLEDfu::begin() and wraps the library's: flush under both locks, then drop
the LittleFS mutex before handing over, because the library reloads the bond
keys through LittleFS on its way to the jump and the mutex is not recursive.
spiLock stays held across the handler: every LittleFS writer on the BLE task
takes it first, the loop task cannot preempt the callback task, and the
handler never blocks after the flush, so nothing can dirty flash before
bootloader_util_app_start(). If the handler returns, nothing jumped, and the
lock is released.

The library callback is a file-static, so it is read back out of the
characteristic through a pointer-to-member obtained via a using-declaration;
that is well-formed C++ and compiles under the pinned GCC 9.3 with LTO.

* test(nodedb): pin the save-failure contract of saveProto and saveToDisk

A failed rename must come back as false from saveProto, a one-off unsafe
rail reading during a write must be retried and land, and a rail still
unsafe at the retry gate must make saveToDisk return false with the
filesystem untouched. The rail is scripted through a strong
powerHAL_isPowerLevelSafe() over the weak native default; on Windows the
default is strong, so only the rename case runs there. The format branch
itself is unreachable natively (a FLASH_CORRUPTION critical error exits
the portduino process), which is what the survival assertions pin.

* fix(nodedb): only format when the filesystem itself is unreadable

Making saveProto honest about write failures gave the recovery path a new way in:
any persistent write failure now reached fsFormat(), which takes every file with
it. A busy or lock-protected nRF52 flash fails every write for as long as it lasts,
so two retries are not enough to tell that apart from a corrupt filesystem, and
guessing wrong costs the node its config, keys and bonds.

Reads settle it. They never touch the SoftDevice write path that a busy flash
fails on, so if /prefs still walks and a stored proto still opens and reads, the
metadata chain is intact and the write failure was transient - return false and
let the caller try again later. Genuine corruption is not silently tolerated: lfs
asserts on it, and the nRF52 handler reboots and formats on the way back up.

Covered by a test that fails without this: a save whose rename cannot succeed,
against an otherwise healthy filesystem, must leave devicestate untouched.

* trunk: exempt Unity test entry points from trufflehog

trufflehog's Lob detector matches "test_" followed by alphanumerics, which
describes every Unity test function name. It fired on a new test in
test_nodedb_save_retry and will fire again on the next suite added. Scoped to
test/**/test_main.cpp, alongside the existing gitleaks exemption for the
synthetic node-DB fixtures.

* fix(nodedb): feed the RP2040 watchdog around the format and the resave

saveToDisk() only feeds the watchdog at the top of each retry. The last
retry, the readable probe, fsFormat() and the five-segment resave then
share one 8 s budget (watchdog_enable in main-rp2xx0.cpp) with no loop
left to feed it. A timeout during the resave leaves the filesystem empty
and the node boots on defaults with a new identity - the exact outcome
this PR exists to prevent, reached by a different road.

Feed once before the probe and again before the resave. Both feeds sit
outside any lock: filesystemStillReadable() takes spiLock itself, and
the format has already released it. ARCH_RP2040 covers rp2040 and rp2350
alike, and the blocks compile out everywhere else, so no other platform
and no native test changes.

Raised by @caveman99 in review.

* fix(nodedb): narrow the save-probe comment and name the full-filesystem case

The comment on filesystemStillReadable() claimed "real corruption asserts
in lfs and formats on reboot". That does hold on nRF52 - nrf52.ini builds
with -DLFS_NO_ASSERT and force-includes cpp_overrides/lfs_util.h, whose
LFS_NO_ASSERT arm routes LFS_ASSERT to the lfs_assert() in main-nrf52.cpp,
which stamps NRF52_MAGIC_LFS_IS_CORRUPT and resets into the format - but
NodeDB.cpp compiles for ESP32, RP2040 and portduino too, where nothing of
the sort is wired up. It is also not true on nRF52 under POFWARN, where
lfs_assert() deliberately skips the stamp. Drop the claim rather than
qualify it three ways.

The log line now names what a field log actually needs to tell apart: a
filesystem that still reads but cannot be written is either busy or full.

Raised by @caveman99 in review.

* fix(sx128x): quiesce flash before the 2.4GHz region reset

reinitChip() saves the region, waits 2 s and resets. On nRF52 that was
the last software reset still going straight to NVIC_SystemReset with a
page program possibly in flight, so "every software reset" in the earlier
commit did not quite hold.

The quiesce stays inside the ARCH_NRF52 arm on purpose. The #else arm
logs and falls through to lora.setCRC() further down, which re-enters
spiBeginTransaction(); a quiesce hoisted above the #if would take spiLock
and never give it back, self-deadlocking portduino and stm32wl. Routing
this through Power::reboot() is wrong for the same class of reason:
setupModules() runs before initLoRa, so its notifyReboot observers and
waypointStore.saveToFlash() are live and would add a flash write to an
aborted radio init.

Raised by @caveman99 in review.

* fix(nrf52): only quiesce on the DFU control write that actually resets

QuiescingBLEDfu wrapped every control-point write, so a write that was
never going to reset still blocked the Bluefruit callback task on spiLock,
forced an early page-cache commit and held back the GATT authorize reply.
Only START_DFU resets; gate on that.

Deliberately no "request->len &&" term. The library's own test is
`request->data[0] == START_DFU` with no length check (BLEDfu.cpp:110 in
both the nRF52 and nRF54 cores), and Bluefruit hands the callback a copy
of a reused event buffer, so a zero-length write carrying a stale 0x01
still resets inside the library. A len term here would let exactly that
reset run unquiesced, which is the case this wrapper exists for. Reading
data[0] is always in bounds: ble_gatts_evt_write_t declares uint8_t
data[1] and the copy covers it.

Raised by @caveman99 in review.
2026-09-17 22:31:32 +00:00
Jonathan BennettandClaude Opus 5 6f3f0bd7c2 Prove explicit acks with Routing.ack_proof (#11877)
* Prove explicit acks with Routing.ack_proof

Explicit acks are ROUTING_APP packets, and ROUTING_APP is excluded from PKC, so
an ack travels under channel encryption alone - and the default channel key is
public. Anyone in range can forge one, and the client grants its strongest
delivery claim on the strength of the ack's unauthenticated `from`.

Where the acknowledged packet was PKI encrypted the endpoints already share a
Curve25519 secret, so the recipient can prove receipt in ~10 encoded bytes:

    ack_proof = HMAC-SHA256(shared_key,
                            "ack" | LE32(from) | LE32(to) | LE32(request_id)
                                  | routing)[0..8)

where `routing` is the encoded Routing message without the ack_proof field,
taken as received with that byte range removed rather than re-encoded. Excising
keeps the value a function of the received bytes alone, so it does not depend on
two implementations' encoders agreeing and does not drop fields this build has
never heard of. For the same reason the sender appends the field rather than
setting it on a decoded struct and re-encoding.

What this does and does not buy. It buys an authenticated delivery receipt from
the actual recipient, which is the property a forged ack costs a user and which
matters where people act on a delivery confirmation. It does NOT protect the
retransmission loop, and must not be described as if it does:
perhapsGenerateImplicitAckForOwnOverheard clears a pending retransmission on any
overheard rebroadcast of our own (from, id), header-only and keyless, so
replaying the originator's own ciphertext stops their retries more cheaply than
forging an ack. No ack authentication of any kind closes that path.

Advisory only, and deliberately not a step toward enforcement. A rule requiring
a proof once a peer has sent one would make a missing proof destroy the only
delivery signal we have, on state the user cannot see: the proof needs the peer
to hold our key, and peer-side eviction, a downgrade or a factory reset are all
invisible to us. A valid proof marks the ack verified; anything else behaves
exactly as today. The client renders the difference.

Verification costs one X25519 and nothing caches the shared secret, so
ackProofPermitsAction gates on findPendingPacket first - otherwise a forged ack
naming any packet id, which is visible in the cleartext header, would force a DH.

Depends on meshtastic/protobufs#1094 for the generated field.

* test: correct a comment that predates ack_proof being generated

The wire-roundtrip test described ack_proof as an unknown field, which was true
while the prototype hand-encoded it. The field is generated now, so this build
understands it - but older firmware does not, which is the case the assertion
actually covers.

* Fix the EXCLUDE_PKI test build, and format

Two copies of test_proof_binds_error_reason and test_proof_binds_direction were
sitting in the #else branch of test_ack_proof, where Identity, makeIdentity,
makeAck, becomeNode, crypto and ACK_PROOF_SIZE do not exist. They were also
unregistered, so they were dead code that only served to break the build. The
native suite never compiles that branch, so a local run could not see it.

Also apply clang-format to a declaration that had been wrapped by hand.

* Exempt test_ack_proof from trufflehog's Lob false positive

Same detector and same shape as the three suites already listed here: it
stitches nearby hex literals into one candidate string, and this suite's node
numbers and request ids (0x0A0A0A0A, 0x0B0B0B0B, 0xABCD1234) happen to match a
Lob API key. The suite holds no literal key material - every key it uses comes
from crypto->generateKeyPair at runtime.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012hLcVif8GDEmA2k77hmFG8

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-17 22:24:29 +00:00
54334ff936 Initial firmware support for Axiometa Genesis Mini (#11852)
* Initial firmware support for Axiometa Genesis Mini

* Report the real AXIOMETA_GENESIS_MINI hardware model

The protobufs now carry AXIOMETA_GENESIS_MINI = 148, so the board no
longer has to masquerade as private hardware.

On ESP32 the -D PRIVATE_HW flag never selected the model on its own -
architecture.h has no arm for it, so the board fell through to the
PRIVATE_HW default at the end of the chain. Give it its own arm and drop
the flag.

* fix(input): sample the encoder button after light-sleep wake

The edge that wakes the device lands while beforeLightSleep() has the
interrupts detached, and a button that is still held produces no further edge
until it is released. The thread stayed parked at INT32_MAX, so the press was
never sampled - no event was emitted, PowerFSM's GPIO-wake branch reads
BUTTON_PIN rather than the encoder pin, and the node dropped straight back into
light sleep with the press swallowed entirely. The second press worked, the
first did not.

Sample once on wake, and only when the button is asserted, so a timer or radio
wake leaves the thread alone. The press then follows the ordinary path and
InputBroker drops the event because the screen was off, so it wakes the screen
and does nothing more - the same behaviour every other input device has.

Rotation stays deliberately non-waking: only the button pin is armed in
doLightSleep(), and the abState re-seed discards a shaft moved during sleep
rather than replaying it as detents.

Reported by CodeRabbit on #11852.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: rcarteraz <robert.l.carter2@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 22:22:05 +00:00
Simplycissmus ff3cc66827 fix(rp2xx0): log and reset on a failed assert instead of hanging (#11853)
RP2xx0 had no __assert_func, so newlib's ran: it prints to stdio and abort()
reaches arduino-pico's _exit, a breakpoint loop. Before rp2040Loop() arms the
watchdog that hangs the node until power is removed; afterwards it costs a
silent stall of up to 8 s.

Install one along the lines of the nRF52 handler: log the failed expression
and reboot through watchdog_reboot().

Refs #11795
2026-09-17 20:52:32 +00:00
github-actions[bot]andjp-bennett 7c3c730a50 Update protobufs (#11887)
Co-authored-by: jp-bennett <5630967+jp-bennett@users.noreply.github.com>
2026-09-17 18:34:04 +00:00
Jason P a4e8b9444f baseui_fixfavoritesonsmalllcds (#11883) 2026-09-17 15:46:42 +00:00
Thomas Göttgens f36a1ea821 nRF52: reclaim flash to bring rak4631 back under its size budget (#11873)
* build(nrf52): drop unused TinyUSB classes and assert function-name strings

Only the CDC class is used on nRF52. Disable the MSC, HID, MIDI, vendor and
video class drivers in the Adafruit TinyUSB config, and pass an empty
__ASSERT_FUNC so assert() no longer embeds __PRETTY_FUNCTION__ strings.
File and line are still reported.

rak4631 estimate: ~7.4 KB flash, ~2.3 KB RAM.

* fix(nrf52): link only the secp256r1 cc310 curve domain

CRYS_ECPKI_GetEcDomain indexes ecDomainsFuncP, which references the
parameter tables of all eleven cc310 curves. Bluefruit LESC pairing only
requests secp256r1, so override the lookup to return that domain alone.

rak4631 estimate: ~7.4 KB flash.

* fix(airtime): replace powf in the channel-utilization EMA fold

foldChannelUtil was the only powf caller on nRF52. The exponent is an
integer step count, so raise the EMA factor by squaring instead; a
multi-day sleep still folds in at most 32 multiplications.

rak4631 estimate: ~1.9 KB flash.

* fix(graphics): use double sin/cos in the compass renderers

The compass renderers were the only sinf/cosf callers on nRF52 screen
builds, pulling in the float trig kernels next to the double ones GeoCoord
already links. Call the double variants instead.

rak4631 estimate: ~3.2 KB flash.

* fix(motion): use double atan2 for magnetometer heading fallbacks

MMC5983MA, QMC6309 and the InkHUD map centre were the remaining
application atan2f callers. The double atan2 is already linked, so the
float variant only added atan2f, __ieee754_atan2f and atanf. The saving
lands once meshtastic/Fusion#1 removes the library's atan2f as well.

rak4631 estimate: ~0.8 KB flash with Fusion#1.

* fix(hopscale): trim diagnostic logging to state changes and anomalies

Drop the save/restore confirmations, the hourly histogram and trend dumps,
the denominator step logs and the per-packet hop_limit log (printPacket
already reports HopLim). Keep the save-failure and histogram-full warnings,
the congestion on/off transition and a single periodic status line, and
remove lastScaledPerHop, which only fed the logs.

* fix(hopscale): silence cppcheck uselessAssignmentArg on restored count

* perf(crypto): use full-schedule AES128/AES256 for AES-CCM

aesSetKey used AESSmall128/AESSmall256, which re-derive round keys for
every block. AES128/AES256 precompute the schedule, encrypt faster and are
already linked by encryptAESCtr, so the AESSmall*/AESTiny* code drops out.
No change on ESP32, where AESSmall* already aliases AES128/AES256.

rak4631 estimate: ~3.9 KB flash; cipher object up to 184 bytes larger.

* perf(nrf52): use the shared software CTR for AES-256 and remove tiny-aes

CryptoCell only accelerates AES-128, which stays on hardware. AES-256 CTR
now calls CryptoEngine::encryptAESCtr (rweather CTR<AES256>, already
linked) instead of the in-tree tiny-aes copy, whose sources were removed
in the previous commit. Output is identical.

rak4631 estimate: ~0.8 KB flash.

* perf(mesh): use std::map for pending retransmissions and API port timestamps

NextHopRouter::pending and PhoneAPI::lastPortNumToRadio were the only
unordered_map instances linked on nRF52. Switching them to std::map, which
is already linked, drops the libstdc++ hashtable, rehash policy and prime
table. GlobalPacketId gains operator<; the unused hash functor is removed.

rak4631 estimate: ~2.1 KB flash.

* perf: parse sensor decimals without strtod

The WS85 serial parser (strtof) and DFRobotLarkSensor (String::toFloat)
were the only callers of newlib's strtod. Add parseDecimalFloat to
meshUtils for plain [+-]digits[.digits] fields and use it at both sites.
Covered by test_type_conversions against strtof.

rak4631 estimate: ~4.5 KB flash.

* perf(gps): compute tan from sin/cos in UTM and OSGR conversion

latLongToUTM and latLongToOSGR were the only tan callers. sin and cos are
already linked, so deriving tan from them drops tan and __kernel_tan.

rak4631 estimate: ~1.1 KB flash.

* fix(graphics): only dispatch the theme menu when TFT coloring is enabled

The Theme option is only offered with GRAPHICS_TFT_COLORING_ENABLED, but
handleMenuSwitch dispatched ThemeMenu unconditionally, linking kThemes and
the theme accessors into monochrome builds where the menu is unreachable.

rak4631 estimate: ~1 KB flash.

* fix(senxx): trim diagnostic logging to errors and user-visible actions

Keep all errors and warnings and a single version line; shorten the admin
action messages; drop progress chatter, state save/restore confirmations and
the per-reading and VOC-state debug dumps. The nested VOC restore branch
collapses to one condition with the same behaviour.

rak4631 estimate: ~2 KB flash.

* build(nrf52): define CRYPTO_AES_NO_DECRYPT

CTR and CCM only encrypt, so the AES inverse tables and round helpers are
dead code on nRF52. Takes effect once the Crypto dependency includes
meshtastic/Crypto#5.

rak4631 estimate: ~1.0 KB flash.
2026-09-17 09:36:41 +00:00
Thomas Göttgens 67e8aafef7 fix(http): hold spiLock only for filesystem calls in the HTTP file handlers (#11870)
* fix(http): hold spiLock only for filesystem calls in the static and upload handlers

* fix(http): hold spiLock only for filesystem calls in the browse and delete handlers

* fix(http): abort an upload when a write comes up short
2026-09-16 22:01:55 +00:00
ae8dee9582 Fix: MQTT topic not updated when LoRa region changes (#10565)
* Initial plan

* Fix: update MQTT topics when LoRa region changes

When the LoRa region is changed via AdminModule::handleSetConfig,
moduleConfig.mqtt.root is updated (e.g. from msh/US to msh/EU_868)
but the running MQTT instance kept using the stale topic strings
(cryptTopic / jsonTopic / mapTopic) that were set at construction time.

Introduce MQTT::reinitTopics() which:
- resets the topic strings to their base values and prepends the
  current moduleConfig.mqtt.root, and
- disconnects from the broker so the next reconnect re-subscribes
  under the new topic prefix.

Call reinitTopics() from MQTT's constructor (replacing the inline
block) so the logic lives in one place, and call it from
AdminModule::handleSetConfig right after moduleConfig.mqtt.root is
rewritten on a region change.

Add a unit test (test_reinitTopicsUpdatesOnRegionChange) that verifies
both the updated subscriptions and the updated publish topic after a
simulated region change.

* Fix: call mqtt->reinitTopics() on region change via menuhandler

* Format MQTT test with trunk style

* fix(mqtt): drop undeclared jsonTopic refs in reinitTopics()

reinitTopics() assigned to a jsonTopic member that does not exist on this
branch (the MQTT class only has cryptTopic and mapTopic), so MQTT.cpp failed
to compile ("'jsonTopic' was not declared in this scope") and broke every
build that compiles it. Remove the jsonTopic lines so reinitTopics() rebuilds
exactly the topics the original constructor did.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* refactor(mqtt): simplify reinitTopics and its call sites

* fix(mqtt): rebuild topics in runOnce when the root changes

Replaces the per-call-site reinitTopics() calls. Also keep the device state and node database segments when the EU clamp swaps the region.

* fix(mqtt): refresh topics in onSend when the root changed

Rename the region change test to underscore-separated segments.

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Ben Meadors <benmmeadors@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Thomas Göttgens <tgoettgens@gmail.com>
2026-09-16 21:47:27 +00:00
Thomas Göttgens ce7e6e448d Separate the nRF54 platform code from src/platform/nrf52 (#11867)
* Move the nRF54L platform code into src/platform/nrf54l15 and drop the ARCH_NRF54L branches from src/platform/nrf52

* Name the platform directory nrf54 so future nRF54 variants can share it

* Rename the nRF54 platform base to nrf54_base in variants/nrf54l15/nrf54.ini

* Leave the SoftDevice random seed to the Bluefruit core, which seeds in begin() and answers NRF_EVT_RAND_SEED_REQUEST

* Seed the SoftDevice from checkSDEvents() when it pops NRF_EVT_RAND_SEED_REQUEST
2026-09-16 21:21:41 +00:00
github-actions[bot]andjp-bennett 7469d52900 Update protobufs (#11874)
Co-authored-by: jp-bennett <5630967+jp-bennett@users.noreply.github.com>
2026-09-16 23:19:47 +02:00
Jonathan Bennett 19dfa1385b Use new store-and-forward original_id field (#11849) 2026-09-16 17:28:24 +00:00