mirror of
https://github.com/meshtastic/firmware.git
synced 2026-09-19 20:39:09 -04:00
thinknode-m9-v2
12725
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ae7e186dfa | Merge branch 'develop' into thinknode-m9-v2 | ||
|
|
2a398fb2fd |
Cancel superseded runs of Check PR Labels and Semgrep Differential Scan (#11906)
* Cancel superseded runs of Check PR Labels and Semgrep Differential Scan Both workflows trigger on pull_request without a concurrency group, so every push, label or edit on a PR queues a fresh run while the earlier ones keep their place in the org's shared runner queue. Check PR Labels listens to six event types, so opening, labelling and editing one PR queued three identical runs within a minute today. Give each the same head_ref-keyed group with cancel-in-progress that CI, Tests and the trunk check already use. Develop pushes are unaffected: they fall through to run_id, as before. * Key the new concurrency groups by PR number, not head_ref Fork PRs often come from a branch named develop, so two of them share github.head_ref and one would cancel the other's run. The PR number is unique per PR. |
||
|
|
5455e38be5 | update device-ui commit reference | ||
|
|
eaf647137e | Merge branch 'develop' into thinknode-m9-v2 | ||
|
|
c47f85f2ff | thinknode v2 keyboard and GPS | ||
|
|
a5dce941dd |
fix(telemetry): follow AS3935Config rename to AS3935State (#11903)
meshtastic/protobufs#1045: admin.proto's AS3935_config and telemetry.proto's AS3935Config collide after name mangling. The flash-persisted message is renamed to AS3935State upstream. Field numbers are unchanged, so existing /prefs/as3935.dat files still decode. Depends on the protobufs rename and the regenerated sources landing first. Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
60f82b4488 |
Fix Mesh Node T1 device image filename (#11893)
custom_meshtastic_images pointed at heltec-mesh-node-t1.svg, which does not exist in web-flasher. The file there is heltec-meshnode-t1.svg. The wrong path returns HTTP 200 with a 7 KB HTML error page rather than the 203 KB SVG, so it does not surface as a 404. |
||
|
|
f13faa6aba |
fix(esp32s3): bound the SerialConsole idle sleep on hardware USB CDC (#11901)
HWCDC::isPlugged() is a SOF watchdog that reads false transiently while USB is connected and working. runOnce() answered that with a 20 s sleep, and nothing wakes the thread on RX, so host traffic sat in the CDC RX ring and reached the API as a burst. Cap the sleep at 250 ms, the rate readStream() already idles at. IS_USB_SERIAL only tested ARDUINO_USB_CDC_ON_BOOT, so ARDUINO_USB_MODE=0 boards ran the same check against a USB-Serial/JTAG peripheral that is not attached to the PHY and never sees a SOF. Gate the check on IS_USB_HWCDC. Measured on tlora-t3s3-v1, 900 s of 1 Hz ToRadio/FromRadio round trips: before 12 stalls, rtt_max 19.96 s, console asleep 26.4% of wall time. After 0 stalls, rtt_max 0.147 s, p50 unchanged at 0.028 s. Fixes #11864 |
||
|
|
11550fa3bd |
Update protobufs (#11904)
Co-authored-by: caveman99 <25002+caveman99@users.noreply.github.com> |
||
|
|
d96c690a91 |
fix(detect): check LPS22HB before SFA30 at 0x5D and CRC-validate SFA30 probe (#11881)
The SFA30 probe only compared the requestFrom() length, which equals the requested length for any device that ACKs, so an LPS33HW/LPS35HW at 0x5D was reported as SFA30. Probe WHO_AM_I first so ST sensors never receive the SFA30 command, and require valid Sensirion CRC-8 on every word of the device marking response. Fixes #11880 |
||
|
|
c29bd00971 |
Rewrite the MQTT region root topic only on the default broker (#11899)
* fix(mqtt): rewrite the region root topic only on the default broker A region change rewrote any root starting with "msh", on any broker. Custom roots such as msh/home were clobbered, and private brokers had their topics moved even though a regional broker is regional already. An empty root, which MQTT treats as the default, was never updated. Region changes now go through MQTT::applyRegionRootTopic(), which rewrites the root only on the default broker and only when the root is empty, the default, or a msh/<region> the firmware wrote itself. * fix(mqtt): parse the broker address before the default-server check Persist module config only when the root actually changed. Replace a stray NUL byte in the test with the \0 escape. * fix(mqtt): count the regional roots as the default root topic |
||
|
|
3aac179397 | fix(phoneapi): wake clients after config sync (#11818) | ||
|
|
0dafcc90fe |
fix(api): retain the unwritten tail on a short TCP API write (#11890)
* fix(api): retain the unwritten tail on a short TCP API write ServerAPI closed the session whenever stream->write() returned fewer bytes than requested. A short write is transmit-buffer backpressure, not a dead socket, and it is most likely during the back-to-back frames of the initial NodeDB dump, so a node at its node cap dropped clients on effectively every connect. Route TCP frames through StreamFrameWriter, the retained-tail path the USB CDC console already uses: the remainder is re-offered on the next pass and the session is closed only when the link itself is gone. Poll at 25ms while output is still undelivered, since nothing wakes the thread when the socket frees transmit space. Fixes #11822 * fix(api): block log re-encoding while a TCP frame is retained emitLogRecord() writes into txBufLog and StreamFrameWriter can now hold that buffer as a retained tail, so a second log record would overwrite bytes the transport has not sent yet. Gate encoding on the retained-frame state, matching SerialConsole. No caller reaches this today (emitLogRecord() is only used by SerialConsole), but retaining the buffer at all is new here. --------- Co-authored-by: Ben Meadors <benmmeadors@gmail.com> |
||
|
|
d4166d0426 |
chore(deps): update libch341-spi-userspace digest to eaaef01 (#11895)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> |
||
|
|
585ce17f59 |
fix(nrf52): stop concurrent flash writers corrupting LittleFS, and stop a failed save formatting it (#11872)
* fix(nrf52): serialise the warm-node ring against LittleFS on the shared flash cache On nRF52840 the warm-node store writes its 3-page record ring straight through flash_nrf5x_write/erase/flush, holding only spiLock. Every LittleFS writer instead holds Adafruit_LittleFS's own mutex, and two of them run on other tasks entirely: Bluefruit's bond saves on the callback task, and - since phone config writes moved into BLE context - a whole saveToDisk on the BLE task. Neither takes spiLock. Both writers share one 4 KB page cache, one SoftDevice flash semaphore and one result word. flash_cache_write repoints that cache when the requested page differs from the cached one, so a second writer arriving mid-write flushes the first writer's page and re-points the buffer; the first writer's remaining memcpy then lands in the wrong page's image. Ring records end up inside LittleFS metadata, or the reverse. The collision also exhausts the flash layer's 20 x 1 ms busy-retry budget against an 85 ms page erase, and flash_cache_flush discards the failure, so 32 LittleFS blocks vanish with no error reaching the filesystem. What the user sees is a torn directory pair on the next mount, a format, and critical error 13. Take the filesystem mutex in the five ring entry points that reach flash, after spiLock and never before - the order every existing path already uses. The ring touches no LittleFS call itself, so the non-recursive mutex is never re-entered. Longest new hold is a page rotation at roughly half a second, against a 2 s supervision timeout and a 90 s watchdog. Non-nRF52840 backends are untouched. * fix(nodedb): make saveProto report a failed readback or rename SafeFile::close() already verifies the .tmp by hash and renames it over the live file, and saveProto captured that result, logged it, and then returned the pb_encode status alone. A torn or half-programmed page therefore counted as a successful save: for the fullAtomic files the old contents silently survived, for nodes.proto (written in place) the file was simply gone, and saveToDisk's recovery path never fired for the one failure it exists for. * fix(nodedb): retry a failed save before formatting, and never format on a low rail saveToDisk answered any failed write with an immediate fsFormat(), which is where most "critical error 12/13" reports and the total config wipe behind them come from. A write that fails once is far more often a busy SoftDevice or a VDD dip mid-save than a corrupt filesystem, so: - retry twice, 150 ms apart, re-checking powerHAL_isPowerLevelSafe() before each attempt and before the format; on a low rail return false and leave the filesystem alone (the next save lands once the rail recovers, and boot already waits for a safe level) - check fsFormat()'s result instead of assuming it worked - after a successful format rewrite every segment, not only the ones this call asked for: the format took config.proto and the node identity with it, so a nodes-only save that ended in a format used to come back up as a new node - with encrypted storage a format also destroys the DEK; skip the resave rather than land the private key and PSKs on flash in plaintext RP2040 feeds its watchdog across the delays, as the neighbouring code does. * fix(nrf52): quiesce flash before every software reset and power-off The Adafruit flash layer keeps one 4 KB page image and one SoftDevice flash semaphore for the whole chip. Every reset path we own - Power::reboot(), enterDfuMode() (admin enter_dfu_mode_request, which arrives on the BLE task since #10967), cpuDeepSleep()'s reset and system-off arms, and the wio-t1000-s secure DFU handler - went straight to NVIC_SystemReset or sd_power_system_off while another task could be half-way through a page program or erase. A reset in that window leaves the page erased or partly programmed; LittleFS finds the torn metadata on the next mount and the corruption handler formats the filesystem. nrf52FlashQuiesce() takes spiLock and the LittleFS mutex, waits out whatever write is in flight, flushes the page cache, and keeps both locks because the caller resets next. The corruption-reboot handler and __assert_func are left alone: they run inside the filesystem call stack or a fault, where taking the mutex would deadlock. nRF54L is a second copy of these paths since #11867 and still defines ARCH_NRF52, so it gets the same function on the same core flash layer. * fix(nrf52): quiesce flash before the library BLE DFU handler jumps to the bootloader On every board except wio-t1000-s the Nordic DFU service is the framework's BLEDfu, whose START_DFU handler runs on the callback task and jumps to the bootloader with no regard for a flash write in progress on the loop task. That is the OTA path the Apple app and nRF Connect use (Android sends enter_dfu_mode_request instead, which the previous commit covers). QuiescingBLEDfu re-installs the control-point write callback after BLEDfu::begin() and wraps the library's: flush under both locks, then drop the LittleFS mutex before handing over, because the library reloads the bond keys through LittleFS on its way to the jump and the mutex is not recursive. spiLock stays held across the handler: every LittleFS writer on the BLE task takes it first, the loop task cannot preempt the callback task, and the handler never blocks after the flush, so nothing can dirty flash before bootloader_util_app_start(). If the handler returns, nothing jumped, and the lock is released. The library callback is a file-static, so it is read back out of the characteristic through a pointer-to-member obtained via a using-declaration; that is well-formed C++ and compiles under the pinned GCC 9.3 with LTO. * test(nodedb): pin the save-failure contract of saveProto and saveToDisk A failed rename must come back as false from saveProto, a one-off unsafe rail reading during a write must be retried and land, and a rail still unsafe at the retry gate must make saveToDisk return false with the filesystem untouched. The rail is scripted through a strong powerHAL_isPowerLevelSafe() over the weak native default; on Windows the default is strong, so only the rename case runs there. The format branch itself is unreachable natively (a FLASH_CORRUPTION critical error exits the portduino process), which is what the survival assertions pin. * fix(nodedb): only format when the filesystem itself is unreadable Making saveProto honest about write failures gave the recovery path a new way in: any persistent write failure now reached fsFormat(), which takes every file with it. A busy or lock-protected nRF52 flash fails every write for as long as it lasts, so two retries are not enough to tell that apart from a corrupt filesystem, and guessing wrong costs the node its config, keys and bonds. Reads settle it. They never touch the SoftDevice write path that a busy flash fails on, so if /prefs still walks and a stored proto still opens and reads, the metadata chain is intact and the write failure was transient - return false and let the caller try again later. Genuine corruption is not silently tolerated: lfs asserts on it, and the nRF52 handler reboots and formats on the way back up. Covered by a test that fails without this: a save whose rename cannot succeed, against an otherwise healthy filesystem, must leave devicestate untouched. * trunk: exempt Unity test entry points from trufflehog trufflehog's Lob detector matches "test_" followed by alphanumerics, which describes every Unity test function name. It fired on a new test in test_nodedb_save_retry and will fire again on the next suite added. Scoped to test/**/test_main.cpp, alongside the existing gitleaks exemption for the synthetic node-DB fixtures. * fix(nodedb): feed the RP2040 watchdog around the format and the resave saveToDisk() only feeds the watchdog at the top of each retry. The last retry, the readable probe, fsFormat() and the five-segment resave then share one 8 s budget (watchdog_enable in main-rp2xx0.cpp) with no loop left to feed it. A timeout during the resave leaves the filesystem empty and the node boots on defaults with a new identity - the exact outcome this PR exists to prevent, reached by a different road. Feed once before the probe and again before the resave. Both feeds sit outside any lock: filesystemStillReadable() takes spiLock itself, and the format has already released it. ARCH_RP2040 covers rp2040 and rp2350 alike, and the blocks compile out everywhere else, so no other platform and no native test changes. Raised by @caveman99 in review. * fix(nodedb): narrow the save-probe comment and name the full-filesystem case The comment on filesystemStillReadable() claimed "real corruption asserts in lfs and formats on reboot". That does hold on nRF52 - nrf52.ini builds with -DLFS_NO_ASSERT and force-includes cpp_overrides/lfs_util.h, whose LFS_NO_ASSERT arm routes LFS_ASSERT to the lfs_assert() in main-nrf52.cpp, which stamps NRF52_MAGIC_LFS_IS_CORRUPT and resets into the format - but NodeDB.cpp compiles for ESP32, RP2040 and portduino too, where nothing of the sort is wired up. It is also not true on nRF52 under POFWARN, where lfs_assert() deliberately skips the stamp. Drop the claim rather than qualify it three ways. The log line now names what a field log actually needs to tell apart: a filesystem that still reads but cannot be written is either busy or full. Raised by @caveman99 in review. * fix(sx128x): quiesce flash before the 2.4GHz region reset reinitChip() saves the region, waits 2 s and resets. On nRF52 that was the last software reset still going straight to NVIC_SystemReset with a page program possibly in flight, so "every software reset" in the earlier commit did not quite hold. The quiesce stays inside the ARCH_NRF52 arm on purpose. The #else arm logs and falls through to lora.setCRC() further down, which re-enters spiBeginTransaction(); a quiesce hoisted above the #if would take spiLock and never give it back, self-deadlocking portduino and stm32wl. Routing this through Power::reboot() is wrong for the same class of reason: setupModules() runs before initLoRa, so its notifyReboot observers and waypointStore.saveToFlash() are live and would add a flash write to an aborted radio init. Raised by @caveman99 in review. * fix(nrf52): only quiesce on the DFU control write that actually resets QuiescingBLEDfu wrapped every control-point write, so a write that was never going to reset still blocked the Bluefruit callback task on spiLock, forced an early page-cache commit and held back the GATT authorize reply. Only START_DFU resets; gate on that. Deliberately no "request->len &&" term. The library's own test is `request->data[0] == START_DFU` with no length check (BLEDfu.cpp:110 in both the nRF52 and nRF54 cores), and Bluefruit hands the callback a copy of a reused event buffer, so a zero-length write carrying a stale 0x01 still resets inside the library. A len term here would let exactly that reset run unquiesced, which is the case this wrapper exists for. Reading data[0] is always in bounds: ble_gatts_evt_write_t declares uint8_t data[1] and the copy covers it. Raised by @caveman99 in review. |
||
|
|
6f3f0bd7c2 |
Prove explicit acks with Routing.ack_proof (#11877)
* Prove explicit acks with Routing.ack_proof
Explicit acks are ROUTING_APP packets, and ROUTING_APP is excluded from PKC, so
an ack travels under channel encryption alone - and the default channel key is
public. Anyone in range can forge one, and the client grants its strongest
delivery claim on the strength of the ack's unauthenticated `from`.
Where the acknowledged packet was PKI encrypted the endpoints already share a
Curve25519 secret, so the recipient can prove receipt in ~10 encoded bytes:
ack_proof = HMAC-SHA256(shared_key,
"ack" | LE32(from) | LE32(to) | LE32(request_id)
| routing)[0..8)
where `routing` is the encoded Routing message without the ack_proof field,
taken as received with that byte range removed rather than re-encoded. Excising
keeps the value a function of the received bytes alone, so it does not depend on
two implementations' encoders agreeing and does not drop fields this build has
never heard of. For the same reason the sender appends the field rather than
setting it on a decoded struct and re-encoding.
What this does and does not buy. It buys an authenticated delivery receipt from
the actual recipient, which is the property a forged ack costs a user and which
matters where people act on a delivery confirmation. It does NOT protect the
retransmission loop, and must not be described as if it does:
perhapsGenerateImplicitAckForOwnOverheard clears a pending retransmission on any
overheard rebroadcast of our own (from, id), header-only and keyless, so
replaying the originator's own ciphertext stops their retries more cheaply than
forging an ack. No ack authentication of any kind closes that path.
Advisory only, and deliberately not a step toward enforcement. A rule requiring
a proof once a peer has sent one would make a missing proof destroy the only
delivery signal we have, on state the user cannot see: the proof needs the peer
to hold our key, and peer-side eviction, a downgrade or a factory reset are all
invisible to us. A valid proof marks the ack verified; anything else behaves
exactly as today. The client renders the difference.
Verification costs one X25519 and nothing caches the shared secret, so
ackProofPermitsAction gates on findPendingPacket first - otherwise a forged ack
naming any packet id, which is visible in the cleartext header, would force a DH.
Depends on meshtastic/protobufs#1094 for the generated field.
* test: correct a comment that predates ack_proof being generated
The wire-roundtrip test described ack_proof as an unknown field, which was true
while the prototype hand-encoded it. The field is generated now, so this build
understands it - but older firmware does not, which is the case the assertion
actually covers.
* Fix the EXCLUDE_PKI test build, and format
Two copies of test_proof_binds_error_reason and test_proof_binds_direction were
sitting in the #else branch of test_ack_proof, where Identity, makeIdentity,
makeAck, becomeNode, crypto and ACK_PROOF_SIZE do not exist. They were also
unregistered, so they were dead code that only served to break the build. The
native suite never compiles that branch, so a local run could not see it.
Also apply clang-format to a declaration that had been wrapped by hand.
* Exempt test_ack_proof from trufflehog's Lob false positive
Same detector and same shape as the three suites already listed here: it
stitches nearby hex literals into one candidate string, and this suite's node
numbers and request ids (0x0A0A0A0A, 0x0B0B0B0B, 0xABCD1234) happen to match a
Lob API key. The suite holds no literal key material - every key it uses comes
from crypto->generateKeyPair at runtime.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012hLcVif8GDEmA2k77hmFG8
---------
Co-authored-by: Claude <noreply@anthropic.com>
|
||
|
|
54334ff936 |
Initial firmware support for Axiometa Genesis Mini (#11852)
* Initial firmware support for Axiometa Genesis Mini * Report the real AXIOMETA_GENESIS_MINI hardware model The protobufs now carry AXIOMETA_GENESIS_MINI = 148, so the board no longer has to masquerade as private hardware. On ESP32 the -D PRIVATE_HW flag never selected the model on its own - architecture.h has no arm for it, so the board fell through to the PRIVATE_HW default at the end of the chain. Give it its own arm and drop the flag. * fix(input): sample the encoder button after light-sleep wake The edge that wakes the device lands while beforeLightSleep() has the interrupts detached, and a button that is still held produces no further edge until it is released. The thread stayed parked at INT32_MAX, so the press was never sampled - no event was emitted, PowerFSM's GPIO-wake branch reads BUTTON_PIN rather than the encoder pin, and the node dropped straight back into light sleep with the press swallowed entirely. The second press worked, the first did not. Sample once on wake, and only when the button is asserted, so a timer or radio wake leaves the thread alone. The press then follows the ordinary path and InputBroker drops the event because the screen was off, so it wakes the screen and does nothing more - the same behaviour every other input device has. Rotation stays deliberately non-waking: only the button pin is armed in doLightSleep(), and the abState re-seed discards a shaft moved during sleep rather than replaying it as detents. Reported by CodeRabbit on #11852. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: rcarteraz <robert.l.carter2@gmail.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
683b3533fd |
setup-base: Move python dependencies into a requirements.txt, pin versions for caching (#11876)
Moves python dependencies declared in `setup-base` to a separate requirements.txt file (which renovate can keep updated) and aligns it with gh-action-firmware. Also removes `pio upgrade`, this is pointless (we just installed the latest platformio in the previous step) |
||
|
|
ff3cc66827 |
fix(rp2xx0): log and reset on a failed assert instead of hanging (#11853)
RP2xx0 had no __assert_func, so newlib's ran: it prints to stdio and abort() reaches arduino-pico's _exit, a breakpoint loop. Before rp2040Loop() arms the watchdog that hangs the node until power is removed; afterwards it costs a silent stall of up to 8 s. Install one along the lines of the nRF52 handler: log the failed expression and reboot through watchdog_reboot(). Refs #11795 |
||
|
|
7c3c730a50 |
Update protobufs (#11887)
Co-authored-by: jp-bennett <5630967+jp-bennett@users.noreply.github.com> |
||
|
|
a4e8b9444f | baseui_fixfavoritesonsmalllcds (#11883) | ||
|
|
7bfa062f27 | fix(esp32): raise T-Watch S3 to support level 1 (#11884) | ||
|
|
dbbd42687b |
chore(deps): update meshtastic/crypto digest to 1c817c2 (#11879)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> |
||
|
|
d1cefc8382 |
chore(deps): update fusion digest to 2051197 (#11878)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> |
||
|
|
f36a1ea821 |
nRF52: reclaim flash to bring rak4631 back under its size budget (#11873)
* build(nrf52): drop unused TinyUSB classes and assert function-name strings Only the CDC class is used on nRF52. Disable the MSC, HID, MIDI, vendor and video class drivers in the Adafruit TinyUSB config, and pass an empty __ASSERT_FUNC so assert() no longer embeds __PRETTY_FUNCTION__ strings. File and line are still reported. rak4631 estimate: ~7.4 KB flash, ~2.3 KB RAM. * fix(nrf52): link only the secp256r1 cc310 curve domain CRYS_ECPKI_GetEcDomain indexes ecDomainsFuncP, which references the parameter tables of all eleven cc310 curves. Bluefruit LESC pairing only requests secp256r1, so override the lookup to return that domain alone. rak4631 estimate: ~7.4 KB flash. * fix(airtime): replace powf in the channel-utilization EMA fold foldChannelUtil was the only powf caller on nRF52. The exponent is an integer step count, so raise the EMA factor by squaring instead; a multi-day sleep still folds in at most 32 multiplications. rak4631 estimate: ~1.9 KB flash. * fix(graphics): use double sin/cos in the compass renderers The compass renderers were the only sinf/cosf callers on nRF52 screen builds, pulling in the float trig kernels next to the double ones GeoCoord already links. Call the double variants instead. rak4631 estimate: ~3.2 KB flash. * fix(motion): use double atan2 for magnetometer heading fallbacks MMC5983MA, QMC6309 and the InkHUD map centre were the remaining application atan2f callers. The double atan2 is already linked, so the float variant only added atan2f, __ieee754_atan2f and atanf. The saving lands once meshtastic/Fusion#1 removes the library's atan2f as well. rak4631 estimate: ~0.8 KB flash with Fusion#1. * fix(hopscale): trim diagnostic logging to state changes and anomalies Drop the save/restore confirmations, the hourly histogram and trend dumps, the denominator step logs and the per-packet hop_limit log (printPacket already reports HopLim). Keep the save-failure and histogram-full warnings, the congestion on/off transition and a single periodic status line, and remove lastScaledPerHop, which only fed the logs. * fix(hopscale): silence cppcheck uselessAssignmentArg on restored count * perf(crypto): use full-schedule AES128/AES256 for AES-CCM aesSetKey used AESSmall128/AESSmall256, which re-derive round keys for every block. AES128/AES256 precompute the schedule, encrypt faster and are already linked by encryptAESCtr, so the AESSmall*/AESTiny* code drops out. No change on ESP32, where AESSmall* already aliases AES128/AES256. rak4631 estimate: ~3.9 KB flash; cipher object up to 184 bytes larger. * perf(nrf52): use the shared software CTR for AES-256 and remove tiny-aes CryptoCell only accelerates AES-128, which stays on hardware. AES-256 CTR now calls CryptoEngine::encryptAESCtr (rweather CTR<AES256>, already linked) instead of the in-tree tiny-aes copy, whose sources were removed in the previous commit. Output is identical. rak4631 estimate: ~0.8 KB flash. * perf(mesh): use std::map for pending retransmissions and API port timestamps NextHopRouter::pending and PhoneAPI::lastPortNumToRadio were the only unordered_map instances linked on nRF52. Switching them to std::map, which is already linked, drops the libstdc++ hashtable, rehash policy and prime table. GlobalPacketId gains operator<; the unused hash functor is removed. rak4631 estimate: ~2.1 KB flash. * perf: parse sensor decimals without strtod The WS85 serial parser (strtof) and DFRobotLarkSensor (String::toFloat) were the only callers of newlib's strtod. Add parseDecimalFloat to meshUtils for plain [+-]digits[.digits] fields and use it at both sites. Covered by test_type_conversions against strtof. rak4631 estimate: ~4.5 KB flash. * perf(gps): compute tan from sin/cos in UTM and OSGR conversion latLongToUTM and latLongToOSGR were the only tan callers. sin and cos are already linked, so deriving tan from them drops tan and __kernel_tan. rak4631 estimate: ~1.1 KB flash. * fix(graphics): only dispatch the theme menu when TFT coloring is enabled The Theme option is only offered with GRAPHICS_TFT_COLORING_ENABLED, but handleMenuSwitch dispatched ThemeMenu unconditionally, linking kThemes and the theme accessors into monochrome builds where the menu is unreachable. rak4631 estimate: ~1 KB flash. * fix(senxx): trim diagnostic logging to errors and user-visible actions Keep all errors and warnings and a single version line; shorten the admin action messages; drop progress chatter, state save/restore confirmations and the per-reading and VOC-state debug dumps. The nested VOC restore branch collapses to one condition with the same behaviour. rak4631 estimate: ~2 KB flash. * build(nrf52): define CRYPTO_AES_NO_DECRYPT CTR and CCM only encrypt, so the AES inverse tables and round helpers are dead code on nRF52. Takes effect once the Crypto dependency includes meshtastic/Crypto#5. rak4631 estimate: ~1.0 KB flash. |
||
|
|
67e8aafef7 |
fix(http): hold spiLock only for filesystem calls in the HTTP file handlers (#11870)
* fix(http): hold spiLock only for filesystem calls in the static and upload handlers * fix(http): hold spiLock only for filesystem calls in the browse and delete handlers * fix(http): abort an upload when a write comes up short |
||
|
|
ae8dee9582 |
Fix: MQTT topic not updated when LoRa region changes (#10565)
* Initial plan
* Fix: update MQTT topics when LoRa region changes
When the LoRa region is changed via AdminModule::handleSetConfig,
moduleConfig.mqtt.root is updated (e.g. from msh/US to msh/EU_868)
but the running MQTT instance kept using the stale topic strings
(cryptTopic / jsonTopic / mapTopic) that were set at construction time.
Introduce MQTT::reinitTopics() which:
- resets the topic strings to their base values and prepends the
current moduleConfig.mqtt.root, and
- disconnects from the broker so the next reconnect re-subscribes
under the new topic prefix.
Call reinitTopics() from MQTT's constructor (replacing the inline
block) so the logic lives in one place, and call it from
AdminModule::handleSetConfig right after moduleConfig.mqtt.root is
rewritten on a region change.
Add a unit test (test_reinitTopicsUpdatesOnRegionChange) that verifies
both the updated subscriptions and the updated publish topic after a
simulated region change.
* Fix: call mqtt->reinitTopics() on region change via menuhandler
* Format MQTT test with trunk style
* fix(mqtt): drop undeclared jsonTopic refs in reinitTopics()
reinitTopics() assigned to a jsonTopic member that does not exist on this
branch (the MQTT class only has cryptTopic and mapTopic), so MQTT.cpp failed
to compile ("'jsonTopic' was not declared in this scope") and broke every
build that compiles it. Remove the jsonTopic lines so reinitTopics() rebuilds
exactly the topics the original constructor did.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* refactor(mqtt): simplify reinitTopics and its call sites
* fix(mqtt): rebuild topics in runOnce when the root changes
Replaces the per-call-site reinitTopics() calls. Also keep the device state and node database segments when the EU clamp swaps the region.
* fix(mqtt): refresh topics in onSend when the root changed
Rename the region change test to underscore-separated segments.
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Ben Meadors <benmmeadors@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Thomas Göttgens <tgoettgens@gmail.com>
|
||
|
|
ce7e6e448d |
Separate the nRF54 platform code from src/platform/nrf52 (#11867)
* Move the nRF54L platform code into src/platform/nrf54l15 and drop the ARCH_NRF54L branches from src/platform/nrf52 * Name the platform directory nrf54 so future nRF54 variants can share it * Rename the nRF54 platform base to nrf54_base in variants/nrf54l15/nrf54.ini * Leave the SoftDevice random seed to the Bluefruit core, which seeds in begin() and answers NRF_EVT_RAND_SEED_REQUEST * Seed the SoftDevice from checkSDEvents() when it pops NRF_EVT_RAND_SEED_REQUEST |
||
|
|
7469d52900 |
Update protobufs (#11874)
Co-authored-by: jp-bennett <5630967+jp-bennett@users.noreply.github.com> |
||
|
|
19dfa1385b | Use new store-and-forward original_id field (#11849) | ||
|
|
ff9a2b1b86 |
fix(http): give a TLS session the contiguous heap it actually needs (#6960) (#11838)
* fix(http): give a TLS session the contiguous heap it actually needs (#6960) An HTTPS connection to a no-PSRAM ESP32-S3 fails its handshake with MBEDTLS_ERR_SSL_ALLOC_FAILED and the client gets a TCP reset. Four things were in the way, and the first two are the bug. mbedtls_ssl_setup() makes two independent calloc() calls, one inbound record buffer and one outbound, each needing its own contiguous block of 16717 bytes. The admission gate compared ESP.getFreeHeap(), which is a sum over every free block and says nothing about the largest one, so on a fragmented heap it waved the connection through and let the handshake be the thing that discovered there was no room. Size the two record buffers separately. Inbound stays at the RFC maximum because browsers and our own MQTT and OTA clients receive full-size records; outbound only ever carries what this device sends, 4 kB at a time at most. That is 12288 bytes of contiguous heap back per session, which is the difference between one usable HTTPS connection and two. These are the ESP-IDF defaults - Espressif's arduino-lib-builder is what turns the asymmetric split off. It replaces CONFIG_MBEDTLS_SSL_MAX_CONTENT_LEN, which no longer applies: that symbol is `depends on !MBEDTLS_ASYMMETRIC_CONTENT_LEN`, and mbedTLS dropped it in 3.0. Then ask the allocator the question mbedTLS is about to ask it, rather than a different one. The probe holds both record buffers at once, as setup() does, through the same heap that CONFIG_MBEDTLS_*_MEM_ALLOC pins mbedTLS to - a plain malloc would be served from PSRAM that mbedTLS never touches. Not heap_caps_get_largest_free_block(), which walks every block under the allocator lock and tripped the interrupt watchdog in #11666. Closed connections are reaped before the measurement, so the library cannot free a slot behind it and accept into it unmeasured, and so the probe sees the memory a finished session just returned. The answer is advisory: the heap can still turn between the probe and setup(), which makes the failure rare rather than routine. A smaller outgoing record makes an existing bug much easier to hit. mbedtls_ssl_write() accepts at most one record per call and returns a short count for the rest, and neither the server library nor Print::print() loops on that, so any response body larger than one record was quietly cut off - /json/nodes has been truncating at 16 kB already. Loop at the bulk write sites. Looping on the returned count also covers a client that negotiates a smaller maximum fragment length, which a fixed-size chunker would not. /json/nodes also built its whole body in one std::string. A full node DB is 200 entries of ~240 bytes, and std::string grows by doubling, so it asked for a 64 kB contiguous block while still holding a 32 kB one, on a board with ~50 kB of heap, through operator new, which aborts rather than throws here. Stream it a node at a time instead. Finally, stop every node in the fleet paying for geofencing. The crossing table reserved all 256 slots at construction, 4 kB, though a node with no geofenced waypoints never tracks a crossing at all. Allocate on the first one instead, stepping the capacity explicitly so there is a size to probe for, and hand the block back once the last waypoint is gone. The probe matters because this allocation has moved off the pristine boot heap onto a live one, where a reserve() that cannot find the block would abort; a refusal now degrades into the bounded-drop path that already exists. heltec-v3 builds clean from scratch with the IDF recompiled, and comes out 20,656 bytes smaller in flash and 172 bytes smaller in static RAM. Full native suite is green at 1414 cases. * style(http,geofence): cut the added comments to the repo's one-or-two-line rule The rationale belongs in the commit message and the PR, not in block comments above every new symbol. No behaviour change; the binary is byte-identical. * fix(geofence): index the capacity probe write to clear cppcheck uninitdata * fix(http): stop the /json/nodes body at the first failed write * fix(geofence): grow crossing state with realloc so a refused allocation drops instead of aborting * refactor(http): probe through esp_mbedtls_mem_calloc and fold the free-slot scan into reaping * refactor(geofence): allocate the crossing table once on first use and free it on store changes * fix(http): probe through mbedtls_calloc so the calls keep C linkage * fix(http): name the check that refused a TLS session, and sum the right heap Both failure modes logged the same line, so a refusal on the free-heap threshold read as fragmentation. The probe returns a verdict now and the warning names the check that failed. That threshold also read ESP.getFreeHeap(), the internal heap, while the probe below it goes through mbedtls_calloc() - PSRAM wherever EXTERNAL_MEM_ALLOC sends mbedTLS, which [device-ui_base] sets. There is no free-size query on the allocator, and the handshake's remaining small blocks want a sum rather than a block to probe for, so the sum asks heap_caps for the same heap. * fix(http): probe the TLS heap only when a client is waiting for a slot The probe holds ~21 kB across two callocs and ran whenever a slot was free, which on a board with nobody connecting is every 50 ms while the web server is active and every second when it is idle - on exactly the low-heap boards this branch is for. It now runs only when a zero-timeout select() on the listen socket says a client is already in the backlog, which is the question loop() asks a few microseconds later. With nothing knocking the server drives its open connections and accepts nothing, so a connection is never accepted past an unmeasured gate; one that lands between the two select()s waits a tick. * perf(http): flush /json/nodes a couple of kB at a time One writeAll() per node is one mbedtls_ssl_write() per node, about 200 of them for a full node DB. Filling a 2 kB buffer first cuts that by roughly 9x and still bounds the allocation, which is why the whole-body string went away in the first place. * fix(geofence): keep the table-full warning for a full table A refused malloc shared the one-shot table-full warning, so a momentary heap dip spent the message that means the 256-entry cap was reached, and the real thing then never printed. They are separate now, and the allocation failure throttles rather than latching, because the next position may well find the 4 kB the table needs. * fix(http): close a waiting client the TLS heap check refuses * fix(http): hold spiLock only for filesystem calls in the browse and delete handlers * refactor(http): leave the browse and delete spiLock scoping to #11870 --------- Co-authored-by: Thomas Göttgens <tgoettgens@gmail.com> |
||
|
|
4b3f2a2beb | fix drop_stale_sdkconfig_defaults() (#11861) | ||
|
|
a70fe9ea10 |
Actions: Build MacOS 27, Drop 15 (#11851)
MacOS 15 will still be built in homebrew, but not tested as part of CI here. |
||
|
|
ee76117835 |
fix(router): relay opaque packets in CORE_PORTNUMS_ONLY (#11844)
* fix(router): relay opaque packets per rebroadcast_mode, not only in ALL |
||
|
|
0caf3a09f8 |
ci(size-budget): raise the rak4631 flash budget 746 000 -> 752 000 (#11866)
develop has been over the 746 000 line since #11826 (748 088 on the
CI runner at
|
||
|
|
2d6dad9ee9 |
Portduino: Fix LR2021 switch tables, power ceilings and IRQ handling (#11382)
* fix(portduino): recognise the LR2021 power ceilings in --check loadConfig() has read Lora.LR2021_MAX_POWER and Lora.LR2021_MAX_POWER_HF since LR2021 support landed, but neither was listed in the config checker's schema. --check therefore reported both as "unknown key ... ignored by meshtasticd" -- false, and actively misleading: it tells the user to delete a key that is doing exactly what they wanted. This breaks the contract stated above schema(), that a key taught to loadConfig() is added there too. CI enforces that by running --check over bin/config.d/**, but no shipped config sets either key -- or mentions lr2021 at all -- so nothing ever tripped over the omission. It could only surface for someone hand-writing an LR2021 config. LR20x0 is the only module with two power ceilings, one per band, selected at runtime by region; every other family expresses the split as separate module names and needs a single key. That is the likely reason the pair was missed while every other *_MAX_POWER key was added. Also adds both to valueSpecs(), so a wrong-typed value is reported rather than silently replaced by the default. * feat(portduino): configurable IRQ DIO and a chip-neutral RF switch table for LR20x0 Two gaps found bringing an LR2021 up under meshtasticd on a Luckfox Lyra Zero W. Both sit in the LR20x0 support added in #11252, and they interact: the switch table has to be written slightly wrong to pass validation, and the interrupt lands on a pin that table is driving. Every symptom is silent, because begin() only exercises SPI and BUSY -- the radio reports init success and then receives nothing. IRQ DIO could not be set on Portduino ------------------------------------- LR20x0Interface picked the IRQ DIO purely at compile time, and neither LR2021_IRQ_DIO_NUM nor IRQ_DIO_NUM exists for a Portduino target, so meshtasticd always fell through to RadioLib's default of DIO5 -- which is also the first RF switch line on carriers using the DIO5-DIO8 table. A variant says this with a #define (the pro-micro DIY board uses DIO9); a carrier has only the YAML, and had no way to say it. Adds Lora.IRQ_DIO_NUM, and an ARCH_PORTDUINO branch after the two existing #define branches, so a variant that already sets one still wins. The switch table was parsed as LR11xx-only ------------------------------------------ Pin names resolved to RADIOLIB_LR11X0_DIOn whatever the radio, and the mode set was the LR11xx's, so MODE_RX_HF -- a mode the LR20x0 really has -- was rejected as an unknown key and had to be omitted. The two families are not interchangeable: an LR11xx has no DIO9, so its fifth switch slot is DIO10, while an LR20x0's fifth slot is DIO9 and DIO10 is its sixth. A table naming DIO10 was therefore driving the wrong pin on an LR20x0. The YAML layer now stores what was written -- a DIO number and a neutral mode id -- and each interface supplies its own DIO constants and OpMode_t map to a shared builder. Neither family's constants are assumed to coincide with the other's. This also fixes a round trip in the config writer, which decoded pins by comparing against RADIOLIB_LR11X0_* and always emitted five values per mode row: for a four-pin table it produced YAML that --check would reject for mismatched row lengths. --check ------- Findings are now judged against the resolved module rather than a fixed list, so a mode or pin the part does have can no longer be rejected, and one it does not have is named instead of silently accepted. The claim that the table "is only applied to LR11xx radios" was stale and is corrected, and the missing-table warning now covers both families. "auto" is excluded throughout: the module has not been probed yet, so absence cannot be judged. The IRQ/switch-pin collision is reported in both directions, including the harder case where no key is set and the radio default collides -- nothing in the file looks wrong. Note that listing a pin is what breaks it, not driving it: setRfSwitchTable() reassigns the DIO function for every pin in the list whatever the levels say, so an all-LOW column is still a collision. Seven fixtures cover these, including a false-positive guard: DIO5 as the interrupt is normal, and must stay silent when the table is elsewhere. * feat(portduino): let the YAML ask for a TCXO probe, across every family that has one A variant declares "a TCXO may or may not be fitted" at compile time with TCXO_OPTIONAL, because the board is known when the image is built. A Portduino carrier cannot: the same meshtasticd binary runs on hardware populated either way, so the statement has to arrive as YAML and be answered at runtime. Adds Lora.TCXO_OPTIONAL, and TCXO_OPTIONAL_ENABLED in RadioLibInterface.h to unify the two, so each driver asks the question once rather than growing a second, Portduino-shaped code path. On an embedded target it stays a compile-time constant, so `if (TCXO_OPTIONAL_ENABLED)` folds away exactly as the old `#if` did: the nrf52_promicro_diy_tcxo image, which defines TCXO_OPTIONAL and so exercises the converted branches, still ends at 0xDF1D0 -- the same address as before this change. It is defined there rather than in a header of its own because InterfacesTemplates.cpp includes all three interface .cpp files into one translation unit, where a per-file definition would collide. Covers every family that has a TCXO reference to probe for: SX126x (sx1262/sx1268/LLCC68), LR11xx and LR20x0. With no DIO3_TCXO_VOLTAGE given, the TCXO attempt uses RadioLib's own 1.6 V default rather than being skipped -- otherwise there is nothing to fall back FROM and the flag would silently do nothing. This is also the FIXME that sat on the Portduino branch in LR20x0Interface: an unset voltage now means "no TCXO" explicitly. Two things are deliberately left alone: Each family keeps its own probe order. LR11xx tries XTAL first, because a TCXO-first attempt hangs RadioLib's unbounded calibration wait on a module with no TCXO fitted, whereas XTAL fails fast and cleanly on a module that has one; LR20x0 and SX126x try the TCXO first. A carrier therefore behaves the same way in a Portduino build as in an embedded one, and changing an order stays a hardware-behaviour decision rather than a tidying-up one. The SX126x retry is Portduino-only. An embedded TCXO_OPTIONAL board already gets this from initLoRa(), which constructs a second SX126x interface with no Vref when the first fails; retrying inside init() as well would leave that ladder step unreachable and change how every existing t-echo-class board reports its oscillator. A Portduino build has no ladder to fall through, because the module is named in YAML rather than probed. --check learns the key, reports which Vref will actually be tried, and warns when it is set on a radio with no TCXO reference, where it is read, stored and inert. * docs(portduino): condense the comments on this branch The repo asks for one or two lines and no multi-paragraph blocks, on the grounds that the diff and the commit message carry the rationale while the code carries the behaviour. What landed here was well past that: 163 added comment lines, including a 26-line block above a single macro. Removes the rhetoric, the issue numbers and the before-and-after asides, and the notes on where a thing used to live. No added block is longer than three lines now. Two facts needed stating and are stated once each rather than repeated at every use: the slot/DIO divergence between the families, in PortduinoGlue.h, and the per-family TCXO probe order, in RadioLibInterface.h. The longest surviving explanation is why an all-LOW switch column still collides with the interrupt, which sits in the fixtures README because without it that pair of fixtures reads as contradictory. Comments only; no functional change. * address CodeRabbit review on #11382 - SX126xInterface: distinguish an explicit DIO3_TCXO_VOLTAGE from the TCXO_OPTIONAL probing default in the debug log instead of always claiming the config field was set. - ConfigCheck: modesFor() now reports an unresolved use_autoconf against the union of both radio families' modes, not the LR11xx subset - fixes a false "not a mode this part has" warning for valid LR2021-only modes (e.g. RFSW_RX_HF) before autodetection resolves the module. - PortduinoGlue loadConfig: build rfswitch_mode_high[m] as a fresh per-row bitmask instead of OR-accumulating onto a stale value, so a config re-parse can clear a slot back to LOW. - PortduinoGlue YAML serialization: gate rfswitch_table emission on has_rfswitch_table rather than rfswitch_dio_num[0] >= 0 (missed sparse pin lists), and track each emitted pin's original slot so row values line up correctly instead of shifting when a low slot is absent. - config-dist.yaml: document the per-family TCXO/XTAL probe order (SX126x/LR20x0 TCXO-first, LR11x0 XTAL-first). - Trim three overlong comments per the coding-guideline nitpicks. Left the SX126x XTAL-retry-on-oscillator-failure nitpick alone - RadioLib's begin() already does its own XOSC_START_ERR recovery internally, and narrowing our wrapper's retry condition on top of that needs hardware to verify it doesn't regress a real failure path. Verified: bin/test-config-check.sh GREEN 69/69 against an isolated native build; pio test -e native -f test_rtc PASSED. * fix CI: cppcheck duplicateValueTernary, harden kRfSwitchModes init LR11x0Interface::init(): work around cppcheck's duplicateValueTernary on `TCXO_OPTIONAL_ENABLED ? 0 : tcxoVoltage` (both branches fold to 0 on a board with no ARCH_PORTDUINO, no TCXO_OPTIONAL, and no explicit Vref, since tcxoVoltage already reduces to 0 via the same macro chain) by splitting it into a plain assignment + if, rather than suppressing the warning. The other TCXO_OPTIONAL_ENABLED ternaries in this PR (LR11x0Interface.cpp:75, SX126xInterface.cpp:76, LR20x0Interface.cpp: 86,240) pick between TCXO_OPTIONAL_DEFAULT_VOLTAGE (1.6f) and 0, which can never coincide, so they're unaffected and left as-is. ConfigCheck.cpp: kRfSwitchModes was a namespace-scope global with dynamic initialization (a lambda IIFE) reading kRfSwitchModeNames, which is defined in a different translation unit (PortduinoGlue.cpp). Currently safe only because kRfSwitchModeNames's initializer is constant-expression-only (string literals + enum constants), which the standard guarantees completes before any TU's dynamic initializers - but that safety is silent and would break if PortduinoGlue.cpp's array initializer ever stopped being a constant expression, with nothing to warn a future editor. Converted to a function-local static (Meyers' singleton), which is correct by construction regardless of the other TU's initializer, updating all 4 call sites (definition + 3 uses) from kRfSwitchModes to kRfSwitchModes(). Verified: pio test -e native -f test_radio PASSED; bin/test-config- check.sh GREEN 69/69 against an isolated native build. * refactor: simplify TCXO voltage handling across interfaces and improve comments * fix rfswitch_table cross-file merge; drop now-stale checker warning Three CodeRabbit findings on |
||
|
|
3468af94aa |
fix(hopscale): gate hop scaling on measured channel congestion (#11826)
* fix(hopscale): gate hop scaling on measured channel congestion * fix(hopscale): rename test tripping trufflehog and release congestion at the threshold * test(hopscale): pin the unscaled precondition in the busy-channel gate test * fix(hopscale): floor infrastructure roles and engage below the polite gate * test(hopscale): name the symbol under test and reset the gate in clear() * fix(hopscale): drop the unneeded congestion reset in clear() * refactor(hopscale): drop the unused utilization accessor and duplicated log fields * refactor(hopscale): drive the politeness extension from measured congestion (#11831) * refactor(hopscale): drive the politeness extension from measured congestion * refactor(airtime): own the smoothed channel utilization (#11832) * refactor(airtime): own the smoothed channel utilization * refactor(airtime): cut comments to the two-line limit * fix(airtime): fold the smoothed utilization once per crossed bucket * fix(hopscale): compare the congestion thresholds at whole-percent resolution |
||
|
|
bc2528b005 | Show the full Bluetooth pairing PIN on tiny OLED panels by drawing it full-screen with all lines spread evenly. (#11855) | ||
|
|
e0cf782131 | Build xiao_nrf54l15_lr2021 on every PR as the nrf54l15 canary and mark it community supported (#11856) | ||
|
|
2fe711115e |
feat(nrf52): add RAK3401 + LR2021 (RAK13700) variant (#11819)
* feat(nrf52): add RAK3401 + LR2021 (RAK13700) variant WisBlock core with IO-slot LR2021: DIO RF switch, board LF PA table, and 1.6 V TCXO. Keep extra/unsupported until it has its own hw_model. * fix(lr2021): log custom PA setOutputPower in fullBegin Match init(): a calibration miss stays a warning so band-hop keeps the begin() PA config. * refactor(lr2021): share custom LF PA table helper init() and fullBegin() both re-install the board table after begin(); keep the warn-only setOutputPower miss. |
||
|
|
ea7d4aa410 |
Port nRF54L15 to the s145 SoftDevice Arduino core (#11842)
* Remove the Zephyr based nRF54L15 port * Add nRF54L15 port on the s145 Arduino core: nrf54l15dk and xiao_nrf54l15 variants * nRF54L: errno-style nrfx results, flush console before assert reset * nRF54L: log the SoftDevice status on Bluefruit failure, ignore the seed request event * Support the Wio-LR2021 LoRa Plus expansion board with OLED and K1 on the XIAO nRF54L15 variant * Consume the nRF54L15 platform, core and bootloader from their repositories * Pin the nRF54L15 platform to v0.2.0 * nRF52: forward SoftDevice flash events taken by the main loop to the flash driver, log the pairing failure status * Pin the nRF54L15 platform to v0.2.1 * Pin the nRF54L15 platform to v0.3.0 * Split the XIAO nRF54L15 variant into SX1262 and LoRa Plus environments, seed the SoftDevice on request, add the nrf54l15 CI build script * Pin the nRF54L15 platform to meshtastic/platform-nordicnrf54 v0.3.1 * NRF54: Fix mtjson generation --------- Co-authored-by: vidplace7 <vidplace7@gmail.com> |
||
|
|
9d27b276aa |
NRF54: Update to toolchain-gccarmnoneeabi@1.90301.200702, align NRF52840 (#11850)
Version 1.90201.191206 supports arm64 MacOS, but not arm64 Linux (1.90301.200702 supports both) Also change the fuzzy match for gccarmnoneeabi to an exact match for nRF52840 (for reproducible builds), this is effectively a no-op change, today. |
||
|
|
2826c6712b | Pin nrf54 platform, update toolchain-gccarmnoneeabi for arm64 build hosts (#11848) | ||
|
|
f5158f50be |
feat(rp2040/rp2350): Update earlephilhower/arduino-pico to 6.1.0, bump maxgerhardt/platform-raspberrypi to latest (#11814)
* Update earlephilhower/arduino-pico to 6.1.0 * feat(rp2040/rp2350) bump platform-raspberrypi * fix(rp2040/rp2350) recover printf and scanf |
||
|
|
32eb1a1237 |
Honor an explicit -c config path when -s is given (#11348)
The simradio flag (-s) is the first branch of an if/else-if chain that also handles config loading, so it short-circuits every later branch -- including the one for an explicit -c <path>. Skipping config discovery under -s is intended, but a config path the user passed by hand is not discovery, and it is silently ignored today. Move the -s check after the -c branch so an explicit path is always parsed, and skip only the implicit discovery (./config.yaml, /etc/meshtasticd/config.yaml) when -s is given without -c. The radio override then runs after every config source, since -c and its ConfigDirectory entries can both set Lora.Module and -s has to win over them. Doing it there also fixes --check and --output-yaml, which reported the configured module rather than the simulated one because the old override sat behind an early return. Behaviour with a bare -s is unchanged: no YAML is loaded and the radio is the simulator. Co-authored-by: Jonathan Bennett <jbennett@incomsystems.biz> |
||
|
|
ee15508494 |
time: arm the remaining 0-means-unset stamps through the helpers (#11830)
* time: add skipZero/safeMillis/timerEndsAtMillis helpers
skipZero() steps a millis value past 0, since stored stamps and deadlines
conventionally use 0 for "unset" and the one tick per ~49.7-day wrap that
lands on 0 would otherwise read as never-set.
safeMillis() covers a bare stamp; timerEndsAtMillis(delayMs) covers a
deadline, where the sum is what has to dodge 0 - a non-zero read plus a
delay lands there once per wrap - so it is not safeMillis() + delayMs.
* time: replace hand-rolled zero-dodging with the UptimeClock helpers
PacketHistory rxTimeMsec, EncryptedStorage s_lastFailMillis (stamps), and
SGM41562 lastRefreshMs_ / NextHopRouter learnedAtMsec (ternary stamps) each
hand-rolled skipZero() in place; swap in safeMillis()/skipZero() directly.
HapticFeedback pulseOffAt/delayedPulseAt and GPS fixHoldEnds hand-rolled the
deadline form - millis() + delay, then remap a 0 result to 1 - swap in
timerEndsAtMillis(delay).
No behavior change; each site keeps the value it already computed.
* time: guard the remaining 0-means-unset deadline/stamp writes
rebootAtMsec, shutdownAtMsec, and NotificationRenderer::alertBannerUntil are
all read back with a bare == 0 / != 0 check for 'not scheduled', but every
write site computed millis() + delay (or a bare millis() stamp) with no
guard against landing exactly on 0 - the same wrap hazard skipZero() exists
for, just never applied here.
Route every rebootAtMsec/shutdownAtMsec/alertBannerUntil write through
timerEndsAtMillis()/safeMillis(); RadioLibInterface's reboot-on-stuck-tx
sums an already-captured stamp rather than "now", so it goes through
skipZero() directly instead.
No behavior change outside the ~1-in-2^32 wrap window each site was
already exposed to.
* time: guard three more 0-means-unset deadline writes
ntp_renew (ethClient.cpp), suppressTouchTapUntilMs (Events.cpp), and tx_after
(RadioLibInterface.cpp) all read back 0 as a real state - forced NTP renewal,
no suppress window active, no TX delay armed, respectively - but each arm
site wrote a bare millis()/getMillis() + delay with no guard against the sum
landing exactly on 0.
Route each through Time::timerEndsAtMillis(). No behavior change outside the
wrap window each site was already exposed to.
Refresh the Throttle.h TODO list to note ntp_renew is converted too.
* motion: guard the calibration deadline and use Throttle::deadlinePassed
endCalibrationAt's arm site wrote millis() + calibrateFor with no guard
against landing on 0, the same value finishCalibrationIfExpired()/
drawFrameCalibration() treat as "not calibrating". Route it through
Time::timerEndsAtMillis().
Also swap finishCalibrationIfExpired()'s hand-rolled (int32_t)(now - deadline)
< 0 for Throttle::deadlinePassed(): same wrap-safe comparison the codebase
already provides, without the signed-cast pattern Throttle.h documents as
implementation-defined past INT32_MAX, and it drops the file's last direct
millis() call in favor of the Time:: wrapper the rest of it already uses.
* time: fix Throttle::execute()'s own zero-dodging
Both places execute() writes *lastExecutionMs - the first-ever-run branch
and the regular update - used bare Time::getMillis() with no guard against
landing on 0, which is the exact sentinel this function reads back as
"never run" one line above. A hit there makes the next call re-fire
immediately instead of respecting minumumIntervalMs.
Capture now via Time::safeMillis() once; every use downstream (the elapsed
comparison, the stored value) is then safe by construction instead of
needing the guard reapplied at each write.
* revert some safeMillis cases where overflow is a bad thing
* test(uptime): pin skipZero/safeMillis/timerEndsAtMillis at the wrap boundary
Covers the zero case, an ordinary nonzero value, and a sum that lands
exactly on 0 from a nonzero start - the case timerEndsAtMillis() exists
for, and the one the prior suite had no direct coverage of.
* time: restore the route-health write normalization and put it on one clock
noteRouteLearned()/noteRouteSuccess() lost their `now ? now : 1` normalization,
leaving learnedAtMsec able to store 0 - which getOrAllocRouteHealth() reads as an
ever-growing age, making the slot the first eviction candidate and permanently
stale. Normalize at the write, where the block comment already says it happens,
so every caller is covered rather than just today's two.
Both callers, the two isRouteStale() sites and doRetransmissions() now read
Time::getMillis(), so the stamp and every comparison against it share a clock.
doRetransmissions() goes back to getMillis(): its `now` feeds only comparisons,
never a 0-sentinel field, so skipping zero there only cost accuracy.
* time: read the haptic, InkHUD and calibration deadlines on the write's clock
These three deadlines were converted to Time::timerEndsAtMillis() on the write
side while their reads stayed on millis(), so each spanned two clocks and would
fire immediately or never under an injected test clock. Convert the reads to
match: HapticFeedback::scheduleNext()/runOnce(), the InkHUD tap-suppression
window, and the calibration countdown's read-back of screen->getEndCalibration().
MotionSensor's sampledAtMs is left alone - its write and read are both millis()
and consistent already.
* time: correct the sentinel notes to match what the code actually does
The Throttle.h enumeration claimed the remaining timerEndsAtMillis() callers
"already dodge the sentinel", which reads as a completeness claim the same branch
contradicts: RadioLibInterface's tx_after and activeReceiveStart are both 0=unarmed
and both still arm from bare millis(). Name them instead, so the deadline-type
conversion has the real list. The ntp_renew entry now separates a deliberate 0
("due now", forced at link-up) from a computed one, which is what changed there.
The three TODO(elapsed-stamp) blocks ran four and five lines against the repo's
one-or-two rule, and two of them argued their case wrongly. Throttle.cpp implied
safeMillis() simply doesn't help; in fact neither store is safe on the wrap tick -
the 1 underflows a same-instant read, the 0 re-takes the never-run branch - which
is the symmetry worth recording. PacketHistory.cpp called its dodge "reflecting
the previous pattern" when it is load-bearing: rxTimeMsec 0 means "empty slot"
(PacketHistory.h:21) and insert() drops a record stamped 0 outright, so without it
a packet arriving on the wrap tick is never stored and loses its dedup.
Also picks up trunk fmt's trailing-whitespace fix in Throttle.cpp and the comment
realignment in SGM41562.cpp that this branch's added comment knocked out.
* test(nexthop): pin the route-health stamp against the 0 sentinel
The uptime suite covers skipZero/safeMillis/timerEndsAtMillis themselves, but
nothing covered a call site, so the branch deleted noteRouteLearned()'s
normalization and stayed green. None of the existing route-health tests pass 0 as
`now` - they use 1000, learnAt, or millis() - (TTL + 5000) - which is exactly the
gap the regression went through.
Both new tests fail with "Expected 0 to be not equal to 0" when the skipZero() is
backed out of NextHopRouter, and pass with it. noteRouteSuccess() only refreshes
an existing record, so its twin learns a route first to reach the write.
Also drops a self-referential assertion in the uptime suite: comparing
getMillis() against safeMillis() passes even if safeMillis() does no dodge at
all, so it now asserts the literal.
* discard safemillis for skipzero (better semantics and therefore maintainability) and make consistent use of getmillis where it is called (to permit testing)
* more wrapzero safety
* STM gets some too
* time: stop the next 0-means-unset deadline being armed from raw millis()
The fields this branch armed through Time::timerEndsAtMillis() / Time::skipZero()
are the kind that get added by copy-paste: `rebootAtMsec = millis() + N` appears
at twenty-odd sites across six files, and the next module to defer a reboot will
be written from one of them. Nothing catches the mistake afterwards - the sum
lands on 0 for one tick per ~49.7-day wrap, so a test run, a soak and a bench
session all pass while a pending reboot, shutdown, DFU jump or banner expiry is
silently dropped.
Two guards, at the two places it can go wrong.
The helpers themselves: skipZero() is constexpr, so its contract is now pinned by
static_assert in the header rather than only by test_uptime_clock. The asserts are
chosen against the two plausible rewrites - `ms | 1` perturbs every even value and
`ms + 1` turns the last tick of the wrap into the 0 the function exists to avoid.
Both compile, and both pass a test that only checks skipZero(0); each trips a
distinct assert here, naming the failure mode.
The call sites: bin/lint-unset-sentinel-millis.sh flags a sentinel field in src/
assigned from a raw millis()/getMillis() read, and names the helper to use. It is
name-driven because the 0 contract is declared in src/main.h and enforced in six
other files, so no single-file scan can infer it; every one of the thirteen fields
was checked to actually test against 0 before being listed. nagCycleCutoff and
LinuxJoystick's nextRepeatX/nextRepeatY are deliberately absent - their unset state
is a separate bool - and the nine remaining `millis() + x` sites in src/ are locals
that never store 0 for anything to misread.
Blocking, unlike its note-level neighbours: there is no run-time enforcer to pair
with, and the tree has zero violations today, so gating costs nothing. Scoped to
src/ so test_uptime_clock can keep building raw wrap values on purpose.
bin/test-lint-unset-sentinel-millis.sh pins the scanner against 23 fixtures -
reads, disarms, shadowing locals, comments, string literals and the already-fixed
forms all have to stay quiet.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* time: guard the four 0-means-unset stamps this branch had missed
Sweeping src/ for the `if (stamp && <deadline check>)` idiom - the shape that
makes 0 mean "unset" - turned up four stamps still armed from a raw clock read,
so the new lint rule would have had to either ignore them or go red on checkout.
Each is the same one-tick-per-wrap hole the rest of the branch closes:
* TrackballInterruptBase lastInterruptTime, armed in all four ISR handlers and
explicitly disarmed to 0 at the threshold reset. getMillis() is the ISR-safe
read by construction - it compiles to millis() outside PIO_UNIT_TESTING - and
skipZero() is pure, so neither adds anything to interrupt context.
* NeighborInfoModule lastSentReply, read as `if (lastSentReply && ...)` before
the 3-minute reply throttle. Needed the UptimeClock.h include.
* PositionModule lastSentReply, same throttle; already on the injectable clock
but still missing the guard.
* NodeDB lastSort, whose own read spells the sentinel out as `lastSort == 0 ||`.
On the wrap tick each would read as never-stamped: a trackball debounce window
lost, a neighbour or position reply sent inside the throttle it was meant to
respect, one extra NodeDB sort. Cheap individually, which is why they were missed.
All four are now listed in bin/lint-unset-sentinel-millis.sh, so the rule covers
every field in the tree that actually tests against 0 rather than a subset, and
the header records the eight stamps left off for the opposite reason - their unset
state is a separate flag (isNagging, busyTx, heldX/heldY, formatted_this_boot,
heartbeat, gotwind, haveSample, lastIaqValid), so 0 is a value they may legally
hold. The rule is silent across src/ on this tree.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* lint: let a site opt out of the sentinel rule, with its reason on the record
The rule is blocking, so it needs an escape hatch for the site where 0 genuinely
is a legal timestamp - and the hatch should cost something, or it becomes the
first thing anyone reaches for. `unset-sentinel-ok: <reason>` in a comment on the
write, or on a comment line above it, suppresses that one statement:
// unset-sentinel-ok: busyTx carries the armed state, so 0 is a legal stamp here
lastTxStart = Time::getMillis();
The reason is mandatory. A bare `unset-sentinel-ok`, or a colon with nothing
after it, is reported instead of honoured - with a message saying so - so the
only way to silence a site is to write down why it is safe. trunk-ignore still
works, but this states the justification at the write and also applies when the
script runs outside trunk.
The marker is read from comment text collected during the same character-level
pass that strips comments and literals, not by re-scanning the raw line. That is
what keeps it out of reach of data: LOG_DEBUG("unset-sentinel-ok: ...") mutes
nothing, because a string literal is not a comment. It is also consumed by the
statement it was written for, so it cannot leak onto the next write - while still
carrying across any number of intervening comment lines to the statement below,
which is where a real justification wants to be written.
Twelve fixtures added for the new behaviour: both comment styles, block and
multi-line block comments, the bare form, the marker-in-a-string cases, and three
leak cases. 35 total, all green, under bash 3.2 as well.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* lint: watch the separate-flag stamps too, with their exemption stated at the write
The nine stamps whose armed state lives in a companion boolean were previously
just absent from the rule's list, which meant the reasoning for leaving them out
existed only as prose in a shell script. They are now listed and individually
opted out at the write, naming the flag that actually carries the armed state:
// unset-sentinel-ok: haveSample carries the armed state, so 0 is a legal stamp
lastSampleMs = Time::getMillis();
The point is what happens later. If someone rewrites `if (haveSample && ...)` as
`if (lastSampleMs && ...)`, the field has silently acquired the 0 contract; with
the opt-out sitting at the write, the claim to re-examine is in front of whoever
makes that edit instead of buried in bin/.
Every exemption was checked against its real read sites before being written, and
three candidates did not survive that check. They stay off the list, because
listing one would mean stamping an opt-out over a claim that does not hold:
* nagCycleCutoff. handleInputEvent reads `if (nagCycleCutoff != UINT32_MAX)`
without consulting isNagging, so at that read the field is its own armed flag
with UINT32_MAX as the sentinel - and the arm at ExternalNotificationModule
.cpp:521 can land exactly there. skipZero() cannot help: it lifts 0 to 1 and
leaves UINT32_MAX alone, which UptimeClock.h's own static_assert pins. There
is also a live boot-state bug behind this - the in-class initializer is 1
while isNagging starts false - and fixing the read is a behaviour change that
belongs in its own PR.
* TouchScreenBase::_start. Overloaded as an event stamp AND a `+ 30000`
suppression deadline compared by signed subtraction, so a near-zero value
reads as "long ago" rather than "armed 30s out" and LONG_PRESS re-fires.
skipZero() does not fix this one either: 1 reads as long-ago exactly as 0
does. It needs the stamp and the deadline held separately.
* StoreForwardModule::retry_delay. No reads at all today, so nothing misbehaves
yet; exempting it now would pre-approve the raw arm for whoever implements the
retry its own comment promises.
The rule is silent across src/ on this tree, and the header records all three
rejections so the next person does not have to re-derive them. The self-test's
negative fixture no longer uses nagCycleCutoff as its example of a safely
unlisted field - that would have encoded the opposite of what the header says.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(time,lint): guard the recomputed tx_after, and judge one write at a time
Two review findings, both real.
setTransmitDelay() recomputes p->tx_after from a clamp of three candidates, and
that recomputation was still raw. Two lines above it, `if (p->tx_after)` is the read
that takes 0 as "no delay wanted", so a clamp landing on 0 drops the CSMA backoff
and the packet goes out immediately instead of after its computed delay. The first
arm site in this function was already guarded; this one was missed because the
value is not a plain `now + delay` and so does not fit timerEndsAtMillis() - it
takes skipZero() instead.
The narrowing order matters here and is spelled out at the site: add_delay is
unsigned long, 64-bit on the portduino host, so the clamp can exceed UINT32_MAX
there. skipZero() on the wide value would pass 0x100000000 through as non-zero and
the store to this uint32_t field would then truncate it back to the 0 being
avoided, so the cast comes first.
The lint rule judged each write by the wrong text. rhs was taken from the write to
the end of the accumulated statement, so a neighbour on the same line decided the
verdict - and it was wrong in both directions:
rebootAtMsec = millis() + 5; shutdownAtMsec = Time::timerEndsAtMillis(10);
the later helper call suppressed a genuine raw arm
rebootAtMsec = otherDeadline; shutdownAtMsec = millis();
the later millis() reported a safe copy
rhs is now cut at its own semicolon. Six fixtures cover it, including both cases
above, two raw writes on one line, two helper writes on one line, and a statement
split across lines, which must still see its whole right-hand side. 41 fixtures
total, green under bash 3.2.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* time: arm the remaining 0-means-unset stamps through the helpers
The follow-up sweep to the sixteen fields the previous commits covered. A field
here uses 0 to mean "unset" - some read spells `if (f)`, `f != 0`, `f == 0 ||` or
`f > 0`, or a site disarms it with `f = 0` - but it was armed from a raw clock read,
so once per ~49.7-day wrap it stores the value its own readers treat as never-set.
34 fields, 58 arm sites.
The rule could not have found most of them first. It treated `field = <variable>`
as inheriting whatever that variable did, which made the commonest shape in the tree
invisible: one `now = millis()` at the top of a runOnce(), then several
`xStartTime = now` below it. Listing those names would have bought no protection at
all, so the scanner now tracks a local assigned from a clock and treats a write from
it as the raw arm it is. One hop, one function, name-based, and it forgets a local
reassigned from anything else; taint is dropped at each function boundary. Twelve
fixtures pin it, including the negative cases - no leak across functions, `now` does
not match `nowMs`, and neither `==` nor `+=` records anything.
That pass immediately found a site the previous commits missed: setTransmitDelay()
recomputes p->tx_after from a tainted `now`, two lines under the `if (p->tx_after)`
read that takes 0 as "no delay wanted".
Three of the fields are worth naming because the consequence is not cosmetic:
* UpDownInterruptBase press/up/downStartTime - xDetected is only cleared INSIDE
the block guarded by `xDetected && xStartTime > 0`, so a stored 0 makes both the
entry and the exit condition unreachable and that button is dead for the rest of
the boot, not for one tick.
* PhoneAPI lastContactMsec - ServerAPI reads `lastContactMsec > 0` before the TCP
idle close, and the field stays 0 until the next inbound packet, so a client that
never speaks again leaks the socket for the life of the connection.
* EInkDisplay lastDrawMsec - `if (lastDrawMsec)` gates every plain display() call
on a keyframe having been shown, so a stored 0 stops the screen updating until
something calls Screen::forceDisplay() again.
TransmitHistory needed more than its arm sites. getLastSentToMeshMillis() returns 0
to mean "module has never sent", and besides the two stores, both reconstruction
helpers end in `millis() - msAgo`, which can produce a 0 of their own. All three
computed returns are guarded; the deliberate `return 0;` sentinels are untouched.
Judged and deliberately not changed:
* nRF54L15 connect_time_ms is armed from k_uptime_get_32(), not millis(). It is
guarded with skipZero() but keeps its own clock - swapping in Time::getMillis()
would have it compared against a k_uptime now at the watchdog read. The rule now
recognises that clock too, so listing the field is not an empty gesture.
* RotaryEncoderInterruptBase pressStartTime shares a name with the UpDown field and
has a different contract: no read here tests the stamp against 0, pressDetected
is the only armed flag. Opted out at the write. Its lastPressLongEventTime
sibling IS a `== 0` latch and is fixed.
* PositionModule line 38 copies a value the enclosing `if (restored != 0)` has
already proven non-zero. Opted out.
* pmMeasureStarted, adminKeyFallbackRefillMs, the two autosave stamps and
scrollStartDelay are lazy initialisations whose wrap behaviour costs at most one
interval and drops nothing. Left alone, and not listed.
47 lint fixtures green, the rule silent across src/, full native suite 1421/1421.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* time: dodge the wrap at the clock read, not only at the store
Review on #11830 made a point that was right and that this branch had wrong.
Applying skipZero() at the STORE while a reader measures elapsed time against a
raw clock splits the two sides apart for one tick per ~49.7-day wrap: the stamp
becomes 1 while `now` is still 0, so `now - stamp` is UINT32_MAX and a brand new
stamp reads as about 49.7 days old. Every elapsed-since guard then fires when it
must not. Concretely, UpDownInterruptBase computed `now - pressStartTime` and
emitted a long press for a fresh press, and TraceRouteModule read
`now - lastTraceRouteTime < cooldownMs` as false and bypassed its cooldown.
So the dodge moves to the read. Time::stampMillis() is getMillis() with the one 0
tick called 1; a site that both stores a stamp and measures against stamps reads
the clock once through it and stores that value directly. Nine files, and the 1 ms
skew is the same one skipZero() already documents.
Where the clock arrives as a PARAMETER the store keeps its own skipZero() as well,
because the function cannot assume the caller dodged anything. Removing that was a
real regression and test_nexthop_routing caught it: noteRouteLearned() and
noteRouteSuccess() are called with a literal 0 by
test_health_learn_never_stores_zero_sentinel and
test_health_success_never_stores_zero_sentinel, which assert the store normalises
it - 0 is the empty-slot marker getOrAllocRouteHealth() evicts on. The two guards
compose without shifting twice, since skipZero() of a non-zero value is itself.
trySmartBroadcast() and directResponseAllowed() have the same parameter shape and
keep their store-side guard for the same reason. Only stores fed by a stampMillis()
local in the same function are bare.
EInkParallelDisplay was missed the first time: the third class in the family, still
storing skipZero(getMillis()) while rate-limiting against a raw millis() local.
Normalised like its siblings.
The lint rule gained three false positives with the class-scope tracking, all of
them shapes that are not class bodies at all:
template <class T> void f(T x) { uint32_t lastSort = millis(); }
class Foo { void tick() { uint32_t lastSort = millis(); } };
void g(struct Bar *b) { uint32_t lastSort = millis(); }
Two causes. pending_class matched class/struct anywhere on the line, so a template
parameter list and an elaborated type in a parameter list both marked the following
FUNCTION body as class scope; it is anchored to the start of the line now. And
update_scope() runs at the end of a line, so a body opened earlier on the same line
had not been counted when the statement was judged; is_declaration() now also
counts unmatched braces earlier in the statement. The rule is blocking and
`template <class T>` is ordinary C++, so these would have reddened files nobody
touched.
note_taint() also never received the per-write `;` cut the judging path was given
earlier in review, so on a line holding two statements it learned taint from the
neighbour. Same cut applied.
Five tests in test_uptime_clock pin the contract, including one that asserts the
old store-only shape really does produce UINT32_MAX, and one that pins the 1 ms
skew at 399 rather than 400 so nobody "corrects" it back into a raw read. Lint
fixtures 53 -> 65. Full native suite 1426/1426, rule silent across src/.
Known residual, deliberately not changed: a store that dodges zero while its reader
measures through a Throttle:: helper still splits for that one tick, because those
helpers read the clock internally and raw. About eight sites tree-wide, including
PositionModule trySmartBroadcast and the lastContactMsec TCP idle check. Closing it
means making Throttle read through the dodge, which was proposed on #11692 and
declined there pending a caller audit, so fixing one site here would only make the
tree inconsistent. The direction is also the same one the un-dodged code already
took: a fresh stamp reads as old, and the guards involved were already passing on a
0 stamp.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* lint: a dodged value is safe to copy, not to do arithmetic on
Two more review findings on the rule, both real, both false negatives.
Arithmetic on an already-dodged value was excused. stampMillis() guarantees only
its own result, so `now + 5000` can carry a non-zero stamp straight back onto the
sentinel - 0xFFFFEC78 + 5000 is exactly 0. That sum is precisely what
Time::timerEndsAtMillis() exists to dodge, and the rule was waving it through
because a helper name appeared somewhere in the expression. Worse, a fixture
asserted that behaviour was correct, so the self-test was pinning the hole open.
A local holding a dodged value is now tracked separately from a tainted one: it may
be stored or copied straight through, but + or - applied at the OUTERMOST level is
reported and the message points at timerEndsAtMillis(). Depth-aware, so the operator
inside Time::skipZero(getMillis() - msAgo) is still fine, and so is the
`(d == 0) ? 0 : timerEndsAtMillis(d)` arming form, which has no top-level operator
at all. The wrong fixture is replaced by four: store-through, copy one more hop,
arithmetic on a dodged local, and arithmetic on a direct helper call.
A class body that opens and closes on one line was never recognised. The header
check rejected it because the line ends in a semicolon, which a one-liner body
always does, and even once armed the class brace counted as a function body and
excused the member. Both halves fixed: the header arms on the brace rather than on
the absence of a semicolon, and when the body opened on the statement being judged,
one unmatched brace is class scope while two is a method body inside it. Getting
that wrong first broke every multi-line class, because setting the per-statement
flag without also arming pending_class meant update_scope() never registered the
body - the three existing class fixtures caught it.
72 fixtures, green under bash 3.2, shellcheck clean, rule silent across src/.
No src/ or test/ file changes, so the native suite is untouched by this commit.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* style(time): trim comments to the house limit
---------
Co-authored-by: Tom <116762865+Nestpebble@users.noreply.github.com>
Co-authored-by: nomdetom <nomdetom@protonmail.com>
Co-authored-by: Tom <116762865+NomDeTom@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
be449b525f |
fix(esp32): rebuild IDF libs when the HybridCompile cache is stale (#11834)
The platform decides whether framework-arduinoespressif32-libs matches the current env from a hash written into sdkconfig.defaults at the start of the IDF-libs pass, before those libs are compiled, so an interrupted or metadata-only pass leaves a hash describing libs that were never built. Later builds match that hash, skip the recompile and link the previously compiled board's IDF configuration; on tlora-t3s3-v1 this linked a t-connect-pro build's CONFIG_SPIRAM_MODE_OCT libs into a quad-PSRAM ESP32-S3FH4R2, which aborts in esp_psram_init() before the console exists and boot loops with no serial output. Drop sdkconfig.defaults when the package's own <mcu>/sdkconfig, rewritten as the last step of a completed compile, predates it. |
||
|
|
31f05ab057 |
fix(touch): stop LONG_PRESS repeating when the suppression deadline wraps (#11829)
* fix(touch): stop LONG_PRESS repeating when the suppression deadline wraps TouchScreenBase::_start was one field doing two incompatible jobs. It held the press-down timestamp, and then the LONG_PRESS handler overwrote it with `millis() + 30000` to stop the event repeating for the rest of the hold. Every read was a hand-rolled signed subtraction on time_t, and suppression worked only because `time_t(millis()) - _start` came out around -30000. Where time_t is 64 bits - the portduino host - that uint32_t sum wraps to a small number while millis() is still just under 0xFFFFFFFF. The subtraction then goes hugely positive instead of negative, the threshold test passes on every 20ms poll, and each pass re-arms to another wrapped value. It keeps firing until millis() itself wraps, up to ~30 s later: about 1500 TOUCH_ACTION_LONG_PRESS events injected into InputBroker for one finger that never moved. Modelling the old expression across press-start offsets puts the worst case at exactly 1500 for a 60 s hold, where three is correct. On a 32-bit time_t build the signed wrap happens to keep suppressing, so this is host-and-variant dependent rather than universal. The zero-dodging helpers in src/UptimeClock.h are no use here: they map 0 to 1, and 1 reads as "long ago" exactly as 0 does. The defect is the overload, not the zero, so the field is split by what it is actually asked: _pressStartMs a past event time - how long has the finger been down _longPressSuppressed is repeat suppression armed _longPressSuppressUntilMs when it expires, read only while the bool is set Two fields for the suppression rather than one, for the reason Throttle.h's TODO(deadline-type) gives: armed has to stay a separate question from passed. No single value can stand in for "unarmed" here either, since deadlinePassed() reads 0 as long past below ~24.8 days of uptime and as far future above it. Nothing new uses 0 as a sentinel, so bin/lint-unset-sentinel-millis.sh needs no entry. All three comparisons now go through Throttle - hasElapsed() for the two elapsed-since-press questions, which also buys the full ~49.7 day range that a stored event time gets, and deadlinePassed() for the suppression window. Behaviour is preserved deliberately, including the part that is easy to miss: the old `+ 30000` made a held finger re-report LONG_PRESS once every 30 s, not once per touch. A bool latch would have been simpler and quietly narrowed that, so the window is kept as LONG_PRESS_REPEAT_SUPPRESS_MS. Old and new were compared across five wrap scenarios and agree everywhere except the wrap window the old code got wrong. The tap-on-release suppression the old write also provided is not needed: a hold long enough to reach here has duration >= TIME_LONG_PRESS, so the tap branch already takes its else and clears _tapped. One guard added while here. The RAK14014 deferred-tap window is TIME_LONG_PRESS - 50 and that subtraction is unsigned now, so a variant lowering TIME_LONG_PRESS below 50 would underflow it into a ~49.7 day wait and the deferred TAP would never fire. The only override in the tree is t5s3_epaper at 500; a static_assert fails the build instead of the touch panel. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * style(touch): trim comments to the house limit --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: nomdetom <nomdetom@protonmail.com> |
||
|
|
bef289ef42 |
fix(extnotif): make isNagging the only armed flag for the nag cycle (#11828)
* fix(extnotif): make isNagging the only armed flag for the nag cycle
ExternalNotificationModule kept the nag cycle's armed state in two places that
could disagree: the isNagging bool, and nagCycleCutoff reserving UINT32_MAX for
"not armed". handleInputEvent() read only the second one:
if (nagCycleCutoff != UINT32_MAX) { stopNow(); return 1; }
The field is declared `= 1`, while isNagging starts false, so at boot that test
said "armed" when nothing was nagging. The first input event of every boot was
therefore answered with stopNow() and a non-zero return - and a non-zero return
ends the observer chain (Observable::notifyObservers in src/Observer.h returns on
the first one), so that event was swallowed from every later observer. The handler
is registered whenever external_notification.enabled, and InputBroker only
short-circuits while nagging() is true, so the event does reach it.
The same read had a second failure mode once per ~49.7-day wrap: armNagCycle()
computes `millis() + durationMs`, which can land exactly on UINT32_MAX. When it
does, a real nag is running with isNagging true, but this read says "not armed" and
the module's own handler never stops it. Time::skipZero() cannot help here - it
lifts 0 to 1 and leaves UINT32_MAX alone, which src/UptimeClock.h static_asserts.
So the fix is not a zero guard, it is removing the second opinion. isNagging is
the armed flag - which is what the comment above the expiry check already claimed,
and what the other four reads already use - and nagCycleCutoff is now only ever a
deadline, read after isNagging has been checked. Nothing reserves a value, which
matters because an arm site spelled `millis() + interval` can produce any value
there is, so no value is safe to reserve. That is the shape the TODO(deadline-type)
note in src/mesh/Throttle.h is aiming at, and that note is updated to match rather
than keep describing the sentinel this removes.
Worth knowing for review, though not changed here: InputBroker::handleInputEvent
already calls stopNow() itself when nagging() is true, and returns without
notifying observers. Every path that starts a notification calls armNagCycle()
first, so isNagging is true for the whole life of any real nag. That makes this
handler reachable only when there is nothing to stop - its stopNow() was never
doing useful work. Gated rather than deleted, because removing a public handler
and its observer registration is a bigger call than fixing the defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* style(extnotif): trim comments to the house limit
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: nomdetom <nomdetom@protonmail.com>
|
||
|
|
644a43ca9b |
Update protobufs (#11841)
Co-authored-by: jp-bennett <5630967+jp-bennett@users.noreply.github.com> |