mirror of
https://github.com/alexhopeoconnor/firmware.git
synced 2026-10-04 03:18:10 +10:00
ee48094ea8309858cd579100b8aaf14560bcc4b2
100
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
bfd1e1a231 |
Add Heltec RC32, RC52 and RCC6 boards, and LC760CA GNSS support (#11572)
* refactor(graphics): select Arduino_GFX panels with a capability flag TFTDisplay tested `defined(HACKADAY_COMMUNICATOR)` in a dozen places to mean "this panel is driven by Arduino_GFX rather than LovyanGFX". Every new Arduino_GFX board had to be appended to all of them. Move the decision into the variant as USE_ARDUINO_GFX so the display code stops naming individual boards. No behaviour change: the Hackaday Communicator is still the only board that sets it. * feat(boards): add Heltec RC32, RC52 and RCC6 Three boards around the same 128x220 NV3001B panel: RC32 (ESP32-S3), RCC6 (ESP32-C6) and RC52 (nRF52840). They differ only in how the panel bus is wired, so they share one branch in TFTDisplay behind TFT_NV3001B. RC32 and RC52 also carry a rotary encoder on a TCA6408 I2C expander. That lands as its own input source rather than as board conditionals inside i2cButton, which is the M5Stack UnitC6L button driver and stays untouched. On RC52 and RCC6 the panel is an add-on module, so probe it before reporting a screen. The probe reuses the bit-banged SPI helper that already backs the T114 ST7789 check. Arduino_GFX is pinned to the upstream commit that added the NV3001B driver; it has not shipped in a tagged release yet. Co-Authored-By: Quency-D <55523105+Quency-D@users.noreply.github.com> * feat(gps): detect and configure the LC760CA GNSS module The LC760CA is another Unicore part, so it joins the $PDTINFO probe family and reuses the CM121 message-rate setup. It answers with CC1161W. GNSS_MODEL_LC760CA goes immediately before GNSS_MODEL_GENERIC_NMEA: the sentinel has to stay last because isValidGnssModel() uses it as the exclusive upper bound on values the probe cache may hold. Placing the new model after it would leave LC760CA permanently uncacheable. Co-Authored-By: Quency-D <55523105+Quency-D@users.noreply.github.com> * fix(graphics): re-init the NV3001B after the panel rail comes back DISPLAYOFF de-asserts VTFT_CTRL, which cuts power to the panel, so the controller loses MADCTL, COLMOD and gamma. displayOn() only sends sleep-out and cannot restore them, leaving the panel dark or in the wrong format after wake. Re-run begin() once the rail has settled, and repaint in full since the re-init leaves display RAM undefined. Also stop the TCA6408 rotary polling from two threads at once. Registering as an InputPollable meant InputBroker's pollSoon task could call pollOnce() while runOnce() was mid-transfer on the main thread, with nothing serialising Wire or the decoder state. Drop InputPollable and have the interrupt wake the thread instead, the way ButtonThread does, so the bus and the decode stay on one thread. * fix(graphics): skip the NV3001B wake when re-init fails begin() reports whether the bus came up. Ignoring it meant a failed re-init still lit the backlight and drove a full-screen repaint at a panel that was never initialised. * chore(boards): ship the Heltec RC boards at release level release is the normal level for a variant; the matrix generator still builds each of these in this PR because they add a new platformio.ini. --------- Co-authored-by: Quency-D <55523105+Quency-D@users.noreply.github.com> |
||
|
|
73f7b35bea |
Report the right hardware model on four boards (#11570)
Four variants declare a custom_meshtastic_hw_model that the build never reaches, so the device announces something else in NodeInfo and the apps cannot match it for OTA. Mini ePaper S3 (125) and Heltec V4 R8 (132) had no arm in the esp32 HW_VENDOR chain at all, so both fell through to #else and reported PRIVATE_HW. Heltec Mesh Node T096 (127) had none in the nrf52 chain and reported NRF52_UNKNOWN. WisMesh Tap V2 defines both RAK3312 and RAK_WISMESH_TAP_V2, and the generic RAK3312 arm sat first, so the board reported RAK3312 (106) instead of WISMESH_TAP_V2 (116). Order the specific arm ahead of the generic one, the same way the nrf52 chain already keeps custom RAK4630 boards ahead of the generic RAK4630. Verified by preprocessing each platform's HW_VENDOR chain with the env's full define set - build flags resolved through extends, the board JSON's build.extra_flags, and the bare #defines in the variant's own variant.h. All four now match their manifest, and rak3312, heltec-v4, heltec-v4-tft and the ThinkNode M9 arm added in #11567 are unchanged. |
||
|
|
f6f116a39d |
Fill in device registry metadata for recently added hardware (#11567)
Audit of the custom_meshtastic_* manifest on the variants backing the newest boards, against the protobuf HardwareModel enum, the compiled HW_VENDOR, the board flash size and the artwork actually published by the web flasher. No support flag changes here - actively_supported is left exactly as each variant already had it. ThinkNode M9 had no HW_VENDOR arm, so every M9 has been reporting PRIVATE_HW while its manifest advertised 131; add the mapping and rename the slug to the enum name (THINKNODE_M9) it is meant to mirror. Seeed SenseCAP Mesh-Tracker X1 moves from the PR matrix to release, and its images entry now points at seeed_mesh_tracker_x1.svg, which is what the flasher actually ships - the hyphenated name resolved to nothing. T-Beam BPF, T-Beam 1W and Heltec Wireless Tracker V2 declared the architecture as "esp32s3"; the value is copied verbatim into the manifest, and the flash flow matches on the normalized "esp32-s3". T-Beam BPF and M5Stack Unit C6L both build default_16MB.csv on 16 MB flash but declared no partition scheme, which leaves the flasher on the 4 MB fallback offsets for a legacy clean install. Meshnology W10 and W12 gain the artwork and vendor tag that already exist for them. |
||
|
|
74119c088b |
fix(mesh): don't reference the position module on MESHTASTIC_EXCLUDE_GPS builds
The event-channel position-request reply added in #11545 calls positionModule-> replyOnPositionChannel() guarded only by USERPREFS_BLOCK_POSITION_ON_EVENT_CHANNEL. Targets that set MESHTASTIC_EXCLUDE_GPS (repeaters such as rak_wismesh_repeater_mini_hp) never construct PositionModule in Modules.cpp, so an event build for one of those fails to link: undefined reference to `PositionModule::replyOnPositionChannel(...)' undefined reference to `positionModule' Guard the call, the include and the isEventChannelPositionRequestForUs() helper with !MESHTASTIC_EXCLUDE_GPS, matching how AdminModule guards its positionModule use. A node with no position module has nothing to answer a position request with, so skipping the reply is the correct behavior there. Not reachable on develop, where the userpref defaults off and the whole block compiles out - it only breaks builds that enable it, which is why #11545 was green. Verified by building rak_wismesh_repeater_mini_hp with the pref enabled. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
9fcb289643 |
fix(thinknode_m9): define SPI_FREQUENCY for the non-MUI build (#11546)
The M9's variant.h defines ST7789_CS, so TFTDisplay.cpp compiles its ST7789 LGFX branch, which reads SPI_FREQUENCY for the panel write clock (SPI_READ_FREQUENCY, its pair, is already in variant.h). The flag was only set in the -tft env, so `build (thinknode_m9, esp32s3)` has failed on develop since the board landed in #10908: src/graphics/TFTDisplay.cpp:504:30: error: 'SPI_FREQUENCY' was not declared in this scope; did you mean 'SD_SPI_FREQUENCY'? Move the flag up into thinknode_m9_base, keeping the 75 MHz the -tft env already used for the same panel and matching the SD card's 75 MHz on the bus they share. The -tft env inherits the base flags, so device-ui's LGFX_GENERIC.h - which falls back to 20 MHz when the macro is absent - still sees the identical value. |
||
|
|
a5fc95f774 |
fix(mesh): coerce coordinate traffic to the position channel on event builds (#11545)
Under USERPREFS_BLOCK_POSITION_ON_EVENT_CHANNEL every coordinate packet a client aimed at the event channel was rejected with the "Location sharing is disabled on this channel" notification - including the phone's own location feed. Both apps hand a GPS-less node its fix as a POSITION_APP packet addressed to the node itself on channel 0; that packet never leaves the device (Router::sendLocal delivers it locally) but resolved to the event channel and was dropped before PositionModule saw it. Result: the toast on every location tick, and nodes without a GPS never learned a position to share on their private channel. Position traffic now converges on the position channel - findPositionChannel(), the first channel with non-zero on-wire precision, which is never the event channel: - From-us-to-us coordinate packets are exempt from the event block. - Local coordinate sends aimed at the event channel (phone share-location, request-position, waypoints, any module/UI originator) are moved onto the position channel in Router::sendLocal and PhoneAPI instead of rejected. The client notification is only sent when no channel carries positions at all. - A position request DM'd to us on the event channel is answered on the position channel at that channel's precision (request_id preserved, same reply throttle); the requester's coordinates are still not stored, forwarded, relayed or published. want_response from the bitfield is merged before the event-channel decode short-circuit so such requests are seen. - PositionModule::sendOurPosition, positionUnchangedSinceLastSend and MeshService::trySendPosition use the shared helper instead of three copies of the same walk. Non-event builds are unaffected: the coercion compiles out and the helper matches the previous walk. Tests: coverage-event-policy (test_event_channel_phone_api, test_event_channel_router, test_position_precision, test_mqtt, test_nexthop_routing) and the same suites with the policy off. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
48699a7a48 |
fix(http): keep reaping open TLS connections under low heap so the heap can recover (#11539)
* fix(http): keep reaping open TLS connections under low heap so the heap can recover Once free heap dropped below MIN_HEAP_FOR_SSL (40 KB) with HTTPS connections open, the node's heap never came back and every later HTTPS or TCP-API connection failed until a reset - node alive, on WiFi, unusable. handleWebResponse() skipped secureServer->loop() entirely under low heap so no new TLS handshake would be attempted on a heap that can't hold its context. But HTTPServer::loop() is the only place already-accepted connections are serviced and reaped: its first pass calls ->loop() on each open one (where the 20 s idle timeout and the SSL close-notify state machine run) and deletes the closed ones. Skipping the whole loop froze the up-to-MAX_HTTPS_CONNECTIONS TLS sessions already open. Never looped, they never timed out, their mbedTLS contexts and pbufs were never freed, so free heap never climbed back over 40 KB, so the loop was skipped forever. The guard's own precondition was what kept it from clearing. Split the two halves. Under low heap keep driving and reaping the connections we already hold, and only skip the accept. HTTPServer keeps its connection table protected, so a thin MeshHTTPSServer subclass exposes serviceExistingConnections(), the first half of HTTPServer::loop() verbatim. Log line reworded to say what now happens: not accepting, not skipping. Verified on a Heltec V3 (Endor AP) against a control build with #11537 (so the node survives the squeeze instead of aborting first): - Recipe: held sockets on 80/4403 + pending TLS, 100 s of HTTPS pokes, repeat. Control: Low heap pins at 6-17 KB, HTTPS dead, and 3 min after all pressure is released heap is still ~12 KB with Low heap firing every 30 s - permanent until reset. Fix: never dips under 40 KB, both pressure rounds 3/3, 65 KB after. - Branch driven deliberately (verify-only heap hog pinning free heap at ~28 KB with a real idle TLS session held open): under the guard the fix logs open=1 -> reaped=1 at the 20 s idle timeout, and heap goes 26 -> 65 KB before the hog is even released. On the control logic that session stays frozen for the whole window. Fixes #11538. * fix(http): trim the low-heap comments to the two-line guideline The mechanism is in the commit message and PR; the source keeps the one-line why. No code change. (CodeRabbit) |
||
|
|
fe15786dc1 |
fix(api): stop rebooting ESP32 nodes when a client connects to a fragmented heap (#11537)
* fix(api): stop rebooting ESP32 nodes when a client connects to a fragmented heap
Connecting a client to an ESP32 node over WiFi/TCP rebooted the node. Two
allocations on the accept + config path use operator new, and on ESP32 that is
fatal when it fails: the framework builds with CONFIG_COMPILER_CXX_EXCEPTIONS=n
(esp32-common.ini), and ESP-IDF's cxx component then --wraps __cxa_throw and
every unwinder entry point straight to abort(). libstdc++'s operator new throws
std::bad_alloc on a NULL from malloc, so any new that cannot get its block is a
reboot with no chance to recover. Both hit on a Meshnology W12 running develop
|
||
|
|
83fd62b756 |
test(native): add 14 suites for routing, persistence, parsing and identity gaps (#11515)
* test(native): add 14 suites for routing, persistence, parsing and identity gaps Coverage audit of the native test tree; adds the highest-value untested logic as 11 new suites and extends 3 existing ones (200 test functions). New: test_stream_framing, test_nodedb_boot_recovery, test_nodedb_legacy_migration, test_nodedb_v25_roundtrip, test_nodedb_identity_hygiene, test_channel_keys, test_reliable_ack_matrix, test_hop_start_policy, test_routing_response_hops, test_phone_api_config_dump, test_observer. Extended: test_rtc, test_mqtt, test_xmodem. Two source changes the audit produced: - StreamAPI::handleRecStream copied stream->read()'s `cInt < 0` EOF check into the buffer-fed path, where there is no EOF sentinel; with signed char any byte >= 0x80 (START1 is 0x94) aborted the parse. Read the byte as uint8_t directly. Latent on develop (no callers), pinned by test_stream_framing. - Extract the post-decode pre-hop predicate from Router::handleReceived into shouldSkipHandleForPostDecodeHop() (NodeDB.h) so test_hop_start_policy drives the exact expression the router calls. No behavior change. test/state-manifest.tsv declares the suites that construct a NodeDB. Full 68-suite Docker coverage run matches the pre-change baseline. * test(native): address review - harden observer dispatch, trim comments Review follow-ups on the coverage-audit suites: - Observable::notifyObservers() erased list nodes while holding an iterator into them, so an observer that unobserves itself from onNotify corrupted the dispatch. Today the only self-detacher (PhoneAPI::onNotify -> checkConnectionTimeout -> close -> unobserve) survives solely because it returns -1 and aborts the chain before the increment; that unwritten contract is now gone. Removal during a dispatch nulls the entry and the outermost notify sweeps afterwards, which keeps self-detach, next-detach and destruction-during-notify all safe without an allocation. Hoisting the next iterator instead would have inverted the hazard and broken the existing next-detach case. Two regression tests added. - Correct the documented caller of shouldSkipHandleForPostDecodeHop: the call is in Router::dispatchReceived, not handleReceived. - Cast hop fields to unsigned at the %u call site in test_hop_start_policy. - Trim the new suites' file headers to the one-or-two-line rule in AGENTS.md. - Rename eight test functions whose names were exactly `test_` + 35 chars: that is the shape of a Lob API key, so trufflehog flagged them as secrets and failed the Trunk CI check. Full 68-suite Docker coverage run matches the pre-change baseline. * test(native): revert the observer dispatch change, keep the contract test Backs out the notifyObservers() deferred-removal hardening from the previous commit. It was reviewer-driven scope creep: nothing in the coverage audit needed it, no test required it, and it changes dispatch semantics in a header with ~76 observe() call sites on native verification alone. The hazard it addressed is not reachable today. The only observer that unobserves itself from onNotify is PhoneAPI (onNotify -> checkConnectionTimeout -> close -> unobserve), and it returns -1, which aborts the chain before the iterator is advanced past the erased node. test_self_detach_with_abort_during_notify stays: it passes against the unmodified dispatch and pins that the -1 is load-bearing, so a later cleanup that "simplifies" it away goes red. The unsafe variant (self-detach returning 0) is documented in a comment rather than tested, since asserting it would be asserting UB. * fix(serial): recover the frame behind a stray framing marker A byte that failed the START2 check was discarded rather than re-tested as a possible START1, so 0x94 0x94 0xc3 ... lost the real frame: one corrupted byte on a noisy UART silently dropped the frame behind it. Re-test the byte in place instead. Applied to both copies of the receive state machine. readStream() is the one that matters in the field - it is the serial path every phone client uses - while handleRecStream() still has no callers on develop. Strictly widens what the parser accepts; no frame that parsed before parses differently. test_stream_framing covers it on both receive paths, plus a run of stray markers and a START1-then-unrelated-byte resync. This was originally documented as a known gap in the framing suite. Fixing it instead was NomDeTom's call on review: a passing test asserting the bad behavior is what makes it hard to change later, and it is the same defect shape as the signedness fix three functions away. Also: use Throttle::deadlinePassed() in test_reliable_ack_matrix rather than a bare millis() compare, matching the house deadline rule. * test(native): cover the stray-marker resync on the buffer path too The stray-marker fix went into both copies of the receive state machine, but only test_stray_start1_before_frame_still_delivers drove both. The repeated- marker and unrelated-byte cases drove readStream() alone, so a regression in handleRecStream() would have gone unnoticed by two of the three. Verified load-bearing: reverting only the handleRecStream() half of the fix turns test_repeated_stray_start1_before_frame_still_delivers red on the new assertion. test_start1_then_unrelated_byte_resyncs stays green under that mutation by design - its failing byte is 0x00, where both branches reset to 0 - and covers the other half of the ternary. Also drops the stale header on test_stray_start1_before_frame_still_delivers, which still described the gap as pinned-as-is after the fix landed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(native): make the hop-start truth table assert the rows it prints test_truth_table_summary was six TEST_MESSAGE lines and no assertion, so it reported as a case that could not fail - the anti-pattern #11517 names in its unfinished assertion-presence lint, and the one exception to NomDeTom's "no RUN_TEST without an assertion" pass over this PR. The printed row and the checked expectation now come from one struct, so the summary cannot narrate a table the predicates no longer implement. It also covers the consequence columns the per-row tests do not assert together: classifyHopStart, shouldDropPacketForPreHop and shouldSkipHandleForPostDecodeHop for the same packet, with the expectations gated on MESHTASTIC_PREHOP_DROP. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e1ea653a45 |
fix(graphics): crash and leak fixes across display drivers (#11455)
* fix(graphics): crash and leak fixes across display drivers - TFTDisplay (portduino): _touch_instance was an uninitialized member, and the touch-config block only assigns it for xpt2046/stmpe610/ ft5x06 while the guard accepts any configured module. A gt911 entry in config.yaml (supported by the color-UI path) reached _touch_instance->config() through an indeterminate pointer. Initialize to nullptr and guard the config block. - Screen: the destructor freed normalFrames but leaked the owned dispdev (driver + framebuffer) and ui objects. Screen is genuinely destroyed on the portduino reboot path (screen = nullptr in Power.cpp). - EInkDisplay2: GxEPD2_BW's constructor takes the low-level driver by value and stores a copy, so the 'new EINK_DISPLAY_MODEL' at nine sites was orphaned the moment connect() returned. Pass temporaries, as GxEPD2Multi already does. - EInkParallelDisplay: the async full-refresh task cleared asyncFullRunning before nulling asyncTaskHandle, so the destructor could observe running==false with a stale handle and vTaskDelete a freed TCB. Null the handle first (same ordering fix the eink/Drivers/EInkParallel.cpp sibling already carries). - Panel_sdl: initFrameBuffer only null-checked the first of its three allocations and returned true regardless, leaving the line array full of null+offset garbage on failure; later redraws would write through those. Check all three, release partial allocations, and return false. * fix(graphics): propagate Panel_sdl framebuffer allocation failure from init() Per review: initFrameBuffer() can now fail cleanly, so init() must not register the monitor and report success when it does. |
||
|
|
f57ee0bd71 |
fix(mesh): restore the implicit ACK for our own overheard PKI DMs (#11502)
* fix(mesh): restore the implicit ACK for our own overheard PKI DMs A DM we originate is PKI-encrypted to the recipient, so when we overhear it being rebroadcast we cannot decrypt it. perhapsHandleReceived() classifies it DECODE_OPAQUE and returns before shouldFilterReceived() runs, which is where the implicit ACK for our own transmission is generated. The client therefore never receives the ROUTING_APP ack it renders as "Delivered to mesh" for a DM, and the message sits in "sending" until it either succeeds outright or times out as max retransmissions. The ACK only needs the packet header (from/id), not the decoded payload, so split it out of shouldFilterReceived() into perhapsGenerateImplicitAckForOwnOverheard() and also call it from the opaque short-circuit for packets that are from us. Behavior on the decodable path is unchanged. Broadcasts on a PSK channel decode normally and always reached the generator, which is why channel messages were unaffected and only DMs showed the symptom. * test: rename implicit-ack tests to avoid a trufflehog false positive The camelCase identifiers tripped trunk's trufflehog/Lob secret detector. * test: shorten one test name past trunk's Lob secret-detector pattern trufflehog's Lob rule matches test_ followed by exactly 35 word characters, which both new test names happened to hit. Unrelated to the fix. |
||
|
|
905482ccce |
fix(serial): don't sleep forever with pending PhoneAPI output on UART consoles (#11500)
* fix(serial): don't sleep forever with pending PhoneAPI output on UART consoles Since #11164 bounded the stream drain, a config dump can end a dispatch with output still queued. On UART-console ESP32 boards runOnce() then returns INT32_MAX with no RX pending, and neither rxInt() nor onNowHasData() fires for the remaining output, so the download wedges mid nodeinfo stream until the client happens to send a byte. Add StreamAPI::hasPendingOutput() (transport-retained frame or queued PhoneAPI data) and have SerialConsole::runOnce() short-poll (<=25ms) while it holds instead of sleeping INT32_MAX. The #11164 write budget is unchanged; idle sleep behavior with a drained queue is unchanged. The retained-frame probe also covers the ESP32-S2 USB-CDC branch, which takes the same INT32_MAX path. * test(serial): restore scratch NodeDB via tearDown, trim comments to house style A failed TEST_ASSERT longjmps out of a Unity test without running destructors, so RAII cannot restore the swapped nodeDB pointer; install the scratch NodeDB explicitly and restore/delete it in tearDown(), which runs after every test outcome. Also shorten the new comments to the two-line house limit. |
||
|
|
fb6a212b44 |
fix(graphics): make on-screen keyboard lifecycle safe and RAII-managed (#11460)
- VirtualKeyboard::handleLongPress VK_ESC invoked the onTextEntered member std::function directly, but that callback path reaches OnScreenKeyboardModule::stop(), which destroys the keyboard - and with it the std::function whose invocation is still on the stack. handlePress and submitText already deliberately copy-and-clear before invoking for exactly this reason (CannedMessageModule documents the same hazard); do the same here. - OnScreenKeyboardModule's keyboard becomes unique_ptr, replacing the delete-in-destructor / delete-then-new-in-start / delete-in-stop bookkeeping that runs on every keyboard open/close. The NotificationRenderer legacy hook receives a non-owning raw pointer, as before. |
||
|
|
0ff10318ad |
refactor(net): unique_ptr for connection-lifecycle objects (#11459)
- WiFiServerAPI/ethServerAPI apiPort and ethApiServer's listener are create/destroy cycles that repeat across WiFi teardown and W5500 chip resets; the manual delete+null bookkeeping becomes reset(). (ethTlsApiServer's listener is left for a follow-up: that file is already touched by the partial-init fix PR and converting it here would conflict.) - ContentHandler::handleFormUpload held its body parser raw with delete on four separate exit paths of a per-request handler; any future early return was a silent leak. unique_ptr removes all four. - The portduino ch341Hal global becomes unique_ptr. The LoRa-error recovery loop's delete/null/new sequence was correct only by hand-preserved ordering; it becomes reset()/make_unique. RadioLibHAL keeps a non-owning raw pointer, as before. No behavior change. |
||
|
|
5e54262fe1 |
refactor(io): unique_ptr ownership for motion sensors and I2C keyboard (#11458)
* refactor(io): unique_ptr ownership for motion sensors and I2C keyboard - AccelerometerThread / MagnetometerThread: the owned MotionSensor becomes unique_ptr, removing the manual delete/null bookkeeping in clean(). Deletion behavior is unchanged (MotionSensor's destructor is virtual). - KbI2cBase: the TCA keyboard was a reference member bound to an anonymous heap allocation - ownership was invisible and nothing could ever free it. It becomes unique_ptr with an out-of-line destructor (the base type is only forward-declared in the header). - GeoCoord::pointAtDistance returned shared_ptr with no shared ownership anywhere (and no callers); return by value instead. No behavior change. * fix(io): make TCA8418KeyboardBase destructor public for unique_ptr ownership * refactor(gps): delete dead pointAtDistance instead of converting it Per review: zero callers in this repo or device-ui, and the math was wrong at both ends (rangeMetersToRadians multiplies meters by 1852, treating meters as nautical miles). Remove it, its now-unused helper, and the <memory> include the old shared_ptr signature pulled in. |
||
|
|
b565a07a83 |
Remove proprietary Bosch BSEC blob; open in-tree IAQ estimator for BME680 (#11381)
* Remove proprietary Bosch BSEC blob; open in-tree IAQ estimator for BME680 BSEC2 cost ~37-39 KB flash and ~4-5 KB static RAM on ~190 of ~240 build targets, linked whether or not a BME680 was attached, and was a no-source proprietary archive inside GPLv3 release binaries. The firmware consumed exactly one BSEC-exclusive output: the IAQ value. - New BME680IaqEstimator: clean-room log-domain baseline tracker (humidity-compensated gas resistance vs a rise-fast/decay-slow ceiling, 0-500 scale matching the existing UI bands), pure math, unit-tested on native (test_bme680_iaq, 15 tests incl. a deep-sleep reboot simulation). Warm-up/burn-in progress persists to /prefs/bme680.dat via SafeFile so one-sample-per-wake SENSOR nodes converge across reboots; stale /prefs/bsec.dat is removed once. - BME680Sensor: single-path rewrite on Adafruit_BME680 with async once-per-minute sampling (~20x lower heater duty than BSEC LP mode), a hard 2-minute publish-freshness bound (a dead sensor stops reporting instead of freezing its last reading on the wire), and suppression of bogus gas_resistance=0 points from heater-unstable cycles. - platformio.ini: environmental_extra_common/_extra/_no_bsec collapsed into one section; Bosch BSEC2 + BME68x deps deleted; per-variant BSEC link-path hacks and the TEMPORARY promicro lib_ignore removed. nrf52_promicro_diy_tcxo regains BME680 support at 36 KB clear of the warm-store cap; rak4631 lands at 75 KB clear. - EnvironmentTelemetry: iaq rendering gates on has_iaq (a genuine IAQ of 0 now displays); stale BSEC comments rewritten. - rak4631 size budgets tightened (113000->108000 RAM, 786000->746000 flash) to lock in the reclaimed headroom. - bin/bme680_iaq_replay.cpp: host-side replay harness for tuning the estimator against captured BSEC traces (mean abs error + band agreement), no reflashing needed. Measured (develop -> this branch): rak4631 -38.8 KB flash / -4.9 KB RAM; heltec-v3 -36.4 KB / -4.0 KB; tlora-v2-1-1_6 +1.3 KB (its IAQ approximation had been dead code since #9663 due to an inverted isfinite check and now actually runs). Note: gas_resistance stays kOhm on the wire for fleet compatibility; the proto comment claiming MOhm gets a separate meshtastic/protobufs docs PR. * Address CodeRabbit review feedback - Use Throttle::isWithinTimespanMs for all elapsed-time predicates in BME680Sensor per coding guidelines (deadline math for the async reading completion stays raw, as it targets an absolute timestamp) - Make the state file name members static constexpr - Replay tool: cast uint16_t before %u (default argument promotion), report malformed input lines instead of silently skipping, and fail non-zero on stream read errors * Address CodeRabbit nitpicks - Replace the local clampf helper with std::clamp (meshUtils.h's clamp drags in Arduino.h, which would break the estimator's standalone host build that the replay harness depends on) - Trim the replay tool's file header to a two-line summary; the full build, capture, and tuning workflow moves to docs/bme680_iaq_replay.md |
||
|
|
230da77642 |
fix(time): convert the millis() rollover sites #11291's CI guard cannot see (#11483)
* MeshPacketQueue: fix millis() rollover in the late-packet drop test replaceLowerPriorityPacket() read `backPacket->tx_after < now`, with `now` taken from millis() on the line above. tx_after is an absolute deadline, so that comparison inverts while the deadline sits on the far side of the 32-bit wrap: a queued late packet reads as not-yet-due for the rest of the wrap window, or every late packet reads as droppable at once. The same statement ordered two deadlines against each other with `backPacket->tx_after > p->tx_after`, which has the same problem. #11291 swept every site where millis() sits next to the comparison operator, and its CI guard matches that shape. Stashing the clock in a local first is the same bug written so the guard cannot see it. Both tests now subtract before comparing: the due test through Throttle::deadlinePassedAt(), and the ordering through the elapsed-since-now form already used in AdminModule's oldest-slot scan. The snapshot comes from Time::getMillis() so the deadlines and the test read one clock, per the convention deadlinePassedAt() documents. The `dt` the log line reports is now derived from the same elapsed value rather than recomputed. Behaviour is otherwise unchanged, save the boundary: deadlinePassedAt() is inclusive, so a deadline landing exactly on `now` reads as due rather than one millisecond early. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * RadioLibInterface: don't widen a uint32_t deadline delta into a 64-bit long TRANSMIT_DELAY_COMPLETED tested whether the front packet was still waiting with long delay_remaining = txp->tx_after ? txp->tx_after - millis() : 0; if (delay_remaining > 0) ... The subtraction is uint32_t. Where long is 32-bit - every embedded target - an already-due deadline lands negative and the packet transmits, which is why this has never been visible on device. Where long is 64-bit (portduino, and the native test build) the same value zero-extends to ~4.29e9, reads as positive, and the packet is rescheduled 49.7 days out. It stays parked until some later notifyLater() with overwrite happens to reset the timer. That is not an edge case. notifyLater() schedules through setIntervalFromNow(), so the thread wakes at or after the deadline; being a millisecond past due is the ordinary path through this branch. Ask Throttle instead. deadlinePassedAt() is the unsigned half-range test, so there is no signed conversion to get wrong at any width, and the remaining delay handed to notifyLater() is computed from the same snapshot. On 32-bit the behaviour is identical, including at the boundary: a deadline equal to now transmitted before and still does. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * ExpressLRSFiveWay: convert the two remaining raw window checks to Throttle runOnce() dismissed the alert frame with `now > alertingSinceMs + 2000` and chose its poll rate with `now < keyDownStart + 20000`, both against a millis() snapshot in a local. Same rollover inversion as any other naive compare, and invisible to the millis-deadline-check guard because millis() is not adjacent to the operator. update() in the same file was already on Throttle. hasElapsed()/isWithinTimespanMs() with the stored event give the full ~49.7 day range and need no snapshot. Sentinels are unchanged in meaning: `alerting` is the armed flag for alertingSinceMs and is tested first, and keyDownStart == 0 reads as "recent" for the first 20s of uptime exactly as `now < 0 + 20000` did - a poll rate either way. The arm sites move to Time::getMillis() so the writes land on the clock Throttle reads, which also puts them within reach of Time::setTestMillis(). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * GPSUpdateScheduling: record whether a search is running, don't infer it elapsedSearchMs() answered "am I searching?" by ordering two raw millis() stamps: searchStartedMs > searchEndedMs. Whichever stamp lands on the far side of the 32-bit wrap reads as the larger one, so the answer inverts once per wrap cycle, in both directions: - a search that started before the wrap and ended after it keeps reading as "searching". elapsedSearchMs() then grows without bound and searchedTooLong() aborts a search that is not running. - a search that started after the wrap, following one that ended before it, reads as "idle". elapsedSearchMs() returns 0, so an unproductive search is never aborted and the receiver stays powered until it locks. Both self-heal at the next informSearching(), which bounds the damage to one GPS cycle - but the ordering test cannot be made wrap-correct, because the two stamps carry no information about which wrap they belong to. It does not need to be. Whether a search is in progress is a fact the three inform*() calls already have in hand; the ordering was only ever standing in for it. Add the flag and set it there. elapsedSearchMs() keeps its unsigned subtraction, which was always the correct part. The file's clock reads move to Time::getMillis() so the suite can drive them across the wrap. Behaviour-preserving in production - Time::getMillis() is millis() unless a test injects a clock. test_gps_update_scheduling/ gains seven cases: the idle/searching/ended states, elapsed exactness across the wrap, both inversion directions above, and reset(). The two wrap cases fail on the old predicate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * MessageStore: date boot-relative messages in uptime seconds A message received before the wall clock is trustworthy is stamped boot-relative and healed by upgradeBootRelativeTimestamps() once the RTC arrives. Both the stamp and the "same boot?" test were millis() / 1000, which wraps every 49.7 days: a stamp taken before the wrap reads as newer than `bootNow` afterwards, so `m.timestamp <= bootNow` declines to heal it and the message shows "???" until it ages out. MessageRenderer's own copy of the test falls the same way and prints invalidTime. Neither produces a wrong time - the guard is what fails safe - but Time::getUptimeSecs() landed in #11291 for exactly this, and does not wrap for 136 years. Both sites take it, which makes the comparison exact rather than merely fail-safe. While here, the autosave tick had its own hand-rolled deadline helper - `reachedMs(now, target)` as `(int32_t)(now - target) >= 0`. Wrap-correct, but a competing idiom for what Throttle::isWithinTimespanMs() already answers, and the signed cast is the form #11291 replaced everywhere else. Deleted; the stamps read Time::getMillis() so the whole path is on one clock. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * WebServer: drop the hand-rolled millis() wrap branch getAdaptiveInterval() special-cased the wrap by hand: if (currentTime >= lastActivityTime) timeSinceActivity = currentTime - lastActivityTime; else timeSinceActivity = (UINT32_MAX - lastActivityTime) + currentTime + 1; Those two expressions are the same number - unsigned subtraction already computes the difference modulo 2^32 - so this is not a bug, just eight lines reimplementing what Throttle does. It also reads like a site that has thought about the wrap and settled it, which makes it a bad example to copy. Two isWithinTimespanMs() calls against the stored activity stamp, matching ethApiServer's shape for the same adaptive-interval decision. The stamps move to Time::getMillis() so the writes and the reads share a clock. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * MeshPacketQueue: only order elapsed times once both deadlines have passed The late-packet eviction I rewrote compared how long ago each deadline passed: backElapsed < (uint32_t)(now - p->tx_after) That is only an ordering when both deadlines are in the past. An incoming packet whose tx_after is still in the future subtracts to a near-2^32 elapsed, which reads as the most overdue packet in the queue rather than the least - so a full queue would drop the overdue packet it was about to transmit in favour of one that is not ready yet. The comparison it replaced, `backPacket->tx_after > p->tx_after`, got this right away from the wrap; I lost it in the conversion. Classify before ordering: p->tx_after must be unset, or passed, before its elapsed time means anything. Two expired deadlines still order by which is further overdue, which is what the branch is for. Caught by CodeRabbit on #11483. test/test_meshpacket_queue/ pins the branch: the future-dated arrival that started this, both directions of the both-expired ordering, the undelayed arrival, and all of it again with the deadlines and `now` on opposite sides of the wrap. maxLen is 1 so the suite reaches the branch without dragging in CompareMeshPacketFunc and a NodeDB. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * ExpressLRSFiveWay: treat "no key pressed yet" as no activity keyDownStart is 0 until the first press of a boot, and the fast-poll window read that as a press at time zero: 100ms polling for the first 20s of uptime with no activity at all, re-triggering once per millis() wrap. The arithmetic this replaced (`now < keyDownStart + 20000`) did the same, so it is not a regression - but the sentinel is exactly what the conventions say to test before the elapsed comparison, and "has there been recent key activity" has an honest answer here. 250ms is the documented floor for not missing presses, so an idle node simply starts there and moves to 100ms on the first press. Also trims the wrap-cases comment in test_gps_update_scheduling to the two-line house limit. Both from CodeRabbit review on #11483. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
fa031c95dc |
perf(crypto): stop heap-allocating a cipher object per packet (#11462)
* perf(crypto): stop heap-allocating a cipher object per packet encryptAESCtr() constructed a fresh CTR<AES128/256> on the heap for every call - once per encrypted transmit and once per channel decrypt attempt on every received encrypted packet. On the platforms that use this base implementation (STM32WL, RP2040, nRF54L15, portduino) that is avoidable per-packet malloc/free churn on small heaps. Reuse lazily-created singletons instead. Safe for the same reason the function's static scratch buffer already is: every caller serializes under cryptLock, and setKey/setIV reinitialize the cipher state each call. Lazy heap pointers rather than static objects so ESP32/nRF52 (which override this method) never reserve the RAM. * Improve comments in encryptAESCtr function Refactor comments for clarity and conciseness in AES-CTR encryption implementation. --------- Co-authored-by: Thomas Göttgens <tgoettgens@gmail.com> |
||
|
|
4296d5d584 |
fix(eth): free partially initialized TLS contexts on init failure (#11452)
* fix(eth): free partially initialized TLS contexts on init failure initTlsContext() inits the four mbedtls contexts and populates them step by step, but every failure return left the already-parsed material (X.509 chain, EC key, ssl config) allocated. Since tlsReady is only set after full success, deInitEthTlsApiServer()'s cleanup - guarded by if (tlsReady) - could never reclaim that partial state, and runOnce() hard-fails with the contexts stranded for the life of the process. Also reachable via the cert-regeneration path when a DHCP lease change bumps the cert generation and the rebuild fails. Factor the four frees into freeTlsContexts() (safe on init-but- unpopulated contexts) and call it from every initTlsContext failure return; deInit now uses the same helper. * docs(eth): shorten freeTlsContexts comment per review |
||
|
|
d4de61362b |
fix(mesh): plug pooled-object and driver leaks in core paths (#11450)
- MeshService::sendQueueStatusToPhone: release the pooled QueueStatus when the toPhone queue enqueue fails, matching what sendMqttMessageToClientProxy and sendClientNotification already do. The full-queue guard makes this failure rare, but the check/enqueue sequence is not atomic and this path is reachable concurrently from the main loop and the nRF52 BLE write callback; each failure permanently lost one of the four pool slots, and after four losses the phone never receives QueueStatus again until reboot. - MessageStore::storeTextInPool: bail out when the boot-time pool allocation failed instead of memcpy'ing through a null pointer. The read side (getTextFromPool) already guards and maps offset 0 to an empty string. - RF95Interface: hold the RadioLibRF95 driver in a unique_ptr. It is constructed in init(), and when init() subsequently fails (e.g. chip probe NOT_FOUND) initLoRa() destroys the interface, leaking the driver; every sibling interface holds its driver by value so nothing else needed a destructor here. |
||
|
|
22079ca863 |
fix(platform): OOM null-write in stm32wl File and exception leak in portduino GPIO init (#11456)
- STM32_LittleFS File::_open_dir: the _dir_path allocation was the only unchecked malloc in the file, followed immediately by strcpy - an OOM became a NULL write, and the half-initialized state (open dir, null path) would later feed strlen(NULL) in openNextFile(). Check it and unwind the already-opened dir, matching the sibling failure path. - PortduinoGlue initGPIOPin: if setSilent()/gpioBind() threw after the LinuxGPIOPin was constructed, the pointer was lost in the catch block. Hold it in a unique_ptr and release only after the gpio table takes ownership. |
||
|
|
0739d2a68c |
fix(raspihttp): memory safety and error-path leaks in PiWebServer (#11451)
- handleAPIv1ToRadio: the toRadio handler memcpy'd a fixed 512 bytes out of the request body regardless of its actual size. ulfius allocates binary_body at exactly binary_body_length bytes and leaves it NULL for a body-less PUT, so short bodies caused a heap overread and empty ones dereferenced NULL. The unclamped length (ulfius accepts up to 1024) was also handed to handleToRadio, reading past the 512-byte stack buffer. Clamp both directions, matching the ESP32 ContentHandler behavior. - callback_static_file: the stream-free callback only runs when ulfius_set_stream_response succeeds, so the failure branch leaked the FILE (and its fd) on every failed request. Close it and return 500. - CheckSSLandLoad: free cert_pem before the missing-key return; the constructor retries after regenerating certs and the reload overwrote (leaked) the first buffer. - generate_rsa_key / CreateSSLCertificate: free the EVP_PKEY_CTX, EVP_PKEY, and X509 on their error returns; previously only the success paths released them. |
||
|
|
d27d0c4a7f |
fix(detectionsensor): release unsent packets and drop heap scratch buffer (#11449)
Both sendDetectionMessage() and sendCurrentStateMessage() allocate a packet from packetPool and then, when the primary channel is the public/default channel, log and return without sending or releasing it. Each refused send permanently leaks one packet from the fixed-size pool; with state_broadcast_secs configured, the heartbeat path repeats this on a timer until the pool is exhausted and the device can no longer allocate packets at all. Release the packet in the refusal branch, matching the pattern used in PositionModule. Also replace the per-message 'new char[40]' scratch buffer with a stack buffer (and sprintf with snprintf), removing the manual delete[] bookkeeping on every exit path. |
||
|
|
a41ddec1a7 |
fix(nrf54l15): don't write past String buffer when a grow fails (#11454)
* fix(nrf54l15): don't write past String buffer when a grow fails reserve() correctly keeps the old buffer when realloc returns NULL, but returned void, and assign()/concat() proceeded to memcpy with n >= _cap anyway - a heap overflow of up to n+1-_cap bytes into adjacent allocations. On this Zephyr target allocation failure is a realistic condition, and the result was heap corruption instead of a clean no-op. reserve() now reports success and the callers leave the string unchanged when the grow fails. * fix(nrf54l15): guard String length arithmetic against wraparound Per review: reject size requests whose n+1 / _len+n arithmetic would wrap before they reach the capacity check, and make reserve(0) fail without calling realloc (realloc(p, 0) would free the buffer and return NULL, leaving _buf dangling). |
||
|
|
5172e58526 |
fix(nimble): stop leaking BLE2904 descriptor on BLE re-setup (#11453)
setupService() re-runs on every Bluetooth re-enable cycle, and every other callback object in it is deliberately function-local static for exactly that reason - the battery level descriptor was the one heap allocation that was missed. The framework never frees descriptors (~BLECharacteristic has an empty body and BLEDevice::deinit only deletes server/advertising/scan), so each cycle leaked one BLE2904. Make it static like its neighbors. |
||
|
|
87a009da63 |
fix(nrf52): LTO was dropping the board variant's weak hook overrides (#11415)
* fix(nrf52): LTO was dropping the board variant's weak hook overrides Whole-image LTO (enabled arch-wide for nrf52840 in #10655) inlines the empty weak body of earlyInitVariant()/lateInitVariant()/variant_shutdown()/ variant_nrf52LoopHook()/variantDefault*Config() at the call site, because the weak default and the call site live in the SAME translation unit. The strong override in variants/<arch>/<board>/variant.cpp is then never linked, and the board's hardware setup silently does not run. nrf52_lto.py's -fno-lto variant recompile does not help here: the caller is the problem, not the variant object. Needs both ingredients, so this only affects 2.8: the earlyInitVariant() indirection landed in #9438 and is present in v2.7.26 too, but v2.7.26 has no -flto, so the override linked normally. Found on the muzi R1 Neo, whose earlyInitVariant() drives DCDC_EN_HOLD (P0.13, the DC-DC hold after the user button) and NRF_ON (P0.29, "tells IO controller device is on"). Both were dropped from the image, so the companion MCU never saw the nRF application come up and stayed in its DFU indication (purple LED). Verified in the ELF: pre-fix setup() runs straight from waitUntilPowerLevelSafe() to the LED_NOTIFICATION block with no earlyInitVariant symbol in the binary and no pinMode/digitalWrite on P0.13 or P0.29 anywhere; post-fix it calls the real override. HW-confirmed on an R1 Neo. Also affected on nrf52840: earlyInitVariant() on 10 variants (incl. t-echo-card, which sequences its RT9080 3V3 rail there), variant_shutdown() on 18 variants (t114, t-echo, ThinkNode M1-M8, meshlink, wio-tracker-L1 ... - sleep pin parking, so deep-sleep leakage), variant_nrf52LoopHook() on 3 RAK variants. Confirmed dropped on heltec-mesh-node-t114 by build, not just by inspection. Fix is __attribute__((noinline)) on both the weak declaration and definition - the same guard already carried by loopCanSleep(), preFSBegin(), PowerHAL and variant_enableBatteryLpcompWake(), whose comment in main-nrf52.cpp already documents this exact failure mode. Also extend _VARIANT_OVERRIDES in extra_scripts/nrf52_lto.py from just _Z11initVariantv to all eight hooks. That post-link guard already had the right logic and would have caught this on every PR - it simply was not listing Meshtastic's own weak variant hooks, only the core's. With the list extended it goes red on both r1-neo and heltec-mesh-node-t114 when the noinline is reverted, and green with it. Its failure message now names both possible causes. * review: trim the noinline rationale comments to two lines Per AGENTS.md ("keep code comments minimal - one or two lines, max"), the incident detail and extended background belong in the PR description, not the source. Keeps the LTO/noinline rationale and the pointer to the guard. --------- Co-authored-by: Jonathan Bennett <jbennett@incomsystems.biz> |
||
|
|
25996e9330 |
Fix W12 battery reading: add ADC_CTRL and correct the divider ratio (#11323)
* Fix W12 battery reading: add ADC_CTRL and correct the divider ratio
The W12 battery config was taken from the vendor's Demo_06_ADC_Read.ino,
which reads GPIO1, multiplies by 2.0 under a literal "Assumption: 2:1
voltage divider" comment, and never touches the ADC enable at all. All
three of those details are wrong.
Schematic W12-MB-V0.2 sheet 1 has the divider behind a P-MOSFET high-side
switch so it only draws from the cell during a reading:
BAT --S[Q6 AO3401A]D-- R50 390K --IO1_ADC_IN-- R51 100K -- GND
|G R49 1K to BAT (gate pull-up: Q6 off by default)
+-- R48 1K -- C[Q7 S8050 NPN]E -- GND, base <- R52 1K <- IO2
So GPIO2 is ADC_CTRL, not "a second (solar/VUSB) divider" as the variant
claimed, and the NPN inverts it, making it active HIGH. Left undriven, Q6
stays off and GPIO1 sits at ground through R51 - a hard 0 raw rather than
the 100-250mV of noise a floating pin gives - so every boot reported
"battery hardware absent (USB-only)" and battery_level 101.
The divider is 390K/100K, so the multiplier is 4.9, not 2.0. That puts a
4.2V cell at only ~857mV on the pin, so drop the attenuation from the
12dB default (0-3100mV) to 2.5dB (0-1250mV) to use the range properly.
This matches the Heltec V3/V4 network, but their ADC_CTRL 37 cannot be
reused here: GPIO33-37 are consumed by this board's octal PSRAM.
Verified on hardware: reports 4067mV / 91%, stable to the mV across
consecutive samples, where it previously read 0mV with a cell attached.
* Trim the battery comment block to house style
Per the repo guideline that code comments stay to one or two lines and
avoid multi-paragraph blocks, drop the ASCII schematic from the header.
The full circuit trace lives in the previous commit message and the PR
description, which is where that rationale belongs.
Comment-only; both define values are unchanged.
|
||
|
|
e2460b546e |
MUI: stop holding the SPI bus across the whole UI cycle (#11278)
tft_task_handler held spiLock for the entire LVGL cycle. Most of that cycle is timer work and rendering into the draw buffer, which issues no SPI at all - but on boards where the TFT, SD card and LoRa radio share one bus (T-Deck), every radio operation on the main loop still waited it out. That is tens to hundreds of milliseconds whenever the UI animates, felt as mesh RX/TX latency. device-ui now takes the lock around its own transfers instead (meshtastic/device-ui#356), so the coarse hold here can go and the bus is contended only during real traffic. Lend it spiLock through a reentrant adapter. device-ui nests its guards - SdFsCard::usedBytes() calls cardSize() and freeBytes(), each of which takes the lock - while spiLock is a plain binary semaphore that would self-deadlock on the second take, so track the owning task and only touch the underlying lock on the outermost acquire. Requires the device-ui pin bump included here. Tested on T-Deck: flush, touch, panel init, SD detect, powersave sleep and wake all exercised; LoRa RX decoding under a live UI, no deadlocks, no watchdog resets. Also builds seeed-sensecap-indicator-tft. Co-authored-by: Manuel <71137295+mverch67@users.noreply.github.com> |
||
|
|
a8623a60c5 |
Stream our own position to the phone/UI while mesh position sharing is opt-in (#11270)
* Position: stream our own position to the phone/UI while mesh sharing is opt-in Position broadcasts became opt-in in 2.8 (#10929): with every public channel at position_precision 0, sendOurPosition() finds no eligible channel and returns without queueing anything, so the connected phone or on-device UI never sees the node's own GPS fix ("GPS looks dead" on standalone MUI devices even though the receiver has a lock). Mirror device telemetry's local delivery: once a minute, when the toPhone queue is idle, stream our own position to the connected client at full precision. The packet is handed straight to sendToPhone() and never touches the mesh, so the per-channel opt-in and the public-channel precision clamp still govern everything on the air. Also stop logging "Send pos ... to mesh" before the channel scan has found an eligible channel; when sharing is disabled everywhere the skip is now logged explicitly instead of pretending a send happened. The cadence gate is a pure static (shouldSendPositionToPhone) alongside the existing broadcast-policy helpers, with unit tests covering the first-send, cadence, gating, and millis() rollover cases. * Review: only advance phone cadence on a queued packet; drop the ms==0 sentinel sendOurPositionToPhone() now reports whether a packet actually reached the phone queue, and runOnce() stamps the cadence only on success, so a guard or allocation failure retries on the next tick instead of waiting out a minute. The never-sent state is a dedicated hasSentPositionToPhone flag rather than lastPhoneSendMs == 0, so a send stamped exactly at millis() == 0 still holds the cadence. New regression test covers that case; existing cases updated to the explicit flag. * Review: align the rollover test fixture with its documented elapsed times lastSent now sits exactly 30,000 ms before the uint32 wrap, so the two cases are precisely 70,000 ms (sends) and 40,000 ms (held) - the previous comments claimed 70s/20s against actual elapsed values of 70,001/40,001 ms. |
||
|
|
2a20ee5954 |
ESP32: link via @response files across the family, not just P4
The rak_wismesh_tap_v2-tft link failed on CI with "sh: Argument list too long": the device-ui bump in #11214 grew the MUI object list enough that the linker command line crossed the kernel exec arg limit (this env's long name lengthens every object path, so it tipped first). extra_scripts/ld_response_file.py already solves this for esp32p4; promote it to esp32_common so every ESP32 target links through a response file, and drop the now-inherited entry from esp32p4.ini. Verified by building rak_wismesh_tap_v2-tft at the failing commit plus this fix: links clean (macOS's exec arg limit is smaller than the Linux one that tripped CI, so an unwrapped link could not have passed). |
||
|
|
134fe5ec3e |
Stop BLE log streaming from starving the NimBLE notify pool (#11253)
sendLog() notified once per firmware log line with no backpressure. With debug_log_api_enabled and a subscribed client, a logging burst drains the msys_1 mbuf pool that every GATT notification and ATT response allocates from, so ble_gatts_notify_custom starts returning BLE_HS_ENOMEM (rc=6) -- for the fromNum doorbell too, not just the log line that exhausted it. Each failure then prints an error over serial at ~12ms apiece, which is itself enough to throttle the main loop during a burst. Register a callback on the log characteristic so the host's own ERROR_GATT verdict pauses log notifies for 250ms and lets the pool refill. Keying off the failure rather than a free-block count keeps this correct however the pools are sized. Also restore CONFIG_BT_NIMBLE_MSYS_1_BLOCK_COUNT to the IDF default of 12. Pool selection is by block size and a notify request always resolves to the smallest pool, so msys_1 is the only pool GATT draws from; 8 was thin once anything chains or fragments. msys_2 and the ACL/EVT transport pools are left trimmed -- neither is on the notify path, and restoring the full defaults would re-spend the contiguous allocation that #10741 was fixing on heap-tight boards. sendLog() also lacked the null check onNowHasData() already has, so a log line arriving during BLE teardown dereferenced a freed characteristic. Fixes #11245 |
||
|
|
ae16c052d5 |
Time out an abandoned admin edit transaction (#11254)
begin_edit_settings sets a bool that only the matching commit ever clears. If the commit never arrives -- the client dropped its link partway through a bulk config import, or went away entirely -- the transaction stays open indefinitely, and every later config write from any client is applied to RAM, acknowledged, and then never saved. The writes look successful and are visible in get_config, but the node reverts all of them at the next boot, and only a reboot clears it. Give the transaction a one minute idle timeout. Each deferred write restarts the clock, so the window bounds the gap between writes rather than the length of the edit; a bulk import sends them milliseconds apart. The next admin message after it lapses retires the transaction, persisting the segments it had deferred and flushing the warnings it was holding. The check runs after the auth gates and before the switch, so every case sees consistent state and the recovery's flash write happens in the main loop rather than a disconnect callback. It saves rather than rolls back because there is nothing to roll back to: each write already took effect in RAM and was acknowledged, so the transaction only ever deferred the save. That also makes an early expiry cheap -- the worst case for a client that really was still going is that its remaining writes are saved individually instead of batched. Closing on disconnect instead was tempting, but the reporter's captures show iOS dropping and reconnecting within half a second several times per session, which would abort imports that currently survive the blip. Also drops the file-scope hasOpenEditTransaction, shadowed by the class member at every use site and referenced nowhere else in the tree. Reported in #11245 |
||
|
|
f61f65891c |
fix(promicro): exclude emoji to fit under the warm-store region (#11256)
nrf52_promicro_diy_tcxo has been failing the nrf52 warm-region guard on develop: the image ends at 0xEA0C0, 192 bytes past the 12 KB WarmNodeStore record-ring reserved at 0xEA000. This variant compiles four radio driver families (SX126x/LLCC68, SX127x/RF95, LR11x0/LR1121, LR2021) so any module can be soldered on, which makes it the largest nrf52 image we ship. Building it with -D EXCLUDE_EMOJI saves 6,808 bytes and puts the image at 0xE8618, 6.6 KB clear of the warm region. EXCLUDE_EMOJI was not previously usable: graphics::emotes[] becomes empty, but the canned-message emote picker never checked for that. Entering the picker clamped emotePickerIndex to numEmotes - 1 (i.e. -1), and selecting read emotes[-1].label into a String - an out-of-bounds read of whatever precedes the array in flash. Guard both entry points instead: refuse to open the picker when there are no emotes, and bounce back to freetext if the picker state is somehow reached anyway. Under whole-image LTO those guards fold to constants, so the picker draw and input paths dead-strip entirely on builds that set EXCLUDE_EMOJI - which is where the savings come from beyond the bitmap data itself. Received messages containing emoji render as text on this variant, and the emote-list key is a no-op. Verified rak4631 (emoji enabled) is unaffected: 0xE46D8, 22 KB clear. |
||
|
|
5c5cb094e6 |
fix(admin): stop the admin dispatcher from overflowing the ESP32 loopTask stack (#11239)
AdminModule::handleReceivedProtobuf carried a 3008-byte stack frame - 37% of the 8 KB Arduino loopTask stack. GCC inlines the eight handleGet* helpers into it, each of which builds a whole meshtastic_AdminMessage (480 B), so the dispatcher reserved a slot for every one of them at once even though only one case ever runs. Two more full AdminMessages lived directly in the function. That left any admin write short of stack for what runs beneath it: NodeDB::saveToDisk -> saveProto -> pb_encode -> SafeFile -> LittleFS -> flash. On heltec-v4 the canary fired at 7824 of 8192 bytes while writing /prefs/config.proto, so no setting ever persisted and the device rebooted. Mark the response builders NOINLINE and move the module-API and observer responses into their own functions. Pure refactor; measured on heltec-v4 the dispatcher frame drops 3008 -> 384 bytes, taking the crashing path from 368 bytes of headroom to 2992. Fixes #11237 |
||
|
|
3f4e7cc622 |
Fix nRF52 freeze + watchdog reset when saving config over BLE (#11202)
* Fix nRF52 freeze + watchdog reset when saving config over BLE Since #10967 phone-originated admin messages are handled synchronously on Bluefruit's BLE event task. NRF52Bluetooth::disconnect() busy-waited for BLE_GAP_EVT_DISCONNECTED, which only that same task can process, so any config save that requires a reboot (e.g. position) deadlocked the device until the 90s watchdog fired. Bound the wait to 1s and sleep instead of spinning so lower-priority tasks (including the watchdog feed) keep running; the SoftDevice completes the link termination on its own. * Address review: use Throttle helper, tighten comment * Name the disconnect timeout constant * Log unconfirmed BLE disconnect at WARN with elapsed time |
||
|
|
becfb770c1 |
nrf52: fix BLE-task stack overflow crashing pairing and wiping LittleFS (#11190)
Since #10967 made Router::sendLocal handle self-addressed packets synchronously, the entire phone-API chain for a BLE client runs inline in the Bluefruit characteristic write callback: toRadioWriteCb -> PhoneAPI::handleToRadio -> admin set-config -> radio reconfigure -> NodeDB::saveToDisk. That callback executes on the Bluefruit BLE FreeRTOS task, whose stock stack is 5 KB (CFG_BLE_TASK_STACKSIZE = 256*5 words) - not the Arduino loop task that #10944 already raised to 8 KB. The loop-task fix therefore protects the wrong task for BLE-originated writes. On a Seeed Wio Tracker L1 the 5 KB stack overflows during pairing first-sync, resetting the device mid-LittleFS-write, every single time. Repeated mid-write resets tear the LittleFS metadata, lfs_assert fires on the next boot, and the corruption handler formats the whole filesystem: region, channels, module config, and the node's keypair are all lost (critical fault #13, new node identity on next region set). Reproduced end-to-end tonight on stock develop 6908d27; with this change the same device pairs, serves config screens, and survives back-to-back config.proto saves over BLE. Raise the BLE task to the same 2048 words (8 KB) as LOOP_STACK_SZ, for the same reason. bluefruit.cpp's #ifndef guard makes the -D take effect with no framework patch. Costs 3 KB of RAM on nrf52840 targets only. Credit where due: Ixitxachitl independently established in #11155 testing that the save-path crash persists after #11185 and that re-queueing sendLocal (moving the pipeline back to the Router thread) makes it go away - which corroborates this diagnosis from the other direction. This commit is the minimal capacity-side fix; #11155's relocation of the pipeline off the BLE task remains the right architectural follow-up, and this guard stays correct even after it lands. Likely also explains #10905 (L1 display-thread crash when a client requests full configuration) and the 2.8 field reports of idle nodes losing region and keys after a BLE session. |
||
|
|
01d1873bfc |
Deliver locally-generated replies addressed to us to the phone (#11185)
* Deliver locally-generated replies addressed to us to the phone Config get/set from the phone times out on every device: the client sends an admin request, the node handles it, and the response is silently dropped before it reaches the phone queue. #10967 changed Router::sendLocal's isToUs branch from enqueueReceivedMessage() to handleReceived(p, src), so a local packet keeps its RxSource instead of being relabeled RX_SRC_RADIO by the queue round-trip. That is the right call for the new policy gates, but module replies go out through MeshService::sendToMesh() with the default RX_SRC_LOCAL, and a reply to a phone-originated request is addressed to our own node (setReplyTo resolves from == 0 to ourNodeNum). Those replies now re-enter callModules as RX_SRC_LOCAL, where the loopback gate skips every module whose loopbackOk is false - including RoutingModule, whose promiscuous sniff is the only path that moves a received packet into toPhoneQueue. The reply is released, never sent. Requests still work, because the phone's own packets arrive as RX_SRC_USER and pass the gate, so a set_config is applied and only its acknowledgement is lost. That is why a client can connect and download config but times out on every config screen and every setter. Deliver the phone's copy from sendToMesh() instead: for a local packet addressed to us, the loopback gate is doing its job in keeping the packet away from module re-dispatch, and the phone copy is exactly what is missing. Setting loopbackOk on RoutingModule would instead echo every locally-generated broadcast back to the phone, and relabeling replies RX_SRC_RADIO would undo the origin separation #10967 added. Also stop reporting ERRNO_SHOULD_RELEASE (35) to the phone in the QueueStatus for these packets. It means "caller frees", not a send failure, and the same hunk changed it from the 0 the phone used to see. * Address review: trim comments, assert the QueueStatus count Condense the added comments to the one-or-two-line house style; the rationale lives in the commit message and PR. The reply test drained QueueStatus records in a while loop, which would have passed just as happily on an empty queue. Count them and require both the request's and the reply's. |
||
|
|
b08c04f264 |
Add Meshnology W12 (WiFi LoRa 32 V5) LR2021 board support (#10910)
* Add Meshnology W12 (WiFi LoRa 32 V5) variant with LR2021 radio ESP32-S3R8 + 16 MB flash + Semtech LR2021 dual-band LoRa + SSD1315 OLED, Heltec V3-style pinout. Pin map comes from the vendor pin sheet and vendor Arduino demos, confirmed on hardware: I2C scan finds the OLED at 0x3C on SDA17/SCL18, and the LR2021 answers on NSS=8 SCK=9 MISO=11 MOSI=10 with BUSY=13/NRESET=12 verified by reset-pulse probing. The IRQ line needs LR2021_IRQ_DIO_NUM 8. RadioLib defaults irqDioNum to 5, but DIO5 is not bonded out on this board - the two IRQ lines are DIO7 -> IO7 and DIO8 -> IO14, established by driving each radio DIO as a GPIO output and watching which ESP32 pin followed. Left at the default, every interrupt is routed to a floating pin: RX/TX still appear to work because RadioLibInterface::pollMissedIrqs() falls back to polling the IRQ status register, but the blocking scanChannel() waits on the pin itself and never returns. A shadow pins_arduino.h is needed because the generic esp32s3 one defines RGB_BUILTIN, which pulls the Arduino RMT HAL into the link and fails against this build's FreeRTOS config. * Address review: add W12 variant metadata, trim comments Declares the custom_meshtastic_* metadata the build manifest expects. Without it the emitted .mt.json carried no hwModel, slug, display name or support level, so the board was invisible to the flasher - and board_level = extra kept it out of PR builds entirely, which is why it never appeared in this PR's board list. It now matches the W10 sibling at board_level = pr with support level 1. hwModel stays 255/PRIVATE_HW until a W12 enum lands in the protobufs. Also trims the header comments to the one-or-two-line limit in AGENTS.md; the vendor pin sheet and probing evidence live in the PR description. |
||
|
|
876962e8f2 | Protos | ||
|
|
d9150e8dc7 |
ci: warn and skip the pio check job on transient network failures (#11145)
* ci: warn and skip the pio check job on transient network failures
The `check` matrix job intermittently fails before it analyses anything:
`pio check` has to fetch the platform package first, and that download gets
reset mid-flight by the CDN often enough to be a nuisance. Three of the last
~200 CI runs died this way, all identical and all ~1s into the run:
Checking station-g3 > cppcheck (board: station-g3; platform:
https://github.com/pioarduino/platform-espressif32/.../platform-espressif32.zip)
requests.exceptions.ConnectionError: ('Connection aborted.',
ConnectionResetError(104, 'Connection reset by peer'))
cppcheck never ran, so there is no signal at all - just a red X someone has to
re-run by hand.
`pio check` exits 1 for both a dead socket and a real defect, so only the output
can tell them apart. check-all.sh now tees the run and classifies the failure:
* cppcheck reported something (summary table or a `file:line: [severity]`
defect line) -> real, propagate the exit status. This is checked first, so
no amount of network noise elsewhere in the log can mask a genuine defect.
* transport-level error and nothing else -> log a ::warning:: annotation
naming the board and the underlying error, and exit 0 under CI (exit 2
locally, so a local run never quietly reports itself clean).
* anything else -> propagate the exit status.
404 / 403 / UnknownPackageError are deliberately not treated as transient: a
missing or forbidden package is a reproducible break from a bad pin in a .ini,
not weather.
No workflow change is needed - gh-action-firmware's entrypoint runs
/workspace/bin/check-all.sh for MT_TARGET=check, so the whole fix lives in this
repo and applies to local runs too.
Verified by replaying the captured CI logs through the script with a stubbed
`pio`: the station-g3 connection-reset log skips with a warning under CI and
exits 2 locally, the rak11200 `Total 1 0 0` defect log still fails, that same
defect log with the reset traceback appended still fails, and a 404 still fails.
* ci: pass board names to pio as an argument array
Addresses CodeRabbit review feedback on the previous commit (and the SC2086 the
line has carried since it was written): BOARDS was a space-joined string and
$CHECK was expanded unquoted, so a board override containing whitespace or a
glob character could turn into extra pio arguments.
BOARDS and CHECK are arrays now, and pio is invoked with "${CHECK[@]}", so board
names reach it as literal arguments. The two skip messages use "${BOARDS[*]}" -
a bare "${BOARDS}" would have quietly named only the first board.
shellcheck -x is clean on the file. Re-verified with a stubbed pio that records
its argv: 'tb*' and 'tbeam extra' each arrive as one literal argument, the
default list still expands to 15 -e pairs, a two-board skip names both boards,
and all six classification cases (network/CI, network/local, real defects,
defects with a reset traceback appended, 404, clean pass) are unchanged.
|
||
|
|
efc9708236 |
Add userpref for the packet signature policy default (#11144)
Co-authored-by: Austin <vidplace7@gmail.com> |
||
|
|
8a9980576f |
Merge pull request #11154 from RCGV1/codex/fix-packet-policy-excluded-test
test: cover packet policy coercion on excluded builds |
||
|
|
1804fd1886 |
Merge pull request #11150 from meshtastic/always-exclude-rangetest
Always exclude range test |
||
|
|
cad3a3a5c7 |
Always exclude range test
Range test is excluded on every target as of 2.8, so set MESHTASTIC_EXCLUDE_RANGETEST=1 globally in [env] build_flags and drop the now-redundant per-variant flags from esp32, rak11200 and stm32. Guard RangeTestModule.cpp with the same condition used in Modules.cpp so the translation unit compiles to nothing rather than relying on --gc-sections to strip it, and always report RANGETEST_CONFIG in the device metadata excluded_modules bitmask so clients hide the config. |
||
|
|
9662c818b5 |
Merge pull request #11148 from NomDeTom/flash-clawback
claw back some space from signing policies |
||
|
|
f90c1340de |
Merge pull request #11117 from NomDeTom/nailed-test-number
Nailed native test numbers so PRs don't forget to increment |
||
|
|
c263a9e4a3 |
Merge pull request #11136 from meshtastic/renovate/rak13800-w5100s-1.x
Update RAK13800-W5100S to v1.0.3 |
||
|
|
1d6f8a8f34 |
Merge pull request #11130 from meshtastic/ci-native-tests-by-area
Run native PlatformIO tests by area for readable failures |
||
|
|
5e3d1ea606 |
Merge pull request #11129 from RCGV1/agent/include-statusmessage-config
phoneapi: send status message config |
||
|
|
c6669a4ed8 |
Merge pull request #11128 from RCGV1/agent/keyverification-allocation-safety
keyverification: handle response allocation failure |
||
|
|
6e37b53324 |
Merge pull request #11134 from RCGV1/codex/verify-packet-auth-config-docker
fix(security): round-trip packet policy defaults |
||
|
|
72aac12b8d | Merge branch 'develop' into agent/keyverification-allocation-safety | ||
|
|
df971c2921 | Merge branch 'develop' into agent/include-statusmessage-config | ||
|
|
6e48fac741 | Merge pull request #11125 from meshtastic/baseui_fixindicator | ||
|
|
5548bd3195 | Merge pull request #10967 from RCGV1/codex/packet-auth-policy | ||
|
|
fdb7668c15 | Merge pull request #11124 from meshtastic/fix/admin-pin-tests-simradio | ||
|
|
6fc7c1b1db |
test: unset force_simradio so admin request pinning is exercised
The native test harness boots Portduino in simulated mode, and
wouldEncryptWithPKC() short-circuits to false whenever
portduino_config.force_simradio is set. noteOutgoingAdminRequest() derives its
pin from that predicate:
keyValid = haveDestKey && wouldEncryptWithPKC(&p, p.channel, haveDestKey);
so under the harness no outgoing admin request is ever key-pinned, and
responseIsSolicited() admits a plaintext response to a PKC-pinned request.
test_pinned_request_keeps_its_key_after_an_unpinned_request and
test_request_to_keyed_node_pins_the_stored_key have therefore failed since
#11092 added them, on develop and on every branch that merges it.
Instrumenting the predicate shows every other term already satisfied
(haveDestKey=1, isFromUs=1, private_key=32, unicast, ADMIN_APP, channel
LongFast) with wouldEncryptWithPKC=0, leaving only the simradio guard.
Clear the flag in setUp so the fixture models a real device. Skipping PKC under
force_simradio is correct for a simulated radio - there is no PKC to pin - so
the predicate is left alone rather than relaxed to make a test pass. The only
sim-mode path in this suite is AdminModule's exit_simulator intercept, which no
test here exercises.
Native suite: 38 suites, 723/723 cases, no sanitizer findings.
|
||
|
|
a58daf015d |
Don't misdetect the ES7210 audio codec as an INA219 (#11120)
The T-Deck's ES7210 audio ADC straps to 0x40, which is INA_ADDR. It answers none of the INA226/INA260/SHT2X checks, so it fell into the bare "assume INA219" fallback and PowerTelemetry then drove an audio codec as a power monitor, panicking the board into a boot loop. The INA219 has no manufacturer or die ID register, so it can only ever be inferred. Positively identify the ES7210 instead, via its chip ID registers 0x3D/0x3E (0x72/0x10), and leave the device unclaimed. Also guard the INA219 fallback on the type rather than chaining it to the SHT2X check: that else bound to the SHT2X if, so a positively identified INA226/INA260 was overwritten with INA219 right after being found. Fixes #11115 |
||
|
|
e11635baa2 | Merge branch 'develop' into codex/packet-auth-policy | ||
|
|
7d954e668d |
Merge branch 'develop' into codex/packet-auth-policy
Resolve conflicts against the NodeDB signer/key primitives (#11050) and the admin-key PKI decrypt budget (#11100). - NodeDB: drop this branch's hasSeenXeddsaSigner in favour of develop's isKnownXeddsaSigner. They answer the same question, but develop's reads the dedicated warm signer bit (warmSignerOf) rather than the WarmProtected category, and TrafficManagementModule already depends on it. Keep develop's copyPublicKey/copyPublicKeyAuthoritative, isVerifiedSignerForKey and commitRemoteKey/KeyCommitTrust. - checkXeddsaReceivePolicy: keep this branch's Strict/Balanced/Compatible policy, which is a superset of develop's balanced-only downgrade gate, and call isKnownXeddsaSigner from it. develop's !pki_encrypted term is dropped because the policy returns early for PKI packets before that check. - perhapsDecode: keep develop's key resolution (NodeDB then pending-key, only for real PKI candidates) plus its admin-key token bucket, and re-apply this branch's pkiAttempted flag feeding the DECODE_OPAQUE verdict. Keep both passesRoutingAuthGate and adminKeyFallbackAllowed/Refund. - test_A17: model eviction the way NodeDB actually does it, passing the warm signer bit as well as the XeddsaSigner category, since isKnownXeddsaSigner reads the former. Native suite: 38 suites, 743/743 cases, no sanitizer findings. |
||
|
|
e247f70c6d |
test: drain toPhoneQueue in MQTT test to fix ASan leak
sendLocal() now dispatches local packets through handleReceived() directly instead of the mock-overridden enqueueReceivedMessage(). The MQTT implicit ACK-to-self therefore runs the full receive pipeline (RoutingModule -> MeshService::handleFromRadio -> sendToPhone), enqueuing pooled MeshPacket copies into toPhoneQueue. Production drains that queue via PhoneAPI, but the test has no phone reader, so the copies leaked at teardown (LeakSanitizer: 42824 bytes / 101 objects across test_receiveFuzzServiceEnvelope and test_receiveAcksOwnSentMessages). Give MockMeshService a destructor that drains toPhoneQueue like the phone would. Test-only; the real firmware does not leak here. |
||
|
|
0165e914a9 | fix: gate module replies by request port (#11094) | ||
|
|
8b0f29026a | fix: include board PSRAM class in the HybridCompile sdkconfig cache key (#11112) | ||
|
|
1317176f78 |
test: accept DECODE_OPAQUE/DECODE_POLICY_REJECT in perhapsDecode fuzz assertion
The packet-auth-policy change extends the DecodeState enum with DECODE_OPAQUE and DECODE_POLICY_REJECT. test_E1_perhaps_decode_fuzz drives arbitrary ciphertext through perhapsDecode and asserted the verdict was one of the original three states; random ciphertext that matches no channel with no PKI attempt now returns DECODE_OPAQUE, tripping the assertion. Broaden the check to accept all five valid verdicts, matching the test's stated 'any verdict is fine' contract. |
||
|
|
beb9a3095d | Merge branch 'develop' into codex/packet-auth-policy | ||
|
|
5cf346311c |
fix: don't wipe admin keys when regenerating the keypair (#11088)
A client's "regenerate keys" action sends a blank SecurityConfig carrying only the new private key rather than the config it read from the device, so assigning it wholesale cleared admin_key, is_managed, serial_enabled, debug_log_api_enabled, admin_channel_enabled and packet_signature_policy. Losing the admin keys locks the owner out of remote admin with no recourse but a physical connection to the node. Detect that bare-rotation shape and swap in just the keypair, leaving the rest of the security config intact. Deliberately clearing admin keys still works through a SET that leaves the private key alone. Fixes #11073 |
||
|
|
829ff80a09 |
fix(t5s3-epaper): 40MHz octal PSRAM; guard EInk init; esp32s3 sw-atomic shim (#11062)
- t5s3-epaper: the board's 3.3V octal (AP_3v3) PSRAM is unreliable at the qio_opi
default of 80MHz (Total PSRAM reads 0, display dead / boot-loop); clock it at 40MHz.
- EInkParallelDisplay: skip EInk init and no-op display ops when the PSRAM framebuffer
is unavailable, so the node boots headless instead of hard-faulting on a NULL buffer.
- src/platform/esp32: weak software __atomic_*_{1,2,4} so esp32s3 links on toolchains
(e.g. macOS) that lack the sized libcalls GCC emits under -mdisable-hardware-atomics.
|
||
|
|
4726375282 | Update protos | ||
|
|
b20d89974a |
Stop accelerometer thread when double-tap/wake-on-motion disabled at runtime (#11025)
* Stop accelerometer thread when double-tap/wake-on-motion disabled at runtime double_tap_as_button_press and wake_on_tap_or_motion are applied live: the OFF->ON edge calls accelerometerThread->start(), but there was no ON->OFF branch, so turning either flag off left the sensor thread running (polling I2C, drawing power) until reboot. Worse, because enabled stayed true, a later OFF->ON edge was a no-op (the enabled==false guard blocked re-start), leaving the feature un-restartable without a reboot. Add the symmetric ON->OFF branch in both handlers. When a flag goes true->off and the other consumer of the shared thread is also off, call accelerometerThread->disable() (stops runOnce polling and clears enabled so a later re-enable can start() again). Each branch checks the other flag first so disabling one feature never stops the sensor while the other still needs it. * Keep accelerometer thread running when its sensor drives the compass * Guard accelerometer thread config toggles against a null thread pointer * AdminModule: factor shared accelerometer start/stop into a helper The device and display config handlers had mirror-image blocks reconciling the shared accelerometer thread. Extract reconcileAccelerometerThread(wasOn, nowOn, otherFeatureOn) so the null guard, edge logic, compass (providesHeading) guard, and rationale live in one place; each call site is now a single call. Behavior is unchanged. Also drops the redundant per-field assignment that the whole-struct `config.device = ...` / `config.display = ...` overwrites anyway. --------- Co-authored-by: Thomas Göttgens <tgoettgens@gmail.com> |
||
|
|
6e0e9c174d | Protobufs | ||
|
|
f6f3284cfc |
fix(lockdown): re-lock serial console admin auth on USB link drop (#11028)
On lockdown builds (MESHTASTIC_PHONEAPI_ACCESS_CONTROL, nRF52) the SerialConsole is a process-lifetime singleton, so the per-connection admin-auth slot keyed by its inherited PhoneAPI* is reused for every USB/serial client for the whole boot. An operator's admin unlock stayed latched across serial client swaps: an attacker plugging into the USB/serial port before the prior session's 15-minute inactivity timeout inherited admin authorization -- the serial analog of the BLE stale-session reuse bug closed by resetting state in onConnect()/onDisconnect(). Sample the USB-CDC host link (DTR/mount) state each runOnce(); on the link-drop edge call close(), which frees the auth slot and resets PhoneAPI state so whoever connects next re-locks via handleStartConfig()'s !isConnected() branch on their first want_config -- the same physical-link boundary BLE enforces in onConnect(). On the nRF52 TinyUSB (Adafruit) core, (bool)Port == tud_cdc_n_connected(), which goes false on cable unplug or host port-close. Console transports without a real DTR line fall back to the existing inactivity timeout, no worse than before. Entirely nRF52-lockdown-gated; non-lockdown builds are byte-identical. |
||
|
|
5cf85d523a |
fix: heap leaks on display wake and BLE re-enable; remove dead PacketCache (#11019)
* fix(display): don't leak a TFT driver instance on every screen wake On variants whose Screen::handleSetOn() re-runs ui->init() on each display wake (Heltec Tracker V1.x/V2, VTFT_LEDA and ST7796 boards, MUZI/Cardputer via dispdev->init()), OLEDDisplay::init() unconditionally re-invokes connect(), and TFTDisplay::connect() allocated a fresh driver instance each time without freeing the old one. Each OFF->ON transition orphaned a full LGFX device (measured: exactly 1,260 bytes per wake on a Heltec Wireless Tracker V2). Since PowerFSM wakes the screen on every received text message, a device on a busy channel leaked tens of KB per day - matching field reports of heap climbing from 78% to 92% within a day. Null-guard the driver allocations (matching the existing linePixelBuffer/ repaintChunkBuffer guards below them) so re-entry re-runs tft->init() on the existing instance. Bench-validated on a Tracker V2: free heap now flat across 13 consecutive wake cycles, with the panel still re-initializing and rendering on every wake. * fix(ble): don't leak BluetoothPhoneAPI and callbacks on BLE re-enable NimbleBluetooth::setup()/setupService() re-run when Bluetooth is re-enabled after deinit(), but BLEDevice::deinit() only frees the GATT objects - the caller-owned allocations were re-created unguarded on every cycle: - bluetoothPhoneAPI: ~3.5KB PhoneAPI object per cycle, and the orphaned instance is an OSThread that stays registered with the scheduler, so a duplicate thread kept servicing the same static queues - toRadioCallbacks / fromRadioCallbacks / security + server callbacks: one small object each per cycle - the BLESecurity shim was heap-allocated and never freed even on first boot; it only forwards to static setters, so use a stack instance Reuse the existing objects on re-setup; the setCallbacks() calls still run every time since the characteristics themselves are new. * chore: remove unused PacketCache PacketCache landed in #8341 but no consumer was ever wired up - repo-wide, the only references to PacketCache/packetCache are in its own two files. As designed it also malloc()s per cached packet with no eviction or size cap, so it should be re-reviewed for bounds before any future use. Remove the dead code; git history preserves it if a bounded revival is wanted. * fix(ble): reset stale session state when reusing BluetoothPhoneAPI Review follow-up: reusing bluetoothPhoneAPI across BLE enable cycles could hand the next session the previous session's dirty state. The only full cleanup path (onDisconnect) is skipped when deinit()'s bounded disconnect wait expires before the event lands (its 2s cap matches the connection supervision timeout, so a phone that walked out of range makes this a coin flip): the reused object then enters the next session mid-config, with stale queue contents served to the new phone and a stale connection handle that defeats the checkConnectionTimeout self-heal. Factor onDisconnect's cleanup into resetBleSessionState() and run it from setupService() when reusing the instance, restoring the old fresh-object invariant. Also switch the four stateless callback objects to function- local statics (the resolved framework BLE wrapper stores plain pointers and never frees them, so static instances are safe and avoid the guarded heap allocations), correct the comment that claimed deinit() frees the GATT objects (it deletes only the BLEServer itself; the services and characteristics it created remain a small library-side leak per cycle), and fix the inverted connection check in getRssi() that made BLE RSSI always read 0 on ESP32-S3/C6. * refactor(display): construct-once guard around the whole driver ladder Review follow-up: one if (!tft) around the #if/#elif/#else construction block instead of a guard per branch, so a future display family can't reintroduce the per-wake leak by copying an unguarded branch. Make the HACKADAY bus pointer local to the construction (it was a write-only global), and point the comment at the Screen::handleSetOn gates instead of hand-listing boards that would go stale. * fix(ble): use-after-free notifying a freed characteristic after deinit() deinit() nulled bleServer and BatteryCharacteristic but left fromNumCharacteristic dangling after BLEDevice::deinit(true) freed the GATT graph, and it never detached the PhoneAPI fromNumChanged observer. When deinit()'s bounded 2s disconnect wait expires before onDisconnect runs (the same stale-bond / host-reset race resetBleSessionState was built for), close() is skipped, so the observer stays attached with state == STATE_SEND_PACKETS. After BLE is off, the next mesh packet drives MeshService fromNumChanged -> PhoneAPI::onNotify -> onNowHasData -> fromNumCharacteristic->notify() on freed memory (the framework BLE wrapper's notify() dereferences the freed server via getConnectedCount). Unlike sendLog(), onNowHasData() had no isConnected guard. deinit() now calls resetBleSessionState() to detach the observer and reset session state unconditionally (also forcing the conn handle to NONE so checkConnectionTimeout can't be fooled by the stale handle), nulls fromNumCharacteristic/logRadioCharacteristic like the other freed pointers, and onNowHasData() bails on a null characteristic. Reachable on all ESP32/S3/C6 NimBLE boards via admin disable-bluetooth or a sleep transition while a phone is connected. Found by a follow-up lifecycle audit of the re-enable changes in this PR. |
||
|
|
43084873a6 |
nRF52 BLE: restore pairing passkey callback on re-enable (#11027)
shutdown() installs onUnwantedPairing to actively refuse pairing (used on the factory-reset / BT-disable teardown path), but the correct callback (onPairingPasskey) is installed only in setup(). Re-enabling BLE on an already-constructed nrf52Bluetooth goes through resumeAdvertising(), not setup() (main-nrf52.cpp), which only re-armed advertising and never restored the pairing callback. So a shutdown()->resumeAdvertising() cycle without a reboot left the device silently refusing all pairing until the next reboot. Restore the correct pairing passkey callback in resumeAdvertising(), guarded by the same config.bluetooth.mode check setup() uses. Only PIN modes drive a passkey-display callback; NO_PIN (Just Works) never invokes it, so no restore is needed there. Reachable today only via the PowerStress module's BT_OFF/BT_ON opcodes (normal runtime BT-disable paths reboot), so low severity, but a real teardown-latches / re-enable-doesn't-restore asymmetry. |
||
|
|
f7fd058308 |
Fix packet-pool slot leak in canned message destination picker (#11017)
updateDestinationSelectionList() allocated a MeshPacket via allocDataPacket() that was never sent or released, permanently consuming one packetPool slot every time the destination-selection picker was rebuilt. On non-PSRAM targets packetPool is a static 70-slot BSS pool, so repeated picker use exhausts it and eventually blocks all packet allocation (TX/RX failures). On PSRAM/portduino (MemoryDynamic) targets it is a true heap leak of ~424B per rebuild. The allocation and its two field writes (pki_encrypted, channel) were a copy/paste artifact of the PKI setup in sendText() and had no effect in this function. Remove the dead allocation. |
||
|
|
0902c13473 |
feat: build-flagged frame injection into the RX pipeline for testing (#11011)
MeshService::injectAsReceived (gated by MESHTASTIC_ENABLE_FRAME_INJECTION, off by default) extends the portduino SimRadio SIMULATOR_APP path to real hardware: a client-supplied frame, wrapped in a Compressed envelope on the SIMULATOR_APP portnum, is delivered through the real receive pipeline (router->enqueueReceivedMessage) as if it arrived off the LoRa chip. It therefore exercises from!=0 enforcement, channel/PKC decryption, remote-admin authorization, and hop/dedup/module dispatch - paths the toRadio API cannot reach (it forces from=0). from==0 is dropped to match real RX. This forges over-the-air traffic, so the flag must never ship enabled. Drive it with the meshtastic-mcp inject_frame tool / cli/meshinject.py. Documents the technique in the agent-facing copilot-instructions.md + AGENTS.md. |
||
|
|
c5355641d3 |
Add MEDIUM_TURBO modem preset (#10988)
* Protobufs * Wire up MEDIUM_TURBO modem preset MEDIUM_TURBO (500 kHz, SF9, CR 4/5) already existed in the protobuf enum but was never wired into firmware, so selecting it silently fell through to the LONG_FAST default and rendered an "Invalid" display name. Add its bw/sf/cr mapping (modemPresetToParams), display name (MediumTurbo/MedT), PRESETS_STD membership (standard regions only — 500 kHz does not fit EU868's 250 kHz band, so it stays out of PRESETS_EU_868 and is rejected/clamped there), and the MEDIUM SNR-grading bucket. Includes positive coverage in test_radio, EU868-reject + US-accept coverage in test_admin_radio and test_mesh_beacon, the STD preset count 9->10, an extended fuzz range, and the client-spec doc. * Address review feedback on MEDIUM_TURBO tests - test_mesh_beacon: assert has_mesh_beacon before checking the invalid preset was cleared, so the EU868-cleared test can't pass on a dropped message (matches the existing SHORT_TURBO test). - test_fuzz_packets: draw modem presets from _ModemPreset_ARRAYSIZE instead of a hard-coded 17 so the fuzz range tracks future enum additions automatically. |
||
|
|
8abae90838 |
Give the 4MB T3-S3 boards an app partition that fits (#10978)
* fix(device-install): flash littlefs at the offset from firmware metadata device-install.sh read the spiffs offset out of the .mt.json metadata into SPIFFS_OFFSET, but flashed the littlefs image at $OFFSET -- a separate variable still holding the hardcoded 0x300000 default. The metadata value was never used, so any board with a non-default partition table had its filesystem written to the wrong place. Unify on SPIFFS_OFFSET, and guard both metadata overrides so a table that omits ota_1 or spiffs falls back to the built-in default instead of handing esptool an empty offset. Quote both offsets at the call site: the spiffs one is newly consumed from jq, and a multi-line result would otherwise word-split into esptool's argv and silently mis-pair address with file. device-install.bat already does all of this correctly; this brings the shell script to parity. * fix(tlora-t3s3): give the 4MB T3-S3 boards an app partition that fits tlora-t3s3-v1 and tlora-t3s3-epaper have overflowed the shared 4MB partition-table.csv app slot (ota_0 = 0x250000 = 2,424,832 bytes) on every develop build since |
||
|
|
b3ddeca9d0 |
ci: nightly develop build published to github.io firmware-nightly/ (#10957)
Adds a scheduled (09:00 UTC) run of the CI build off develop that does everything a workflow_dispatch build does except create a GitHub release. Instead it refreshes a single, stable firmware-nightly/ folder on meshtastic.github.io with the current develop build, leaving that folder's hand-maintained release_notes.md untouched so it can be edited manually. On a nightly run (the schedule, or a manual nightly=true dispatch) only the firmware build + gather-artifacts + the new publish-nightly job run; the tests/docker/wasm/macOS/debian-src/size jobs and the three release jobs are skipped, so no release or tag is created. publish-nightly downloads the per-board artifacts, generates the firmware-<version>.json release manifest and an index.json version pointer (read by the web-flasher), then publishes to firmware-nightly/ via the same peaceiris/actions-gh-pages action used for releases. keep_files:false is scoped to destination_dir, so it clears stale nightly binaries while leaving sibling release folders untouched; the hand-maintained release_notes.md is fetched and carried forward first (fail-closed) so it is never clobbered. |
||
|
|
515fe8b94f | Clamp bw for malformed use_preset false with empty bandwidth | ||
|
|
168f669e61 |
esp32: fix NimBLE use-after-free crash when disabling bluetooth while connected (#10950)
* esp32: fix NimBLE use-after-free crash when disabling bluetooth while connected
Saving certain configs over BLE (e.g. MQTT module config) calls
disableBluetooth() -> NimbleBluetooth::deinit() while the phone is still
connected. deinit() called BLEDevice::deinit(true) straight away, which
deletes the BLEServer before nimble_port_stop(). Stopping the NimBLE
port with a live connection makes the host synthesize unsubscribe events
via ble_gatts_connection_broken() and dispatch them into the freed
BLEServer -> LoadProhibited panic in BLEServer::handleGATTServerEvent
(iterating server->m_notifyChrVec on freed memory).
Backtrace tail:
BLEServer::handleGATTServerEvent (SUBSCRIBE)
<- ble_gatts_subscribe_event <- ble_gatts_connection_broken
<- ble_gap_conn_broken <- ble_gap_rx_disconn_complete
It was reliably preceded by our onRead callback pinning the NimBLE host
task in its up-to-20s polling loop ('BLE onRead: timeout after 19920 ms,
4000 tries'), so the host task couldn't process the teardown until long
after the objects under it were freed.
Fix: drain before demolition in deinit(). Clear
onReadCallbackIsWaitingForData so an in-flight onRead returns
immediately; disconnect the peer and wait (bounded, ~2s) for the host
task to run onDisconnect (which clears nimbleBluetoothConnHandle) so the
unsubscribe/GATT teardown lands on the still-live server; only then
BLEDevice::deinit(true). isDeInit is raised after the drain (onDisconnect
early-returns when set and would otherwise never clear the handle), the
deferred advertising restart is dropped (the stack is going away), and
the now-dangling bleServer pointer is nulled.
deinit() runs on the main task, so waiting on the host task's callbacks
is safe; the waits are bounded regardless.
Upstream ordering bug (delete BLEServer before nimble_port_stop) to be
reported to espressif/arduino-esp32 separately.
* esp32: gate onRead during BLE teardown and use Throttle for drain wait
Addresses review feedback on the NimBLE deinit UAF fix:
- Set a bleDraining flag before disconnecting so onRead bails immediately
instead of arming the up-to-20s wait. Without this, a read that re-arms
after deinit() clears onReadCallbackIsWaitingForData could pin the NimBLE
host task, block onDisconnect from clearing the conn handle, and make the
2s drain expire. onRead now skips the wait when draining, and the wait
loop also observes the flag so an in-flight read is released.
- Use Throttle::isWithinTimespanMs() for the bounded drain wait instead of
a raw millis() subtraction, per coding guidelines.
* esp32: reset BLE teardown guards on re-init
deinit() latches bleDraining (new) and isDeInit (pre-existing) as teardown
guards, but setup() never cleared them. AdminModule's disable-bluetooth admin
command calls deinit() directly without a reboot, and PowerFSM can re-enable
via setBluetoothEnable(true) -> setup() on the same boot (isActive() is false
because deinit() nulls bleServer). Without resetting the flags, the re-initialized
stack has onRead permanently bailing on the drain path (empty reads) and
onDisconnect early-returning without clearing the connection handle.
Clear both flags at the top of setup() so a re-init recovers cleanly.
|
||
|
|
b4b1ece8f1 |
nrf52: fix loop-task stack overflow causing reset on BLE pairing/first-sync (#10944)
The Arduino loop task runs Meshtastic's entire cooperative OSThread scheduler on the Adafruit core's stock 4 KB (1024-word) stack. The 2.8 first-sync path - AdminModule::handleSetConfig -> saveChanges -> configChanged -> RadioInterface::reloadConfig -> LR11x0 reconfigure -> SPI Lock::lock, with vsnprintf/USB-CDC logging frames stacked on top - overflows it. SWD fault capture on a T1000-E (pyocd vector_catch=h) proved it: the loop task's SP was driven 40 bytes BELOW its own pxStack base, the stack paint was consumed to the floor, and the crash was a BusFault in xQueueSemaphoreTake dereferencing a semaphore handle that adjacent-heap corruption had overwritten with log text (0x3f3f207c = "| ??"). The core's HardFault handler is a bare NVIC_SystemReset, so in the field this presents as a silent reboot ~seconds after a phone pairs (or a 30 s zombie hang -> supervision timeout 0x8 when the corruption lands on task TCBs instead). Also explains the set_time_only -> immediate-NodeInfo crash variant reported on empty-DB nodes: same task, different deep chain. 2.7.26 is unaffected because the deep frames (XEdDSA, satellite map conversion, replay engine) did not exist. Fix: set -DLOOP_STACK_SZ=2048 (words = 8 KB) for all nrf52840 targets. Requires the #ifndef guard from meshtastic/Adafruit_nRF52_Arduino#7; until that merges the flag is a harmless redefinition warning. Validated on hardware: pairing + full sync completes, stack paint shows healthy margin under load. Also fix the boot-time reset-reason log: the core's init() caches RESETREAS and W1C-clears the register before setup() runs, so the raw read here has always printed 0. Use readResetReason() instead. (preFSBegin's RESETREAS==0 gate is intentionally left as-is - its degenerate GPREGRET-only behavior is what makes lfs recovery work.) |
||
|
|
9b5b2889e7 | Remove triage workflows | ||
|
|
8cff9e0ac3 | Setup renovate for nrf52 arduino platform package (#10946) | ||
|
|
8d20606203 |
feat(config): make position & telemetry broadcast opt-in (#10929)
Position sharing was opt-out (the default primary channel shipped position_precision=13) while device telemetry was already opt-in, and a normal firmware upgrade preserves saved config, so existing nodes stayed position-on. This makes both broadcasts opt-in, both on a fresh flash and via a one-time migration for upgrading nodes. - Fresh default: initDefaultChannel now sets position_precision=0. - One-time migration in loadFromDisk, gated on a dedicated POSITION_TELEMETRY_OPTIN_VER (26) watermark kept separate from DEVICESTATE_CUR_VER (bumping that would re-run the NodeDatabase v24 legacy decoder on already-migrated v25 DBs): disables position broadcast on PUBLIC/default-PSK channels and forces every device-telemetry mesh-broadcast flag plus the MQTT map-report location off. - Private-PSK channels are left untouched: a channel with a custom key is a deliberate trusted-group setup, so its configured precision is preserved. - Idempotent: ordinary saves never re-stamp .version, so a user who re-enables sharing is not re-disabled on the next boot. - Added pure channelFileUsesPublicKey() (operates on the raw ChannelFile, since the channels singleton isn't initialized during loadFromDisk) and refactored Channels::usesPublicKey() to delegate to it. - New test/test_optin_migration native suite (14 cases). |
||
|
|
97302f1845 |
Add Meshnology W10 AIOT Dev Kit (SX1262 via MCP23017 I2C expander) (#10911)
* Add software LoRa IRQ polling and an MCP23017 expander HAL for radios behind an I2C expander
Some boards route the SX126x control lines (NRESET/DIO1/BUSY) through an I2C
GPIO expander instead of native GPIOs, and don't wire the expander's /INT to
the MCU, so RadioLib cannot take a hardware DIO1 interrupt.
Two reusable pieces, both gated behind USE_MCP23017 / LORA_DIO1_SOFTWARE_POLL
so builds that don't define them are unaffected:
- ExtensionIOMCP23017 + MCP23017LockingArduinoHal: a RadioLib HAL that maps
virtual pins (MCP23017_VPIN_BASE .. +15) onto an MCP23017's GPIO, so
RESET/BUSY/DIO1 reads and writes become I2C transactions while the real SPI
GPIOs (SCK/MOSI/MISO/CS) pass straight through. A failed I2C read skips the
read-modify-write rather than clobbering the rest of the bank.
- LORA_DIO1_SOFTWARE_POLL: with no DIO1 interrupt available, the SX126x
interface polls the radio IRQ status register from the radio thread and
synthesizes the ISR_TX/ISR_RX events, filtering noisy preamble/header IRQs so
they can't starve TX. A 1 ms ISR_POLL_TICK drives the poll; TX timers may
overwrite the pending tick (pollMissedIrqs() is the backup that bounds the
latency of the busy-Rx contention window). Behaviour is unchanged on all
boards that keep the hardware interrupt.
* Add Meshnology W10 AIOT Dev Kit variant (SX1262 via MCP23017 expander)
ESP32-S3R8 + EBYTE E22-900MM22S (SX1262) + AXP2101 PMIC + Quectel L76KB GPS +
SPI TFT, with the radio's RESET/DIO1/BUSY and the LCD reset routed through an
MCP23017 I2C expander at 0x20. The expander's /INT is not wired to the ESP32,
so DIO1 uses the software poll. Pins come from the board schematic and the
vendor firmware; the expander is an MCP23017 despite the V1.1 schematic's stale
'TCA9555' label (the V1.2 placement diagram and all vendor code use the MCP23017
register map).
A local pins_arduino.h shadows the generic esp32s3 variant to drop RGB_BUILTIN,
which otherwise pulls the Arduino RMT HAL into the link and fails against this
build's trimmed FreeRTOS config.
Reports HardwareModel 140 (MESHNOLOGY_W10, already present in the protobufs).
Hardware-verified on a real W10 AIOT Dev Kit: AXP2101 PMU, MCP23017, SX1262
init + a full over-the-air TX/RX round trip through the expander, PCF85063 RTC,
SHT41 and QMI8658 auto-detected, and an L76KB GPS fix.
* meshnology-w10: enable ES8311 speaker for notification tones
Bring up the ES8311 codec (I2C 0x18) -> NS4150 amp -> speaker so the board can
play notification tones / ringtones over the I2S buzzer path (the
use_i2s_as_buzzer external-notification option). Codec2 voice stays out of
scope: it is SX1280-only and this is a sub-GHz board.
- variant.h: HAS_I2S + DAC_I2S pins (MCLK=1, BCK=2, WS=4, DOUT=5, DIN=3)
- platformio.ini: arduino-audio-driver + ESP8266Audio + ESP8266SAM
- extra_variants/meshnology_w10/variant.cpp: lateInitVariant() configures the
ES8311, mirroring the other Meshtastic ES8311 boards. Kept codec-only (no
main.h) so the audio driver's 'using namespace audio_driver' does not pull a
conflicting GpioPin into scope alongside Meshtastic's class GpioPin.
- AudioThread.h: toggle the NS4150 amp (MCP23017 EXIO_PA_CTRL) around playback.
Verified on hardware: a test tune played audibly through the speaker.
* meshnology-w10: address review feedback on the MCP23017 driver
- ExtensionIOMCP23017: serialize register access with a mutex, since the radio
HAL (BUSY/DIO1/RESET) and AudioThread (amp enable) now reach the expander from
different threads and the read-modify-write paths are not atomic.
- digitalRead: return a fail-safe HIGH on a failed read instead of LOW, so a
transient I2C error on the LoRa BUSY line can't look like 'ready' and let
RadioLib start an SPI transaction early. (DIO1 is polled via the radio IRQ
register, not this pin.)
- enablePinChangeInterrupt: use the checked readReg() overload and skip on a
failed read, matching the other read-modify-write helpers.
- meshnology_w10 variant.cpp: log a warning if an ES8311 register write NACKs
instead of silently leaving the codec half-configured.
- RadioInterface.cpp: guard the MCP23017 HAL branch with ARCH_ESP32 to match the
include and driver-file guards.
* Replace board-model audio/sleep ifdefs with reusable variant capability macros
Two shared-code spots keyed off specific hardware models; move the board-specific
detail into opt-in macros the variants define, so the core code stays generic and
new boards can opt into the behavior class without touching shared files.
- AudioThread amp control: drop the T_LORA_PAGER / MESHNOLOGY_W10 ifdefs. A board
with an I2S amp now defines AUDIO_AMP_ENABLE(on) in its variant.h (T-LoRa Pager
and Meshnology W10 do), and AudioThread just calls it around playback. The
expander includes are keyed on USE_XL9555 / USE_MCP23017 (capabilities) instead
of the model name.
- Light-sleep / DIO1 wakeup: drop the SENSECAP_INDICATOR check. Boards whose
LORA_DIO1 is an expander pin (not an ESP32 GPIO) define LORA_DIO1_EXTENDED_IO
(SenseCAP Indicator and Meshnology W10), and sleep.cpp keys off that for both
the light-sleep skip and the DIO1 GPIO-wakeup skip. This also fixes a latent
issue on the SenseCAP: it previously reached the GPIO-wakeup path and called
gpio_pulldown_en() on a virtual expander pin; it now skips that cleanly.
Builds verified on meshnology_w10, tlora-pager, and seeed-sensecap-indicator.
* meshnology-w10: silence cppcheck constParameterPointer on the ISR-callback helper
isIsrTxCallback() takes a function pointer it only compares; cppcheck's
constParameterPointer wants it 'pointer to const', which is meaningless for a
function pointer (the codebase already globally suppresses the sibling
constParameterCallback). Inline-suppress it to keep the check green.
* HopScaling: qualify the member assignment as this->count
cppcheck (CI's version) flags 'count = newCount;' at the end of trimIfNeeded()
as uselessAssignmentArg ('assignment of function parameter has no effect') - a
false positive, since count is a member that outlives the call. Write it as
this->count (matching the Step-1 assignment above) so the analyzer sees a member
write. This finding comes in via the develop merge, not this board; fixing it
here to unblock the PR's cppcheck check.
|
||
|
|
f250cf3cdc | Update protos | ||
|
|
24c1dccf00 |
CI: track RAM (.data+.bss) in size reports and gate on per-env budgets (#10899)
* CI: track RAM (.data+.bss) in size reports and gate on per-env budgets
The 2.8.0 nRF52840 heap regression (99% heap in field reports) shipped
invisibly because CI only tracked flash. On nRF52840 the heap arena is the
linker gap after .bss, so every byte of static .data+.bss growth shrinks the
usable heap 1:1 - RAM needs the same guardrails flash already has.
What's added:
- bin/platformio-custom.py emits ram_bytes (.data + .bss from the ELF, via
the toolchain size tool) into each .mt.json manifest. Heap/stack
placeholder sections are deliberately excluded.
- bin/collect_sizes.py records {flash_bytes, ram_bytes} per env;
bin/size_report.py grows RAM and RAM-delta columns in the PR size report.
Older artifacts without ram_bytes (and legacy int-schema baselines)
degrade to "n/a" instead of crashing.
- bin/ram_budgets.json: per-env RAM/flash budgets, enforced only for envs
listed there. Seeded with rak4631: ram 113,000 (current 110,948 + ~2 KB
slack), flash 786,000 (current 765,192 + ~20 KB; the app region is
0x27000..0xEA000 = 798,720 and the image must stay clear of the
warm-store ring guard in extra_scripts/nrf52_warm_region.py).
- New size-budget-gate CI job runs size_report.py --enforce-budgets and
fails the build on violation; the informational firmware-size-report job
now also renders budget usage into the PR comment.
- src/main.cpp: opt-in boot heap watermark (-DMESHTASTIC_HEAP_WATERMARK_CHECK)
logs LOG_ERROR when less than 20% of the heap is free at the end of
setup(). Off by default; skipped on platforms without heap accounting.
How budgets are raised: deliberately, never automatically. If a change needs
more headroom, bump the env's limit in bin/ram_budgets.json in the same PR
and justify the increase in the PR description.
Verified: python3 bin/test_size_scripts.py (23/23 pass, including ram_bytes
parsing, n/a fallback, and over/under/missing-env budget-gate cases).
* Address review: fail the budget gate closed, fix RAM section matching
- size-budget-gate workflow: drop continue-on-error on the manifest
download and the empty-dir fallback, so the job fails when the data it
gates on cannot be fetched.
- size_report.py --enforce-budgets now fails closed on every
missing-data path instead of trivially passing: empty collected sizes,
a budgeted env that was not built, or a manifest without the budgeted
metric. Report-only mode keeps rendering those as n/a.
- load_budgets() rejects zero/negative/non-integer budgets with a clear
error (a typo'd budget could previously crash budget_markdown with
ZeroDivisionError or silently skip the check); the percentage render
keeps a defensive guard for direct callers.
- compute_ram_bytes(): count RISC-V small-data sections (.sdata/.sbss)
and exclude ESP-IDF .rtc.* sections, which live outside the
heap-competing SRAM.
- Trim the heap-watermark comment in main.cpp to two lines.
bin/test_size_scripts.py: 27/27 - the two fail-open assertions are
flipped to fail-closed, with new report-only counterparts plus cases for
empty sizes under enforcement and malformed budgets.
|
||
|
|
12b2d973a4 |
Add central memory-class ladder (MemClass.h) with fail-safe-small defaults (#10901)
* Right-size nRF52 heap tiers after 2.8.0 heap-exhaustion field reports Field reports on 2.8.0 show nRF52840 devices at 99% heap (114/115 KB) within minutes of boot; operator new asserts on OOM, so these devices are one allocation from a reboot. The 2.8.0 cache sizing ladders gave nRF52 the largest non-PSRAM tiers on the assumption that a BLE-only part has a roomy heap - the arena is actually ~125 KB shared with the FreeRTOS task stacks. Per-target retiers (nRF52840 unless noted): - Traffic Management cache 1000 -> 250 entries (10 KB -> 2.5 KB); the unclassified fallthrough drops 1000 -> 400 to match the classic-ESP32 tier (also affects RP2040/RP2350) - Warm node store 200 -> 100 entries (8 KB -> 4 KB); the non-XXAA fallthrough drops 320 -> 100 so an unclassified RAM-constrained part can't boot-allocate 12.8 KB - MESSAGE_HISTORY_LIMIT 20 -> 10 (text pool 4.4 KB -> 2.2 KB), the tier classic ESP32 already ships - MAX_RX_TOPHONE 32 -> 16, shrinking the static packet pool 70 -> 54 slots (~6.6 KB of .bss returned to the heap arena) - PacketHistory hash index off arch-wide (1 KB); O(n) over 240 records is negligible at LoRa packet rates - OLEDDISPLAY_REDUCE_MEMORY arch-wide (~1 KB OLED back buffer); the five TFT variants -U it because TFTDisplay.cpp needs buffer_back for dirty-window diffing - Drop the stale "for testing" 1024-entry TMM override on T1000-E Measured on rak4631: heap arena grows 124,572 -> 131,180 B and boot allocations drop ~15.7 KB, roughly +22 KB free heap on the field-report device class. Migration: the nRF52840 warm flash ring replays through place() (LRU), so the newest 100 identities survive the shrink; the file backend rejects oversized snapshots cleanly (new test covers this). Native suites pass (536/536 Docker, 13/13 native-macos warm store); rak4631, heltec-mesh-node-t114 (TFT) and tracker-t1000-e build green. * Add central memory-class ladder (MemClass.h) with fail-safe-small defaults The 2.8.0 nRF52840 heap exhaustion happened because each RAM-sized cache picked its per-platform tier from its own chip #ifdef ladder, and every ladder's fallthrough default was its largest non-PSRAM tier - nRF52 was never named, so it silently got 1000-entry caches on a ~115 KB arena. This introduces src/memory/MemClass.h: a single MESHTASTIC_MEM_CLASS (TINY / SMALL / MEDIUM / LARGE) ranked by usable app heap after platform overheads, with the deliberate property that an unclassified chip lands in SMALL - a new target boots with small caches until someone opts it up in one visible place. The TMM cache, warm store, MAX_RX_TOPHONE and MAX_SATELLITE_NODES ladders in mesh-pb-constants.h now key off the class; branches pinned by something other than RAM stay explicit and say why (nRF52840's SoftDevice arena, RP2040's warm.dat watchdog bound). MAX_NUM_NODES intentionally stays separate - it is flash-shaped (nodes.proto vs LittleFS), not heap-shaped. A per-class MESHTASTIC_BOOT_CACHE_BUDGET static_assert now covers the three big boot-allocated caches, so the next cache-adding PR that would blow a small platform's budget fails to compile instead of exhausting heap in the field. No values change for any existing target: rak4631, tbeam, rak11310 and wio-e5 build byte-identical before/after; all ladders remain #ifndef-guarded so variant overrides keep working. * Address review: fix RP2350 class-table doc, add PacketRecord static_assert - MemClass.h's class table claimed RP2350 was MEDIUM while the mapping ladder classifies it SMALL (with RP2040) - the table now matches the ladder, with a note that RP2350 is a MEDIUM candidate whenever someone wants to tune it up (kept SMALL here so this header stays a behavioral no-op). - The boot-cache budget comment referenced a static_assert pinning PacketHistory::PacketRecord at 20 B that did not exist (only a layout comment). Add the real static_assert so the budget math in mesh-pb-constants.h fails to compile if the record layout changes. Also merges develop (the base #10898 landed there as a squash, which is what made this stacked branch conflict); develop's mesh-pb-constants.h is byte-identical to this branch's base, so the resolution keeps the MemClass ladder unchanged. rak4631 and wio-e5 build green; test_packet_history 47/47. * Address review: share PACKETHISTORY_MAX, trim policy comments - Hoist PACKETHISTORY_MAX from PacketHistory.cpp into mesh-pb-constants.h (next to the MAX_NUM_NODES it derives from) so the constructor clamp and the boot-cache budget static_assert use one definition instead of hand-mirrored arithmetic that could drift. The expression stays valid where MAX_NUM_NODES resolves at runtime (ESP32-S3, portduino); the pointless 2.0 double math becomes integer. - Trim the MemClass.h header (36 -> 16 comment lines) and the budget / sizing-policy comments per the repo comment-length guideline, keeping the class table, the fail-safe-small rule, the override mechanism, and the include-order constraint. rak4631 (compile-time MAX_NUM_NODES) and heltec-v3 (runtime) build green; test_packet_history 47/47. |
||
|
|
00685a4e61 |
nrf52840: right-size the SoftDevice RAM reservation (+8 KB heap arena) (#10903)
The app RAM ORIGIN in both nrf52840 linker scripts has been a hard-coded
0x20006000 (24 KB reserved for S140) since they were introduced, but
sd_ble_enable() with our fixed Bluefruit configuration (one peripheral
link, BANDWIDTH_MAX / ATT MTU 247, 0x1000 attribute table) requires
substantially less. The SoftDevice cannot use the gap and the app image
is linked above it, so every byte between the true requirement and the
ORIGIN is simply unusable RAM - on a part where 2.8.0 field reports
showed the heap arena at 99% use.
Lower the ORIGIN to 0x20004000 in both s140 v6 and v7 scripts. The heap
arena is the linker gap on this platform, so the change is worth exactly
+8,192 B of free heap on every nRF52840 board (verified: rak4631 .heap
section 124,572 -> 132,764 B; both v6- and v7-script boards link with
.data at 0x20004000).
Safety net: Bluefruit.begin()'s return value - previously discarded -
is now checked. If a future SoftDevice or Bluefruit config change raises
the requirement past the reservation, the node logs a critical error
instead of silently running without BLE, with instructions to re-measure
via CFG_DEBUG=1 ("SoftDevice's RAM requires: 0x...").
|
||
|
|
6c7ee8afc7 |
Add MemAudit: per-subsystem heap accounting in the boot log (#10900)
* Add MemAudit: per-subsystem heap accounting in the boot log The 2.8.0 nRF52840 heap-exhaustion field reports had to be diagnosed by hand, reconstructing each subsystem's heap footprint from source and build flags one report at a time. This makes every future report self-diagnosing from the serial log: a tiny fixed-size registry (src/memory/MemAudit.*) that big long-lived allocations report into, printed as one line at the end of setup() and alongside the periodic "Heap free:" log: MemAudit[boot]: tmm=2500 warm=4000 pkthist=5824 nodedb=13440 msgstore=2200 pktpool(live)=3270 total=31234 Instrumented: NodeDB hot vector (nodedb) + satellite maps (satmaps, rb-tree overhead estimated), WarmNodeStore (warm), PacketHistory records and hash index (pkthist), TrafficManagement caches (tmm/tmm_ni), MessageStore text pool (msgstore), TFT line/repaint buffers (display), and live in-flight packets (pktpool(live)) via an optional audit tag on the packet pool allocator - the one hot path, counted with a relaxed 32-bit atomic add (single instructions on Cortex-M, no locks). Cost: 128 B RAM for the 16-tag table, well under 1 KB flash on rak4631. MESHTASTIC_MEM_AUDIT=0 compiles it out to inline no-op stubs (call sites need no ifdefs); STM32WL, the tightest flash target, defaults off. New native suite test_mem_audit covers add/set/snapshot arithmetic, tag reuse (pointer and cross-TU strcmp fallback), null/unknown tags, and table-full behavior; test/native-suite-count bumped to 28. * native-wasm: add src/memory/ to the curated source filter The wasm env denies all sources and adds an explicit file list; MemAudit callers (main, MeshService, NodeDB, PacketHistory) are in that list but src/memory/MemAudit.cpp was not, so wasm-ld failed on undefined memaudit:: symbols. |
||
|
|
ed03a69555 |
Right-size nRF52 heap tiers after 2.8.0 heap-exhaustion field reports (#10898)
* Right-size nRF52 heap tiers after 2.8.0 heap-exhaustion field reports Field reports on 2.8.0 show nRF52840 devices at 99% heap (114/115 KB) within minutes of boot; operator new asserts on OOM, so these devices are one allocation from a reboot. The 2.8.0 cache sizing ladders gave nRF52 the largest non-PSRAM tiers on the assumption that a BLE-only part has a roomy heap - the arena is actually ~125 KB shared with the FreeRTOS task stacks. Per-target retiers (nRF52840 unless noted): - Traffic Management cache 1000 -> 250 entries (10 KB -> 2.5 KB); the unclassified fallthrough drops 1000 -> 400 to match the classic-ESP32 tier (also affects RP2040/RP2350) - Warm node store 200 -> 100 entries (8 KB -> 4 KB); the non-XXAA fallthrough drops 320 -> 100 so an unclassified RAM-constrained part can't boot-allocate 12.8 KB - MESSAGE_HISTORY_LIMIT 20 -> 10 (text pool 4.4 KB -> 2.2 KB), the tier classic ESP32 already ships - MAX_RX_TOPHONE 32 -> 16, shrinking the static packet pool 70 -> 54 slots (~6.6 KB of .bss returned to the heap arena) - PacketHistory hash index off arch-wide (1 KB); O(n) over 240 records is negligible at LoRa packet rates - OLEDDISPLAY_REDUCE_MEMORY arch-wide (~1 KB OLED back buffer); the five TFT variants -U it because TFTDisplay.cpp needs buffer_back for dirty-window diffing - Drop the stale "for testing" 1024-entry TMM override on T1000-E Measured on rak4631: heap arena grows 124,572 -> 131,180 B and boot allocations drop ~15.7 KB, roughly +22 KB free heap on the field-report device class. Migration: the nRF52840 warm flash ring replays through place() (LRU), so the newest 100 identities survive the shrink; the file backend rejects oversized snapshots cleanly (new test covers this). Native suites pass (536/536 Docker, 13/13 native-macos warm store); rak4631, heltec-mesh-node-t114 (TFT) and tracker-t1000-e build green. * Drop stale OLEDDISPLAY_REDUCE_MEMORY -U on t114 / mesh-solar-tft Only USE_TFTDISPLAY variants (t1, t096, wismeshtap) compile TFTDisplay.cpp and need the lib's buffer_back; t114 and mesh-solar-tft render through the meshtastic-st7789 driver, which handles the reduced-memory configuration fine - as #10894 (merged from develop) already established by defining the flag there. Remove the -U guard and the per-variant -D (redundant with the arch-wide define in nrf52_base on this branch). Both variants verified building. |
||
|
|
b4dd76a4db |
Harden against crafted-packet crashes + adversarial fuzzing (#10862)
Audit and fuzzing of the RF-packet decode -> dispatch -> display/phone paths for the "crash a node or phone with a crafted packet" surface, beyond the XEdDSA authenticity work. Crash fixes (reproduced under AddressSanitizer / UBSan): - GeoCoord::latLongToUTM/latLongToMGRS read fixed letter tables out of bounds on extreme latitude_i/longitude_i from a received Position, and narrowed out-of-range easting/northing doubles to unsigned (float-cast-overflow UB). Clamp the UTM zone, the easting/northing narrowing, and the band/col/row indices. Regression: test_geocoord_extreme_coords_no_oob. - EnvironmentTelemetry/AirQualityTelemetry render attacker floats via String(float), which on nRF52/RP2040/STM32/portduino formats into a fixed char[33] (dtostrf) and overflows near FLT_MAX. Clamp the rendered metrics via UnitConversions::displaySafeFloat (finite + magnitude <= 1e9), unit-tested in test_type_conversions. Defense-in-depth + robustness: - TraceRouteModule::printRoute: fix an snr_back[-1] OOB read (wrong count in the guard) and stop formatting the INT8_MIN "unknown SNR" sentinel as a dB value. - WaypointModule/NodeDB: sanitize untrusted strings before the OLED renderer and the phone-facing ClientNotification (belt-and-suspenders vs PB_VALIDATE_UTF8). - MeshService::sendToPhone: withhold NODEINFO/WAYPOINT packets whose nested string won't cleanly decode, protecting strict phone protobuf decoders without affecting mesh relay. Tests: new test_fuzz_decode (protobuf decode + UTF-8 sanitizer fuzz) and test_fuzz_packets (perhapsDecode / module-handler / traceroute / phone-gate fuzz), all under AddressSanitizer; native-suite-count 25 -> 27. Full suite 515/515 green. |
||
|
|
510e9796f9 |
Extract mcp-server to its own repo (meshtastic/meshtastic-mcp) (#10861)
The Python MCP server + hardware test harness that lived under mcp-server/ now has its own home at https://github.com/meshtastic/meshtastic-mcp (published, versioned independently). Remove the in-tree copy and wire the firmware repo to the standalone server externally. - Delete mcp-server/ (96 files) and the 8 harness-coupled AI workflow files under .claude/commands/ and .github/prompts/ that drove ./mcp-server/ run-tests.sh — those workflows now ship with meshtastic-mcp as skills. - .mcp.json: register the server via `uvx --from git+https://github.com/meshtastic/meshtastic-mcp meshtastic-mcp` instead of a local ./mcp-server/.venv, keeping MESHTASTIC_FIRMWARE_ROOT="." so the MCP tools still work from this checkout with no local build. - Repoint the remaining references (AGENTS.md, CLAUDE.md, .github/copilot-instructions.md, bin/regen-*.sh, docs, Screen.h, userPrefs.jsonc, test/fixtures/nodedb/README.md, .trunk/configs/.bandit) at the standalone repo. The MCP tool surface is unchanged — only the pytest harness moves out; run it from a meshtastic-mcp checkout with MESHTASTIC_FIRMWARE_ROOT pointed here. No build/CI/platformio coupling existed, so nothing in the firmware build changes. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0e84c1a827 |
Harden XEdDSA unsigned-packet policy and add coverage (#10858)
Audit of the XEdDSA packet-signing implementation (#10478) surfaced several issues in when unsigned packets are accepted on receive or emitted on send. This fixes them and adds regression coverage. - Unicast NodeInfo exchange no longer breaks against signer nodes: the NodeInfoModule downgrade drop is gated to broadcasts, since senders never sign unicast (want_response replies, directed exchanges). - Replace the payload-size sign heuristic with an exact encoded-size gate (signedDataFits) and mirror it on the receive side, removing a dead band where 167-168 B broadcasts were signed then failed TOO_LARGE. - Extract the receive policy into checkXeddsaReceivePolicy() and apply it to plaintext-MQTT decoded downlink, which previously skipped signature verification and downgrade protection entirely. - Reject signatures whose length is neither 0 nor 64 as malformed, so a crafted partial signature can't inflate the size estimate and dodge the unsigned-downgrade drop. - Hold cryptLock on the MQTT verify path (shared Ed25519 key cache). - Clear any client-preset signature on packets we originate, on all builds. - Randomized (hedged) signing per the Signal XEdDSA spec: bump the meshtastic/Crypto pin to the build where XEdDSA::sign mixes 32 bytes of caller randomness into the nonce as Z (meshtastic/Crypto#3), and seed those bytes in xeddsa_sign from HardwareRNG (checked, with a seeded-CSPRNG fallback). test_crypto pins that repeated signs differ and both verify. Adds test coverage: test_packet_signing groups A-E (receive matrix, send policy, NodeInfo backstop, encoding invariants, decoded-ingress policy), test_mqtt end-to-end downlink cases, and a test_crypto randomization check. |
||
|
|
23bb55d965 |
chore(deps): forward-port LovyanGFX + audio-driver renovate bumps from master (#10826)
Two more renovate updates landed on master (2.7) after #10824; reconcile develop to their values (value reconcile, consistent with #10824): - LovyanGFX 1.2.21 -> 1.2.24 (#10783) -- 18 lovyan03 custom.pio pins - pschatzmann arduino-audio-driver v0.2.1 -> v0.3.0 (#10770) -- tlora-pager, tracksenger Left untouched (deliberate develop-only pins; master's renovate didn't bump them): - lovyan03/LovyanGFX@1.2.19 (single variant pin) - mverch67/LovyanGFX fork (seeed-sensecap-indicator variant) |
||
|
|
d8884125b6 |
chore(deps): forward-port renovate updates from master to develop (#10824)
Reconciles develop to master's latest renovate values for deps bumped on the 2.7 line but never back-merged, so the develop->master 2.8 promotion (#10777) doesn't regress them. Done as a value reconcile, not a cherry-pick: several master commits are superseded (5 device-ui bumps, 2 ststm32 bumps), and develop's stale archive/refs/tags/ URL form actually blocked renovate from bumping these (the regex expects archive/<version>.zip). GitHub Actions: - actions/checkout v6 -> v7 (35 refs) - actions/cache v5 -> v6 - actions/github-script v8 -> v9 (3 refs) - actions/stale v10.2.0 -> v10.3.0 Build / platform: - alpine 3.23 -> 3.24 (alpine.Dockerfile) - platformio/ststm32 19.5.0 -> 19.7.0 - platformio/nordicnrf52 10.11.0 -> 10.12.0 Libraries: - Adafruit SSD1306 2.5.16 -> 2.5.17 - SparkFun MMC5983MA v1.1.4 -> v1.1.5 - Sensirion I2C SCD30 1.0.0 -> 1.1.1 - meshtastic esp8266-oled-ssd1306 6bfd1f1 -> 2e26010 Deliberately excluded: - meshtastic/device-ui digest: coupled to firmware protobuf/API and develop has diverged hard (NodeDB v25); left for a separate maintainer bump + visual check. - libpax: develop uses the mverch67 fork (pinned by Arduino-3.x migration #9122), master uses dbinfrago -- a fork divergence, not a version bump; not reconciled. - platform-native digest: develop already at 61067ac (equal to master). |
||
|
|
8345c2b6a9 | Protos | ||
|
|
0488a46a3c |
fix: stop unexpected NodeNum regeneration from PKI key loss (#10808)
* fix: stop unexpected NodeNum regeneration from PKI key loss A node's NodeNum is derived from its PKI public key (my_node_num == crc32(public_key)), so it changes only when the keypair is regenerated -- which happens only when a keygen path runs while config.security.private_key.size != 32. Two paths could trigger that without the user intending a new identity: 1. Boot: loadFromDisk() collapsed every non-success loadProto() result for the config file into installDefaultConfig() (preserveKey=false), which memset()s config and wipes the private key. loadProto() distinguishes DECODE_FAILED (file present but undecodable/undecryptable) from OTHER_FAILURE (absent), so a transient/corrupt read silently minted a new identity. A new configDecodeFailed flag now gates the boot keygen and the boot config-save directly (not via region, so it survives the userprefs/region overrides that run later in loadFromDisk): identity is frozen, the unreadable file is left intact for a later clean boot to recover, and the node boots radio-silent (region UNSET, TX off). A genuinely-absent config still gets defaults + keygen. 2. AdminModule security SET: config.security = c.payload_variant.security wholesale-clobbered the keypair, then regenerated whenever the incoming private_key.size != 32. A client editing an unrelated security field without round-tripping the private key would re-mint identity. The existing keypair is now preserved when the incoming config omits it; first provisioning and key import are unchanged. Intentional reset goes through factory_reset. Native PlatformIO unit suite (Docker): 499/499 test cases pass. * test: cover security keypair preservation; clarify load-failure comment Addresses PR review feedback: - Add native unit tests (test_admin_radio) asserting handleSetConfig preserves the existing identity keypair when a security SET omits the private key, and applies a full supplied keypair (key import). Guards against reintroduced NodeNum/keypair regeneration. AdminModuleTestShim (now a friend) defers saves so the lightweight harness doesn't saveToDisk an uninitialized node database. - Clarify the non-DECODE_FAILED config-load comment: OTHER_FAILURE / NO_FILESYSTEM cover an absent or unopenable file, not just first boot. |
||
|
|
0baf8b1db6 |
Portduino WASM hello world (still experimental) (#10793)
* Add ARCH_PORTDUINO_WASM build: meshtasticd in WebAssembly over WebUSB
Compile the full portduino firmware to WebAssembly (Emscripten) so a real node runs in a browser tab or headless Node, driving a LoRa radio over WebUSB through a CH341 — the desktop Ch341Hal path with its libusb backend swapped for a WebUSB one. The native/desktop portduino build is unchanged.
New build env under src/platform/portduino/wasm/ (excluded from the native build_src_filter): WebUSB libpinedio backend, config/FS/region/MAC/PhoneAPI glue, wasm setup/loop, JS WebUSB runtime, and build stubs. bin/build-portduino-wasm.sh runs a standalone cached emcc build to build/wasm/meshnode.{mjs,wasm}.
Six firmware sources gain #ifdef ARCH_PORTDUINO_WASM guards (single-threaded cooperative emscripten_sleep loop, continuous RX, US region default, std RNG, no-popen exec); none affect non-wasm builds.
* PortDuino WASM: CI build job, cross-platform libdeps, configurable adapter
bin/build-portduino-wasm.sh: auto-detect native-macos (macOS) vs native (Linux/CI) libdeps with a NATIVE_ENV override, and use the SAME env's Crypto with an XEdDSA-present guard. Drops the heltec-v3/Crypto borrow — the meshtastic/Crypto pin already ships XEdDSA; the native libdeps cache was just stale. No env pin change.
Add .github/workflows/build_portduino_wasm.yml: build the ARCH_PORTDUINO_WASM target in CI (ubuntu + emsdk, native libdeps) so it can't silently bit-rot; asserts build/wasm/meshnode.{mjs,wasm}.
src/platform/portduino/wasm: add wasm_set_lora_* setters so the JS host can configure any CH341 LoRa adapter (module, USB ids, DIO/TCXO, SPI speed, pins); wasm_config_apply falls back to the MeshToad defaults when unset.
* WASM: build as a first-class [env:wasm] via platform-wasm (retire emcc script)
Replace the standalone emcc build (bin/build-portduino-wasm.sh) with a normal
PlatformIO env, `pio run -e wasm`, using the new meshtastic/platform-wasm
platform (emcc/em++ + Asyncify, the WASM sibling of platform-native). The
portduino WebAssembly node is now built the same way as every other target.
- variants/native/portduino/platformio.ini: add [env:wasm] (platform pinned to
platform-wasm, board wasm). Translates the script's source set into a curated
build_src_filter, the ~30 EXCLUDE_* defines, and the lib set; defines
ARCH_PORTDUINO_WASM in-repo so correctness doesn't hinge on the platform's
board.json.
- extra_scripts/wasm_link_flags.py: the firmware-specific emcc *link* settings
(EXPORT_NAME, EXPORTED_RUNTIME_METHODS, EXPORTED_FUNCTIONS, ASYNCIFY_IMPORTS).
PlatformIO feeds build_flags to compile only, so these must ride LINKFLAGS;
without it the WebUSB Asyncify seam and the JS host's runtime methods are
dropped (the _wasm_* exports survive only via EMSCRIPTEN_KEEPALIVE).
- PortduinoGlue.{h,cpp}: guard the yaml-cpp dependency out of the WASM build
(#ifndef ARCH_PORTDUINO_WASM around the include, emit_yaml/loadConfig/
readGPIOFromYaml). The browser node configures via the wasm_set_lora_* setters
and dead-strips the YAML path; this drops the host yaml-cpp build dependency
entirely. Native is unchanged (guards are inert there).
- portduino_glue_wasm.cpp / portduino_main_wasm.cpp: repair EM_ASM JS that a
formatter had mangled (!== -> != =, regex split) in the prior landing; the
emcc link succeeds regardless, so CI now runs `node --check meshnode.mjs`.
- .github/workflows/build_portduino_wasm.yml: build via `pio run -e wasm`
(artifacts under .pio/build/wasm/), trigger on the shared inputs the env
inherits (root platformio.ini, bin/platformio-*.py).
- NodeDB.cpp: drop the dead ARCH_PORTDUINO_WASM region-default branch (region
now defaults the same as native).
- Crypto renovate pins: add the missing gitBranch so they track upstream.
Output: .pio/build/wasm/meshnode.{mjs,wasm} (ES module, factory createMeshNode).
Verified: pio run -e wasm (against the published platform archive), node --check,
module instantiates in Node with all exports; native-macos + Docker native unit
tests (450/450) still pass.
* Fix name on the Piggystick
* wasm: pin platform-wasm at the GPL-3.0-relicensed commit
platform-wasm's LICENSE was always GPLv3 (matching this firmware), but its
platform.json/README still declared Apache-2.0 (mis-copied from platform-native).
That's fixed upstream in b83fa5b; bump the [env:wasm] pin to it. Build output is
unchanged (license metadata only). Verified: pio run -e wasm against b83fa5b.
* wasm: make reboot() actually restart the node (was a no-op)
In wasm the reboot path is live (main.cpp -> Power::powerCommandsCheck ->
Power::reboot), but Power::reboot's ARCH_PORTDUINO arm tore down SPI/Wire/Serial
and then called the no-op ::reboot() stub — leaving the node running with a dead
radio until the tab was manually reloaded. Triggers include an admin/phone
reboot, factory reset, the "reconfigure failed" path, and the 60 s stuck-TX
hardware watchdog (RadioLibInterface).
- Power::reboot(): add an ARCH_PORTDUINO_WASM arm (before ARCH_PORTDUINO, since
the wasm build defines both) that skips the host teardown and just calls
::reboot(). notifyReboot already let modules persist.
- ::reboot() (glue): hand off to the JS host — browser reloads the tab (NodeDB
state survives via IDBFS, same identity returns); headless calls Module.onReboot
if provided, else logs. Loose !=/== so clang-format doesn't mangle the EM_ASM JS.
- README: document the reboot handoff + the Module.onReboot hook.
Verified: pio run -e wasm + node --check (EM_ASM intact); native-macos unaffected.
* wasm: rename env to native-wasm and run it in the main CI matrix
Rename [env:wasm] -> [env:native-wasm] for consistency with the portduino
native family (native, native-macos, native-tft). The build dir follows to
.pio/build/native-wasm/ (artifact is still meshnode.{mjs,wasm}); the PIOENV
guard in extra_scripts/wasm_link_flags.py, the README, and the companion wrapper
move with it. The board stays `wasm`.
Also wire the build into normal CI: build_portduino_wasm.yml becomes a reusable
workflow (workflow_call) invoked as the `build-wasm` job of main_matrix.yml, so
the WebAssembly node is built like every other platform instead of on a separate
path trigger.
* native-wasm: auto-locate the Emscripten SDK (pre-build script)
`pio run -e native-wasm` failed with "emcc not found" whenever it was invoked
from a shell that hadn't sourced emsdk_env.sh — a VS Code task, an IDE build
button, a bare terminal. Add a pre: extra script that probes the usual emsdk
locations ($EMSDK_ENV, $EMSDK, ~/emsdk, ./.emsdk, the sibling companion
checkout), sources emsdk_env.sh, and imports the resulting environment so the
platform builder and emcc see PATH/EMSDK/EM_CONFIG. No-op when emcc is already
reachable (CI), silent when no SDK is found (the platform emits its own error).
* wasm: address PR review feedback
- js/bridge.js: import CH341 from "./ch341.js" (sibling in this layout), not
"../src/ch341.js" which doesn't resolve here.
- js/ch341.js: a zero-length transferIn while MISO bytes are still outstanding
now throws instead of breaking out with a partially-filled buffer — silent SPI
corruption becomes a loud error, matching the comment above it.
- libpinedio_webusb.c: webusb_set_auto_cs honors the AUTO_CS option (? 1 : 0)
instead of the always-on ? 1 : 1. Runtime behavior is unchanged — Ch341Hal sets
AUTO_CS=0 right after pinedio_init (RadioLib drives the active-low NSS); the
option just isn't set yet at init, so this now correctly defaults off.
- SX126xInterface.cpp: the RX-start error log now names the method actually
called (startReceive vs startReceiveDutyCycleAuto) instead of hardcoding the
duty-cycle name in the WASM branch.
* native-wasm: drop the emsdk bootstrap shim (now in platform-wasm)
The Emscripten SDK auto-location moved into the platform-wasm builder, so the
firmware no longer needs its own pre: extra script. Remove
extra_scripts/wasm_emsdk_env.py and bump the platform pin to the build that
carries the bootstrap. The wasm_link_flags.py post script stays — those exported
fns / runtime methods / Asyncify import seam are firmware-app-specific.
* wasm: use the canonical companion name (meshtasticd-wasm-node)
The companion repo was renamed meshtastic-web-node -> meshtasticd-wasm-node; fix
the stale name in the wasm README and bump the platform pin to the build that
promotes the canonical name in its emsdk auto-location.
* wasm: re-entrancy guard for the API/region entry points + flaky-open retry
Two robustness fixes for the browser node:
- Re-entrancy guard. The node is single-threaded + Asyncify: while setup()/loop()
is suspended inside a WebUSB transfer, the JS event loop is free, so a stray
DOM/timer callback that re-enters a wasm_* entry point starts a second Asyncify
unwind ("async operation already in flight" abort) or clobbers shared PhoneAPI
state (observed as a "PhoneAPI::available unexpected state" flood). Add a
g_wasm_in_firmware flag set around setup()/loop() (portduino_main_wasm.cpp); the
wasm_set_region / wasm_api_to_radio / wasm_api_from_radio / wasm_api_available
entry points now reject a mid-tick call (return busy) instead of corrupting or
aborting. The host must still call them between ticks — this is the safety net
the design lacked, not a substitute for the JS queue.
- CH341 open retry (js/bridge.js). First-connect WebUSB is flaky — the interface
is briefly unclaimable right after the grant, or held by a prior session,
giving a transient "Could not open SPI: -1". Retry the open with a short
backoff, closing the device between attempts so claimInterface starts clean.
* wasm: exclude emscripten-only sources from cppcheck
The `check` board matrix runs `pio check` (cppcheck) over all of src/,
including src/platform/portduino/wasm/. cppcheck can't parse the EM_ASYNC_JS/
EM_JS macros (Syntax Error: AST broken at libpinedio_webusb.c:39,
internalAstError) and these sources are not part of any checked board build
([env:native-wasm] is board_level=extra, compiled by the build-wasm CI job).
Suppress the wasm dir in suppressions.txt, the same way generated/ and .pio/
are already excluded.
* wasm: coalesce FS.syncfs so two never run at once
IDBFS syncfs is async; the explicit wasm_fs_sync (5s timer + post-save +
beforeunload) could overlap a prior in-flight sync, warning "2 FS.syncfs
operations in flight at once". Serialize: if a sync is running, mark a pending
re-sync and let the in-flight one chain it on completion — at most one in flight,
trailing writes still flushed. (Companion drops IDBFS autoPersist so this is the
single persistence path.)
* wasm: silence false-positive SAST on the emscripten glue
- extra_scripts/wasm_link_flags.py: restore the trunk-ignore-all(ruff/F821,
flake8/F821) header every other SCons extra_script carries; Import/env are
SConscript-injected globals, so ruff/flake8 flag them as undefined.
- .semgrepignore: exclude src/platform/portduino/wasm/js/ (browser WebUSB glue,
not part of the firmware binary). The unsafe-formatstring rule false-positives
on its benign retry/diagnostic console logs.
* Update .github/workflows/build_portduino_wasm.yml
Co-authored-by: Austin <vidplace7@gmail.com>
---------
Co-authored-by: Austin <vidplace7@gmail.com>
|
||
|
|
ad132e9540 | Protobufs update |