* docs(nodedb): make the native node cap unambiguous The native node cap was stated in four places that disagreed, and the disagreement already caused a wrong diagnosis: a saturated 200-node database looked arithmetically impossible because the cap had been read as 248, computed from a header that does not apply on this platform. The real value is 198. On portduino MAX_NUM_NODES is not a compile-time constant at all - the variant defines it as `portduino_config.MaxNodes`, resolved at runtime, default 200 and settable per host with `General: MaxNodes`. variant.h is reached before mesh-pb-constants.h, so that header's ARCH_PORTDUINO branch never fires and its plausible-looking 250 is dead code. - #error-guard the dead branch rather than leave a wrong number where people grep. The guard found a real defect: seven translation units reach mesh-pb-constants.h without configuration.h (SerialConsole.cpp, StreamAPI.cpp, PacketAPI.cpp, ServerAPI.cpp, PiWebServer.cpp, ServiceEnvelope.cpp, MeshtasticOTA.cpp, and test/TestUtil.cpp), so each was compiling with a different MAX_NUM_NODES - and therefore a different PACKETHISTORY_MAX - than the rest of the build. Each now includes configuration.h first. It cannot be included from mesh-pb-constants.h itself: that reaches SerialConsole.h through DebugConfiguration.h and closes a cycle. - Name the bare 250 in getMaxNodesAllocatedSize() NODEDB_MIGRATION_LOAD_CEILING. It is a decode allowance for files written by larger-cap firmware, not a cap, and it read like one. - Fix docs/node_info_stores.md, which named the wrong source and a "10-250" range that is wrong for native, and the copilot-instructions tunables line that said "portduino 250". * test(harness): give each suite its own scratch HOME and report leftovers Native suites shared one directory. Every suite that constructs a NodeDB loads and saves ~/.portduino/default/prefs/ - nodes.proto, config.proto, channels.proto, module.proto, device.proto, warm.dat, transmit_history.dat - and nothing cleared it, so state leaked suite -> suite within a run and run -> every run after it. A test run could also rewrite a real meshtasticd node database on the same machine. Per-run isolation does not fix this: the leak is generated inside a single run, so the boundary has to be per suite. bin/pio-test-isolate.sh runs each suite in its own scratch $HOME, registered as test_testing_command for env:native and env:coverage so a bare `pio test` and CI get the same boundary, not just bin/run-tests.sh. It runs the binary unchanged and exits with its exit code, so PlatformIO's pass/fail is untouched. Overriding HOME here rather than around `pio` also sidesteps the blocker that a bare HOME= breaks pio's own ~/.platformio/penv/bin/pio lookup. Leftovers are reported as a second axis, PASS/FAIL x CLEAN/DIRTY, because an unintended write has no matching assertion by definition - nobody writes TEST_ASSERT for a save they do not know is happening. The harness asserts it from outside, so it applies to every suite without the author opting in. - Only the *set of changed paths* is asserted, never contents. Hashes answer the boolean "did this change?" and nothing more; content baselines over protobuf bytes would churn on every NodeInfoLite field added, which is how snapshot suites become noise. - Deliberate writes are declared in test/state-manifest.tsv - one central file, suite / flags / mandatory reason. run-tests.sh prints the opt-out count on every run. - Granularity follows the state flag, so the two ship together: per-test by default (TestUtil redefines RUN_TEST to checkpoint after each test, naming the exact test that dirtied things), suite boundary for state=per-suite, where carrying state across test cases is the declared behaviour. - A declared write that does NOT happen is reported as MISSING, not folded into DIRTY. It catches silently broken persistence; a warning for now, since some are conditional. - Graded AMBER, not RED. With isolation in place DIRTY means "undeclared", not "dangerous", and a check that lands red on day one gets switched off. Guard the guard, both halves: state_assert_empty() refuses to run a suite against a sandbox that is not empty (otherwise the after-diff measures against the wrong baseline and reports CLEAN while meaning nothing), and bin/test-state-check.sh drives the real wrapper with fixtures asserting CLEAN / CLEAN / DIRTY / MISSING plus both directions of the empty assertion. A checker that silently matches everything would otherwise pass forever. --write-manifest proposes entries for a human to paste and justify; it never applies them, and neither does CI. * test(harness): stop reporting Unity's exit code as a signal A native suite ends in exit(UNITY_END()), and UNITY_END() returns the failure count. PlatformIO's native runner reads that non-zero exit code as a POSIX signal number, so four failures print "Program received signal SIGILL", five print "SIGTRAP", and the suite is classified [ERRORED] rather than [FAILED]. There is no crash. The signal name tracks the failure count and nothing else - it moved SIGILL -> SIGTRAP when a diagnostic probe added a fifth failure - and it cost hours of hunting a memory bug that did not exist, on an env (native) that carries no sanitizer at all. It also explains the phantom extra test case in the totals: the runner adds a synthetic entry for the signal it thinks it saw. run-tests.sh now says so inline whenever a signal line appears, and the three agent-facing docs say it too. * test(admin): isolate NodeDB and globals per test setUp() did `if (!nodeDB) nodeDB = new NodeDB();` and never deleted it, so 83 of the 85 tests shared one never-reset database and never restored config, owner, devicestate or channelFile. The fixture that does restore them was opt-in and armed by exactly two tests. The setUp comment claiming the rest "set their own config/region state and are unaffected" was not true - the admin handlers under test write all four globals. Route every test through the fixture instead: setUp saves the globals and installs a fresh NodeDB, tearDown restores and deletes it. The two tests that armed it themselves no longer need to. All 85 pass, so nothing was silently relying on the shared state. It costs about 7% of the suite's runtime (a NodeDB construction is a loadFromDisk plus, with a region set, key generation) - worth paying to write the phase 3 tests against a clean fixture rather than 83 tests' residue. Also cap the per-test attribution in the run summary at five entries; the full list stays in the suite's sandbox. * test(fs): cover the bounded file-manifest walk getFiles() runs on every phone sync via STATE_SEND_FILEMANIFEST, and nothing asserted any of its bounding behaviour. It does execute unasserted from test_stream_api's handshakes, but the cap, the depth limit, the wasLimited paths, overlong-path rejection and capacity release were all unguarded. Eight tests, all describing what the code does today: today's code is already correct here, since #10778 landed the by-reference collectFiles(), the 64-entry cap, the strlcpy bounds and the swap-idiom release. They pass on arrival, which is the point - this is the baseline a later change has to leave alone. Two things they do not cover, and cannot: - Moving reserve() outside the __cpp_exceptions guard. Exceptions are on natively, so the #else branch is not compiled. The suite's job there is to prove that change alters nothing observable. - The file.name() null guard. No in-tree backend returns null; the guard is defensive. The manifest-release test pins the swap idiom rather than calling PhoneAPI's releaseFilesManifest(), which is file-local. It asserts capacity() == 0, not just size() == 0 - a size-only check passes on clear(), which is the bug #7924 shipped. Suite count 43 -> 44, recounted against the directories rather than copied. * test(admin): assert node-DB metadata saves skip the radio reload set_favorite_node, set_ignored_node and toggle_muted_node each persist a NodeInfoLite bit and nothing else. MeshService::reloadConfig() gates its region re-derivation and configChanged notification on saveWhat & (SEGMENT_CONFIG | SEGMENT_CHANNELS), so a SEGMENT_NODEDATABASE-only save already skips the live radio reconfigure. Pure characterization - all three pass on develop. Worth pinning because that reconfigure is the path implicated in the WisMesh Tag favourite-node crash, and develop asserts nothing about it: widening the saveWhat mask or reordering the check would currently go unnoticed. Ported from the config-save series along with ConfigChangedCounter (an Observer<void *> counting configChanged notifications, the only externally visible signal that the reload branch was taken) and TEST_NODE_NUM. They join the existing suite, so no suite-count change. * refactor(menu): extract the mute toggle into a named function The node menu's mute action was inline in a banner-callback lambda, and that lambda only ever runs via screen->showOverlayBanner() - which is why nothing in MenuHandler.cpp was reachable from a test. Lift the `selected == Mute` branch into menuHandler::toggleNodeMuted(uint32_t) and call it from the lambda. Behaviour-neutral by construction: same statements, same order, same bare saveToDisk(). The null check moves into the function, so the call site no longer needs its own lookup. Verified by the native build and suite; the byte-identical-image check on a headroom-constrained nRF52 board was not run locally - CI's firmware-size comment covers it. Three tests come with it, all describing today's behaviour: - the bit flips both ways and no configChanged fires (develop never calls reloadConfig on this path); - an unknown node is a no-op rather than a write; - and the segment mask. Flipping one NodeInfoLite bit currently rewrites all five segments via bare saveToDisk(). That is asserted deliberately, with the comment naming it as characterization of a known defect: a pending fix narrows it to SEGMENT_NODEDATABASE, and when it lands this assertion is expected to change, which makes the improvement visible in the diff instead of silent. saveToDisk() is not virtual, so the mask is observed through its effect - remove the five prefs files, toggle, and see which reappear. * docs(test): make every suite count a pointer to the canonical one test/native-suite-count is the registered total and is machine-checked against test/test_* on every full run and by the suite-count-check CI job. Every other statement of the count is a copy that drifts: copilot-instructions said 12, AGENTS.md said 19, and the real number is 44. Replace both literals with a pointer to the file, say explicitly that no document should state the count as a literal, and reframe the two suite listings as descriptions rather than inventories - they carry per-suite information the count does not, so they stay, but nothing should infer completeness from their length. Register the new FS suite in both. * test(harness): randomise suite order, reproducibly Landed last, deliberately. Randomising an order-dependent suite set does not find bugs so much as convert a silent pass into intermittent red, and the first instinct is to revert the randomisation rather than fix the coupling. Phases 1-2 removed the coupling; this keeps it removed. Both runners previously hid order dependence behind a fixed order that happened to differ between them, and neither order was chosen: CI's area rules put admin first, PlatformIO's local discovery is reverse alphabetical and put it last. CI was green by accident. - bin/run-tests.sh --shuffle / --seed <n>. The seed defaults to HEAD's short SHA: one order per commit, so a red is replayable and attributable to the diff instead of flaky, while the project keeps exploring orders. Printed at the start and carried into the RESULT line, so a verdict is replayable from that line alone; the full order is printed on failure, because for an order-dependent failure the order is the diagnostic. - The shuffle is a Fisher-Yates over a MINSTD generator rather than awk's rand(), whose sequence differs between gawk and mawk. A seed that does not reproduce the same order on another machine is not a seed. - Shuffling needs one `pio test -f <suite>` invocation per suite - PlatformIO orders by its own os.walk() over test/ and filters only select - which measures at about 4.7s per suite of extra startup. - CI shuffles its area order, seeded from GITHUB_SHA and printed with the command to replay it locally. Intra-area order stays PlatformIO's; controlling it there would mean per-suite invocations, which is a cost worth deciding separately. Also records the 16 measured entries in test/state-manifest.tsv, each with its reason, taken from a full run's --write-manifest output rather than guessed. * test(default): cover the region-throttle interval overload getConfiguredOrDefaultMsScaled(configured, default, nodes, TrafficType) is the overload every telemetry and position module actually calls, and nothing referenced TrafficType anywhere under test/. All four of its behaviours were unguarded: the no-region guard, the throttle <= 1 short-circuit, the multiply, and the 64-bit overflow clamp. The throttles are real, not hypothetical - EU_866 carries PROFILE_LITE, which sets both positionThrottle and telemetryThrottle to 10, so a change here moves broadcast spacing in that region by an order of magnitude. Each test pins numOnlineNodes at the congestion threshold and uses ROUTER, which never congestion-scales, so the coefficient is 1 and the throttle is the only variable. The overflow case needs a base above INT32_MAX/10, hence three days rather than one. * ci(test): keep pull-request suite order fixed, seed the rest Shuffling the area order on every run - including pull_request - would turn a contributor's PR red for an ordering they did not choose, which is how a randomisation gets reverted instead of the coupling being fixed. That is the exact dynamic the ordering work was sequenced last to avoid, and the previous commit walked straight into it. - pull_request keeps the fixed declared area order. - push and schedule shuffle, seeded from the commit SHA: deterministic per commit, printed, attributable, and never blocking someone else's PR. - A suite_order_seed input on workflow_call and workflow_dispatch overrides both, so a specific failing order can be replayed anywhere, including on a PR. The run log prints which mode it took, the resulting order, and the local command to replay it. * ci(test): satisfy CKV_GHA_7 and yamllint on the seed input The seed is reachable through workflow_call, which callers can pass programmatically. The workflow_dispatch copy tripped checkov's "workflow_dispatch inputs MUST be empty" rule, and suppressing it was not worth it: replaying a specific order is a local operation, and the run log already prints the exact bin/run-tests.sh command to do it. * style(menu): apply the node-ID format convention RadioInterface.cpp documents the rule: 0x%08x in logs, !%08x in user-facing display. MenuHandler held every remaining exception - seven logs printing bare %08X, and two display labels doing the same. Repo-wide there are now no bare %08X node IDs left in log calls. * ci(test): pass workflow inputs through env, not shell interpolation suite_order_seed and github.event_name were spliced into the run: script as ${{ }} text, so a value carrying shell metacharacters would execute as code on the runner rather than being read as data. semgrep (run-shell-injection) and zizmor (template-injection) both flag it. Both now arrive as environment variables and are read as "$VAR". * refactor(test): share the seeded shuffle between the harness and CI bin/run-tests.sh and test_native.yml each carried a byte-identical copy of the MINSTD Fisher-Yates awk. The workflow prints "replay locally: ./bin/run-tests.sh --shuffle --seed $seed" after a shuffled CI run, and that instruction is only true while the two agree - drift would be announced by a replay quietly reproducing a different order than the one that failed. Extract shuffle_suites() to bin/lib/shuffle.sh and source it from both. Permutations verified identical across seeds before and after the move. * fix(test): correct the shared-state MISSING check and summary join Three defects in the new harness: state_classify() matched declarations two different ways - state_path_declared() for "undeclared", a hand-rolled regex for "missing". Interpolating an entry into an ERE also let a metacharacter in a manifest name match a file that is not the declared one. Both directions now go through the one helper. `paste -sd'; '` does not join with "; ": with -s, paste cycles through a multi-character delimiter one character per join, so paths rendered as "a;b c;d e". Replaced with an awk join. test-state-check.sh ran on after a failed cd instead of stopping (SC2164). ./bin/test-state-check.sh: 6/6 fixtures pass, MISSING included. * fix(portduino): bound General.MaxNodes MaxNodes was validated only for <= 0. Any positive value, including a typo'd or pasted-in one, propagates to MAX_NUM_NODES and scales both the node DB and the nodes.proto decode ceiling - failing at boot with no obvious cause. The ceiling is a sanity bound, not a capability limit; raise it if a host genuinely needs more. * docs(nodedb): reconcile the capacity tables The property matrix omitted the ESP32-S3 100-node flash tier that the platform table above it lists, and neither mentioned that the WASM build overrides MaxNodes to 80 in wasm_config_apply(). * fix(nodedb): make mesh-pb-constants.h self-sufficient on portduino The ARCH_PORTDUINO #error assumed it was unreachable in a normal build. It is not: the vendored device-ui sources include this header without configuration.h, which broke both native-tft docker builds. Include configuration.h here instead, ahead of every compile-time default - variant.h overrides MAX_RX_TOPHONE as well as MAX_NUM_NODES, so placing it lower in the file just moves the divergence to a redefinition. The #error stays as a backstop for the case where that include genuinely stops providing the cap. Verified with the native env's own flags: a TU including only this header now compiles, normal-order use of both macros compiles, and NodeDB.cpp compiles. * fix(portduino): raise the MaxNodes ceiling to 16000 Marked artificial: nothing in the node DB fails at 16001. 16000 sits just under the 16384 (128 x 128) population where HopScalingModule saturates its sampling denominator and starts dropping nodes, so a host inside the bound still gets meaningful hop recommendations. * lint(trunk): advise on node IDs logged as bare %08x RadioInterface.cpp documents the convention - 0x%08x in logs, !%08x in display - but nothing enforced it, which is how the MenuHandler cluster drifted. 22 call sites in PacketHistory, NodeInfoModule and PositionModule are still off it. A trunk linter rather than a CI grep job, because trunk checks changed files: new violations get flagged without a 22-site cleanup landing in an unrelated PR. Modelled on the existing too-many-defined definition. Scoped to values it can tell are IDs - an ID-shaped argument (->num, .from, getNodeNum) or message text naming one. A 32-bit hex that is not an ID is out of scope, so the CRC32 logs in ethOTA.cpp are correctly ignored. Emits "note", trunk's only non-blocking level: "warning" and "info" both exit non-zero and would gate CI, which is not what a log-format nit deserves. The pre-existing sites are line-scoped in the allowlist, so a new bad call in those same files is still caught. * lint(trunk): stop exempting the known node-id-format sites The seeded allowlist made the rule green by declaring the backlog acceptable. Empty it instead, so the 22 pre-existing sites are reported and get cleaned up by whoever next edits those files. Costs nothing to do: the rule emits "note", so these are non-blocking either way. The allowlist stays for its real purpose - a value the linter misreads as an ID. * style: log node and packet IDs as 0x%08x Clears the 22 sites the node-id-format linter reports, so the rule starts from zero rather than from a backlog nobody can see - trunk suppresses pre-existing findings by default, so left alone these would not have surfaced on edit the way an empty allowlist implies. Format strings only; no argument or control flow changes. The !%08x user-facing display forms are deliberately untouched - that is the other half of the same convention. * test(harness): build once up front, so suite timings mean something run-tests.sh fused build and run in a single pio invocation, so whichever suite PlatformIO's directory walk reached first absorbed the entire src compile and reported it as its own duration. On a real run that made a 0.03s suite report 13m21s, and hid the build cost from every other number in the summary. Do what .github/workflows/test_native.yml already does: one --without-testing build pass, then run with --without-building. Measured on a full 44-suite run - the build is now a single reported figure and 968 test cases execute in 1.9s, with no suite above 0.084s. Build output goes to its own log rather than $LOG: the outcome regexes match "error:" and "[ERRORED]", so a compiler diagnostic sharing that file would read as a test failure. Both red paths now keep the log they quote from. $LOG and the build log are mktemps the EXIT trap removes, so the three grepped lines were previously all anyone ever saw - and the cause is usually further up than the first [FAILED]. * test(harness): keep the run log on every red path bin/pio-test-isolate.sh already keeps a failing or DIRTY suite's sandbox and log under .pio/test-state/<suite>/. What was missing is the cross-suite view: $LOG is a mktemp the EXIT trap deletes, so run-tests.sh quoted three grepped lines from a file that no longer existed by the time anyone looked. Preserve it as .pio/build/<env>/test-failure.log from both red paths - including "no success summary found", which said "see log" while preserving nothing, and which is exactly the case where the build died before any suite ran and so left no per-suite sandbox either. Cleared at the start of every run, so a green run cannot leave a red one's log lying around looking current. * fix(test): report the real failure count on a shuffled red A shuffled run is one `pio test` invocation per suite, all appending to the same log, so the log carries one PlatformIO "N test cases:" summary per suite. verdict_red() took `tail -1`, which reports whatever the LAST suite did: a failure in suite 3 printed a "0 failed" summary from suite 44 directly under "RED - failures detected:". Sum the summaries instead. A single summary line - every unshuffled run - is passed through verbatim, so the familiar output is byte-identical. The patterns are passed to the awk helper as strings rather than /regex/ literals: awk evaluates a regex literal in argument position as `$0 ~ /re/`, so the callee would receive 0 or 1 and silently sum garbage. * fix(test): do not emit an empty suite name for an empty shuffle `printf '%s\n' "$@"` with no arguments still writes one empty line, and both callers read shuffle_suites through mapfile, so an empty suite list arrived as a single suite named "". Return before the printf when there is nothing to shuffle. * test(harness): state and enforce the Linux host requirement The native harness is a Linux tool: bash 4+ (mapfile), GNU coreutils and GNU find (-printf, md5sum, -executable). Most of that predates this branch - mapfile and both find predicates are already on develop - but none of it was written down, so the requirement was there to be discovered rather than read. Refuse to start on a non-Linux uname instead of degrading. On a BSD userland this would not fail cleanly: it would mis-hash the sandbox and mis-read the suite list, and still print a verdict. A state check that silently measures the wrong thing is worse than one that declines to run. Carrying a per-host fallback was the alternative, and it buys a second code path that nothing in CI exercises. bin/test-native-docker.sh already exists for macOS and non-Linux hosts, and the native-macos PlatformIO env is a build target for meshtasticd, not a test host - the isolation wrapper is registered for env:native and env:coverage only. Documented in the script header, test/README.md, and both agent docs. * fix(test): terminate every suite with exit(UNITY_END()) Two sites across two suites ended on a bare UNITY_END(). That ends the reporting, not the suite: setup() returns, the runtime goes on calling loop(), and the process runs forever. PlatformIO does not notice - it reports a suite from its Unity output, not from process exit - so the suite passes, the run goes green, and the binary stays resident. Thirteen of them had accumulated on one dev box, the oldest 19 hours old. The costs are quiet by construction: - the per-suite sandbox is deleted underneath a live process, so its CLEAN/DIRTY verdict describes what the suite had written when the harness stopped looking, not what it left behind; - .gcda coverage and LeakSanitizer's report both flush from atexit handlers, so a suite that never exits contributes no coverage and gets no leak check; - each survivor pins its own deleted 94 MB binary, which du cannot see. One of the two is the #else of an architecture guard, which is the easiest one to get wrong - it looks like there is nothing to clean up. test_mqtt has a correct exit(UNITY_END()) in its live branch, so a "does this file call exit() anywhere" check passes the file whole. test_serial had two more. develop's serial-config validation rework restructured that suite - the architecture guard is gone and both remaining branches now exit correctly - so this commit no longer has anything to change there; bin/lint-unity-exit.sh, added later on this branch, is what keeps it that way. test/README.md gets a section on it, since the skeleton showing the right shape had not stopped this happening. * test(harness): detect and reap suites that outlive their run A suite that never exits was invisible: PlatformIO reports a suite from its Unity output, so the run stayed green while the binary kept running. Two checks, because they fail differently. Runtime, in bin/pio-test-isolate.sh: the sandbox $HOME is mktemp-unique per suite, so any process still holding it is a survivor of that suite. Matching on the environment rather than a remembered PID identifies one whatever its parentage - a fork, a grandchild, a process already reparented to init - none of which a $! comparison catches. Reaped before the after-fingerprint is taken, so that fingerprint measures a tree nobody is still writing to, and so a run cannot leave processes accumulating on the host. Recorded as a sixth summary column and graded AMBER: the tests did pass, but the CLEAN verdict and the coverage were measured under a false assumption. Author-time, as bin/lint-unity-exit.sh, wired into trunk at "note" like node-id-format: every UNITY_END() must be wrapped in exit(). The rule is per occurrence, and that is the point - a file-level "calls exit() somewhere" check passes test_serial and test_mqtt, which have a correct one in their live branch and a bare one in the #else. Running it over the tree turned up test_mqtt, which the file-level pass had missed. It allows `int rc = UNITY_END(); ...; exit(rc)`, used by test_packet_signing to restore globals between the summary and the exit. That is where the rule gives ground: capturing and never exiting would leak and is not flagged. Flagging a correct idiom would push someone to "fix" working code. bin/test-state-check.sh gains a survivor fixture, asserting the wrapper both reports and reaps - a detector that only reports leaves the host accumulating processes, which is half the harm. 8/8. * fix(lint): make the unity-exit scanner statement-aware The rule judged one physical line at a time, which reports two kinds of correct code as bare: /* a comment that happens to mention UNITY_END() */ <- interior lines were never stripped exit( UNITY_END()); <- exit( and the macro never met On a probe of both, two of three findings were wrong. This is a note-level rule whose whole job is advice, and bin/lint-node-id-format.sh already says why that matters: a false positive costs more than a miss. One that cries wolf gets ignored, and the real finding goes with it. Carry /* ... */ state across lines and accumulate logical statements before testing, with a 12-line cap so one unclosed call cannot swallow the rest of the file - the same structure lint-node-id-format.sh uses, so the two custom linters in bin/ work alike rather than each having its own idea. Verified both directions: the develop-era sources still produce the same four findings, the fixed tree produces none, and a probe covering block-comment interiors, wrapped exit(), line comments, return UNITY_END() and capture-then- exit reports only the genuinely bare calls - including a complete block comment followed by real bare code on the same line, which the state machine has to keep live. Reported by CodeRabbit on #11322. * fix(lint): tokenise instead of pattern-matching, and self-test it Second round of review findings on the same scanner, all confirmed by direct test before changing anything. Six defects, one root cause: layered regexes cannot tokenise C++. False positives (correct code reported): - UNITY_END() inside a string literal read as code False negatives (real leaks missed): - a string containing "/*" opened comment state and swallowed later lines - greedy .* removed everything between two block comments on one line, taking a bare call with it - myexit(UNITY_END()) matched the exit() exemption as a substring - x == UNITY_END() and total += UNITY_END() matched the assignment exemption Replaced with a character-level scan carrying comment state, and token-bounded exemptions: exit must be a whole identifier, and the capture form must be a plain `=`. Raw string literals are still not modelled - there are none under test/, and delimiter tracking for a case that does not occur would be untested code guarding untested code, so it is documented rather than guessed at. Also drops the `return UNITY_END()` exemption. It only terminates from main(), there is no main() under test/, and from a helper it just returns a count. bin/test-lint-unity-exit.sh pins all fifteen cases, every false positive and false negative found in review among them. The rule has been wrong twice in a way that looked fine by inspection; it needed a self-test more than it needed another careful reading. Two further findings in the same review: - bin/run-tests.sh dropped PASSTHRU in shuffled mode, so `--shuffle -vvv` built verbosely and then ran quietly. The shuffled loop now forwards EXTRA_ARGS, which is PASSTHRU minus the -f pair it supplies per suite. - bin/run-tests.sh did not guard `cd "$ROOT_DIR"`. And one that did not reproduce: the survivor fixture's glob does find the pid file (verified with the lookup instrumented - the earlier failure was an artifact of running the script from /tmp, where SCRIPT_DIR cannot resolve). The assertion was still weak, because an empty pid took the "not running" branch and passed vacuously. It now fails if the pid was never recorded, and finds the file by search rather than assuming a directory depth. Reported by CodeRabbit on #11322. * fix(lint): report each UNITY_END occurrence at its own location The self-test only asked "did the linter say anything", so it could not have caught a wrong line, a wrong column, or a missing second finding. Fixtures now assert the exact diagnostics as line:col, and the first run of that assertion found two real problems. The caret pointed at the wrong occurrence. For `exit(UNITY_END()); UNITY_END();` the verdict was right but the column was 17 - the wrapped call - because the scanner stripped terminating forms out of the whole statement and then reported the first occurrence it had seen. Two bare calls on one line reported once. Judged per occurrence now, by looking back through whitespace at what wraps it, so both the count and the caret are right. That also needed a position map from strip_noncode(): removing a comment or collapsing a literal shifts every later column, and counting occurrences in the raw line does not recover it either - TEST_MESSAGE("... UNITY_END() ..."); UNITY_END(); has two occurrences in the raw text and one in the code. Four of the expected columns I wrote by hand were also wrong, off by one. The linter was right in every case; the assertions were not. They are computed from the fixture text now rather than pasted from output, because a baseline accepted from the tool it is testing asserts nothing. 17 fixtures, including the two-on-one-line case from review and its mirror. Reported by CodeRabbit on #11322.
81 KiB
Meshtastic Firmware - Copilot Instructions
TL;DR
Local tests ./bin/run-tests.sh(exit 0 GREEN · 1 RED · 2 AMBER · 3 FILTERED)Hardware tests meshtastic/meshtastic-mcp ( MESHTASTIC_FIRMWARE_ROOT→ this checkout)Format trunk fmtMirror docs AGENTS.md(short pointer for agents that don't read this file) ·CLAUDE.md(Claude Code)Need this? It's here.
General helpers (clamp, UTF-8, string fmt…) src/meshUtils.hLogging macros (LOG_DEBUG / INFO / WARN…) src/DebugConfiguration.hNew module skeleton inherit ProtobufModule<T>insrc/mesh/ProtobufModule.hObserver / event wiring src/Observer.h
This document provides context and guidelines for AI assistants working with the Meshtastic firmware codebase.
Project Overview
Meshtastic is an open-source LoRa mesh networking project for long-range, low-power communication without relying on internet or cellular infrastructure. The firmware enables text messaging, location sharing, and telemetry over a decentralized mesh network. The project uses C++17 as its language standard across all platforms.
Supported Hardware Platforms
- ESP32 (ESP32, ESP32-S3, ESP32-C3, ESP32-C6) - Most common platform
- nRF52 (nRF52840, nRF52833) - Low power Nordic chips
- RP2040/RP2350 - Raspberry Pi Pico variants
- STM32WL - STM32 with integrated LoRa
- Linux/Portduino - Native Linux builds (Raspberry Pi, etc.)
- macOS native - Headless
meshtasticdon Apple Silicon / x86_64; seevariants/native/portduino/platformio.inifor Homebrew prereqs + CH341 LoRa setup
Supported Radio Chips
- SX1262/SX1268 - Sub-GHz LoRa (868/915 MHz regions)
- SX1280 - 2.4 GHz LoRa
- LR1110/LR1120/LR1121 - Wideband radios (sub-GHz and 2.4 GHz capable, but not simultaneously)
- RF95 - Legacy RFM95 modules
- LLCC68 - Low-cost LoRa
MQTT Integration
MQTT provides a bridge between Meshtastic mesh networks and the internet, enabling nodes with network connectivity to share messages with remote meshes or external services.
Key Components
src/mqtt/MQTT.cpp- Main MQTT client singleton, handles connection and message routingsrc/mqtt/ServiceEnvelope.cpp- Protobuf wrapper for mesh packets sent over MQTTmoduleConfig.mqtt- MQTT module configuration
MQTT Topic Structure
Messages are published/subscribed using a hierarchical topic format:
{root}/{channel_id}/{gateway_id}
root- Configurable prefix (default:msh)channel_id- Channel name/identifiergateway_id- Node ID of the publishing gateway
Configuration Defaults (from Default.h)
#define default_mqtt_address "mqtt.meshtastic.org"
#define default_mqtt_username "meshdev"
#define default_mqtt_password "large4cats"
#define default_mqtt_root "msh"
#define default_mqtt_encryption_enabled true
#define default_mqtt_tls_enabled false
Key Concepts
- Uplink - Mesh packets sent TO the MQTT broker (controlled by
uplink_enabledper channel) - Downlink - MQTT messages received and injected INTO the mesh (controlled by
downlink_enabledper channel) - Encryption - When
encryption_enabledis true, only encrypted packets are sent; plaintext JSON is disabled - ServiceEnvelope - Protobuf wrapper containing packet + channel_id + gateway_id for routing
- JSON Support - Optional JSON encoding for integration with external systems (disabled on nRF52 by default)
PKI Messages
PKI (Public Key Infrastructure) messages have special handling:
- Accepted on a special "PKI" channel
- Allow encrypted DMs between nodes that discovered each other on downlink-enabled channels
Encryption & Key Management
Meshtastic packets on the air are typically encrypted one of two ways: the per-channel symmetric layer (AES-CTR with a shared PSK) for broadcasts and channel traffic, and the per-peer PKI layer (X25519 ECDH → AES-256-CCM) for direct messages and remote admin. A channel with a 0-byte PSK (or Ham mode, which wipes PSKs) transmits cleartext - see the size table below. Both are implemented in src/mesh/CryptoEngine.cpp; the send/receive dispatch lives in src/mesh/Router.cpp; admin authorization lives in src/modules/AdminModule.cpp.
High-level model
- Channels are symmetric rooms: anyone with the PSK can read any message on the channel. Channel 0 is the "primary" channel and ships with the short-form default PSK on factory devices, forming the public mesh most users join. (The LoRa modem preset
LONG_FASTlives onconfig.lora.modem_presetand is an independent field - don't conflate "channel 0 default PSK" with the modem preset name.) - DMs addressed to a single node require PKI so that other holders of the channel PSK can't read them. Outside Ham mode, Meshtastic does not fall back to channel-symmetric encryption when the destination public key is unknown.
- Remote admin is a DM carrying an
AdminMessage. The receiver only acts on it if the sender's public key is on its allowlist (config.security.admin_key[0..2]). - Ham mode (
owner.is_licensed=true, whereowneris the localmeshtastic_Userrecord) disables PKI entirely and sends cleartext - FCC Part 97 prohibits encryption on amateur bands. - No ratchet, no session. Every packet is encrypted from scratch - a stateless design that matches the high-loss, store-and-forward nature of LoRa.
Symmetric channel encryption (AES-CTR)
CryptoEngine::encryptPacket / decrypt / encryptAESCtr in src/mesh/CryptoEngine.cpp.
- Cipher: AES-CTR, AES-128 or AES-256 depending on key length. Same routine in both directions (CTR is a stream cipher, so encrypt == decrypt).
- Key:
ChannelSettings.pskbytes. Size semantics:- 0 bytes → no encryption, cleartext on the air
- 1 byte → short-form index into the well-known
defaultpsk[]insrc/mesh/Channels.h. Index 0 = cleartext; 1 = defaultpsk unchanged; 2..255 = defaultpsk with its last byte incremented by (index − 1). This is what the CLI's--ch-set psk defaultproduces. - 16 bytes → raw AES-128 key
- 32 bytes → raw AES-256 key
- 2..15 bytes → zero-padded to 16 and used as AES-128 (with a warn log); 17..31 bytes → zero-padded to 32 and used as AES-256 (with a warn log). Defensive fallback for malformed PSK input, not something to rely on.
- Nonce (128 bit):
packet_id(u64 LE) ‖from_node(u32 LE) ‖block_counter(u32, starts at 0). Built inCryptoEngine::initNonce. - No AEAD: channel packets carry no MAC, so the channel-hash byte is not an integrity or authenticity check.
Channels::getHashis a 1-byte XOR-derived hint over the channel name bytes and PSK bytes that helps receivers pick a candidate channel/PSK for decryption. Because it is only a small hint and collisions are easy to find, it should be described purely as a PSK-selection aid, not as a security filter an attacker cannot bypass. - Channel 0 is special in one way only: it's the channel the Router attempts PKI decryption on before falling through to AES-CTR. Non-zero channels always go straight to AES-CTR.
PKI encryption for DMs (X25519 ECDH + AES-256-CCM)
CryptoEngine::encryptCurve25519 / decryptCurve25519 in src/mesh/CryptoEngine.cpp.
- Keypair: Curve25519 (aka X25519), 32-byte public + 32-byte private. Stored in
config.security.public_key/private_key; the public half is mirrored intoowner.public_keyso it rides along in NodeInfo broadcasts and propagates through the mesh like any other identity field. - Key generation (
generateKeyPair): stirsHardwareRNG::fill()(64 B from platform TRNG when available), the 16-bytemyNodeInfo.device_id, and a call torandom()into the rweather/Crypto library's software RNG, thenCurve25519::dh1.regeneratePublicKeyrecomputes the public half from a known private (used when restoring from backup). - Keygen entry points: at boot,
NodeDBcallsgenerateKeyPair(orregeneratePublicKeywhen a stored private key is present and passes a low-entropy check) directly when!owner.is_licensedandconfig.lora.region != UNSET.ensurePkiKeyswraps the same logic for runtime/admin flows - it's the pathAdminModule::handleSetConfigruns when first assigning a valid region or when security config is written; do not assume it's the universal boot-time gate, because the NodeDB path bypasses it. - Handshake:
Curve25519::dh2(local_private, remote_public) → 32-byte shared secret → SHA-256 → 32-byte AES-256 key. Recomputed per packet. The SHA-256 step is effectively a KDF over the raw ECDH output. - Cipher: AES-256-CCM via
aes_ccm_ae/aes_ccm_ad(src/mesh/aes-ccm.cpp). MAC length (theMparameter) is 8 bytes. No AAD - the MAC covers ciphertext only. - Nonce (13 bytes / 104 bit):
aes_ccm_ae/aes_ccm_aduse a 13-byte CCM nonce (L = 2is hardcoded insrc/mesh/aes-ccm.cpp), not a 16-byte nonce. For PKI packets,CryptoEngine::initNonce(fromNode, packetNum, extraNonce)starts from the usual packet-derived nonce material, then overwrites nonce bytes4..7with a fresh 32-bitextraNonce = random(). The effective nonce bytes are therefore: bytes0..3=packet_id, bytes4..7= transmittedextraNonce, bytes8..11=from_node, byte12=0x00. The receiver reconstructs the same 13-byte nonce from the packet metadata plus the appendedextraNonce. - Wire overhead: 12 bytes appended to the ciphertext = 8-byte MAC ‖ 4-byte extraNonce. Defined as
MESHTASTIC_PKC_OVERHEAD = 12insrc/mesh/RadioInterface.h. Only the 4-byteextraNonceis sent; the rest of the 13-byte CCM nonce is reconstructed from packet fields as described above. The Router's send path checks this overhead againstMAX_LORA_PAYLOAD_LENbefore committing to PKI. - Send selection (
Router::send): the sender enters the PKI path when all hold - we're the originator AND not Ham mode AND not Portduino simradio AND not on theserial/gpiochannels (unless the packet is already markedpki_encrypted) ANDconfig.security.private_key.size == 32AND destination is a single node (not broadcast) AND the portnum isn't infrastructure.TRACEROUTE_APP,NODEINFO_APP,ROUTING_APP, andPOSITION_APPare routed through channel encryption even when DMed (these need to be readable by relaying peers). Once on the PKI path, if the destination's public key isn't in our NodeDB the send fails withPKI_SEND_FAIL_PUBLIC_KEY- it does not silently fall back to channel encryption. If the client explicitly setpki_encrypted=trueand any condition blocks PKI, the send fails withPKI_FAILED. - Receive selection (
Router::perhapsDecode): try PKI decrypt first whenchannel == 0ANDisToUs(p)AND not broadcast AND both peers have public keys in NodeDB ANDrawSize > MESHTASTIC_PKC_OVERHEAD. On success the packet getspki_encrypted=truestamped and the sender's public key copied intop->public_keyfor downstream authorization.
Remote admin authorization
Implemented in src/modules/AdminModule.cpp → handleReceivedProtobuf. The authorization check runs in this order:
- Response messages - if
messageIsResponse(r)is true (the payload is a response to one of our earlier admin requests), it's accepted without any further check. The in-file comment flags this as a known-untightened gap: a stricter implementation would remember whichpublic_keywe last queried and reject responses that don't match. - Local admin -
mp.from == 0(phone app over BLE, serial CLI, internal module); never travels over the air. Rejected ifconfig.security.is_managedis true, because managed devices expect admin to arrive over the air through an authorized remote path. - Legacy admin channel (deprecated) - the packet arrived on a channel named literally
"admin". Gated byconfig.security.admin_channel_enabled; returnsNOT_AUTHORIZEDif the flag is false. Kept for backward compatibility; new deployments should use PKI admin. - PKI admin (preferred for remote) -
mp.pki_encrypted == trueANDmp.public_keymatches one ofconfig.security.admin_key[0..2](up to three authorized 32-byte Curve25519 public keys, typically copied from the admin node's ownuser.public_key). - Fallthrough →
NOT_AUTHORIZED.
On top of authorization, any remote admin message that mutates state (not a request, not a response) also has to pass a session-key check (checkPassKey): the client must first pull a fresh 8-byte session_passkey via get_admin_session_key_request, then echo that passkey back in the mutating message. The device rotates the passkey after 150 s and rejects values older than 300 s - a narrow anti-replay window on top of the PKI layer.
config.security.is_managed = true disables local admin writes (mp.from == 0 is rejected). It does not by itself force every admin action through PKI - the legacy "admin" channel still authorizes remote admin when config.security.admin_channel_enabled == true. The AdminModule refuses to persist is_managed=true unless at least one admin_key is populated - a deliberate guard against operators locking themselves out.
Key-rotation hazards (actions that invalidate peers)
factory_reset_device(the "full" variant, callsNodeDB::factoryReset(eraseBleBonds=true)) → wipes the X25519 private key; a fresh keypair is generated on the next region-set. Every existing peer holds the old public key, so DMs to this node silently fail PKI decrypt until every peer re-exchanges NodeInfo.factory_reset_config(the "partial" variant, callsNodeDB::factoryReset()witheraseBleBonds=false) → preserves the X25519 private key ininstallDefaultConfig(preserveKey=true); the public key is zeroed and gets rebuilt from the preserved private key on the next boot via the NodeDB path'sregeneratePublicKeycall. Identity is preserved and the mesh does not need to re-exchange keys.region=UNSET → valid region→ensurePkiKeysruns inside the samehandleSetConfigpath; missing keys get generated at that moment.- Ham mode transitions - entering Ham mode (
user.is_licensed=true) runsChannels::ensureLicensedOperation, which wipes every channel PSK (all traffic becomes cleartext) and disables the legacy admin channel. The X25519 private key is preserved on the device but not used becauseRouter::sendskips PKI whenowner.is_licensedis true. Leaving Ham mode re-enables PKI with the preserved keypair but does not restore the wiped channel PSKs - the operator has to re-set them. - Channel 0 PSK change → every peer must re-learn the channel hash; cached NodeInfo becomes temporarily unreachable until the next broadcast.
security.private_keyblanked via admin → regenerates both halves (unless in Ham mode) and propagates the new public key via NodeInfo.
NodeDB Layout (v25)
DEVICESTATE_CUR_VER = 25, DEVICESTATE_MIN_VER = 24. The on-device NodeDB was split in v25 into a slim header table plus four optional satellite stores. Older v24 saves auto-migrate at boot. Old training-data instincts (node->user.long_name, node->position.latitude_i, node->is_favorite, node->device_metrics.battery_level) are wrong now - the fields aren't there. Read this section before touching anything that walks nodeDB->meshNodes.
Slim NodeInfoLite
UserLite is flattened onto NodeInfoLite (no nested sub-message); position and device_metrics are removed entirely (tags reserved). MAC address is dropped. Long names are capped at 25 chars (max_size:25 in deviceonly.options); hw_model and role are int_size:8. Encoded size dropped from ~166 B → ~105 B per node.
Booleans are bit-packed into NodeInfoLite.bitfield. Do not read or write the bits directly - use the inline helpers in src/mesh/NodeDB.h:
nodeInfoLiteHasUser(n) // bit 5 - user fields populated
nodeInfoLiteIsFavorite(n) // bit 3
nodeInfoLiteIsIgnored(n) // bit 4
nodeInfoLiteIsMuted(n) // bit 1
nodeInfoLiteIsLicensed(n) // bit 6 - Ham mode peer
nodeInfoLiteIsKeyManuallyVerified(n) // bit 0
nodeInfoLiteHasIsUnmessagable(n) // bit 8 - "is_unmessagable was sent"
nodeInfoLiteIsUnmessagable(n) // bit 7
// via_mqtt is bit 2 (mask exposed; predicate uses the mask directly)
nodeInfoLiteSetBit(n, NODEINFO_BITFIELD_IS_FAVORITE_MASK, true); // setter
Satellite stores
Four std::unordered_map<NodeNum, …> members on NodeDB, each gated by its own build flag:
| Map | Value type | Build flag |
|---|---|---|
nodePositions |
meshtastic_PositionLite |
MESHTASTIC_EXCLUDE_POSITIONDB |
nodeTelemetry |
meshtastic_DeviceMetrics |
MESHTASTIC_EXCLUDE_TELEMETRYDB |
nodeEnvironment |
meshtastic_EnvironmentMetrics |
MESHTASTIC_EXCLUDE_ENVIRONMENTDB |
nodeStatus |
meshtastic_StatusMessage |
MESHTASTIC_EXCLUDE_STATUSDB |
Defaults are ON (i.e., maps excluded) for STM32WL only - see src/mesh/mesh-pb-constants.h. On every other arch all four maps are present. When excluded, the map member is absent and the corresponding accessors return false.
All four maps are guarded by mutable concurrency::Lock satelliteMutex - concurrent access from receive threads, the phone API state machine, and the renderer is the rule, not the exception.
Accessor convention
Never hand out pointers into the maps. Use the copy-out accessors on NodeDB:
bool copyNodePosition(NodeNum, meshtastic_PositionLite &out) const;
bool copyNodeTelemetry(NodeNum, meshtastic_DeviceMetrics &out) const;
bool copyNodeEnvironment(NodeNum, meshtastic_EnvironmentMetrics &out) const;
bool copyNodeStatus(NodeNum, meshtastic_StatusMessage &out) const;
Each takes the lock, copies the value if present, returns false if the entry is absent or the DB is excluded. Pass-by-out-param is deliberate - pointer-style accessors would invite UAF and lock-leak bugs across the renderer. The "has any X" convenience predicates (hasValidPosition etc.) are implemented in terms of these.
Writers go through setNodeStatus, updatePosition, updateTelemetry (which dispatches on which_variant for device vs environment metrics) - these own the lock and the eviction hooks.
Eviction
Every code path that drops a node from the header table must also evict the satellites. The single chokepoint is eraseNodeSatellites(NodeNum); it's already called from getOrCreateMeshNode's oldest-boring eviction, demoteOldestHotNodesToWarm (the over-cap warm-tier migration), removeNodeByNum, both branches of resetNodes, cleanupMeshDB, addFromContact's ignored-branch, and AdminModule's set_ignored_node. Add new eviction sites here, not by calling .erase() directly. (Note: enforceSatelliteCaps/evictSatelliteOverCap call .erase() directly on purpose - that's a satellite-only cap trim where the node stays in the header, a different operation from this chokepoint.)
Warm tier (long-tail identity)
On every arch except STM32WL and bare nRF52832 (WARM_NODE_COUNT > 0), a node evicted from the header table is not forgotten outright: WarmNodeStore (src/mesh/WarmNodeStore.{h,cpp}) keeps a 40 B {num, last_heard, public_key} record per evicted node - primarily so PKI DMs to/from a long-tail node keep decrypting without re-running a NodeInfo exchange (the rest of NodeInfoLite rebuilds from traffic in seconds).
- Write:
getOrCreateMeshNode's eviction anddemoteOldestHotNodesToWarm(the over-cap boot migration) callwarmStore.absorb(num, last_heard, key)before the node leaves the header. - Read-back:
getOrCreateMeshNodecallswarmStore.take()to rehydratelast_heard+ key when a warm node is re-admitted;copyPublicKey()falls back to the warm tier so the PKI send path finds keys for evicted peers. - Persistence: nRF52840 uses a 12 KB raw-flash record-ring at
0xEA000(below LittleFS; append + replay + compact-on-rotate, link-guarded bynrf52840_s140_v7.ldandextra_scripts/nrf52_warm_region.py). Everywhere else: a/prefs/warm.datsnapshot flushed bysaveIfDirty()on the node-DB save cadence. - Tunables (
mesh-pb-constants.h):WARM_NODE_COUNT(per-arch;0disables the tier) andMAX_NUM_NODES(hot cap - 120 on nRF52840/generic ESP32 to fit the 28 KB LittleFS; ESP32-S3 picks 100/200/250 at boot from its flash size). Verbose migration/self-care tracing routes throughLOG_MIGRATION, gated byMESHTASTIC_NODEDB_MIGRATION_VERBOSE. MAX_NUM_NODESon native is not in that header and is not a constant.variants/native/portduino{,-buildroot}/variant.hdefine it asportduino_config.MaxNodes- resolved at runtime, default 200, overridable per-host withGeneral: MaxNodesin the portduino YAML.variant.his reached first, so theARCH_PORTDUINObranch inmesh-pb-constants.hnever fires; it is now#error-guarded rather than holding a plausible-looking250. Reading 250 there yields a protected-node cap of 248 when the real one is 198 (numProtectedNodes() < MAX_NUM_NODES - 2), which has already produced one wrong diagnosis. The separate 250 inNodeDB::getMaxNodesAllocatedSize()isNODEDB_MIGRATION_LOAD_CEILING, a decode allowance for files from larger-cap firmware - not a cap.
Satellite caps
Only the freshest MAX_SATELLITE_NODES nodes keep satellite payloads; the rest of the header table carries just the NodeInfoLite. The cap is per-platform: 40 on RAM-constrained parts (nRF52840, generic ESP32) since the four maps live in internal SRAM (not PSRAM, ~408 B/node across the four), and 250 on flash-rich hosts (ESP32-S3, portduino) so every hot node can carry rich data as before the cap existed. enforceSatelliteCaps() trims each map to the cap on load (returns whether it trimmed); evictSatelliteOverCap() trims before each insert. Eviction is by the owning node's hot last_heard (stalest first, demoted/absent nodes rank as last_heard==0); self is never trimmed.
On-boot self-care
NodeDB::nodeDBSelfCare() runs once identity is established (the constructor after key (re)gen, and reloadFromDisk() - not inside loadFromDisk, where getNodeNum() is still 0). It confirms self is present (warns if a non-empty DB is missing us - a foreign/over-cap file), pins self to index 0, demotes/trims only non-self overflow into the warm tier, then rewrites nodes.proto once and only if it healed something - and never while encrypted storage is locked (it would persist placeholder defaults). loadFromDisk deliberately leaves the loaded store untrimmed for this pass.
Sync flow: thin NodeInfo + post-COMPLETE_ID replay (no opt-in)
There is no capability flag and no special "gradient" nonce. The default sync flow is:
- Config / module-config / channel / metadata segments (same as before).
STATE_SEND_OWN_NODEINFO- our own NodeInfo, still bundled with our position and device_metrics (because the replay snapshot excludes our own NodeNum). Emitted viaConvertToNodeInfo(lite).STATE_SEND_OTHER_NODEINFOS- every other peer's NodeInfo, always thin (noposition, nodevice_metrics). Emitted viaConvertToNodeInfoThin(lite).STATE_SEND_FILEMANIFEST→STATE_SEND_COMPLETE_ID- the phone seesconfig_complete_idand treats sync as done.STATE_SEND_PACKETS- live mesh packets, with a trailing replay drain interleaved. The replay drain walks four cached satellite stores in order (positions → telemetry → environment → status) and emits each cached entry as an ordinaryMeshPacketon the matching portnum (POSITION_APP,TELEMETRY_APPdevice + environment variants,NODE_STATUS_APP). These are indistinguishable on the wire from live mesh traffic, so clients need no special handling - any code that already updates UI onPOSITION_APPetc. works.
PhoneAPI::sendConfigComplete() arms replayPhase = REPLAY_PHASE_POSITIONS for default/full sync and SPECIAL_NONCE_ONLY_NODES, while SPECIAL_NONCE_ONLY_CONFIG skips replay. The drain runs inside STATE_SEND_PACKETS via popReplayPacket(), lower priority than live traffic. When all four phases drain, replayPhase flips back to REPLAY_PHASE_IDLE and the snapshot vectors get shrink_to_fited.
STM32WL and any other build with all four MESHTASTIC_EXCLUDE_*DB flags set produces zero replay packets - popReplayPacket advances through each phase in microseconds without emitting anything.
Special nonces that still mean something:
SPECIAL_NONCE_ONLY_CONFIG(69420) - skip node sync entirely, just config.SPECIAL_NONCE_ONLY_NODES(69421) - skip config segments, jump straight toSTATE_SEND_OWN_NODEINFO. Still gets the post-COMPLETE_ID replay drain.
There are no other reserved nonces; everything else is a fresh random want_config_id from the client.
v24 → v25 migration
The legacy migration code lives in src/mesh/NodeDBLegacyMigration.cpp, not in NodeDB.cpp. It owns the meshtastic_NodeDatabase_Legacy callback and NodeDB::migrateLegacyNodeDatabase(). The legacy proto descriptor is protobufs/meshtastic/deviceonly_legacy.proto (only included by the migration TU). The boot path peeks the file's leading version tag, runs the migration if version < 25, then re-saves in v25 layout. The legacy descriptor is scheduled for removal once DEVICESTATE_MIN_VER is bumped.
Read-site rules of thumb
- Never
node->position.X/node->device_metrics.X- those fields no longer exist. Pull from the satellite map viacopyNodePosition/copyNodeTelemetry. - Never
node->user.long_name-long_name,short_name,public_key,hw_model,role,macaddr(gone),is_licensed,is_unmessagableare flat onNodeInfoLite. - Never
node->is_favorite/node->is_ignored/node->via_mqtt/node->is_key_manually_verified- use the bitfield helpers. - Never assume
nodeDB->getMeshNode(num)->position.time- callcopyNodePositionand check the return. - Don't lock
satelliteMutexyourself in renderer code; the copy-out accessors already do.
Unit tests for the conversion layer live in test/test_type_conversions/test_main.cpp (Unity) - bitfield round-trips, long_name truncation, thin-vs-full conversions. Add cases there when extending the schema.
Project Structure
firmware/
├── src/ # Main source code
│ ├── main.cpp # Application entry point
│ ├── mesh/ # Core mesh networking
│ │ ├── NodeDB.* # Node database management
│ │ ├── Router.* # Packet routing
│ │ ├── Channels.* # Channel management
│ │ ├── CryptoEngine.* # AES-CTR (channels) + X25519 ECDH→AES-256-CCM (PKI for DMs/admin)
│ │ ├── *Interface.* # Radio interface implementations
│ │ ├── api/ # WiFi/Ethernet server APIs (ServerAPI, PacketAPI)
│ │ ├── http/ # HTTP server (WebServer, ContentHandler)
│ │ ├── wifi/ # WiFi support (WiFiAPClient)
│ │ ├── eth/ # Ethernet support (ethClient)
│ │ ├── udp/ # UDP multicast
│ │ ├── compression/ # Message compression (unishox2)
│ │ └── generated/ # Protobuf generated code
│ ├── modules/ # Feature modules (Position, Telemetry, etc.)
│ │ └── Telemetry/ # Telemetry subsystem
│ │ └── Sensor/ # 50+ I2C sensor drivers
│ ├── gps/ # GPS handling
│ ├── graphics/ # Display drivers and UI
│ │ └── niche/ # Specialized UIs (InkHUD e-ink framework)
│ ├── platform/ # Platform-specific code (esp32, nrf52, rp2xx0, stm32wl, portduino)
│ ├── input/ # Input device handling (InputBroker, keyboards, buttons)
│ ├── detect/ # I2C hardware auto-detection (80+ device types)
│ ├── motion/ # Accelerometer drivers (BMA423, BMI270, MPU6050, etc.)
│ ├── mqtt/ # MQTT bridge client
│ ├── power/ # Power HAL
│ ├── nimble/ # BLE via NimBLE
│ ├── buzz/ # Audio/notification (buzzer, RTTTL)
│ ├── serialization/ # JSON serialization, COBS encoding
│ ├── watchdog/ # Hardware watchdog thread
│ ├── concurrency/ # Threading utilities (OSThread, Lock)
│ ├── PowerFSM.* # Power finite state machine
│ └── Observer.h # Observer/Observable event pattern
├── variants/ # Hardware variant definitions
│ ├── esp32/ # ESP32 variants
│ ├── esp32s3/ # ESP32-S3 variants
│ ├── esp32c3/ # ESP32-C3 variants
│ ├── esp32c6/ # ESP32-C6 variants
│ ├── nrf52840/ # nRF52 variants
│ ├── rp2040/ # RP2040/RP2350 variants
│ ├── stm32/ # STM32WL variants
│ └── native/ # Linux/Portduino variants
├── protobufs/ # Protocol buffer definitions
├── boards/ # Custom PlatformIO board definitions
├── test/ # Native unit-test suites (count: test/native-suite-count)
└── bin/ # Build and utility scripts
Coding Conventions
Formatting & the trunk toolchain
trunk fmt is the project formatter (trunk_check CI rejects unformatted code). For Claude Code users, .claude/settings.json ships a PostToolUse hook that runs trunk fmt --force on every file the agent writes or edits. The hook is pure sh/grep/sed - no python or jq required - but trunk itself must be able to run:
- Trunk's launcher (
~/.cache/trunk/launcher/trunk, ortrunkon PATH) downloads the CLI version pinned in.trunk/trunk.yamlon first use and again whenever that pin is bumped. The launcher needscurlorwget; without one it fails with "Cannot download… please install curl or wget", and the hook surfaces that as a warning on every write. - No curl/wget available (e.g. a minimal WSL image)? Bootstrap by hand with any Python (PlatformIO bundles one at
~/.platformio/penv/bin/python): downloadhttps://trunk.io/releases/<ver>/trunk-<ver>-linux-x86_64.tar.gzand place thetrunkbinary at~/.cache/trunk/cli/<ver>-linux-x86_64/trunk(chmod +x), where<ver>is thecli.versionfrom.trunk/trunk.yaml. - The hook fails loudly by design (visible warning, non-blocking). Silent no-op formatting hooks hide real breakage - don't re-add
2>/dev/null || truearound the whole thing. - More generally: don't assume a stock Linux userland in hooks or helper scripts - minimal WSL/container images may lack
python3,curl,wget, andjq. Prefer plain sh + coreutils, or PlatformIO's bundled Python for anything heavier.
General Style
- Follow existing code style - run
trunk fmtbefore commits - Prefer
LOG_DEBUG,LOG_INFO,LOG_WARN,LOG_ERRORfor logging - Format node IDs and packet IDs as
0x%08xin logs. This coversNodeNum/PacketIdand theuint32_tpacket fieldsfrom,to,id,dest,source,request_id, andnode_id. They are 32-bit, so 8 hex digits is exact -%08xnever truncates or leaves a value ragged. Do not use%x(variable width) or%0x(a no-op typo for%08x- the0flag does nothing without a width). User-facing display uses!%08x(the!xxxxxxxxconvention), e.g.Applet::hexifyNodeNum. - Do not zero-pad one-byte values to 8.
next_hop,relay_node, and the next-hop hint areuint8_tlast-byte route hints, andchannelis a one-byte hash/index - log these as0x%x(or%d). Padding a byte to0x000000abfalsely implies a full node number. The same goes for I2C addresses, register values, flags/bitmasks, and error/reason codes: they are not IDs, so leave them0x%x. - Use
assert()for invariants that should never fail - C++17 features are available (
std::optional, structured bindings,if constexpr, etc.) - Keep code comments minimal - one or two lines, max. Comment only when the why isn't obvious from the code; never restate what the next line does. No multi-paragraph block comments explaining straightforward changes. The diff and commit message carry the rationale; the code carries the behavior.
- Use
Throttlefor time-based rate limiting, not rawmillis()math.src/mesh/Throttle.hprovidesThrottle::isWithinTimespanMs(lastMs, intervalMs)(returns true while inside the cooldown) andThrottle::execute(&lastMs, intervalMs, func)(function-pointer form that updates the timestamp on fire). Use these for any "did N ms pass since X" check - rawmillis() > lastMs + Nis rollover-unsafe (breaks after ~49.7 days) and inconsistent with the rest of the codebase. The helpers computenow - lastMswith unsigned subtraction, which wraps correctly.
Naming Conventions
- Classes:
PascalCase(e.g.,PositionModule,NodeDB) - Functions/Methods:
camelCase(e.g.,sendOurPosition,getNodeNum) - Constants/Defines:
UPPER_SNAKE_CASE(e.g.,MAX_INTERVAL,ONE_DAY) - Member variables:
camelCase(e.g.,lastGpsSend,nodeDB) - Config defines:
USERPREFS_*for user-configurable options
Key Patterns
Module System
Modules use a three-tier class hierarchy:
MeshModule- Base class. ImplementwantPacket()andhandleReceived(). ReturnsProcessMessage::STOPorProcessMessage::CONTINUE.SinglePortModule- Handles a single portnum. SimplifiedwantPacket()that checksdecoded.portnum.ProtobufModule<T>- Template for protobuf-based modules. Handles encoding/decoding automatically.
Most modules also inherit from OSThread for periodic tasks (the "mixin" pattern):
class MyModule : public ProtobufModule<meshtastic_MyMessage>, private concurrency::OSThread
{
public:
MyModule();
protected:
virtual bool handleReceivedProtobuf(const meshtastic_MeshPacket &mp, meshtastic_MyMessage *msg) override;
virtual meshtastic_MeshPacket *allocReply() override; // Generate response packets
virtual int32_t runOnce() override; // Periodic task (returns next interval in ms)
virtual bool alterReceivedProtobuf(meshtastic_MeshPacket &mp, meshtastic_MyMessage *msg); // Modify in-flight
virtual bool wantUIFrame(); // Request a UI display frame
};
Modules are registered in src/modules/Modules.cpp guarded by MESHTASTIC_EXCLUDE_* flags.
Observer/Observable Pattern
Event-driven communication between subsystems uses src/Observer.h:
// Observable emits events
Observable<const meshtastic::Status *> newStatus;
newStatus.notifyObservers(&status);
// Observer receives events via callback
CallbackObserver<MyClass, const meshtastic::Status *> statusObserver =
CallbackObserver<MyClass, const meshtastic::Status *>(this, &MyClass::handleStatusUpdate);
Configuration Access
config.*- Device configuration (LoRa, position, power, etc.)moduleConfig.*- Module-specific configurationchannels.*- Channel configuration and managementowner- Device owner infomyNodeInfo- Local node info
Default Values
Use the Default class helpers in src/mesh/Default.h:
Default::getConfiguredOrDefaultMs(configured, default)- Returns ms, using default if configured is 0Default::getConfiguredOrDefault(configured, default)- Generic configured/default getterDefault::getConfiguredOrMinimumValue(configured, min)- Enforces minimum valuesDefault::getConfiguredOrDefaultMsScaled(configured, default, numNodes)- Scales based on network size
Thread Safety
- Use
concurrency::Lockandconcurrency::LockGuardfor mutex protection - Radio SPI access uses
SPILock - Prefer
OSThreadfor background tasks
Hardware Detection
src/detect/ScanI2C automatically enumerates 80+ I2C device types at boot including displays, sensors, RTCs, keyboards, PMUs, and touch controllers. This drives automatic initialization of the correct drivers.
Graphics/UI System
Multiple display driver families in src/graphics/:
- OLED: SSD1306, SH1106, ST7567
- TFT: TFTDisplay (LovyanGFX-based)
- E-Ink: EInkDisplay2, EInkDynamicDisplay, EInkParallelDisplay
InkHUD (src/graphics/niche/InkHUD/) is an event-driven e-ink UI framework:
- Applet-based architecture - modular display tiles
- Read-only, static display optimized for minimal refreshes and low power
- Configured per-variant via
nicheGraphics.h - Separate PlatformIO config:
src/graphics/niche/InkHUD/PlatformioConfig.ini
Input System
src/input/InputBroker is the centralized input event dispatcher. Supports multiple input sources: buttons, keyboards (BBQ10, Cardputer, TCA8418), touch screens, rotary encoders, and matrix keyboards.
Power Management
src/PowerFSM.* implements a finite state machine with states: stateON, statePOWER, stateSERIAL, stateDARK. Key events: EVENT_PRESS, EVENT_WAKE_TIMER, EVENT_LOW_BATTERY, EVENT_RECEIVED_MSG, EVENT_SHUTDOWN. Conditionally excluded with MESHTASTIC_EXCLUDE_POWER_FSM (falls back to FakeFsm).
Motion Sensors
src/motion/AccelerometerThread provides background motion monitoring with automatic screen wake and double-tap button press detection. Supports 10+ accelerometer/gyroscope chips (BMA423, BMI270, MPU6050, LIS3DH, LSM6DS3, STK8XXX, QMA6100P, ICM20948, BMX160).
Telemetry Sensor Library
src/modules/Telemetry/Sensor/ contains 50+ I2C sensor drivers organized by category:
- Power monitoring: INA219/226/260/3221, MAX17048
- Environmental: BME280/680, SCD4X (CO₂), SEN5X (particulate)
- Humidity/Temperature: SHT3X/4X, AHT10, MCP9808, MLX90614
- Light: BH1750, TSL2561/2591, VEML7700, LTR390UV, OPT3001
- Air quality: PMSA003I, SFA30
- Specialized: CGRadSens (radiation), NAU7802 (weight scale)
API/Networking
src/mesh/api/ provides a template-based ServerAPI for client communication over WiFi (WiFiServerAPI) and Ethernet (ethServerAPI). Default port: 4403. HTTP server in src/mesh/http/. JSON serialization in src/serialization/MeshPacketSerializer.
Hardware Variants
Each hardware variant has:
variant.h- Pin definitions and hardware capabilitiesplatformio.ini- Build configuration- Optional:
pins_arduino.h,rfswitch.h,nicheGraphics.h(for InkHUD variants)
Key defines in variant.h:
#define USE_SX1262 // Radio chip selection
#define HAS_GPS 1 // Hardware capabilities
#define HAS_SCREEN 1 // Display present
#define LORA_CS 36 // Pin assignments
#define SX126X_DIO1 14 // Radio-specific pins
Protobuf Messages
- Defined in
protobufs/meshtastic/*.proto(~32 proto files) - Generated code in
src/mesh/generated/meshtastic/ - Regenerate with
bin/regen-protos.sh - Message types prefixed with
meshtastic_ - Nanopb
.optionsfiles control field sizes and encoding - Never edit or commit files under
src/mesh/generated/. They are regenerated from themeshtastic/protobufssubmodule by theupdate_protobufs.ymlGitHub Action and any hand edits will be overwritten - guaranteed merge conflict on the next sync. To change a wire format, open a PR against the protobufs repo first; the workflow then re-runsbin/regen-protos.shand opens a PR here with the regenerated sources.
Conditional Compilation
#if !MESHTASTIC_EXCLUDE_GPS // Feature exclusion
#if !MESHTASTIC_EXCLUDE_WIFI // Network feature exclusion
#if !MESHTASTIC_EXCLUDE_BLUETOOTH // BLE exclusion
#if !MESHTASTIC_EXCLUDE_POWER_FSM // Power FSM exclusion
#ifdef ARCH_ESP32 // Architecture-specific
#ifdef ARCH_NRF52 // Nordic platform
#ifdef ARCH_RP2040 // Raspberry Pi Pico
#ifdef ARCH_PORTDUINO // Linux native
#if defined(USE_SX1262) // Radio-specific
#ifdef HAS_SCREEN // Hardware capability
#if USERPREFS_EVENT_MODE // User preferences
Build System
Agent Tooling Baseline
Mirror counterpart: AGENTS.md under Agent Tooling Baseline.
To reduce avoidable agent mistakes, assume these tools are available (or install them before significant repo work):
- Required CLI basics:
bash,git,find,grep,sed,awk,xargs - Strongly recommended:
rg(ripgrep) for fast file/text search,jqfor JSON processing - Build/test tools:
python3,pip, virtualenv (python3 -m venv),platformio(pio) - Containerized native testing:
docker(fallback for non-Linux hosts; macOS can also build natively viapio run -e native-macos)
Fallback expectations for agents:
- If
rgis unavailable, usefind+grepinstead of failing. - For native tests on hosts without Linux deps, prefer
./bin/test-native-docker.sh. - The simulator helper script is
./bin/test-simulator.sh.
Uses PlatformIO with custom scripts:
bin/platformio-pre.py- Pre-build scriptbin/platformio-custom.py- Custom build logic, manifest generation
Build commands:
pio run -e tbeam # Build specific target
pio run -e tbeam -t upload # Build and upload
pio run -e native # Build native/Linux version
pio run -e native-macos # Build headless macOS meshtasticd (Homebrew prereqs in variants/native/portduino/platformio.ini)
Build Manifest
bin/platformio-custom.py emits a build manifest with metadata:
hasMui,hasInkHud- UI capability flags (overridable viacustom_meshtastic_has_mui,custom_meshtastic_has_ink_hud)- Architecture normalization (e.g.,
esp32s3→esp32-s3for API compatibility)
Common Tasks
Adding a New Module
- Create
src/modules/MyModule.cppand.h - Inherit from appropriate base class (
MeshModule,SinglePortModule, orProtobufModule<T>) - Mix in
concurrency::OSThreadif periodic work is needed - Register in
src/modules/Modules.cppguarded by#if !MESHTASTIC_EXCLUDE_MYMODULE - Add protobuf messages if needed in
protobufs/meshtastic/ - Add test suite in
test/test_mymodule/if applicable
Adding a New Hardware Variant
- Create directory under
variants/<arch>/<name>/ - Add
variant.hwith pin definitions and hardware capability defines - Add
platformio.iniwith build config - useextendsto reference common base (e.g.,esp32s3_base) - Set
board_level(required -releasefor a normal variant; see "Build Matrix Generation") - Set
custom_meshtastic_support_level(1-3) and the othercustom_meshtastic_*metadata - For e-ink displays, add
nicheGraphics.hfor InkHUD configuration
Adding a New Telemetry Sensor
- Create driver in
src/modules/Telemetry/Sensor/following existing sensor pattern - Register I2C address in
src/detect/ScanI2Cfor auto-detection - Integrate with the appropriate telemetry module (Environment, Health, Power, AirQuality)
- Add proto fields in
protobufs/meshtastic/telemetry.protoif new data types are needed
Modifying Configuration Defaults
- Check
src/mesh/Default.hfor default value defines - Check
src/mesh/NodeDB.cppfor initialization logic - Consider
isDefaultChannel()checks for public channel restrictions
Important Considerations
Traffic Management
The mesh network has limited bandwidth. When modifying broadcast intervals:
- Respect minimum intervals on default/public channels
- Use
Default::getConfiguredOrMinimumValue()to enforce minimums - Consider
numOnlineNodesscaling for congestion control
Power Management
Many devices are battery-powered:
- Use
IF_ROUTER(routerVal, normalVal)for role-based defaults - Check
config.power.is_power_savingfor power-saving modes - Implement proper
sleep()methods in radio interfaces
Channel Security
channels.isDefaultChannel(index)- Check if using default/public settings- Default channels get stricter rate limits to prevent abuse
- Private channels may have relaxed limits
GitHub Actions CI/CD
The project uses GitHub Actions extensively for CI/CD. Key workflows are in .github/workflows/:
Core CI Workflows
-
main_matrix.yml- Main CI pipeline, runs on push tomaster/developand PRs- Uses
bin/generate_ci_matrix.pyto dynamically generate build targets - Builds all supported hardware variants
- PRs build a subset (
--level pr) for faster feedback
- Uses
-
trunk_check.yml- Code quality checks on PRs- Runs Trunk.io for linting and formatting
- Must pass before merge
-
tests.yml- End-to-end and hardware tests- Runs daily on schedule
- Includes native tests and hardware-in-the-loop testing
-
test_native.yml- Native platform unit tests- Runs
pio test -e native
- Runs
Release Workflows
-
release_channels.yml- Triggered on GitHub release publish- Builds Docker images
- Packages for PPA (Ubuntu), OBS (openSUSE), and COPR (Fedora)
- Handles Alpha/Beta/Stable release channels
-
nightly.yml- Nightly builds from develop branch -
docker_build.yml/docker_manifest.yml- Docker image builds
Build Matrix Generation
The CI uses bin/generate_ci_matrix.py to dynamically select which targets to build:
# Generate full build matrix
./bin/generate_ci_matrix.py all
# Generate PR-level matrix (subset for faster builds)
./bin/generate_ci_matrix.py all --level pr
Every variant env must declare a board_level in its platformio.ini; the matrix
generator exits non-zero if any env is missing it or uses an unrecognized value:
board_level = pr- Smallest subset, built on every PR (and in every larger matrix)board_level = release- The full release matrix, built on push / schedule /workflow_dispatchboard_level = extra- Opt-in only, built when explicitly requested via--level extra
custom_meshtastic_support_level (1-3) is not part of this filtering. It is variant
metadata that bin/platformio-custom.py emits as supportLevel in the generated
hardware list; changing it does not change which targets CI builds.
Running Workflows Locally
Most workflows can be triggered manually via workflow_dispatch for testing.
Testing
Native unit tests (C++)
Unit tests in test/ directory. The canonical suite count is in test/native-suite-count, cross-checked against test/test_* on every full run and by the suite-count-check CI job. Never state the count as a literal anywhere else - point at that file. The list below is a partial description of what suites cover, not an inventory:
test_admin_radio/- LoRa region/config validation, AdminModule dispatch, node-DB metadata savestest_fscommon_getfiles/- bounded file-manifest walk (cap, depth, truncation reporting)test_atak/- ATAK integrationtest_crypto/- Cryptographytest_default/- Default configurationtest_hop_scaling/- Hop scaling histogram and required-hop logictest_http_content_handler/- HTTP handlingtest_mac_from_string/- MAC address parsingtest_mesh_module/- Module frameworktest_meshpacket_serializer/- Packet serializationtest_mqtt/- MQTT integrationtest_nexthop_routing/- Next-hop routing logictest_nodedb_blocked/- NodeDB blocked-node handlingtest_packet_history/- Packet history trackingtest_packet_signing/- Packet signingtest_position_module/- Position module behaviourtest_position_precision/- Position precision helperstest_radio/- Radio interfacetest_rtc/- RTC / time handlingtest_serial/- Serial communicationtest_tak_config/- TAK (ATAK) team/role value fidelity through set/save/load/gettest_module_config/- every ModuleConfig submessage survives admin set -> save -> load -> gettest_traffic_management/- Traffic management (dedup, rate-limit, hop-trim, role exceptions)test_transmit_history/- Retransmission trackingtest_type_conversions/- NodeDB v25 type conversion (bitfield round-trips, NodeInfoLite)test_utf8/- UTF-8 utilitiestest_warm_store/- Warm-tier node store
Preferred run command - bin/run-tests.sh (defaults to the coverage env; emits a machine-readable verdict on the final line; update test/native-suite-count when adding or removing suites):
./bin/run-tests.sh # all suites
./bin/run-tests.sh -f test_traffic_management # single suite (yields FILTERED, not GREEN)
The harness is Linux-only, and rejects anything else. bin/run-tests.sh needs bash 4+ and GNU coreutils/find (find -printf, md5sum, -executable), so it exits 2 on a non-Linux uname rather than degrade quietly - a state check that silently mis-hashes a sandbox still prints a verdict, and that verdict would be worthless. The native-macos PlatformIO env is a build target for meshtasticd, not a test host. On macOS or Windows use ./bin/test-native-docker.sh.
Sanitizer coverage is per env, and only one env has any. coverage (the default) adds gcov + ASan/LSan on top of native. native itself has none - verified, zero ASan symbols in the built binary. A -e native run is not sanitized, so do not reason from "run-tests.sh uses ASan" when you passed -e native.
A signal name from the runner is not a crash. exit(UNITY_END()) returns the failure count, and PlatformIO's native runner renders a non-zero exit code as a POSIX signal - 4 failures prints Program received signal SIGILL, 5 prints SIGTRAP, and the suite is reported [ERRORED] instead of [FAILED]. Check the exit code against the failure count before theorising about memory bugs; confirm any real crash under a debugger.
Suite order is randomisable. ./bin/run-tests.sh --shuffle runs suites in a seeded random order; --seed <n> replays one. The seed defaults to the commit SHA (deterministic per commit, varied across commits), is printed at the start and on the RESULT: line, and the full order is printed on failure. CI shuffles its area order the same way, seeded from GITHUB_SHA. A single green seed is not evidence of order independence.
-f is not a gate. A filtered run can pass while a full run fails, because filtering removes the suites that create the state a later suite trips over. Iterate with -f; gate on a full run.
Exit codes and verdicts (exact counts will vary; examples below are illustrative):
| Exit | Verdict | Meaning |
|---|---|---|
| 0 | GREEN |
All canonical suites ran, all passed, no ignored test cases |
| 1 | RED |
At least one failure, build error, or sanitizer fault |
| 2 | AMBER |
All that ran passed, but something was lost or unexplained: a suite silently went missing on a full run, individual test cases were skipped (TEST_IGNORE), test/native-suite-count disagrees with the test/ directory count, or a suite left behind shared state it does not declare |
| 3 | FILTERED |
A -f run completed cleanly; suites outside the filter were intentionally not run |
Examples - exact counts will vary by suite count and env:
# GREEN: all suites ran and passed
RESULT: GREEN N/N suites passed [canonical: N/N]
# RED: real test failure
RESULT: RED 1 failed
# RED: sanitizer exit-time abort (all tests passed but process aborted at exit)
RESULT: RED exit-time abort (tests passed; likely sanitizer - see hint above)
# AMBER: native-suite-count disagrees with test/ directory count (too low)
RESULT: AMBER test/ has 24 suite directories but native-suite-count says 5 - update test/native-suite-count after registering new suites
# AMBER: native-suite-count disagrees with test/ directory count (too high)
RESULT: AMBER test/ has 24 suite directories but native-suite-count says 99 - update test/native-suite-count after removing suites
# FILTERED: single suite run completed cleanly
RESULT: FILTERED 1/24 suites ran (not run: test_admin_radio test_atak …) - filtered: test_serial [canonical: 1/24]
Copilot interface note: When running tests via the Copilot chat interface, edits made through the chat may not be reflected in the on-disk files that the test binary reads. If tests pass in chat but fail locally (or vice versa), verify the files on disk match what you expect before trusting the result. Always confirm with a local terminal run.
Raw pio test (no sanitizers, no verdict logic) - use only when you need to override the env:
~/.platformio/penv/bin/python -m platformio test -e native -f test_your_suite > /tmp/test_out.txt 2>&1
grep -E 'error:|PASS|FAIL|succeeded|failed' /tmp/test_out.txt
tail -15 /tmp/test_out.txt
Do not pipe pio test - line-buffering makes the terminal appear hung and hides build errors.
Simulation testing: bin/test-simulator.sh
Quick entry point for new test modules: test/README.md (native unit-test authoring guide, skeleton, pitfalls, and setup checklist).
Shared state: every suite gets a clean sandbox
Each suite runs inside its own scratch $HOME (bin/pio-test-isolate.sh, wired in per env as test_testing_command, so a bare pio test and CI get it too). State never crosses a suite boundary. Mutation inside a suite is free; carrying state out of one is impossible by construction, not by policy.
The state in question lives in ~/.portduino/default/prefs/ - nodes.proto, config.proto, channels.proto, module.proto, device.proto, warm.dat, transmit_history.dat. NodeDB's constructor calls loadFromDisk(), so any suite that constructs one reads it, and several NodeDB paths (removeNodeByNum(), resetNodes(), nodeDBSelfCare(), and the constructor when the file is absent) write it without being asked.
Two orthogonal axes: PASS/FAIL x CLEAN/DIRTY.
- CLEAN - nothing changed, or everything that changed is declared.
- DIRTY - an undeclared path changed. Graded AMBER: with isolation in place it means "undeclared", not "dangerous".
- MISSING - a declared write did not happen. A warning only; it catches persistence that silently stopped working.
Declare deliberate writes in test/state-manifest.tsv - one central file, <suite> / <flags> / <reason>, with the reason mandatory and reviewed on change. Central so every opt-out is visible in one diffable list; per-suite files hide growth. run-tests.sh prints how many suites declare non-default handling on every run.
| Flag | Meaning |
|---|---|
| (no entry) | the default: fresh state in, contents discarded out |
writes=<a,b> |
files this suite mutates on purpose; matched on the path relative to the sandbox $HOME or just the basename |
state=per-suite |
state persists across this suite's own test cases (persistence round-trips, migration ladders). Only the suite boundary is checked; the default is per-test, which names the exact test that dirtied things |
No flag grants cross-suite carry. A suite that needs another suite's output needs an explicit fixture, not inheritance.
./bin/run-tests.sh --write-manifest prints the entries a run would need, for a human to paste and justify - it never applies them, and neither does CI. bin/test-state-check.sh is the checker's own self-test: fixtures asserting CLEAN / CLEAN / DIRTY / MISSING, plus the before-empty assertion.
Hardware-in-the-loop tests (meshtastic-mcp)
Separate pytest suite that exercises real USB-connected Meshtastic devices. It now lives in the standalone meshtastic-mcp repo, run against a firmware checkout via MESHTASTIC_FIRMWARE_ROOT. See the MCP Server & Hardware Test Harness section below for invocation, tier layout, and agent usage rules.
MCP Server & Hardware Test Harness
The firmware-aware MCP server plus its pytest-based integration suite now live in the standalone meshtastic-mcp repo. AI agents that speak MCP get a well-defined tool surface for flashing, configuring, and inspecting physical Meshtastic devices - use it instead of hand-rolling pio or meshtastic --port calls where possible. The meshtastic-mcp repo's README is the operator-facing setup doc; this section is the agent-facing usage contract.
The repo registers the server via .mcp.json at the repo root - Claude Code / Copilot pick it up automatically and run it through uvx --from git+https://github.com/meshtastic/meshtastic-mcp meshtastic-mcp, so the MCP tools work with no local build. To run the pytest hardware harness instead, clone meshtastic-mcp and point MESHTASTIC_FIRMWARE_ROOT at this firmware checkout.
When to use which surface
| Goal | Tool |
|---|---|
| Find a connected device | mcp__meshtastic__list_devices |
| Read a live node's config/state | mcp__meshtastic__device_info, list_nodes, get_config |
| Mutate a device (owner, region, channels, reboot) | set_owner, set_config, set_channel_url, reboot, shutdown, factory_reset - all require confirm=True |
| Flash firmware to a variant | pio_flash (any arch) or erase_and_flash (ESP32 factory install) |
| Stream serial logs while debugging | serial_open → serial_read loop → serial_close |
Administer userPrefs.jsonc build-time constants |
userprefs_get, userprefs_set, userprefs_reset, userprefs_manifest |
| Run the regression suite | ./run-tests.sh from a meshtastic-mcp checkout (or /test slash command) |
| Diagnose a specific device | /diagnose [role] slash command (read-only) |
| Triage a flaky test | /repro <node-id> [count] slash command |
One MCP call per port at a time. SerialInterface holds an exclusive OS-level lock on the serial port for its lifetime. If a serial_* session is open on /dev/cu.usbmodem101, calling device_info on the same port will fail fast pointing at the active session. Sequence calls: open → read/mutate → close, then next device. Never parallelize tool calls on the same port.
MCP tool surface (44 tools)
Grouped by purpose. Full argument shapes in the meshtastic-mcp repo's README; a few high-value signatures are called out here.
- Discovery & metadata:
list_devices,list_boards,get_board - Build & flash:
build,clean,pio_flash,erase_and_flash(ESP32 only),update_flash(ESP32 OTA),touch_1200bps - Serial sessions (long-running, 10k-line ring buffer):
serial_open,serial_read,serial_list,serial_close - Device reads:
device_info,list_nodes - Device writes:
set_owner,get_config,set_config,get_channel_url,set_channel_url,send_text,send_input_event(inject a button/key press via the firmware's InputBroker),inject_frame(inject an over-the-air-style frame into the RX pipeline - see below),set_debug_log_api; destructive/power-state writes requireconfirm=True:reboot,shutdown,factory_reset - userPrefs admin (build-time constants, not runtime config):
userprefs_get,userprefs_set,userprefs_reset,userprefs_manifest,userprefs_testing_profile - Vendor escape hatches:
esptool_chip_info,esptool_erase_flash,esptool_raw,nrfutil_dfu,nrfutil_raw,picotool_info,picotool_load,picotool_raw - USB power control (via
uhubctl, per-port PPPS toggle):uhubctl_list(read-only),uhubctl_power(action='on'|'off', confirm=True),uhubctl_cycle(delay_s, confirm=True). Target by raw(location, port)or byrole("nrf52","esp32s3"); role lookup checksMESHTASTIC_UHUBCTL_LOCATION_<ROLE>+_PORT_<ROLE>env vars first, falls back to VID auto-detection. - Observability (UI tier + operator ad-hoc):
capture_screen(role, ocr=True)- grabs a USB-webcam frame of the device OLED and optionally OCRs it. Requiresmeshtastic-mcp[ui]extras (opencv-python-headless,easyocr) andMESHTASTIC_UI_CAMERA_DEVICE_<ROLE>env var; falls through to a 1×1 black PNGNullBackendwhen unconfigured.
confirm=True is a tool-level gate on top of whatever permission prompt your MCP host shows. Don't bypass it by asking the host to auto-approve - it exists specifically because MCP hosts sometimes remember "always allow this tool" and that's dangerous for factory_reset, erase_and_flash, uhubctl_power(action='off'), and uhubctl_cycle.
TCP / native-host nodes. Setting MESHTASTIC_MCP_TCP_HOST=<host[:port]> makes list_devices surface a meshtasticd daemon (e.g. the native-macos build) as a synthetic tcp://host:port entry, and connect() routes through meshtastic.tcp_interface.TCPInterface instead of SerialInterface. Every read/write/admin tool that flows through connect() works against the daemon transparently. USB-only tools (pio_flash, erase_and_flash, update_flash, touch_1200bps, serial_open, esptool_*, nrfutil_*, picotool_*) raise a clear ConnectionError when handed a tcp:// port; pio_flash against a native* env raises a FlashError (no upload step - use build and run the binary directly). The pytest harness still assumes USB-attached devices per role; TCP-aware fixtures are deferred. See the meshtastic-mcp repo's README § "TCP / native-host nodes".
Frame injection: testing the off-air receive path
The toRadio API can only inject locally-originated traffic - the firmware forces p.from = 0 in MeshService::handleToRadio, which bypasses the from != 0 receive path and everything gated on it (remote admin authorization, the admin session-passkey check, hop handling, promiscuous sniffing). To exercise those paths you either need a second transmitting radio, or frame injection: a build-flagged seam that delivers a client-supplied frame into the real RX pipeline as if it arrived off the LoRa chip.
- Firmware: build with
-D MESHTASTIC_ENABLE_FRAME_INJECTION=1(src/configuration.h, off by default - it forges over-the-air traffic and must never ship enabled).MeshService::injectAsReceivedextends the existing portduinoSimRadioSIMULATOR_APPpath to real hardware: it unwraps aCompressedenvelope (portnum == UNKNOWN_APP→ verbatim ciphertext the router decrypts; else → decoded payload for that portnum), then callsrouter->enqueueReceivedMessage()- the exact entry pointRadioLibInterface::handleReceiveInterruptuses. Injection is reached before thep.from = 0line, so a forged sender survives;from == 0is dropped to match real RX. - Host: drive it with the meshtastic-mcp
inject_frametool (orcli/meshinject.py). The crafter replicates channel crypto (default-PSK expansion,xorHashchannel hash, AES-CTR with thepacketId|from|0nonce), so an encrypted frame decrypts on-device as if received. Modes:text,raw,admin(+pki/public_keyfor the PKC-admin path),ciphertext,fuzz(malformed-frame decode-path/crash testing). - Example - remote-admin session-key repro: set the target's
admin_key[0]to a key you hold, then inject anadminset_owner withpki=true, that key, and a stalesession_hex. The node logsPKC admin payload with authorized sender key→Expected session key: 00…→Admin message without session_key!- the exact ndoo scenario, on real silicon. Capture logs viaset_debug_log_apion the same connection. - nRF52 gotcha: the USB CDC wedges under rapid
SerialInterfaceopen/close churn (unrelated to injection) - keep setup + inject + log-capture on one connection; recover a hung board via a 1200 bps-touch DFU reflash.
Hardware test suite (run-tests.sh, from a meshtastic-mcp checkout)
The wrapper auto-detects connected devices (VID → role map: 0x239A → nrf52, 0x303A/0x10C4 → esp32s3), maps each role to a PlatformIO env (nrf52 → rak4631, esp32s3 → heltec-v3, overridable via MESHTASTIC_MCP_ENV_<ROLE>), then invokes pytest. Zero pre-flight config needed from the operator.
Suite tiers (collected + run in this order via pytest_collection_modifyitems):
tests/unit/- pure Python (boards parse, pio wrapper, userPrefs parse, testing profile, uhubctl parser). No hardware.tests/test_00_bake.py- flashes each detected device with currentuserPrefs.jsoncmerged with the session's test profile. Has its own skip-if-already-baked check comparing region + primary channel to the session profile; skips cheaply on warm devices.tests/mesh/- multi-device mesh: bidirectional send, broadcast delivery, direct-with-ACK, mesh formation within 60s. Parametrized[nrf52->esp32s3]and[esp32s3->nrf52]. Includestest_peer_offline_recoverywhich uses uhubctl to physically power off one peer mid-conversation (requires uhubctl; skips without).tests/telemetry/-DEVICE_METRICS_APPbroadcast timing.tests/monitor/- boot-log panic check.tests/recovery/-uhubctlpower-cycle round-trip + NVS persistence across hard reset. Requiresuhubctlinstalled and a PPPS-capable hub; entire tier auto-skips otherwise.tests/ui/- input-broker-driven screen navigation with camera + OCR evidence.tests/fleet/- PSK seed session isolation.tests/admin/- channel URL roundtrip, owner persistence across reboot.tests/provisioning/- region + modem + slot bake, admin key presence,UNSETregion blocks TX, userPrefs survive factory reset.
Invocation patterns:
# run from a meshtastic-mcp checkout, with MESHTASTIC_FIRMWARE_ROOT=/path/to/firmware
./run-tests.sh # full suite (auto-bake-if-needed)
./run-tests.sh --force-bake # reflash before testing
./run-tests.sh --assume-baked # skip bake (caller vouches for device state)
./run-tests.sh tests/mesh # one tier
./run-tests.sh tests/mesh/test_direct_with_ack.py # one file
./run-tests.sh -k telemetry # name filter
No hardware detected? The wrapper auto-narrows to tests/unit/ only and prints detected hub : (none) in the pre-flight header. Agents interpreting the output should call this out explicitly - a 52-test green run without hardware is qualitatively different from a 12-unit-test green run.
Artifacts every run produces:
tests/report.html- self-contained pytest-html. Each test gets aMeshtastic debugsection with the tail of firmware log + device state dump. Open this first on failures; it's the canonical evidence source.tests/junit.xml- CI-parseable.tests/reportlog.jsonl- pytest-reportlog stream ($report_typekeyed JSONL). Consumed by the live TUI.tests/fwlog.jsonl- firmware log mirror from themeshtastic.log.linepubsub topic. Populated by the_firmware_log_streamautouse session fixture.
Live TUI (meshtastic-mcp-test-tui)
A Textual-based live view that wraps run-tests.sh. Tails reportlog for per-test state, streams firmware logs, polls device state at startup + post-run (gated out of the active run because hub_devices holds exclusive port locks). Key bindings:
| Key | Action |
|---|---|
r |
re-run focused test (leaf → that node id; internal node → directory or -k) |
f |
filter tree by substring |
d |
failure detail modal (pulls longrepr + captured stdout from the reportlog) |
g |
export reproducer bundle (tar.gz with README, test_report.json, time-filtered fwlog, devices.json, env.json) |
l |
toggle firmware log pane |
x |
tool coverage modal |
c |
cross-run history sparkline |
q |
quit (SIGINT → SIGTERM → SIGKILL escalation, 5-s windows each) |
Launch:
# from a meshtastic-mcp checkout (MESHTASTIC_FIRMWARE_ROOT set)
.venv/bin/meshtastic-mcp-test-tui # full suite
.venv/bin/meshtastic-mcp-test-tui tests/mesh # args pass through to pytest
The plain CLI stays primary; the TUI is for operators who want a live dashboard. Both consume the same run-tests.sh.
Slash commands (Claude Code + Copilot)
Three AI-assisted workflows wrap the test harness. Claude Code operators get /test, /diagnose, /repro; Copilot operators get /mcp-test, /mcp-diagnose, /mcp-repro. Bodies:
.claude/commands/{test,diagnose,repro}.md.github/prompts/mcp-{test,diagnose,repro}.prompt.md
.claude/commands/README.md is the index.
House rules for agents running these prompts:
- Interpret failures, don't just echo them. Pull firmware log tails from
report.htmland classify each failure as transient / environmental / regression. Use the exact format in.claude/commands/test.md. - No destructive writes without operator approval. Any skill that could reflash, factory-reset, or reboot a device must describe the action and stop. The operator authorizes.
- Sequential MCP calls per port. See above.
- "Unknown" is a valid classification. If evidence doesn't support a root cause, say so and list what would disambiguate. Do not invent.
Key fixtures (test authors + agents debugging)
tests/conftest.py (in the meshtastic-mcp checkout) provides:
_session_userprefs(autouse session) - snapshotsuserPrefs.jsoncat session start, merges the session test profile viauserprefs.merge_active(test_profile), restores at teardown. Four layers of safety: pytest teardown +atexit+ sidecar file (userPrefs.jsonc.mcp-session-bak) + startup self-heal inrun-tests.sh. Do not edituserPrefs.jsoncfrom inside a test._firmware_log_stream(autouse session) - subscribes tomeshtastic.log.linepubsub on every connectedSerialInterfaceand mirrors lines totests/fwlog.jsonl. Drives the TUI firmware-log pane._debug_log_buffer(autouse per-test) - captures last 200 firmware log lines + device state for attachment to the pytest-htmlMeshtastic debugsection on failure.hub_devices(session) -dict[role, SerialInterface]with session-long exclusive port locks. Reason the TUI's device poller is gated to startup + post-run only.baked_mesh- parametrized mesh-pair fixture; depends ontest_00_bake.pytest_generate_testsinconftest.pyauto-generates[nrf52->esp32s3]and[esp32s3->nrf52]variants.test_profile- session-scoped dict: region, primary channel, admin key, PSK seed. Derived fromMESHTASTIC_MCP_SEED(defaults tomcp-<user>-<host>).
Firmware integration points tied to the test harness
Two firmware changes exist specifically so the test harness works reliably. Keep these in mind when touching related code.
src/mesh/StreamAPI.cpp+StreamAPI.h-emitLogRecorduses a dedicatedfromRadioScratchLog+txBufLogpair and aconcurrency::Lock streamLock. Before this fix,debug_log_api_enabled=truewould tearFromRadioprotobufs on the serial transport becauseemitTxBufferandemitLogRecordshared a single scratch buffer. The conftest enables the log stream session-wide; without this fix the device would corrupt its own FromRadio replies mid-session.src/mesh/PhoneAPI.cpp-ToRadioHeartbeat(nonce=1)triggersnodeInfoModule->sendOurNodeInfo(NODENUM_BROADCAST, true, 0, true)for serial clients, mirroring the pre-existing behavior for TCP/UDP clients inPacketAPI.cpp. The mesh tests rely on this to force a NodeInfo broadcast right after connect so the peer discovers them before the test's first assertion.
If you're modifying StreamAPI, PhoneAPI, NodeInfoModule, or userPrefs flow, run ./run-tests.sh (from a meshtastic-mcp checkout, with MESHTASTIC_FIRMWARE_ROOT pointed here) at minimum before asking for review.
Recovery playbooks
| Symptom | First check | Fix |
|---|---|---|
userPrefs.jsonc dirty after test run |
git status --porcelain userPrefs.jsonc |
If non-empty, re-run ./run-tests.sh (from a meshtastic-mcp checkout) once - the pre-flight self-heal restores from sidecar. If still dirty, git checkout userPrefs.jsonc. |
| Port busy / wedged CP2102 on macOS | lsof /dev/cu.usbserial-0001 |
Kill the holder. USB replug if the kernel still reports busy. Often a stale pio device monitor or zombie meshtastic_mcp process. |
| nRF52 appears unresponsive | list_devices shows VID 0x239A but device_info times out |
touch_1200bps(port=...) drops it into the DFU bootloader → pio_flash re-installs. |
| Device fully wedged (Guru Meditation, frozen CDC, no DFU) | list_devices shows the VID but every admin call times out |
uhubctl_cycle(role="nrf52", confirm=True) hard-power-cycles the port via USB hub PPPS. baked_single's auto-recovery hook does this once automatically if uhubctl is installed. Falls back to physical replug if no PPPS hub. |
| Multiple MCP server processes | ps aux | grep meshtastic_mcp shows >1 |
Kill all but the one your MCP host spawned. Zombies hold ports and break tests. |
| Mesh formation fails, one side sees peer but other doesn't | /diagnose (or list_nodes on both sides) |
Asymmetric NodeInfo. test_direct_with_ack has a heal path; /repro it a few times. If persistent, both devices' clocks may be out of sync with their NodeInfo cooldown. |
| "role not present on hub" in skip reasons | list_devices |
Expected if a device is unplugged. Reconnect before re-running the tier. |
Entire tests/recovery/ tier skipped |
command -v uhubctl |
Expected if uhubctl isn't on PATH. Install via brew install uhubctl (macOS) or apt install uhubctl (Debian/Ubuntu). Also skips if no hub advertises PPPS. |
Entire tests/ui/ tier skipped ("firmware not baked with USERPREFS_UI_TEST_LOG") |
reportlog.jsonl for the skip reason | Re-run with --force-bake so the UI-log macro gets compiled into the fresh firmware. First run after the Round-3 landing always re-bakes. |
tests/ui/ runs but captures are all 1×1 black PNGs |
MESHTASTIC_UI_CAMERA_DEVICE_ESP32S3 |
Env var not set → NullBackend. Point a USB webcam at the heltec-v3 OLED and set the device index; .venv/bin/python -c "import cv2; [print(i, cv2.VideoCapture(i).read()[0]) for i in range(5)]" discovers it. |
| Tests fail only on first attempt then pass on rerun | - | State leak from a prior session. Run with --force-bake to reset to a known state. |
Never do these without asking
factory_reset- wipes node identity; regenerates PKI keypair. Mesh peers will reject old DMs until re-exchange. Legitimate only when the operator explicitly wants it.erase_and_flash- full chip erase; destroys all on-device state.esptool_erase_flash/esptool_rawwrite/erase - bypasses pio's safety chain.set_configonlora.region- changes regulatory domain; requires physical-location context the operator has and the agent doesn't.reboot/shutdownmid-test - breaks fixture invariants.push -f,rebase -i,reset --hard, or any history-rewriting git operation.- Clicking computer-use tools on web links in Mail/Messages/PDFs - open URLs via the claude-in-chrome MCP so the extension's link-safety checks apply.