A stall spanning the 32-bit millis() rollover zero-extends to a value smaller
than the pre-wrap start, so the now_ms >= start_ms guard blocked the bail until
the counter lapped the ~49.7-day start value. Use 32-bit unsigned subtraction,
which measures true elapsed time across the wrap. Regression: a stall starting
near the counter max and wrapping bails in bounded ticks.
Three findings on the transient-stall bail, all valid:
- Reload invalidation: a bail recorded BYPASS, but NomadNetCacheFlow::service()
accepts only MISS as a successful invalidation, so a bail during an admitted
reload reported 'Page cache invalidation failed' instead of falling through
to a live fetch. The bail now records MISS.
- Open-resource leak: the bail cleared read_open_/write_open_ (and abandoned an
open directory enumeration) without calling endRead()/abortWrite()/endList(),
leaking SD handles. The bail now releases each in-flight resource via the
seam's own teardown (bounded best-effort; a still-transient close is
accepted rather than re-pinning the op).
- Tick-vs-time: the not-ready UNAVAILABLE path returns immediately (no bus
wait), so a pure 500-tick budget could expire during a legitimate SD mount
window and disable caching for the whole session. The bail is now gated on
BOTH the tick floor AND a 10s wall-time window (service() takes a monotonic
ms clock; production passes millis(), 0 is a safe default for tests).
Regression tests: reload-invalidation bail -> NEED_LIVE (flow), flat-clock does
not bail, healable transient keeps authority, list-open (RECOVERY_END) stall
bails and releases the handle (cache).
The SD page cache (16af3b5) made the SD card a soft dependency of
NomadNet page loads, but its step machine retries a transient storage
result (BUSY = SPI-mutex timeout, UNAVAILABLE = card not mounted)
forever with no budget. On a persistently unhealthy seam the boot-time
recovery pins operation_ != NONE, the flow stays LOOKUP, and the UI
freezes at 'Checking SD page cache...' (the CACHE state has no
deadline, unlike every other NomadNet state).
NomadNetCache::service() now compares each call's entry state to the
previous call's. Any advance (op, offset, scan/cleanup index, scan
count, open-flags) resets a stall counter, so slow-but-progressing
steps (chunked 64 KB transfers, up to 96-record directory scans) never
false-trip; a no-progress tick is a transient stall. Past 500
consecutive no-progress ticks (far beyond any real SPI contention or
SD mount window) the cache bails: mark the namespace non-authoritative
for the session (lookups/commits bypass) and clear the op, so the flow
falls through to a live fetch -- the pre-cache page-load behavior.
Regression tests: cache-level (persistent UNAVAILABLE recovery bails;
lookup bypasses after bail) and flow-level (a permanently BUSY
beginList no longer parks the flow in LOOKUP; it reaches NEED_LIVE in
bounded ticks). Both fail on the pre-fix code and pass with it.
Verified: tdeck firmware build SUCCESS; tests/native 121 passed (3
pre-existing env failures, fail identically on origin/main baseline);
tests/build_scripts 184 passed.
Long-press to delete a conversation, then release: the row fired
on_conversation_long_pressed (confirm dialog) AND, on release,
on_conversation_clicked, navigating into the conversation and hiding
the dialog. Root cause: LVGL 8.4 lv_indev.c indev_proc_release()
(lines 973-980) sends LV_EVENT_CLICKED on every pointer release without
scrolling, including long-press release — only LV_EVENT_SHORT_CLICKED
is gated on long_pr_sent == 0. The row had both CLICKED and LONG_PRESSED
bound.
Selection is now bound to LV_EVENT_SHORT_CLICKED. Trackball selection
is unaffected: the keypad path sends SHORT_CLICKED + CLICKED on a
plain enter release and suppresses both when long_pr_sent is set
(lv_indev.c:526-531, 678-692). No programmatic CLICKED sends exist
in the app. This row is the only object with the CLICKED +
LONG_PRESSED double-bind (bubble/textarea long-press sites are
single-bound, verified).
Contract pins: selection must bind SHORT_CLICKED, no CLICKED binding
on the row, long-press delete binding unchanged. 184/184 contracts,
tdeck + tdeck-release green.
Greptile round on ba5af76 (4/5) correctly rejected the first attempt:
the submitted-text marker was assigned in ChatScreen::on_send_clicked
AFTER the mailbox publish returned, so the main loop could take() +
admit the send and enter clear_composer() while the marker was still
empty — neither clearing the submitted text nor associating the commit
with its submission.
The marker is now recorded by the send callback itself
(UIManager::on_send_message_from_chat) immediately after the mailbox
accept, in the same LVGL lock section as the publish. The click handler
runs on the LVGL task with the LVGL mutex held (LVGLInit.cpp:160-179
wraps the whole lv_task_handler in the recursive mutex), so the marker
is visible to the main loop only after the mailbox entry is — the
take() + admit + clear sequence can never observe an empty marker for
an accepted send. clear_composer() additionally no-ops on an empty
marker, which is the retained-text path for rejected/retry sends.
The contract test is tightened to assert the marker is NOT assigned in
the click handler and IS assigned in the callback, so the race cannot
silently regress.
Verification: 181/181 contracts, tdeck + tdeck-release green.
Greptile P1 remediation on the exact head (round: 1a34c55):
1. Send Completion Erases Draft (UIManager.cpp:1780). The async send
deferral (1c68860) leaves the composer un-cleared between the send
click and the main-loop's ADDED commit, so input typed into the
composer while persistence/admission is in flight was wiped by the
unconditional clear_composer(). ChatScreen now captures the exact
submitted text when the send is accepted into the mailbox, and
clear_composer() only clears when the composer still holds that
text. A rejected send still retains input (unchanged), and a fresh
draft can no longer be erased by a late completion.
2. Same-Peer History Stays Stale (ChatScreen.cpp:177). The same-peer
early-return (ce92e80) skipped the store re-read, so a message for
this peer that persisted while the chat was hidden (
on_message_received only appends to the visible chat) never surfaced
on re-open. The early-return now compares the store's in-memory
conversation count (get_messages_for_conversation — pure slot
lookup, no LittleFS, safe under the LVGL lock) against the count at
the last prepare commit and falls through to the peer-change path on
a mismatch, which resets the list and re-arms prepare so the main
loop re-gathers off-lock and rebuilds with the new message.
Verification: 181/181 build-script contracts (5 new pins), tdeck +
tdeck-release green. Compose path audited and unaffected: the single
send slot makes a second send a no-op until the first commits, and its
clear rides on the route replacement (render_route).
[SENDT] send-pipeline timing and the microReticulum [PG] path-store
call-site counters now compile to no-ops unless explicitly enabled:
- DPYXIS_SEND_DIAG / -DRNS_PATHGET_DIAG added to env:tdeck base flags
- both removed by env:tdeck-release build_unflags (no-op in release)
So production tdeck/tdeck-release builds are merge-clean, while the
instrumented build stays one env/flag away for the path-request
spammer hunt. Bumps microReticulum pin to 921b3aa (same endpoint
hot-path read gate, counters gated behind RNS_PATHGET_DIAG).
Contract suite: 176/176.
Bumps microReticulum pin to e2c9d4d (diag/path-get-caller on cd0338e):
Transport::inbound() and path_request() no longer perform a full
microStore get() (flash write + read) for every inbound packet / path
request on endpoint-only nodes. The read is gated on the exact
conditions where destination_entry is consumed, and the local-destination
path-request answer uses the in-memory _destinations table, so the
device stays discoverable. The build still carries the temporary [PG]
counters for the live before/after capture; counters are stripped in a
follow-up before anything merges.
TEMPORARY DIAGNOSTIC PIN. ef07187 = cd0338e + per-call-site
_new_path_table.get() counters ([PG] summary every 30s on serial).
Purpose: identify which Transport call site drives the ~1.5s full
FileStore get() on the offline propagation node (6b9f6601...).
Revert this commit (back to cd0338e) after the capture.
Bumps microLXMF d7e05fd -> 82d2e54 (fork branch
fix/sync-path-wait-backoff): LXMRouter::process_sync() PR_PATH_REQUESTED
now re-polls Transport::has_path() at most once per PATH_REQUEST_WAIT
instead of every main-loop iteration. In the CBA microStore-backed
microReticulum fork has_path() is a FileStore exists() = a LittleFS
read, so the old code issued a path-table read every ~2s for the whole
60s sync window while the propagation node was unreachable. Measured on
the T-Deck: one destination fetched 293x in 600s (~every 2s), all on
the same 2MB partition as message storage, with [DISP] flush stalls
interleaved. Worst-case sync cycle cost drops from ~30 reads to ~3-4.
The periodic message sync was already 4h default with a user setting
(Settings 'Prop Sync Interval (hrs)', sync_int=14400) — unchanged.
Adds tests/build_scripts/test_propagation_sync_poll_contract.py:
pins the re-poll gate ordering, the PATH_REQUEST_WAIT window, the
untouched one-shot check in request_messages_from_propagation_node,
and enforces platformio.ini <-> audit-tool pin agreement.
[SENDT] marks: queue_wait (mailbox wait), identity_recall,
router_lock, save_message (isolated), admission_done, ui_commit_done.
Instrumentation only; to be removed before merge. Loosens one contract
pin to the save call expression (invariant preserved: save runs in
persistOutgoingMessage, not service_pending_sends).
Opening a conversation crashed the same way sending did. load_conversation()
(LVGL task, under the LVGL lock held by replace_route) ran the full open
pipeline synchronously: identity recall (ustore), display-name read, the
message-index read, and the per-message metadata reads. On this device's
degraded LittleFS each op is 0.4-2s, so a cold open of a dozen-message
conversation held the LVGL mutex past the 5s deadlock guard and asserted at
LVGLLock.h:45. The send path already got the mailbox fix; the open path never
did.
Restructure with the same pattern:
- load_conversation() (LVGL task) now only navigates + resets the list and
shows the truncated hash in the header. Same-peer re-opens return early
with zero store I/O (rows are still built).
- prepare_conversation() (main loop, called from update()) does the store
I/O between a short guard lock and a short commit lock, then commits the
header name + initial bubbles + background-fill arming under a brief
LVGL_LOCK. A generation counter discards a stale in-flight prepare when
the conversation changes mid-I/O.
- refresh() re-arms the prepare instead of re-reading under the lock.
The 1Hz store 'not found in index' fetch is pre-existing (present on
2527c6d) and is being tracked separately as a flash-wear follow-up.
Build tdeck SUCCESS, 170/170 contract tests pass.
Every message send on the device was deterministically rebooting it:
send_message() ran the full pipeline (identity recall, message
construction, RouterLock-scoped router admission, and LittleFS
persistence) synchronously on LVGL's 8 KiB task while holding the LVGL
mutex. On this device's degraded filesystem a single save takes ~7s of
400ms-2s per-op gaps, tripping the 5s LVGL deadlock guard and asserting
at LVGLLock.h:45 (assert failed: LVGL mutex timeout (5s)). The receive
path already carries the fix pattern for exactly this failure class
(see on_message_received); the send path never got it.
Restructure the send path as a mailbox handoff, following the existing
CallStartMailbox / LocationShareCommandMailbox precedent:
- send_message() (LVGL task) now only validates and publishes
(destination, content, source) into a mutex-guarded single-slot
OutgoingSendMailbox. No router lock, no I/O, no message construction.
- update() services the mailbox in service_pending_sends() on the main
loop, before the big LVGL_LOCK() — the only place in the send path
that may take the router lock, block on admission, or wait on
LittleFS.
- On acceptance, a brief LVGL_LOCK in apply_outbound_result() commits
the UI (add_message / clear_composer / compose->chat navigation,
route-guarded). The admitted packed form is unpacked for display
with incoming/state flags restored.
- On rejection (storage error, router busy, queue full) the user's
input is retained for retry, matching the old behavior.
The 500-char UI cap bounds the mailbox payload.
Build tdeck SUCCESS, 170/170 contract tests pass.
Replace the flat Settings screen with a card hub + dedicated sub-views.
- Hub: 8 navigation cards (Status, Network, Identity, Radio, Delivery,
Appearance, Advanced, Transport) reusing the Network-screen card widget;
Transport stays last (danger invariant preserved).
- Tap a card -> dedicated sub-view holding that area's controls; no accordion.
- Network sub-view gains a LoRa interface toggle alongside TCP/Auto/BLE;
it two-way-mirrors the Radio page's canonical lora_enabled switch.
- Save model: simple controls apply immediately; a Save button appears only
on the form sub-views (Network, Radio, Identity).
- Identity sub-view adds a View Identity row routing to the existing lxma://
QR screen (the Status Share button already reached it).
- Focus group rebuilt per view (only-visible objects) so the auto-scroll-to-
bottom class of bug cannot recur; each sub-view entry scrolls to top.
Contracts updated for the new structure; 170/170 build-script tests pass.
tdeck build green: RAM 23.1% (75,848 B, unchanged), Flash 92.7%
(2,914,713 B, +2,640 B). Not flashed, not PR'd.
Settings held live status (GPS fix, storage/RAM/identity) that belongs
on the Status screen, plus a per-second tick() doing SPI flash stat
reads and label churn mid-scroll — the main cause of laggy scrolling.
- GPS section (sats/location/altitude/HDOP/time) -> StatusScreen
- System Info (firmware build, storage, RAM) -> StatusScreen, with
storage/RAM stat reads throttled to ~5s and stack-buffer snprintfs
instead of Arduino String concatenation
- Settings gains a Status link row (trackball-reachable) that opens
Route::STATUS; the per-second SettingsScreen tick/refresh is deleted
- Reordered sections by frequency of use: General (name/brightness/
timeout/kb-light), Notifications, Network (now includes the
TCP/Auto/BLE interface switches), Radio (LoRa + params), Delivery,
Advanced, DANGER: Transport Mode (still final)
- Identity/LXMF hashes shown in Settings were truncated duplicates of
the Status screen's full display; removed
- main.cpp publishes firmware build + GPS to the Status screen
- Contract test for the storage readout follows the code to StatusScreen
Tracks microLXMF PR #11 head d7e05fd: archived-branch cache sync,
bulk-delete cache invalidation, and the uniform capped-content
contract for load_message_metadata(). No Pyxis code change — the
chat/list paths already rely on the capped preview and the full-view
accessor.
Uses the microLXMF in-process message-metadata cache (bumped pin
6bea23c -> 58a6eb0, branch feat/conversation-preview-cache):
- Reopening a conversation (and background page fill / paging) is now
O(1) in-memory once warmed instead of re-reading each message file
from SPI LittleFS (~230ms/read, ~2.3s per open measured on the T-Deck).
- The cache table is PSRAM-allocated on ESP32; a static .bss placement
starved internal DRAM and the LVGL task's 8 KiB stack allocation
failed at boot ("Failed to create LVGL task", hang at startup logo).
- Long-press full-message view now defers to the main loop: the LVGL
event handler only records the hash; tick_pending_full_message() does
load_message_content() (uncapped, no msgpack unpack) off the LVGL
task, then builds the modal. Fixes the capped (600-char) text that
the in-memory rows carry.
- A peer change cancels an in-flight background fill from the previous
conversation (it would otherwise prepend the old conversation's
rows into the new one).
Validated on the T-Deck (T-Deck Plus, 8MB PSRAM): warm opens
reads=3 sync=0-2ms total=15-35ms (was ~2.3s); cold first open per
conversation still pays ~0.6-1.2s disk while warming the cache; no
panics/watchdogs over a ~6-minute interaction session.
Host gates: microLXMF conformance 6/6 (incl. 32-assertion
test_message_metadata_cache), native reference test against pinned
58a6eb0, tdeck build SUCCESS.
On-device [PERF] capture decomposed the remaining ~1s tap latency:
each uncached conversation costs one SPI-LittleFS metadata read
(~230ms). Repeat taps still paid 2 perpetual fallbacks (empty-
content tails never got cached), and the first tap after every
boot paid all 9 (the in-memory repop was never committed to the
index).
- bump microLXMF pin c8d3156 -> 6bea23c (feat/conversation-
preview-cache): preview_valid index flag so 'cached empty'
differs from 'unpopulated'; bounded preview copy (the old
strncpy read past the non-terminated content Bytes); public
commit_index().
- refresh() re-pops empty-content tails as a valid cached
preview and arms a one-shot deferred index commit;
UIManager::update() drains it out-of-lock (flush_pending_
index_commit, between drops and mark-read) so the warmed
previews persist and the next cold boot reads them from the
index.
- fix the [PERF] skip-log total= wrap (printed p_t_diff - p_t0,
a uint32 underflow; total was already p_t_diff).
[PERF] instrumentation stays in this commit (temporary,
marked); it is removed before merge once the fix is validated
on-device.
refresh() previously opened + parsed each conversation's newest message
file (load_message_metadata) on every list refresh, which dominated
list-load time on SPI LittleFS. It now reads the per-conversation
last-message preview + timestamp from the store's in-memory index
(O(1), zero I/O) and falls back to load_message_metadata only when the
index has no cached preview — the first refresh after a firmware
upgrade or after a corrupt-tail drop — then writes the preview back
through so the fallback happens at most once per conversation per
firmware generation. The write-through is skipped when drops were
queued (the drained deletes move the tail, so a preview written for the
old tail would be stale for one refresh).
Host benchmark (x86 + POSIX fs, 24 msgs/conv): the per-conversation
store-load work drops from ~0.149ms (9 convs) / ~0.342ms (20 convs) to
~0.002ms / ~0.005ms, and the cold-boot path (store reconstructed from
disk) matches the warm path because the preview now persists in the
index.
Bumps the microLXMF pin to c8d3156 (feat/conversation-preview-cache)
in platformio.ini, the release audit, and the native reference test.
Adds tests/microlxmf/bench_conversation_list_load.cpp (baseline vs
index warm/cold) and the bench target in tests/microlxmf/CMakeLists.txt.
Greptile note: the derive-before-apply position check could pass even if
the color call regressed to the wrong branch. Tighten the contract test
to assert exactly one label-color call, located after the marker loop's
hidden-marker continue (i.e. in the main render branch, not the hide
path), and that the once-per-frame derivation precedes it.
Marker (pin) labels previously inherited the app's default text color,
which reads white and disappears on light basemaps. Set each label's
text color in applyFrame() against the active style: black by default
(light basemaps osm-bright/positron/toner) and white only on the one
dark basemap (dark-matter). Re-evaluated every frame so a style switch
re-colors visible labels on the next applied frame.
Add a contract test pinning the dark-basemap detection and the
black-by-default / white-on-dark ternary.
A message whose source identity is KNOWN but whose signature fails to
validate is spoofed or malicious and must not be rendered. The
opportunistic (on_packet) and direct (on_resource_concluded) router
paths already reject these, but the propagated (store-and-forward) path
in process_propagated_lxmf queues them without a signature check, so
UIManager::on_message_received is the single choke point that covers
all three inbound routes.
Drop the message at the top of on_message_received — before the key
request, location ingest, persistence, chat render, and notification
beep — when !signature_validated() && reason == SIGNATURE_INVALID.
SOURCE_UNKNOWN (first contact) is untouched: those still render and
trigger the bounded key request from PR #92. Validated messages are
unaffected.
Add a source-level contract test locking in the drop gate's ordering
relative to every side effect and its enum specificity.
Greptile P2 on PR #92: with 64 tracked sources, evicting the oldest
entry discarded its open cooldown, so a flood of distinct bogus
identities reset other sources' windows and forced unbounded path
requests (each answered by every peer holding the announce).
- Per-identity cooldown 5 min -> 30 min.
- The table no longer evicts an open window: while all 64 slots hold
unexpired windows, never-before-seen identities are deferred until a
slot frees (at most one 30-minute window) instead of dropping
someone else's cooldown. Open windows are only ever pruned after
they expire.
- Aggregate bound: at most kMaxTrackedSources automatic path requests
per rolling 30-minute window, capping the worst-case network cost of
a rotating-identity flood.
Host tests extended: saturated-table deferral, 512-identity flood bound
(one window and across consecutive windows), lazy expiry/prune,
boundary checks at the 30-minute cooldown.
Move the rate-limit/cooldown/cap decision out of the UIManager.cpp
anonymous namespace into a pure, header-only policy
(UI/LXMF/UnknownSourceKeyRequest.h) so it is host-testable without the
ESP/microReticulum stack. UIManager keeps only the side effect
(Transport::request_path).
Adds tests/native/test_unknown_source_key_request.{cpp,py} (24 checks,
ASan+UBSan): new-source request, 5-min cooldown boundary, re-record
resets the window, per-source independence, 64-entry cap with oldest
eviction, and steady-state cost.
Greptile P2 (round 1): a post-rename parent-fsync failure during activation now reports that the record is committed and visible with only durability uncertain, instead of the generic 'activation failed' message; and the subprocess lock-holder helper repr-escapes the pyxis-map path so a path containing a quote still yields valid child Python.
Raise MapTileStore::PATH_CAPACITY from 64 to 80 so the mount buffer
(PATH_CAPACITY + 4) holds the 80-character mounted tile path produced by
a maximum-length (31-char) pack ID with full-width tile coordinates.
Pre-fix, any tile with a two-digit x or y returned INVALID_ARGUMENT from
the mount-prefix snprintf, which MapTilePack maps to IO_ERROR and the UI
shows as 'Tile I/O error' even though the file exists and is intact.
The host regression test added in the prior commit now passes under both
strict C++11 and ASan/UBSan; reverting this change fails it.
A 31-character pack ID makes mounted tile paths
(/sd/pyxis-map/packs/<id>/tiles/<z>/<x>/<y>.png) exceed the 68-byte mount
buffer once either x or y is two digits, so beginRead returns
INVALID_ARGUMENT and the UI shows 'Tile I/O error' at z4+ even though the
file exists and is intact.
Compiles the unmodified MapTilePack/MapTileStoreSD/SDAccess/codec/manifest
sources against small Arduino/FreeRTOS host shims. The core section (no
card needed) drives the real store with a model of the mount-prefix
arithmetic and asserts the 31-char-ID boundary loads at z4/z5; an optional
end-to-end section runs the real MapTilePack over a bound card root when
/sd is writable.
Fails on the current 64-byte PATH_CAPACITY: red by design.