Commit Graph
318 Commits
Author SHA1 Message Date
Torlando fabd3713df fix(nomadnet): make the transient-stall time window wrap-safe across millis()
A stall spanning the 32-bit millis() rollover zero-extends to a value smaller
than the pre-wrap start, so the now_ms >= start_ms guard blocked the bail until
the counter lapped the ~49.7-day start value. Use 32-bit unsigned subtraction,
which measures true elapsed time across the wrap. Regression: a stall starting
near the counter max and wrapping bails in bounded ticks.
2026-09-15 14:03:25 +00:00
Torlando 32a09938d8 fix(nomadnet): address Greploop round 1 on the cache transient-stall guard
Three findings on the transient-stall bail, all valid:

- Reload invalidation: a bail recorded BYPASS, but NomadNetCacheFlow::service()
  accepts only MISS as a successful invalidation, so a bail during an admitted
  reload reported 'Page cache invalidation failed' instead of falling through
  to a live fetch. The bail now records MISS.

- Open-resource leak: the bail cleared read_open_/write_open_ (and abandoned an
  open directory enumeration) without calling endRead()/abortWrite()/endList(),
  leaking SD handles. The bail now releases each in-flight resource via the
  seam's own teardown (bounded best-effort; a still-transient close is
  accepted rather than re-pinning the op).

- Tick-vs-time: the not-ready UNAVAILABLE path returns immediately (no bus
  wait), so a pure 500-tick budget could expire during a legitimate SD mount
  window and disable caching for the whole session. The bail is now gated on
  BOTH the tick floor AND a 10s wall-time window (service() takes a monotonic
  ms clock; production passes millis(), 0 is a safe default for tests).

Regression tests: reload-invalidation bail -> NEED_LIVE (flow), flat-clock does
not bail, healable transient keeps authority, list-open (RECOVERY_END) stall
bails and releases the handle (cache).
2026-09-15 06:27:14 +00:00
Torlando feb7f0574b fix(nomadnet): bound cache transient-stall so a failing SD seam cannot freeze the UI
The SD page cache (16af3b5) made the SD card a soft dependency of
NomadNet page loads, but its step machine retries a transient storage
result (BUSY = SPI-mutex timeout, UNAVAILABLE = card not mounted)
forever with no budget. On a persistently unhealthy seam the boot-time
recovery pins operation_ != NONE, the flow stays LOOKUP, and the UI
freezes at 'Checking SD page cache...' (the CACHE state has no
deadline, unlike every other NomadNet state).

NomadNetCache::service() now compares each call's entry state to the
previous call's. Any advance (op, offset, scan/cleanup index, scan
count, open-flags) resets a stall counter, so slow-but-progressing
steps (chunked 64 KB transfers, up to 96-record directory scans) never
false-trip; a no-progress tick is a transient stall. Past 500
consecutive no-progress ticks (far beyond any real SPI contention or
SD mount window) the cache bails: mark the namespace non-authoritative
for the session (lookups/commits bypass) and clear the op, so the flow
falls through to a live fetch -- the pre-cache page-load behavior.

Regression tests: cache-level (persistent UNAVAILABLE recovery bails;
lookup bypasses after bail) and flow-level (a permanently BUSY
beginList no longer parks the flow in LOOKUP; it reaches NEED_LIVE in
bounded ticks). Both fail on the pre-fix code and pass with it.

Verified: tdeck firmware build SUCCESS; tests/native 121 passed (3
pre-existing env failures, fail identically on origin/main baseline);
tests/build_scripts 184 passed.
2026-09-07 18:52:03 +00:00
Torlando cab2ad8a83 fix(ui): conversation-row selection on SHORT_CLICKED, not CLICKED
Long-press to delete a conversation, then release: the row fired
on_conversation_long_pressed (confirm dialog) AND, on release,
on_conversation_clicked, navigating into the conversation and hiding
the dialog. Root cause: LVGL 8.4 lv_indev.c indev_proc_release()
(lines 973-980) sends LV_EVENT_CLICKED on every pointer release without
scrolling, including long-press release — only LV_EVENT_SHORT_CLICKED
is gated on long_pr_sent == 0. The row had both CLICKED and LONG_PRESSED
bound.

Selection is now bound to LV_EVENT_SHORT_CLICKED. Trackball selection
is unaffected: the keypad path sends SHORT_CLICKED + CLICKED on a
plain enter release and suppresses both when long_pr_sent is set
(lv_indev.c:526-531, 678-692). No programmatic CLICKED sends exist
in the app. This row is the only object with the CLICKED +
LONG_PRESSED double-bind (bubble/textarea long-press sites are
single-bound, verified).

Contract pins: selection must bind SHORT_CLICKED, no CLICKED binding
on the row, long-press delete binding unchanged. 184/184 contracts,
tdeck + tdeck-release green.
2026-09-07 01:59:11 +00:00
Torlando a1c2ec8569 fix(lxmf): set send marker in the LVGL lock section (close completion race)
Greptile round on ba5af76 (4/5) correctly rejected the first attempt:
the submitted-text marker was assigned in ChatScreen::on_send_clicked
AFTER the mailbox publish returned, so the main loop could take() +
admit the send and enter clear_composer() while the marker was still
empty — neither clearing the submitted text nor associating the commit
with its submission.

The marker is now recorded by the send callback itself
(UIManager::on_send_message_from_chat) immediately after the mailbox
accept, in the same LVGL lock section as the publish. The click handler
runs on the LVGL task with the LVGL mutex held (LVGLInit.cpp:160-179
wraps the whole lv_task_handler in the recursive mutex), so the marker
is visible to the main loop only after the mailbox entry is — the
take() + admit + clear sequence can never observe an empty marker for
an accepted send. clear_composer() additionally no-ops on an empty
marker, which is the retained-text path for rejected/retry sends.

The contract test is tightened to assert the marker is NOT assigned in
the click handler and IS assigned in the callback, so the race cannot
silently regress.

Verification: 181/181 contracts, tdeck + tdeck-release green.
2026-09-07 01:16:34 +00:00
Torlando ba5af761fb fix(lxmf): preserve draft on async send; re-gather hidden-peer history (Greptile P1s)
Greptile P1 remediation on the exact head (round: 1a34c55):

1. Send Completion Erases Draft (UIManager.cpp:1780). The async send
   deferral (1c68860) leaves the composer un-cleared between the send
   click and the main-loop's ADDED commit, so input typed into the
   composer while persistence/admission is in flight was wiped by the
   unconditional clear_composer(). ChatScreen now captures the exact
   submitted text when the send is accepted into the mailbox, and
   clear_composer() only clears when the composer still holds that
   text. A rejected send still retains input (unchanged), and a fresh
   draft can no longer be erased by a late completion.

2. Same-Peer History Stays Stale (ChatScreen.cpp:177). The same-peer
   early-return (ce92e80) skipped the store re-read, so a message for
   this peer that persisted while the chat was hidden (
   on_message_received only appends to the visible chat) never surfaced
   on re-open. The early-return now compares the store's in-memory
   conversation count (get_messages_for_conversation — pure slot
   lookup, no LittleFS, safe under the LVGL lock) against the count at
   the last prepare commit and falls through to the peer-change path on
   a mismatch, which resets the list and re-arms prepare so the main
   loop re-gathers off-lock and rebuilds with the new message.

Verification: 181/181 build-script contracts (5 new pins), tdeck +
tdeck-release green. Compose path audited and unaffected: the single
send slot makes a second send a no-op until the first commits, and its
clear rides on the route replacement (render_route).
2026-09-07 00:46:07 +00:00
Torlando 3c6275d9e6 chore(diag): gate [SENDT] and [PG] instrumentation behind build flags
[SENDT] send-pipeline timing and the microReticulum [PG] path-store
call-site counters now compile to no-ops unless explicitly enabled:
  - DPYXIS_SEND_DIAG / -DRNS_PATHGET_DIAG added to env:tdeck base flags
  - both removed by env:tdeck-release build_unflags (no-op in release)

So production tdeck/tdeck-release builds are merge-clean, while the
instrumented build stays one env/flag away for the path-request
spammer hunt. Bumps microReticulum pin to 921b3aa (same endpoint
hot-path read gate, counters gated behind RNS_PATHGET_DIAG).

Contract suite: 176/176.
2026-09-06 22:42:16 +00:00
Torlando f0e705e529 diag(send): temporary SENDT pipeline timing for save-latency measurement
[SENDT] marks: queue_wait (mailbox wait), identity_recall,
router_lock, save_message (isolated), admission_done, ui_commit_done.
Instrumentation only; to be removed before merge. Loosens one contract
pin to the save call expression (invariant preserved: save runs in
persistOutgoingMessage, not service_pending_sends).
2026-09-06 01:20:09 +00:00
Torlando 7bd6e6712e fix(lxmf): move background-fill metadata reads off the LVGL lock
load_more_messages() (background fill, main loop) ran each batch's
metadata reads while holding the LVGL lock. A cold batch is 2+ LittleFS
reads; on the degraded device a fill batch measured ~4s in earlier
captures, within a hair of the 5s deadlock guard. The fill starts
immediately after any conversation open (first page > INITIAL_RENDER),
so an open triggered a second under-lock stall right after the first
was fixed.

Same restructure: guard lock copies the batch's hashes, metadata I/O
runs off-lock, the bubble prepend commits under a brief lock. A
_fill_generation counter (bumped on every list rebuild: open, prepare
commit, refresh) discards a batch whose conversation changed mid-I/O.

Display order is preserved (newest-first read order + push_front gives
oldest-to-newest, matching the old loop).
2026-09-05 15:00:32 +00:00
Torlando ce92e8073e fix(lxmf): move conversation-open store I/O off the LVGL task
Opening a conversation crashed the same way sending did. load_conversation()
(LVGL task, under the LVGL lock held by replace_route) ran the full open
pipeline synchronously: identity recall (ustore), display-name read, the
message-index read, and the per-message metadata reads. On this device's
degraded LittleFS each op is 0.4-2s, so a cold open of a dozen-message
conversation held the LVGL mutex past the 5s deadlock guard and asserted at
LVGLLock.h:45. The send path already got the mailbox fix; the open path never
did.

Restructure with the same pattern:
- load_conversation() (LVGL task) now only navigates + resets the list and
  shows the truncated hash in the header. Same-peer re-opens return early
  with zero store I/O (rows are still built).
- prepare_conversation() (main loop, called from update()) does the store
  I/O between a short guard lock and a short commit lock, then commits the
  header name + initial bubbles + background-fill arming under a brief
  LVGL_LOCK. A generation counter discards a stale in-flight prepare when
  the conversation changes mid-I/O.
- refresh() re-arms the prepare instead of re-reading under the lock.

The 1Hz store 'not found in index' fetch is pre-existing (present on
2527c6d) and is being tracked separately as a flash-wear follow-up.

Build tdeck SUCCESS, 170/170 contract tests pass.
2026-09-05 14:39:32 +00:00
Torlando 1c688608b5 fix(lxmf): move outgoing-send persistence off the LVGL task
Every message send on the device was deterministically rebooting it:
send_message() ran the full pipeline (identity recall, message
construction, RouterLock-scoped router admission, and LittleFS
persistence) synchronously on LVGL's 8 KiB task while holding the LVGL
mutex. On this device's degraded filesystem a single save takes ~7s of
400ms-2s per-op gaps, tripping the 5s LVGL deadlock guard and asserting
at LVGLLock.h:45 (assert failed: LVGL mutex timeout (5s)). The receive
path already carries the fix pattern for exactly this failure class
(see on_message_received); the send path never got it.

Restructure the send path as a mailbox handoff, following the existing
CallStartMailbox / LocationShareCommandMailbox precedent:

- send_message() (LVGL task) now only validates and publishes
  (destination, content, source) into a mutex-guarded single-slot
  OutgoingSendMailbox. No router lock, no I/O, no message construction.
- update() services the mailbox in service_pending_sends() on the main
  loop, before the big LVGL_LOCK() — the only place in the send path
  that may take the router lock, block on admission, or wait on
  LittleFS.
- On acceptance, a brief LVGL_LOCK in apply_outbound_result() commits
  the UI (add_message / clear_composer / compose->chat navigation,
  route-guarded). The admitted packed form is unpacked for display
  with incoming/state flags restored.
- On rejection (storage error, router busy, queue full) the user's
  input is retained for retry, matching the old behavior.

The 500-char UI cap bounds the mailbox payload.

Build tdeck SUCCESS, 170/170 contract tests pass.
2026-09-05 05:23:02 +00:00
Torlando 2527c6d96b fix(settings): Greploop remediation — three Greptile P1s
Address the three P1 findings from the Greptile review of PR #96:

1. Network LoRa toggle reverted on save. on_interface_switch_changed
   saved without propagating the Network switch state, so the save path
   read the stale Radio-page twin and wrote it back over the mirror.
   The handler now mirrors the flipped switch to its twin (and toggles
   the LoRa params visibility) before saving, mirroring
   on_lora_enabled_changed in the other direction.

2. Immediate controls committed unsaved drafts from other sub-views.
   Every immediate handler ran the full-screen update_settings_from_ui,
   so typing a TCP host and then adjusting brightness silently persisted
   the host. update_settings_from_ui is now view-scoped: save_settings
   reads only the sub-view that just changed (LVGL events only fire for
   visible widgets, so _view is always the touched view). The lora
   handle merges as Radio-page switch wins, Network mirror when the
   Radio page is unbuilt.

3. Hidden controls retained focus. hide() removed only the header
   buttons and hub cards, and the focus_group_for cleanup list was
   missing the three Radio dropdowns — so hidden sub-view controls
   stayed in the default group (and Radio dropdowns lingered after
   leaving the Radio view). Extracted the full removal list into
   remove_all_from_focus_group (now including _dropdown_lora_bandwidth/
   _sf/_cr) and use it from both focus_group_for and hide.

Verification: tdeck build SUCCESS (RAM 23.1%, Flash 92.7%),
tests/build_scripts 170/170. All save_settings() call sites confirmed
event-driven so _view always matches the touched control.

Signed-off-by: Torlando <tyler@torlando.tech>
2026-09-04 23:42:15 +00:00
Torlando 4fffed3c1b fix(settings): Network view polish — slider overlap, row order, button clarity
- Move the LoRa TX Power slider from x=75 to x=86: the 'TX Power:' label
  in Montserrat 14 is ~78px wide, so the slider track started on top of
  the colon.
- Network sub-view: the Reticulum interface toggles (TCP/AUTO/LoRa/BLE)
  are now row 1, relabeled 'Reticulum Interfaces:', with WiFi SSID,
  WiFi password, TCP host and port following in their original order.
- Rename the ambiguous 'Reconnect' button to 'WiFi Reconnect' — it
  re-applies the WiFi SSID/password fields to the running stack and has
  nothing to do with the TCP host. Sized 110px to fit the label.

Signed-off-by: Torlando <tyler@torlando.tech>
2026-09-04 22:06:24 +00:00
Torlando e6ff73ec30 fix(settings): move LoRa TX power slider clear of the label
'TX Power:' in Montserrat 14 is ~78px wide but the slider track started
at x=75, tucking the colon under the track. Start the track at x=86
(8px clear; still 51px from the right-aligned value label).

Signed-off-by: Torlando <tyler@torlando.tech>
2026-09-04 21:57:42 +00:00
Torlando 72130e07c6 fix(settings): build sub-views lazily to fix LVGL task create at boot
Pre-building all seven settings sub-views at construction created
~100 LVGL objects that (being <256B) allocated into internal DRAM via
the hybrid allocator. On the T-Deck that fragmented internal heap to
an 8180B largest block at LVGL task-creation time, and the 8KB task
(needs ~8.5KB with TCB) failed: xTaskCreatePinnedToCore returned
errQUEUE_FULL, the device parked on the boot splash, and
confirm_running_firmware never ran.

Root cause proven with an on-device heap probe (largest internal
block 8180B, 8192-stack task FAIL, 4096-stack OK), then fixed by
building each sub-view on first navigation. Boot-time internal DRAM
rises to a 47KB largest block and the task starts cleanly.

Signed-off-by: Torlando <tyler@torlando.tech>
2026-09-04 21:15:04 +00:00
Torlando aef7b7f39d fix(settings): create sec_label with lv_label_create, not lv_obj_create
Boot-loop regression from the hub-and-spoke rewrite: lv_label_set_text()
on a plain lv_obj object wrote label internals into a non-label and
freed a bogus pointer (heap_caps_free assert 'outside heap areas') during
UIManager::init -> SettingsScreen::create_advanced_view, crashing every
boot at 'Initializing UIManager'. Verified by decoding the on-device
backtrace (ELF 62cf1d0e...) against the exact flashed firmware.
2026-09-04 18:52:43 +00:00
Torlando dfb2a73876 feat(settings): hub-and-spoke card layout (Columba-inspired)
Replace the flat Settings screen with a card hub + dedicated sub-views.
- Hub: 8 navigation cards (Status, Network, Identity, Radio, Delivery,
  Appearance, Advanced, Transport) reusing the Network-screen card widget;
  Transport stays last (danger invariant preserved).
- Tap a card -> dedicated sub-view holding that area's controls; no accordion.
- Network sub-view gains a LoRa interface toggle alongside TCP/Auto/BLE;
  it two-way-mirrors the Radio page's canonical lora_enabled switch.
- Save model: simple controls apply immediately; a Save button appears only
  on the form sub-views (Network, Radio, Identity).
- Identity sub-view adds a View Identity row routing to the existing lxma://
  QR screen (the Status Share button already reached it).
- Focus group rebuilt per view (only-visible objects) so the auto-scroll-to-
  bottom class of bug cannot recur; each sub-view entry scrolls to top.

Contracts updated for the new structure; 170/170 build-script tests pass.
tdeck build green: RAM 23.1% (75,848 B, unchanged), Flash 92.7%
(2,914,713 B, +2,640 B). Not flashed, not PR'd.
2026-09-04 18:26:39 +00:00
Torlando 9ccadba207 refactor(settings): strip SCROLLDIAG instrumentation after physical confirmation
The explicit scroll-to-top in show() was confirmed working on the device
(open-at-top on every entry, including repeat entries after scrolling to
the bottom). Remove the temporary Serial logging and the 30ms deviation
watcher.
2026-09-04 14:41:37 +00:00
Torlando 5a32b9c0a3 diag(settings): instrument scroll position at entry and watch for post-entry movement
Temporary [SCROLLDIAG] build: show() logs _content scroll_y before and
after the explicit reset, and a 30ms LVGL timer logs the first deviation
from 0 after entry. Captures the actual mechanism behind the bottom-open
instead of guessing. Remove before PR.
2026-09-04 14:16:15 +00:00
Torlando d33ae5bb6f fix(settings): recreate entry scroll-reset timer after one-shot delete
The one-shot LVGL timer's callback deleted the timer but left
_entry_scroll_reset_timer dangling. On the 2nd+ screen entry show()
took the 'else' path and called lv_timer_reset() on freed memory, so
the view was never reset to the top — it stayed wherever the group
re-focus had auto-scrolled it (the bottom). Null the pointer before
lv_timer_del so each entry recreates a fresh timer.
2026-09-04 13:40:45 +00:00
Torlando f62f465958 fix(settings): don't auto-scroll to bottom on screen entry
LVGL re-focuses the default input group when the screen is shown; the
focused member is the last focusable object (the transport switch), so
the view landed at the bottom of the list. A one-shot 50ms LVGL timer
after show() scrolls the content back to the top without touching the
focus state, and leaves subsequent user scrolling alone.
2026-09-04 05:52:56 +00:00
Torlando 28ea218a18 refactor(settings): move status readouts to Status screen, reorganize by use
Settings held live status (GPS fix, storage/RAM/identity) that belongs
on the Status screen, plus a per-second tick() doing SPI flash stat
reads and label churn mid-scroll — the main cause of laggy scrolling.

- GPS section (sats/location/altitude/HDOP/time) -> StatusScreen
- System Info (firmware build, storage, RAM) -> StatusScreen, with
  storage/RAM stat reads throttled to ~5s and stack-buffer snprintfs
  instead of Arduino String concatenation
- Settings gains a Status link row (trackball-reachable) that opens
  Route::STATUS; the per-second SettingsScreen tick/refresh is deleted
- Reordered sections by frequency of use: General (name/brightness/
  timeout/kb-light), Notifications, Network (now includes the
  TCP/Auto/BLE interface switches), Radio (LoRa + params), Delivery,
  Advanced, DANGER: Transport Mode (still final)
- Identity/LXMF hashes shown in Settings were truncated duplicates of
  the Status screen's full display; removed
- main.cpp publishes firmware build + GPS to the Status screen
- Contract test for the storage readout follows the code to StatusScreen
2026-09-04 04:56:07 +00:00
Torlando 8c0d38ee2e Merge pull request #95 from torlando-tech/perf/conversation-list-load
Build and Deploy Firmware / build-and-deploy (push) Canceled after 0s
Test / Pyxis pytest suite (build_scripts + native) (push) Canceled after 0s
Test / microReticulum native unit tests (PlatformIO native17) (push) Canceled after 0s
perf(messages): fast conversation-list load + real unread badges
2026-09-03 23:00:32 -04:00
Torlando 362db2c805 fix(ui): fail closed when the queue mutex is unavailable
Greptile follow-up on 67c1083: if xSemaphoreCreateMutex() failed,
QueueGuard silently locked nothing and the deferred queues still ran
unsynchronized. The constructor now logs the allocation failure and
every guarded site refuses mutation when _queue_mutex is null:
requesters stop enqueuing, flushers stop draining (nothing queued can
exist), and the one-shot preview commit is skipped. The screen keeps
rendering and chatting; only the deferred list mutations degrade.
2026-09-04 02:49:26 +00:00
Torlando 67c1083f22 fix(ui): mutex the conversation-list deferred queues
Greptile P1 on the list-perf commit: the deferred-queue pair
(requester task -> main-loop flush) was unsynchronized. Producers
push from the LVGL task (click handlers, refresh() during navigation)
while flush_*() swaps the queues on the main loop without the LVGL
lock, so concurrent push/swap is a data race.

A single FreeRTOS mutex (QueueGuard, fail-closed) now guards all
four shared queues: _pending_mark_reads, _pending_drops,
_pending_name_writes, and the _index_commit_pending flag. Critical
sections are bounded vector operations only — no store I/O, no LVGL
lock, so acquisition cannot deadlock and the LVGL task is never
stalled on LittleFS. Lock order is LVGL-lock -> queue-mutex only;
queue sections never take the LVGL lock.

The mutex is created before the screen's LVGL_LOCK section and
deleted after it in the destructor.
2026-09-04 02:32:57 +00:00
Torlando 403035f115 perf(messages): instant chat open/reopen + deferred full-message view
Uses the microLXMF in-process message-metadata cache (bumped pin
6bea23c -> 58a6eb0, branch feat/conversation-preview-cache):

- Reopening a conversation (and background page fill / paging) is now
  O(1) in-memory once warmed instead of re-reading each message file
  from SPI LittleFS (~230ms/read, ~2.3s per open measured on the T-Deck).
- The cache table is PSRAM-allocated on ESP32; a static .bss placement
  starved internal DRAM and the LVGL task's 8 KiB stack allocation
  failed at boot ("Failed to create LVGL task", hang at startup logo).
- Long-press full-message view now defers to the main loop: the LVGL
  event handler only records the hash; tick_pending_full_message() does
  load_message_content() (uncapped, no msgpack unpack) off the LVGL
  task, then builds the modal. Fixes the capped (600-char) text that
  the in-memory rows carry.
- A peer change cancels an in-flight background fill from the previous
  conversation (it would otherwise prepend the old conversation's
  rows into the new one).

Validated on the T-Deck (T-Deck Plus, 8MB PSRAM): warm opens
reads=3 sync=0-2ms total=15-35ms (was ~2.3s); cold first open per
conversation still pays ~0.6-1.2s disk while warming the cache; no
panics/watchdogs over a ~6-minute interaction session.
Host gates: microLXMF conformance 6/6 (incl. 32-assertion
test_message_metadata_cache), native reference test against pinned
58a6eb0, tdeck build SUCCESS.
2026-09-04 01:30:16 +00:00
Pike 09bf277aa4 chore(messages): remove temporary [PERF] instrumentation
On-device capture confirmed the fix (cold-boot tap: fallbacks=0,
gather=2ms, total=48ms vs 2086ms pre-fix; repeat taps no longer
perpetually fall back). All [PERF] timing markers and stage
variables are removed from refresh()/show()/render_route; the
branch is now clean of diagnostic code.
2026-09-03 17:46:38 +00:00
Pike 8204506b3b perf(messages): persist preview warm-up; cache empty-content tails
On-device [PERF] capture decomposed the remaining ~1s tap latency:
each uncached conversation costs one SPI-LittleFS metadata read
(~230ms). Repeat taps still paid 2 perpetual fallbacks (empty-
content tails never got cached), and the first tap after every
boot paid all 9 (the in-memory repop was never committed to the
index).

- bump microLXMF pin c8d3156 -> 6bea23c (feat/conversation-
  preview-cache): preview_valid index flag so 'cached empty'
  differs from 'unpopulated'; bounded preview copy (the old
  strncpy read past the non-terminated content Bytes); public
  commit_index().
- refresh() re-pops empty-content tails as a valid cached
  preview and arms a one-shot deferred index commit;
  UIManager::update() drains it out-of-lock (flush_pending_
  index_commit, between drops and mark-read) so the warmed
  previews persist and the next cold boot reads them from the
  index.
- fix the [PERF] skip-log total= wrap (printed p_t_diff - p_t0,
  a uint32 underflow; total was already p_t_diff).

[PERF] instrumentation stays in this commit (temporary,
marked); it is removed before merge once the fix is validated
on-device.
2026-09-03 17:12:52 +00:00
Pike af4019f933 diag(messages): [PERF] stage timings for tap-to-visible path
Temporary, removable instrumentation (marked [PERF] throughout):
- refresh(): gather / diff / rebuild stage ms + metadata-fallback count,
  one [PERF] convlist line per call (skip vs build)
- show(): unhide + focus-group ms
- render_route(MESSAGES): whole route window incl. hide_all_screens
All INFO-level so they survive DEBUG-off and land in the serial capture.
2026-09-03 16:02:35 +00:00
Pike ff79a87e42 perf(messages): revalidate conversation rows, skip LVGL rebuild when unchanged
refresh() unconditionally ran lv_obj_clean(_list) and recreated 5-7 LVGL
objects per row on every call - every navigation back to Messages, every
750ms coalesced inbound batch, and the periodic name-resolution sweep -
even when nothing had changed. With the store read now O(1) (index
preview cache), that widget churn was the remaining gap vs NomadNet /
Network / Maps, which build once and just unhide.

refresh() now gathers row data first (index preview + hash + unread, no
message-file I/O) and diffs it against the rendered rows; the rebuild
only happens when data actually changes. Rows keep focus-group
membership and the screen keeps its scroll position on the no-change
path. The per-refresh 'Found N conversations' log drops from INFO to
DEBUG (serial output under the LVGL lock stalls the render task).

get_conversations() ordering is deterministic (last_activity desc with
peer-hash tie-break), so the index-aligned comparison is stable.
2026-09-03 15:20:24 +00:00
Pike cb6a9d38e6 perf(messages): read conversation preview from store index cache
refresh() previously opened + parsed each conversation's newest message
file (load_message_metadata) on every list refresh, which dominated
list-load time on SPI LittleFS. It now reads the per-conversation
last-message preview + timestamp from the store's in-memory index
(O(1), zero I/O) and falls back to load_message_metadata only when the
index has no cached preview — the first refresh after a firmware
upgrade or after a corrupt-tail drop — then writes the preview back
through so the fallback happens at most once per conversation per
firmware generation. The write-through is skipped when drops were
queued (the drained deletes move the tail, so a preview written for the
old tail would be stale for one refresh).

Host benchmark (x86 + POSIX fs, 24 msgs/conv): the per-conversation
store-load work drops from ~0.149ms (9 convs) / ~0.342ms (20 convs) to
~0.002ms / ~0.005ms, and the cold-boot path (store reconstructed from
disk) matches the warm path because the preview now persists in the
index.

Bumps the microLXMF pin to c8d3156 (feat/conversation-preview-cache)
in platformio.ini, the release audit, and the native reference test.
Adds tests/microlxmf/bench_conversation_list_load.cpp (baseline vs
index warm/cold) and the bench target in tests/microlxmf/CMakeLists.txt.
2026-09-03 14:30:03 +00:00
Torlando 918f89d63e fix(messages): drop corrupt last message instead of hiding the row
When a conversation's newest message is unreadable, walk the index
newest-to-oldest for the newest readable preview and queue the
unreadable messages for deletion (deferred out of the LVGL lock,
drained by UIManager::update()). The store's delete_message() commits
the index and updates last_message_hash, so the next refresh converges
to the same preview and the row never disappears over one bad message.
2026-09-03 04:14:52 +00:00
Torlando 8ed9551a02 perf(messages): fast conversation-list load + real unread badges
refresh() did heavy unnecessary work per conversation, on the LVGL
render task under the render lock:
  - get_messages_for_conversation() copied the full 256-slot hash array
    (8KB) just to read the newest hash (messages.back());
  - load_message() on that hash read the payload file ~3x, JSON-parsed
    it twice (including the large hex 'packed' blob), hex-decoded and
    msgpack-unpacked the whole message — to extract a 30-char preview
    and a timestamp;
  - the unread badge was rendered from a hardwired 0 even though the
    store maintains and persists unread_count.

Now:
  - newest hash via the new O(1) MessageStore::get_last_message_hash
    (index tail — always hot-tier, no I/O, no array copy);
  - preview/timestamp via load_message_metadata (single open + filtered
    parse, the fast path ChatScreen already uses for the same fields);
  - unread badge from MessageStore::get_conversation_unread_count;
  - badge cleared + mark-read on open (click) and when a message lands
    in the currently-viewed chat, with the LittleFS index commit
    deferred out of the LVGL lock (UIManager::update), matching the
    existing deferred display-name write-through pattern.

Pins microLXMF 59ca70a (PR #10, temporary branch head) for the two new
accessors.
2026-09-03 02:38:03 +00:00
Pike 902df6554c ui(map): contrast pin labels against active basemap
Marker (pin) labels previously inherited the app's default text color,
which reads white and disappears on light basemaps. Set each label's
text color in applyFrame() against the active style: black by default
(light basemaps osm-bright/positron/toner) and white only on the one
dark basemap (dark-matter). Re-evaluated every frame so a style switch
re-colors visible labels on the next applied frame.

Add a contract test pinning the dark-basemap detection and the
black-by-default / white-on-dark ternary.
2026-09-02 21:20:24 +00:00
Torlando 40c1869765 fix(lxmf): drop SIGNATURE_INVALID inbound messages
A message whose source identity is KNOWN but whose signature fails to
validate is spoofed or malicious and must not be rendered. The
opportunistic (on_packet) and direct (on_resource_concluded) router
paths already reject these, but the propagated (store-and-forward) path
in process_propagated_lxmf queues them without a signature check, so
UIManager::on_message_received is the single choke point that covers
all three inbound routes.

Drop the message at the top of on_message_received — before the key
request, location ingest, persistence, chat render, and notification
beep — when !signature_validated() && reason == SIGNATURE_INVALID.
SOURCE_UNKNOWN (first contact) is untouched: those still render and
trigger the bounded key request from PR #92. Validated messages are
unaffected.

Add a source-level contract test locking in the drop gate's ordering
relative to every side effect and its enum specificity.
2026-09-02 17:51:52 +00:00
Torlando 11fd91bf03 fix(lxmf): make key-request policy C++11-compatible
The tdeck toolchain rejects aggregate initialization of Entry (NSDMI
makes the class non-aggregate under C++11). Use a default constructor
and field assignment in record_request.
2026-09-02 17:06:02 +00:00
Torlando 97fc8b4562 chore(lxmf): drop temporary location-ingest diagnostic from release path
Greptile P2 on PR #92: the per-inbound-location-frame INFOF logged peer
bytes, precise coordinates, and source-clock skew unconditionally in the
release build. The missing-map-pin investigation is closed (announce
timing, verified physically), so remove the temporary instrumentation
and its ingest_diag_result plumbing, and update the feature comment to
describe the final rate-limit semantics instead of 'remove with the
fix'.
2026-09-02 17:05:23 +00:00
Torlando 3de986daed fix(lxmf): bound automatic path requests to 30 min per identity
Greptile P2 on PR #92: with 64 tracked sources, evicting the oldest
entry discarded its open cooldown, so a flood of distinct bogus
identities reset other sources' windows and forced unbounded path
requests (each answered by every peer holding the announce).

- Per-identity cooldown 5 min -> 30 min.
- The table no longer evicts an open window: while all 64 slots hold
  unexpired windows, never-before-seen identities are deferred until a
  slot frees (at most one 30-minute window) instead of dropping
  someone else's cooldown. Open windows are only ever pruned after
  they expire.
- Aggregate bound: at most kMaxTrackedSources automatic path requests
  per rolling 30-minute window, capping the worst-case network cost of
  a rotating-identity flood.

Host tests extended: saturated-table deferral, 512-identity flood bound
(one window and across consecutive windows), lazy expiry/prune,
boundary checks at the 30-minute cooldown.
2026-09-02 17:05:14 +00:00
Torlando 2eec34dfe1 test: extract unknown-source key-request policy + host tests
Move the rate-limit/cooldown/cap decision out of the UIManager.cpp
anonymous namespace into a pure, header-only policy
(UI/LXMF/UnknownSourceKeyRequest.h) so it is host-testable without the
ESP/microReticulum stack. UIManager keeps only the side effect
(Transport::request_path).

Adds tests/native/test_unknown_source_key_request.{cpp,py} (24 checks,
ASan+UBSan): new-source request, 5-min cooldown boundary, re-record
resets the window, per-source independence, 64-entry cap with oldest
eviction, and steady-state cost.
2026-09-02 16:22:56 +00:00
Torlando e9fe63e9e4 diag: request network keys for unknown LXMF sources (Sideband request-keys equivalent)
on_message_received now fires a rate-limited RNS path request when an
LXMF message arrives from a SOURCE_UNKNOWN sender, mirroring Sideband's
'Query Network For Keys' button (RNS.Transport.request_path). Any peer
that already knows the source's announce (hub, phone, other node) can
answer with the cached identity, letting the NEXT message from that
peer validate and reach location ingest.

5-minute per-source cooldown, 64-entry cap. App-only diagnostic;
remove with the other ingest diagnostics once the root cause closes.
2026-09-02 16:22:23 +00:00
Torlando 0f6309c5a6 fix: widen MapTileStore path buffer for 31-char pack IDs
Raise MapTileStore::PATH_CAPACITY from 64 to 80 so the mount buffer
(PATH_CAPACITY + 4) holds the 80-character mounted tile path produced by
a maximum-length (31-char) pack ID with full-width tile coordinates.
Pre-fix, any tile with a two-digit x or y returned INVALID_ARGUMENT from
the mount-prefix snprintf, which MapTilePack maps to IO_ERROR and the UI
shows as 'Tile I/O error' even though the file exists and is intact.

The host regression test added in the prior commit now passes under both
strict C++11 and ASan/UBSan; reverting this change fails it.
2026-08-28 22:44:27 +00:00
Torlando 7c96ffac82 fix(maps): cancel stale tile rendering safely 2026-08-20 04:53:00 +00:00
Torlando d9cb64a1ea feat(maps): resolve visible tiles without span indexes 2026-08-20 04:53:00 +00:00
torlando-agent[bot] 26de7c3e41 fix(nomadnet): keep partial refresh status in chrome 2026-08-18 01:30:14 +00:00
torlando-agent[bot] 1027667f5a feat(nomadnet): integrate dynamic partial rendering 2026-08-17 23:49:13 +00:00
torlando-agent[bot] b6faec8104 feat(nomadnet): add bounded partial core 2026-08-17 19:54:41 +00:00
torlando-agent[bot] 27ac127f5e fix: preserve eight-column NomadNet tables 2026-08-17 16:50:21 +00:00
torlando-agent[bot] ca73ff2e3f fix: show NomadNet navigation progress promptly 2026-08-17 15:47:41 +00:00
torlando-agent[bot] 5b76abfdbe fix: align NomadNet cache response policy 2026-08-17 14:50:00 +00:00
torlando-agent[bot] 2f0f7900ca feat: enforce NomadNet limits and observability 2026-08-17 09:11:10 +00:00