Region scopes had to be hand-edited in config.ini, and the radio's own default
region could not be set from the bot at all. Both now live on the Radio page.
Node Settings gains a Default Region Scope field beside the path hash size,
which is the firmware's own setting — NodePrefs.default_scope_name and
default_scope_key, the region the radio falls back to for any send the bot does
not scope itself. Firmware that has no such command is reported as not having
answered rather than shown as an empty field, and a stored key that is not the
stored name's hash is flagged, because the radio routes by the key and the name
beside it is only a label. A build-flag default stores the bare name while
hashing the '#' form, so the comparison normalizes first.
Clearing it does not go through the meshcore library, because none of its reset
paths can: set_default_flood_scope(None) raises on len(None), "" earns
ILLEGAL_ARG from the firmware, and "*" only works because the frame is padded by
character count rather than by the name it wrote, landing one byte short of the
length the firmware reads as "a scope follows". The firmware's own contract for
clearing is a bare CMD_SET_DEFAULT_FLOOD_SCOPE, so that is what goes out. The
same padding bug is why a device scope name must be ASCII: a multi-byte
character displaces the transport key out of its field.
A separate Region Scopes card covers the bot's side — [Channels] flood_scopes as
an explicit choice between "reply whatever the scope" and an allowlist, and
outgoing_flood_scope_override. It writes config.ini, queues a hot reload, then
polls that reload and reports what the bot did with it, including saying plainly
when nothing picked it up. Per-channel flood_scope.<channel> entries are listed
read-only, since they are the reason a channel can ignore the default.
The two settings interact in a way worth stating on the page: once the bot sends
a scoped message it restores with set_flood_scope("*"), which leaves the radio in
forced-unscoped mode, so the device default stops applying until the bot scopes
another send.
Also from building it:
- flood_scopes left in [Bot] is surfaced rather than hidden. CommandManager
still honours it when [Channels] is empty, so a page that ignored it would
show "replies to every scope" while the bot enforced an allowlist — and
saving "off" hands the allowlist straight back to [Bot], which the banner now
says instead of claiming the old key stops mattering.
- outgoing_flood_scope_override = none is read as global flood on the send path,
as it already was everywhere else. send_channel_message tested the raw value
against a fixed tuple, so the lowercase spelling became the region "#none".
- Scope normalization moved to modules/flood_scope.py so the viewer, a separate
process, can validate a typed name without importing the bot's command
machinery. CommandManager delegates to it.
- Dark mode named only .border, so a .border-start divider drew Bootstrap's
light #dee2e6 onto a dark card. All four directional utilities now match.
The known-contact gate for DM delivery returns before anything is recorded, so
a bot that keeps no contacts dropped every warning and wrote no rows: the page
showed "No warnings decided yet" indefinitely beside a status card saying it
was sending, and nothing distinguished a clean mesh from one where every
warning was being discarded. Only a debug log said otherwise.
Withheld warnings are now counted per local day in bot_metadata — a counter
rather than event rows, because this fires ahead of the cooldowns that would
rate-limit rows, so logging each one would bury the decisions worth reading.
The status card shows it, and the empty log explains itself and points at
channel delivery.
Also from the verification pass:
- The percent test built its value with config.set and a doubled %, which
set() is the one input that cannot produce the bug. It now uses read_string
with a bare %, and I checked it fails against raw=False by substituting
DEFAULT_MESSAGE — the actual failure mode.
- docs still said the bot identifies itself by public key, contradicting the
section forty lines above. On this path it is name-only, and that means the
self-exemption is spoofable; the doc says so.
- delivered_today folded dry runs in with real sends, so a morning's preview
read as afternoon transmissions. previewed_today is now separate and the
figure follows the current mode.
- TRIM(MIN(channel)) so a whitespace-padded historical row cannot become the
display name for a merged channel.
**A `global` verdict now requires RF correlated to the message.** It was also
reachable through `_is_confirmed_global_flood`'s argument-from-absence route
("no scope-eligible packet in the window, therefore unscoped"), so a row that
literally said TC_FLOOD, or a scoped ADVERT that the GRP_TXT filter excluded,
came back as "no region code". That inference is fine for deciding whether a
`*` in flood_scopes authorizes a reply; it is not fine for accusing someone of
a misconfiguration. Channel messages still correlate through the payload match
(#255), so the ordinary case is unaffected.
**A `%` in the warning message no longer wedges config reload.** `_get` read
without `raw=True`, so configparser's interpolation raised, the error was
swallowed, and the bot transmitted the default wording instead. Worse, the
bot's own `_validate_config_snapshot` iterates `config.items(section)` and
would reject every hot reload until the file was hand-edited. The save endpoint
now rejects `%` outright and caps the message at 500 characters, and the read
is raw so a hand-edited value is still honored.
**One unreachable sender no longer eats the daily cap forever.** The cap counted
failed attempts but the per-sender cooldown did not, so a node the radio cannot
reach was retried every `min_unscoped_messages` messages indefinitely and no
real offender was ever warned. Both count attempts now. The regression test
fails with three sends against the old filter.
**The sender is a display name, not an identity.** MeshCore's CHANNEL_MSG_RECV
carries no public key, so `sender_pubkey` was always empty on this path and the
only identity is a prefix anyone with the channel key can forge. DM delivery
now requires a contact the radio already holds, which bounds the bot to nodes
it knows and stops failed sends spending cap slots. Two tests asserted the
opposite because the fixture supplied a pubkey the radio never sends; they now
run with what the call site actually passes, and the docs no longer claim
pubkey identity.
Also: a negative `max_warnings_per_day` fell back instead of clamping to the
"unlimited" sentinel; tallies group case-insensitively so a `#` or case change
does not split a channel; retention uses the same clock the rows are written
in; the status figure shows delivered with attempts beside it, rather than a
count of four next to "last warning: none yet"; and the channel bars are scaled
over classified traffic so their width equals the percentage printed beside
them (33.7% was drawn at 28.7%).
An empty daily-volume chart rendered as an 84px hole with a date range under
it, which reads as broken rather than as no data; it is hidden until there is
something to plot, and the summary says so in words.
"Count only" disabled the rest of the form, which implied those values would
not be saved — they are — and stopped anyone drafting their wording before
turning warnings on. A sentence under the mode selector says it instead.
Bootstrap's .text-warning is #ffc107, which is 1.6:1 on a white card — the
headline percentage and the "daily cap reached" line were close to invisible
in light mode. The chart fills had the same shape of problem in reverse: the
grey "couldn't tell" band sat at 2.1:1 against a white card and 2.9:1 against
a dark one.
Each color is now a token chosen per theme against the surface it lands on:
text clears 4.5:1, fills clear 3:1. Measured in the browser in both themes.
An unselected chip looked identical to a selected one, so clicking it was a
coin flip between adding and removing. The chips now reflect the field, in
both directions, and stay in step when the list is typed by hand.
Also re-read after a save rather than patching the form in place (the status
strip's budget is derived from the cap that was just changed), and stop the
reset-to-default button from blanking the message when the page never loaded.
A channel message with no "Name: " prefix falls back to a stand-in name.
That is not a node: every such message shares the identity, so they would
accumulate a run together and then get a DM addressed to a contact that
does not exist. The hook now passes no sender for those, so they are still
counted but can never earn a warning. The literal is a named constant and a
test pins the hook's guard to it.
Also folds the channel body-budget formula into models.channel_body_limit
rather than letting the web viewer keep a fourth copy of
"max(130, 160 - len(name) - 2)"; BaseCommand and CommandManager now call it
too, and 158 becomes models.DM_BODY_LIMIT.
Closes#279.
A MeshCore client with no region configured sends every channel message as
a plain FLOOD, which every repeater on the mesh rebroadcasts. This adds two
things: free observation of how much of that the bot hears, and an opt-in
warning to the senders.
Observation classifies each channel message as scoped, global or unknown and
tallies it per channel per local day. It costs one upsert and no airtime, and
it is what lets an operator see the size of the problem before deciding to
spend airtime on it. The classification runs ahead of the flood_scopes
allowlist, because an unscoped message is exactly what that allowlist drops.
Warnings only ever fire on positive RF evidence of an unscoped FLOOD. Absence
of correlation is not proof that a sender omitted a region, so it classifies
as unknown and stays quiet. They are off by default, start in dry run, and are
fenced by min_unscoped_messages, a per-sender cooldown, a mesh-wide cooldown
and a daily cap on attempts. Dry run consumes the same budget it previews, so
the log is what going live would put on the air, not an upper bound. All three
limits read from the database, so a restart cannot release a burst.
Channel-delivered warnings go out at global scope on purpose: the recipient is
by definition outside any region the bot replies under.
New Settings -> Region Warnings page in the web viewer covers all of it. Its
dark-mode striped rows exposed a pre-existing base.html bug where Bootstrap's
light-theme text color survived on a dark row background (~1.3:1); fixed there
for every table in the app.
Discover local commands and services in the Plugins UI and route their settings to the local config overlay. Upgrade GitHub Actions to Node 24-capable majors and migrate frontend linting to ESLint 10 flat config. Keep command and service discovery namespaces isolated and align duplicate-name handling with runtime loading.
Add a password type to settings_schema so plugin secrets render as masked inputs and never reach the browser. Stored values stay in config.ini; an empty box on save leaves an existing secret unchanged.
- Addresses #267
- Added `original_content` attribute to `MeshMessage` to store the on-air body of messages, ensuring it remains unchanged during command processing.
- Updated command matching methods across various commands to utilize `_cleaned_content_matches`, which restores original content when necessary, preventing mention stripping from altering the message context.
- Implemented tests to verify that original content is preserved and correctly utilized in command execution and message handling.
Carto now requires an API key for basemaps.cartocdn.com and is retiring
raster tiles, so the mesh map's dark theme rendered an "API KEY REQUIRED"
watermark. OSM hosts no dark tiles of its own -- its Standard layer is
light only -- so switch the dark basemap to OpenFreeMap's vector 'dark'
style, rendered through MapLibre GL via maplibre-gl-leaflet. Leaflet and
the light OSM raster layer are unchanged, and the bridge renders into
tilePane so marker and overlay z-order is untouched.
Fall back to inverted OSM raster tiles when WebGL 2 or the bridge is
unavailable, so a failure degrades to a filtered map rather than a blank
one. The filter targets .leaflet-tile rather than .leaflet-tile-container:
the container is a 0x0 element that its absolutely-positioned tiles
overflow, so a filter there has an empty reference box and paints nothing.
CSP: drop cartocdn, allow tiles.openfreemap.org, and permit blob: workers.
MapLibre spawns its renderer in a worker from a blob: URL; without
worker-src it falls back to default-src 'self' and never starts.
PUT verified original_schedule outside the lock, so a concurrent delete
between the check and the write resurrected the entry as a new one, and
two concurrent renames of the same original left both results present.
The existence check now runs inside _save_scheduled_message_locked
against the same snapshot the duplicate check uses, so create, update and
delete are each a single critical section.
Added a concurrency test: two simultaneous creates of one schedule now
produce exactly one 200 and one 409 with a single entry on disk.
A flood_scopes of "*" alone leaves scope_keys empty while setting
flood_scope_allow_global, and the loader already logs that as an active
allowlist. The handler gated on scope_keys alone, so that configuration
skipped authorisation entirely and admitted absent, uncorrelated and
TRANSPORT_FLOOD traffic. The gate now fires when either is set, so "*"
means global-only rather than everything.
The scheduled-message duplicate check ran outside the write lock inside
update_ini_values, so two concurrent creates for the same schedule could
both pass and the second silently replace the first instead of getting
the 409. Check and write are now one critical section, and delete is too.
Repaired the delete path's error handling while moving it, so an OSError
during the write is still a 500 rather than escaping.
Correctness:
- {path_distance} was always blank in production. The resolution code that
builds repeater_info dropped latitude/longitude in every branch, so the
calculator could never find a coordinate. My tests passed hand-built
dicts straight to the calculator and never exercised the builder, which
is why they stayed green. Coordinates are now carried through all four
construction sites, and the new tests drive _lookup_repeater_names via
its lookup_func hook so the real builder runs.
- The #80 route guard was defeated two ways in the channel handler. When
the RF data was an uncorrelated fallback, control fell through to the
raw-hex and routing_info fallbacks below, which took the route from the
unrelated packet anyway; the guard had actually made that path
reachable. message.routing_info was also assigned unconditionally, and
the path command reads it. "Not attributable" is now a terminal branch
and the routing_info hand-off checks provenance.
- Same fix was incomplete for DMs: routing_info was captured and turned
into path_info before the provenance check ran, so the later check only
declined to overwrite an already-wrong value. Guarded at the source.
- Rendering could transmit for real. Capture only intercepts
send_response, but advert calls send_advert() directly and
send_response_chunked never checked capture_sink. Chunked sends are now
captured, and rendering is opt-in via BaseCommand.render_safe (default
False) instead of a denylist that cannot be complete. This also closes
the DM-only leak: schedule is not marked safe, so {cmd:schedule} can no
longer broadcast configuration to a channel.
- Multi-part rendered output was rejoined into one oversized send.
Scheduled messages are now split to the RF body budget and sent through
send_channel_messages_chunked, on character boundaries so multi-byte
text is not corrupted.
- _last_path_distance_km is instance state that was only set on success,
so an invalid path request could show the previous request's distance.
Reset at the start of every execute().
- Stale-contact retries were only counted on non-OK results, so timeouts
and exceptions left a contact eligible forever and could recreate the
storm. All failed attempts count now.
Web viewer:
- A failed config reload was reported as a successful save. The API and
UI now distinguish "saved and active" from "saved, restart needed".
- Preview count was unbounded; clamped to 1-20.
- update_ini_values is a read-modify-replace with no locking, so two
concurrent viewer saves could lose one. Serialised behind a lock.
The bar had grown to twelve items at full expansion, over half of them
configuration surfaces. Radio, Scheduled Messages, Greeter, Feeds,
Plugins and Configuration now sit behind one gear, so the top level is
about what the mesh is doing (Dashboard, Real-time, Contacts, Mesh
Graph, Multibyte, Logs) and the gear is about how the bot is set up.
Logs stays top level: it is what you reach for while watching behaviour,
not while configuring.
Also added active-page highlighting, which the bar never had. A settings
page lights up both its dropdown entry and the gear, so the current
location is still obvious once a page is one level down.
Conditional entries keep their guards, so Greeter and Feeds appear in
the menu only when enabled.
Adds a Schedule page that lists every [Scheduled_Messages] entry with its
next run time and supports add, edit and delete. Writes go to config.ini
and queue a config reload, so schedules change without restarting the
bot, which was the actual request.
It edits the config section the bot already reads rather than
introducing a database table. reload_config() already re-runs
setup_scheduled_messages() with rollback, so there is nothing to keep in
sync and the schedule command lists exactly what the page shows.
Validation runs through the same parsers the scheduler uses, so the UI
cannot accept a schedule the bot would later reject. The builder
composes cron from plain-language options and previews the next five
runs; entries the bot cannot run are listed as "Not scheduled" with the
reason instead of being hidden, since a typo that silences a message is
what an operator most needs to see. The 15-minute floor for {cmd:...}
messages is enforced at save time too.
Also relabels the radio Disconnect button to "Stop Bot" behind a
confirmation (#240). It was never a radio-only disconnect: the main loop
runs while self.connected is true, so disconnecting exits the process.
That surprised an operator running under tmux with nothing to restart
it. disconnect_radio()'s docstring now says so as well.
mypy cannot track non-Noneness through the intermediate `corroborated`
boolean, so `float(snr)` in the discover-only neighbor branch was flagged
as `Any | None`. Test the value directly; `corroborated` still backs the
`signal_corroborated` field.
This commit enhances the handling of zero-hop neighbors in the dashboard and database. It ensures that the **One-hop neighbours** section accurately reflects radios heard directly (MeshCore hop count 0) rather than originators of relayed adverts. The `observed_paths` table now includes nullable `snr` and `rssi` columns for zero-hop advert rows, allowing for better signal reporting. Additionally, a one-time backfill process copies recent zero-hop ADVERTs from the `packet_stream` to `observed_paths`. Documentation and tests have been updated to reflect these changes.
Removes 26 console.log calls, including contacts.html dumping whole data
structures and sample contact records on every render. Variables and
callback parameters that existed only to feed those logs go with them
(deviceTypes, anyNewDevices, an availableEdges debug block, and a
realtime status handler whose body was nothing but a log), so no new
no-unused-vars warnings are introduced — the count drops 60 to 59.
console.warn and console.error are kept. One console.log survives in
realtime.html: the decoder key-count line is a one-time startup message
with real diagnostic value, and removing it would mean deleting two
counters and their increments inside a loop.
Also normalizes the mesh page's user-visible "Neighbours" strings to
match the project's American spelling. The option value stays
"neighbors", so evidence-mode filtering is untouched.
login.html loads DM Sans and Outfit from fonts.googleapis.com with the
font files on fonts.gstatic.com, but neither host was in style-src or
font-src, so the browser blocked both and the login page silently fell
back to system fonts with CSP violations in the console.
Only consider observed_paths rows with bytes_per_hop in {1,2,3} so NULL
or invalid encodings fall back to out_bytes_per_hop instead of classifying
as 1-byte.
Four review findings on this branch, all confirmed against the code and the
installed meshcore 2.3.8:
Single-flight the discovery cycle. The command guarded only its own task, so
the scheduler's independent call could overlap a manual cycle and each round
would collect into the other's discover window. run_neighbors_cycle is now a
guard around the cycle body, refusing whichever trigger arrives second.
Make the 15-minute cooldown per node rather than per sender. The base class
rations per user, but the cost here is mesh airtime: users could take turns and
keep the radio discovering continuously. Measured from the last cycle that
produced a result, so the scheduler's cycles count and a cycle that bailed out
without transmitting does not start the clock.
Window the neighbours evidence label. The combined viewer applied `days` only
to mesh_connections while reading every lifetime row from neighbor_links, which
is deliberately never pruned — so a link last heard years ago kept claiming a
recent path-derived edge was a current direct neighbour.
Match that label on full public keys too. MeshGraph.add_edge does not promote a
1-byte edge that has no public key, but still fills in the keys discovery
supplied, so the 3-byte prefix comparison alone left confirmed neighbours
labelled singlebyte. Truncating our keys to 2 chars instead would relabel every
other node sharing that byte.
Restore a contact's flood path after an interrupted scope request. send_anon_req
pins a path-less contact to zero-hop and restores it after the send with no
try/finally, so our own budget cancelling the request left the contact pinned
and every later message to it sent direct-only.
The meshcore>=2.3.8 pin needed no change: PyPI publishes 2.3.8 now.
Port the observer firmware's neighbours feature into the bot's packet capture
service, by way of meshcore-packet-capture (upstream PRs #42/#43). On a long
interval the bot asks which repeaters it hears directly and records each
confirmed link with its measured SNR.
This is the strongest link evidence the bot collects: a first-party RF
measurement between two full 32-byte public keys. Path inference works from
1-3 byte prefixes with no keys, and complete_contact_tracking.hop_count
over-claims zero-hop (800 claimed vs 68 corroborated on the live database).
modules/neighbors_discovery.py keeps upstream's public names so its fixes and
tests stay portable. Two deliberate divergences:
- No command_lock plumbing. _SerializedCommands in modules/core.py already
serialises and paces every radio command, strictly more than upstream's
reentrant lock did.
- neighbors_collect_scopes defaults off. Upstream's zero-hop scope probe
relies on a neighbour not being a known contact; this bot tracks contacts,
and for a repeater with no stored path the library reaches zero-hop by
calling change_contact_path() then reset_path() -- mutating the device's
contact table per neighbour. Scope requests also hold the radio lock for
their whole round trip (~25s), stalling bot replies. The default cycle is
one command plus a passive listen window, during which the bot stays
responsive.
Evidence lands in neighbor_links and neighbor_observations (migration 22)
rather than mesh_connections, which cannot persist provenance. The viewer
exposes it as evidence=neighbors on /api/mesh/edges and a Neighbours Only
mode on the mesh page, with populated public keys and real SNR; confirmed
neighbours also relabel edges in the combined view and count as
provenance-trusted when framing the initial map.
neighbors_enabled is the single switch. Every enabled broker publishes once
it is on (mqttN_neighbors defaults true; set false to hold one back). The
topic derives from each broker's packets topic with the last segment swapped,
so a templated broker gets meshcore/{IATA}/{PUBLIC_KEY}/neighbors -- the
topic the firmware uses -- instead of an unrelated flat one. A derived
location-routed topic is skipped with a warning when no iata is set, rather
than publishing into meshcore/XYZ/... on a shared namespace. Snapshots are
non-retained: heard_secs_ago is relative to publish time, so a retained copy
would read as current days later.
Also adds a DM-gated `neighbors` command (the 12h interval floor makes
waiting for the scheduler impractical), which acks immediately and reports in
a second message once the window closes.
Requires meshcore >= 2.3.8 for send_node_discover_req / req_regions_sync.
- Replaced inline timestamp formatting logic with a call to `format_relative_timestamp` for improved readability and maintainability.
- Modularized feed formatting functions by moving them to `feed_format`, enhancing code organization and reusability.
- Updated references in `feed_manager` and `web_viewer` to utilize the new modular functions, ensuring consistent behavior across modules.
- Adjusted tests to reflect changes in function imports and ensure proper functionality of the new structure.
- Introduced a locking mechanism to ensure journal mode is initialized once per configuration section, preventing redundant database operations.
- Added a validation function to check for required repeater tables, raising an error with actionable messages if any are missing.
- Updated the DBManager to handle journal modes more efficiently, ensuring that only persistent modes are applied across connections.
- Refactored the RepeaterManager to utilize the new validation function, improving error handling during initialization.
- Enhanced tests to cover new journal mode behaviors and validate repeater table existence, ensuring robust database management.
- Added a new feature to track and display the multibyte share of packets by payload type in the dashboard.
- Introduced a new database migration to store per-payload-type multibyte encoding data in the daily rollup.
- Updated the dashboard to visualize the multibyte share as a stacked bar chart, reflecting the share of each day's packets that took a multibyte path.
- Enhanced the API to provide raw counts for each payload type, ensuring accurate representation in the dashboard.
- Adjusted the frontend to maintain consistent color coding for payload types and improve the overall user experience.
- Updated tests to validate the new multibyte share functionality and ensure data integrity.
- Implemented caching for multi-byte mesh aggregation, allowing concurrent requests to share computation and reducing CPU load.
- Updated the web viewer to coalesce live edge events and refresh the mesh graph every 30 seconds, improving responsiveness on busy networks.
- Adjusted logging behavior in the web viewer to respect the configured log level, ensuring efficient logging without excessive duplicate entries.
- Added configuration options for mesh graph caching duration, enhancing user control over performance settings.
- Implemented chunked deletion for data retention, allowing for smoother cleanup processes without monopolizing SQLite's writer lock.
- Configured retention settings to delete in batches with pauses, improving performance on SD-card installations.
- Updated Linux service installers to allocate 1GB of memory and 200% CPU, providing better resource management for Raspberry Pi workloads.
- Added detailed configuration examples for Raspberry Pi in documentation to guide users on optimal settings.
- Updated mesh graph path splitting and aggregation to run in SQLite, improving performance by avoiding Python materialization.
- Defaulted graph persistence to batched writes for new installations, reducing WAL churn and SD-card writes.
- Enhanced data retention to execute shortly after startup, independent of the nightly maintenance schedule.
- Added a table-specific index for `mesh_connections` to support window and retention queries, ensuring efficient data access.
The flood tail decays over roughly twenty hops in bars under a pixel
tall. Buckets holding less than 0.1% of the series are no longer drawn,
which on the live mesh takes the axis from 64 buckets to 44.
What is left out is reported, not dropped: a line under the chart reads
"1,467 further flood packets (0.9%) sit in 20 hop buckets below 0.1%
each, and are not drawn." An unannounced cut would be the same lie as a
silently truncated category list.
Two details that matter for honesty:
Percentages still divide by the whole series, never by the drawn subset,
so removing the tail cannot inflate the bars that remain. The smallest
surviving bar on live data is 0.118%, which is what it was before.
A withheld bucket inside the axis is null rather than zero. Zero would
claim no packets travelled that far, which is a different and false
statement; null draws nothing and says nothing. Padding gaps stay zero,
because there it is true.
The threshold applies to the flood series only — the node series is
small enough to draw in full — and the underlying computation still
covers the whole 64-hop protocol range.
Three layout and cost changes.
Hops away and Roles now grow into their cards instead of leaving dead
space under a fixed-height canvas. Chart.js needs a positioned parent
with a real height, so the body becomes a flex column and the chart takes
the slack via flex-basis 0 — height: 100% would resolve against an
auto-height parent and collapse.
The hop axis now ends at the last hop that carries an observation rather
than at the last non-empty bucket of the padded union. On a quiet window
that collapses 64 buckets to 13; on a full one it changes nothing,
because the flood series really does have packets at every hop out to 63
— 14 of them at hop 63, and 14.1% of all flood traffic beyond hop 20.
That tail is real data, so it is drawn rather than truncated.
Dropped the Live Activity card. It opened three SocketIO subscriptions
and re-rendered on every packet to duplicate /realtime, which is a page
that already does it better. The dashboard now costs one snapshot read
per poll and holds no streaming subscriptions. A test asserts it stays
that way rather than merely that the markup is gone.
MeshCore carries up to a 64-byte path, so a route can be 64 hops long at
one byte per hop — and proportionally fewer with 2- or 3-byte hashes,
which is where the observed ceilings of 32 and 21 come from.
The chart capped at 32. That number was inherited from the old
dashboard's `BETWEEN 0 AND 32` filters and carried forward without
checking it against the protocol. On the live database it silently
discarded 5,654 flood packets arriving from as far as 63 hops.
Worse for the node series: the cap runs after the per-node MIN(), so a
node whose *closest* route exceeded 32 hops disappeared from the chart
altogether instead of appearing at the far end. The live history holds
one such node, reachable only by a 48-hop path.
The axis can now span 64 grouped categories, so bars drop their rounded
corners and padding above 24 buckets and the x ticks auto-skip. 64 is
kept as a hard ceiling: beyond it the path field could not have held the
route, so the value is corrupt rather than distant.
Keeps the advert series — nodes by their closest observed path — and
adds arriving flood packets by how far they had already travelled.
The two answer different questions and, on the live mesh, disagree
usefully: nodes peak at 2-3 hops and fall away quickly, while flood
traffic peaks at 5 and holds a long tail past 16. A close-in
neighbourhood absorbing flood from well beyond it.
One series counts nodes (2.8k) and the other packets (156k), so raw
counts on a shared axis would flatten the node series into the baseline.
Both are drawn as a share of their own total, with absolute counts in
the tooltip, and padded onto one contiguous hop range so the bars line
up. They also cover different spans — 7 days of adverts against
whatever packet_stream retains — so each is labelled with its own
window instead of being presented as one period.
Watch the units. observed_paths.path_length is a BYTE count, so hops are
path_length / bytes_per_hop. packet_stream.path_len is already a HOP
count, with the byte length carried separately as path_byte_length. A
17-hop 3-byte path is path_length 51 in one table and path_len 17 in the
other. Applying either rule to the other silently rescales the axis and
the only symptom is a chart that looks a bit off, so both conventions
are now pinned by tests against the shapes real rows take.
Flood packets carry no sender identity and observed_paths holds only
adverts, so the flood series cannot be reduced to a shortest path per
node the way the advert series is. It is a per-packet distribution, and
the tooltip says so.
Moving the hop histogram out into its own card left the routing mix as a
26px bar alone in a col-lg-5, next to a full-height chart. Fill it with
something that belongs there: what the traffic is, beside how it is
routed, over exactly the same packets — the totals agree because both
read the dimensioned rows.
This also gives payload_type_name a reason to exist. Migration 0019
added the column, integration.py writes it at capture time, the
refresher backfills it and an index covers it, and until now nothing
read it.
Category lists now roll their tail into "Other" instead of being
truncated at eight rows in the client. Silently dropping the tail left
bars that no longer summed to the total printed beside them; on the live
database that would have hidden 1,015 packets across four payload types.
307 direct neighbours was not plausible, and it was not real.
complete_contact_tracking.hop_count claims 800 zero-hop contacts. Only
68 of them have any one-hop path in observed_paths to corroborate that.
Their stored SNR piles up in a 1.5 dB band — 655 of 800 between 11.25
and 12.75 dB — and their RSSI clusters at -39..-48 dBm. Hundreds of
radios at different distances and terrain cannot land in a 10 dB window.
That is the signature of one strong local link being recorded against
every node whose traffic happened to arrive through it. Their return
paths agree: these "direct" contacts have out_path_len of 3 to 11.
The writer's intent is sound — repeater_manager only stores RSSI/SNR
when signal_info reports hops == 0 — so the field being fed to it does
not mean what the surrounding code assumes. Left as is; this change
stops the dashboard depending on it.
Neighbour membership now comes from path evidence: an advert whose
path_length equals its bytes_per_hop travelled exactly one hop. That
yields 38 nodes in 24h and 124 in 7d, with a plausible spread. Signal is
shown only where the path evidence and the stored hop count agree, which
is 5 and 12 nodes respectively; the rest read "no signal reading" rather
than borrowing a measurement taken on somebody else's link. A 24h/7d
selector bounds the window, capped well under observed_paths' 90-day
retention because a month-old link says nothing about today.
Separately, this fixes a bug I introduced. path_length is a BYTE count,
and with 2- or 3-byte hop encoding a 3-hop path is 6 or 9 bytes long. The
path-length histogram plotted that raw value on an axis readers would
take as hops, overstating distance two- to threefold on a mesh that is
~95% multibyte. It is replaced by a single hops-away chart computed as
path_length / bytes_per_hop, which also retires the histogram built on
the untrustworthy stored hop count. The result is unimodal, peaking at 3
hops and decaying — the shape a mesh should have, and not the bimodal
one the old chart drew.
Three adjustments from review.
Role and device type are the same field twice. Measured on the live
database they disagree on 16 of 11,028 contacts (ten roomservers and a
handful of bots and gateways reporting device Companion); every other
row is repeater/Repeater, companion/Companion, type11/Type11 and so on.
Charting both filled half a card with a copy of the other half. Keep the
role mix, which also carries the type0..type15 bucketing, and move it
into the routing row where the old signal card was.
Drop the tracked-contacts tile. is_currently_tracked does not describe
anything a reader can act on, and node activity is already covered by
nodes-heard and gone-quiet. The known-contacts total moves onto the
coverage tile, which leaves five tiles splitting the row evenly.
Rebuild the signal panel around zero-hop neighbours. Percentiles over
every received message answered no question anyone has: SNR on a relayed
packet measures the last hop into this radio, not the link to the node
that sent it, so averaging across hop counts describes nothing in
particular.
The panel now shows the nodes heard with no repeater in between — how
many, their SNR distribution, and the weakest links named, worst first,
on a fixed -12..+14 dB scale so bars mean the same thing between
refreshes. On the live database that is 307 neighbours, median 12.0 dB,
with eight marginal links surfaced from -9.0 dB down.
This also corrects the source. The plan rejected
complete_contact_tracking.snr as a badly biased 7% sample; in fact it is
populated on exactly the 800 hop_count=0 contacts and NULL on all 10,228
others. It is not a sample of the network, it is a complete census of
the neighbours — which is precisely the population the metric applies
to. message_stats.hops=0 covers only 34 senders by comparison.
The landing page re-ran ~50 aggregate queries five times per load, then
repeated the whole sequence every 30 seconds forever — including in
backgrounded tabs. Against the live 1.44 GB database that was roughly 20
seconds of SQLite work per page load.
Move the work off the request path. A refresher thread in the viewer
process (which already runs migrations, so it works for a split-DB
install) writes two tables: daily_rollup, one row per local date, and
dashboard_snapshot, a single JSON row. A page load now reads one row.
Measured on the live database: first paint 6 requests -> 2,
/api/dashboard/summary p50 1.1 ms (304 in 0.8 ms), /api/stats 130 ms,
and a 0.32 s refresh once a minute in the background.
Make the numbers mean what they say:
- Window selectors are built from each source's retention. The page
offered "30d" and "All" against tables pruned at 7 days, so three of
four choices returned the same figure under a label that denied it.
- The incoming-packet chart reports its measured window instead of
claiming 7 days for a table pruned at 3 — it sat beside a genuine
7-day contacts chart inviting an invalid comparison.
- Days with no source data store NULL and render as gaps. Writing 0
would put a fake cliff at every retention boundary.
- Signal metrics are stored as sums and counts, never means, so any
window re-aggregates correctly.
- Delta chips compare the last two complete calendar days and say so;
the headline above them is a rolling 24 hours.
- Unmapped role ordinals (type0..type15) bucket into "Unknown".
- SNR comes from message_stats, where it is populated on every row, not
from complete_contact_tracking, where it is populated on 7%.
Kill the json_extract scans: packet_stream gains denormalized
route_type_name, payload_type_name, path_len and bytes_per_hop, written
at capture time. Aggregating those from JSON cost 3-6 s per query.
Existing rows convert a bounded batch per tick rather than in one
migration that would rewrite ~180 MB into the WAL and stall bot startup.
A partial index serves as the backfill worklist — without it the "any
rows left?" probe is a full scan costing 4.6 s per tick, and it costs
that after the backfill finishes, because finding nothing still means
reading everything.
Also: replace the per-contact hop-prefix scan with the existing bucketed
matcher and memoize the 7-day chunk set (264 ms -> 35 ms on a synthetic
100k-row database, regression-locked by a test); move the dashboard's JS
and CSS to static files, which removes the CSP nonce requirement for the
bulk of the page; and give cleanup_old_stats a future-timestamp guard,
without which rows dated 2103 are never older than the cutoff and so
live forever.
Deletes the orphaned /stats page, unreachable from the nav and rendering
stub charts that never populated. /api/stats stays as a shim with every
key name intact plus Deprecation and Sunset headers.
All schema changes are additive, so a downgraded codebase can read the
data; it would however need the new schema_version rows removed, since
MigrationRunner rejects versions it does not know.
Changed the default setting for `auto_manage_contacts` from `false` to `device` across configuration files and updated related documentation. This ensures that the device handles auto-addition of contacts while the bot manages capacity. Adjusted comments and documentation to reflect this change for clarity and consistency.
Remaining findings from the same triage pass. None are security
relevant; each produces a wrong user-visible result.
- greeter: treat rollout_started_at as UTC. It is written by SQLite
CURRENT_TIMESTAMP, but datetime.timestamp() read the naive value as
local time, shifting the backfill cutoff by the host's offset (7h on
PT). Users who posted inside that window were never marked as already
greeted and could be sent a welcome they should not have received.
- greeter: write greeted_at in SQLite's own format. The rollout backfill
used isoformat() while every other path used CURRENT_TIMESTAMP. "T"
sorts above a space, so ORDER BY greeted_at interleaved the two
formats wrongly, corrupting duplicate cleanup and the web viewer's
recently-greeted list.
- greeter: let an empty `channels =` disable the command. BaseCommand
reads an empty value as disabled-on-channels, but the greeter
collapsed "key absent" and "key present but empty" into one fallback
and kept greeting via monitor_channels.
- greeter: stop comma-splitting greeting text. channel_greetings split
entries on ",", so "Public:Welcome to the mesh, {sender}!" was stored
as "Welcome to the mesh" with the placeholder silently dropped. A
fragment now starts a new entry only when the text before its first
colon looks like a channel name, which keeps commas, URLs and clock
times attached to the greeting they belong to.
- wxsim_parser: match condition abbreviations longest-first. Substring
matching in dict order let RAIN shadow CHNC. RAIN, so five conditions
lost their "chance" qualifier (rain, snow, drizzle, t-storm, and
FAIR-P.C.).
- wxsim_parser: re-anchor "now" on each parse. current_date and
current_year were fixed at construction and wx_command builds one
parser at startup, so after a few days forecast dates rolled back a
year and staleness checks read permanently true.
- thesportsdb_client: hold the rate-limit read/sleep/write under a lock.
Concurrent callers read the same last_request_time, slept the same
interval and fired together, bursting past the 2.1s throttle. This
exposes an unrelated request fan-out problem in fetch_league_scores;
filed in TODO.md rather than fixed here.
- alert_command: stop duplicating the first incident. The tail treated
any single-line buffer as "header only" and appended incidents[0],
but messages after the first carry no header, so a final chunk holding
one incident had incident 0 pasted onto it. The same confusion inside
the loop also let a second incident be appended past the 130-character
limit.
- feed: reject non-positive poll intervals. -1 is truthy and was stored,
and the poller's `now - last_check >= interval` then treats the feed
as permanently due and re-fetches the URL every cycle; 0 was ignored
while still reporting success. The web viewer, which is the primary
editor, had no validation at all, and a JSON null there raised a
TypeError that aborted the poll cycle for every feed rather than one.
feed_manager falls back to the default for rows written before this.
- transmission_tracker: age out confirmed transmissions that have
repeats. Cleanup removed only repeat_count == 0, so repeated records
accumulated for the lifetime of the process, which matters on a Pi
Zero. Repeat counts are already persisted to packet_stream, so the
longer 30-minute retention loses nothing.
- mesh_graph: weighted-merge avg_hop_position when promoting an edge.
Promotion overwrote the average with the single new observation, so an
edge averaging 2.0 over 4 observations became 9.0 instead of 3.4 and
skewed subsequent path scoring.
- multitest_command: show the path that ends exactly at the display LCP.
The suffix helpers are correct in isolation; the loss happens in the
cluster formatters, where _shrink_display_lcp refuses to shrink a
single-token LCP and the trunk then rendered as "96 ┐", meaning
"everything continues past here" and dropping the bare 96 route while
the header still counted it. Now uses the file's existing "├ common"
marker. Empty suffixes are also no longer handed to the nested
renderer, which mapped them onto the same [] as a route ending at the
inner LCP and drew a row for a route that did not exist.
- sports_mappings: stop shadowing seven unique team nicknames. In one
flat dict a repeated key silently drops the earlier team, so hawks
resolved to the NBA Hawks rather than the Seahawks alias it was added
as, blazers to Kamloops rather than Portland, and rockets to Kelowna
rather than Houston. First definition now wins for hawks, giants,
jets, rangers, kings, blazers and rockets; every shadowed team keeps
its unambiguous full-name alias. The 55 city and abbreviation
collisions (chicago, sf, la, ...) are genuinely ambiguous and stay
last-wins, now pinned by a test so a new collision fails loudly
instead of passing unnoticed.
Four defects in how configuration is written and displayed. All are
reachable without authentication when web_viewer_password is unset,
which is an explicitly supported setup.
- ini_writer: reject sections, keys and values that cannot survive the
round-trip. update_ini_values() wrote free text straight into the
file, so a newline in any value ended the key's line and everything
after it was re-parsed as INI on the next load. Saving a greeting of
"Welcome!\n[Injected]\npwned = true" through the plugin settings
endpoint created a real new section; a repeated section name bricked
startup with DuplicateSectionError. The check lives at the writer
because three separate callers reach it, and raises IniValueError so a
bad payload is a 400 instead of a half-written file. DEFAULT is
refused as well: its keys apply to every section, and
ConfigParser.add_section('DEFAULT') raises.
- settings_store, web_viewer: persist to disk before mirroring into the
in-memory config. The mirror ran first, so a value rejected by the
writer stayed live in the running process despite never reaching
disk. Both routes taking free-form input (the zombie and offline
alert emails) had the same ordering. The zombie route also caught
only OSError, so an IniValueError there escaped as a 500 rather than
the intended 400.
- config_snapshot: redact Discord webhook URLs. Neither
discord_webhook_urls nor the [DiscordBridge] bridge.<channel> keys
matched any redaction rule, so --show-config and /admin/config
printed live webhook secrets in full; anyone holding one can post to
the channel. Telegram's api_token was already redacted. The
bridge. prefix needs its own rule because the varying part is the
channel name, leaving no fixed stem for the substring match.
- mqtt_weather, packet_capture: verify TLS certificates by default.
Both called tls_set(cert_reqs=ssl.CERT_NONE) unconditionally, so the
broker username and password sent immediately afterwards were
readable by anyone able to intercept the connection.
BREAKING: brokers presenting self-signed certificates now fail to
connect until tls_insecure (mqtt_weather) or mqttN_tls_insecure
(packet_capture) is set to true. Both log a warning while enabled.
- Added a tracked `LICENSE` file and updated `pyproject.toml` to include license metadata.
- Enhanced `CHANGELOG.md` with recent changes and clarifications.
- Updated `config.ini.example` and related documentation to reflect clamping behavior for numeric limits in `[Feed_Manager]`.
- Improved startup validation to suggest corrections for unknown and misspelled keys.
- Refactored geocoding and HTTP request handling to run off-thread, preventing event loop stalls.
- Added thread safety to cache management in geocoding functions to avoid race conditions.
- Enhanced the `/api/contacts` endpoint to support pagination, allowing users to specify `page` and `page_size` parameters.
- Added search functionality to filter contacts based on a search term.
- Implemented sorting options for the contacts list, enabling users to sort by various fields.
- Updated the frontend to reflect these changes, including a new pagination UI and improved loading states.
- Introduced caching for multibyte hop chunks to optimize performance when retrieving contact badge evidence.
- Added tests to ensure the new features work as expected and maintain existing functionality.
Resolve the feeds.html conflict by keeping the branch's XSS-safe DOM
construction of the feed-details view and re-adding dev's per-feed and
reset-all error buttons via addEventListener instead of inline onclick,
since /feeds is a nonce-CSP page where inline handlers are blocked.
Update tests/test_feed_manager_extended.py::TestPollFeedPosting to patch
the SafeUrlPolicy.validate_async path (this branch's SSRF refactor) rather
than the removed module-level validate_external_url symbol, matching the
idiom already used throughout that test file.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Added max_posts_per_check to limit the number of items posted per check, defaulting to max_items_per_check for backward compatibility.
- Updated logic in FeedManager to process new items, ensuring that only the specified number of posts are made while examining a larger set of items.
- Enhanced emoji selection to prioritize per-item emojis from the API, falling back to heuristics based on feed names when absent.
- Introduced new truncation and substring functions for formatting messages, improving text handling in feed outputs.
- Added API endpoints to reset feed error counts globally or for individual feeds, enhancing error management capabilities in the web viewer.