159 Commits
Author SHA1 Message Date
agessaman 31e56b3766 chore(release): prepare 1.1.0
Fold the Unreleased section into 1.1.0 and date it. The 1.1.0 heading was
written on 2026-09-13 but never tagged, so everything merged since belongs in
it; the section also carried two separate "### Added" headings from earlier
merges, now consolidated. All 79 entries preserved.

Extend the v1.0.0 -> v1.1.0 upgrade guide with the meshcore 2.3.14 requirement,
migration 24, and the configuration and behavior changes that landed after the
heading was first written.
2026-09-20 11:44:34 -07:00
agessaman cc6ca00b4e docs(keywords): say the newline escape is a backslash, not a slash (#277)
The report used /n and got two literal characters. The escape works the same in
every keyword path — I checked the test command's template engine too — so the
defect is that the instructions never said "backslash" and used `test` as their
example, which is the one keyword a built-in command owns. That command renders
through a different engine and silently drops the mesh-info placeholders listed
in the same block.

Also documents the configparser continuation form, which needs no escape, and
corrects two packet-capture examples that still showed a naive local timestamp
where the bot has published UTC with a Z suffix since v1.0.0 (#276).
2026-09-20 10:58:17 -07:00
agessaman b5e61306b2 Merge remote-tracking branch 'origin/dev' into feat/region-code-warnings
# Conflicts:
#	CHANGELOG.md
2026-09-20 10:43:24 -07:00
Gerard Hickey 0fe22d2b4f docs: Documentation for developing local commands (#259)
* docs: Documentation for developing local commands

Signed-off-by: Gerard Hickey <hickey@kinetic-compute.com>

* Updated for developing using the local directory

Signed-off-by: Gerard Hickey <hickey@kinetic-compute.com>

* Minor updates and corrections

Signed-off-by: Gerard Hickey <hickey@kinetic-compute.com>

* Implementing recommendations from agessaman

Signed-off-by: Gerard Hickey <hickey@kinetic-compute.com>

* Update CHANGELOG.md

Signed-off-by: Gerard Hickey <hickey@kinetic-compute.com>

---------

Signed-off-by: Gerard Hickey <hickey@kinetic-compute.com>
2026-09-20 10:42:04 -07:00
Gerard Hickeyandagessaman af05c2ef7f feat(website): Include local commands in generated HTML (#287)
* feat(website): Include local commands when generating website

Signed-off-by: Gerard Hickey <hickey@kinetic-compute.com>

* feat(website): Disabled commands are excluded from generated HTML

Signed-off-by: Gerard Hickey <hickey@kinetic-compute.com>

* Updated CHANGELOG.md

Signed-off-by: Gerard Hickey <hickey@kinetic-compute.com>

* fix(website): resolve enablement through the bot's own config helpers

The generated page is meant to list what the bot actually answers, so it
has to read the config the same way the bot does. Three places where it
did not:

The local `config.ini` overlay was never read. `read_config` parsed only
the base file, while the bot overlays `<local_dir_path>/config.ini` on
top of it, and that overlay is exactly where the web viewer saves a local
plugin's settings. A local command disabled from the settings page still
showed up on the site.

`is_command_enabled` derived the section name and the legacy `enabled`
spellings by hand, duplicating `BaseCommand._derive_config_section_name`
and guessing that a legacy key lives in the command's own section. Half
of them do not: `[Jokes] joke_enabled` disables `Joke_Command`. It now
calls `command_section_name` and `read_enabled`, which share the alias
table with the runtime and the settings UI.

That table was missing the same-section `[Joke_Command] joke_enabled` and
`[DadJoke_Command] dadjoke_enabled` spellings, which both commands accept
through their own fallback. The settings page read only the `[Jokes]`
form, so a bot disabled the same-section way displayed as enabled.

Also fold the duplicated local-commands-dir resolution into
`resolve_local_commands_dir`, take `bot_root` as a `MinimalBot` argument
instead of assigning it post-construction twice, and restore the `hidden`
attribute check that `generate_samples` lost when it moved to
`filter_commands`.

---------

Signed-off-by: Gerard Hickey <hickey@kinetic-compute.com>
Co-authored-by: agessaman <adam@gessaman.com>
2026-09-19 22:33:54 -07:00
Gerard Hickeyandagessaman f9b1cbddb8 feat(command): Add contact command (#293)
* feat(command): Add contact command

The contact command allows the user to ask the bot to provide its
public key as a clickable contact. This allows the user to then DM the
bot outside of a channel.

Signed-off-by: Gerard Hickey <hickey@kinetic-compute.com>

* fix(contact): harden matching and self_info handling, add tests and docs

Review follow-ups on the contact command:

- Match through _cleaned_content_matches so a non-matching message keeps its
  original content. The raw cleanup_message_for_matching call rewrote
  message.content in place during the keyword scan, which is the #267
  regression every other command already avoids.
- Match against self.keywords so a configured alias works.
- Read self_info defensively (dict or object, may be absent) and validate the
  public key instead of sending a literal "<None:1:None>" over the air.
- Drop the hardcoded '!' prefix strip in execute; matching already normalizes
  the content, and prefixes are configurable.
- Forward skip_channel_check to the base can_execute.
- Remove copy-paste docstring leftovers from the roll command.
- Move [Contact_Command] into the alphabetical run in config.ini.example and
  fix the "chanels" typo.
- Add tests, a docs/command-reference.md entry, and flesh out the changelog.

---------

Signed-off-by: Gerard Hickey <hickey@kinetic-compute.com>
Co-authored-by: agessaman <adam@gessaman.com>
2026-09-19 20:19:38 -07:00
Adam Gessaman 9939e22575 fix(core): release the radio lock while waiting for ACKs (#291)
Radio commands were serialized per method call, so a composite command held the lock for everything it awaited. A DM's send_msg_with_retry kept every other command waiting through all of its ACK timeouts (up to ~36 s with three attempts), stalling channel replies, other DMs, scheduled sends, and automatic message fetching. req_regions_sync did the same for each neighbor scope reply.

Serialize at the frame instead: wrap CommandHandler.send() on the handler instance, which is where every meshcore command writes its frame and waits for the radio's immediate reply. There is still at most one in-flight companion frame, paced as before, and library methods that call self.send are covered without a proxy. Waits for ACKs and remote responses now run outside the lock.

Add MeshCoreBot.radio_session() for short sequences that must not interleave, and use it for region-scoped channel sends so set scope, send, and restore stay together. The per-call lock never guaranteed that: a DM retry loop queued between set_flood_scope and send_chan_msg sent its flood attempts under the channel's scope.
2026-09-19 09:12:07 -07:00
Gerard Hickey 6686564a25 feat(website): add --link-css and --embed-css for custom CSS (#288)
generate_website.py can layer custom CSS on top of the built-in style chosen with --style. --embed-css appends a local file to the page's <style> block so the page stays a single file, and --link-css adds a stylesheet link after it, so both override built-in rules of equal specificity. The flags can be combined (the linked stylesheet loads last), apply to --sample pages, and an unreadable --embed-css file stops generation with an error. The website docs gain a custom CSS guide and a reference for the generated page's structure, classes, and theme variables.
2026-09-17 21:47:31 -07:00
agessaman 16cce92c46 feat(region): set region scopes from the Radio page (#283)
Region scopes had to be hand-edited in config.ini, and the radio's own default
region could not be set from the bot at all. Both now live on the Radio page.

Node Settings gains a Default Region Scope field beside the path hash size,
which is the firmware's own setting — NodePrefs.default_scope_name and
default_scope_key, the region the radio falls back to for any send the bot does
not scope itself. Firmware that has no such command is reported as not having
answered rather than shown as an empty field, and a stored key that is not the
stored name's hash is flagged, because the radio routes by the key and the name
beside it is only a label. A build-flag default stores the bare name while
hashing the '#' form, so the comparison normalizes first.

Clearing it does not go through the meshcore library, because none of its reset
paths can: set_default_flood_scope(None) raises on len(None), "" earns
ILLEGAL_ARG from the firmware, and "*" only works because the frame is padded by
character count rather than by the name it wrote, landing one byte short of the
length the firmware reads as "a scope follows". The firmware's own contract for
clearing is a bare CMD_SET_DEFAULT_FLOOD_SCOPE, so that is what goes out. The
same padding bug is why a device scope name must be ASCII: a multi-byte
character displaces the transport key out of its field.

A separate Region Scopes card covers the bot's side — [Channels] flood_scopes as
an explicit choice between "reply whatever the scope" and an allowlist, and
outgoing_flood_scope_override. It writes config.ini, queues a hot reload, then
polls that reload and reports what the bot did with it, including saying plainly
when nothing picked it up. Per-channel flood_scope.<channel> entries are listed
read-only, since they are the reason a channel can ignore the default.

The two settings interact in a way worth stating on the page: once the bot sends
a scoped message it restores with set_flood_scope("*"), which leaves the radio in
forced-unscoped mode, so the device default stops applying until the bot scopes
another send.

Also from building it:

- flood_scopes left in [Bot] is surfaced rather than hidden. CommandManager
  still honours it when [Channels] is empty, so a page that ignored it would
  show "replies to every scope" while the bot enforced an allowlist — and
  saving "off" hands the allowlist straight back to [Bot], which the banner now
  says instead of claiming the old key stops mattering.
- outgoing_flood_scope_override = none is read as global flood on the send path,
  as it already was everywhere else. send_channel_message tested the raw value
  against a fixed tuple, so the lowercase spelling became the region "#none".
- Scope normalization moved to modules/flood_scope.py so the viewer, a separate
  process, can validate a typed name without importing the bot's command
  machinery. CommandManager delegates to it.
- Dark mode named only .border, so a .border-start divider drew Bootstrap's
  light #dee2e6 onto a dark card. All four directional utilities now match.
2026-09-16 07:59:10 -07:00
agessaman 10b6b6b01d fix(region): report warnings withheld for want of a contact
The known-contact gate for DM delivery returns before anything is recorded, so
a bot that keeps no contacts dropped every warning and wrote no rows: the page
showed "No warnings decided yet" indefinitely beside a status card saying it
was sending, and nothing distinguished a clean mesh from one where every
warning was being discarded. Only a debug log said otherwise.

Withheld warnings are now counted per local day in bot_metadata — a counter
rather than event rows, because this fires ahead of the cooldowns that would
rate-limit rows, so logging each one would bury the decisions worth reading.
The status card shows it, and the empty log explains itself and points at
channel delivery.

Also from the verification pass:

- The percent test built its value with config.set and a doubled %, which
  set() is the one input that cannot produce the bug. It now uses read_string
  with a bare %, and I checked it fails against raw=False by substituting
  DEFAULT_MESSAGE — the actual failure mode.
- docs still said the bot identifies itself by public key, contradicting the
  section forty lines above. On this path it is name-only, and that means the
  self-exemption is spoofable; the doc says so.
- delivered_today folded dry runs in with real sends, so a morning's preview
  read as afternoon transmissions. previewed_today is now separate and the
  figure follows the current mode.
- TRIM(MIN(channel)) so a whitespace-padded historical row cannot become the
  display name for a merged channel.
2026-09-16 00:28:33 -07:00
agessaman 4c0f43499e fix(region): close four defects found in adversarial review
**A `global` verdict now requires RF correlated to the message.** It was also
reachable through `_is_confirmed_global_flood`'s argument-from-absence route
("no scope-eligible packet in the window, therefore unscoped"), so a row that
literally said TC_FLOOD, or a scoped ADVERT that the GRP_TXT filter excluded,
came back as "no region code". That inference is fine for deciding whether a
`*` in flood_scopes authorizes a reply; it is not fine for accusing someone of
a misconfiguration. Channel messages still correlate through the payload match
(#255), so the ordinary case is unaffected.

**A `%` in the warning message no longer wedges config reload.** `_get` read
without `raw=True`, so configparser's interpolation raised, the error was
swallowed, and the bot transmitted the default wording instead. Worse, the
bot's own `_validate_config_snapshot` iterates `config.items(section)` and
would reject every hot reload until the file was hand-edited. The save endpoint
now rejects `%` outright and caps the message at 500 characters, and the read
is raw so a hand-edited value is still honored.

**One unreachable sender no longer eats the daily cap forever.** The cap counted
failed attempts but the per-sender cooldown did not, so a node the radio cannot
reach was retried every `min_unscoped_messages` messages indefinitely and no
real offender was ever warned. Both count attempts now. The regression test
fails with three sends against the old filter.

**The sender is a display name, not an identity.** MeshCore's CHANNEL_MSG_RECV
carries no public key, so `sender_pubkey` was always empty on this path and the
only identity is a prefix anyone with the channel key can forge. DM delivery
now requires a contact the radio already holds, which bounds the bot to nodes
it knows and stops failed sends spending cap slots. Two tests asserted the
opposite because the fixture supplied a pubkey the radio never sends; they now
run with what the call site actually passes, and the docs no longer claim
pubkey identity.

Also: a negative `max_warnings_per_day` fell back instead of clamping to the
"unlimited" sentinel; tallies group case-insensitively so a `#` or case change
does not split a channel; retention uses the same clock the rows are written
in; the status figure shows delivered with attempts beside it, rather than a
count of four next to "last warning: none yet"; and the channel bars are scaled
over classified traffic so their width equals the percentage printed beside
them (33.7% was drawn at 28.7%).
2026-09-16 00:14:20 -07:00
agessaman 0c1aa4fd15 fix(region): never warn the synthetic "Channel User" sender
A channel message with no "Name: " prefix falls back to a stand-in name.
That is not a node: every such message shares the identity, so they would
accumulate a run together and then get a DM addressed to a contact that
does not exist. The hook now passes no sender for those, so they are still
counted but can never earn a warning. The literal is a named constant and a
test pins the hook's guard to it.

Also folds the channel body-budget formula into models.channel_body_limit
rather than letting the web viewer keep a fourth copy of
"max(130, 160 - len(name) - 2)"; BaseCommand and CommandManager now call it
too, and 158 becomes models.DM_BODY_LIMIT.
2026-09-15 22:39:34 -07:00
agessaman bb3bfe321b feat(region): warn senders whose channel messages carry no region code
Closes #279.

A MeshCore client with no region configured sends every channel message as
a plain FLOOD, which every repeater on the mesh rebroadcasts. This adds two
things: free observation of how much of that the bot hears, and an opt-in
warning to the senders.

Observation classifies each channel message as scoped, global or unknown and
tallies it per channel per local day. It costs one upsert and no airtime, and
it is what lets an operator see the size of the problem before deciding to
spend airtime on it. The classification runs ahead of the flood_scopes
allowlist, because an unscoped message is exactly what that allowlist drops.

Warnings only ever fire on positive RF evidence of an unscoped FLOOD. Absence
of correlation is not proof that a sender omitted a region, so it classifies
as unknown and stays quiet. They are off by default, start in dry run, and are
fenced by min_unscoped_messages, a per-sender cooldown, a mesh-wide cooldown
and a daily cap on attempts. Dry run consumes the same budget it previews, so
the log is what going live would put on the air, not an upper bound. All three
limits read from the database, so a restart cannot release a burst.

Channel-delivered warnings go out at global scope on purpose: the recipient is
by definition outside any region the bot replies under.

New Settings -> Region Warnings page in the web viewer covers all of it. Its
dark-mode striped rows exposed a pre-existing base.html bug where Bootstrap's
light-theme text color survived on a dark row background (~1.3:1); fixed there
for every table in the app.
2026-09-15 22:34:46 -07:00
agessaman 98893a4e4f chore(release): prepare v1.1.0 2026-09-14 20:47:09 -07:00
agessaman 5f89708461 Merge remote-tracking branch 'origin/dev' into pr262
# Conflicts:
#	CHANGELOG.md
2026-09-07 20:29:26 -07:00
agessaman ae4fd55051 Merge pull request #254 from ComchanNet/add_support_for_shlink_message_filter
Add shlink support and rework response templates onto a state machine.

Merged via this branch rather than the PR head: the fork is org-owned, so
"allow edits by maintainers" does not grant push access and the review fixes
could not be pushed to #254 directly. Roger Fedor's four commits are included
with authorship intact, followed by two commits addressing the review.
2026-09-07 16:24:32 -07:00
agessaman 4c76ee3dbe fix: address review findings on shlink shortener and template parser
Security: a shlink deployment with short_url_website unset POSTed the operator's
API key to v.gd. _normalize_base falls back to the public default and the shlink
branch guarded only on a missing key, never on a missing base. shlink now requires
an explicit base and refuses a v.gd/is.gd host outright rather than sending an
X-Api-Key there.

Correctness: _shorten_url_with_gd lost its response.ok guard when the backends were
split out, so a 502 whose body starts with http was returned as the short URL and
transmitted over RF. Restored, with a warning. The regression test that should have
caught this passed only because it left mock_resp.text as a MagicMock; it now uses
a realistic proxy maintenance body.

_shorten_url_with_shlink never checked the status either. Shlink reports failures as
RFC 7807 problem details, which parse as JSON and simply lack shortUrl, so a bad API
key was indistinguishable from an unshortenable URL at DEBUG. It now warns on a
non-OK status, a non-JSON body, and a missing shortUrl. Dropped the shortUrlSlug
fallback: that is a request field, not a response field, and returning it puts a
bare slug where a link belongs.

Timeouts and connection errors are back on their own DEBUG handler. They had fallen
through to the broad handler, whose level rose to ERROR in the same diff, so a
routine intermittent uplink logged "Unexpected error shortening URL" on every reply.

Parser: _GREEDY_ARG_FILTERS was tested before the quoted-argument branch, so
prefix_if_nonempty, the one filter already in shipped configs, could not use the
quoted syntax this PR adds. {path_distance|prefix_if_nonempty:"Dist {sender}: "}
rendered raw template text. A quote immediately after the ':' now selects the quoted
grammar; anything else stays greedy, so config.ini.example's
`prefix_if_nonempty: | Path Dist: ` keeps its pipe and its whitespace.

Blocking HTTP on the event loop: the path command rendered its reply prefix inline
from async code, so a slow shortener stalled radio RX, MQTT and every other handler
for the full 5s timeout. Added format_piped_template_async and switched the path
command to it. The test command still renders synchronously through the sync
check_keywords dispatcher; making that async is a separate change, so the filter now
warns once when it runs on the loop.

Naming: one operation should not have two names in an operator-facing DSL. shorten
and shorten_url are aliases in both response templates and feed formats, and
if_nonempty is canonical with if_notempty as an alias, so a filter chain copied
between a feed format and a response_format works either way.

Also: reduced _build_create_shlink_url to the one parameter it uses and dropped its
dead query/startswith lines, renamed the shlink POST callable from `get`, documented
the config argument on format_piped_template, stopped gating the render trace on an
unrelated parameter and logging field values (sender IDs, user phrases) with it, and
moved the changelog entry from Fixed to Added and Changed.
2026-09-07 15:45:49 -07:00
Gerard Hickey 97dc2bbf91 docs: Update docs concerning waev.app broker authentication
Signed-off-by: Gerard Hickey <hickey@kinetic-compute.com>
2026-09-04 21:49:53 -04:00
Adam Gessaman 52b76d319c Merge branch 'dev' into docs/mqtt-update 2026-08-29 14:06:34 -07:00
Roger Fedor 31fc107e5f Fixed url shortener test case and updated linting errors. Updated documentation and changelog to reflect changes. 2026-08-29 16:00:46 -05:00
Adam Gessaman 1740ca8072 fix(mqtt): resolve reconnect storm and improve token renewal handling
This commit addresses several issues related to MQTT connections. It prevents MQTT brokers from entering a reconnect storm by ensuring that the packet-capture watchdog only triggers a reconnect when necessary, avoiding duplicate CONNACKs and spurious disconnects. Additionally, it ensures that renewed MQTT auth tokens are properly enforced by implementing a clean reconnect process.

New configuration options are introduced: `mqttN_keepalive` for setting the PINGREQ interval per broker, and `mqttN_jwt_reconnect_on_renew` to control reconnect behavior after token renewal. Documentation has been updated to reflect these changes and to clarify the importance of unique client IDs for brokers in the same cluster.
2026-08-28 22:13:42 -07:00
agessaman 7debb92e9e feat(template): add a hops_min filter for gating on route length
{firstlast_distance|prefix_if_nonempty: | F/L Dist: } prints "| F/L Dist: N/A"
on a direct message: the distance helpers return the literal "N/A" when there
is no path to measure, prefix_if_nonempty only asks whether the value is
non-empty, and "N/A" is. Suppressing that needed a gate, and pathbytes_min was
the only one available—so it was being used as a stand-in for "did this
message take any hops", which is not what it asks. It asks how the path is
encoded, so pathbytes_min:2 also discards a one-byte multi-hop path whose
distance is real and measurable.

hops_min:N asks about the route itself. hops_min:1 drops a clause on a direct
message and nothing else:

    hops_min:1   pathbytes_min:2
    direct         cleared      cleared
    1-byte 2 hops  kept         cleared   <- the difference
    2-byte 2 hops  kept         kept
    unknown        cleared      cleared

An unknown hop count clears the value: a gate that cannot confirm the route
should suppress rather than guess, matching pathbytes_min.

The hop-count logic moves to utils.message_hop_count so the filter and
BaseCommand.get_hops_display_values share one implementation instead of two.
It returns None for unknown rather than 0, since a gate must not read
"cannot tell" as "direct".
2026-08-25 10:03:30 -07:00
agessaman d2b83481f7 feat(placeholders): expose {packet_hash} through the placeholder service
Adds the 16-char MeshCore packet identity hash to the three formatters that
render user-facing templates: [Keywords] responses, the test command's
response_format, and the path command's reply_prefix. It is read only from
routing_info, which is attached solely from an RF packet correlated to the
message, so a hash from an unrelated transmission is never shown. Missing,
empty, and the 0000000000000000 error sentinel all render empty.

get_packet_hash_placeholder returns "" rather than a False sentinel. Empty
gives byte-identical output on every path—prefix_if_nonempty already tests
falsiness and str.format renders "" as nothing—whereas False would print
the literal string "False" into a mesh reply from any consumer of
get_standard_placeholder_fields that formatted it directly.

TestCommand.format_response now builds on get_standard_placeholder_fields
instead of duplicating it. That also honors the extra argument its docstring
already promised but silently ignored, and picks up the path-string hop
fallback, so {hops} reads "0" rather than "?" for a message whose hops and
routing_info are both unset but whose path says "Direct (0 hops)".

Documents the placeholder in config.ini.example, the default config that
core.py writes on a fresh install, the README and docs/path-command-config.md.
The generated config was also missing {elapsed}; added that too.
2026-08-25 09:16:09 -07:00
Gerard Hickey dcc7cfb665 docs: Update docs concerning waev.app broker authentication
Signed-off-by: Gerard Hickey <hickey@kinetic-compute.com>
2026-08-23 16:59:16 -04:00
agessaman fc63039a48 fix: address Codex review of tonight's work
Correctness:

- {path_distance} was always blank in production. The resolution code that
  builds repeater_info dropped latitude/longitude in every branch, so the
  calculator could never find a coordinate. My tests passed hand-built
  dicts straight to the calculator and never exercised the builder, which
  is why they stayed green. Coordinates are now carried through all four
  construction sites, and the new tests drive _lookup_repeater_names via
  its lookup_func hook so the real builder runs.

- The #80 route guard was defeated two ways in the channel handler. When
  the RF data was an uncorrelated fallback, control fell through to the
  raw-hex and routing_info fallbacks below, which took the route from the
  unrelated packet anyway; the guard had actually made that path
  reachable. message.routing_info was also assigned unconditionally, and
  the path command reads it. "Not attributable" is now a terminal branch
  and the routing_info hand-off checks provenance.

- Same fix was incomplete for DMs: routing_info was captured and turned
  into path_info before the provenance check ran, so the later check only
  declined to overwrite an already-wrong value. Guarded at the source.

- Rendering could transmit for real. Capture only intercepts
  send_response, but advert calls send_advert() directly and
  send_response_chunked never checked capture_sink. Chunked sends are now
  captured, and rendering is opt-in via BaseCommand.render_safe (default
  False) instead of a denylist that cannot be complete. This also closes
  the DM-only leak: schedule is not marked safe, so {cmd:schedule} can no
  longer broadcast configuration to a channel.

- Multi-part rendered output was rejoined into one oversized send.
  Scheduled messages are now split to the RF body budget and sent through
  send_channel_messages_chunked, on character boundaries so multi-byte
  text is not corrupted.

- _last_path_distance_km is instance state that was only set on success,
  so an invalid path request could show the previous request's distance.
  Reset at the start of every execute().

- Stale-contact retries were only counted on non-OK results, so timeouts
  and exceptions left a contact eligible forever and could recreate the
  storm. All failed attempts count now.

Web viewer:

- A failed config reload was reported as a successful save. The API and
  UI now distinguish "saved and active" from "saved, restart needed".

- Preview count was unbounded; clamped to 1-20.

- update_ini_values is a read-modify-replace with no locking, so two
  concurrent viewer saves could lose one. Serialised behind a lock.
2026-08-22 00:16:03 -07:00
agessaman 29a2e7b412 feat(web): manage scheduled messages from the viewer (#174)
Adds a Schedule page that lists every [Scheduled_Messages] entry with its
next run time and supports add, edit and delete. Writes go to config.ini
and queue a config reload, so schedules change without restarting the
bot, which was the actual request.

It edits the config section the bot already reads rather than
introducing a database table. reload_config() already re-runs
setup_scheduled_messages() with rollback, so there is nothing to keep in
sync and the schedule command lists exactly what the page shows.

Validation runs through the same parsers the scheduler uses, so the UI
cannot accept a schedule the bot would later reject. The builder
composes cron from plain-language options and previews the next five
runs; entries the bot cannot run are listed as "Not scheduled" with the
reason instead of being hidden, since a typo that silences a message is
what an operator most needs to see. The 15-minute floor for {cmd:...}
messages is enforced at save time too.

Also relabels the radio Disconnect button to "Stop Bot" behind a
confirmation (#240). It was never a radio-only disconnect: the main loop
runs while self.connected is true, so disconnecting exits the process.
That surprised an operator running under tmux with nothing to restart
it. disconnect_radio()'s docstring now says so as well.
2026-08-21 23:32:45 -07:00
agessaman 133a3bb595 docs: correct prefix API guidance and document pipx installs
Two documentation gaps behind open questions.

The [External_Data] repeater_prefix_api_url comment claimed that leaving
it empty "disables prefix command functionality" (#70). That is not what
happens: the prefix command answers from the bot's own database of heard
repeaters, and the API only augments it with node counts from a wider
dataset. The wrong comment is a plausible reason the question was asked
at all. Corrected, and the JSON contract is now documented in the command
reference for anyone serving their own endpoint, since map.w0z.is is
defunct and has no drop-in replacement.

Also documented the pipx path (#222). The installer already grew a
virtualenv in July, which covered the PEP 668 half of that report, but
the unanswered part was where config.ini, the database and local/ live
under a pipx install. They are all resolved relative to the directory
containing config.ini, which means a bare `meshcore-bot` picks up
whatever is in the current directory — so the guidance is to pass an
absolute --config. Both console scripts are already smoke-tested from an
installed wheel in CI, so this path is supported rather than incidental.
2026-08-21 23:00:55 -07:00
agessaman a764186eb2 feat(scheduler): airtime guards for {cmd:...}, plus missed docs
Two guards on command placeholders in scheduled messages, neither
configurable, because both exist to protect a shared medium:

A 15-minute floor. A schedule containing {cmd:...} that fires more often
is rejected at startup with an error rather than quietly running slower.
The interval is sampled and measured by the tightest gap between
firings, so "0,1 * * * *" is correctly treated as every 60 seconds and
not as hourly. Schedules without a command placeholder are unaffected.

The command's own cooldown_seconds now applies to a render. Scheduling
is not a way around the rate a command was configured to run at. The
execution is recorded before the command runs, matching execute_commands,
so a slow or failing render cannot be retried straight past the cooldown.

Also fills documentation gaps from the preceding commits:

- --install-extras was only in the script's own --help; documented in
  service-installation.md (including alongside --update-venv) and
  upgrade.md.
- weather-service.md documented weather_alarm's once-a-day scheduling
  with no route to more than one forecast a day, which is exactly what
  people go there looking for. It now points at {cmd:wx ...} and notes
  that sunrise/sunset still belong to weather_alarm.
2026-08-21 21:33:54 -07:00
agessaman b0148506a6 feat(scheduler): broadcast any command's output with {cmd:...}
A scheduled message can now embed the reply of any bot command:

    [Scheduled_Messages]
    0 6,12,18 * * * = Public:{cmd:wx Seattle}

The schedule side was never the missing piece — [Scheduled_Messages]
keys have been 5-field cron with @presets for a while, so multiple
times, intervals and @hourly were all already expressible. What was
missing was any way to get a service's output into a scheduled message:
placeholders were limited to mesh info (contact counts). That gap is
why services were growing their own schedule parsers.

CommandManager.render_command_output runs a command for its text alone.
MeshMessage.capture_sink makes that safe: when set, send_response
collects the reply and transmits nothing, checked before the
_last_response bookkeeping so a background render cannot overwrite the
response captured for a real user's command.

Degrades quietly in every failure mode — unknown, disabled, admin-only,
non-renderable, timed out or silent commands expand to nothing, so raw
{cmd:...} text is never put on the air, and a message left empty is not
sent at all. Command output is not re-scanned, so a reply containing
{cmd:...} cannot recurse. Commands that transmit directly rather than
returning text (announcements) are refused outright, since rendering
them would broadcast for real.

Bounded by [Bot] scheduled_command_timeout_seconds (default 30).
2026-08-21 21:25:16 -07:00
agessaman 4d8624494c feat(path): expose {path_distance} in the path command reply prefix
Total path distance was already a solved problem here: the test command
computes it and exposes {path_distance} through the piped-template
engine, with pipe filters and docs. The path command just had no way to
reach it, since reply_prefix went through plain str.format.

Rather than add a second distance implementation, this reuses what
exists. BaseCommand.get_standard_placeholder_fields now returns the
shared placeholder set so subclasses can extend it, the path command
renders reply_prefix through format_piped_template with an added
{path_distance}, and the sum itself uses utils.calculate_distance over
the nodes the path command has already resolved.

Distance covers sender -> each hop -> bot and is deliberately blank when
the chain cannot be measured end to end (unresolved hop, prefix
collision, missing or 0,0 coordinates, unknown sender, no configured
bot position), so a partial sum is never reported as the real distance.

Because the prefix now supports pipe filters, an empty value takes its
label with it: {path_distance|prefix_if_nonempty:📏 }

Supersedes #198, which added a parallel haversine, a show_path_distance
flag, and a distance_traveled key across 10 locales.
2026-08-21 21:10:03 -07:00
agessaman 25671f28de Merge pull request #241 from ajquick/dev
Add configurable PacketCapture observer name

Adds an optional [PacketCapture] observer_name used as the MQTT `origin`
for packet and status reporting, letting the observer identity differ
from the connected MeshCore device name. Unset keeps the previous
behavior. The public key backing origin_id, authentication, and topic
resolution is unchanged.

CHANGELOG conflict resolved by keeping both Unreleased entries.
2026-08-21 20:33:42 -07:00
agessaman 5bdd6302c1 fix(airplanes): switch Airplanes API to adsb.lol default 2026-08-21 09:57:23 -07:00
agessaman a594b72a85 fix(neighbors): update zero-hop neighbor discovery and data handling
This commit enhances the handling of zero-hop neighbors in the dashboard and database. It ensures that the **One-hop neighbours** section accurately reflects radios heard directly (MeshCore hop count 0) rather than originators of relayed adverts. The `observed_paths` table now includes nullable `snr` and `rssi` columns for zero-hop advert rows, allowing for better signal reporting. Additionally, a one-time backfill process copies recent zero-hop ADVERTs from the `packet_stream` to `observed_paths`. Documentation and tests have been updated to reflect these changes.
2026-08-12 18:01:28 -07:00
AJ Quick 9eeb226def Add observer_name configuration to packet-capture
Add optional observer name configuration for MQTT.
2026-08-08 15:37:46 +00:00
agessaman 2fea97b21d fix(venv): rewrite shebangs for console scripts after virtualenv relocation
This update ensures that console scripts in the virtual environment do not retain outdated shebangs pointing to temporary build paths. The `--update-venv` command now rewrites shebangs in place, allowing broken installs to recover without a full rebuild. Additionally, all pip invocations now use `python -m pip` to prevent issues with stale shebangs. This change addresses issue #229 and improves the overall reliability of the service installation process.
2026-08-07 21:38:12 -07:00
agessaman 1777805090 merge(main): resolve conflicts preferring dev for #235 2026-08-07 13:56:42 -07:00
agessaman c6b8940b23 docs: surface solar provenance in the nav, drop internal integration log
solar-conditions-provenance.md was in the tree but absent from the
mkdocs nav and unlinked from anywhere, so it never appeared on the docs
site. Added under a Project section and linked from the solar command
entry, where a reader would look for it.

Removes the kg7qin PR integration log — an internal development record
that was excluded from the docs build but still shipped in the repo —
along with the now-dead exclude glob.
2026-08-07 13:45:20 -07:00
Adam Gessaman 9d6dc6fa2d Merge pull request #234 from agessaman/feat/neighbors-mqtt
Implement zero-hop neighbor discovery and improve cycle management
2026-08-06 21:03:36 -07:00
agessaman ae3a3c66fa fix(neighbors): move the airtime cooldown into the service
Move the 15-minute gap between cycles from the DM command to the service, as
MIN_CYCLE_GAP_SECONDS. The command was one caller among several: the scheduler
retried a failed cycle on its own 300s backoff, keyed off last_neighbors_publish
which a failed cycle never stamps, so a discover round whose acknowledgement was
lost went back on the air every five minutes. run_neighbors_cycle now refuses a
cycle inside the gap whoever asks, and the scheduler's retry backoff waits out
whatever remains of it. Failures that never reached the radio still stamp
nothing, so re-checking a disconnected radio stays on the short backoff.

The command asks the service for the remaining wait instead of computing its own,
so the two cannot drift, and rewinds the sender's per-user cooldown to expire
with the shared one. The command manager records an execution before calling
execute(), so a refusal had been consuming the sender's full 15 minutes: told to
wait one more minute, they would retry and be refused for another fourteen.
Rewound rather than cleared — a refusal reply is airtime too.
2026-08-05 19:38:36 -07:00
agessaman ce68e7ab7e fix(neighbors): close three gaps in the cycle guards
Follow-up to 7cffe1d; all three confirmed against meshcore 2.3.8.

Charge the node-wide cooldown for the transmission, not the result.
last_neighbors_publish is stamped only by a cycle that completes, so a discover
request whose acknowledgement was lost spent the airtime and left the clock at
zero -- another sender could immediately start a second round. The service now
stamps last_neighbors_attempt before the request, and the command rations on
whichever stamp is later. A cycle that bails out before touching the radio still
records nothing, so retries stay possible.

Restore the contact path when req_regions_sync returns None. send_anon_req gives
up early when change_contact_path reports an error -- which is also what a lost
acknowledgement for an applied path change looks like -- and that return skips
the library's own reset_path. req_regions_sync collapses it to None, previously
treated as a plain no-response. On the common "neighbour did not answer" path
the extra reset is a redundant device command: no airtime, idempotent, and it
re-syncs the contact cache.

Stop reporting a rejected restore as a success. reset_path returns an ERROR
event for a device rejection or its own response timeout rather than raising, so
the helper logged "restored flood path" either way and hid a contact left pinned
to zero-hop. It now inspects the event and warns with the reason.
2026-08-05 19:01:23 -07:00
agessaman 7cffe1d7c0 fix(neighbors): bound cycle airtime and correct evidence labelling
Four review findings on this branch, all confirmed against the code and the
installed meshcore 2.3.8:

Single-flight the discovery cycle. The command guarded only its own task, so
the scheduler's independent call could overlap a manual cycle and each round
would collect into the other's discover window. run_neighbors_cycle is now a
guard around the cycle body, refusing whichever trigger arrives second.

Make the 15-minute cooldown per node rather than per sender. The base class
rations per user, but the cost here is mesh airtime: users could take turns and
keep the radio discovering continuously. Measured from the last cycle that
produced a result, so the scheduler's cycles count and a cycle that bailed out
without transmitting does not start the clock.

Window the neighbours evidence label. The combined viewer applied `days` only
to mesh_connections while reading every lifetime row from neighbor_links, which
is deliberately never pruned — so a link last heard years ago kept claiming a
recent path-derived edge was a current direct neighbour.

Match that label on full public keys too. MeshGraph.add_edge does not promote a
1-byte edge that has no public key, but still fills in the keys discovery
supplied, so the 3-byte prefix comparison alone left confirmed neighbours
labelled singlebyte. Truncating our keys to 2 chars instead would relabel every
other node sharing that byte.

Restore a contact's flood path after an interrupted scope request. send_anon_req
pins a path-less contact to zero-hop and restores it after the send with no
try/finally, so our own budget cancelling the request left the contact pinned
and every later message to it sent direct-only.

The meshcore>=2.3.8 pin needed no change: PyPI publishes 2.3.8 now.
2026-08-05 18:32:42 -07:00
agessaman 035ec963b6 docs: clarify APScheduler cron usage and day-of-week numbering
- Updated `config.ini.example`, `command-reference.md`, and `configuration.md` to clarify the use of APScheduler for cron scheduling, emphasizing the difference in day-of-week numbering compared to Vixie cron.
- Added examples and notes to prevent confusion regarding the interpretation of cron expressions, particularly for users transitioning from Vixie cron conventions.
- Enhanced documentation to guide users in using the correct syntax for scheduling messages, ensuring better understanding and usability.
2026-08-04 20:33:26 -07:00
agessaman ebc66992dd feat(neighbors): zero-hop neighbour discovery in packet capture
Port the observer firmware's neighbours feature into the bot's packet capture
service, by way of meshcore-packet-capture (upstream PRs #42/#43). On a long
interval the bot asks which repeaters it hears directly and records each
confirmed link with its measured SNR.

This is the strongest link evidence the bot collects: a first-party RF
measurement between two full 32-byte public keys. Path inference works from
1-3 byte prefixes with no keys, and complete_contact_tracking.hop_count
over-claims zero-hop (800 claimed vs 68 corroborated on the live database).

modules/neighbors_discovery.py keeps upstream's public names so its fixes and
tests stay portable. Two deliberate divergences:

- No command_lock plumbing. _SerializedCommands in modules/core.py already
  serialises and paces every radio command, strictly more than upstream's
  reentrant lock did.
- neighbors_collect_scopes defaults off. Upstream's zero-hop scope probe
  relies on a neighbour not being a known contact; this bot tracks contacts,
  and for a repeater with no stored path the library reaches zero-hop by
  calling change_contact_path() then reset_path() -- mutating the device's
  contact table per neighbour. Scope requests also hold the radio lock for
  their whole round trip (~25s), stalling bot replies. The default cycle is
  one command plus a passive listen window, during which the bot stays
  responsive.

Evidence lands in neighbor_links and neighbor_observations (migration 22)
rather than mesh_connections, which cannot persist provenance. The viewer
exposes it as evidence=neighbors on /api/mesh/edges and a Neighbours Only
mode on the mesh page, with populated public keys and real SNR; confirmed
neighbours also relabel edges in the combined view and count as
provenance-trusted when framing the initial map.

neighbors_enabled is the single switch. Every enabled broker publishes once
it is on (mqttN_neighbors defaults true; set false to hold one back). The
topic derives from each broker's packets topic with the last segment swapped,
so a templated broker gets meshcore/{IATA}/{PUBLIC_KEY}/neighbors -- the
topic the firmware uses -- instead of an unrelated flat one. A derived
location-routed topic is skipped with a warning when no iata is set, rather
than publishing into meshcore/XYZ/... on a shared namespace. Snapshots are
non-retained: heard_secs_ago is relative to publish time, so a retained copy
would read as current days later.

Also adds a DM-gated `neighbors` command (the 12h interval floor makes
waiting for the scheduler impractical), which acks immediately and reports in
a second message once the window closes.

Requires meshcore >= 2.3.8 for send_node_discover_req / req_regions_sync.
2026-08-04 19:24:18 -07:00
agessaman 75fda424c2 feat(web-viewer): implement per-payload-type multibyte share tracking
- Added a new feature to track and display the multibyte share of packets by payload type in the dashboard.
- Introduced a new database migration to store per-payload-type multibyte encoding data in the daily rollup.
- Updated the dashboard to visualize the multibyte share as a stacked bar chart, reflecting the share of each day's packets that took a multibyte path.
- Enhanced the API to provide raw counts for each payload type, ensuring accurate representation in the dashboard.
- Adjusted the frontend to maintain consistent color coding for payload types and improve the overall user experience.
- Updated tests to validate the new multibyte share functionality and ensure data integrity.
2026-07-31 21:23:51 -07:00
agessaman 40307bd8cb feat(mesh): enhance mesh graph performance and caching
- Implemented caching for multi-byte mesh aggregation, allowing concurrent requests to share computation and reducing CPU load.
- Updated the web viewer to coalesce live edge events and refresh the mesh graph every 30 seconds, improving responsiveness on busy networks.
- Adjusted logging behavior in the web viewer to respect the configured log level, ensuring efficient logging without excessive duplicate entries.
- Added configuration options for mesh graph caching duration, enhancing user control over performance settings.
2026-07-30 13:16:16 -07:00
agessaman b8b1647b27 feat(retention): enhance data retention and resource limits for Raspberry Pi
- Implemented chunked deletion for data retention, allowing for smoother cleanup processes without monopolizing SQLite's writer lock.
- Configured retention settings to delete in batches with pauses, improving performance on SD-card installations.
- Updated Linux service installers to allocate 1GB of memory and 200% CPU, providing better resource management for Raspberry Pi workloads.
- Added detailed configuration examples for Raspberry Pi in documentation to guide users on optimal settings.
2026-07-30 08:06:21 -07:00
agessaman 07d6737008 feat(graph): optimize mesh graph persistence and data retention
- Updated mesh graph path splitting and aggregation to run in SQLite, improving performance by avoiding Python materialization.
- Defaulted graph persistence to batched writes for new installations, reducing WAL churn and SD-card writes.
- Enhanced data retention to execute shortly after startup, independent of the nightly maintenance schedule.
- Added a table-specific index for `mesh_connections` to support window and retention queries, ensuring efficient data access.
2026-07-30 07:23:06 -07:00
agessaman 7ea34a2672 feat(web-viewer): hide flood hop buckets under 0.1% of the series
The flood tail decays over roughly twenty hops in bars under a pixel
tall. Buckets holding less than 0.1% of the series are no longer drawn,
which on the live mesh takes the axis from 64 buckets to 44.

What is left out is reported, not dropped: a line under the chart reads
"1,467 further flood packets (0.9%) sit in 20 hop buckets below 0.1%
each, and are not drawn." An unannounced cut would be the same lie as a
silently truncated category list.

Two details that matter for honesty:

Percentages still divide by the whole series, never by the drawn subset,
so removing the tail cannot inflate the bars that remain. The smallest
surviving bar on live data is 0.118%, which is what it was before.

A withheld bucket inside the axis is null rather than zero. Zero would
claim no packets travelled that far, which is a different and false
statement; null draws nothing and says nothing. Padding gaps stay zero,
because there it is true.

The threshold applies to the flood series only — the node series is
small enough to draw in full — and the underlying computation still
covers the whole 64-hop protocol range.
2026-07-29 23:16:51 -07:00
agessaman 196f42789c refactor(web-viewer): fill card height, trim hop axis, drop live feed
Three layout and cost changes.

Hops away and Roles now grow into their cards instead of leaving dead
space under a fixed-height canvas. Chart.js needs a positioned parent
with a real height, so the body becomes a flex column and the chart takes
the slack via flex-basis 0 — height: 100% would resolve against an
auto-height parent and collapse.

The hop axis now ends at the last hop that carries an observation rather
than at the last non-empty bucket of the padded union. On a quiet window
that collapses 64 buckets to 13; on a full one it changes nothing,
because the flood series really does have packets at every hop out to 63
— 14 of them at hop 63, and 14.1% of all flood traffic beyond hop 20.
That tail is real data, so it is drawn rather than truncated.

Dropped the Live Activity card. It opened three SocketIO subscriptions
and re-rendered on every packet to duplicate /realtime, which is a page
that already does it better. The dashboard now costs one snapshot read
per poll and holds no streaming subscriptions. A test asserts it stays
that way rather than merely that the markup is gone.
2026-07-29 23:07:38 -07:00
agessaman 3a4d7aab33 feat(web-viewer): add flood-packet distance to the hops chart
Keeps the advert series — nodes by their closest observed path — and
adds arriving flood packets by how far they had already travelled.

The two answer different questions and, on the live mesh, disagree
usefully: nodes peak at 2-3 hops and fall away quickly, while flood
traffic peaks at 5 and holds a long tail past 16. A close-in
neighbourhood absorbing flood from well beyond it.

One series counts nodes (2.8k) and the other packets (156k), so raw
counts on a shared axis would flatten the node series into the baseline.
Both are drawn as a share of their own total, with absolute counts in
the tooltip, and padded onto one contiguous hop range so the bars line
up. They also cover different spans — 7 days of adverts against
whatever packet_stream retains — so each is labelled with its own
window instead of being presented as one period.

Watch the units. observed_paths.path_length is a BYTE count, so hops are
path_length / bytes_per_hop. packet_stream.path_len is already a HOP
count, with the byte length carried separately as path_byte_length. A
17-hop 3-byte path is path_length 51 in one table and path_len 17 in the
other. Applying either rule to the other silently rescales the axis and
the only symptom is a chart that looks a bit off, so both conventions
are now pinned by tests against the shapes real rows take.

Flood packets carry no sender identity and observed_paths holds only
adverts, so the flood series cannot be reduced to a shortest path per
node the way the advert series is. It is a per-packet distribution, and
the tooltip says so.
2026-07-29 22:45:21 -07:00