Commands are bare by default; `command_prefix` is empty unless configured, so
`!wx` misrepresents a stock install. Shipped release sections are left alone.
Fold the Unreleased section into 1.1.0 and date it. The 1.1.0 heading was
written on 2026-09-13 but never tagged, so everything merged since belongs in
it; the section also carried two separate "### Added" headings from earlier
merges, now consolidated. All 79 entries preserved.
Extend the v1.0.0 -> v1.1.0 upgrade guide with the meshcore 2.3.14 requirement,
migration 24, and the configuration and behavior changes that landed after the
heading was first written.
The report used /n and got two literal characters. The escape works the same in
every keyword path — I checked the test command's template engine too — so the
defect is that the instructions never said "backslash" and used `test` as their
example, which is the one keyword a built-in command owns. That command renders
through a different engine and silently drops the mesh-info placeholders listed
in the same block.
Also documents the configparser continuation form, which needs no escape, and
corrects two packet-capture examples that still showed a naive local timestamp
where the bot has published UTC with a Z suffix since v1.0.0 (#276).
waev.app rejects a token whose exp is more than an hour past its iat, so the
project-wide 24-hour default could never authenticate there — the operator saw
only a bare auth failure with no hint that the lifetime was the problem, while
the same public key worked from meshcoretomqtt.
waev.app hosts now resolve to a 3600s TTL with renewal at 3500s when no value
was configured for them, per broker or globally; a configured value still wins,
and the substitution is logged. Matching is the apex and its subdomains only, so
a lookalike host like waev.app.example.com does not qualify.
Classify every channel message's flood scope, tally it per channel per day,
and optionally warn senders whose messages carry no region code. Warnings are
off by default and start in dry run; they fire only on positive RF evidence of
an unscoped FLOOD, and are fenced by a per-sender cooldown, a mesh-wide
cooldown and a daily cap read from the database.
Adds Region Scopes and Default Region Scope cards to the Radio page so the
bot's own scopes and the radio's firmware default can be set from the web
viewer.
# Conflicts:
# CHANGELOG.md
* feat(website): Include local commands when generating website
Signed-off-by: Gerard Hickey <hickey@kinetic-compute.com>
* feat(website): Disabled commands are excluded from generated HTML
Signed-off-by: Gerard Hickey <hickey@kinetic-compute.com>
* Updated CHANGELOG.md
Signed-off-by: Gerard Hickey <hickey@kinetic-compute.com>
* fix(website): resolve enablement through the bot's own config helpers
The generated page is meant to list what the bot actually answers, so it
has to read the config the same way the bot does. Three places where it
did not:
The local `config.ini` overlay was never read. `read_config` parsed only
the base file, while the bot overlays `<local_dir_path>/config.ini` on
top of it, and that overlay is exactly where the web viewer saves a local
plugin's settings. A local command disabled from the settings page still
showed up on the site.
`is_command_enabled` derived the section name and the legacy `enabled`
spellings by hand, duplicating `BaseCommand._derive_config_section_name`
and guessing that a legacy key lives in the command's own section. Half
of them do not: `[Jokes] joke_enabled` disables `Joke_Command`. It now
calls `command_section_name` and `read_enabled`, which share the alias
table with the runtime and the settings UI.
That table was missing the same-section `[Joke_Command] joke_enabled` and
`[DadJoke_Command] dadjoke_enabled` spellings, which both commands accept
through their own fallback. The settings page read only the `[Jokes]`
form, so a bot disabled the same-section way displayed as enabled.
Also fold the duplicated local-commands-dir resolution into
`resolve_local_commands_dir`, take `bot_root` as a `MinimalBot` argument
instead of assigning it post-construction twice, and restore the `hidden`
attribute check that `generate_samples` lost when it moved to
`filter_commands`.
---------
Signed-off-by: Gerard Hickey <hickey@kinetic-compute.com>
Co-authored-by: agessaman <adam@gessaman.com>
Adds a `[Hello_Command] include_sender` setting (default off) that names the user the hello reply is answering, so a busy channel can tell whose greeting the bot picked up.
The mention takes the place of the random human descriptor, keeping the translated sentence intact, and is dropped when it would push the reply past the channel body budget. DMs are unaffected, and channel sender names are sanitized before being echoed.
* feat(command): Add contact command
The contact command allows the user to ask the bot to provide its
public key as a clickable contact. This allows the user to then DM the
bot outside of a channel.
Signed-off-by: Gerard Hickey <hickey@kinetic-compute.com>
* fix(contact): harden matching and self_info handling, add tests and docs
Review follow-ups on the contact command:
- Match through _cleaned_content_matches so a non-matching message keeps its
original content. The raw cleanup_message_for_matching call rewrote
message.content in place during the keyword scan, which is the #267
regression every other command already avoids.
- Match against self.keywords so a configured alias works.
- Read self_info defensively (dict or object, may be absent) and validate the
public key instead of sending a literal "<None:1:None>" over the air.
- Drop the hardcoded '!' prefix strip in execute; matching already normalizes
the content, and prefixes are configurable.
- Forward skip_channel_check to the base can_execute.
- Remove copy-paste docstring leftovers from the roll command.
- Move [Contact_Command] into the alphabetical run in config.ini.example and
fix the "chanels" typo.
- Add tests, a docs/command-reference.md entry, and flesh out the changelog.
---------
Signed-off-by: Gerard Hickey <hickey@kinetic-compute.com>
Co-authored-by: agessaman <adam@gessaman.com>
The ACK fix I wrote for the radio-lock work is now upstream and released
(meshcore_py#108, v2.3.13). Before it, send_msg_with_retry reported ACKed DMs as
failures: the ACK subscription was registered only after send_msg returned, so
an ACK queued right behind MSG_SENT was dispatched with no listener, and each
attempt accepted only its own ACK code, so a late ACK answering an earlier
attempt was ignored. The bot logged "no ACK received after retries" and skipped
the delivery bookkeeping for messages the recipient had in fact received. Pin
2.3.14, which also fixes send_cmd's destination-type handling.
Two fixes from re-reading the merged radio-lock change:
A region-scoped channel message now restores global flood even when
set_flood_scope raises. The restore only ran in the finally around the send, so
a set that raised left the device pinned to that region and every later send
went out under it.
Drop the __setattr__ left behind when _SerializedCommands became
_serialize_command_frames; it sat after the function's return, unreachable, and
still referenced the proxy's _commands slot.
Radio commands were serialized per method call, so a composite command held the lock for everything it awaited. A DM's send_msg_with_retry kept every other command waiting through all of its ACK timeouts (up to ~36 s with three attempts), stalling channel replies, other DMs, scheduled sends, and automatic message fetching. req_regions_sync did the same for each neighbor scope reply.
Serialize at the frame instead: wrap CommandHandler.send() on the handler instance, which is where every meshcore command writes its frame and waits for the radio's immediate reply. There is still at most one in-flight companion frame, paced as before, and library methods that call self.send are covered without a proxy. Waits for ACKs and remote responses now run outside the lock.
Add MeshCoreBot.radio_session() for short sequences that must not interleave, and use it for region-scoped channel sends so set scope, send, and restore stay together. The per-call lock never guaranteed that: a DM retry loop queued between set_flood_scope and send_chan_msg sent its flood attempts under the channel's scope.
_channel_secrets() assigned dict_items and enumerate to the same
variable, and mypy inferred Any | None for the channel_idx .get() with
a mixed Any | int default. Declare items as Iterable[tuple[int, Any]]
and hold the raw index as Any. Runtime behavior is unchanged.
Also cover the list layout meshcore actually uses, where entries
without channel_idx fall back to their position.
generate_website.py can layer custom CSS on top of the built-in style chosen with --style. --embed-css appends a local file to the page's <style> block so the page stays a single file, and --link-css adds a stylesheet link after it, so both override built-in rules of equal specificity. The flags can be combined (the linked stylesheet loads last), apply to --sample pages, and an unreadable --embed-css file stops generation with an error. The website docs gain a custom CSS guide and a reference for the generated page's structure, classes, and theme variables.
Region scopes had to be hand-edited in config.ini, and the radio's own default
region could not be set from the bot at all. Both now live on the Radio page.
Node Settings gains a Default Region Scope field beside the path hash size,
which is the firmware's own setting — NodePrefs.default_scope_name and
default_scope_key, the region the radio falls back to for any send the bot does
not scope itself. Firmware that has no such command is reported as not having
answered rather than shown as an empty field, and a stored key that is not the
stored name's hash is flagged, because the radio routes by the key and the name
beside it is only a label. A build-flag default stores the bare name while
hashing the '#' form, so the comparison normalizes first.
Clearing it does not go through the meshcore library, because none of its reset
paths can: set_default_flood_scope(None) raises on len(None), "" earns
ILLEGAL_ARG from the firmware, and "*" only works because the frame is padded by
character count rather than by the name it wrote, landing one byte short of the
length the firmware reads as "a scope follows". The firmware's own contract for
clearing is a bare CMD_SET_DEFAULT_FLOOD_SCOPE, so that is what goes out. The
same padding bug is why a device scope name must be ASCII: a multi-byte
character displaces the transport key out of its field.
A separate Region Scopes card covers the bot's side — [Channels] flood_scopes as
an explicit choice between "reply whatever the scope" and an allowlist, and
outgoing_flood_scope_override. It writes config.ini, queues a hot reload, then
polls that reload and reports what the bot did with it, including saying plainly
when nothing picked it up. Per-channel flood_scope.<channel> entries are listed
read-only, since they are the reason a channel can ignore the default.
The two settings interact in a way worth stating on the page: once the bot sends
a scoped message it restores with set_flood_scope("*"), which leaves the radio in
forced-unscoped mode, so the device default stops applying until the bot scopes
another send.
Also from building it:
- flood_scopes left in [Bot] is surfaced rather than hidden. CommandManager
still honours it when [Channels] is empty, so a page that ignored it would
show "replies to every scope" while the bot enforced an allowlist — and
saving "off" hands the allowlist straight back to [Bot], which the banner now
says instead of claiming the old key stops mattering.
- outgoing_flood_scope_override = none is read as global flood on the send path,
as it already was everywhere else. send_channel_message tested the raw value
against a fixed tuple, so the lowercase spelling became the region "#none".
- Scope normalization moved to modules/flood_scope.py so the viewer, a separate
process, can validate a typed name without importing the bot's command
machinery. CommandManager delegates to it.
- Dark mode named only .border, so a .border-start divider drew Bootstrap's
light #dee2e6 onto a dark card. All four directional utilities now match.
The known-contact gate for DM delivery returns before anything is recorded, so
a bot that keeps no contacts dropped every warning and wrote no rows: the page
showed "No warnings decided yet" indefinitely beside a status card saying it
was sending, and nothing distinguished a clean mesh from one where every
warning was being discarded. Only a debug log said otherwise.
Withheld warnings are now counted per local day in bot_metadata — a counter
rather than event rows, because this fires ahead of the cooldowns that would
rate-limit rows, so logging each one would bury the decisions worth reading.
The status card shows it, and the empty log explains itself and points at
channel delivery.
Also from the verification pass:
- The percent test built its value with config.set and a doubled %, which
set() is the one input that cannot produce the bug. It now uses read_string
with a bare %, and I checked it fails against raw=False by substituting
DEFAULT_MESSAGE — the actual failure mode.
- docs still said the bot identifies itself by public key, contradicting the
section forty lines above. On this path it is name-only, and that means the
self-exemption is spoofable; the doc says so.
- delivered_today folded dry runs in with real sends, so a morning's preview
read as afternoon transmissions. previewed_today is now separate and the
figure follows the current mode.
- TRIM(MIN(channel)) so a whitespace-padded historical row cannot become the
display name for a merged channel.
**A `global` verdict now requires RF correlated to the message.** It was also
reachable through `_is_confirmed_global_flood`'s argument-from-absence route
("no scope-eligible packet in the window, therefore unscoped"), so a row that
literally said TC_FLOOD, or a scoped ADVERT that the GRP_TXT filter excluded,
came back as "no region code". That inference is fine for deciding whether a
`*` in flood_scopes authorizes a reply; it is not fine for accusing someone of
a misconfiguration. Channel messages still correlate through the payload match
(#255), so the ordinary case is unaffected.
**A `%` in the warning message no longer wedges config reload.** `_get` read
without `raw=True`, so configparser's interpolation raised, the error was
swallowed, and the bot transmitted the default wording instead. Worse, the
bot's own `_validate_config_snapshot` iterates `config.items(section)` and
would reject every hot reload until the file was hand-edited. The save endpoint
now rejects `%` outright and caps the message at 500 characters, and the read
is raw so a hand-edited value is still honored.
**One unreachable sender no longer eats the daily cap forever.** The cap counted
failed attempts but the per-sender cooldown did not, so a node the radio cannot
reach was retried every `min_unscoped_messages` messages indefinitely and no
real offender was ever warned. Both count attempts now. The regression test
fails with three sends against the old filter.
**The sender is a display name, not an identity.** MeshCore's CHANNEL_MSG_RECV
carries no public key, so `sender_pubkey` was always empty on this path and the
only identity is a prefix anyone with the channel key can forge. DM delivery
now requires a contact the radio already holds, which bounds the bot to nodes
it knows and stops failed sends spending cap slots. Two tests asserted the
opposite because the fixture supplied a pubkey the radio never sends; they now
run with what the call site actually passes, and the docs no longer claim
pubkey identity.
Also: a negative `max_warnings_per_day` fell back instead of clamping to the
"unlimited" sentinel; tallies group case-insensitively so a `#` or case change
does not split a channel; retention uses the same clock the rows are written
in; the status figure shows delivered with attempts beside it, rather than a
count of four next to "last warning: none yet"; and the channel bars are scaled
over classified traffic so their width equals the percentage printed beside
them (33.7% was drawn at 28.7%).
An empty daily-volume chart rendered as an 84px hole with a date range under
it, which reads as broken rather than as no data; it is hidden until there is
something to plot, and the summary says so in words.
"Count only" disabled the rest of the form, which implied those values would
not be saved — they are — and stopped anyone drafting their wording before
turning warnings on. A sentence under the mode selector says it instead.
Bootstrap's .text-warning is #ffc107, which is 1.6:1 on a white card — the
headline percentage and the "daily cap reached" line were close to invisible
in light mode. The chart fills had the same shape of problem in reverse: the
grey "couldn't tell" band sat at 2.1:1 against a white card and 2.9:1 against
a dark one.
Each color is now a token chosen per theme against the surface it lands on:
text clears 4.5:1, fills clear 3:1. Measured in the browser in both themes.
An unselected chip looked identical to a selected one, so clicking it was a
coin flip between adding and removing. The chips now reflect the field, in
both directions, and stay in step when the list is typed by hand.
Also re-read after a save rather than patching the form in place (the status
strip's budget is derived from the cap that was just changed), and stop the
reset-to-default button from blanking the message when the page never loaded.
A channel message with no "Name: " prefix falls back to a stand-in name.
That is not a node: every such message shares the identity, so they would
accumulate a run together and then get a DM addressed to a contact that
does not exist. The hook now passes no sender for those, so they are still
counted but can never earn a warning. The literal is a named constant and a
test pins the hook's guard to it.
Also folds the channel body-budget formula into models.channel_body_limit
rather than letting the web viewer keep a fourth copy of
"max(130, 160 - len(name) - 2)"; BaseCommand and CommandManager now call it
too, and 158 becomes models.DM_BODY_LIMIT.
Recording the decision after the send left an await between the cap check
and the row that check reads. Two channel messages in flight both passed a
cap of one and both transmitted; the new test fails with two sends against
the old ordering.
Every gate and the reservation now run with no await between them, so on the
single event loop they are atomic. A failed send corrects its own row to
'failed' and releases the mesh cooldown, but the row stays and still counts
against the daily cap: a send that reported failure may have put something on
the air before it did.
Also from a self-review pass: bound the per-sender run table (an unbounded
dict on the channel-message path is a slow leak), identify the bot's own node
by public key rather than only by the configured name, and respect
channelpause.
Closes#279.
A MeshCore client with no region configured sends every channel message as
a plain FLOOD, which every repeater on the mesh rebroadcasts. This adds two
things: free observation of how much of that the bot hears, and an opt-in
warning to the senders.
Observation classifies each channel message as scoped, global or unknown and
tallies it per channel per local day. It costs one upsert and no airtime, and
it is what lets an operator see the size of the problem before deciding to
spend airtime on it. The classification runs ahead of the flood_scopes
allowlist, because an unscoped message is exactly what that allowlist drops.
Warnings only ever fire on positive RF evidence of an unscoped FLOOD. Absence
of correlation is not proof that a sender omitted a region, so it classifies
as unknown and stays quiet. They are off by default, start in dry run, and are
fenced by min_unscoped_messages, a per-sender cooldown, a mesh-wide cooldown
and a daily cap on attempts. Dry run consumes the same budget it previews, so
the log is what going live would put on the air, not an upper bound. All three
limits read from the database, so a restart cannot release a burst.
Channel-delivered warnings go out at global scope on purpose: the recipient is
by definition outside any region the bot replies under.
New Settings -> Region Warnings page in the web viewer covers all of it. Its
dark-mode striped rows exposed a pre-existing base.html bug where Bootstrap's
light-theme text color survived on a dark row background (~1.3:1); fixed there
for every table in the app.
Discover local commands and services in the Plugins UI and route their settings to the local config overlay. Upgrade GitHub Actions to Node 24-capable majors and migrate frontend linting to ESLint 10 flat config. Keep command and service discovery namespaces isolated and align duplicate-name handling with runtime loading.
Add a password type to settings_schema so plugin secrets render as masked inputs and never reach the browser. Stored values stay in config.ini; an empty box on save leaves an existing secret unchanged.
[Test_Command] distance_unit (auto/km/mi) controls {path_distance} and {firstlast_distance}. Auto follows the reply language (miles for en/en-US, kilometres otherwise, including en-GB). Elapsed times of 1s and above render as seconds to keep the ack short.
The published packet payload mixed clocks: "timestamp" was UTC ISO 8601
with a Z suffix, but the "time" and "date" fields beside it were rendered
from a naive datetime.now(), i.e. the host's local wall clock. A consumer
reading those two off a bot in a non-UTC zone sees a skew of exactly that
zone's UTC offset and flags the observer's clock as wrong.
Local time was never the intended reading. The original script took time
and date off the firmware log line, and mctomqtt sets the device clock
from calendar.timegm(), so upstream both fields are already UTC.
All three now render one aware UTC instant, so they cannot disagree with
each other or straddle a second boundary. _utc_iso_timestamp() takes an
optional instant to make that possible; its no-arg behavior is unchanged.
`test_mesh_refresh_coordinator_handles_bursts_visibility_and_failures` sampled
the coordinator after fixed `setTimeout` waits. The `scheduledRetry` phase is the
tightest: a 10ms scheduled refresh, a failing load, a 10ms retry and a 2ms load
is ~22ms of nominal work against a 35ms budget, so barely 13ms of slack. On a
contended runner a 10ms timer fires well late, the wait then samples a
half-finished retry, and the assertion reports a timing artifact as a logic
failure. It flaked twice in six PRs, once on a docs-only branch that cannot
affect this code.
Each of those trailing waits now polls until the coordinator is genuinely at
rest -- nothing pending, running, timed or in flight -- with a 2s ceiling that
names the phase that failed to settle. Two consecutive idle observations are
required so the synchronous gap between a load finishing and its retry being
armed is not mistaken for rest. The `hiddenLoads` wait is left as a sleep: it
asserts that nothing happens, which polling cannot establish, and extra time
there can only ever make it stricter.
Verified by shimming `setTimeout` to fire every timer late, which is what a
loaded scheduler does without altering the logic under test. The old test starts
failing at 20ms of overshoot -- on the same `scheduledRetry` assertion CI hit --
and this one still passes at 200ms.
Also confirmed the rewrite keeps the test's teeth: across seven coordinator
mutations, detection is identical before and after. The three neither version
catches are equivalent mutants, since the two hidden-tab guards mask each other;
removing both is caught by both versions.
Raised the node timeout from 5s to 60s so a real stall surfaces as the phase
name rather than an opaque subprocess timeout. Runtime is unchanged in practice
(~0.36s), since the waits now end as soon as the work is done.
The local translation path defaulted to a bare `local/translations/`, resolved
against the process cwd and unrelated to `[Bot] local_dir_path`. Since
`local_dir_path` already selects where an operator's commands, service plugins
and config overlay live, the catalog belongs in that same tree. It now defaults
to `<local_dir_path>/translations`, resolved absolute against the bot root, so
relocating `local_dir_path` moves the catalog with it and the lookup no longer
depends on the working directory. An explicit `local_translation_path` still
wins.
Only the constructor at startup was passing the local path. `get_translator()`
builds and caches a translator per detected language, and `reload_config()`
builds a fresh one, and both were still constructing `Translator` with the
distributed path alone. So local overrides silently vanished from any reply that
used the sender's language, and did not survive a config reload. Both now pass
it, `reload_config()` republishes it alongside `translation_path`, and it is
snapshotted and restored on rollback. The DummyTranslator fallback sets it too,
since `get_translator()` would otherwise raise AttributeError there.
Collapsed `_deep_merge_translations` into the existing `_merge_translations`,
which already merged deeply with the primary winning. Rewrote that one to stop
mutating its input: it shallow-copied `fallback` and then recursed into the
shared sub-dicts, so merging a base language over English corrupted the cached
English catalog at every level below the first. Also replaced the two `print()`
calls on the error paths with logger calls.
Registered `local_translation_path` in the config schema next to
`translation_path`, and reverted a trailing-whitespace reflow that touched 84
lines of config.ini.example for a one-key addition.
Tests: local overlay (single-string override, added keys, local-only catalog,
missing directory), the merge no longer mutating its fallback, the default
following `local_dir_path`, an explicit setting overriding it, the overlay
reaching `bot.translator` end to end, and a reload picking up a moved path.
test_neighbor_evidence_edges.py pinned RECENT to 2026-08-01 and used it as the
default last_seen for the seeded links. The two days=30 tests compare that against
wall-clock now, so once 2026-08-31 passed, the "recent" link filtered out as stale
and both assertions saw an empty set. RECENT, EARLIER and ANCIENT now derive from
datetime.now(timezone.utc), which fixes their relationship to the window no matter
when CI runs. Nothing else in the file depended on the literal values—the other
assertions compare against the same constants, and the filter only ever reads
last_seen, never first_seen.
Swept the rest of tests/ for the same shape by running the suite under a clock
shifted forward 400 days, then 10 years. That turned up one more:
test_packet_capture_neighbors.py stored 1774482900.0 (2026-03-25) as the persisted
neighbors timestamp, and _load_neighbors_timestamp rejects anything older than
now-400d, so those two round-trip tests would have started failing on 2027-04-29.
That value and the far-future one (2**31, i.e. 2038) are both relative now.
The failures that remain under a shifted clock are all fixtures seeding through
SQLite's datetime('now') or time.strftime() while the code under test reads
Python's clock. Those two move together on a real runner, so they are artifacts of
the sweep rather than time bombs.
Addresses the review on #243. The original change interpolated the last
traceback frame into an `err_desc` string used in two places, and the second
one was the problem: `errors.execution_error` is sent back over the mesh, so
an absolute install path went out over RF and spent airtime on the error
path, where retries are most likely.
logger.exception() puts the full traceback in the log instead of just one
frame, which is strictly more than the original wanted, and the RF reply goes
back to str(e). That also removes the `list(tb.stack)[-1]` IndexError on an
empty stack, since the frame handling is gone entirely.
Tests cover both halves: the failure is logged via exception(), and the reply
carries only the exception text with no filesystem path.
CHANGELOG entry moves under the existing `## [Unreleased]` / `### Fixed`.