Files
meshcore-analyzer/.github/workflows/deploy.yml
T
efitenandClaude Opus 5.5 de364a8bb6 feat: mail notifications for watched nodes (user management part E) (#2140)
Part E of #2128: a logged-in user watches nodes and gets one mail when a
watched node goes offline, comes back, or reports a low battery. Admins
can add instance events: a new foreign node, and an observer going
offline or back. This covers the per-user mail part of #775.

Stacked on #2139 (part D, which builds on #2138). Review the commits
after that branch.

## The situation

A sysop learns that a repeater went silent, or that its battery is
running down, only by opening CoreScope. #775 asks for notifications.
Accounts (part A) now give a verified address to mail and a mailer with
delivery status.

## What this PR adds

**Events.** They use the instance thresholds, so a mail does not
disagree with the node page:
- `node.offline`: silent for the role's `healthThresholds` window. Last
heard comes from the packet store, and repeaters and rooms count relayed
traffic (#1598).
- `node.battery`: advert telemetry below `batteryThresholds.lowMv`,
recovered at `lowMv + 100`.
- Admins only: `foreign.new` (once per node) and `observer.offline`.

**Store.** Schema v5 with `notification_prefs`, `notification_watches`
and `notification_state`.

**Server** (only with `userManagement.notifications.enabled`)
- A notifier next to the janitor, every `intervalMinutes` (5). The
evaluator is a pure function of watches, prefs, stored states and node
snapshots.
- No burst after a restart or a new watch: a subject's first evaluation
stores its state without mailing. The loop waits for the startup load.
While ingest is stale (newest packet older than 30 minutes) the offline
checks pause, and after recovery they wait one silent window.
- One mail per user per check. Limits: 20 per user and 100 per instance
per rolling 24 hours, 50 watches per user. A change over a limit is
recorded and never mailed later.
- Every mail has a one-click unsubscribe (`List-Unsubscribe` and
`List-Unsubscribe-Post`). The GET only redirects to a confirm page, so
link scanners cannot unsubscribe anyone. Node names are cleaned of
control and bidi characters before they go into a mail.
- Routes: `GET`/`PUT /api/account/notifications`, `PUT`/`DELETE
/api/account/notifications/watches/{pubkey}`, `POST
/api/account/notifications/watch-my-nodes` (copies the synced "my nodes"
list), `GET`/`POST /api/notifications/unsubscribe`. `mailer.Message`
gains `Headers`.

**Frontend.** A "Notify me" toggle on the node pages, a Notifications
section on the account page (`#/account?section=notifications`), an
unsubscribe page, and the notification mail count on the admin overview.

Spec:
[`docs/specs/2026-10-07-node-notifications-design.md`](https://github.com/efiten/CoreScope/blob/feat/notifications/docs/specs/2026-10-07-node-notifications-design.md).

## Performance

Each check reads watches, prefs and states from `users.db` (one query
each), the watched nodes from the analyzer DB in chunks of 500, and
last-heard times from the packet store.

| Measurement | Result |
|---|---|
| One check on our staging instance, 130,000 packets in memory, 1
watcher | 6.9 ms, of which 0.93 ms holding the store read lock |
| Evaluator benchmark, 100 users with 50 watches each, 2,000 nodes |
1.75 ms |

No work is added to ingest, broadcast or any request path.

## Verification

- `internal/users` and `internal/mailer` with `-race`; the `cmd/server`
suite plus `-race` on the notifier tests; 74 new Go tests. The same
Windows-only failure as noted in #2138 applies.
- `sh test-all.sh` exits 0; the XSS gate in diff mode passes.
- User-management E2E: 26 of 26 steps locally. With notifications off
the new steps fail, so they do test the flag.
- On our staging and production instance since 7 October 2026.

## Not in this PR

- Other channels (Discord, Telegram, webhooks). Detection is separate
from delivery, so one can be added.
- The topology, RF and anomaly alerts of #775.
- If the whole server starts after a feed outage that already ended, the
grace window is unknown and a watcher can get one wrong offline mail.
The user guide says so.

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 22:02:39 +02:00

1281 lines
63 KiB
YAML
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
name: CI/CD Pipeline
# Documentation-only changes still trigger this workflow, on purpose. Skipping
# it at the trigger level leaves every required check in a pending state that
# can never arrive, and the pull request can then never merge. GitHub documents
# exactly this: a workflow skipped by path filtering keeps its checks pending
# and blocks the merge, while a JOB skipped by an if: conditional reports
# Success and does not. So the filtering lives on the jobs below, not here.
#
# The scope check only applies to pull requests. A push to master always runs
# the full pipeline, which keeps :edge built for every master commit. That
# matters: release-fast-path.yml re-tags :edge to :vX.Y.Z only when the :edge
# revision label matches the tagged commit, so a master commit without an image
# breaks tagging.
on:
push:
branches: [master]
pull_request:
branches: [master]
workflow_dispatch:
inputs:
images_published:
description: 'Release fast path already published the tag images'
type: boolean
default: false
permissions:
contents: read
packages: write
concurrency:
group: ci-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: ${{ github.event_name == 'pull_request' }}
env:
FORCE_JAVASCRIPT_ACTIONS_TO_NODE24: true
STAGING_COMPOSE_FILE: docker-compose.staging.yml
STAGING_SERVICE: staging-go
STAGING_CONTAINER: corescope-staging-go
# Pipeline. The test jobs and the image build run side by side; only the GHCR
# push waits for all of them:
# changes ─┬─ go-test ─────────────────────────┐
# ├─ race-test (ingestor changes only) ─┤
# ├─ e2e-shard ×3 ── e2e-test (gate) ───┼─ build-and-publish → deploy, publish
# └─ image-check (two-arch build) ──────┘
# PRs stop after build-and-publish (no GHCR push). Master continues to deploy + badges.
jobs:
# ───────────────────────────────────────────────────────────────
# 0. Change scope — decides whether the expensive jobs below run.
# ───────────────────────────────────────────────────────────────
changes:
name: "🔎 Change scope"
runs-on: ubuntu-latest
outputs:
code: ${{ steps.scope.outputs.code }}
ingestor: ${{ steps.scope.outputs.ingestor }}
steps:
- name: Checkout code
uses: actions/checkout@v5
with:
fetch-depth: 0
- name: Decide whether anything but documentation changed
id: scope
run: |
set -euo pipefail
# Only pull requests are scoped. Pushes and dispatches always count as
# code, so master keeps producing an :edge image and tag builds are
# never skipped.
if [ "${{ github.event_name }}" != "pull_request" ]; then
echo "code=true" >> "$GITHUB_OUTPUT"
echo "ingestor=true" >> "$GITHUB_OUTPUT"
echo "event=${{ github.event_name }} -> full pipeline"
exit 0
fi
CHANGED=$(git diff --name-only "${{ github.event.pull_request.base.sha }}" "${{ github.event.pull_request.head.sha }}")
# Empty diff means something is off; run everything rather than guess.
if [ -z "$CHANGED" ]; then
echo "code=true" >> "$GITHUB_OUTPUT"
echo "ingestor=true" >> "$GITHUB_OUTPUT"
echo "empty diff -> full pipeline"
exit 0
fi
echo "Changed files:"
echo "$CHANGED" | sed 's/^/ /'
if echo "$CHANGED" | grep -qvE '(^docs/|[.]md$|^LICENSE$)'; then
echo "code=true" >> "$GITHUB_OUTPUT"
echo "-> non-documentation files changed, full pipeline"
else
echo "code=false" >> "$GITHUB_OUTPUT"
echo "-> documentation only, heavy jobs skip (and report Success)"
fi
# The race detector is its own job, and it only earns its ten minutes
# when the code it inspects changed. A frontend or docs PR cannot
# introduce a data race in the ingestor.
if echo "$CHANGED" | grep -qE '^cmd/ingestor/.*[.]go$'; then
echo "ingestor=true" >> "$GITHUB_OUTPUT"
echo "-> ingestor Go files changed, race detector runs"
else
echo "ingestor=false" >> "$GITHUB_OUTPUT"
echo "-> no ingestor Go changes, race detector skips"
fi
# ───────────────────────────────────────────────────────────────
# 1a. Race detector (ingestor) — its own job on purpose
# ───────────────────────────────────────────────────────────────
# Runs beside "Go Build & Test" rather than inside it, so its ten-odd
# minutes overlap the E2E job instead of extending the critical path, and
# only when ingestor Go files changed. A frontend or documentation PR cannot
# introduce a data race here, and paying for the check on every PR is what
# gets a check switched off again.
#
# The server has -race in its own step already; this closes the same gap for
# the ingestor, which carries an atomic.Pointer snapshot (the region key
# set) whose safety is an argument until something checks it.
race-test:
name: "🏁 Race detector (ingestor)"
runs-on: ubuntu-latest
needs: [changes]
if: needs.changes.outputs.ingestor == 'true'
steps:
- name: Checkout code
uses: actions/checkout@v5
- name: Set up Go 1.27
uses: actions/setup-go@v6
with:
go-version: '1.27'
# All modules: a -race build of a cgo package is the slowest thing in
# this pipeline, and setup-go caches ~/.cache/go-build as well as the
# module cache, so a key that misses on an unrelated module hurts most
# here. Same reasoning as go-test below.
cache-dependency-path: "**/go.sum"
- name: go test -race
run: |
set -e -o pipefail
cd cmd/ingestor
go test -timeout 30m -race ./...
# ───────────────────────────────────────────────────────────────
# 1. Go Build & Test
# ───────────────────────────────────────────────────────────────
go-test:
name: "✅ Go Build & Test"
runs-on: ubuntu-latest
needs: [changes]
if: needs.changes.outputs.code == 'true'
steps:
- name: Checkout code
uses: actions/checkout@v5
with:
fetch-depth: 0
- name: Clean Go module cache
run: rm -rf ~/go/pkg/mod 2>/dev/null || true
- name: Set up Go 1.27
uses: actions/setup-go@v6
with:
go-version: '1.27'
# All modules, not just two: the cgo SQLite build is expensive to
# redo, and setup-go caches ~/.cache/go-build as well as the module
# cache, so a key that misses on an unrelated module hurts.
cache-dependency-path: "**/go.sum"
- name: Enforce gofmt + go vet (issue #1859)
run: |
set -e
# gofmt: every tracked *.go file except the misnamed Dockerfile.go,
# which is a Dockerfile (FROM ...) that gofmt cannot parse.
unformatted=$(gofmt -l $(git ls-files '*.go' | grep -vi 'Dockerfile.go'))
if [ -n "$unformatted" ]; then
echo "::error::gofmt required on these files:"
echo "$unformatted"
exit 1
fi
echo "gofmt: clean"
# go vet: multi-module repo, so vet each module in turn.
for mod in $(git ls-files '*go.mod' | sed 's#/go.mod##'); do
echo "== go vet $mod =="
( cd "$mod" && go vet ./... )
done
- name: Build and test Go server (with coverage)
run: |
set -e -o pipefail
cd cmd/server
go build .
# -race gates PR #1208's atomic.Pointer migration: the race-detector
# is what makes path_inspect_atomic_race_test.go actually assert.
go test -timeout 20m -race -coverprofile=server-coverage.out ./... 2>&1 | tee server-test.log
# The e2etest build (fake mailer + /__e2e/last-mail) has its own tests.
go test -tags e2etest -count=1 -run 'E2E' .
echo "--- Go Server Coverage ---"
go tool cover -func=server-coverage.out | tail -1
- name: Build and test Go ingestor (with coverage)
run: |
set -e -o pipefail
cd cmd/ingestor
go build .
go test -timeout 20m -coverprofile=ingestor-coverage.out ./... 2>&1 | tee ingestor-test.log
echo "--- Go Ingestor Coverage ---"
go tool cover -func=ingestor-coverage.out | tail -1
- name: Build and test channel library + decrypt CLI
run: |
set -e -o pipefail
cd internal/channel
go test ./...
echo "--- Channel library tests passed ---"
cd ../../cmd/decrypt
go build -ldflags="-s -w" -o corescope-decrypt .
go test ./...
echo "--- Decrypt CLI tests passed ---"
- name: Test user-management modules (users, mailer)
run: |
set -e -o pipefail
cd internal/users && go test -race ./...
cd ../mailer && go test -race ./...
echo "--- users + mailer tests passed ---"
- name: Test migrate CLI + dbschema
run: |
set -e -o pipefail
# Neither module had CI test execution before the mattn/go-sqlite3
# migration; both open the database, and dbschema owns the migrations.
cd internal/dbschema
go test ./...
cd ../../cmd/migrate
go build -o /dev/null .
go test ./...
echo "--- migrate + dbschema tests passed ---"
- name: Verify Dockerfile COPY invariants (issue #1316)
run: bash scripts/check-dockerfile-internal-pkgs.sh
- name: Staging disk-monitor unit tests (issue #1684)
run: bash scripts/staging/test-disk-monitor.sh
- name: QA SQL parameter-binding unit tests (issue #1977)
run: bash qa/scripts/test-blacklist-sql.sh
- name: Lint CSS variables (issue #1128)
run: |
set -e
node scripts/check-css-vars.js
node scripts/test-check-css-vars.js
- name: Run JS unit tests
run: sh test-all.sh
- name: 🛡️ Preflight XSS gate — actual --diff check (PR only)
# The fixture self-test above (test-preflight-xss-gate.js) only
# asserts the script's behavior against fixtures. It does NOT scan
# the PR's own changes. This step closes that gap by running the
# gate against added lines in public/**/*.{js,html} on the PR.
# Gate is PR-scoped only (per djb finding: merge commits would
# slip an opt-out otherwise). Master pushes skip this step.
if: github.event_name == 'pull_request'
env:
PR_BODY: ${{ github.event.pull_request.body }}
PREFLIGHT_PR_LABELS: ${{ join(github.event.pull_request.labels.*.name, ' ') }}
run: |
set -e
git fetch origin master --depth=50 2>&1 | tail -3 || true
# Materialize PR body to a file for the opt-out parser.
printf '%s' "$PR_BODY" > /tmp/pr-body.md
PREFLIGHT_PR_BODY=/tmp/pr-body.md bash scripts/check-xss-sinks.sh --diff origin/master
- name: 🧹 Frontend lint (eslint no-undef) — issue #1342
run: |
set -e
# Use eslint@8 (legacy .eslintrc.json). Don't migrate to flat-config / eslint@9.
# --no-save: avoid touching package.json / no committed node_modules.
npm install --no-save --no-audit --no-fund eslint@8
npx eslint public/*.js
- name: Verify proto syntax
run: |
set -e
sudo apt-get update -qq
sudo apt-get install -y protobuf-compiler
for proto in proto/*.proto; do
echo " ✓ $(basename "$proto")"
protoc --proto_path=proto --descriptor_set_out=/dev/null "$proto"
done
echo "✅ All .proto files are syntactically valid"
- name: Generate Go coverage badges
if: success()
run: |
mkdir -p .badges
SERVER_COV="0"
if [ -f cmd/server/server-coverage.out ]; then
SERVER_COV=$(cd cmd/server && go tool cover -func=server-coverage.out | tail -1 | grep -oP '[\d.]+(?=%)')
fi
SERVER_COLOR="red"
if [ "$(echo "$SERVER_COV >= 80" | bc -l 2>/dev/null)" = "1" ]; then SERVER_COLOR="green"
elif [ "$(echo "$SERVER_COV >= 60" | bc -l 2>/dev/null)" = "1" ]; then SERVER_COLOR="yellow"; fi
echo "{\"schemaVersion\":1,\"label\":\"go server coverage\",\"message\":\"${SERVER_COV}%\",\"color\":\"${SERVER_COLOR}\"}" > .badges/go-server-coverage.json
INGESTOR_COV="0"
if [ -f cmd/ingestor/ingestor-coverage.out ]; then
INGESTOR_COV=$(cd cmd/ingestor && go tool cover -func=ingestor-coverage.out | tail -1 | grep -oP '[\d.]+(?=%)')
fi
INGESTOR_COLOR="red"
if [ "$(echo "$INGESTOR_COV >= 80" | bc -l 2>/dev/null)" = "1" ]; then INGESTOR_COLOR="green"
elif [ "$(echo "$INGESTOR_COV >= 60" | bc -l 2>/dev/null)" = "1" ]; then INGESTOR_COLOR="yellow"; fi
echo "{\"schemaVersion\":1,\"label\":\"go ingestor coverage\",\"message\":\"${INGESTOR_COV}%\",\"color\":\"${INGESTOR_COLOR}\"}" > .badges/go-ingestor-coverage.json
echo "## Go Coverage" >> $GITHUB_STEP_SUMMARY
echo "| Module | Coverage |" >> $GITHUB_STEP_SUMMARY
echo "|--------|----------|" >> $GITHUB_STEP_SUMMARY
echo "| Server | ${SERVER_COV}% |" >> $GITHUB_STEP_SUMMARY
echo "| Ingestor | ${INGESTOR_COV}% |" >> $GITHUB_STEP_SUMMARY
- name: Upload Go coverage badges
if: success()
uses: actions/upload-artifact@v6
with:
name: go-badges
path: .badges/go-*.json
retention-days: 1
if-no-files-found: ignore
include-hidden-files: true
# ───────────────────────────────────────────────────────────────
# 2. Playwright E2E Tests (against Go server with fixture DB)
# ───────────────────────────────────────────────────────────────
# Three shards on three runners, each with its own server and fixture copy,
# so no shard shares a browser, a CPU or server state with another. One
# runner per shard rather than parallel suites on one runner: the gesture and
# animation suites are timing-sensitive and would flake under CPU contention.
#
# It needs only `changes`, not go-test: this job builds its own binaries and
# uses nothing go-test produces. build-and-publish still needs both, so
# nothing reaches GHCR unless every test job passed.
e2e-shard:
name: "🎭 Playwright E2E shard ${{ matrix.shard }}/3"
needs: [changes]
runs-on: ubuntu-latest
strategy:
matrix:
shard: [1, 2, 3]
env:
SHARD: ${{ matrix.shard }}
E2E_SHARDS: 3
defaults:
run:
shell: bash
if: needs.changes.outputs.code == 'true' && !(startsWith(github.ref, 'refs/tags/v') && inputs.images_published)
steps:
- name: Checkout code
uses: actions/checkout@v5
with:
fetch-depth: 0
- name: Set up Node.js 22
uses: actions/setup-node@v5
with:
node-version: '22'
- name: Clean Go module cache
run: rm -rf ~/go/pkg/mod 2>/dev/null || true
- name: Set up Go 1.27
uses: actions/setup-go@v6
with:
go-version: '1.27'
cache-dependency-path: "**/go.sum"
- name: Build Go server
run: |
cd cmd/server
go build -o ../../corescope-server .
echo "Go server built successfully"
# Only test-user-management-e2e.js uses the e2etest build and its server,
# and it runs on shard 3. Moving that suite means moving these two too;
# it fails loudly (nothing on port 13582) rather than skipping.
- name: Build Go server (e2etest tag, user management)
if: matrix.shard == 3
run: |
cd cmd/server
go build -tags e2etest -o ../../corescope-server-e2e .
- name: Build Go migrate tool
run: |
cd cmd/migrate
go build -o ../../corescope-migrate .
echo "Go migrate tool built successfully"
- name: Install npm dependencies
run: npm ci --production=false
- name: Install Playwright browser
run: |
npx playwright install chromium 2>/dev/null || true
npx playwright install-deps chromium 2>/dev/null || true
- name: Instrument frontend JS for coverage
run: sh scripts/instrument-frontend.sh
- name: Freshen fixture timestamps
run: bash tools/freshen-fixture.sh test-fixtures/e2e-fixture.db
- name: Seed grouped-packet row for #1486 collapse test
# The committed fixture has 499 packets, each with exactly ONE
# observation, so the packets-page renders only flat
# (select-hash) rows. The #1486 repro needs at least one grouped
# (toggle-select) row. Insert a NEW transmission with 3
# observations.
#
# The server's async hash-migrate (cmd/server/hash_migrate.go)
# recomputes `transmissions.hash` from `raw_hex` via
# ComputeContentHash(), so the inserted hash MUST equal that
# function's output for the chosen raw_hex — otherwise the row
# gets relabelled and the E2E can't find it.
#
# raw_hex 15000102030405060708090a0b0c0d0e0f
# → header=0x15 (route_type=1, payload_type=5)
# → ComputeContentHash(...) = fae0c9e6d357a814
#
# Fixed old timestamps keep the row behind the freshly shifted
# fixture observations and outside the default 15-min UI window.
# The #1486 and #2104 tests reach it via an explicit hash +
# ?timeWindow=0 deep-link.
run: |
sqlite3 test-fixtures/e2e-fixture.db <<'SQL'
-- Sort the seeded row LAST in BOTH default packets views:
-- • flat view sorts by transmissions.id DESC → id=0 puts it last
-- • grouped view (#default for the packets page) sorts by
-- MAX(observations.timestamp) DESC → we must keep our obs
-- timestamps OLDER than every other fixture observation.
-- freshen-fixture.sh shifts fixture observations forward.
-- Keep this row at 2026-05-15 so it sorts LAST in the grouped
-- view too; the #1486 and #2104 tests use the explicit deep-link.
INSERT INTO transmissions(id,raw_hex,hash,first_seen,route_type,payload_type,payload_version,decoded_json,channel_hash,from_pubkey)
VALUES (0,'15000102030405060708090a0b0c0d0e0f','fae0c9e6d357a814','2026-05-15T00:00:00Z',1,5,0,'{"type":"CHAN","channel":"#test","text":"#1486 fixture"}',NULL,NULL);
INSERT INTO observations(transmission_id,observer_idx,direction,snr,rssi,score,path_json,timestamp,resolved_path) VALUES
(0,1,'rx',5.0,-95,0,'["AA"]',CAST(strftime('%s','2026-05-15T00:00:00Z') AS INTEGER),'["aa00000000000000000000000000000000000000000000000000000000000000"]'),
(0,2,'rx',5.5,-92,0,'["BB"]',CAST(strftime('%s','2026-05-15T00:00:00Z') AS INTEGER),'["bb00000000000000000000000000000000000000000000000000000000000000"]'),
(0,3,'rx',6.0,-90,0,'["CC"]',CAST(strftime('%s','2026-05-15T00:00:00Z') AS INTEGER),'["cc00000000000000000000000000000000000000000000000000000000000000"]');
-- #1791 fixture: a single GRP_DATA (payload_type=6) packet so the
-- E2E "Group Data filter" test has at least one row to filter on.
-- Use an obs timestamp within the default UI window so the row
-- appears with no time-window override.
--
-- raw_hex header byte 0x19 = bits 5-2 (payload)=0110=6 (GRP_DATA),
-- bits 1-0 (route)=01=1 (FLOOD).
-- path_len byte 0x00 = hash_size=1, hash_count=0 (zero-hop on-wire,
-- typical GRP_DATA going FLOOD). path_json/resolved_path are kept
-- EMPTY so the rendered hop-row count matches the hex-path byte
-- count (a prior fixture used path_json=["AA"] but raw_hex
-- path_len=0, which broke the "hex strip Path range matches hop
-- row count" E2E).
--
-- Note: id=-1000000 is a deliberately out-of-band sentinel id so
-- this synthetic fixture row cannot collide with real ingested
-- transmissions (real ids are positive autoincrement values).
INSERT INTO transmissions(id,raw_hex,hash,first_seen,route_type,payload_type,payload_version,decoded_json,channel_hash,from_pubkey)
VALUES (-1000000,'19000102030405060708090a0b0c0d0e0f','17910000deadbeef',strftime('%Y-%m-%dT%H:%M:%SZ','now'),1,6,0,'{"type":"GRP_DATA","channel":"#test","raw":"deadbeef"}',NULL,NULL);
INSERT INTO observations(transmission_id,observer_idx,direction,snr,rssi,score,path_json,timestamp,resolved_path) VALUES
(-1000000,1,'rx',7.0,-88,0,'[]',CAST(strftime('%s','now') AS INTEGER),'[]');
SQL
- name: Migrate fixture DB to current schema (#1287)
# Server now ASSERTs schema is migrated and refuses to start
# otherwise (cmd/server/main.go: dbschema.AssertReady). In prod
# the ingestor owns dbschema.Apply, but CI starts only the
# server against the committed e2e fixture — so we run the
# standalone migrate tool here to bring the fixture up to the
# required shape before the server boots.
run: ./corescope-migrate -db test-fixtures/e2e-fixture.db
- name: 'Seed distinct observation bytes for #2104'
# raw_hex exists only after migration. Keep the #1486 payload/hash
# while encoding each observation's existing one-hop path.
run: |
sqlite3 test-fixtures/e2e-fixture.db <<'SQL'
UPDATE observations SET raw_hex = CASE observer_idx
WHEN 1 THEN '1501aa0102030405060708090a0b0c0d0e0f'
WHEN 2 THEN '1501bb0102030405060708090a0b0c0d0e0f'
WHEN 3 THEN '1501cc0102030405060708090a0b0c0d0e0f'
END WHERE transmission_id = 0 AND observer_idx IN (1,2,3);
SQL
- name: Seed deterministic Path Inspector routes (#2060)
run: sqlite3 test-fixtures/e2e-fixture.db < test-fixtures/path-inspector.sql
- name: Start Go server with fixture DB
run: |
fuser -k 13581/tcp 2>/dev/null || true
sleep 1
./corescope-server -port 13581 -db test-fixtures/e2e-fixture.db -public public-instrumented &
echo $! > .server.pid
for i in $(seq 1 30); do
if curl -sf http://localhost:13581/api/healthz > /dev/null 2>&1; then
echo "Server ready after ${i}s"
break
fi
if [ "$i" -eq 30 ]; then
echo "Server failed to start within 30s"
exit 1
fi
sleep 1
done
- name: Start user-management E2E server (fake mailer)
if: matrix.shard == 3
run: |
fuser -k 13582/tcp 2>/dev/null || true
rm -rf /tmp/cs-um
mkdir -p /tmp/cs-um
# Own copy of the (already migrated and seeded) fixture; fresh users.db every run.
cp test-fixtures/e2e-fixture.db /tmp/cs-um/on.db
cat > /tmp/cs-um/config.json <<'JSON'
{
"port": 13582,
"userManagement": {
"enabled": true,
"dbPath": "/tmp/cs-um/users.db",
"adminEmails": ["admin@e2e.test"],
"publicBaseUrl": "http://localhost:13582",
"mail": { "provider": "fake", "fromEmail": "noreply@e2e.test" },
"channelProposals": { "enabled": true },
"notifications": { "enabled": true }
}
}
JSON
./corescope-server-e2e -config-dir /tmp/cs-um -port 13582 -db /tmp/cs-um/on.db -public public-instrumented &
echo $! > .server-um.pid
for i in $(seq 1 30); do
if curl -sf http://localhost:13582/api/healthz > /dev/null 2>&1; then echo "UM server ready after ${i}s"; break; fi
if [ "$i" -eq 30 ]; then echo "UM server failed to start within 30s"; exit 1; fi
sleep 1
done
- name: Run Playwright E2E tests (fail-fast)
run: |
# suite <shard> <file> [VAR=value ...]: runs tests/e2e/<file> with
# exactly those variables, but only on its own shard. The shard numbers
# balance the measured suite times; give a new suite the shard whose
# total is lowest. Anything but a number in 1..E2E_SHARDS (a missing
# number, 01, 4) fails the step, so a typo cannot silently drop a suite.
B=http://localhost:13581
suite() {
local shard=$1 file=$2
shift 2
if ! [[ "$shard" =~ ^[1-9][0-9]*$ ]] || [ "$shard" -gt "$E2E_SHARDS" ]; then
echo "::error::$file is assigned to shard $shard, outside 1..$E2E_SHARDS"
exit 1
fi
[ "$shard" -eq "$SHARD" ] || return 0
echo "=== E2E SUITE: $file ===" | tee -a e2e-output.txt
env "$@" node "tests/e2e/$file" 2>&1 | tee -a e2e-output.txt
}
suite 2 test-e2e-playwright.js BASE_URL=$B
# M5+M6 of #1668 — axe-core CI gate.
# M5: color-contrast on desktop dark+light.
# M6: expanded ruleset (image-alt, label, aria-required-attr,
# aria-valid-attr, aria-valid-attr-value, landmark-one-main,
# region, button-name, link-name, document-title, html-has-lang,
# duplicate-id) AND adds 375x812 mobile viewport (with
# color-contrast on mobile too).
# Allowlist: tests/a11y-allowlist.yaml (0 entries — hard pass policy).
# Per-viewport summary printed at the end; any net>0 fails the build.
suite 1 test-a11y-axe-1668.js BASE_URL=$B AXE_SCREENSHOT_DIR=/tmp/axe-1668
suite 2 test-filter-ux-e2e.js BASE_URL=$B
suite 1 test-channel-issue-1087-e2e.js BASE_URL=$B
suite 1 test-channel-issue-1111-e2e.js BASE_URL=$B
suite 2 test-map-modal-fluid-e2e.js BASE_URL=$B
suite 2 test-map-nodes-pagination-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-observer-iata-1188-e2e.js BASE_URL=$B
suite 2 test-issue-1639-observers-sort-e2e.js BASE_URL=$B
suite 2 test-issue-1758-ng-filter-rerenders-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 2 test-nav-fluid-1055-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-nav-priority-1102-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-nav-priority-1311-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 2 test-nav-priority-1391-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 1 test-issue-1413-nav-overlap-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 2 test-issue-1400-nav-vertical-clip.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 2 test-nav-more-floor-1139-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-bottom-nav-1061-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-gestures-1062-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 1 test-gestures-1185-scroll-discriminator-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-gesture-hints-1065-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-touch-gestures-coverage-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 2 test-channel-fluid-e2e.js BASE_URL=$B
suite 2 test-table-fluid-e2e.js BASE_URL=$B
suite 3 test-charts-fluid-1058-e2e.js BASE_URL=$B
suite 3 test-slideover-1056-e2e.js BASE_URL=$B
suite 3 test-issue-1692-packets-init-parallel-e2e.js BASE_URL=$B
suite 2 test-slideover-1168-munger-e2e.js BASE_URL=$B
suite 3 test-logo-pulse-1173-e2e.js BASE_URL=$B
suite 2 test-issue-1122-packets-filter-ux-e2e.js BASE_URL=$B
suite 3 test-issue-1128-packets-layout-e2e.js BASE_URL=$B
suite 3 test-issue-1128-multi-viewport-e2e.js BASE_URL=$B
suite 2 test-issue-1136-live-region-e2e.js BASE_URL=$B
suite 1 test-live-multibyte-only-e2e.js BASE_URL=$B
suite 3 test-issue-1150-404-state-e2e.js BASE_URL=$B
suite 2 test-issue-1146-path-link-contrast-e2e.js BASE_URL=$B
suite 1 test-issue-1705-subpath-contrast-e2e.js CHROMIUM_REQUIRE=1
suite 2 test-issue-1147-section-order-e2e.js BASE_URL=$B
suite 3 test-issue-1151-orphan-separators-e2e.js BASE_URL=$B
suite 3 test-issue-1486-collapse-reopens-detail-e2e.js BASE_URL=$B
suite 1 test-logo-rebrand-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 2 test-logo-theme-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 2 test-logo-default-sage-teal-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-issue-1109-hamburger-dropdown-visible-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 2 test-live-layout-1178-1179-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 2 test-issue-1205-live-controls-anchor-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-live-mql-leak-1180-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 2 test-issue-1204-live-panel-structure-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-issue-1234-live-chrome-pass2-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 1 test-issue-1206-vcr-overlap-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 1 test-issue-1833-legend-toggle-clickable-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-issue-1244-live-vcr-row-hints-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 2 test-issue-1510-live-nav-pin-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-live-fullscreen-1572-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-issue-1599-replay-freeze-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 2 test-issue-1648-m1-icons-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-issue-1648-m2-icons-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 2 test-issue-1648-m3-icons-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 1 test-issue-1648-m4-icons-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 1 test-issue-1657-analytics-channels-group-sprites-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 2 test-issue-1224-channels-mobile-ux-e2e.js BASE_URL=$B
suite 3 test-issue-1367-channels-chat-app-e2e.js BASE_URL=$B
suite 3 test-issue-1236-map-mobile-e2e.js BASE_URL=$B
suite 3 test-issue-1329-map-controls-accordion-e2e.js BASE_URL=$B
suite 1 test-issue-1273-qr-overlay-height-e2e.js BASE_URL=$B
suite 3 test-issue-1281-location-row-e2e.js BASE_URL=$B
suite 1 test-issue-1279-legend-p2-e2e.js BASE_URL=$B
suite 2 test-issue-1799-label-vocab-e2e.js BASE_URL=$B
suite 2 test-home-coverage-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 2 test-path-inspector-coverage-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-issue-1206-resize-observer-leak-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-nav-drawer-1064-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-audio-live-1297-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-audio-lab-1297-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 1 test-channel-decrypt-e2e.js BASE_URL=$B
suite 2 test-channel-qr-e2e.js BASE_URL=$B
suite 3 test-channel-color-picker-e2e.js BASE_URL=$B
suite 1 test-customize-theme-e2e.js BASE_URL=$B
suite 3 test-customize-branding-e2e.js BASE_URL=$B
suite 2 test-customize-display-e2e.js BASE_URL=$B
suite 3 test-customize-export-e2e.js BASE_URL=$B
suite 2 test-drag-manager-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 1 test-issue-1567-corner-clears-drag-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-issue-1306-collisions-terminology-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 2 test-issue-1374-route-map-a11y-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 2 test-channels-list-render-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 1 test-channels-selection-flow-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-channels-add-modal-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 1 test-channels-share-color-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 2 test-channels-ws-batch-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-channels-ws-race-1498-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-issue-1487-byop-modal-layout-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 2 test-issue-1630-reach-mobile-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 1 test-node-reach-coverage-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-issue-1640-compare-discovery-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 2 test-neighbor-map-btn-clip-e2e.js CHROMIUM_REQUIRE=1
suite 3 test-rx-coverage-viewport-e2e.js BASE_URL=$B
suite 1 test-rx-coverage-noise-e2e.js BASE_URL=$B
# Issue #2037: these eight sat in tests/e2e/ with no runner at all.
# Measured on upstream master in CI run 35340835178: all eight pass,
# 45 seconds together. test-rx-coverage-mobile-nav-e2e.js is the
# ninth that passed and is deliberately NOT here: it exits 0 with
# "SKIP (clientRxCoverage disabled on this deployment)", which is the
# default, so wiring it in would add a vacuous pass rather than a guard.
suite 2 test-1110-live-filter.js BASE_URL=$B
suite 3 test-analytics-fluid-charts.js BASE_URL=$B
suite 3 test-e2e-1267-mobile-vcr.js BASE_URL=$B
suite 2 test-issue-1274-legend-coverage-e2e.js BASE_URL=$B
suite 3 test-issue-1648-m5-icons-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 2 test-issue-1997-distance-building-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 1 test-channel-modal-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 1 test-issue-1522-trace-url-sync-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 2 test-marker-outline-weight.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 2 test-node-reach-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-packets-compact-columns-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-packets-scope-column.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-path-inspector-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-pr-1490-live-map-gpu-animations-e2e.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-table-sort.js
suite 1 test-touch-targets.js CHROMIUM_REQUIRE=1 BASE_URL=$B
suite 3 test-live-dedup.js BASE_URL=$B
suite 1 test-nodes-export-e2e.js BASE_URL=$B
suite 2 test-show-neighbors.js BASE_URL=$B
suite 3 test-user-management-e2e.js BASE_URL=http://localhost:13582 BASE_URL_OFF=$B
# #1616: slide-over focus-restore flake-gate. Runs the slide-over
# E2E 3 more times against the SAME backend instance (it already
# ran once in the list above) so the Chromium-headless focus race
# documented in #1172/#1616 gets repeated shots at firing. Any
# single non-zero exit aborts. This is the architectural-fix gate:
# if it ever turns red post-merge, the focused-but-hidden state has
# crept back in.
#
# PERMANENT step, on shard 3 (the shard that runs the slide-over
# suite). Adds ~70s to that shard in exchange for closing out a
# flake family that was blocking ~8 unrelated PRs at a time. If
# profiling pressures the budget later, drop repeat count first;
# do not delete.
- name: Slide-over E2E flake-gate (#1616, --repeat-each=3)
if: matrix.shard == 3
run: |
set -e
for i in $(seq 1 3); do
echo "--- slide-over E2E run $i/3 ---"
echo "=== E2E SUITE: test-slideover-1056-e2e.js ==="
BASE_URL=http://localhost:13581 node tests/e2e/test-slideover-1056-e2e.js 2>&1 | tee -a slideover-repeat-output.txt
done
echo "3 passed"
# Shard 2 runs test-e2e-playwright.js, the one suite that writes
# .nyc_output, so the coverage collection and its badge live there too.
- name: Collect frontend coverage (parallel)
if: success() && github.event_name == 'push' && matrix.shard == 2
run: |
BASE_URL=http://localhost:13581 node scripts/collect-frontend-coverage.js 2>&1 | tee fe-coverage-output.txt || true
- name: Generate frontend coverage badge
if: success()
run: |
mkdir -p .badges
if [ -f .nyc_output/frontend-coverage.json ] || [ -f .nyc_output/e2e-coverage.json ]; then
npx nyc report --reporter=text-summary --reporter=text 2>&1 | tee fe-report.txt
FE_COVERAGE=$(grep 'Statements' fe-report.txt | head -1 | grep -oP '[\d.]+(?=%)' || echo "0")
FE_COVERAGE=${FE_COVERAGE:-0}
FE_COLOR="red"
[ "$(echo "$FE_COVERAGE > 50" | bc -l 2>/dev/null)" = "1" ] && FE_COLOR="yellow"
[ "$(echo "$FE_COVERAGE > 80" | bc -l 2>/dev/null)" = "1" ] && FE_COLOR="brightgreen"
echo "{\"schemaVersion\":1,\"label\":\"frontend coverage\",\"message\":\"${FE_COVERAGE}%\",\"color\":\"${FE_COLOR}\"}" > .badges/frontend-coverage.json
echo "## Frontend: ${FE_COVERAGE}% coverage" >> $GITHUB_STEP_SUMMARY
fi
- name: Stop test server
if: always()
run: |
if [ -f .server.pid ]; then
kill $(cat .server.pid) 2>/dev/null || true
rm -f .server.pid
fi
if [ -f .server-um.pid ]; then
kill $(cat .server-um.pid) 2>/dev/null || true
rm -f .server-um.pid
fi
- name: Upload shard results
if: success()
uses: actions/upload-artifact@v6
with:
name: e2e-shard-${{ matrix.shard }}
path: |
e2e-output.txt
.badges/
retention-days: 1
if-no-files-found: ignore
include-hidden-files: true
# The gate the rest of the pipeline (and branch protection) sees under the
# old single-job name. It folds the shard outputs into the same e2e-badges
# artifact the publish job reads.
#
# !cancelled() is what makes it a gate. Without it a failed shard would SKIP
# this job, and branch protection counts a skipped required check as
# passing. With it, the job runs and fails on its first step instead.
e2e-test:
name: "🎭 Playwright E2E Tests"
needs: [e2e-shard, changes]
runs-on: ubuntu-latest
if: ${{ !cancelled() && needs.changes.outputs.code == 'true' && !(startsWith(github.ref, 'refs/tags/v') && inputs.images_published) }}
steps:
- name: Require every shard to have passed
run: |
result="${{ needs['e2e-shard'].result }}"
echo "e2e-shard: $result"
[ "$result" = success ] || { echo "::error::E2E shards did not all pass (result: $result)"; exit 1; }
- name: Checkout code
uses: actions/checkout@v5
- name: Download shard results
uses: actions/download-artifact@v6
with:
pattern: e2e-shard-*
path: shards/
- name: Generate E2E badges
run: |
set -e
ls shards
[ "$(ls -d shards/e2e-shard-* | wc -l)" -eq 3 ] || { echo "::error::expected results from 3 shards"; exit 1; }
cat shards/e2e-shard-*/e2e-output.txt > e2e-output.txt
mkdir -p .badges
cp shards/e2e-shard-*/.badges/*.json .badges/ 2>/dev/null || true
# Aggregate per-suite PASS/FAIL across every test-*-e2e.js summary.
# The previous regex (grep -oP '[0-9]+(?=/)' | tail -1) caught a
# stray digits-before-slash like the '2' in '2/3 tests passed' from
# some sub-output and stamped the badge as '2 passed'. See #1296.
eval "$(bash scripts/aggregate-e2e-pass.sh e2e-output.txt)"
E2E_PASS=${PASS:-0}
E2E_FAIL=${FAIL:-0}
if [ "${E2E_FAIL:-0}" -gt 0 ]; then
E2E_MSG="${E2E_PASS:-0} passed, ${E2E_FAIL} failed"
E2E_COLOR="red"
else
E2E_MSG="${E2E_PASS:-0} passed"
E2E_COLOR="brightgreen"
fi
echo "{\"schemaVersion\":1,\"label\":\"e2e tests\",\"message\":\"${E2E_MSG}\",\"color\":\"${E2E_COLOR}\"}" > .badges/e2e-tests.json
echo "E2E: ${E2E_MSG} ($(grep -c '^=== E2E SUITE' e2e-output.txt) suite runs)" | tee -a "$GITHUB_STEP_SUMMARY"
- name: Upload E2E badges
uses: actions/upload-artifact@v6
with:
name: e2e-badges
path: .badges/
retention-days: 1
if-no-files-found: ignore
include-hidden-files: true
# ───────────────────────────────────────────────────────────────
# 3a. Two-arch image build — beside the tests, never publishes
# ───────────────────────────────────────────────────────────────
# Since the SQLite driver became cgo, the container build depends on a zig
# cross-toolchain and static musl linking. Builds both architectures and
# actually runs the arm64 binaries under QEMU, on every event: for a PR it is
# the only image check (the GHCR push below is push/tag-only, so without it
# the first signal would arrive on master); for a push it fills the layer
# cache that the push then reuses.
#
# It needs only `changes`, so its minutes overlap the test jobs instead of
# following them. It cannot publish anything, so running it before the tests
# pass is safe. The build metadata is computed here once and handed to
# build-and-publish: BUILD_TIME is a build arg of the Go layers, so a second
# timestamp would miss the cache and rebuild everything.
image-check:
name: "🧱 Two-arch image build"
needs: [changes]
runs-on: ubuntu-latest
if: needs.changes.outputs.code == 'true' && !(startsWith(github.ref, 'refs/tags/v') && inputs.images_published)
outputs:
build_time: ${{ steps.meta.outputs.build_time }}
git_commit: ${{ steps.meta.outputs.git_commit }}
app_version: ${{ steps.meta.outputs.app_version }}
steps:
- name: Checkout code
uses: actions/checkout@v5
- name: Compute build metadata
id: meta
run: |
BUILD_TIME=$(date -u '+%Y-%m-%dT%H:%M:%SZ')
GIT_COMMIT="${GITHUB_SHA::7}"
if [[ "$GITHUB_REF" == refs/tags/v* ]]; then
APP_VERSION="${GITHUB_REF#refs/tags/}"
else
APP_VERSION="edge"
fi
echo "build_time=$BUILD_TIME" >> "$GITHUB_OUTPUT"
echo "git_commit=$GIT_COMMIT" >> "$GITHUB_OUTPUT"
echo "app_version=$APP_VERSION" >> "$GITHUB_OUTPUT"
echo "Build: version=$APP_VERSION commit=$GIT_COMMIT time=$BUILD_TIME"
# The staging compose file used to be built here as well. That image was
# never used: the runner is discarded, and the deploy job pulls :edge from
# GHCR on its own runner. It rebuilt the same Dockerfile as the step
# below for two minutes. Validating the compose file is what is left.
- name: Validate staging compose file
run: docker compose -f "$STAGING_COMPOSE_FILE" -p corescope-staging config --quiet
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@v3
- name: Set up QEMU (arm64 runtime stage + arm64 smoke)
uses: docker/setup-qemu-action@v3
- name: Two-arch build (no publish)
uses: docker/build-push-action@v6
with:
context: .
push: false
platforms: linux/amd64,linux/arm64
tags: corescope:check-${{ github.run_id }}
build-args: |
APP_VERSION=${{ steps.meta.outputs.app_version }}
GIT_COMMIT=${{ steps.meta.outputs.git_commit }}
BUILD_TIME=${{ steps.meta.outputs.build_time }}
cache-from: type=gha
cache-to: type=gha,mode=max
- name: Smoke the arm64 binaries under QEMU
run: |
set -e
# Export the arm64 binaries so the runner's own `file` can confirm
# they are static, then execute one under QEMU. Cheap: the build above
# already populated the layer cache.
docker buildx build . --platform linux/arm64 \
--build-arg APP_VERSION=smoke --cache-from type=gha \
--output type=local,dest=./arm64-image
for b in corescope-server corescope-ingestor corescope-decrypt; do
file "./arm64-image/app/$b" | tee /dev/stderr | grep -q 'statically linked' \
|| { echo "::error::$b is not statically linked"; exit 1; }
file "./arm64-image/app/$b" | grep -q 'ARM aarch64' \
|| { echo "::error::$b is not an arm64 binary"; exit 1; }
done
docker run --rm --platform linux/arm64 \
-v "$PWD/arm64-image/app:/bin-under-test:ro" alpine:3.20 \
/bin-under-test/corescope-decrypt --version
rm -rf ./arm64-image
echo "arm64 static binaries run under QEMU ✅"
# ───────────────────────────────────────────────────────────────
# 3b. Publish Docker Image — only after every test job passed
# ───────────────────────────────────────────────────────────────
# The GHCR steps below publish on a push to master (the :edge image) and
# on a tag ref. The tag half is not decoration: release-fast-path.yml re-tags
# :edge to :vX.Y.Z when the :edge revision label matches the tagged commit,
# and dispatches THIS workflow when it does not. That fallback previously
# produced no image at all, because a workflow_dispatch is not a push and
# every publishing step was gated on github.event_name == 'push'. It built
# locally, reported success, and pushed nothing.
#
# That went unnoticed until v3.10.0, where the tagged commit was
# documentation-only. At that time a trigger-level paths-ignore skipped this
# workflow for such commits, so no :edge image was ever built for it, the label
# comparison failed, and the fallback ran for the first time. That filter is
# gone: the `changes` job now forces code=true for anything that is not a pull
# request, so every master commit gets an image and a documentation commit is
# safe to tag. See the comment on the `on:` block.
#
# On a pull request only the first step runs; the job stays so that its
# check still means "tests and image build all passed".
#
# !cancelled() plus the first step make that true for failures too: without
# them a failed needed job would SKIP this job, and branch protection counts
# a skipped required check as passing. race-test is skipped when no ingestor
# Go file changed, so for it skipped is accepted; a master push always runs it.
build-and-publish:
name: "🏗️ Build & Publish Docker Image"
needs: [go-test, race-test, e2e-test, image-check, changes]
runs-on: ubuntu-latest
if: ${{ !cancelled() && needs.changes.outputs.code == 'true' && !(startsWith(github.ref, 'refs/tags/v') && inputs.images_published) }}
steps:
- name: Require every test job and the image check to have passed
run: |
fail=0
require() {
echo "$1: $2"
case " $3 " in
*" $2 "*) ;;
*) echo "::error::$1 result is $2"; fail=1 ;;
esac
}
require go-test "${{ needs['go-test'].result }}" "success"
require race-test "${{ needs['race-test'].result }}" "success skipped"
require e2e-test "${{ needs['e2e-test'].result }}" "success"
require image-check "${{ needs['image-check'].result }}" "success"
exit $fail
- name: Checkout code
if: ${{ github.event_name == 'push' || startsWith(github.ref, 'refs/tags/v') }}
uses: actions/checkout@v5
- name: Set up Docker Buildx
if: ${{ github.event_name == 'push' || startsWith(github.ref, 'refs/tags/v') }}
uses: docker/setup-buildx-action@v3
- name: Set up QEMU (arm64 runtime stage)
if: ${{ github.event_name == 'push' || startsWith(github.ref, 'refs/tags/v') }}
uses: docker/setup-qemu-action@v3
- name: Log in to GHCR
if: ${{ github.event_name == 'push' || startsWith(github.ref, 'refs/tags/v') }}
uses: docker/login-action@v3
with:
registry: ghcr.io
username: ${{ github.actor }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Extract Docker metadata
if: ${{ github.event_name == 'push' || startsWith(github.ref, 'refs/tags/v') }}
id: docker-meta
uses: docker/metadata-action@v5
with:
images: ghcr.io/kpa-clawbot/corescope
tags: |
type=semver,pattern=v{{version}}
type=semver,pattern=v{{major}}.{{minor}}
type=semver,pattern=v{{major}}
type=raw,value=latest,enable=${{ startsWith(github.ref, 'refs/tags/v') }}
type=edge,branch=master
# Same build args as image-check, so every layer is a cache hit.
- name: Build and push to GHCR
if: ${{ github.event_name == 'push' || startsWith(github.ref, 'refs/tags/v') }}
uses: docker/build-push-action@v6
with:
context: .
push: true
platforms: linux/amd64,linux/arm64
tags: ${{ steps.docker-meta.outputs.tags }}
labels: ${{ steps.docker-meta.outputs.labels }}
build-args: |
APP_VERSION=${{ needs.image-check.outputs.app_version }}
GIT_COMMIT=${{ needs.image-check.outputs.git_commit }}
BUILD_TIME=${{ needs.image-check.outputs.build_time }}
# No cache-to: image-check exported these same layers minutes ago.
cache-from: type=gha
# ───────────────────────────────────────────────────────────────
# 4. Release Artifacts (tags only)
# ───────────────────────────────────────────────────────────────
release-artifacts:
name: "📦 Release Artifacts"
needs: [go-test, changes]
runs-on: ubuntu-latest
permissions:
contents: write
if: startsWith(github.ref, 'refs/tags/v') && needs.changes.outputs.code == 'true'
steps:
- name: Checkout code
uses: actions/checkout@v5
- name: Set up Go 1.27
uses: actions/setup-go@v6
with:
go-version: '1.27'
cache-dependency-path: |
cmd/decrypt/go.sum
internal/channel/go.sum
# The SQLite driver is cgo now (github.com/mattn/go-sqlite3), so the Go
# toolchain alone cannot cross-compile these. zig cc is the C compiler
# that can target both architectures; musl makes the result static.
- name: Set up Zig
uses: mlugg/setup-zig@v2
with:
version: 0.16.0
- name: Build corescope-decrypt (static, linux/amd64)
run: |
cd cmd/decrypt
CGO_ENABLED=1 GOOS=linux GOARCH=amd64 CC="zig cc -target x86_64-linux-musl" go build -trimpath -tags netgo,osusergo,sqlite_omit_load_extension -ldflags="-s -w -extldflags '-static -Wl,-s' -X main.version=${{ github.ref_name }}" -o ../../corescope-decrypt-linux-amd64 .
- name: Build corescope-decrypt (static, linux/arm64)
run: |
cd cmd/decrypt
CGO_ENABLED=1 GOOS=linux GOARCH=arm64 CC="zig cc -target aarch64-linux-musl" go build -trimpath -tags netgo,osusergo,sqlite_omit_load_extension -ldflags="-s -w -extldflags '-static -Wl,-s' -X main.version=${{ github.ref_name }}" -o ../../corescope-decrypt-linux-arm64 .
- name: Verify release binaries are static and runnable
run: |
set -e
for f in corescope-decrypt-linux-amd64 corescope-decrypt-linux-arm64; do
file "$f" | grep -q 'statically linked' || { echo "::error::$f is not statically linked"; exit 1; }
done
# amd64 runs natively on the runner; arm64 is only checked structurally
# here (the container image job exercises arm64 under QEMU).
./corescope-decrypt-linux-amd64 --version
- name: Locate the release notes for this tag
id: notes
run: |
set -euo pipefail
# docs/release-notes/<tag>.md is the convention: v3.10.1.md, v3.12.0.md.
F="docs/release-notes/${GITHUB_REF_NAME}.md"
if [ -f "$F" ]; then
echo "path=$F" >> "$GITHUB_OUTPUT"
echo "found $F ($(wc -c < "$F") bytes)"
else
# Not an error. A release without a notes file still gets the
# generated pull request list below, which is more than the empty
# body v3.10.1, v3.11.0 and v3.12.0 were published with.
echo "path=" >> "$GITHUB_OUTPUT"
echo "::warning::no $F, publishing with generated notes only"
fi
- name: Upload release assets
# Standard releases upload both assets to a draft before publishing.
# Keep one writer: published releases are immutable in this repository.
#
# body_path and generate_release_notes together: our own notes first,
# then GitHub's generated pull request list under them. Without either,
# this step publishes an empty release body, which is what happened to
# every release from v3.10.1 to v3.12.0 (v3.12.0's body was filled in by
# hand afterwards). The hand-written file is the substance; the generated
# list is the "what changed" index that a reader expects to find.
uses: softprops/action-gh-release@v2
with:
fail_on_unmatched_files: true
body_path: ${{ steps.notes.outputs.path }}
generate_release_notes: true
files: |
corescope-decrypt-linux-amd64
corescope-decrypt-linux-arm64
# ───────────────────────────────────────────────────────────────
# 4b. Deploy Staging (master only)
# ───────────────────────────────────────────────────────────────
deploy:
name: "🚀 Deploy Staging"
# DISABLED. Re-enable by setting the repository variable
# ENABLE_STAGING_DEPLOY to 'true' (Settings > Secrets and variables >
# Actions > Variables). No code change needed.
#
# Why: this job runs on [self-hosted, meshcore-runner-2] and that runner
# has not picked up a job since at least 2026-08-31. It has no
# timeout-minutes, so it sat queued for 22+ hours, held its run open, and
# through the concurrency group ci-refs/heads/master caused GitHub to
# cancel every subsequent master push. The result was that master produced
# no completed pipeline result at all: 30+ runs cancelled or stuck while
# go-test, e2e-test and build-and-publish were passing inside them.
#
# timeout-minutes alone would unblock the queue but leave master
# permanently red on a job that cannot succeed while the runner is absent,
# so the deploy is gated instead. It is added here as well, so that if the
# variable is set while the runner is still missing the job fails in ten
# minutes rather than blocking the branch again.
if: |
vars.ENABLE_STAGING_DEPLOY == 'true'
&& (github.event_name == 'push' || github.event_name == 'workflow_dispatch')
&& github.ref == 'refs/heads/master'
needs: [build-and-publish]
runs-on: [self-hosted, meshcore-runner-2]
timeout-minutes: 10
steps:
- name: Checkout code
uses: actions/checkout@v5
- name: Pull latest image from GHCR
run: |
# Try to pull the edge image from GHCR and tag for docker-compose compatibility
if docker pull ghcr.io/kpa-clawbot/corescope:edge; then
docker tag ghcr.io/kpa-clawbot/corescope:edge corescope-go:latest
echo "Pulled and tagged GHCR edge image ✅"
else
echo "⚠️ GHCR pull failed — falling back to locally built image"
fi
- name: Deploy staging
run: |
# Force-remove the staging container regardless of how it was created
# (compose-managed OR manually created via docker run)
docker stop corescope-staging-go 2>/dev/null || true
docker rm -f corescope-staging-go 2>/dev/null || true
docker compose -f "$STAGING_COMPOSE_FILE" -p corescope-staging down --timeout 30 2>/dev/null || true
# Wait for container to be fully gone and OS to reclaim memory (3GB limit)
for i in $(seq 1 15); do
if ! docker ps -a --format '{{.Names}}' | grep -q 'corescope-staging-go'; then
break
fi
sleep 1
done
sleep 5 # extra pause for OS memory reclaim
# Ensure staging data dir exists (config.json lives here, no separate file mount)
STAGING_DATA="${STAGING_DATA_DIR:-$HOME/meshcore-staging-data}"
mkdir -p "$STAGING_DATA"
# If no config exists, copy the example (CI doesn't have a real prod config)
if [ ! -f "$STAGING_DATA/config.json" ]; then
echo "Staging config missing — copying config.example.json"
cp config.example.json "$STAGING_DATA/config.json" 2>/dev/null || true
fi
docker compose -f "$STAGING_COMPOSE_FILE" -p corescope-staging up -d staging-go
- name: Healthcheck staging container
run: |
for i in $(seq 1 120); do
HEALTH=$(docker inspect corescope-staging-go --format '{{.State.Health.Status}}' 2>/dev/null || echo "starting")
if [ "$HEALTH" = "healthy" ]; then
echo "Staging healthy after ${i}s"
break
fi
if [ "$i" -eq 120 ]; then
echo "Staging failed health check after 120s"
docker logs corescope-staging-go --tail 50
exit 1
fi
sleep 1
done
- name: Smoke test staging API
run: |
PORT="${STAGING_GO_HTTP_PORT:-80}"
if curl -sf "http://localhost:${PORT}/api/stats" | grep -q engine; then
echo "Staging verified — engine field present ✅"
else
echo "Staging /api/stats did not return engine field (port ${PORT})"
exit 1
fi
- name: Clean up old Docker images
if: always()
run: |
# Remove dangling images and images older than 24h (keeps current build)
echo "--- Docker disk usage before cleanup ---"
docker system df
docker image prune -af --filter "until=24h" 2>/dev/null || true
docker builder prune -f --keep-storage=1GB 2>/dev/null || true
echo "--- Docker disk usage after cleanup ---"
docker system df
# ───────────────────────────────────────────────────────────────
# 5. Publish Badges & Summary (master only)
# ───────────────────────────────────────────────────────────────
publish:
name: "📝 Publish Badges & Summary"
if: github.event_name == 'push'
# Was needs: [deploy]. The badges are built from the go-test and e2e-test
# artifacts and never needed the deploy step; depending on it meant the
# coverage badges stopped updating whenever the self-hosted runner was
# unavailable. build-and-publish already transitively requires both test
# jobs, so ordering is unchanged.
needs: [build-and-publish]
runs-on: ubuntu-latest
steps:
- name: Checkout code
uses: actions/checkout@v5
- name: Download Go coverage badges
continue-on-error: true
uses: actions/download-artifact@v6
with:
name: go-badges
path: .badges/
- name: Download E2E badges
continue-on-error: true
uses: actions/download-artifact@v6
with:
name: e2e-badges
path: .badges/
- name: Publish coverage badges to repo
continue-on-error: true
env:
GH_TOKEN: ${{ secrets.BADGE_PUSH_TOKEN }}
run: |
# GITHUB_TOKEN cannot push to protected branches (required status checks).
# Use admin PAT (BADGE_PUSH_TOKEN) via GitHub Contents API instead.
for badge in .badges/*.json; do
FILENAME=$(basename "$badge")
FILEPATH=".badges/$FILENAME"
CONTENT=$(base64 -w0 "$badge")
CURRENT_SHA=$(gh api "repos/${{ github.repository }}/contents/$FILEPATH" --jq '.sha' 2>/dev/null || echo "")
if [ -n "$CURRENT_SHA" ]; then
gh api "repos/${{ github.repository }}/contents/$FILEPATH" \
-X PUT \
-f message="ci: update $FILENAME [skip ci]" \
-f content="$CONTENT" \
-f sha="$CURRENT_SHA" \
-f branch="master" \
--silent 2>&1 || echo "Failed to update $FILENAME"
else
gh api "repos/${{ github.repository }}/contents/$FILEPATH" \
-X PUT \
-f message="ci: update $FILENAME [skip ci]" \
-f content="$CONTENT" \
-f branch="master" \
--silent 2>&1 || echo "Failed to create $FILENAME"
fi
done
echo "Badge publish complete"
- name: Post deployment summary
run: |
echo "## Staging Deployed ✓" >> $GITHUB_STEP_SUMMARY
echo "" >> $GITHUB_STEP_SUMMARY
echo "**Commit:** \`$(git rev-parse --short HEAD)\` — $(git log -1 --format=%s)" >> $GITHUB_STEP_SUMMARY