* implement resolving 2LD names by labelhash
* tiny test addition
* simplify language
* report expiry and availability correctly
* compare block instead of wall clock
* simplify language
* simplify language
* map resolver 410 to NAME NOT_FOUND (a lapsed name is an answer, not a failure)
* split error into a machine-readable code and a human message
* resolver: reduce comments
* resolver: reduce comments, remove REVIEW.html
* extend SMP protocol to support name availability queries with accurate and meaningful replies
* review fixes
* more review fixes
* docs shortening and other review fixes
* add catch all for furture variants
* adversarial review (against simplex-chat) fix
* next iteration fixes
* adapt house style
* eth_call guard and cache constants
* revert NAVL command and wrap it all into RSLV
* doc fixes
* Update src/Simplex/Messaging/Server/Names.hs
Co-authored-by: Evgeny <evgeny@poberezkin.com>
* Update src/Simplex/Messaging/Protocol.hs
Co-authored-by: Evgeny <evgeny@poberezkin.com>
* Update src/Simplex/Messaging/Protocol.hs
Co-authored-by: Evgeny <evgeny@poberezkin.com>
* Update src/Simplex/Messaging/Protocol.hs
Co-authored-by: Evgeny <evgeny@poberezkin.com>
* Update src/Simplex/Messaging/Protocol.hs
Co-authored-by: Evgeny <evgeny@poberezkin.com>
* protocol types refactoring
* next iteration on types only
* rentPrices map
* claude answering ep review. to be continued...
* separate labelhash and plaintext names cleanly
* align implementation with latest type changes to test against client
* fix adversarial review findings
* trim diff
* align resolver with contract changes for names v2
* fix reverse compatibility with .testing mainnet
* fix more robustly
* NameQuery simplification
* revert drive-by refactoring
* implement full type change and introduce resolver endpoint versioning
* fix stale docs
* derive NameRegistration JSON the same way on every build
sumTypeJSON switches to the _owsf form on swift builds, but this JSON is
the RNAME payload and the resolver's HTTP contract, so a swift client and
a Linux relay would disagree on every field. taggedObjectJSON is what it
already resolves to everywhere else.
The label modifier keeps the reservedReason_ collision escape out of the
API, as AgentWorkersDetails does: both arms now say reservedReason.
* fix resolver boundary: status before body, cover /v2, drop dead field
httpGet read the response body before checking the status, so an oversized
error page surfaced as a transient "response too large" instead of the
authoritative status. Reverting it to its previous shape restores that and
removes the status test both callers had been re-deriving.
registration() had no tests at all, though it is the endpoint SMP v22
consumes. RegistrationV2Tests covers the three answer shapes, the error
paths and the exact key set of each, which is the wire contract.
auctionUntil was always None with no consumer. The spec claimed a reason
word travels unchanged; a resolver can only send a word it has, and SNRC's
registry records a number.
* fix review findings
* resolver errors say what went wrong, not "no such name"
/v2/resolve answers 200, 400 or 502, and an unregistered name is
NRAvailable, so no status means "not registered". Mapping 400/404/410 to
NOT_FOUND made a misconfigured relay deny every name, and hid a relay
upgraded ahead of its resolver. All three now surface as RESOLVER.
rslvNotFound would have gone dead, so it counts what its name says: an
availability answer the encoder downgrades for a session below v22. The
wire is unchanged.
A hashed query the registrar cannot name is refused with 502 rather than
answered with a record named "unknown", which the client rejects anyway.
NRRUnknown is capped to 32 printable characters again, as the spec says.
A registered name that is also reserved no longer offers a date it will
never free up on. rentPrices is registrationPrices throughout, and
yearPriceUSD is USDCents rather than a bare Int64.
* fix regressions found reviewing the last two commits
resolveNameMsg read the version off thParams', which on the PFWD path is
the proxy's session, not the client's. It takes the version as an argument
now, so each call site passes its own — the forwarded one uses fwdVersion.
reservedReasonOf matched the reason words before capping, so "internal
review" became NRRUnknown "internal", which encodes back as NRRInternal.
Capping precedes the match, so what is kept encodes to what it decoded.
Four agent tests still pinned NAME NOT_FOUND from a 404 stub, and two spec
statements still described the old mapping. A registrar that does not
record labels cannot answer a hashed query, which the resolver README now
says.
* docs sweep
* simplify encoding
* drop resolver caching (#1866)
* rename
* remove trailing_ filter
* fix resolver v2 for subnames (#1867)
* fix resolver v2 for subnames
* simplify doc
* resolver should report response truth freshness (#1868)
* first shot at reporting freshness
* fix review findings
* renaming
---------
Co-authored-by: sh <github.shum@liber.li>
Co-authored-by: Evgeny <evgeny@poberezkin.com>
* feat: add server public information handling in XFTP protocol
* test
---------
Co-authored-by: sh <github.shum@liber.li>
Co-authored-by: Evgeny Poberezkin <evgeny@poberezkin.com>
* agent: fast queue rotation does not requiring the current server to be online
* update plan
* update plan and rfc
* corrections
* update plan
* implementation
* test
* refactor
* test switching notificaitons
* rename
* fix, refactor
* fix possible crash
* simplify
* refactor
* simplify
* refactor
* simplify
* order, avoid re-reading connection
---------
Co-authored-by: Evgeny @ SimpleX Chat <259188159+evgeny-simplex@users.noreply.github.com>
* agent: report PRXY broker errors as proxy-to-relay errors
The proxy returns PROXY BROKER errors only when it fails to connect to the
destination relay, but for PRXY they were mapped to SMP <proxy> (PROXY ...),
losing the relay address, so the clients reported them as errors of the
connection to the forwarding server. Map them to PROXY {proxyServer,
relayServer, ...}, the same shape PFWD errors already use.
* Apply suggestion from @epoberezkin
* agent: keep ProxyProtocolError for PRXY broker errors
Reverts 826a2377.
ProxyResponseError wraps PCEResponseError, a response that failed to
parse, and both clients match protocolError when rendering this error,
so the connect alert falls through to the raw error dump - and released
clients cannot render it at all. temporaryAgentError and serverHostError
also match ProxyProtocolError only, so the failure stops being retried
and proxy fallback no longer engages.
---------
Co-authored-by: Evgeny <evgeny@poberezkin.com>
* smp-server: add serverInfoBytes to THandleParams and related functions
* server information
* fix test
---------
Co-authored-by: Ed Asriyan <service.github@asriyan.me>
* servers: drop support of versions before 01/2025
* fix agent tests
* fix ntf tests
* fix test
* remove code for non-batched transmissions
* fix batching test
* fix ntf
* comment
Service subscription counters (totalServiceSubs, serviceSubsCount,
ntfServiceSubsCount) are TVar (Int64, IdsHash). modifyTVar' only forces
the pair to WHNF, so `n +/- n'` and `idsHash <> idsHash'` stay
unevaluated and accumulate an unbounded thunk chain under subscription
and delivery churn (the IdsHash chain also retains a bytestring per
update) - a space leak proportional to the number of updates.
Force both components in addServiceSubs/subtractServiceSubs. Verified
with the load bench: svc churn drops from +5.4 KiB/iter (linear) to flat.
* smp-server: namespaces resolver scaffolding
* smp-server: Names resolver hardening + cleanup
* smp-server: fuse parallel dispatchers
* smp-server: JSON wire format for NameRecord + Names.hs restructure
* smp-server: redact RpcAuth in Show
* smp-server: JSON wire fixups + spec rewrite + small cleanups
* plan: prepend implementation-diverged banner
* move SimplexName into shared module
* smp-server: name + contract whitelist on RSLV
* smp-server: address audit findings (canonical JSON, INI guards, SSRF, TLD case, shutdown)
* smp-server: round 2 audit fixes (label case, response cap, ipv6 link-local)
* smp-server: round 3 audit fixes (SSRF coverage, drop noop closeManager, CSV order)
* smp-server: round 4 audit fixes (0X-hex host, expanded IPv6 forms, pingEndpoint timeout)
* smp-server: hardcode TldRegistries (drop registry_tld_* INI keys)
* smp-server: round 6 audit fixes (IPv6 SSRF, redirects, ASCII labels)
- Reject IPv6 aliases of 169.254.169.254 (IPv4-compatible / IPv4-mapped /
6to4 / NAT64) via numeric range check on parsed IPv6.
- Disable HTTP redirects on the Eth RPC request.
- Restrict SimplexName labels to ASCII (Cyrillic/Greek/full-width otherwise
hash to different on-chain records and diverge from UTS-46 registrars).
- pingEndpoint: only JsonRpcErr means "reachable"; transport/decode failures
fail startup. boundedIniInt: readMaybe over partial read.
- Add 127.0.0.0/8 and 0.0.0.0 to isLoopback.
- Replace hand-rolled hex helpers with Data.ByteArray.Encoding; raise
managerConnCount to match rpcMaxConcurrency; hex Show for NameOwner.
- Fuse parallel http/https when into unless+case; drop reverse/re-reverse
in mkDomain TLDWeb; first AbiInvariantViolated; Nothing <$ decodeAddress;
forM_ (eitherToMaybe ...); >>= chain in NameOwner FromJSON.
- Drop dead imports/exports/pragmas and two restating comments.
- Tests: factor unsafeOwner/unsafeLink, addr1/2/3, testNamesConfig; add
non-ASCII label rejection coverage.
* namespace: bound parser input to 253 bytes (DoS defense)
The bare-name fallback and bareDomain parser would otherwise consume
arbitrarily many non-space bytes via takeWhile1 before any validation
or length check. A crafted multi-megabyte token would be decoded as
UTF-8 and re-parsed in full before being rejected.
Introduce `boundedNonSpace` (scan with 253-byte cap) at the two
takeWhile1 sites. Inputs longer than 253 bytes leave residue that
parseOnly's implicit endOfInput rejects, so the parser fails fast
without ever allocating the full input.
The bound is the DNS full-domain limit, chosen for being a familiar
ceiling generous enough to cover any realistic SimpleX name (longest
plausible @user.subdomain.simplex stays well under 100 bytes). No
per-label cap — SimpleX names don't go through DNS label resolution
and there's no semantic reason to constrain individual labels.
* namespace: switch to Python HTTP resolver + agent plumbing (#1796)
* namespace: relax resolver_endpoint validation (path prefix, http without auth)
validateUrl gains two operator-friendly relaxations and a regression test:
- Allow a path prefix (e.g. https://gw.example.com:443/snrc) for a resolver
behind a reverse-proxy sub-path; /resolve/<name> and /health are appended
(HttpResolver already strips one trailing slash, so root and sub-path
behave identically). Query/fragment/userinfo stay rejected.
- Off-loopback, reject only http WITH resolver_auth (the Authorization header
would travel in cleartext). http without auth is now allowed (no secret to
leak; resolver data is public — also lets dev setups reach a host resolver
via http://host.docker.internal). https is always allowed, with or without
auth. Plain http has no response integrity; intended for trusted/local
networks only.
Exports validateUrl and adds validateUrlSpec (11 cases) to SMPNamesTests.
* namespace: NameRecord links as arrays (multi-link, cap 5)
* namespace: distinct RSLV error responses
RSLV collapsed every non-hit (no resolver, malformed name, not found,
backing-store failure) to ERR AUTH, so a client iterating its configured
servers could not tell "this router has no resolver, try the next" from
"name not registered, stop", and a transient backend error read as an
authoritative miss.
Names capability is runtime config, orthogonal to the linear SMP version
(a future v21 router without [NAMES] must still advertise v21), so it is
signalled by a command-time error like allowSMPProxy, not by the version
range:
no resolver configured -> ERR CMD PROHIBITED (client skips, tries next)
backing-store failure -> ERR INTERNAL (transient: retry/surface)
not found / malformed -> ERR AUTH (authoritative "no such name")
Update the protocol spec error table and add agent tests for the
no-resolver (CMD PROHIBITED) and backend-failure (INTERNAL) paths.
* refactor(names): server role + one error type
Addresses epoberezkin's review (PR #1784). Name resolution becomes a
server role like proxy; the agent owns resolution + server selection;
one error type flows through the whole stack.
- ServerRoles gains `names`; UserServers gains `nameSrvs` (opt-in list);
resolveSimplexName drops the explicit server arg and picks a
names-capable server via getNextServer.
- RSLV carries SimplexNameDomain (was RslvRequest): no JSON on the wire,
contract dropped, name validated at parse (invalid -> CMD SYNTAX).
- Version check moves from the encoder to Client.hs (no ERR to server).
- ErrorType.NAME {nameErr :: NameErrorType} (+ AgentErrorType.NAME),
wire- and JSON-encoded; resolver errors surface with diagnostics.
Success response renamed NAME -> RNAME to free the collision.
- NameOwner -> EthAddress (record selector); NameRecord derives FromJSON
and gains field-ordered Encoding; per-field caps removed.
- Remove newEnvWithNames / runSMPServerBlockingWithNames test seams;
stub resolver folded into ServerConfig.namesResolverCall_.
* test(server): update stats backup line count
NameResolverStatsData adds 6 lines to the server stats backup (the
"rslvStats:" header plus the reqs/succ/notFound/resolverErrs/disabled
fields), so testRestoreMessages' expected stats-backup line count is
95 -> 101.
* feat(names): public-namespace resolution via RSLV/RNAME
SNRC names resolver role: RSLV command -> HTTP resolver -> RNAME record.
Agent owns server selection (ServerRoles.names); NAME error family; async,
concurrency-bounded resolution; length-prefixed extensible wire; spec.
* remove comments
Co-authored-by: Evgeny <evgeny@poberezkin.com>
* simplify
* move tests name
* simplify: text addresses, Tail JSON, drop admitRslv
* fix
* remove spaghetti
* reduce diff
* async again, refactor
* different threads limit for name resolutions
* remove comment
* FromField instance for SimplexNameInfo
* remove comments
* unStrJSON
* add sameConnShortLink
* remove scheme prefix
* remove unused import
* remove connecttarget tests
* remove comment
* comment
---------
Co-authored-by: Evgeny Poberezkin <evgeny@poberezkin.com>
Co-authored-by: Evgeny @ SimpleX Chat <259188159+evgeny-simplex@users.noreply.github.com>
* docs: plan for service subscriptions memory leak
* fix: remove leaked service delivery subscriptions
The CSAEndServiceSub handler decremented subscription counters but did
not remove the per-queue delivery Sub from the service client's
subscriptions map. Over queue churn a long-lived service connection
accumulated orphaned Sub entries until disconnect, leaking memory.
Mirror CSAEndSub via endServiceQueueSub (reusing endSub) so the entry
is removed and its delivery thread cancelled.
Add a regression test with white-box access to the server Env via
runSMPServerBlocking_; verified failing without the fix.
* test: remove service subs leak regression test
Remove testServiceSubsRemovedOnQueueDelete and the test-only Env
exposure (runSMPServerBlocking_, withSmpServerConfigEnvOn,
serviceSubsMapSize), leaving only the server fix.
* refactor: simplify unsubPrev with applicative
Express the cancel-if-present logic as sequence_ (unsub_ <*> s_).
* tests: add SMP proxy relay reconnection tests
Reproduces the proxy failing to reconnect to a destination relay when the
sender disconnects mid-connection (empty session var left in smpClients).
* fix: bracket session var creation to drop it on interrupt
getSessVar inserts an empty session var that the connect path then fills with
putTMVar. If the connecting thread is killed by an async exception before that
fill (a proxy worker on client disconnect, an agent worker on cancel), the empty
var was left in the map forever and every later request for that server blocked
on it until timing out (permanent PCEResponseTimeout).
Wrap get-or-create with withGetSessVar (bracketOnError) at the call sites, so the
cleanup is established where the var is created and covers the whole connect: on
interrupt before fill the still-empty var is dropped and the next request
reconnects. This closes the window between getSessVar and the fill that a handler
installed inside the connect function cannot cover.
* test: cover session var leak on interrupted connect
UtilTests: tryAllErrors rethrows ThreadKilled/StackOverflow (the mechanism
that skips putTMVar). SMPProxyTests: agent client reconnection after a
cancelled connect, plus a control proving the stalling relay alone does not
cause the failure; refine the relay reconnection tests.
* refactor
---------
Co-authored-by: Evgeny Poberezkin <evgeny@poberezkin.com>
* Disable APNS test provider in production
* refactor(ntf): extract guardPushProvider for test-provider guard
* test(ntf): fix APNS test provider test compilation
* ntf server: use ifM for push provider guard (review)
---------
Co-authored-by: Paul Bottinelli <paul.bottinelli@trailofbits.com>
* crypto: BBS scheme for anonymous credentials with multiple presentations
* verify
* add files to sources
* more files
* more files, use cabal 3.0
* fix path
* extensions
* switch libbbs to fork
* return either from keygen
* use only secret key to sign
* improve FFI
* simplify
* update libbbs to support iOS
* add commoncrypto flag
* bump libbbs
* reject input of wrong length
* ci: get submodules
---------
Co-authored-by: Evgeny @ SimpleX Chat <259188159+evgeny-simplex@users.noreply.github.com>
* agent: split SimplexNameDomain out of SimplexNameInfo
The type now separates the user-supplied type prefix (#/@) from the
domain itself:
data SimplexNameInfo = SimplexNameInfo
{ nameType :: SimplexNameType
, nameDomain :: SimplexNameDomain
}
data SimplexNameDomain = SimplexNameDomain
{ nameTLD :: SimplexTLD
, domain :: Text
, subDomain :: [Text]
}
The domain is independent of the contact-vs-public-group distinction —
the same dotted-labels structure applies to both. Future code that
needs to talk about a domain without committing to a name type (e.g.
server-side TLD-based registry lookup) can use SimplexNameDomain
directly.
fullDomainName now operates on SimplexNameDomain rather than the
full info wrapper. Parser, StrEncoding instance, and aeson derivations
updated accordingly. No external callers needed updating.
* agent: split StrEncoding instance for SimplexNameDomain
* agent: flatten TLD case + use unless guard
* agent: address review - strict domain parser, permissive channel
* smp server: messaging services (#1565)
* smp server: refactor message delivery to always respond SOK to subscriptions
* refactor ntf subscribe
* cancel subscription thread and reduce service subscription count when queue is deleted
* subscribe rcv service, deliver sent messages to subscribed service
* subscribe rcv service to messages (TODO delivery on subscription)
* WIP
* efficient initial delivery of messages to subscribed service
* test: delivery to client with service certificate
* test: upgrade/downgrade to/from service subscriptions
* remove service association from agent API, add per-user flag to use the service
* agent client (WIP)
* service certificates in the client
* rfc about drift detection, and SALL to mark end of message delivery
* fix test
* fix test
* add function for postgresql message storage
* update migration
* servers: maintain xor-hash of all associated queue IDs in PostgreSQL (#1668)
* servers: maintain xor-hash of all associated queue IDs in PostgreSQL (#1615)
* ntf server: maintain xor-hash of all associated queue IDs via PostgreSQL triggers
* smp server: xor hash with triggers
* fix sql and using pgcrypto extension in tests
* track counts and hashes in smp/ntf servers via triggers, smp server stats for service subscription, update SMP protocol to pass expected count and hash in SSUB/NSSUB commands
* agent migrations with functions/triggers
* remove agent triggers
* try tracking service subs in the agent (WIP, does not compile)
* Revert "try tracking service subs in the agent (WIP, does not compile)"
This reverts commit 59e908100d.
* comment
* agent database triggers
* service subscriptions in the client
* test / fix client services
* update schema
* fix postgres migration
* update schema
* move schema test to the end
* use static function with SQLite to avoid dynamic wrapper
* agent: fail when per-connection transport isolation is used with services (#1670)
* agent: service subscription events (#1671)
* agent: use server keyhash when loading service record
* agent: process queue/service associations with delayed subscription results
* agent: service subscription events
* agent: finalize initial service subscriptions, remove associations on service ID changes (#1672)
* agent: remove service/queue associations when service ID changes
* agent: check that service ID in NEW response matches session ID in transport session
* agent subscription WIP
* test
* comment
* enable tests
* update queries
* agent: option to add SQLite aggregates to DB connection (#1673)
* agent: add build_relations_vector function to sqlite
* update aggregate
* use static aggregate
* remove relations
---------
Co-authored-by: Evgeny Poberezkin <evgeny@poberezkin.com>
* add test, treat BAD_SERVICE as temp error, only remove queue associations on service errors
* add packZipWith for backward compatibility with GHC 8.10.7
---------
Co-authored-by: spaced4ndy <8711996+spaced4ndy@users.noreply.github.com>
* servers: service stats and logging, allow services without option (removed), report errors during service message delivery, remove threads when service subscription ended (#1676)
* smp server: always allow services without option
* smp server: maintain IDs hash in session subscription states
* smp server: service message delivery error handling
* ntf server: log subscription count and hash differences
* smp server: remove delivery threads when service subscription ended/client disconnected
* agent: remove service queue association when service ID changed, process ENDS event, test migrating to/from service (#1677)
* agent: remove service queue association when service ID changed
* agent: process ENDS event
* agent: send service subscription error event
* agent: test migrating to/from service subscriptions, fixes
* agent: always remove service when disabled, fix service subscriptions
* ntf server: use different client certs for each SMP server, remove support for store log (#1681)
* ntf server: remove support for store log
* ntf server: use different client certificates for each SMP server
* smp protocol: fix encoding for SOKS/ENDS responses (#1683)
* agent: create user with option to enable client service (#1684)
* agent: create user with option to enable client service
* handle HTTP2 errors
* do not catch async exceptions
* agent: minor fixes
* docs: update protocol (#1705)
* docs: agent threat model
* update protocol docs
* update RFCs (#1730)
* update RFCs
* update
* update overview
* update terminology
* original language in threat model
---------
Co-authored-by: Evgeny @ SimpleX Chat <259188159+evgeny-simplex@users.noreply.github.com>
* docs: fix minor issues in protocols
* docs: add e2e encrypted message wire encoding to PQDR spec
* docs: add missing encodings and other protocol corrections
* docs: move implemented rfcs
* smp: service fixes (#1737)
* smp: deliver service subscription to correct client
* tests: more resilient to concurrency
* optimize PostgreSQL query
* fix service re-association after server "downgrade"
* correctly handle service removed from server (and ID changed)
* remove unused
---------
Co-authored-by: Evgeny @ SimpleX Chat <259188159+evgeny-simplex@users.noreply.github.com>
* prometheus: fix metrics names (#1747)
* test: rcv service re-association on restart (#1746)
* agent: correct log message
* docs: update whitepaper
* smp: fix messaging client service issues (#1751)
* services: fix minor issues
* fix accounting for subscribed service queues, add prometheus stats
* fix uncorrelated subquery
* fix potential race condition when inserting service defensively, as it is also prevented by how client is created
---------
Co-authored-by: Evgeny @ SimpleX Chat <259188159+evgeny-simplex@users.noreply.github.com>
* agent: refactor cleanup if no pending subs (#1757)
* smp server: batch processing of subscription messages (#1753)
* smp server: batch processing of subscription messages
* refactor
* empty line
* fix
---------
Co-authored-by: Evgeny @ SimpleX Chat <259188159+evgeny-simplex@users.noreply.github.com>
* smp: batch queue association updates on subscriptions (#1760)
* smp: batch queue association updates on subscriptions
* refactor to fused batching
* simpler
* batch assoc functions
* clean up
* fix
---------
Co-authored-by: Evgeny @ SimpleX Chat <259188159+evgeny-simplex@users.noreply.github.com>
* agent: use primary key index in setRcvServiceAssocs (#1783)
* agent: use primary key index in setRcvServiceAssocs
Previous WHERE rcv_id = ? did not match the (host, port, rcv_id)
primary key prefix and fell back to a table scan via
idx_rcv_queues_client_notice_id. With ~390k rows per queue, each
update in a 1350-row batch scanned the whole table, yielding ~290s
per batch and a multi-hour rcv-services migration.
* agent: pass SMPServer explicitly to setRcvServiceAssocs
Avoid extracting host/port from the first queue inside setRcvServiceAssocs.
The caller already has SMPServer in scope (from tSess) and the call chain
is short, so threading it through is simpler than inspecting the list.
Removes the empty-list guard from setRcvServiceAssocs (it remains in
processRcvServiceAssocs).
---------
Co-authored-by: spaced4ndy <8711996+spaced4ndy@users.noreply.github.com>
Co-authored-by: Evgeny @ SimpleX Chat <259188159+evgeny-simplex@users.noreply.github.com>
Co-authored-by: sh <37271604+shumvgolove@users.noreply.github.com>
The per-(srvHost, provider) worker shards added in #1779 still funnel
all APNS sends through one HTTP2Client's reqQ, where a single process
thread calls sendRequest serially - one in-flight HTTP/2 stream at a
time, capping APNS throughput at 1/RTT.
sendRequestDirect bypasses the queue and invokes sendReq directly from
the calling worker, so concurrent workers open parallel HTTP/2 streams
on the shared APNS connection and the multiplexing happens on the wire.
* ntf-server: carry retry reason in PPRetryLater, log retries
Change PPRetryLater from nullary to PPRetryLater Text so the cause
(503 / 410-reason) propagates to the retry call site. Log a warning
at every retry attempt with provider, token id and reason.
* ntf-server: parallel push delivery via forkIO + per-srvHost lock
Fork delivery per notification, taking an MVar keyed by srvHost_ so
notifications from the same SMP server serialize while different
servers proceed concurrently. Switch APNS to sendRequestDirect so
concurrent deliveries share one HTTP/2 connection via stream
multiplexing rather than serializing through the client reqQ.
* ntf-server: single-flight push client creation via SessionVar
Match the take/create/wait pattern in Agent/Client.hs
(newProtocolClient / waitForProtocolClient). pushClients now wraps
clients in SessionVar (Either SomeException PushProviderClient) so
concurrent first-time access and concurrent retries collapse to a
single mkClient call; waiters observe the winner's result via
readTMVar (or its error). retryDeliver evicts the failing client by
SessionVar identity before re-fetching.
* use multiple queues and workers, remove semaphores and threads per notification
* fix
* retry connecting client
* fix
* move config
* fix
---------
Co-authored-by: Evgeny @ SimpleX Chat <259188159+evgeny-simplex@users.noreply.github.com>
Co-authored-by: Evgeny Poberezkin <evgeny@poberezkin.com>