mirror of
https://github.com/simplex-chat/simplexmq.git
synced 2026-10-06 03:27:20 +00:00
244 lines
13 KiB
Markdown
244 lines
13 KiB
Markdown
# SMP server memory at production scale
|
|
|
|
Production: ~40k client connections per server, 32 GB RAM shared with PostgreSQL, pure PostgreSQL
|
|
store, 14 servers proxying to each other, GHC 9.6.3, `+RTS -N -F1.2 -A16m -I0.01 -Iw15 -s -RTS`.
|
|
With `-F1.2` the copying GC needs about 2.2x the live heap plus the nursery, so the server's live
|
|
heap has to stay well under ~10 GB.
|
|
|
|
A busy connection (20 batch-subscribed queues, 2 created with NEW, message traffic) cost 237 KiB of
|
|
server live heap, 9.0 GiB at 40k connections. With the changes below it costs 43 KiB, 1.6 GiB.
|
|
|
|
Earlier findings on leaks and PostgreSQL load are in `leak-findings.md` and `smp-server-db-load.md`.
|
|
|
|
## Method
|
|
|
|
All numbers are server-only: the bench client runs in a child process, so the measuring process holds
|
|
only the server. PostgreSQL store, production RTS flags, 16-core machine. Live heap is measured after
|
|
two major GCs with finalizers allowed to run in between (`liveBytesMiB` in `bench/MemBench.hs`).
|
|
|
|
| Phase | Measures |
|
|
| --- | --- |
|
|
| `prodmix` | 2000 connections, each SUBs 20 queues in one block and creates 2 with NEW, then 30 s of SEND/MSG/ACK traffic |
|
|
| `conns N` | N idle connections after the SMP handshake |
|
|
| `subslice N` | per subscription: `SUBMODE=sub` (`SUBBATCH`, `SUBKEEP`), `new`, `newthensub` |
|
|
| `load` | server CPU per operation under a fixed mixed workload |
|
|
| `ntfloop`, `ntfdeliver` | notification store and its delivery loop |
|
|
| `mesh` | proxy and relay in separate processes: relay sessions, forwards in flight, forward rate (`MESH=`) |
|
|
| `pfwdbig` | oversized forwards left in the proxy's `sentCommands` |
|
|
|
|
```sh
|
|
cabal build -fserver_postgres exe:smp-mem-bench
|
|
BENCHID=1 PROD_CONNS=2000 PROD_QUEUES=20 PROD_NEW=2 PROD_SEC=30 \
|
|
$(cabal list-bin -fserver_postgres exe:smp-mem-bench) prodmix 0 \
|
|
+RTS -N -F1.2 -A16m -I0.01 -Iw15 -T -ki2k -RTS
|
|
```
|
|
|
|
`BENCHID` offsets ports and the database, so runs do not collide with each other or with tests.
|
|
|
|
## Results
|
|
|
|
`prodmix`, live heap per connection after subscribing and peak memory in use under traffic. Each row
|
|
adds to the previous one.
|
|
|
|
| Change | KiB/conn | Peak in use | 40k conns |
|
|
| --- | --- | --- | --- |
|
|
| base (`sh/fix-leak`) | 236.6 | 1075 MiB | 9.0 GiB |
|
|
| `-ki2k` | 111.5 | 719 MiB | 4.3 GiB |
|
|
| connection commits (`sh/fix-conn-mem`) | 89.8 | 635 MiB | 3.4 GiB |
|
|
| unpinned subscription keys, proxy commits (`sh/mem-combined`) | 80.0 | 604-622 MiB | 3.1 GiB |
|
|
| patched tls 1.9 | 42.9 | 469 MiB | 1.6 GiB |
|
|
|
|
Idle connections (`conns 2000`): 153 KiB at default flags, 68.4 KiB on `sh/mem-combined` with
|
|
`-ki2k`, 34.8 KiB with the patched tls.
|
|
|
|
Message throughput did not drop in any row. `-ki2k` costs some GC time (below).
|
|
|
|
## Findings
|
|
|
|
### 1. Thread stacks: 125 KiB per busy connection
|
|
|
|
Each connection has several threads that block in shallow loops. A thread starts with a 1 KiB stack
|
|
(`-ki1k`). On first overflow the RTS copies up to `-kb` (1 KiB) of frames into a new 32 KiB chunk
|
|
(`-kc32k`), which with a 1 KiB initial stack moves the loop frame itself, so the thread never returns
|
|
to its first chunk and keeps the 32 KiB chunk while the connection lives.
|
|
|
|
| Flags | idle KiB/conn | prodmix KiB/conn |
|
|
| --- | --- | --- |
|
|
| default | 153 | 236.6 |
|
|
| `-kc8k` | 100 | 135.0 |
|
|
| `-ki2k` | 91 | 111.5 |
|
|
|
|
`-ki2k` leaves the loop frame in the first chunk, so the overflow chunk is freed on return. Cost, 3
|
|
alternating 60 s `load` runs: GC CPU +15% (4.8-5.2 s to 5.7-5.8 s per run), major GCs +43% (the live
|
|
heap is smaller, so `-F1.2` triggers major GCs sooner), CPU per operation +0-5% (2.74-3.02 ms to
|
|
2.91-3.35 ms, one noisy run). One `load` run with `-kc2k` grew to 6.8 GB RSS; not investigated, so
|
|
`-kc` is left at its default.
|
|
|
|
Branch `sh/rts-ki2k` sets `-with-rtsopts=-ki2k` for smp-server. Command-line `+RTS` flags still apply.
|
|
|
|
### 2. Pinned blocks kept alive by small long-lived objects
|
|
|
|
`ByteString`s and crypton's `ScrubbedBytes` are pinned and share 4 KiB blocks. GHC treats a pinned
|
|
block as live if any object in it is live, and treats every object in a live pinned block as alive,
|
|
including dead `ScrubbedBytes`, whose weak pointers and finalizers then stay allocated (GHC 9.6.3
|
|
`Evac.c:443`, `GCAux.c:73`). A 24-byte value created during request processing, among crypto
|
|
temporaries, therefore keeps ~4 KiB and several dead keys alive for as long as it lives. A standalone
|
|
program reproduces it (pinned key 2742 B live per key, unpinned 80 B).
|
|
|
|
Subscriptions. The subscription key is such a value.
|
|
|
|
| `subslice 5000` | before | unpinned keys |
|
|
| --- | --- | --- |
|
|
| NEW-created | 5.25 KiB | 0.46 KiB |
|
|
| one SUB per block | 1.67-2.42 KiB | 0.47 KiB |
|
|
| 100 SUBs per block | 0.49 KiB | 0.46 KiB |
|
|
| NEW0, then SUB on the same connection | 4.98 KiB | 0.47 KiB |
|
|
|
|
Branch `sh/fix-sub-mem` stores the keys of `subscriptions`, `ntfSubscriptions` and `queueSubscribers`
|
|
as `ShortByteString` and shares one key between the maps. CPU unchanged.
|
|
|
|
TLS session state. tls 1.9 keeps record keys, IVs, TLS 1.3 traffic secrets, verify data and
|
|
handshake state for the connection's lifetime, all allocated during the handshake. A heap census per
|
|
connection before and after re-allocating them together at the end of the handshake:
|
|
|
|
| | tls 1.9 | patched |
|
|
| --- | --- | --- |
|
|
| weak pointers | ~138 | ~45 |
|
|
| memory in pinned blocks, outside the census | ~39 KiB | ~13 KiB |
|
|
| idle KiB/conn | 68.4 | 34.8 |
|
|
| prodmix KiB/conn | 80.0 | 42.9 |
|
|
|
|
The patch (`tls-1.9-compact-state.patch`: `Network.TLS.Handshake.Compact`, called at the end of
|
|
`handshake` under the context locks) re-derives the TLS 1.3 record keys from the kept traffic
|
|
secrets, copies the IVs, secrets, session ID, verify data, randoms, main and resumption secrets and
|
|
the transcript hash, and reseeds the context RNG. It needs a tls fork. Not compacted: the TLS 1.2 bulk
|
|
key, keys installed later by KeyUpdate. The transcript hash copy coerces crypton's hidden `Context`
|
|
newtype, so it must be checked on crypton upgrades.
|
|
|
|
Full test suite with the patched tls: 1431 examples, 0 failures. The same run on unpatched tls had 1
|
|
failure, the XFTP agent's "should resume sending file after restart".
|
|
|
|
SMP handshake state (session secret, chain keys) has the same pattern and is not done yet; the
|
|
remaining ~13 KiB outside the census is likely there.
|
|
|
|
### 3. IDs sliced from the received block: 16 KiB per subscription
|
|
|
|
Parsed `corrId` and `entityId` are slices of the ~16 KB decrypted block, and the entity ID was stored
|
|
as the subscription key, so one live key kept the block. One SUB per block cost 16.43 KiB per
|
|
subscription, one survivor of 100 cost 18.04 KiB. `sh/fix-sub-keys` copies both IDs in
|
|
`tDecodeServer` (16.43 to 1.67-2.42 KiB). Unpinned keys (finding 2) also remove it for the maps.
|
|
|
|
### 4. Connection structure: 21 KiB per busy connection
|
|
|
|
- tls 1.9 stores the receive record state lazily, so each connection kept its last received 16 KB
|
|
record. `recvTLS` forces the state after each read. Not needed with tls 2.x.
|
|
- The send and message-send loops are one thread, and one server-wide thread expires inactive clients
|
|
instead of one per client: 6 threads per connection become 4.
|
|
|
|
`sh/fix-conn-mem`: `prodmix` 111.5 to 89.8 KiB with `-ki2k`, 236.6 to 210.5 KiB without.
|
|
|
|
### 5. Notification store and its delivery loop
|
|
|
|
Keys of notifiers whose notifications were delivered stayed in the store, and the loop in
|
|
`deliverNtfsThread` sends every key to `getQueueNtfServices` (one `notifier_id IN ?` query read in
|
|
full) every 1.5 s.
|
|
|
|
| `ntfloop` (keys, no traffic) | allocation | CPU | peak in use |
|
|
| --- | --- | --- | --- |
|
|
| 0 | 0 MiB/s | 0 | 280 MiB |
|
|
| 100k | 168 MiB/s | 0.42 cores | 530 MiB |
|
|
| 300k | 263 MiB/s | 0.76 cores | 913 MiB |
|
|
|
|
`ntfdeliver` with 20k delivered notifications: 20000 keys kept, idle server allocating 36 MiB/s at
|
|
0.18 cores; with `sh/fix-ntf-store` 0 keys, 0.2 MiB/s, 0.01 cores.
|
|
|
|
### 6. Proxy mesh
|
|
|
|
The relay processed forwarded commands from a proxy one at a time in the connection's command loop,
|
|
with a PostgreSQL round trip each; all users of a proxy share its connection. Forwards the relay has
|
|
not reached wait on the proxy, each with a thread and its ~16 KB block. `mesh`, 1000 forwards per
|
|
second for 20 s, proxy and relay in separate processes, 2 runs each:
|
|
|
|
| | serial | concurrent |
|
|
| --- | --- | --- |
|
|
| forwards ok / failed | 19,968 / 0 | 19,968 / 0 |
|
|
| proxy live heap at end of sending | 192-200 MiB | 13 MiB |
|
|
| proxy memory in use | 653-672 MiB | 325 MiB (idle baseline) |
|
|
| proxy threads | 4,026-4,137 | 420 |
|
|
|
|
Here the backlog drained within the 30 s forward timeout. With a slower relay (a busy database, a
|
|
remote relay over SOCKS) the queueing delay exceeds it and forwards fail; not reproduced.
|
|
|
|
`sh/fix-proxy-mesh` forks each forwarded command under the per-connection concurrency limit, stops
|
|
keeping the forwarded block on the proxy while waiting (53.0 to 36.8 KiB per in-flight forward,
|
|
relay not answering) and caps forwards in flight per relay at 512 (`[PROXY] relay_concurrency`).
|
|
Forwarded commands use per-command nonces and secrets and are matched by correlation ID, so concurrent
|
|
processing is safe; the SimpleX agent keeps one message in flight per queue. The spec requires
|
|
responses in order per queue within a connection (`simplex-messaging.md`, "same order within each
|
|
queue ID"); serializing forwarded commands per queue ID is not done yet.
|
|
|
|
Also in the proxy path (`sh/fix-proxy-leak`): a PFWD of 16260-16266 bytes from a client that declares
|
|
`proxyServer = True` failed with `TELargeMsg` before sending and left its 20.2 KiB request in
|
|
`sentCommands` for the life of the relay session (2001 PFWDs, 2001 entries); timed-out requests were
|
|
kept too; and a relay session that stopped answering was never dropped (10 of 10 forwards timed out).
|
|
|
|
### 7. Concurrency limit and name resolution
|
|
|
|
`forkCmd` released its slot when the thread was forked, not when the command finished, so
|
|
`serverClientConcurrency` and `serverResolverConcurrency` limited nothing: 16 RSLVs with a limit of 4
|
|
all reached the resolver, 64 RSLVs from one connection were 64 concurrent requests, and 5 lookups
|
|
answered with 502 opened 5 connections because error bodies were not read. `sh/fix-rslv-fanout` holds
|
|
the slot until completion, adds a global limit (`[NAMES] resolver_global_concurrency`, 32) and reads
|
|
error bodies. With the limit working, a client that hits it blocks its own command loop.
|
|
|
|
### 8. Smaller findings
|
|
|
|
- Active-queue statistics are `IntSet`s at 64 B per element (15M elements, 916 MiB for one set), six
|
|
of them, reset only when `log_stats` is on.
|
|
- Control port `save` closes the PostgreSQL pool while the server keeps running; every later DB
|
|
operation blocks (`cpsave` phase: NEW after `save` gets no response).
|
|
|
|
### 9. tls 2.x
|
|
|
|
tls 2.1.6 (with tls-session-manager 0.0.6, crypton-connection 0.4.3, http-client-tls 0.3.6.4, an
|
|
index-state bump and version pins) saves 2-5 KiB per connection: `prodmix` 75.4 KiB against 80.0.
|
|
Problems found:
|
|
|
|
- After a failed handshake a tls 2 client and a tls 2 server wait for each other in `bye`; this hung
|
|
`testServerMultipleIdentities`. Fixed on `sh/tls2` by closing the context without `bye`.
|
|
- A tls 2 client that receives a CertificateRequest runs a timed read inside `handshake`; when the
|
|
timeout fires in the middle of a record the stream fails ("bad record mac") or hangs. SMP servers
|
|
always request client certificates. 1-2 failures per 2000 concurrent connections; present in tls
|
|
2.4.3. A `conns 2000` run with tls 2 on both sides stopped at 53 connections.
|
|
|
|
tls 2.x is not recommended; `sh/tls2` is kept for reference.
|
|
|
|
## Open
|
|
|
|
- Serialize forwarded commands per queue ID on the relay (spec response order).
|
|
- A client that sends PFWDs and never reads the responses costs the proxy 6.3 MiB live (11.4 MiB in
|
|
use) until the inactive-client expiry (up to 6 hours), 6 GiB per 1000 such clients.
|
|
- PRXY to many aliases of one relay opens an unbounded number of relay sessions, each 207 KiB on the
|
|
proxy and 149 KiB on the relay: 2.0 GiB and 1.4 GiB per 10k.
|
|
- Compact the SMP handshake state after the handshake, as in finding 2.
|
|
- Publish the tls fork and reference it from `cabal.project`.
|
|
- A SEND right after NEW or SUB can find no subscriber yet (`queueSubscribers` is updated through
|
|
`subQ`), so the message waits for the next SUB or ACK. `deliverIfSame` may leave a subscription
|
|
without a delivery thread. Both seen as occasional stuck steps in `load`; not memory.
|
|
- `[PROXY] relay_concurrency` and `[NAMES] resolver_global_concurrency` are new INI keys.
|
|
|
|
## Branches
|
|
|
|
| Branch | Base | Content |
|
|
| --- | --- | --- |
|
|
| `sh/fix-ntf-store` | master | finding 5 |
|
|
| `sh/fix-sub-keys` | master | finding 3 |
|
|
| `sh/fix-proxy-leak` | master | finding 6, request leak and stuck session |
|
|
| `sh/fix-rslv-fanout` | master | finding 7 |
|
|
| `sh/fix-conn-mem` | master | finding 4 |
|
|
| `sh/fix-sub-mem` | master | finding 2, subscription keys |
|
|
| `sh/fix-proxy-mesh` | `sh/fix-proxy-leak` | finding 6, relay concurrency, per-relay limit |
|
|
| `sh/rts-ki2k` | master | finding 1 |
|
|
| `sh/mem-combined` | `sh/fix-leak` | findings 1-6 together, as measured in Results (local) |
|
|
| `sh/tls2` | `sh/mem-combined` | finding 9, not for merge (local) |
|