13 KiB
SMP server memory at production scale
Production: ~40k client connections per server, 32 GB RAM shared with PostgreSQL, pure PostgreSQL
store, 14 servers proxying to each other, GHC 9.6.3, +RTS -N -F1.2 -A16m -I0.01 -Iw15 -s -RTS.
With -F1.2 the copying GC needs about 2.2x the live heap plus the nursery, so the server's live
heap has to stay well under ~10 GB.
A busy connection (20 batch-subscribed queues, 2 created with NEW, message traffic) cost 237 KiB of server live heap, 9.0 GiB at 40k connections. With the changes below it costs 43 KiB, 1.6 GiB.
Earlier findings on leaks and PostgreSQL load are in leak-findings.md and smp-server-db-load.md.
Method
All numbers are server-only: the bench client runs in a child process, so the measuring process holds
only the server. PostgreSQL store, production RTS flags, 16-core machine. Live heap is measured after
two major GCs with finalizers allowed to run in between (liveBytesMiB in bench/MemBench.hs).
| Phase | Measures |
|---|---|
prodmix |
2000 connections, each SUBs 20 queues in one block and creates 2 with NEW, then 30 s of SEND/MSG/ACK traffic |
conns N |
N idle connections after the SMP handshake |
subslice N |
per subscription: SUBMODE=sub (SUBBATCH, SUBKEEP), new, newthensub |
load |
server CPU per operation under a fixed mixed workload |
ntfloop, ntfdeliver |
notification store and its delivery loop |
mesh |
proxy and relay in separate processes: relay sessions, forwards in flight, forward rate (MESH=) |
pfwdbig |
oversized forwards left in the proxy's sentCommands |
cabal build -fserver_postgres exe:smp-mem-bench
BENCHID=1 PROD_CONNS=2000 PROD_QUEUES=20 PROD_NEW=2 PROD_SEC=30 \
$(cabal list-bin -fserver_postgres exe:smp-mem-bench) prodmix 0 \
+RTS -N -F1.2 -A16m -I0.01 -Iw15 -T -ki2k -RTS
BENCHID offsets ports and the database, so runs do not collide with each other or with tests.
Results
prodmix, live heap per connection after subscribing and peak memory in use under traffic. Each row
adds to the previous one.
| Change | KiB/conn | Peak in use | 40k conns |
|---|---|---|---|
base (sh/fix-leak) |
236.6 | 1075 MiB | 9.0 GiB |
-ki2k |
111.5 | 719 MiB | 4.3 GiB |
connection commits (sh/fix-conn-mem) |
89.8 | 635 MiB | 3.4 GiB |
unpinned subscription keys, proxy commits (sh/mem-combined) |
80.0 | 604-622 MiB | 3.1 GiB |
| patched tls 1.9 | 42.9 | 469 MiB | 1.6 GiB |
Idle connections (conns 2000): 153 KiB at default flags, 68.4 KiB on sh/mem-combined with
-ki2k, 34.8 KiB with the patched tls.
Message throughput did not drop in any row. -ki2k costs some GC time (below).
Findings
1. Thread stacks: 125 KiB per busy connection
Each connection has several threads that block in shallow loops. A thread starts with a 1 KiB stack
(-ki1k). On first overflow the RTS copies up to -kb (1 KiB) of frames into a new 32 KiB chunk
(-kc32k), which with a 1 KiB initial stack moves the loop frame itself, so the thread never returns
to its first chunk and keeps the 32 KiB chunk while the connection lives.
| Flags | idle KiB/conn | prodmix KiB/conn |
|---|---|---|
| default | 153 | 236.6 |
-kc8k |
100 | 135.0 |
-ki2k |
91 | 111.5 |
-ki2k leaves the loop frame in the first chunk, so the overflow chunk is freed on return. Cost, 3
alternating 60 s load runs: GC CPU +15% (4.8-5.2 s to 5.7-5.8 s per run), major GCs +43% (the live
heap is smaller, so -F1.2 triggers major GCs sooner), CPU per operation +0-5% (2.74-3.02 ms to
2.91-3.35 ms, one noisy run). One load run with -kc2k grew to 6.8 GB RSS; not investigated, so
-kc is left at its default.
Branch sh/rts-ki2k sets -with-rtsopts=-ki2k for smp-server. Command-line +RTS flags still apply.
2. Pinned blocks kept alive by small long-lived objects
ByteStrings and crypton's ScrubbedBytes are pinned and share 4 KiB blocks. GHC treats a pinned
block as live if any object in it is live, and treats every object in a live pinned block as alive,
including dead ScrubbedBytes, whose weak pointers and finalizers then stay allocated (GHC 9.6.3
Evac.c:443, GCAux.c:73). A 24-byte value created during request processing, among crypto
temporaries, therefore keeps ~4 KiB and several dead keys alive for as long as it lives. A standalone
program reproduces it (pinned key 2742 B live per key, unpinned 80 B).
Subscriptions. The subscription key is such a value.
subslice 5000 |
before | unpinned keys |
|---|---|---|
| NEW-created | 5.25 KiB | 0.46 KiB |
| one SUB per block | 1.67-2.42 KiB | 0.47 KiB |
| 100 SUBs per block | 0.49 KiB | 0.46 KiB |
| NEW0, then SUB on the same connection | 4.98 KiB | 0.47 KiB |
Branch sh/fix-sub-mem stores the keys of subscriptions, ntfSubscriptions and queueSubscribers
as ShortByteString and shares one key between the maps. CPU unchanged.
TLS session state. tls 1.9 keeps record keys, IVs, TLS 1.3 traffic secrets, verify data and handshake state for the connection's lifetime, all allocated during the handshake. A heap census per connection before and after re-allocating them together at the end of the handshake:
| tls 1.9 | patched | |
|---|---|---|
| weak pointers | ~138 | ~45 |
| memory in pinned blocks, outside the census | ~39 KiB | ~13 KiB |
| idle KiB/conn | 68.4 | 34.8 |
| prodmix KiB/conn | 80.0 | 42.9 |
The patch (tls-1.9-compact-state.patch: Network.TLS.Handshake.Compact, called at the end of
handshake under the context locks) re-derives the TLS 1.3 record keys from the kept traffic
secrets, copies the IVs, secrets, session ID, verify data, randoms, main and resumption secrets and
the transcript hash, and reseeds the context RNG. It needs a tls fork. Not compacted: the TLS 1.2 bulk
key, keys installed later by KeyUpdate. The transcript hash copy coerces crypton's hidden Context
newtype, so it must be checked on crypton upgrades.
Full test suite with the patched tls: 1431 examples, 0 failures. The same run on unpatched tls had 1 failure, the XFTP agent's "should resume sending file after restart".
SMP handshake state (session secret, chain keys) has the same pattern and is not done yet; the remaining ~13 KiB outside the census is likely there.
3. IDs sliced from the received block: 16 KiB per subscription
Parsed corrId and entityId are slices of the ~16 KB decrypted block, and the entity ID was stored
as the subscription key, so one live key kept the block. One SUB per block cost 16.43 KiB per
subscription, one survivor of 100 cost 18.04 KiB. sh/fix-sub-keys copies both IDs in
tDecodeServer (16.43 to 1.67-2.42 KiB). Unpinned keys (finding 2) also remove it for the maps.
4. Connection structure: 21 KiB per busy connection
- tls 1.9 stores the receive record state lazily, so each connection kept its last received 16 KB
record.
recvTLSforces the state after each read. Not needed with tls 2.x. - The send and message-send loops are one thread, and one server-wide thread expires inactive clients instead of one per client: 6 threads per connection become 4.
sh/fix-conn-mem: prodmix 111.5 to 89.8 KiB with -ki2k, 236.6 to 210.5 KiB without.
5. Notification store and its delivery loop
Keys of notifiers whose notifications were delivered stayed in the store, and the loop in
deliverNtfsThread sends every key to getQueueNtfServices (one notifier_id IN ? query read in
full) every 1.5 s.
ntfloop (keys, no traffic) |
allocation | CPU | peak in use |
|---|---|---|---|
| 0 | 0 MiB/s | 0 | 280 MiB |
| 100k | 168 MiB/s | 0.42 cores | 530 MiB |
| 300k | 263 MiB/s | 0.76 cores | 913 MiB |
ntfdeliver with 20k delivered notifications: 20000 keys kept, idle server allocating 36 MiB/s at
0.18 cores; with sh/fix-ntf-store 0 keys, 0.2 MiB/s, 0.01 cores.
6. Proxy mesh
The relay processed forwarded commands from a proxy one at a time in the connection's command loop,
with a PostgreSQL round trip each; all users of a proxy share its connection. Forwards the relay has
not reached wait on the proxy, each with a thread and its ~16 KB block. mesh, 1000 forwards per
second for 20 s, proxy and relay in separate processes, 2 runs each:
| serial | concurrent | |
|---|---|---|
| forwards ok / failed | 19,968 / 0 | 19,968 / 0 |
| proxy live heap at end of sending | 192-200 MiB | 13 MiB |
| proxy memory in use | 653-672 MiB | 325 MiB (idle baseline) |
| proxy threads | 4,026-4,137 | 420 |
Here the backlog drained within the 30 s forward timeout. With a slower relay (a busy database, a remote relay over SOCKS) the queueing delay exceeds it and forwards fail; not reproduced.
sh/fix-proxy-mesh forks each forwarded command under the per-connection concurrency limit, stops
keeping the forwarded block on the proxy while waiting (53.0 to 36.8 KiB per in-flight forward,
relay not answering) and caps forwards in flight per relay at 512 ([PROXY] relay_concurrency).
Forwarded commands use per-command nonces and secrets and are matched by correlation ID, so concurrent
processing is safe; the SimpleX agent keeps one message in flight per queue. The spec requires
responses in order per queue within a connection (simplex-messaging.md, "same order within each
queue ID"); serializing forwarded commands per queue ID is not done yet.
Also in the proxy path (sh/fix-proxy-leak): a PFWD of 16260-16266 bytes from a client that declares
proxyServer = True failed with TELargeMsg before sending and left its 20.2 KiB request in
sentCommands for the life of the relay session (2001 PFWDs, 2001 entries); timed-out requests were
kept too; and a relay session that stopped answering was never dropped (10 of 10 forwards timed out).
7. Concurrency limit and name resolution
forkCmd released its slot when the thread was forked, not when the command finished, so
serverClientConcurrency and serverResolverConcurrency limited nothing: 16 RSLVs with a limit of 4
all reached the resolver, 64 RSLVs from one connection were 64 concurrent requests, and 5 lookups
answered with 502 opened 5 connections because error bodies were not read. sh/fix-rslv-fanout holds
the slot until completion, adds a global limit ([NAMES] resolver_global_concurrency, 32) and reads
error bodies. With the limit working, a client that hits it blocks its own command loop.
8. Smaller findings
- Active-queue statistics are
IntSets at 64 B per element (15M elements, 916 MiB for one set), six of them, reset only whenlog_statsis on. - Control port
savecloses the PostgreSQL pool while the server keeps running; every later DB operation blocks (cpsavephase: NEW aftersavegets no response).
9. tls 2.x
tls 2.1.6 (with tls-session-manager 0.0.6, crypton-connection 0.4.3, http-client-tls 0.3.6.4, an
index-state bump and version pins) saves 2-5 KiB per connection: prodmix 75.4 KiB against 80.0.
Problems found:
- After a failed handshake a tls 2 client and a tls 2 server wait for each other in
bye; this hungtestServerMultipleIdentities. Fixed onsh/tls2by closing the context withoutbye. - A tls 2 client that receives a CertificateRequest runs a timed read inside
handshake; when the timeout fires in the middle of a record the stream fails ("bad record mac") or hangs. SMP servers always request client certificates. 1-2 failures per 2000 concurrent connections; present in tls 2.4.3. Aconns 2000run with tls 2 on both sides stopped at 53 connections.
tls 2.x is not recommended; sh/tls2 is kept for reference.
Open
- Serialize forwarded commands per queue ID on the relay (spec response order).
- A client that sends PFWDs and never reads the responses costs the proxy 6.3 MiB live (11.4 MiB in use) until the inactive-client expiry (up to 6 hours), 6 GiB per 1000 such clients.
- PRXY to many aliases of one relay opens an unbounded number of relay sessions, each 207 KiB on the proxy and 149 KiB on the relay: 2.0 GiB and 1.4 GiB per 10k.
- Compact the SMP handshake state after the handshake, as in finding 2.
- Publish the tls fork and reference it from
cabal.project. - A SEND right after NEW or SUB can find no subscriber yet (
queueSubscribersis updated throughsubQ), so the message waits for the next SUB or ACK.deliverIfSamemay leave a subscription without a delivery thread. Both seen as occasional stuck steps inload; not memory. [PROXY] relay_concurrencyand[NAMES] resolver_global_concurrencyare new INI keys.
Branches
| Branch | Base | Content |
|---|---|---|
sh/fix-ntf-store |
master | finding 5 |
sh/fix-sub-keys |
master | finding 3 |
sh/fix-proxy-leak |
master | finding 6, request leak and stuck session |
sh/fix-rslv-fanout |
master | finding 7 |
sh/fix-conn-mem |
master | finding 4 |
sh/fix-sub-mem |
master | finding 2, subscription keys |
sh/fix-proxy-mesh |
sh/fix-proxy-leak |
finding 6, relay concurrency, per-relay limit |
sh/rts-ki2k |
master | finding 1 |
sh/mem-combined |
sh/fix-leak |
findings 1-6 together, as measured in Results (local) |
sh/tls2 |
sh/mem-combined |
finding 9, not for merge (local) |