Prune user_ips less frequently (less stress on the database) (#20285)

Spawning from @Twi1ightSparkle
[spotting](https://matrix.to/#/!SGNQGPGUwtcPBUotTL:matrix.org/$vbqaJ6fJJBzMe1972Z_Ny8euVPWV8IxIS3YWGxP58r0?via=jki.re&via=element.io&via=matrix.org)
a fresh Synapse [running this query every 5
seconds](https://github.com/element-hq/synapse/blob/373fa7f542d86c1dbf82c4ae87c6bec8390e263a/synapse/storage/databases/main/client_ips.py#L440-L441)
which seemed excessive. My initial sniff test thought it was fine
because the more often we run the query, the smaller number of rows we
need to process at a time but @reivilibre brought up that we still have
to scan over all of the dead tuples each time which is a fixed cost
regardless.

Ideally, we'd instead fix the pagination of the query itself to avoid
re-scanning over the tuples. That's probably also a simple change and
I'm mostly just opening this PR to have a place to chuck my worked
example somewhere (mostly for myself). We can always have another
follow-up to do the proper fix.


### `user_ips` dead tuple calculations from `matrix.org`

For `matrix.org`, our autovacuum triggers after the table has
[~5%](https://github.com/matrix-org/matrix-ansible-private/blob/1b623c950e6b4db87f1dd17fa82c83be1b3b58cb/roles/postgres_role/templates/matrix-postgresql.conf.j2#L64-L66)
dead tuples (Postgres normally has a 20% default for
[`autovacuum_vacuum_scale_factor`](https://www.postgresql.org/docs/current/runtime-config-vacuum.html#GUC-AUTOVACUUM-VACUUM-SCALE-FACTOR))

The `user_ips` table on `matrix.org` has 21.8M rows so that means we
have to wait for `(21.8M * 0.05)` = ~1.1M dead rows to accumulate before
the autovacuum kicks in and cleans up all of the dead tuples.

Upper bound napkin math: If we assume that we've reached a steady state
where we prune just as many rows as we insert over the
`user_ips_max_age` time period (defaults to [28
days](https://github.com/element-hq/synapse/blob/373fa7f542d86c1dbf82c4ae87c6bec8390e263a/synapse/config/server.py#L697));
and if the vacuum only ran because of the prune: `21.8M / 28d ~= 778k` a
day -> takes ~1.4 days to accumulate enough dead tuples.

But most of the churn probably comes from updates to the `user_ips`
since every authenticated request updates `user_ips` [every 2
minutes](https://github.com/element-hq/synapse/blob/373fa7f542d86c1dbf82c4ae87c6bec8390e263a/synapse/storage/databases/main/client_ips.py#L52-L55).

And this matches reality:

For actual metrics of `user_ips` on `matrix.org` looking at the
[Postgres
metrics](https://grafana.matrix.org/d/000000009/postgres?orgId=1&var-data_source=000000001)
we have in Prometheus/Grafana:

- ~6.2M updates per day
(`increase(pg_stat_all_tables_n_tup_upd{schemaname!~"pg_.*",
schemaname!~"information_.*", instance=~"$instance"}[1d])`)
- ~792k inserts per day
(`increase(pg_stat_all_tables_n_tup_ins{schemaname!~"pg_.*",
schemaname!~"information_.*", instance=~"$instance"}[1d])`)
     - (inserts don't create dead tuples)
- ~792k deletes per day
(`increase(pg_stat_all_tables_n_tup_del{schemaname!~"pg_.*",
schemaname!~"information_.*", instance=~"$instance"}[1d])`)
     - 9.17 deletes/second

-> ~7M dead tuples per day

So we wait ~3.77 hours to trigger the next autovacuum for this table
`(1.1M * (7M / 24))`. Since the majority of the dead tuples are from
updates, those dead tuples on the other side of the `last_seen` index
and we probably don't have to scan over those for the prune loop. In
between vacuums, we still end up scanning ~124k dead tuples each time we
query on the prune though `(792k * (3.7/24))`.

So even for `matrix.org` levels of busyness and more aggressive
autovacuum, the `5s` interval overkill.

Given 9.17 deletes/second, if we choose a prune loop interval duration
of `60s`, it will pick-up ~550 rows which is under the query `LIMIT` set
(`5000`) with about an order of magnitude head-room to catch-up from
downtime or peak/heavy traffic.
This commit is contained in:
Eric Eastwood
2026-09-29 12:08:35 -05:00
committed by GitHub
parent 97700856b2
commit c85b1fef9c
2 changed files with 18 additions and 1 deletions
+1
View File
@@ -0,0 +1 @@
Prune `user_ips` less frequently (less stress on the database).
+17 -1
View File
@@ -438,7 +438,23 @@ class ClientIpWorkerStore(ClientIpBackgroundUpdateStore, MonthlyActiveUsersWorke
)
if hs.config.worker.run_background_tasks and self.user_ips_max_age:
self.clock.looping_call(self._prune_old_user_ips, Duration(seconds=5))
self.clock.looping_call(
self._prune_old_user_ips,
# Based on a measured value of ~9.17 average deletes/second on the
# `user_ips` table on `matrix.org`. If we prune every 60 seconds, this
# query will pick-up ~550 rows which is under the query `LIMIT` set
# (5000) with about an order of magnitude head-room to catch-up from
# downtime or peak/heavy traffic. On `matrix.org`, the peak delete
# activity on the `user_ips` table is ~2x the daily low.
#
# Running this more often seems good as that means we pick up a smaller
# number of rows each time (less work), therefore less disruptive the
# queries will be. But in practice, the query still has a fixed cost as
# it has to scan through dead tuples left behind by all of the previous
# prunes until the table is vacuumed which isn't free so just run this a
# reasonable amount.
Duration(seconds=60),
)
if self._update_on_this_worker:
# This is the designated worker that can write to the client IP