From c85b1fef9c0f9e49b3b5034391c4ea0d2fa3b3ce Mon Sep 17 00:00:00 2001 From: Eric Eastwood Date: Tue, 29 Sep 2026 12:08:35 -0500 Subject: [PATCH] Prune `user_ips` less frequently (less stress on the database) (#20285) Spawning from @Twi1ightSparkle [spotting](https://matrix.to/#/!SGNQGPGUwtcPBUotTL:matrix.org/$vbqaJ6fJJBzMe1972Z_Ny8euVPWV8IxIS3YWGxP58r0?via=jki.re&via=element.io&via=matrix.org) a fresh Synapse [running this query every 5 seconds](https://github.com/element-hq/synapse/blob/373fa7f542d86c1dbf82c4ae87c6bec8390e263a/synapse/storage/databases/main/client_ips.py#L440-L441) which seemed excessive. My initial sniff test thought it was fine because the more often we run the query, the smaller number of rows we need to process at a time but @reivilibre brought up that we still have to scan over all of the dead tuples each time which is a fixed cost regardless. Ideally, we'd instead fix the pagination of the query itself to avoid re-scanning over the tuples. That's probably also a simple change and I'm mostly just opening this PR to have a place to chuck my worked example somewhere (mostly for myself). We can always have another follow-up to do the proper fix. ### `user_ips` dead tuple calculations from `matrix.org` For `matrix.org`, our autovacuum triggers after the table has [~5%](https://github.com/matrix-org/matrix-ansible-private/blob/1b623c950e6b4db87f1dd17fa82c83be1b3b58cb/roles/postgres_role/templates/matrix-postgresql.conf.j2#L64-L66) dead tuples (Postgres normally has a 20% default for [`autovacuum_vacuum_scale_factor`](https://www.postgresql.org/docs/current/runtime-config-vacuum.html#GUC-AUTOVACUUM-VACUUM-SCALE-FACTOR)) The `user_ips` table on `matrix.org` has 21.8M rows so that means we have to wait for `(21.8M * 0.05)` = ~1.1M dead rows to accumulate before the autovacuum kicks in and cleans up all of the dead tuples. Upper bound napkin math: If we assume that we've reached a steady state where we prune just as many rows as we insert over the `user_ips_max_age` time period (defaults to [28 days](https://github.com/element-hq/synapse/blob/373fa7f542d86c1dbf82c4ae87c6bec8390e263a/synapse/config/server.py#L697)); and if the vacuum only ran because of the prune: `21.8M / 28d ~= 778k` a day -> takes ~1.4 days to accumulate enough dead tuples. But most of the churn probably comes from updates to the `user_ips` since every authenticated request updates `user_ips` [every 2 minutes](https://github.com/element-hq/synapse/blob/373fa7f542d86c1dbf82c4ae87c6bec8390e263a/synapse/storage/databases/main/client_ips.py#L52-L55). And this matches reality: For actual metrics of `user_ips` on `matrix.org` looking at the [Postgres metrics](https://grafana.matrix.org/d/000000009/postgres?orgId=1&var-data_source=000000001) we have in Prometheus/Grafana: - ~6.2M updates per day (`increase(pg_stat_all_tables_n_tup_upd{schemaname!~"pg_.*", schemaname!~"information_.*", instance=~"$instance"}[1d])`) - ~792k inserts per day (`increase(pg_stat_all_tables_n_tup_ins{schemaname!~"pg_.*", schemaname!~"information_.*", instance=~"$instance"}[1d])`) - (inserts don't create dead tuples) - ~792k deletes per day (`increase(pg_stat_all_tables_n_tup_del{schemaname!~"pg_.*", schemaname!~"information_.*", instance=~"$instance"}[1d])`) - 9.17 deletes/second -> ~7M dead tuples per day So we wait ~3.77 hours to trigger the next autovacuum for this table `(1.1M * (7M / 24))`. Since the majority of the dead tuples are from updates, those dead tuples on the other side of the `last_seen` index and we probably don't have to scan over those for the prune loop. In between vacuums, we still end up scanning ~124k dead tuples each time we query on the prune though `(792k * (3.7/24))`. So even for `matrix.org` levels of busyness and more aggressive autovacuum, the `5s` interval overkill. Given 9.17 deletes/second, if we choose a prune loop interval duration of `60s`, it will pick-up ~550 rows which is under the query `LIMIT` set (`5000`) with about an order of magnitude head-room to catch-up from downtime or peak/heavy traffic. --- changelog.d/20285.misc | 1 + synapse/storage/databases/main/client_ips.py | 18 +++++++++++++++++- 2 files changed, 18 insertions(+), 1 deletion(-) create mode 100644 changelog.d/20285.misc diff --git a/changelog.d/20285.misc b/changelog.d/20285.misc new file mode 100644 index 0000000000..3eb3f49f99 --- /dev/null +++ b/changelog.d/20285.misc @@ -0,0 +1 @@ +Prune `user_ips` less frequently (less stress on the database). diff --git a/synapse/storage/databases/main/client_ips.py b/synapse/storage/databases/main/client_ips.py index 7cd3667a2b..f9e36ea8be 100644 --- a/synapse/storage/databases/main/client_ips.py +++ b/synapse/storage/databases/main/client_ips.py @@ -438,7 +438,23 @@ class ClientIpWorkerStore(ClientIpBackgroundUpdateStore, MonthlyActiveUsersWorke ) if hs.config.worker.run_background_tasks and self.user_ips_max_age: - self.clock.looping_call(self._prune_old_user_ips, Duration(seconds=5)) + self.clock.looping_call( + self._prune_old_user_ips, + # Based on a measured value of ~9.17 average deletes/second on the + # `user_ips` table on `matrix.org`. If we prune every 60 seconds, this + # query will pick-up ~550 rows which is under the query `LIMIT` set + # (5000) with about an order of magnitude head-room to catch-up from + # downtime or peak/heavy traffic. On `matrix.org`, the peak delete + # activity on the `user_ips` table is ~2x the daily low. + # + # Running this more often seems good as that means we pick up a smaller + # number of rows each time (less work), therefore less disruptive the + # queries will be. But in practice, the query still has a fixed cost as + # it has to scan through dead tuples left behind by all of the previous + # prunes until the table is vacuumed which isn't free so just run this a + # reasonable amount. + Duration(seconds=60), + ) if self._update_on_this_worker: # This is the designated worker that can write to the client IP