mirror of
https://github.com/element-hq/synapse.git
synced 2026-10-06 05:57:22 +00:00
Prune user_ips less frequently (less stress on the database) (#20285)
Spawning from @Twi1ightSparkle [spotting](https://matrix.to/#/!SGNQGPGUwtcPBUotTL:matrix.org/$vbqaJ6fJJBzMe1972Z_Ny8euVPWV8IxIS3YWGxP58r0?via=jki.re&via=element.io&via=matrix.org) a fresh Synapse [running this query every 5 seconds](https://github.com/element-hq/synapse/blob/373fa7f542d86c1dbf82c4ae87c6bec8390e263a/synapse/storage/databases/main/client_ips.py#L440-L441) which seemed excessive. My initial sniff test thought it was fine because the more often we run the query, the smaller number of rows we need to process at a time but @reivilibre brought up that we still have to scan over all of the dead tuples each time which is a fixed cost regardless. Ideally, we'd instead fix the pagination of the query itself to avoid re-scanning over the tuples. That's probably also a simple change and I'm mostly just opening this PR to have a place to chuck my worked example somewhere (mostly for myself). We can always have another follow-up to do the proper fix. ### `user_ips` dead tuple calculations from `matrix.org` For `matrix.org`, our autovacuum triggers after the table has [~5%](https://github.com/matrix-org/matrix-ansible-private/blob/1b623c950e6b4db87f1dd17fa82c83be1b3b58cb/roles/postgres_role/templates/matrix-postgresql.conf.j2#L64-L66) dead tuples (Postgres normally has a 20% default for [`autovacuum_vacuum_scale_factor`](https://www.postgresql.org/docs/current/runtime-config-vacuum.html#GUC-AUTOVACUUM-VACUUM-SCALE-FACTOR)) The `user_ips` table on `matrix.org` has 21.8M rows so that means we have to wait for `(21.8M * 0.05)` = ~1.1M dead rows to accumulate before the autovacuum kicks in and cleans up all of the dead tuples. Upper bound napkin math: If we assume that we've reached a steady state where we prune just as many rows as we insert over the `user_ips_max_age` time period (defaults to [28 days](https://github.com/element-hq/synapse/blob/373fa7f542d86c1dbf82c4ae87c6bec8390e263a/synapse/config/server.py#L697)); and if the vacuum only ran because of the prune: `21.8M / 28d ~= 778k` a day -> takes ~1.4 days to accumulate enough dead tuples. But most of the churn probably comes from updates to the `user_ips` since every authenticated request updates `user_ips` [every 2 minutes](https://github.com/element-hq/synapse/blob/373fa7f542d86c1dbf82c4ae87c6bec8390e263a/synapse/storage/databases/main/client_ips.py#L52-L55). And this matches reality: For actual metrics of `user_ips` on `matrix.org` looking at the [Postgres metrics](https://grafana.matrix.org/d/000000009/postgres?orgId=1&var-data_source=000000001) we have in Prometheus/Grafana: - ~6.2M updates per day (`increase(pg_stat_all_tables_n_tup_upd{schemaname!~"pg_.*", schemaname!~"information_.*", instance=~"$instance"}[1d])`) - ~792k inserts per day (`increase(pg_stat_all_tables_n_tup_ins{schemaname!~"pg_.*", schemaname!~"information_.*", instance=~"$instance"}[1d])`) - (inserts don't create dead tuples) - ~792k deletes per day (`increase(pg_stat_all_tables_n_tup_del{schemaname!~"pg_.*", schemaname!~"information_.*", instance=~"$instance"}[1d])`) - 9.17 deletes/second -> ~7M dead tuples per day So we wait ~3.77 hours to trigger the next autovacuum for this table `(1.1M * (7M / 24))`. Since the majority of the dead tuples are from updates, those dead tuples on the other side of the `last_seen` index and we probably don't have to scan over those for the prune loop. In between vacuums, we still end up scanning ~124k dead tuples each time we query on the prune though `(792k * (3.7/24))`. So even for `matrix.org` levels of busyness and more aggressive autovacuum, the `5s` interval overkill. Given 9.17 deletes/second, if we choose a prune loop interval duration of `60s`, it will pick-up ~550 rows which is under the query `LIMIT` set (`5000`) with about an order of magnitude head-room to catch-up from downtime or peak/heavy traffic.
This commit is contained in:
@@ -0,0 +1 @@
|
||||
Prune `user_ips` less frequently (less stress on the database).
|
||||
@@ -438,7 +438,23 @@ class ClientIpWorkerStore(ClientIpBackgroundUpdateStore, MonthlyActiveUsersWorke
|
||||
)
|
||||
|
||||
if hs.config.worker.run_background_tasks and self.user_ips_max_age:
|
||||
self.clock.looping_call(self._prune_old_user_ips, Duration(seconds=5))
|
||||
self.clock.looping_call(
|
||||
self._prune_old_user_ips,
|
||||
# Based on a measured value of ~9.17 average deletes/second on the
|
||||
# `user_ips` table on `matrix.org`. If we prune every 60 seconds, this
|
||||
# query will pick-up ~550 rows which is under the query `LIMIT` set
|
||||
# (5000) with about an order of magnitude head-room to catch-up from
|
||||
# downtime or peak/heavy traffic. On `matrix.org`, the peak delete
|
||||
# activity on the `user_ips` table is ~2x the daily low.
|
||||
#
|
||||
# Running this more often seems good as that means we pick up a smaller
|
||||
# number of rows each time (less work), therefore less disruptive the
|
||||
# queries will be. But in practice, the query still has a fixed cost as
|
||||
# it has to scan through dead tuples left behind by all of the previous
|
||||
# prunes until the table is vacuumed which isn't free so just run this a
|
||||
# reasonable amount.
|
||||
Duration(seconds=60),
|
||||
)
|
||||
|
||||
if self._update_on_this_worker:
|
||||
# This is the designated worker that can write to the client IP
|
||||
|
||||
Reference in New Issue
Block a user