Commit Graph
25995 Commits
Author SHA1 Message Date
Jason Robinson e9b8fd786b Fix deactivation from a worker that is not profile writer
Thanks sytest 🤩
2026-07-08 11:51:00 +03:00
Jason Robinson cb322e497f Ensure our own syncing user always gets their updates, even if they're not in any rooms 2026-07-08 11:45:41 +03:00
Jason Robinson 2ed4201dbc Adjust a few new tests after merging from develop 2026-07-07 22:52:55 +03:00
Jason Robinson 3180924c5a Update Rust side versions endpoint
Since https://github.com/element-hq/synapse/commit/c63d77a79d7157f26f849684520ba9e99f4d07c0 moved things
2026-07-07 22:48:34 +03:00
Jason Robinson cb81bd3a6b Merge remote-tracking branch 'origin/develop' into anoa/msc4429
# Conflicts:
#	synapse/rest/client/versions.py
2026-07-07 22:29:35 +03:00
Jason Robinson ef9d99515a Dispatch deletion of profile upon deactivation over replication to the right instance, if necessary 2026-07-07 22:19:05 +03:00
Jason Robinson ae4e686f6f Ensure delete_profile_upon_deactivation does profile deletion and profile updates stream updates in a single transaction 2026-07-07 22:06:10 +03:00
Jason Robinson 739824217e Refactor user left room / user joined room profile updates
We now do all the actions relating to the stream updates in one transaction.

Additionally ensure we wake up the notifier for the users concerned.
2026-07-07 21:20:27 +03:00
Eric EastwoodandGitHub c63d77a79d Rust database access via Python database connection pool v2 (#19878)
This is a stepping stone before we can go full Rust everywhere. We're
providing a generic interface as we want database access to work in
Synapse and `synapse-rust-apps`. In `synapse-rust-apps`, we will use a
`tokio-postgres` based database connection pool so it's full Rust.

We want to avoid the situation where we have two database connection
pools (one for Python, one for Rust) as we've run into connection
exhaustion problems on Matrix.org before.

As an example of using it and sanity check for all this work (including
tests), I've also ported over the `/versions` handler to the Rust side
with database access. The `/versions` endpoint is the simplest endpoint
I could find that still had some database access. Hopefully the refactor
on `/versions` isn't that controversial as it's not really the point of
this PR. We can always remove it from this PR but it's just here as a
sanity check that all of this works.


### Why `runInteraction(...)`?

Using the same `runInteraction` pattern that we already have in Synapse
means we can port over existing Synapse code/endpoints without much
thought. But this pattern also makes sense because we want[^1]
transactions to have repeatable-read isolation (easy to think about,
less foot-guns). Having everything thappen in a function callback means
we can do retries for serialization/deadlock errors.

[^1]: To note: Ideally, we'd want the least isolation possible but the
problem is that there is no tooling to yell at you when your
queries/logic is wrong so repeatable-read isolation is a great balance.


> When an application receives this error message, it should abort the
current transaction and retry the whole transaction from the beginning.
The second time through, the transaction will see the
previously-committed change as part of its initial view of the database,
so there is no logical conflict in using the new version of the row as
the starting point for the new transaction's update.
> 
> Note that only updating transactions might need to be retried;
read-only transactions will never have serialization conflicts.
>
> *--
https://www.postgresql.org/docs/current/transaction-iso.html#XACT-REPEATABLE-READ*


As a note, this strategy is less of an impedance mismatch (aligns more
closely) with Synapse so the glue code for the `python_db_pool` should
also be simpler.


### How does this interact with logcontext (`LoggingContext`)?

See [docs on log
contexts](https://github.com/element-hq/synapse/blob/4e9f7757f17ba81b8747b7f8f9646d17df145aa3/docs/log_contexts.md)
for more background.

We already support normal logging from Rust -> Python with `pyo3-log`
and `log` but as soon as we pass a thread boundary, everything is logged
against the `sentinel` log context. Normally, we want logs and CPU/DB
usage correlated with the request that spawned the work.

You can see how I took a stab at fixing this in
https://github.com/element-hq/synapse/pull/19846 by capturing the
logcontext in a Tokio task local and re-activating as necessary. For
example, in that PR, I reactivated the logcontext in
`run_python_awaitable(...)` which we use to call `runInteraction(...)`
from the Rust side which means all of the database usage is correlated
with the request as expected. It also means any `log:info!(...)` done in
`run_interaction(...)` is correlated correctly. But there needs to be a
better story for when you want to log everywhere else.

I haven't explored tracking CPU usage on the Rust side.

I've left all of this out of this PR as I think it will be better to
tackle this as a dedicated follow-up. For example, I'm thinking about
instead creating a new `LoggingContext` with the `parent_context` set to
the calling context and try to avoid needing to call
`set_current_context(...)` on the Python side where possible (like
tracking CPU).



### Testing strategy

Added some tests that exercise some `async` Rust handlers for the
`/versions` endpoint:
```
SYNAPSE_TEST_LOG_LEVEL=INFO poetry run trial tests.rest.client.test_versions.VersionsTestCase
```

Real-world:

 1. `poetry run synapse_homeserver --config-path homeserver.yaml`
 1. `GET http://localhost:8008/_matrix/client/versions`
2026-07-07 13:07:35 -05:00
Jason Robinson d7e0d353ff Add some notes on missing tests for future readers
I changed things so deactivation did call these and no tests failed.
2026-07-07 19:18:23 +03:00
Jason Robinson d0aa8d00fc Revert "Make delete_profile_upon_deactivation first delete the profile fields, to update the profile updates stream"
This reverts commit 1937e764ad.
2026-07-07 19:14:04 +03:00
Jason Robinson 905226577f Remove a few unused methods and fix docstrings 2026-07-07 19:13:50 +03:00
Jason Robinson 1937e764ad Make delete_profile_upon_deactivation first delete the profile fields, to update the profile updates stream 2026-07-07 19:06:52 +03:00
Jason Robinson f7d05558b6 Refactor profile updates behind the profile updates stream writer
Setting a profile field (including displayname / avatar_uri) now needs to happen on a profile updates stream writer. There are new "dispatch" methods to do the profile field set/delete, that handle pushing the update over replication, if required.

Profile field updates and profile updates stream writes are now in the same database transaction.
2026-07-07 18:50:59 +03:00
Olivier 'reivilibre ec6c53e7df Merge branch 'master' into develop 2026-07-07 15:01:22 +01:00
Olivier 'reivilibre be8d900438 1.156.0 v1.156.0 2026-07-07 14:20:45 +01:00
f8fefb09bd Re-expose the standalone format_event_* transforms for backwards compat (#19922)
The port of event serialization to Rust (#19837) removed
`format_event_raw`, `format_event_for_client_v1`,
`format_event_for_client_v2` and
`format_event_for_client_v2_without_room_id` from
`synapse.events.utils`, but there are modules in the wild that import
them from there.

Reimplement them as standalone pyfunctions in Rust, operating directly
on the Python dict so the original semantics are preserved exactly
(in-place mutation, returning the same dict, arbitrary non-JSON values
passing through, KeyError on a missing `unsigned` in the v1 format), and
re-export them from `synapse.events.utils`.

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 13:47:03 +01:00
27c3b5394b Port the synchronous event serialization core to Rust (#19837)
### Summary

This moves the synchronous core of client event serialization out of
`synapse/events/utils.py` and into Rust
(`rust/src/events/serialize.rs`).

Event serialization is on the hot path for `/sync`, `/messages`, and
most client-facing endpoints. Previously it was a recursive pure-Python
routine (`_serialize_event` / `_inject_bundled_aggregations` /
`only_fields`) that interleaved
CPU-bound formatting with `async` DB/IO. This PR separates those two
concerns and moves the CPU-bound half to Rust:

- **Async prep stays in Python.**
`EventClientSerializer._prepare_serialization` does all DB/IO up front:
batch-fetching redaction events and running the registered module
`unsigned`-callback hooks, for the top-level events *and* every bundled
sub-event (edits and thread latest events, which are themselves
serialized). The admin/MSC4354 config is resolved once via
`_update_config`, rather than re-checked on every recursive call as the
old code did.
- **The synchronous core moves to Rust.** Given an event plus the
pre-fetched `redaction_map`, `unsigned_additions`, and
`bundle_aggregations`, the Rust code produces the client JSON entirely
in Rust — including the v1/v2 format transforms, `only_event_fields`
filtering, redaction handling, and recursive bundled aggregations.

### Details

- The Rust entry point is a single batch function, `serialize_events`,
taking a list of `(event, membership)` pairs. The three lookup maps are
shared across the whole batch, so they're read out of Python and
converted to Rust
structures **once** per batch rather than once per event.
`EventClientSerializer.serialize_event` (singular) is a thin wrapper
that calls it with a one-element list.
- `SerializeEventConfig` is now a Rust `pyclass`, and the old
`event_format` callable is replaced by the `EventFormat` enum (`Raw` /
`ClientV1` / `ClientV2` / `ClientV2WithoutRoomId`). Call sites in
`rest/admin/events.py` and
`rest/client/{notifications,room,sync}.py` are updated to pass the enum.
`make_config_for_admin` and MSC4354 enablement now go through
`SerializeEventConfig.for_admin()` / `with_msc4354()`.
- New accessors on `EventInternalMetadata` (`redacted_by`, `txn_id`,
`device_id`, `token_id`, `delay_id`, `soft_failed`,
`policy_server_spammy`) expose to Rust the fields the serializer reads.
- The `_split_field` unit tests move from `tests/events/test_utils.py`
to a Rust test in `serialize.rs`, since the implementation moved.

### Behaviour

This is intended to be a behaviour-preserving refactor — the Rust core
mirrors the previous Python output (field ordering, v1 key promotion,
redaction placement per room version, null-`redacts` handling,
transaction-ID gating).

Existing serialization, relations, and sync tests pass unchanged.

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Eric Eastwood <madlittlemods@gmail.com>
Co-authored-by: Eric Eastwood <erice@element.io>
2026-07-07 11:59:27 +01:00
Eric EastwoodandGitHub 026937b251 Update FakeChannel.await_result(...) to drive async Rust (Tokio runtime/thread pool) (#19879)
Split off from https://github.com/element-hq/synapse/pull/19878 (needed
to drive requests in tests that use async Rust)

Follow-up to https://github.com/element-hq/synapse/pull/19871
2026-07-07 04:07:19 -05:00
Andrew MorganandJason Robinson aa0dbdc1cd Update check_schema_delta.py to better parse raw SQL schema deltas 2026-07-07 10:02:28 +03:00
Jason Robinson 2bea8c0657 Revert "Add a prune job for the profile updates stream tables"
Extracted to a follow-up pr.

This reverts commit b657d2f8
2026-07-07 10:01:24 +03:00
Jason Robinson 15466d2621 Use sha256 instead of md5
To make CI happier, even though we just want to compare content has changed.
2026-07-06 18:42:38 +03:00
Jason Robinson 10cccac913 Ensure we bust the lazy loading cache for profile updates, when field values change 2026-07-06 18:36:57 +03:00
Jason Robinson 73986ef24b Add docstring to get_lazy_loaded_profile_fields_cache 2026-07-06 16:41:49 +03:00
Jason Robinson ac1bdfba17 Clarify clear_profile_updates_for_user docstring 2026-07-06 16:39:00 +03:00
963e486a70 Recreate user profile on account reactivation (#19902)
Co-authored-by: Andrew Morgan <1342360+anoadragon453@users.noreply.github.com>
2026-07-06 12:38:24 +00:00
Jason Robinson f431951fdb Add a test to ensure display name changes don't generate extra room_joined profile updates rows, due to membership event changes 2026-07-06 15:23:11 +03:00
Jason Robinson b571fcef4e Clear all old ProfileUpdate rows from profile_updates_per_user on a leave
Regardless of action. This is cleaner and also ensures a join/leave/join/leave case might not get confused.
2026-07-06 15:08:10 +03:00
5f89ff2b31 Add an index on sliding_sync_connections.last_used_ts (#19912)
Speeds up finding and deleting expired sliding sync connections in
`delete_old_sliding_sync_connections`, which previously required a
sequential scan.

On matrix.org I have a suspicion that this might end up blocking some
SSS connections during the delete, which can take minutes. Specifically,
I think the deletion blocks this delete:


https://github.com/element-hq/synapse/blob/ff19c034d300869e64878a15aed9a97f1cec59e4/synapse/storage/databases/main/sliding_sync.py#L233-L241

We didn't previously have an index because we wanted the postgres HOT
updates, however we also limit the update frequency to once every 5
minutes, so hopefully this is fine.

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 12:59:11 +01:00
Jason Robinson 425aa00a9b Add a join → leave → re-join → leave to try and rule out bugs in that aspect 2026-07-06 14:50:04 +03:00
Andrew MorganandGitHub c7a47bc121 Fix flaky 3PID inhibit error unit tests (#19913) 2026-07-06 11:34:41 +00:00
dependabot[bot]andGitHub ff19c034d3 Bump golang.org/x/net from 0.52.0 to 0.55.0 in /complement (#19910)
Signed-off-by: dependabot[bot] <support@github.com>
2026-07-06 09:08:27 +00:00
Jason Robinson 1955fbee22 Ensure falsey values survive the set_field / sync profile updates -cycle 2026-07-03 18:28:49 +03:00
807795e5ee Change the MSC3814 dehydrated device /events endpoint paging to match spec conventions (#19897)
To match [the spec requirements for
pagination](https://spec.matrix.org/v1.9/appendices/#pagination):

* rename the query parameter of
`/org.matrix.msc3814.v1/dehydrated_device/[device_id]/events` to `from`,
and
* omit `next_batch` from the response content if there are no more
events to fetch

This matches [the changes that are coming in
MSC3814](https://github.com/matrix-org/matrix-spec-proposals/pull/3814/files#r1540896531).

The previous change https://github.com/element-hq/synapse/pull/19896
kept the old POST API to support old clients. This preserves that
compatibility, so the POST API still works how it used to.

Depends on https://github.com/element-hq/synapse/pull/19896
Part of https://github.com/element-hq/element-meta/issues/2799


### Pull Request Checklist

* [x] Pull request is based on the develop branch
* [x] Pull request includes a [changelog
file](https://element-hq.github.io/synapse/latest/development/contributing_guide.html#changelog).
The entry should:
* [x] [Code
style](https://element-hq.github.io/synapse/latest/code_style.html) is
correct (run the
[linters](https://element-hq.github.io/synapse/latest/development/contributing_guide.html#run-the-linters))

---------

Co-authored-by: Richard van der Hoff <1389908+richvdh@users.noreply.github.com>
Co-authored-by: Quentin Gliech <quenting@element.io>
2026-07-03 14:48:43 +02:00
Jason Robinson d3dd12f027 Rewrite lazy_loaded_profile_fields_cache 2026-07-03 15:42:57 +03:00
Hugh Nimmo-SmithandGitHub 4104058e83 Make MSC4388 (sign-in with QR code) PUT requests idempotent (#19808) 2026-07-03 12:09:16 +00:00
Jason Robinson 6f6e633b5c Ensure we don't respond with profile fields the client didn't ask for 2026-07-03 14:58:33 +03:00
Eric EastwoodandGitHub 7da21a715d Update HomeserverTestCase.get_success(...) and friends to drive async Rust (Tokio runtime/thread pool) (#19871)
This means you can use `get_success(...)` anywhere regardless
of what kind of work needs to be done.

Spawning from adding some more async Rust things in
https://github.com/element-hq/synapse/pull/19846 and wanting something
more standard instead of the custom `till_deferred_has_result(...)` that
has crept in to a few files.

Alternative to https://github.com/element-hq/synapse/pull/19867 spurred
on by [this
comment](https://github.com/element-hq/synapse/pull/19867#discussion_r3441774685)
from @erikjohnston


### How does this work?

Previously, `get_success(...)` just ran in a hot-loop advancing the
Twisted reactor clock which didn't give any time for other threads to do
some work or acquire the GIL if necessary (whenever there is a hand-off
from Rust to Python, we need the GIL).

Now, `get_success(...)` loops until we see a result (until we hit the
~0.1s real-time timeout). In the loop, we call
[`time.sleep(0)`](https://docs.python.org/3/library/time.html#time.sleep)
which will "Suspend execution of the calling thread [...]" (CPU and GIL)
to allow other threads to do some work. Then like before, we advance the
Twisted reactor clock to run any scheduled callbacks which includes
anything the other threads may have scheduled.


### Does this slow down the entire test suite?

Seems just as fast as before. There is minutes variance in what we had
before and after but both are within the same range of each other.

(see PR for actual before/after timings)
2026-07-02 15:20:38 -05:00
Jason Robinson 2c7b418105 Ensure we filter out federated users in lazy loading sync when collecting profile updates from timeline events 2026-07-02 23:13:05 +03:00
Jason Robinson 182e4c2d10 Clarify variable names and add comments to ReplicationDataHandler 2026-07-02 22:40:41 +03:00
Jason Robinson 56b7de210e Correctly type ProfileUpdatesStreamRow.field_name 2026-07-02 22:27:18 +03:00
Jason Robinson 868513720f Correctly type ProfileUpdatesStreamRow.action 2026-07-02 22:25:43 +03:00
Jason Robinson e68560a8a2 Don't dump to json string in tests 2026-07-02 22:21:00 +03:00
Olivier 'reivilibreandGitHub 5df6d1be65 Remove legacy MSC3861 auth delegation support (#19895)
Modern 'MAS integration' (as we call it now) is, of course, preserved.

Fixes: #19549

---------

Signed-off-by: Olivier 'reivilibre <oliverw@matrix.org>
2026-07-02 12:25:11 +01:00
Erik JohnstonandGitHub cb60c15761 Fix two bugs in the flag_existing_quarantined_media background update (#19901)
The `flag_existing_quarantined_media` background update (added in
#19558, shipped in v1.152.0) back-populates the
`quarantined_media_changes` table with media that was already
quarantined. It has two bugs.

  ### 1. Some quarantined remote media is silently skipped

  The remote-media query paged through `remote_media_cache` with:

  ```sql
  WHERE quarantined_by IS NOT NULL
      AND media_origin >= ? AND media_id > ?
  ```

This ANDs the two key columns independently rather than comparing them
as a tuple. Once an origin has been fully processed (e.g. `media_id`
reaches `zzz` for origin `a.example`), rows in a *later* origin whose
`media_id` is `<=` the last processed `media_id` (e.g. `b.example` /
`aaa`) fail the `media_id > ?` test and are never flagged.

  Fixed by using a proper row-value tuple comparison:

  ```sql
  WHERE quarantined_by IS NOT NULL
      AND (media_origin, media_id) > (?, ?)
  ```

Both the minimum supported SQLite (3.37.2) and PostgreSQL support
row-value comparisons.

  ### 2. Exhausted queries keep re-running every iteration

`flag_quarantined` ran *both* the local and remote queries on every
iteration. When one table was exhausted but the other still had rows,
the update kept returning a positive count, so the finished table's (now
empty) query needlessly re-ran on every subsequent iteration until the
whole update completed.

This could add significant time to the transaction, as since the rows
were deleted it could scan a significant portion of the table each time.

Fixed by tracking per-table completion (`local_done` / `remote_done`) in
the background-update progress and skipping a table's query once it has
returned an empty batch. The progress dict was also restructured into an
incremental build for readability.
2026-07-02 10:42:44 +01:00
dependabot[bot]andGitHub ab277b3e3f Bump the patches group with 2 updates (#19893)
Signed-off-by: dependabot[bot] <support@github.com>
2026-07-01 17:02:39 +00:00
2517db46b3 add before and after filter to redact user method (/_synapse/admin/v1/user/$user_id/redact) (#19802)
Closes #19441 

---------

Co-authored-by: Olivier 'reivilibre <oliverw@matrix.org>
2026-07-01 17:28:22 +01:00
Olivier 'reivilibreandGitHub acb0ae1a15 Remove wall-clock dependency of test_redact_messages_all_rooms test. (#19890)
Introduced in: #17847

This 10-second wall-clock timeout was troublesome as it fails flakily on
slow/struggling CI runners, like the
default ones for private GitHub repositories.

The loop also silently relied on the reactor advance in `make_request`,
whereas we could just deterministically advance the reactor the known
amount of times
instead.

---------

Signed-off-by: Olivier 'reivilibre <oliverw@matrix.org>
2026-07-01 17:12:48 +01:00
Jason Robinson c7f3b79e75 Use deleted_stream_id instead of indexing on row 2026-07-01 19:05:39 +03:00
Jason Robinson 4934fca946 Clean up some SQL for easier reading 2026-07-01 19:03:29 +03:00