Spawning from @Twi1ightSparkle
[spotting](https://matrix.to/#/!SGNQGPGUwtcPBUotTL:matrix.org/$vbqaJ6fJJBzMe1972Z_Ny8euVPWV8IxIS3YWGxP58r0?via=jki.re&via=element.io&via=matrix.org)
a fresh Synapse [running this query every 5
seconds](https://github.com/element-hq/synapse/blob/373fa7f542d86c1dbf82c4ae87c6bec8390e263a/synapse/storage/databases/main/client_ips.py#L440-L441)
which seemed excessive. My initial sniff test thought it was fine
because the more often we run the query, the smaller number of rows we
need to process at a time but @reivilibre brought up that we still have
to scan over all of the dead tuples each time which is a fixed cost
regardless.
Ideally, we'd instead fix the pagination of the query itself to avoid
re-scanning over the tuples. That's probably also a simple change and
I'm mostly just opening this PR to have a place to chuck my worked
example somewhere (mostly for myself). We can always have another
follow-up to do the proper fix.
### `user_ips` dead tuple calculations from `matrix.org`
For `matrix.org`, our autovacuum triggers after the table has
[~5%](https://github.com/matrix-org/matrix-ansible-private/blob/1b623c950e6b4db87f1dd17fa82c83be1b3b58cb/roles/postgres_role/templates/matrix-postgresql.conf.j2#L64-L66)
dead tuples (Postgres normally has a 20% default for
[`autovacuum_vacuum_scale_factor`](https://www.postgresql.org/docs/current/runtime-config-vacuum.html#GUC-AUTOVACUUM-VACUUM-SCALE-FACTOR))
The `user_ips` table on `matrix.org` has 21.8M rows so that means we
have to wait for `(21.8M * 0.05)` = ~1.1M dead rows to accumulate before
the autovacuum kicks in and cleans up all of the dead tuples.
Upper bound napkin math: If we assume that we've reached a steady state
where we prune just as many rows as we insert over the
`user_ips_max_age` time period (defaults to [28
days](https://github.com/element-hq/synapse/blob/373fa7f542d86c1dbf82c4ae87c6bec8390e263a/synapse/config/server.py#L697));
and if the vacuum only ran because of the prune: `21.8M / 28d ~= 778k` a
day -> takes ~1.4 days to accumulate enough dead tuples.
But most of the churn probably comes from updates to the `user_ips`
since every authenticated request updates `user_ips` [every 2
minutes](https://github.com/element-hq/synapse/blob/373fa7f542d86c1dbf82c4ae87c6bec8390e263a/synapse/storage/databases/main/client_ips.py#L52-L55).
And this matches reality:
For actual metrics of `user_ips` on `matrix.org` looking at the
[Postgres
metrics](https://grafana.matrix.org/d/000000009/postgres?orgId=1&var-data_source=000000001)
we have in Prometheus/Grafana:
- ~6.2M updates per day
(`increase(pg_stat_all_tables_n_tup_upd{schemaname!~"pg_.*",
schemaname!~"information_.*", instance=~"$instance"}[1d])`)
- ~792k inserts per day
(`increase(pg_stat_all_tables_n_tup_ins{schemaname!~"pg_.*",
schemaname!~"information_.*", instance=~"$instance"}[1d])`)
- (inserts don't create dead tuples)
- ~792k deletes per day
(`increase(pg_stat_all_tables_n_tup_del{schemaname!~"pg_.*",
schemaname!~"information_.*", instance=~"$instance"}[1d])`)
- 9.17 deletes/second
-> ~7M dead tuples per day
So we wait ~3.77 hours to trigger the next autovacuum for this table
`(1.1M * (7M / 24))`. Since the majority of the dead tuples are from
updates, those dead tuples on the other side of the `last_seen` index
and we probably don't have to scan over those for the prune loop. In
between vacuums, we still end up scanning ~124k dead tuples each time we
query on the prune though `(792k * (3.7/24))`.
So even for `matrix.org` levels of busyness and more aggressive
autovacuum, the `5s` interval overkill.
Given 9.17 deletes/second, if we choose a prune loop interval duration
of `60s`, it will pick-up ~550 rows which is under the query `LIMIT` set
(`5000`) with about an order of magnitude head-room to catch-up from
downtime or peak/heavy traffic.
Part of: MSC4354,
Part of: https://github.com/element-hq/synapse/issues/19641
This PR persists events in redacted form when they already have valid
redactions. That's not directly a sticky events-specific change (and I
feel it makes sense to redact immediately before persistence when we
already have the redaction).
This causes a nice effect which is that the event's stickiness is lost
immediately before persistence, because the `msc4354_sticky` field is
redacted by the redaction algorithm.
---------
Signed-off-by: Olivier 'reivilibre <oliverw@matrix.org>
Co-authored-by: Eric Eastwood <erice@element.io>
To use `reactor.callFromThread` we need to take the GIL. However, the
vast majority of use-cases on the Rust side involve using that from
Tokio reactor threads, where we do not want to take the GIL (as it can
potentially block for a long time).
In #20252 we added a pure Rust alternative to `callFromThread`.
`run_python_awaitable` is the last usage of `callFromThread`, and so we
replace it with `dispatch_to_twisted`.
---
The diff is a lot smaller if you hide pure whitespace changes.
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Our admin API to delete a room has a weird feature where you can supply
`new_room_user_id` and it will create a new room and join all of the
users to that room:
https://github.com/element-hq/synapse/blob/929e3524d189bf7d3402866b930cebde6bb54892/docs/admin_api/rooms.md#L640-L642
But that room can contain [suspended
users](https://spec.matrix.org/v1.19/client-server-api/#account-suspension)
which are unable to join new rooms. Currently, we use
`logger.exception(...)` for this failure which ends up in Sentry but
since this is probably expected outcome and we shouldn't elevate it to
that level of problem as it's just noise.
Spawning from seeing the error in Sentry,
https://sentry.tools.element.io/organizations/element/issues/11266279/?project=2
```
Traceback (most recent call last):
File "synapse/handlers/room.py", line 2532, in shutdown_room
await self.room_member_handler.update_membership(
File "synapse/handlers/room_member.py", line 675, in update_membership
result = await self.update_membership_locked(
File "synapse/handlers/room_member.py", line 794, in update_membership_locked
raise SynapseError(
SynapseError: 403: Joining rooms while account is suspended is not allowed.
```
Fixes#15871
### What was broken
Sending `{"auth": null}` to any endpoint that requires user-interactive
auth (for example `POST /keys/device_signing/upload`, `DELETE
/devices/{deviceId}` or `POST /register`) returned a 500 instead of the
usual 401 with the available flows. A non-object `auth` on endpoints
that do not validate the body with a model (such as `/register`) could
also 500, e.g. `{"auth": ["session"]}`.
### Root cause
`AuthHandler.check_ui_auth` did `clientdict.pop("auth", {})`, so the
`{}` default only applies when the key is missing. An explicit `null`
came through as `None` and the following `"session" in authdict` raised
`TypeError`. `get_session_id` had the same assumption that `auth` is a
dict.
### Pull Request Checklist
* [x] Pull request is based on the develop branch
* [x] Pull request includes a [changelog
file](https://element-hq.github.io/synapse/latest/development/contributing_guide.html#changelog).
* [x] [Code
style](https://element-hq.github.io/synapse/latest/code_style.html) is
correct (run the
[linters](https://element-hq.github.io/synapse/latest/development/contributing_guide.html#run-the-linters))
---------
Signed-off-by: Ankit Jha <ankit.jha@tradomate.one>
Spawning from
https://github.com/element-hq/synapse/pull/20143#discussion_r4080708454
`RestServlet.register` read `PATTERNS` via `getattr`, so it was typed as
`Any` and mypy never checked subclasses' patterns against what
`HttpServer.register_paths` expects.
This declares the attribute so that classes extending `RestServlet`
provide the right type for the `PATTERNS` attribute.
Add a metric for the number of times we see a stall due to being given a
future token.
This would have caught
https://github.com/element-hq/synapse/issues/20080, where lots of sync
streams got stuck due to a stream not getting replicated correctly.
## Summary
Clarify the `--exists-ok` command-line help text for
`register_new_matrix_user`.
The previous help text said:
> Do not fail if user already exists.
This could be interpreted as the existing user's account being updated.
The new wording explicitly states that the existing user account will
not be updated.
This also aligns the CLI help text with the existing Debian man page
documentation.
## Issue
Fixes#20112
## Testing
* `git diff --check`
* Reviewed the staged diff to confirm only the intended help-text change
was included.
Tokio tasks currently complete by taking the GIL and calling
`reactor.callFromThread`. Taking the GIL may block for a period of time,
and we do not want that to happen on the tokio reactor threads (as that
can block other work from happening).
To avoid this, we instead add a work queue that we can add to from Rust
without taking the GIL, which is drained by the reactor. We signal to
the Twisted reactor that it should wake up by using a unix socket pair,
which is exactly how `reactor.callFromThread` works.
---
The self-pipe trick is basically where you create a pair of unix sockets
connected to each other. One end is added to the reactor so the reactor
is woken up when there are bytes to read, and when another thread needs
to wake up the reactor it just needs to write a byte into the other unix
socket.
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Andrew Morgan <1342360+anoadragon453@users.noreply.github.com>
This pull request fixes a few issues with the profile update stream,
when a user leaves a room. This is basically just mimicking what we
already had for when someone leaves a room, and they no longer share any
rooms, but in reverse - ie when we leave a room, and no longer share
rooms with some users. This was missed in the implementation when adding
the profile update stream for legacy and sliding sync.
[Fix missing profile update stream rows on user leaving
room](https://github.com/element-hq/synapse/commit/e953322226731a99718ebe13139b9c5f05ae4b62)
When a user leaves a room, profile update stream rows are generated with
the `LEFT_ROOM` action for each user in the room that no longer shares a
room with the user who left the room.
This also needs to happen in reverse. The user who left the room needs
to have a profile update stream row for each user they no longer share a
room with.
If we don't do this, clients may keep stale data around even after they
don't share a room with a user, which may mean they don't know when to
refetch profiles after re-joining a room with the stale profile data
user.
[Clear out old profile update stream rows when leaving a
room](https://github.com/element-hq/synapse/commit/e7642bf141b098c1eefaf74655e5e7e18415500b)
When we leave a room, ensure profile update stream rows are cleared out
for every user we no longer share a room with. This is the same as what
happens when someone else leaves a room, but in reverse. The only
remaining profile update stream row should be the `LEFT_ROOM` action.
### Pull Request Checklist
<!-- Please read
https://element-hq.github.io/synapse/latest/development/contributing_guide.html
before submitting your pull request -->
* [x] Pull request is based on the develop branch
* [x] Pull request includes a [changelog
file](https://element-hq.github.io/synapse/latest/development/contributing_guide.html#changelog).
The entry should:
- Be a short description of your change which makes sense to users.
"Fixed a bug that prevented receiving messages from other servers."
instead of "Moved X method from `EventStore` to `EventWorkerStore`.".
- Use markdown where necessary, mostly for `code blocks`.
- End with either a period (.) or an exclamation mark (!).
- Start with a capital letter.
- Feel free to credit yourself, by adding a sentence "Contributed by
@github_username." or "Contributed by [Your Name]." to the end of the
entry.
* [x] [Code
style](https://element-hq.github.io/synapse/latest/code_style.html) is
correct (run the
[linters](https://element-hq.github.io/synapse/latest/development/contributing_guide.html#run-the-linters))
This is to prevent the case where URL previews return a MXC that
immediately 404s because the media in question has been quarantined.
This is mostly to help implementations which currently show an ugly
empty preview due to the MXC being sent down, despite being invalid.
### Pull Request Checklist
<!-- Please read
https://element-hq.github.io/synapse/latest/development/contributing_guide.html
before submitting your pull request -->
* [x] Pull request is based on the develop branch
* [x] Pull request includes a [changelog
file](https://element-hq.github.io/synapse/latest/development/contributing_guide.html#changelog).
The entry should:
- Be a short description of your change which makes sense to users.
"Fixed a bug that prevented receiving messages from other servers."
instead of "Moved X method from `EventStore` to `EventWorkerStore`.".
- Use markdown where necessary, mostly for `code blocks`.
- End with either a period (.) or an exclamation mark (!).
- Start with a capital letter.
- Feel free to credit yourself, by adding a sentence "Contributed by
@github_username." or "Contributed by [Your Name]." to the end of the
entry.
* [ ] [Code
style](https://element-hq.github.io/synapse/latest/code_style.html) is
correct (run the
[linters](https://element-hq.github.io/synapse/latest/development/contributing_guide.html#run-the-linters))
Previously, the worker endpoint pattern included only the multi-item
lookup and `/restart`. Adjust the pattern so that it includes the
per-`delay_id` lookup endpoint as well.
Also adjust the worker documentation to use this same endpoint pattern.
Follow-up of #20210
[MSC4140](https://github.com/matrix-org/matrix-spec-proposals/blob/main/proposals/4140-delayed-events-futures.md),
"Account deactivation":
> when an account is deactivated, the homeserver MUST cancel that
account's delayed events which have not yet been added to a room's event
DAG. These cancelled records MAY be removed immediately, including their
stored event content, as an exception to the usual finalised-record
retention policy.
Deactivation currently leaves the user's scheduled delayed events,
including their content, in place, and they are still attempted at their
scheduled time:
| Delayed event scheduled by the user | What happens after deactivation
today |
| --- | --- |
| Message or state event, fires after deactivation has made the user
leave the room | The send fails with a 403 (user not in room) and the
record is dropped |
| Message or state event, fires before deactivation has made the user
leave the room (leaving happens room by room in the background) | The
event is sent |
| `m.room.member` join for themselves in a public room | The deactivated
user re-joins the room |
### What changes
- Deactivating an account (client or admin API) removes the user's
unsent delayed events as its first step, and re-arms the send timer for
whatever is scheduled next.
- Deactivation can run on a worker, while the send timer lives on the
main process, so the cancellation goes through a new replication
request.
- A delayed event whose send has already started is left alone, as with
a normal cancel: it is already on its way into the DAG, and the send
path removes its record itself.
Suspension and locking are unchanged.
### After #19038
Today a cancelled delayed event is simply deleted, so this PR deletes
the user's records too. #19038 changes cancellation to keep the record
and mark it as cancelled ("finalised"), so that clients can look up what
happened to a delayed event. Once it has landed, deactivation should
probably finalise the user's records as cancelled the same way, rather
than delete them. The MSC allows either: it permits removing the records
immediately, content included, as an exception to the usual retention of
finalised records, which is worth doing at least when the user asks to
be erased.
### Pull Request Checklist
<!-- Please read
https://element-hq.github.io/synapse/latest/development/contributing_guide.html
before submitting your pull request -->
* [x] Pull request is based on the develop branch
* [x] Pull request includes a [changelog
file](https://element-hq.github.io/synapse/latest/development/contributing_guide.html#changelog).
The entry should:
- Be a short description of your change which makes sense to users.
"Fixed a bug that prevented receiving messages from other servers."
instead of "Moved X method from `EventStore` to `EventWorkerStore`.".
- Use markdown where necessary, mostly for `code blocks`.
- End with either a period (.) or an exclamation mark (!).
- Start with a capital letter.
- Feel free to credit yourself, by adding a sentence "Contributed by
@github_username." or "Contributed by [Your Name]." to the end of the
entry.
* [x] [Code
style](https://element-hq.github.io/synapse/latest/code_style.html) is
correct (run the
[linters](https://element-hq.github.io/synapse/latest/development/contributing_guide.html#run-the-linters))
---------
Co-authored-by: Andrew Ferrazzutti <af_0_af@hotmail.com>
Co-authored-by: Andrew Ferrazzutti <andrewf@element.io>
As I was working on https://github.com/element-hq/synapse/pull/20218, I
noticed what seemed to be an illegal configuration of synapse.
- `require_auth_for_profile_requests`: blocks profile requests unless
authenticated
- `limit_profile_requests_to_users_who_share_rooms`: blocks profile
requests unless authenticated user share a room with requested user
This, I think, should be an illegal config:
```
require_auth_for_profile_requests = false
limit_profile_requests_to_users_who_share_rooms = true
```
As of now, with such a config the shared-room check is never applied: a
profile can be requested anonymously, and also by an authenticated user
who doesn't share a room.
### Pull Request Checklist
<!-- Please read
https://element-hq.github.io/synapse/latest/development/contributing_guide.html
before submitting your pull request -->
* [x] Pull request is based on the develop branch
* [x] Pull request includes a [changelog
file](https://element-hq.github.io/synapse/latest/development/contributing_guide.html#changelog).
The entry should:
- Be a short description of your change which makes sense to users.
"Fixed a bug that prevented receiving messages from other servers."
instead of "Moved X method from `EventStore` to `EventWorkerStore`.".
- Use markdown where necessary, mostly for `code blocks`.
- End with either a period (.) or an exclamation mark (!).
- Start with a capital letter.
- Feel free to credit yourself, by adding a sentence "Contributed by
@github_username." or "Contributed by [Your Name]." to the end of the
entry.
* [x] [Code
style](https://element-hq.github.io/synapse/latest/code_style.html) is
correct (run the
[linters](https://element-hq.github.io/synapse/latest/development/contributing_guide.html#run-the-linters))
---------
Co-authored-by: Olivier 'reivilibre' <oliverw@element.io>
The count of unread notifications is computed in two phases:
1. From the summaries (`event_push_summary`) when those are up to date
with the user's last read receipt
2. Otherwise from counting the push actions (`event_push_actions`) after
that receipt
The problem in this process is that the threads whose summary is up to
date are identified with only their `thread_id`. But all main timelines
share the same `"main"` thread_id, so as soon as one room has an
up-to-date summary, phase 2 skips the main timeline of every room.
The fix identifies the up-to-date summaries with their `(room_id,
thread_id)` pair.
Authored by @sandhose. Found while investigating the
`TestThreadedReceipts` Complement flake (#15517, #18537), but it is a
bug on its own.
### Pull Request Checklist
<!-- Please read
https://element-hq.github.io/synapse/latest/development/contributing_guide.html
before submitting your pull request -->
* [x] Pull request is based on the develop branch
* [x] Pull request includes a [changelog
file](https://element-hq.github.io/synapse/latest/development/contributing_guide.html#changelog).
The entry should:
- Be a short description of your change which makes sense to users.
"Fixed a bug that prevented receiving messages from other servers."
instead of "Moved X method from `EventStore` to `EventWorkerStore`.".
- Use markdown where necessary, mostly for `code blocks`.
- End with either a period (.) or an exclamation mark (!).
- Start with a capital letter.
- Feel free to credit yourself, by adding a sentence "Contributed by
@github_username." or "Contributed by [Your Name]." to the end of the
entry.
* [x] [Code
style](https://element-hq.github.io/synapse/latest/code_style.html) is
correct (run the
[linters](https://element-hq.github.io/synapse/latest/development/contributing_guide.html#run-the-linters))
---------
Co-authored-by: Quentin Gliech <quentingliech@gmail.com>
Co-authored-by: Devon Hudson <devonhudson@librem.one>
We sometimes see a lot of `ERROR` logs like the following:
```
Closing scope Scope<... master.write_bytes_to_request> which is not the currently-active one None
```
This error is generated by `opentracing.Scope` checking that it is the
"active" one. Synapse tracks "active" spans via logcontexts, so this
indirectly asserts that the scope is closed in the context it was opened
in. However, the producer methods are often called from the reactor and
therefore withing the sentinel logcontext, which produces the error
above.
This specifically happens when the producer tries to write large
responses but gets paused, and then later resumes.
The fix is to simply use `Span` directly, rather than scopes. `Span`
does not perform the checks.
I noticed this when deploying #19979, though it is unrelated.
### Summary
Resolves#18308.
Makes the `TaskScheduler`'s maximum concurrent running tasks
configurable via `task_scheduler.max_concurrent_tasks` in the homeserver
configuration and defaults it to `2` (reduced from the previous
hardcoded limit of `5`).
### Motivation
Twisted's `adbapi.ConnectionPool` defaults to 5 database connections
(`cp_max=5`). When the `TaskScheduler` previously executed up to 5 tasks
concurrently, it could exhaust the entire database connection pool,
starving regular API and synchronization requests on smaller homeserver
instances. Lowering the default to 2 preserves at least 3 connections
for foreground requests while allowing concurrent tasks to progress, and
gives administrators the ability to tune the limit according to their
host and connection pool sizing.
### Changes
- Added `TaskSchedulerConfig` (`task_scheduler.max_concurrent_tasks`)
with validation (must be a positive integer) and registered it in
`HomeServerConfig`.
- Updated `TaskScheduler` to enforce the configured concurrency limit
and updated class default to 2.
- Updated config documentation and JSON schema.
- Added config tests in `tests/config/test_task_scheduler_config.py` and
updated unit tests in `tests/util/test_task_scheduler.py`.
- Added newsfragment in `changelog.d/18308.feature`.
---------
Signed-off-by: nevil06 <nevilansondsouza@gmail.com>
Co-authored-by: Devon Hudson <devonhudson@librem.one>
When a stream advances in the database but stops being replicated to a
process, that process's view of the stream freezes. Requests that wait
for it to catch up to a token issued by another worker then time out and
return empty responses indefinitely (see #20080), and nothing exported
said so.
Report `get_current_token` from every ID generator, on every process.
The value is comparable between processes, so a stream that has stopped
reaching one of them shows up as divergence with no client traffic
needed. It is also the position that `wait_for_stream_token` waits on,
so its divergence is the failure itself rather than a proxy for it.
Being a watermark over gapless runs of persisted IDs, it also catches a
single writer of a sharded stream going quiet, which a maximum across
writers would hide behind the writers still being replicated.
Previously the tokio runtime was stashed in a hidden attribute on the
reactor object, installed lazily by whichever Rust code first needed it,
and started via `callWhenRunning`.
Instead, we create a `RustRuntime` (accessible via
`HomeServer.get_rust_runtime()`) that holds any per-reactor Rust state,
such as the tokio runtime. It is constructed lazily on use. Rust
consumers (`HttpClient`, `VersionsHandler`, the Python DB pool wrapper)
now receive the runtime or reactor handle explicitly, and the
`reactor.run()` / manual-startup workarounds in tests are no longer
needed.
We also add helper wrappers in Rust for `Reactor` and `HomeServer` that
exposes the needed functionality.
The aim is to allow us to have a Rust-side clock (mainly to get the
current time), that respects the unit test per-reactor time management.
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This is an attempt to better cache the cases where there are a large
number of extremities to resolve over, which keep slightly changing.
This spawns from seeing issues on matrix.org.
We already have a cache over the exact state groups being resolved.
However, we can do better by caching the inputs into state res (i.e. the
conflicted sets), which are more likely to be constant across repeated
state res in a room. We key this cache based on a sha256 hash, on the
assumption that this will never conflict.
Also includes a commit that removes needless copying of the state.
Ports the logcontext classes to Rust, and gives tokio tasks a captured
logcontext so that work running in (or spawned from) Rust is attributed
to the request that caused it.
1. **Add characterization tests for logcontext error messages and the
filter** — pins the exact `logcontext_error` message shapes, the
abuse-detection code paths and `LoggingContextFilter`'s observable
behaviour, *against the existing Python implementation* (this commit is
green on its own). These are the behavioural contract the port has to
satisfy.
2. **Port `ContextResourceUsage` to a Rust pyclass** — self-contained:
the new `synapse_rust.logcontext` module, the class, its stub and the
re-export.
3. **Move the logcontext storage and `LoggingContext` to Rust** — the
core change; the commit message carries detailed design notes.
Highlights:
- The slot is typed `Option<Py<LoggingContext>>`, with `None`
representing the sentinel. `_Sentinel`/`SENTINEL_CONTEXT` stay pure
Python (unchanged); thin wrappers on
`current_context`/`set_current_context` convert at the boundary, and
pyo3's extraction enforces the type (`TypeError` otherwise).
- The accounting is native: one `getrusage(RUSAGE_THREAD)` read per
switch via libc, inline `stop`/`start` bookkeeping for base
`LoggingContext`s, Python dispatch only for subclasses
(`BackgroundProcessLoggingContext`) so their overrides run. The thread
id comes from `PyThread_get_thread_ident` (the exact
`threading.get_ident()`
value) without calling into Python.
- The hot paths avoid per-operation allocation: names are `Py<PyString>`
(the per-log-record `str(context)`/`server_name` reads are INCREF-only),
error branches materialise strings only when hit.
4. **Attribute Rust-spawned work to the caller's logcontext** —
`create_deferred` captures the caller's context and scopes it onto the
spawned task via a tokio task-local (`LogContextHandle`);
`current_context()` gives the task-local read precedence, so
`LoggingContextFilter`/`pyo3-log` resolve the right context on worker
threads with no per-record stamping. `run_python_awaitable` restores the
captured context (via a `with_logcontext` helper, the Rust
`PreserveLoggingContext`) around Python called back from Rust, so e.g.
`runInteraction` from the Rust `/versions` handler accounts its DB usage
against the right request. Integration tests exercise both guarantees
through real production code paths.
Follow-up work on top of this (separate PR): porting
`BackgroundProcessLoggingContext` natively and removing further `Py<_>`
indirections. The fact that `BackgroundProcessLoggingContext` is a
subclass is what forces some of the warts in this PR: e.g. having to use
`Py<LoggingContext>` everywhere, etc.
We don't try (yet) to make this pure Rust, instead we see this as simply
maintaining the Python logcontext machinery when crossing, rather than
trying to make a Rust equivalent that can be used by pure Rust
dependencies. We probably do want to do that in future, as well as wire
up e.g. CPU recording on Rust side, but that is unnecessary for now.
For state filters that ask for concrete types. This allows us to cache
the common case of asking for a specific type/state key.
I noticed a bunch of queries in the jaeger traces that could be cached.