P3 benchmark: linearizable authority reads¶
Result: PASS against the P3 criteria as the design states
them. On a real three-member MongoDB 7.0 replica set, using SyRF's own driver and client settings,
uncached single-document _id reads with ReadConcern.Linearizable from the primary:
- worked without error;
- finished or failed within the 5-second deadline in every scenario, with no hangs;
- never returned a revoked role after the revoke was majority-acknowledged, including from a partitioned primary that still believed it was primary;
- recovered on their own after every fault.
Local latency at the realistic concurrency (96 in-flight reads) was p50 5.8–5.9 ms and p99 25.5–26.0 ms.
This benchmark closed the last M0 gate before M1a. The design sets no failover-recovery target: a hard primary loss costs about 10–12 s of denied authority decisions, and each read in that window waits the full 5 s before denying. Plain majority reads behave the same way, so the cost comes from the topology, not from linearizable reads. P3 was accepted on 2026-09-30 under P5 (see Decision); M1a uses this reader.
Test and tooling only. No product runtime code changed, so no runtime feature flag is involved.
What was measured¶
The candidate reader is the design's proposal, implemented in
e2e/authority/benchmark/P3Benchmark/Program.cs (TimedReadAsync):
| Aspect | Setting |
|---|---|
| Query | Find({ _id: CSUUID }), Limit(1): one pmInvestigator-shaped document by immutable _id |
| Read | ReadPreference.Primary, ReadConcern.Linearizable; no causally consistent session; no cache |
| Deadline | maxTimeMS = 5000 and a client CancellationToken at 5 s, so server selection and connection waits are bounded too. maxTimeMS alone does not bound them: the driver's default server-selection timeout is 30 s |
| Writes | Role toggles (Revision, ApplicationRoles) with WriteConcern.WMajority |
| Client | Built through SyRF.Mongo.Common.MongoContext (the services' own client factory): MongoDB.Driver 3.10.0, CSUUID serializer registration and diagnostics subscriber, driver defaults otherwise (pool 100, retryable reads/writes on, server selection 30 s, heartbeat 10 s) |
| Data | 10,000 synthetic investigators, _id stored as BinData subtype 3 (CSUUID, asserted at seed); 64 "hot" documents whose role is toggled continuously |
| Comparison (information only) | The same read with ReadConcern.Majority, primary. The design forbids it as a fallback |
Concurrency: 3 clients × 32 in-flight reads = 96. Production API scales between HPA minReplicas: 3 and
maxReplicas: 10 (cluster-gitops syrf/environments/production/api/values.yaml). Each replica has one
MongoClient, and in the design each protected request makes one to three sequential single-document
authority reads. Thirty-two concurrent in-flight authority reads per replica is well above SyRF's
research-team traffic, while staying inside the driver's default pool of 100. The benchmark therefore runs one
client per replica at the HPA minimum (3 × 32 = 96) as the realistic level. It also runs the HPA maximum
(10 × 32 = 320) as a stress level, and a serial single read (1 × 1) as the unloaded round trip.
Scenarios¶
| Scenario | How the fault is made | What it checks |
|---|---|---|
| Latency | 30 s measured after a 5 s warm-up per profile; two repeats; light background majority writes (~22–37/s) | p50/p95/p99/p99.9/max; errors |
| Graceful step-down | replSetStepDown on the primary |
Reads in flight across an orderly election |
| Primary killed | docker kill -s KILL on the primary |
Hard failure: reads to a dead primary |
| Primary partitioned | docker network disconnect of the primary from the replica-set network |
The primary is unreachable to its peers and to the client, with no TCP reset |
| No majority | docker pause of both secondaries for 20 s, then unpause |
The primary cannot majority-commit, so linearizable reads cannot be confirmed |
| Stale-read check (all fault runs) | Majority-acknowledged grant and revoke toggles on the hot documents, while readers read them (half of all reads) | A read that started after a revoke was acknowledged must never return an earlier revision |
| Dual primary, ×3 | Partition the primary P from its peers only (P stays reachable on a side network), replSetStepUp a secondary S, majority-acknowledge a revoke on S, then read from P directly |
The stale-primary hazard in MongoDB's contract. For this check only, electionTimeoutMillis is raised to 30 s, the adversarial condition: P keeps believing it is primary for longer |
Each fault run has 10 s of load before the fault and 40 s after it. Each read has a watchdog: a read still running 1 s after its deadline counts as a hang, whether or not the driver honours cancellation. A read that completes more than 100 ms after its deadline counts as over-deadline.
Command and environment¶
| Item | Value |
|---|---|
| Recorded run | 2026-09-30, harness commit 3e98d52e5, tree bcf993d1 (git rev-parse <commit>:e2e/authority/benchmark), exit 0, verdict PASS |
| Replica set | 3 members, mongo:7.0 image (mongod 7.0.40), default settings: electionTimeoutMillis 10000, heartbeat 2000 ms, writeConcernMajorityJournalDefault true, default write concern majority; WiredTiger cache 0.25 GB, mem_limit 768 MiB per member, data in tmpfs, no CPU quota |
| Client | The benchmark container on the same internal networks (self-contained .NET 10 publish), 48 visible CPUs |
| Host | Intel Xeon Platinum 8260 @ 2.40 GHz, 48 cores, 187 GB RAM, Docker 29.6.0, Linux 6.8. This host is also a CI runner host; other jobs may have shared it during the run |
| Footprint | Peak ~2.6 GB (members ~750 MiB each at their limit, client ~380 MiB), inside the 4 GB budget. No member was lost to OOM: every scenario began with a verified healthy 1 primary + 2 secondaries |
| Isolation | Compose project syrf-e2e-authority-p3-pr<N>, two internal: true networks, no published port, ownership-labelled containers; teardown verified; docker ps -a before and after identical; syrf-quartz-dev untouched |
Raw results: p3-benchmark-results.json is the complete results.json of the
recorded run, with per-second failover timelines. Its only change is that error examples are trimmed to the
message and the first driver frames. Two earlier full runs also returned PASS, with numbers within about 10%.
They ran before error-example capture was added and before the review tightened the gate (every dual-primary
iteration conclusive; only a server or deadline refusal counts as a denial).
Results¶
Latency (no faults)¶
All values are for successful reads, in milliseconds, from repeats 1 and 2. No read errored in any latency profile: 0 errors, 0 over-deadline and 0 hangs out of 3.38 million reads (1.34 million linearizable).
| Reader | Concurrency | Throughput/s | p50 | p95 | p99 | p99.9 | max |
|---|---|---|---|---|---|---|---|
| linearizable | 1 | 340 / 357 | 2.95 / 2.76 | 3.9 / 3.8 | 4.1 / 4.1 | 7.3 / 7.8 | 9.2 / 19.4 |
| majority (info) | 1 | 1,957 / 1,932 | 0.50 / 0.51 | 0.7 / 0.7 | 0.8 / 0.8 | 5.0 / 5.2 | 22.8 / 20.3 |
| linearizable | 96 | 12,276 / 12,209 | 5.79 / 5.86 | 21.8 / 21.3 | 26.0 / 25.5 | 30.2 / 29.1 | 37.6 / 41.9 |
| majority (info) | 96 | 18,280 / 18,295 | 1.65 / 1.64 | 23.0 / 22.9 | 26.9 / 27.0 | 39.7 / 44.8 | 75.1 / 99.9 |
| linearizable | 320 | 9,761 / 9,584 | 34.9 / 35.3 | 44.9 / 45.6 | 49.4 / 51.6 | 65.0 / 67.2 | 80.8 / 97.7 |
| majority (info) | 320 | 13,980 / 13,611 | 25.2 / 25.9 | 36.1 / 37.2 | 55.6 / 58.3 | 74.9 / 81.7 | 128.1 / 194.3 |
A linearizable read costs one extra majority round trip: the server performs a no-op write and waits for majority acknowledgement before answering. Unloaded, that is about +2.3–2.5 ms at p50. Under load the median is higher (5.8 ms against 1.6 ms), but p99 and above are similar or better at 96 and 320 concurrent reads, and throughput is about two-thirds of majority reads. At 3 × 32 the p95 step from ~6 to ~21 ms appears for both read concerns, so it comes from client and host scheduling, not from linearizability.
In the first trial, each member had a 2-CPU quota. That run showed p95s of 91–200 ms for both readers, which disappeared when the quota was removed. A CFS quota throttles in 100 ms periods, which a VM does not do, so the recorded run has no CPU quota.
Majority writes: alone, 4 writers, p50 3.6 / p95 5.0 / p99 6.5 / max 47.6 ms (1,099/s). Under the 96-read load, background role toggles ran at p50 5.8–5.9 / p99 24–25 ms beside linearizable reads, and at p50 9–13 / p99 38–42 ms beside majority reads.
Failover and deadline (3 × 32 readers throughout)¶
Times are relative to the moment the fault was requested, so they are upper bounds. "Denied window" is when the last failed read ended; "first success" is the first read after that which succeeded.
| Fault | Reader | Reads | Failed | Max read (ms) | Denied window | First success | Stale violations |
|---|---|---|---|---|---|---|---|
| Graceful step-down | linearizable | 523,839 | 0 | 223 | none | +5 ms | 0 / 260,776 |
| Graceful step-down | majority (info) | 867,833 | 0 | 124 | none | +1 ms | 0 / 430,356 |
| Primary killed | linearizable | 257,435 | 192 | 5,045 | +10.19 s | +10.97 s | 0 / 126,643 |
| Primary killed | majority (info) | 728,980 | 192 | 5,041 | +10.18 s | +10.95 s | 0 / 360,517 |
| Primary partitioned | linearizable | 172,378 | 192 | 5,069 | +10.12 s | +12.08 s | 0 / 84,279 |
| Primary partitioned | majority (info) | 694,497 | 192 | 5,044 | +10.14 s | +11.47 s | 0 / 344,560 |
| No majority (20 s) | linearizable | 577,150 | 384 | 5,038 | +20.33 s | +20.45 s | 0 / 287,813 |
| No majority (20 s) | majority (info) | 1,012,173 | 192 | 5,023 | +20.25 s | +21.38 s | 0 / 502,719 |
- Deadline: 0 over-deadline and 0 hangs in every scenario. The longest read of the whole run was 5,069 ms: the client token fires at 5,000 ms and the driver returns within ~70 ms. Across 1.53 million linearizable reads in fault runs, none exceeded the 5,100 ms bound.
- Fail closed: every failure is an error, never data:
ClientDeadline(the 5 s token),ConnectionError,MongoConnectionPoolPausedException,NotWritablePrimaryorInterruptedDueToReplStateChange. The 192/384 failures are one or two 5-second waits per worker: every in-flight read and its successor waits out the deadline while the driver still routes to the lost primary. - Recovery: automatic and without restarts. A graceful step-down costs nothing, because retryable reads absorb it. A hard loss costs one election plus driver rediscovery: about 10 s of denials, and the first success within 10.9–12.1 s. With no majority, reads deny for as long as the majority is missing, and succeed within ~1.4 s of it returning. Linearizable and majority reads behave the same way here.
Stale-read safety after a majority-acknowledged revoke¶
- Continuous check: 0 violations in 759,511 linearizable reads that started after an acknowledged write to the same document. They were spread over the step-down, kill, partition and no-majority runs. Majority reads through the replica-set client also showed 0 in 1,638,152: with a single client that cannot reach a stale primary, they are not exposed to that hazard.
- Dual primary: 3 of 3 iterations were conclusive. In each, the partitioned old primary P still
reported
isWritablePrimary: trueafter the revoke on the new primary had been majority-acknowledged:
| Read of the revoked document on P | #1 | #2 | #3 |
|---|---|---|---|
| linearizable | denied: ClientDeadline at 5,001 ms |
denied at 5,001 ms | denied at 5,002 ms |
| majority (info) | old role (revision before the revoke), 2.6 ms | old role, 2.2 ms | old role, 2.9 ms |
| local (info) | old role | old role | old role |
| linearizable via the replica-set client | new revision, 997 ms (rediscovery) | new revision, 94 ms | new revision, 998 ms |
The gate counts a failed read on P as safe only if the server or the deadline refused it (ClientDeadline,
MaxTimeMSExpired, NotWritablePrimary or NodeIsRecovering), so a broken probe cannot pass as a denial.
It also requires every iteration to be conclusive.
This is the case the design's "no weaker fallback" rule exists for: a majority read on a stale primary returned the revoked role every time, while the linearizable read denied within the deadline.
P3 criteria and verdict¶
| Criterion (design / plan wording) | Evidence | Result |
|---|---|---|
Linearizable single-document _id reads supported on the actual topology and driver |
1.34 M linearizable latency reads with 0 errors, 1.53 M more in fault runs; MongoDB.Driver 3.10.0 via MongoContext, CSUUID _id; Atlas: see below |
PASS (local; Atlas by documentation) |
| Bounded deadline (5 s); failure or timeout denies (503), no hang | 0 hangs, 0 over-deadline; max 5,069 ms; every failure is an exception | PASS |
| Partition/failover/timeout deny, no stale allow | 0 / 759,511 continuous; 3 / 3 conclusive dual-primary iterations denied, while the majority read returned the old role | PASS |
| Failover acceptable | Automatic recovery after all four faults; graceful 0 failures; hard loss ~10–12 s denied | PASS on the design's stated criteria; the recovery figure was accepted under P5 on 2026-09-30 (below) |
| Not "too costly" (return to design review otherwise) | 96-concurrent p50 5.8–5.9 ms, p99 ≤ 26.0 ms locally; ≥ 9.5k reads/s at 320 concurrent | PASS; no SLO was set, so these become the baseline |
Overall: PASS. No weaker fallback is recommended, and none is needed.
Atlas support¶
- Production
Cluster0is M20 (dedicated) according to infrastructure reference and required secrets. A read-only metadata probe of the production cluster was not permitted in this session, so the tier and version come from the repository docs, not a live probe. Previewcluster (staging and PR previews): a read-onlyexplainof a count on a non-existent collection, which reads no documents, returnedserverInfo.version7.0.45 from hostatlas-…-shard-00-01, a replica-set member. Repository docs record it as M20 (index cleanup analysis).- Support:
linearizableis a core server read concern on replica-set primaries. MongoDB documents no tier restriction on it: the Atlas free-cluster limits and unsupported commands pages do not mention read concern, and both clusters are dedicated replica sets in any case. It is not yet executed live on Atlas; the first staging deployment of M1a's dark reader should record it. - Local versus Atlas caveat. Locally, members share one host over a Docker bridge, with sub-millisecond
RTT. Atlas M20 members run on separate 2-vCPU VMs across availability zones, and the API calls them from
GKE
europe-west2over the network. Expect each linearizable read to add at least one client→primary RTT and one primary→majority replication RTT that these figures do not include. Treat the absolute latencies as a lower bound, and re-measure p50/p95/p99 on staging in M1a shadow mode. Failover timing is governed by the same election settings (Atlas defaults:electionTimeoutMillis10 s), so the ~10–12 s figure should carry over. Atlas's rolling maintenance steps primaries down gracefully, which cost nothing here.
Decision¶
Accepted 2026-09-30 (controller), under the already-approved P5 ("accept temporary unavailability over silently restoring revoked access"). The design's criteria pass. The only open point was that, during an unplanned primary kill or partition, protected operations are denied (retryable 503) for about 10–12 s, and each denied request takes up to 5 s to answer. That is P5's temporary unavailability. The same would be true of any read against the primary; only a cache or secondary reads would avoid it, and the design already rejects both. Atlas planned maintenance uses a graceful step-down, which cost 0 failures in this benchmark. M1a (#3858) implements the reader this benchmark measured.
Follow-ups (not part of this PR)¶
- MongoDB.Driver 3.10.0 exception bug (M1b must handle it). When a majority write gets a write-concern
error while the primary loses its majority, the driver throws
ArgumentNullException("connectionId") while constructingMongoBulkWriteOperationException, instead of aMongoException. It reproduced twice per run in the no-majority scenario, in all three full runs and a dedicated diagnostic run. The outcome is still an error, never a false success. M1b's revoke/grant path must treat any exception as an ambiguous outcome needing a fresh read (the design's CAS/409 path), not onlyMongoException. Report upstream or upgrade the driver separately. - Revoke availability after a partition. In the dual-primary runs, the revoke on the new primary took 10–22 s to be majority-acknowledged in the recorded run (0.2 s in one iteration of an earlier run), while the third member switched its sync source away from the partitioned primary (30 s election timeout in that check only). M1b's role mutation should present such a timeout as "outcome unknown, retry with fresh state", which the design already requires.
- Re-measure latency on staging (Atlas M20, cross-zone) with M1a's dark reader; record whether linearizable reads run there.
- CI routing for this benchmark stays opt-in, like the M0a fixture lane: ~17 min and ~2.6 GB.
Reproduce and develop¶
bash e2e/authority/benchmark/run.sh: the recorded command.--quickruns short durations for development (about 7 min).--skip-latency,--skip-dualand--failover kill,partitionrun diagnostic subsets, whose verdict is alwaysINCOMPLETE.--keepleaves the stack up; stop it withbash e2e/authority/benchmark/teardown.sh.- Output:
e2e/test-results/authority-p3/results.jsonandbench.log. - Pure analysis and gate tests (no Docker):
dotnet test e2e/authority/benchmark/P3Benchmark.Tests. They are in neithersyrf.slnnor any workflow, so CI does not run them; run them by hand when the harness changes. Guard checks:bash e2e/authority/test-guards.sh.