Skip to content

P3 benchmark: linearizable authority reads

Result: PASS against the P3 criteria as the design states them. On a real three-member MongoDB 7.0 replica set, using SyRF's own driver and client settings, uncached single-document _id reads with ReadConcern.Linearizable from the primary:

  • worked without error;
  • finished or failed within the 5-second deadline in every scenario, with no hangs;
  • never returned a revoked role after the revoke was majority-acknowledged, including from a partitioned primary that still believed it was primary;
  • recovered on their own after every fault.

Local latency at the realistic concurrency (96 in-flight reads) was p50 5.8–5.9 ms and p99 25.5–26.0 ms.

This benchmark closed the last M0 gate before M1a. The design sets no failover-recovery target: a hard primary loss costs about 10–12 s of denied authority decisions, and each read in that window waits the full 5 s before denying. Plain majority reads behave the same way, so the cost comes from the topology, not from linearizable reads. P3 was accepted on 2026-09-30 under P5 (see Decision); M1a uses this reader.

Test and tooling only. No product runtime code changed, so no runtime feature flag is involved.

What was measured

The candidate reader is the design's proposal, implemented in e2e/authority/benchmark/P3Benchmark/Program.cs (TimedReadAsync):

Aspect Setting
Query Find({ _id: CSUUID }), Limit(1): one pmInvestigator-shaped document by immutable _id
Read ReadPreference.Primary, ReadConcern.Linearizable; no causally consistent session; no cache
Deadline maxTimeMS = 5000 and a client CancellationToken at 5 s, so server selection and connection waits are bounded too. maxTimeMS alone does not bound them: the driver's default server-selection timeout is 30 s
Writes Role toggles (Revision, ApplicationRoles) with WriteConcern.WMajority
Client Built through SyRF.Mongo.Common.MongoContext (the services' own client factory): MongoDB.Driver 3.10.0, CSUUID serializer registration and diagnostics subscriber, driver defaults otherwise (pool 100, retryable reads/writes on, server selection 30 s, heartbeat 10 s)
Data 10,000 synthetic investigators, _id stored as BinData subtype 3 (CSUUID, asserted at seed); 64 "hot" documents whose role is toggled continuously
Comparison (information only) The same read with ReadConcern.Majority, primary. The design forbids it as a fallback

Concurrency: 3 clients × 32 in-flight reads = 96. Production API scales between HPA minReplicas: 3 and maxReplicas: 10 (cluster-gitops syrf/environments/production/api/values.yaml). Each replica has one MongoClient, and in the design each protected request makes one to three sequential single-document authority reads. Thirty-two concurrent in-flight authority reads per replica is well above SyRF's research-team traffic, while staying inside the driver's default pool of 100. The benchmark therefore runs one client per replica at the HPA minimum (3 × 32 = 96) as the realistic level. It also runs the HPA maximum (10 × 32 = 320) as a stress level, and a serial single read (1 × 1) as the unloaded round trip.

Scenarios

Scenario How the fault is made What it checks
Latency 30 s measured after a 5 s warm-up per profile; two repeats; light background majority writes (~22–37/s) p50/p95/p99/p99.9/max; errors
Graceful step-down replSetStepDown on the primary Reads in flight across an orderly election
Primary killed docker kill -s KILL on the primary Hard failure: reads to a dead primary
Primary partitioned docker network disconnect of the primary from the replica-set network The primary is unreachable to its peers and to the client, with no TCP reset
No majority docker pause of both secondaries for 20 s, then unpause The primary cannot majority-commit, so linearizable reads cannot be confirmed
Stale-read check (all fault runs) Majority-acknowledged grant and revoke toggles on the hot documents, while readers read them (half of all reads) A read that started after a revoke was acknowledged must never return an earlier revision
Dual primary, ×3 Partition the primary P from its peers only (P stays reachable on a side network), replSetStepUp a secondary S, majority-acknowledge a revoke on S, then read from P directly The stale-primary hazard in MongoDB's contract. For this check only, electionTimeoutMillis is raised to 30 s, the adversarial condition: P keeps believing it is primary for longer

Each fault run has 10 s of load before the fault and 40 s after it. Each read has a watchdog: a read still running 1 s after its deadline counts as a hang, whether or not the driver honours cancellation. A read that completes more than 100 ms after its deadline counts as over-deadline.

Command and environment

bash e2e/authority/benchmark/run.sh     # ~17 min; tears the stack down on every exit
Item Value
Recorded run 2026-09-30, harness commit 3e98d52e5, tree bcf993d1 (git rev-parse <commit>:e2e/authority/benchmark), exit 0, verdict PASS
Replica set 3 members, mongo:7.0 image (mongod 7.0.40), default settings: electionTimeoutMillis 10000, heartbeat 2000 ms, writeConcernMajorityJournalDefault true, default write concern majority; WiredTiger cache 0.25 GB, mem_limit 768 MiB per member, data in tmpfs, no CPU quota
Client The benchmark container on the same internal networks (self-contained .NET 10 publish), 48 visible CPUs
Host Intel Xeon Platinum 8260 @ 2.40 GHz, 48 cores, 187 GB RAM, Docker 29.6.0, Linux 6.8. This host is also a CI runner host; other jobs may have shared it during the run
Footprint Peak ~2.6 GB (members ~750 MiB each at their limit, client ~380 MiB), inside the 4 GB budget. No member was lost to OOM: every scenario began with a verified healthy 1 primary + 2 secondaries
Isolation Compose project syrf-e2e-authority-p3-pr<N>, two internal: true networks, no published port, ownership-labelled containers; teardown verified; docker ps -a before and after identical; syrf-quartz-dev untouched

Raw results: p3-benchmark-results.json is the complete results.json of the recorded run, with per-second failover timelines. Its only change is that error examples are trimmed to the message and the first driver frames. Two earlier full runs also returned PASS, with numbers within about 10%. They ran before error-example capture was added and before the review tightened the gate (every dual-primary iteration conclusive; only a server or deadline refusal counts as a denial).

Results

Latency (no faults)

All values are for successful reads, in milliseconds, from repeats 1 and 2. No read errored in any latency profile: 0 errors, 0 over-deadline and 0 hangs out of 3.38 million reads (1.34 million linearizable).

Reader Concurrency Throughput/s p50 p95 p99 p99.9 max
linearizable 1 340 / 357 2.95 / 2.76 3.9 / 3.8 4.1 / 4.1 7.3 / 7.8 9.2 / 19.4
majority (info) 1 1,957 / 1,932 0.50 / 0.51 0.7 / 0.7 0.8 / 0.8 5.0 / 5.2 22.8 / 20.3
linearizable 96 12,276 / 12,209 5.79 / 5.86 21.8 / 21.3 26.0 / 25.5 30.2 / 29.1 37.6 / 41.9
majority (info) 96 18,280 / 18,295 1.65 / 1.64 23.0 / 22.9 26.9 / 27.0 39.7 / 44.8 75.1 / 99.9
linearizable 320 9,761 / 9,584 34.9 / 35.3 44.9 / 45.6 49.4 / 51.6 65.0 / 67.2 80.8 / 97.7
majority (info) 320 13,980 / 13,611 25.2 / 25.9 36.1 / 37.2 55.6 / 58.3 74.9 / 81.7 128.1 / 194.3

A linearizable read costs one extra majority round trip: the server performs a no-op write and waits for majority acknowledgement before answering. Unloaded, that is about +2.3–2.5 ms at p50. Under load the median is higher (5.8 ms against 1.6 ms), but p99 and above are similar or better at 96 and 320 concurrent reads, and throughput is about two-thirds of majority reads. At 3 × 32 the p95 step from ~6 to ~21 ms appears for both read concerns, so it comes from client and host scheduling, not from linearizability.

In the first trial, each member had a 2-CPU quota. That run showed p95s of 91–200 ms for both readers, which disappeared when the quota was removed. A CFS quota throttles in 100 ms periods, which a VM does not do, so the recorded run has no CPU quota.

Majority writes: alone, 4 writers, p50 3.6 / p95 5.0 / p99 6.5 / max 47.6 ms (1,099/s). Under the 96-read load, background role toggles ran at p50 5.8–5.9 / p99 24–25 ms beside linearizable reads, and at p50 9–13 / p99 38–42 ms beside majority reads.

Failover and deadline (3 × 32 readers throughout)

Times are relative to the moment the fault was requested, so they are upper bounds. "Denied window" is when the last failed read ended; "first success" is the first read after that which succeeded.

Fault Reader Reads Failed Max read (ms) Denied window First success Stale violations
Graceful step-down linearizable 523,839 0 223 none +5 ms 0 / 260,776
Graceful step-down majority (info) 867,833 0 124 none +1 ms 0 / 430,356
Primary killed linearizable 257,435 192 5,045 +10.19 s +10.97 s 0 / 126,643
Primary killed majority (info) 728,980 192 5,041 +10.18 s +10.95 s 0 / 360,517
Primary partitioned linearizable 172,378 192 5,069 +10.12 s +12.08 s 0 / 84,279
Primary partitioned majority (info) 694,497 192 5,044 +10.14 s +11.47 s 0 / 344,560
No majority (20 s) linearizable 577,150 384 5,038 +20.33 s +20.45 s 0 / 287,813
No majority (20 s) majority (info) 1,012,173 192 5,023 +20.25 s +21.38 s 0 / 502,719
  • Deadline: 0 over-deadline and 0 hangs in every scenario. The longest read of the whole run was 5,069 ms: the client token fires at 5,000 ms and the driver returns within ~70 ms. Across 1.53 million linearizable reads in fault runs, none exceeded the 5,100 ms bound.
  • Fail closed: every failure is an error, never data: ClientDeadline (the 5 s token), ConnectionError, MongoConnectionPoolPausedException, NotWritablePrimary or InterruptedDueToReplStateChange. The 192/384 failures are one or two 5-second waits per worker: every in-flight read and its successor waits out the deadline while the driver still routes to the lost primary.
  • Recovery: automatic and without restarts. A graceful step-down costs nothing, because retryable reads absorb it. A hard loss costs one election plus driver rediscovery: about 10 s of denials, and the first success within 10.9–12.1 s. With no majority, reads deny for as long as the majority is missing, and succeed within ~1.4 s of it returning. Linearizable and majority reads behave the same way here.

Stale-read safety after a majority-acknowledged revoke

  • Continuous check: 0 violations in 759,511 linearizable reads that started after an acknowledged write to the same document. They were spread over the step-down, kill, partition and no-majority runs. Majority reads through the replica-set client also showed 0 in 1,638,152: with a single client that cannot reach a stale primary, they are not exposed to that hazard.
  • Dual primary: 3 of 3 iterations were conclusive. In each, the partitioned old primary P still reported isWritablePrimary: true after the revoke on the new primary had been majority-acknowledged:
Read of the revoked document on P #1 #2 #3
linearizable denied: ClientDeadline at 5,001 ms denied at 5,001 ms denied at 5,002 ms
majority (info) old role (revision before the revoke), 2.6 ms old role, 2.2 ms old role, 2.9 ms
local (info) old role old role old role
linearizable via the replica-set client new revision, 997 ms (rediscovery) new revision, 94 ms new revision, 998 ms

The gate counts a failed read on P as safe only if the server or the deadline refused it (ClientDeadline, MaxTimeMSExpired, NotWritablePrimary or NodeIsRecovering), so a broken probe cannot pass as a denial. It also requires every iteration to be conclusive.

This is the case the design's "no weaker fallback" rule exists for: a majority read on a stale primary returned the revoked role every time, while the linearizable read denied within the deadline.

P3 criteria and verdict

Criterion (design / plan wording) Evidence Result
Linearizable single-document _id reads supported on the actual topology and driver 1.34 M linearizable latency reads with 0 errors, 1.53 M more in fault runs; MongoDB.Driver 3.10.0 via MongoContext, CSUUID _id; Atlas: see below PASS (local; Atlas by documentation)
Bounded deadline (5 s); failure or timeout denies (503), no hang 0 hangs, 0 over-deadline; max 5,069 ms; every failure is an exception PASS
Partition/failover/timeout deny, no stale allow 0 / 759,511 continuous; 3 / 3 conclusive dual-primary iterations denied, while the majority read returned the old role PASS
Failover acceptable Automatic recovery after all four faults; graceful 0 failures; hard loss ~10–12 s denied PASS on the design's stated criteria; the recovery figure was accepted under P5 on 2026-09-30 (below)
Not "too costly" (return to design review otherwise) 96-concurrent p50 5.8–5.9 ms, p99 ≤ 26.0 ms locally; ≥ 9.5k reads/s at 320 concurrent PASS; no SLO was set, so these become the baseline

Overall: PASS. No weaker fallback is recommended, and none is needed.

Atlas support

  • Production Cluster0 is M20 (dedicated) according to infrastructure reference and required secrets. A read-only metadata probe of the production cluster was not permitted in this session, so the tier and version come from the repository docs, not a live probe.
  • Preview cluster (staging and PR previews): a read-only explain of a count on a non-existent collection, which reads no documents, returned serverInfo.version 7.0.45 from host atlas-…-shard-00-01, a replica-set member. Repository docs record it as M20 (index cleanup analysis).
  • Support: linearizable is a core server read concern on replica-set primaries. MongoDB documents no tier restriction on it: the Atlas free-cluster limits and unsupported commands pages do not mention read concern, and both clusters are dedicated replica sets in any case. It is not yet executed live on Atlas; the first staging deployment of M1a's dark reader should record it.
  • Local versus Atlas caveat. Locally, members share one host over a Docker bridge, with sub-millisecond RTT. Atlas M20 members run on separate 2-vCPU VMs across availability zones, and the API calls them from GKE europe-west2 over the network. Expect each linearizable read to add at least one client→primary RTT and one primary→majority replication RTT that these figures do not include. Treat the absolute latencies as a lower bound, and re-measure p50/p95/p99 on staging in M1a shadow mode. Failover timing is governed by the same election settings (Atlas defaults: electionTimeoutMillis 10 s), so the ~10–12 s figure should carry over. Atlas's rolling maintenance steps primaries down gracefully, which cost nothing here.

Decision

Accepted 2026-09-30 (controller), under the already-approved P5 ("accept temporary unavailability over silently restoring revoked access"). The design's criteria pass. The only open point was that, during an unplanned primary kill or partition, protected operations are denied (retryable 503) for about 10–12 s, and each denied request takes up to 5 s to answer. That is P5's temporary unavailability. The same would be true of any read against the primary; only a cache or secondary reads would avoid it, and the design already rejects both. Atlas planned maintenance uses a graceful step-down, which cost 0 failures in this benchmark. M1a (#3858) implements the reader this benchmark measured.

Follow-ups (not part of this PR)

  1. MongoDB.Driver 3.10.0 exception bug (M1b must handle it). When a majority write gets a write-concern error while the primary loses its majority, the driver throws ArgumentNullException ("connectionId") while constructing MongoBulkWriteOperationException, instead of a MongoException. It reproduced twice per run in the no-majority scenario, in all three full runs and a dedicated diagnostic run. The outcome is still an error, never a false success. M1b's revoke/grant path must treat any exception as an ambiguous outcome needing a fresh read (the design's CAS/409 path), not only MongoException. Report upstream or upgrade the driver separately.
  2. Revoke availability after a partition. In the dual-primary runs, the revoke on the new primary took 10–22 s to be majority-acknowledged in the recorded run (0.2 s in one iteration of an earlier run), while the third member switched its sync source away from the partitioned primary (30 s election timeout in that check only). M1b's role mutation should present such a timeout as "outcome unknown, retry with fresh state", which the design already requires.
  3. Re-measure latency on staging (Atlas M20, cross-zone) with M1a's dark reader; record whether linearizable reads run there.
  4. CI routing for this benchmark stays opt-in, like the M0a fixture lane: ~17 min and ~2.6 GB.

Reproduce and develop

  • bash e2e/authority/benchmark/run.sh: the recorded command. --quick runs short durations for development (about 7 min). --skip-latency, --skip-dual and --failover kill,partition run diagnostic subsets, whose verdict is always INCOMPLETE. --keep leaves the stack up; stop it with bash e2e/authority/benchmark/teardown.sh.
  • Output: e2e/test-results/authority-p3/results.json and bench.log.
  • Pure analysis and gate tests (no Docker): dotnet test e2e/authority/benchmark/P3Benchmark.Tests. They are in neither syrf.sln nor any workflow, so CI does not run them; run them by hand when the harness changes. Guard checks: bash e2e/authority/test-guards.sh.