Single-document Study saves with an asynchronous point fold¶
Status: the owner approved the direction on 30 September 2026. Revisions 2-5 (1 October 2026) record the owner's decisions and close four rounds of adversarial review. Build progress, slice by slice, and the as-built deviations are recorded in STATUS. Slices 0 to 2 are merged; slice 3 (fold protocol 2, the reviewer families) is in review in #3922. The decision is recorded in ADR-019.
The design replaces the same-transaction point update for every screening- and annotation-dependent family, behind a new default-off flag. Those families are ProjectScreening, MembershipScreening, ReviewerScreening, StageAnnotation, MembershipStageAnnotation, ReviewerAnnotation, DomainReconciliation and QuestionAnswers. SearchPopulation, and every family with the flag off, keep today's path unchanged.
Code references name symbols rather than lines. They were read on main at 3ffd63394.
Contents¶
- Why
- Owner decisions (1 October 2026)
- Summary of the design
- 1. Pending entry schema
- 2. The save
- 3. The fold worker
- 4. The reader overlay
- 5. Fold mode, the rollback tripwire and version identity
- 6. Interaction with existing mechanisms
- 7. Invariants
- 8. Failure modes
- 9. MVP boundary and acceptance
- 10. Implementation plan
- 11. Deferred items
- 12. Open questions for the owner
- Revision history
Why¶
The idle-host write benchmark of 30 September (PR #3855) closed the evidence gate that the recovered-design reconciliation left open as item 5.
- Command cost. A materialized screening save issues 25 commands inside a multi-document snapshot transaction. A source-only save issues 2. Every cell failed the under-10% p95 gate, by 46% to 1,283%. Even the single-reviewer cells, which have no contention, failed by 575-689%, so the cost is structural.
- Contention. Every save writes four single-per-project documents inside the transaction:
- the control row (three clocks);
- the project-scope current row;
- the publication guard;
- the outbox slots.
The measured contention model showed that any one of them is enough to serialize the path. At ten reviewers on different Studies, 399 of 1,000 materialized submissions exhausted their retries, against zero for source-only.
The annotation, reservation and reviewer-family writers share the same coordinator and the same per-project documents, so the same structural cost applies to them.
Owner decisions (1 October 2026)¶
| # | Decision |
|---|---|
| a | MVP scope. Every screening- and annotation-dependent family migrates in the same MVP. The MVP is done only when every point-saving writer of those families is on the fold path. "Migrated" does not mean "point-maintained" for every family. ReviewerAnnotation and QuestionAnswers stay Stale-on-save: their invalidation travels as an intent in the fold entry, exactly as today's save stales them. In the MVP they are served from the authoritative calculation (scope). Slices land one family group at a time, and each must leave main safe with the flag off (10) |
| b | Write gate. p95 overhead under 10% or at most +2 ms, and zero statistics-caused save failures or exhausted submissions at 1, 2, 5 and 10 reviewers, in both the same-Study and different-Study cells. Ordinary source-level version conflicts that exist without statistics are reported separately and do not count against the gate (9) |
| c | Version bump accepted. The fold bumps the Study's Audit.Version when it removes entries |
| d | Reviewer-tracking mode check accepted. The durable mode is read immediately before the write rather than inside a snapshot. A future mode-switch owner (matrix M15) must add a writer grace period; see the requirement |
| e | Per-stage reviewer tracking deferred. The MVP keeps the durable global mode. Entries whose assumed mode differs from the current mode are discarded and staled, never applied. The per-stage follow-up is #3876 |
| f | Rollback guard. Rollback requires the runbook order (disable, wait for Disabled, then roll back) and a GitOps guard that existing binaries already honour: the fold projects are removed from ProjectStatistics:ProjectAllowlist for the rollback window. See rollback guard |
| g | Staging pilot after slice 2. Enable is allowed on staging after slice 2 (ProjectScreening) behind an explicit staging-only allow setting. Production enable waits for the full MVP (every family and the slice-6 benchmark pass). See enable |
| h | Previous-protocol support (N-1). From production activation onward, readers, the overlay, the fold worker and the rebuild accept entries at the current protocol and the immediately previous one, so an additive protocol upgrade needs no reset or rebuild (a derivation-changing one does). See compatibility |
Summary of the design¶
- Save. The existing single-document Study write also appends a small pending entry: the classified before/after transition, plus any invalidation intents. It uses no statistics transaction and writes no statistics document.
- Fold. A worker in the project-management service holds a per-project lease and turns pending entries into row increments, staling, receipts, delta records, clocks and outbox slots, in batches of up to 32. It removes the applied entries from their Studies. Only the worker runs this transaction.
- Read. The reader serves stored row + derived moves of that project's pending entries, read in the same pinned snapshot. Any pending invalidation intent for a requested scope makes that scope fall back. Values are never behind, and are never shown Fresh while wrong.
- Exactness is arithmetic. Every operation's moves live in exactly one place: the stored row or
the pending set. A rebuild publishes
authoritative(B) - pending(B)from one snapshotB. - Old and lower-protocol code fails closed. A fold-mode project's control carries a
storage-version tripwire that encodes the fold protocol version. Every binary that predates the
fold, and every fold-aware binary at a different protocol, treats the project as incompatible: it
serves nothing current, publishes nothing and repairs nothing for that project. Row
SourceVersionis never overloaded. Old history writers that do not check storage versions are kept away from fold projects by the rollback guard and the enable rollout gate. New ones checkMatchesand the pending set.
1. Pending entry schema¶
Placement on the Study¶
The schema adds two top-level elements to Study (collection pmStudy):
StatisticsFoldSequence int incremented by every fold of this Study; absent = 0 (BsonIgnoreIfDefault)
PendingStatistics document? absent when nothing is pending (BsonIgnoreIfNull)
OldestAtUtc DateTime min(Entries.CreatedAtUtc, Overflow.OverflowedAtUtc); the index key
Entries array bounded (see the cap); each entry carries its own ProtocolVersion
Overflow document? null unless the cap was hit
OverflowedAtUtc DateTime
Dropped array at most 64 { OperationNamespace, OperationId, SubmissionDigest }
UnrecordedCount int saves dropped after Dropped was full
UnrecordedFromRevision int Study revision before the first unrecorded drop
UnrecordedToRevision int Study revision after the last unrecorded drop
The slice-0 contract is unchanged: the field PendingStatistics.OldestAtUtc and the partial index
below are exactly as the pending-index operation builds them.
Every writer maintains the whole sub-document.
- Replace-based writers replace the whole Study, pending set included.
- Update-based writers $set the whole PendingStatistics value. It is computed client-side from the
before-image they hold, or with aggregation expressions in a pipeline update.
- OldestAtUtc is always recomputed as the minimum.
- A bare $push onto PendingStatistics.Entries is forbidden. It would create entries without
OldestAtUtc, which the index, the overlay, the sweep and the rebuild cannot see.
These are persistence state, not domain state. The aggregate carries them as opaque values so that every full-document replace round-trips them. Only the fold-path repository methods append to the pending set, and only the fold removes from it.
StatisticsFoldSequence lets a saver tell that a version change came only from a fold; see
fold-only reloads.
Entry fields¶
| Field | Purpose |
|---|---|
ProtocolVersion |
The fold protocol this entry was written under (5). It is per entry, so a set can hold entries from two protocols during an N-1 window, and each derives under its own rules |
OperationId |
The writer's existing envelope id: screening:{study}:{screener}:{expectedSourceRevision} for a screening save, or {kind}:{study}:{stage}:{session}:{expectedSourceRevision} for an annotation, session-delete or reservation write. It is unique among committed operations because the save suppresses an exact duplicate and fails closed on a changed one before writing (see idempotency). It becomes the receipt and delta RecordId |
OperationNamespace |
stats.source.screening or stats.source.annotation, unchanged. A save that today commits both a screening and an annotation request (ProjectScreeningStatisticsCommit.IncludingAnnotation) appends one entry per namespace in the same write, and every duplicate and unknown-result check examines every namespace the save would write |
ExpectedSourceRevision |
The envelope's expected revision (ProjectStatisticsOperationEnvelope.ExpectedSourceRevision). It is copied into the receipt so that MatchesEnvelope keeps today's replay semantics |
SourceRevisionBefore / SourceRevisionAfter |
The Audit.Version this attempt actually replaced, and the version it wrote. After a rebased retry these may be later than ExpectedSourceRevision, exactly as the receipt allows today |
SubmissionDigest |
The envelope digest (ProjectStatisticsSubmissionDigest.BindScreening or the annotation equivalent). It becomes the receipt's SourceDigest, and duplicate detection compares it |
CreatedAtUtc |
The writing pod's clock. Used only for the age threshold, fold ordering and lag metrics, never for correctness |
Context |
One digest per family group, computed from the Project the attempt classified against. A screening digest: ProjectStatisticsConfigurationDigest.ForThresholds. A membership digest: SHA-256 of the sorted canonical investigator ids, present only when reviewer moves or reviewer-scope intents are encoded. An annotation digest: ForThresholds, sorted stage ids and each stage's effective SessionCountTarget. Membership is deliberately not in the annotation digest: annotation moves take their scopes from the sessions' investigator ids, not the member list, so a member add or remove does not invalidate pending annotation entries. These replace today's in-transaction Project-version admission (see Project admission) |
CatalogueVersion / SourceVersion |
The writer constants, 1 today. Compared with the row and control exactly as today |
AssumedTrackingMode |
The durable active-reviewer mode the attempt verified |
Transitions |
One or more typed transitions, each { Kind, Payload }, not expanded moves (see below). Kind is a closed enum: ProjectScreeningProfile, ReviewerScreeningPoint, AnnotationStage, AnnotationMultiStage, AnnotationReservationClaim. A Kind this binary does not know is never ignored: the entry is poison for the fold and a fallback for the reader |
Invalidations |
Zero or more { Family, Scope? } intents. A null Scope means MarkFamilyStale. These are today's DependentFamilies outputs for anything the transition cannot move exactly, for example more than 100 memberships, missing before-state evidence, ChangesSufficientAllocation, or a question-content change |
TriggerReason |
Today's trigger reason, for history |
Why entries store transitions, not moves¶
Expanded signed moves are small for ProjectScreening, at most 12. They are not small for the
reviewer families. A save that crosses the sufficiency threshold moves the Available and
Unavailable counters of every member, which is up to 100 memberships times several keys,
hundreds of moves. Annotation inclusion changes move stage-wide totals for every stage
(AnnotationStatisticsClassifier.ClassifyPopulation).
The entry therefore stores the compact inputs from which today's classifiers derive moves deterministically:
- Screening.
ProjectScreeningProfileKeybefore and after (NS,Inc, or null). With reviewer evidence it addsReviewerScreeningPointTransitionbefore and after: per-screenerReviewerDecisionCountsfor the Study's screeners plus the persisted sufficient and insufficient flags. This is typically 100-250 bytes. - Annotation, class unchanged (
AnnotationStage).AnnotationStudyProfilebefore and after restricted to the touched stage: tally key, session profile keys and the reservation count, plus the inclusion class and the affected investigator ids. Restricting to one stage is exact only whenbefore.InclusionClass == after.InclusionClass. In that caseAnnotationStatisticsClassifier.Classifymoves keys only on stages whose own profile changed, and only the touched stage changed. This is typically 200-600 bytes. - Annotation, class changed (
AnnotationMultiStage). An include/exclude crossing moves every stage's sessioned, reconciliation, grouped-count and membership-stage keys from the old class key to the new one (Classify→AccumulateStage→AccumulateMembershipStage, over the union ofbefore.Stagesandafter.Stages). It also addsClassifyPopulationover all project stages. The entry therefore stores the full multi-stageAnnotationStudyProfilebefore and after;ClassifyPopulationneeds only stage ids, which come from the context. This is about 300 bytes per stage with a tally; Studies have tallies in 1-3 stages, so 0.5-1.5 KB.
Above 8 KiB encoded, the entry instead carries MarkFamilyStale intents for StageAnnotation,
MembershipStageAnnotation and DomainReconciliation. This is exact but costs a rebuild.
Why not intents always: crossings happen whenever screening completes a Study, which is frequent. Always-intents would keep the three families Stale for most of a screening phase and serve them authoritatively. That is correct, but it gives up the read benefit. Neither choice can fail a save.
The fold, the reader and the rebuild derive moves through one shared function,
PendingEntryMoves.Derive(entry, projectContext, requestedScopes). It calls the existing
ProjectScreeningClassifier, ReviewerScreeningPointClassifier and AnnotationStatisticsClassifier
with thresholds, the canonical membership set and stages taken from the Project read in the same
snapshot. The Context digests prove that this Project context equals the one the saver classified
against. On a mismatch the entry is not derivable for that family group (see the overlay and the fold).
A reader of one scope derives only that scope's moves. For example, the ProjectScreening read never expands reviewer moves.
Size, cap and overflow¶
- Entry size. A typical entry is 400-900 bytes. A multi-stage annotation entry is up to 8 KiB before it falls back to intents.
- Cap. At most 32 entries or 48 KiB of encoded
PendingStatisticsper Study. Production Study documents are median 3.3 KB, p99.9 about 15 KB and at most about 567 KB, against a 16 MB limit. A full set at most adds 48 KiB. In practice the fold drains within seconds. - Overflow. When appending would exceed the cap, the save still succeeds. It writes its source
change and adds
{ namespace, id, digest }toOverflow.Droppedinstead of appending. Past 64 recorded drops it incrementsUnrecordedCountand widensUnrecordedFromRevision..UnrecordedToRevision. Overflow lives on the Study and is found through the same index, so no shared document is written on the hot path.
BSON registration¶
- The pending classes are registered in
StudyRepository.CreateMappingswithAutoMap()andSetIgnoreExtraElements(true), the same convention as the statistics value-object maps (#3145). Ignoring extra elements is not the compatibility mechanism for new content. Any change that adds meaning (a new transition kind or payload field the derivers must read) bumpsFoldProtocolVersion, under the N-1 rules. An entry at a protocol the binary does not accept is poison for the fold and a fallback for the reader. Within an accepted protocol, every reader knows every field. - The
Studymap is unchanged. Unknown top-level elements are preserved by theEntity[BsonExtraElements]capture (SyRF.SharedKernel/BaseClasses/Entity.cs), which is why an older binary keepsPendingStatisticsandStatisticsFoldSequenceintact through a replace. - Removing or re-typing a field needs an ADR-011-style writer floor.
Index¶
IX_Study_PendingStatistics { ProjectId: 1, "PendingStatistics.OldestAtUtc": 1 } with
partialFilterExpression: { "PendingStatistics.OldestAtUtc": { $exists: true } }.
It holds only Studies that have something pending. Every pending Study carries OldestAtUtc, so
the index cannot hide an entry or an overflow marker.
This is a production operation, not a startup side effect. A partial index still scans the whole
pmStudy collection (about 10 GB) while it builds. It is therefore not added to
InitialiseIndexesAsync. Instead:
- Build path. An administrator-only, idempotent
POST api/admin/project-statistics/fold/pending-indexbuilds it withcommitQuorum: "votingMembers", after a separately approved off-peak window. - Timing. Build time is measured first on a restored production-sized copy, and recorded in the fleet runbook.
- Enable gate. Fold-mode enable refuses with
PendingIndexMissinguntillistIndexesshows the index ready. - Precedent. The
ProjectId_1_WorkloadShareBucket_1index rollout.
2. The save¶
Fold-path writers¶
Every Study writer that today hands a fold-family request to
ProjectStatisticsTransactionCoordinator.CommitAsync moves to the fold path:
| Writer | Today | Fold path |
|---|---|---|
Screening saves: SaveScreeningWithCapacityGuardAsync, TrySaveScreeningWithStatisticsAsync (plain and reconciliation), TrySaveActivityReviewAsync (ReviewEligibilityPolicy), including the released-reservation annotation half |
Snapshot transaction carrying the coordinator | Single-document FindOneAndReplace with the entry. TrySaveActivityReviewAsync keeps its eligibility transaction, but no statistics document is written in it |
Annotation session save and complete: SubmitAnnotationSessionService, on both the capacity and SaveWithAnnotationStatisticsAsync branches |
Snapshot transaction (presence rows plus coordinator) | Entry on the Study write. The presence rows are not statistics documents and keep the caller's transaction where one exists today |
Annotation and reconciliation session delete: ReviewController.RemoveSession → TrySaveActivitySessionDeletionAsync / SaveWithAnnotationStatisticsAsync |
Snapshot transaction | Entry on the Study write |
Reservation claims: TryAtomicAssignStudyCoreAsync (atomic claim pipeline), TryAdmitActivityReviewAsync |
Bare FindOneAndUpdate with writes off; a snapshot transaction with writes on (TryAdmitActivityReviewAsync is always transactional, for eligibility) |
The claim stays today's single atomic FindOneAndUpdate, with one prepended pipeline stage that captures the pre-image into the entry. The admission keeps its eligibility transaction. See the claim path |
Reservation releases: TrySaveReservationChangeAsync (hub leave and disconnect, idle and suspended consumers), TryReleaseScreenedReservationAsync |
Snapshot transaction | Entry on the Study write |
Writers that stay off the fold path, with the reason:
| Writer | Why it is safe |
|---|---|
TrySaveAdministrativeScreeningBatchAsync (bulk update file) |
Runs under the BulkStudyUpdate staged fence; families are Stale until rebuilt |
Inclusion recalculation UpdateMany |
Under the source-visibility fence. Its pre-existing lost-update race with concurrent saves (it does not bump Audit.Version) is unchanged and is recorded as deferred item 1 |
StageReviewSettingsController settings save and claim revocation; definition rewrite release; stage SessionCountTarget invalidation |
Low-frequency administrative Project writes. They keep their coordinator transaction, which writes the row directly. That is allowed by invariant 1 |
| Search import, failure and removal; bulk Study update | Under staged fences |
| Presence, idle, schedule-token and bulk-PDF writers | Change no statistics input (matrix §2.13) |
StudyController correction approve and reject |
Writes StudyCorrection, not the Study |
Which attempts take the fold path¶
An attempt takes the fold path when all of the following hold. Otherwise it runs today's path, unchanged.
MaterializedProjectStatisticsWrites, the writer's family flags andMaterializedProjectStatisticsFoldare on, and the project is allowlisted. These are captured once per attempt, as today.- The project control's
FoldModeisEnabled. - The transition classifies to at least one move or one invalidation intent. A no-op attempt appends nothing and writes what a source-only save writes today.
- The attempt is within its pre-write deadline (below).
- The writer can write the project's stamped protocol: its own, or its previous one under the
N-1 rules. Otherwise the binary is outside the project's protocol window and
Matchesalready refuses the project.
A fold-worker outage does not take saves off the fold path. Revision 4 sent writers to the transactional path when the worker heartbeat went stale. That path is the one the benchmark measured at up to 40% exhausted saves, so a project-management outage would have become a statistics-caused save-failure storm. During an outage, then: - saves keep appending; - readers stay exact through the overlay until the age or size threshold, then fall back to the authoritative calculation; - entries that hit the per-Study cap overflow, and the families they touch go Stale.
None of these fails a save. The only reason for the old gate was old history writers. Those are kept away from fold projects by the rollback guard and the enable rollout gate, and new history writers check the pending set themselves (history writers). 6. One path per attempt. A save that carries both a screening and an annotation namespace takes the fold path or the transactional path for both, never one each.
The pre-write reads and the command budget¶
On the fold path the writer does not call BeginOperationAsync's epoch capture. That capture
reads the global control, the project control and one to three family guards for screening, and the
global control, the control and three guards for annotation. Entries carry no epochs; the fold checks
epochs in its own snapshot. The reviewer before-state read (CaptureReviewerBeforeAsync) is replaced
by the attempt's isolated load. The classification context comes from the Project the request already
holds; it may be the cached instance, because the entry's Context digest is re-verified at
consumption. Reservation releases therefore need no extra Project load for statistics.
The remaining statistics reads are issued concurrently, costing one round trip of latency:
- the global control: the existing per-attempt
ResolveVerifiedActiveReviewerTrackingModeAsyncread is reused; - the project control, for
FoldModeand the tripwire. This is the one new read; - on attempts that need it (below), the receipt for the envelope id and the quarantine record for it.
The global and project controls are read, never written, on this path, so they cannot cause write conflicts.
| Save shape | Commands, first attempt | Commands, retry attempt | Source-only today |
|---|---|---|---|
| Capacity-guarded or plain screening | isolated Study find + global find + control find + findAndModify = 4 |
+ reload find + receipt find + quarantine find = 6 |
2 (find + findAndModify); the durable-mode read is skipped with writes off |
ReviewEligibilityPolicy screening |
the above + uncached Project find + the eligibility transaction (its own reads, Project token and commit) |
same + receipt + quarantine | identical eligibility commands |
| Annotation session save or delete | the Study load + global + control + Study write, inside the caller's existing presence transaction where one exists | + receipt + quarantine | the same transaction without statistics |
| Reservation release | Study load + global + control + Study write | + receipt + quarantine | Study load + Study write |
Reservation claim (TryAtomicAssignStudyCoreAsync) |
global + control + one findAndModify (today's claim with one added pipeline stage) = 3, plus one cached Project read on a cache miss when the caller does not pass its Project (4) |
none (the claim has no retry and no envelope reuse; see the claim path) | 1 bare findAndModify (no transaction with writes off) plus the durable-mode read when writes are on |
Eligibility admission (TryAdmitActivityReviewAsync) |
its existing eligibility transaction + global + control; the entry is a $set of the whole pending set inside it |
+ receipt + quarantine | its existing eligibility transaction |
The acceptance test asserts these counts per shape (9).
When the receipt and quarantine checks run. An attempt skips them only when it compare-and-sets
on exactly the envelope's ExpectedSourceRevision. If that id had ever committed, the Study's
version would have moved past that revision (appends, folds and legacy receipts all write this Study),
so the compare-and-set itself would fail. Every other attempt runs both checks:
- any retry;
- a first attempt whose isolated load returned a newer version than the instance the envelope was
minted from.
The envelope is minted from the isolated load wherever the writer controls the load.
Pre-write deadline¶
An attempt records a monotonic timestamp when its pre-write reads complete. It refuses to write, and
reloads instead, if more than FoldPreWriteDeadline (30 seconds) has passed by the time it issues
the Study write. This bounds the window in which a save can append under a stale FoldMode; the
disable finalization grace (5 minutes) is longer.
Durable reviewer-mode check¶
The reviewer-tracking check keeps failing closed with the typed 503 on disagreement, and the capacity filter still follows the durable mode. It is no longer atomic with the write. Owner decision (d) accepts this, on these terms:
- the entry records
AssumedTrackingMode; - the fold compares it with the durable mode in its snapshot and stales the affected families on any disagreement;
- any future owner of the mode transition (matrix M15) must wait at least
FoldPreWriteDeadline+FoldWriterGrace(5 minutes) after committing stage one before it treats the fleet as transitioned. That requirement is recorded in invariant 13.
Every pending entry records the mode it was classified under (AssumedTrackingMode). An entry whose
assumed mode differs from the durable mode at consumption time is never applied:
- the fold discards it and marks every family it touches Stale;
- the overlay falls back for any requested scope while such an entry is pending;
- a rebuild refuses with
PendingNotDraineduntil it is drained.
Future: per-stage reviewer tracking¶
Not in the MVP; tracked as #3876. Today the effective reviewer-tracking mode is a fleet-wide flag. If it became a stage setting (effective mode = the stage setting AND SignalR active), it would be an input carried on the Project. Two things follow:
- The M15 fleet transition becomes an ordinary stage-scoped configuration change. It would be
invalidated like
SessionCountTarget: the annotation context digest changes and the affected families are staled in the settings save. - The fold path needs neither the pre-write global-mode read nor the writer grace of
invariant 13. The mode would be part of the Project the save already loaded, and
the entry's
Contextdigest would cover it.
Project-version admission is replaced, not relocated¶
Today the reviewer-family point path admits ExpectedProjectVersion with
ProjectStatisticsProjectVersionSource.HasVersionAsync. That is a real write of
Project.StatisticsAdmissionToken inside the transaction, and it makes every concurrent save in a
project conflict on the Project document. The fold path drops it:
- The entry's
Contextdigests capture exactly the Project inputs the classification used. - The fold and the reader recompute those digests from the Project in their own snapshot. An entry whose context no longer matches is not derived (see the fold and overlay sections). A concurrent membership or configuration edit is therefore detected at consumption time instead of being serialized at write time.
TrySaveActivityReviewAsync also calls HasVersionAsync, but unconditionally. That call is the
ReviewEligibilityPolicy admission of the eligibility decision against Project settings, and it runs
with statistics flags off. It is a source-level write, so its conflicts appear in the source-only arm
too. The fold path neither adds nor removes it, and the gate counts it as source-level
(9). Whether eligibility should keep that hot write is a separate,
non-statistics question, recorded as deferred item 2.
The Study write¶
The filter is unchanged: BuildScreeningSaveFilter, or the writer's existing { _id, Audit.Version }
guard, plus the capacity and reservation guard where it applies today. The write is still the
writer's existing single-document replace or claim pipeline.
Before the write, the repository appends the entry (or records overflow) on the attempt's isolated copy of the Study:
- Fold-path attempts load under
BeginIsolatedReads()(see.claude/rules/repository-cache.md). The envelope is minted from that load. - The shared cached instance is never mutated, so a failed attempt cannot leak a phantom entry into the repository cache.
- The isolated load is also what guarantees the reviewer before-state is the persisted document at
SourceRevisionBefore. Today that is provided byReviewerScreeningPointSnapshotSource.
The write uses w: majority. Snapshot readers read majority-committed data only after a majority
commit, and a w:1 write rolled back by a failover must not be observable as a pending entry. The
write also carries a client-side operation timeout of FoldPreWriteDeadline (CSOT / maxTimeMS).
Server-side queueing therefore cannot stretch the window beyond the deadline.
The claim path stays one atomic update¶
With statistics writes off, today's claim (TryAtomicAssignStudyCoreAsync when
ReviewEligibilityPolicy is off) is one non-transactional FindOneAndUpdate:
- the filter is
BuildAssignmentCapacityFilter; - the update is the
BuildAssignmentPipelinepipeline; WithCapacityWriteSessionAsyncreturnsWriteWithoutTransactionAsyncwhen writes are off (StudyRepository.cs157-158 and 1808).
Two reviewers claiming one Study serialize on the document lock, and the capacity filter decides each
claim cleanly. Any transaction on the fold path would turn that contention into WriteConflict
aborts, which are statistics-caused failures under gate (b). The fold path therefore keeps the claim
a single atomic update.
What today's pipeline does (BuildAssignmentPipeline, StudyRepository.cs around 2577-2700). It
has four stages:
- append the new
SlotReservation; - seed a zero tally for the stage if none exists:
{NCS 0, NCCS 0, RS false, RC false, reservations 0, …}; - increment the matching tally's slot-reservation count and recompute its
TotalAllocatedandTotalEngagedsums; - bump
Audit.Versionand refreshLastModified.
The filter admits a Study with no tally for the stage (noTallyForStage), so the first claim on a
Study in a stage is the most common claim shape. It moves real statistics.
AnnotationStudyProfile.FromStudy creates a stage entry only when a tally exists, so the transition
is "no stage" → "stage with tally 0:0 (plus that stage's sessions)". AccumulateStage then moves:
GroupedCount(class)at bucket"0:0";- every sessioned and domain-reconciliation predicate that a 0:0 tally holds;
- membership-stage keys for any sessions already on the stage.
The fold-path claim:
- Filter: unchanged. It is exactly
BuildAssignmentCapacityFilter. There is no version filter, no transaction and no retry. - Update: today's four stages, with one stage prepended (stage 0). Stage 0 runs first, so it
sees the pre-image, including the pre-image
Audit.Version, because stage 4's bump comes later. It$sets the wholePendingStatisticsvalue with aggregation expressions, and is written to never fail (S-R4-1): -
Malformed or absent sets. Every read of the existing set is guarded with
$ifNull,$isArrayand$type, and$bsonSizeis applied only to a document.- An absent set is treated as empty.
- A malformed set (
PendingStatisticsis not a document, orEntriesis not an array) is replaced by a well-formed set: no entries, and anOverflowmarking an unrecorded drop over[0, pre-image Audit.Version].
This fails closed: the families go Stale at the next fold, and unknown results on that Study answer
StatisticsOperationUnresolved. The claim itself always succeeds. -OldestAtUtc:{ $min: [ { $ifNull: [ "$PendingStatistics.OldestAtUtc", "$$NOW" ] }, "$$NOW" ] }. -Entries:$concatArraysof the existing entries and oneAnnotationReservationClaimentry. Its payload copies, from the pre-image: - the claimed stage's tally (ornullwhen there is none); - the profile fields of the claimed stage's sessions: id, investigator, status and the reconciliation flag. These are projected with$mapover the sessions filtered to that stage, so they are small; - the claimed stage'sSlotReservations; -InclusionInfo.It also records constants computed client-side: - stage, investigator, regime,
sessionCountTargetandenforceCapacity; - per-entryProtocolVersion, theContextdigests andAssumedTrackingMode.CreatedAtUtcis$$NOW. - Cap. It is evaluated with$sizeand$bsonSize. Past the cap,Overflowis updated instead, and the drop is always recorded as unrecorded, wideningUnrecordedFrom/ToRevision, because a claim has no envelope digest (F-R4-3). Unknown results of enveloped saves whose revision falls in that range then answerStatisticsOperationUnresolved. That is acceptable: it happens only when the cap is reached, with the worker down.
Today's four stages follow unchanged.
3. Derivation. The shared deriver mirrors BuildAssignmentPipeline exactly:
- before is the annotation profile built from the copied pre-image fields. The claimed stage is
absent when the copied tally is null;
- after is the same pre-image fields with stage 2's effect applied (the claimed stage seeded
to the zero tally when absent) and stage 1's effect applied (one more reservation for the
investigator). Stage 3's sums are not profile inputs;
- moves are AnnotationStatisticsClassifier.Classify(before, after). The inclusion class is
unchanged by a claim, so the single-stage rule of N2
is exact. Moves for stages other than the claimed one cancel.
4. ReviewerAnnotation invalidation. This is derived, not stored. Today's reservation path stales
ReviewerAnnotation through AnnotationStatisticsClassifier.DependentFamilies:
- family-wide MarkFamilyStale when ChangesSufficientAllocation(before, after, stage,
sessionCountTarget) holds;
- otherwise MarkStale for the actor's membership-stage scope.
The shared deriver computes exactly that from the payload, which carries sessionCountTarget and
the assumed tracking mode, and returns it as the entry's derived intent. The fold applies
derived intents exactly like stored ones, and the reader falls back on any derived or stored intent
covering a requested scope. A crossing of SufficientlyAllocated therefore can never leave
ReviewerAnnotation served Fresh.
5. Identity. The claim is not an enveloped operation. Its entry id is
claim:{study}:{investigator}:{stage}:{pre-image Audit.Version}, computed in stage 0. The fold
writes a receipt for it like any entry. There is no client duplicate check, because a claim that
retries today simply re-runs the filter.
6. Context. The Context digests need the Project.
- The fold-path overload of TryAtomicAssignStudyCoreAsync takes the caller's Project.
StageReviewService already holds it, and a cached instance is acceptable because the digests
are re-verified at consumption.
- A caller without one loads it through the cached repository, costing one command on a cache
miss.
Mandatory test shapes (slice 5, red-first). The tests run the real pipeline on a replica set and
compare against the authoritative calculation and against FromStudy(returned after-image). Each
shape asserts that derive(entry) equals the classification of the real before and after images
and that the overlay equals the authoritative value:
- first claim, no tally for the stage:
- with no sessions on the stage;
- with existing candidate sessions;
- with existing reconciliation sessions;
- claim on an existing tally (resume after release, another reviewer's tally);
- tracked and untracked mode;
- allocated and unallocated (regime set or not);
- screening and annotation stages;
- a claim that crosses
SufficientlyAllocated(derived family-wide intent); - a claim on a Study whose pending set is absent, well-formed, at the cap, or malformed.
There is no "re-claim by the same reviewer" shape, because the filter excludes an existing reservation or saved session. Timeouts and expiry are releases, which use the release path, not the claim.
Equivalence with today's claim. The filter is identical, so a claim matches exactly when today's
would. Stage 0 writes only PendingStatistics, which neither the filter nor today's stages read, and
it cannot raise an error. The outcomes (the assigned Study, or the ordinary null for "already
reserved, completed or at capacity") are therefore identical, and statistics add no conflict, no abort
and no retry.
TryAdmitActivityReviewAsync (ReviewEligibilityPolicy on) is transactional today for eligibility
reasons, with an explicit version check. It keeps that transaction, and $sets the whole pending set
inside it, classified from its snapshot before-image and the real after-image.
Minimum MongoDB server version¶
The design needs MongoDB 4.4 or later:
| Feature | Minimum |
|---|---|
$bsonSize (claim cap) |
4.4 |
Aggregation expressions with $$NOW in a find projection (heartbeat and fleet freshness) |
4.4 |
Pipeline updates in findAndModify |
4.2 |
$$NOW |
4.2 |
$isArray, $type, $concatArrays, $min |
3.x |
Test fixtures run mongo:8.0, and the benchmarks recorded 8.0.28. The repository already ships 4.2+
pipeline claims with $$NOW in production. The production Atlas cluster version is to be confirmed
by the operator before slice 1 deploys: the review could not read it because the Atlas API returned
401.
Idempotency, retries and unknown results¶
Each attempt runs these checks in order. All but the last use only the loaded Study or one receipt read.
- Pending-set duplicate. If the isolated Study's
PendingStatistics.Entriescontains thisOperationId: - same
SubmissionDigest: answer Saved, a no-op; - different digest: fail closed with
DigestMismatch, as today. - Overflow duplicate. If the id is in
Overflow.Dropped, compare the stored digest: same → Saved, different →DigestMismatch. - Receipt duplicate (when required). Run
ResolveByReceiptAsync, the existingMatchesEnvelopecomparison: - match: answer Saved;
DigestMismatch: fail closed.- Quarantine duplicate (same condition). Look up
pmProjectStatisticsFoldQuarantineby(ProjectId, OperationNamespace, OperationId). - Overflowed and poison entries are moved there by the fold, with their digest when it could be
read. Same digest → Saved; different →
DigestMismatch. - An
Undeserializablerecord has no readable digest. The answer is the typedStatisticsOperationUnresolved(409), which is never Saved.
The fold removes an entry and inserts its receipt or quarantine record in one transaction, and those are read after the Study on the primary. A consumed operation is therefore always visible through check 1, 2, 3 or 4. Quarantine records are kept for 30 days by TTL and are never pruned earlier (poison), which is far beyond any retry or redelivery window. 5. Otherwise, write.
This reproduces today's semantics exactly: an exact duplicate is Saved, and a changed decision under the same envelope fails closed. It closes the two races the review found.
- The id is reused across reloads, and that is deliberate.
ReviewControllermints the envelope once, before the retry loop. Uniqueness is enforced by the checks above, not by the id embedding each attempt's version. - A second request with the same id and a different digest never appends. Two concurrent requests that loaded the same version both carry the same id. The loser's compare-and-set fails, and its reload then sees the winner's entry, receipt or quarantine record and fails closed.
- No defence-in-depth gap. A legacy transactional writer on another pod can still race a fold-path writer to the same id. The fold handles that collision without halting (see poison entries).
Unknown write result. This covers a network error that survives the driver's single retryable-write retry. Reload uncached and run checks 1-4, regardless of the reloaded version, because a fold may already have bumped it. Use the same digest rule:
- the id is found with this attempt's digest: the write committed. Answer Saved without
re-applying, so a released reservation never turns a committed save into
AtCapacity; - the id is found with a different digest: another request with the same envelope committed and
this one did not. Answer
DigestMismatch, never Saved; - the id is not found anywhere: the write did not commit, so continue the attempt loop;
- the envelope's
ExpectedSourceRevisionfalls inside an unrecorded-drop range for this Study: eitherOverflow.UnrecordedFromRevision..UnrecordedToRevisionon the Study, or anUnrecordedDropsquarantine record the fold moved it into. The save may be one of the unrecorded drops, so it cannot be proven either way. Answer the typedStatisticsOperationUnresolved, never Saved and never a retry that could answerAtCapacity. This happens only after 96 saves on one Study while the fold worker is down. - the project's control carries
FoldQuarantineIncompleteSinceUtc(the quarantine hard cap was reached) and nothing was found: answerStatisticsOperationUnresolved, because the id may have been consumed without a record.
Today's transactional path propagates an exhausted unknown commit as a 500. This is a strict improvement.
Fold-only reloads are free retries¶
A fold bumps Audit.Version and StatisticsFoldSequence together. After a VersionStale, the
controller compares the reload with its attempt copy. If
ΔAudit.Version == ΔStatisticsFoldSequence, then only folds changed the document. The attempt is
not charged against the three-attempt budget, up to two free retries per request.
Folds therefore cannot cause a statistics-caused exhaustion; this is gate (b). The benchmark reports fold-caused and source-caused conflicts separately (9).
Receipts and soak evidence¶
- Who writes it. The receipt is written by the fold, once per consumed entry, keyed
(ProjectId, OperationNamespace, OperationId)as today. - Fields.
ExpectedSourceRevisionis the entry'sExpectedSourceRevision.CommittedSourceBefore/AfterRevisionare the entry's actual revisions.SourceDigestisSubmissionDigest.ObservedAtUtcis the fold time. - The #3569 validator needs no change.
scripts/validate-statistics-evidence.pycounts in-window receipts withafter > before >= 0, deduplicated by namespace and id, and fold-written receipts meet all of that. - Two documented differences.
- A receipt is observed at fold time. A save within seconds of a capture boundary can fall into the next capture.
- An overflowed save, and a quarantined entry, have no receipt; they have a quarantine record.
They are counted by
pending_entries.overflowedandfold.quarantined, and the evidence capture reports both next toconfirmedSourceOperations. - The
source_operation_receipts.committedcounter keeps its meaning. The fold records it afterCommitWithUnknownResultRetryAsyncconfirms. - Receipt maintenance and pending entries.
ProjectStatisticsReceiptMaintenanceServicemust refuse to advanceRetiredThroughSourceRevisionfor a Study that has a pending set, which is one indexed probe per Study in its batch. Otherwise a legacy receipt at a higher revision could be retired while an older entry on that Study is still pending, making the entry poison. That outcome would fail closed, but it is avoidable.
3. The fold worker¶
Placement and triggers¶
The worker is the consumer ProjectStatisticsFoldConsumer in the project-management service, with
ConcurrentMessageLimit = 4 per pod. It has two triggers:
- Post-write signal. After a confirmed fold-path write, the API pod enqueues the project id on an
in-process coalescer. The coalescer publishes
IFoldProjectStatisticsCommand { ProjectId }at most once per project every 250 ms per pod, from a background task. The publish is never on the request path and never inside a retryable callback. - Sweep, as the safety net. A Quartz recurring publish,
IRunProjectStatisticsFoldSweepCommand, fires every minute. It follows theProjectStatisticsRepairSchedulepattern: no startup catch-up, one delivery across pods. It enumerates the distinctProjectIds in the partial index whoseOldestAtUtcis more than 10 seconds old, oldest first, up to 256 projects per run.
Change streams are rejected. No product code consumes one today, and they would add resume-token and failover handling that the lease already makes unnecessary.
The project fold lease¶
The fold uses its own document in pmProjectStatisticsRebuildLease, under a reserved scope key
(ScopeKind.Project, canonical fold). It does not share a family scope's rebuild lease, so a
takeover never inherits a rebuild's captured watermark.
- Fields used. The existing
Owner,LeaseGeneration,ExpiresAtUtcandStartedAtUtc. The fold lease records no captured revisions. - Duration. 30 seconds, renewed after each batch.
- Fencing. Every fold transaction updates its lease with filter
{ _id, Owner, LeaseGeneration }and$set: { ExpiresAtUtc }.Tokenis not rotated on takeover, so it is not used for fencing. A paused or partitioned worker whose lease was taken over matches nothing and aborts. - Exclusion with rebuilds. A rebuild or checkpoint build of any fold family takes both its scope lease and the project fold lease, so no fold runs during its capture-to-publication window.
- Yielding. To avoid starving rebuilds on a hot project, a rebuild that finds the fold lease held
sets
YieldRequestedUntilUtcon the lease document (a new field, tolerated byExtraElements). The holder stops after its current batch and does not re-acquire while the request is live. - Missed signals. A holder re-queries once after releasing. The sweep covers the rest.
- Protocol gate. A worker folds a project only when
ProjectStatisticsFleetVersions.Matchesaccepts the control's stamped protocol, which must be its own or, under the N-1 rules, its previous one. A worker outside that window leaves the project's entries untouched. They stay pending and readers at that protocol fall back, so nothing is silently dropped. - Allowlist (S10). The worker folds only allowlisted projects. A non-allowlisted project's entries
stay pending. That is safe because nothing serves a non-allowlisted project, and its entries are
handled by the mandatory
resetthat ends a guard window.
The fold-worker heartbeat¶
The global control carries a map, FoldWorkerHeartbeats: { "<protocol>": { RenewedAtUtc, Host } }.
Renewal. Every project-management pod running the fold consumer renews, every 30 seconds, the keys for every protocol it can fold: its own protocol and, under the N-1 rules, the previous one. It renews with one update:
{ $currentDate: { "FoldWorkerHeartbeats.<P>.RenewedAtUtc": true }, $set: { "FoldWorkerHeartbeats.<P>.Host": host } }
- Server time.
$currentDateuses the server's clock, never the pod's (S2). - No wedge. Each protocol has its own key and there is no protocol filter, so a pod can never be blocked from renewing its own protocol's key after a rollback (S1). A key nobody renews simply goes stale.
- No
Versionbump (S3). The renewal does not touch the document'sAudit.Version, so it never makes an administrative global-control CAS fail. - The cost of that choice: an administrative full-document replace, by a new or old binary, writes back the heartbeat values it read. That can make a heartbeat look up to one interval older, never newer. It fails safe: a stale heartbeat only delays an enable or a stamp advance, and raises the worker-down alert early.
- A slice-1 test asserts that a replace from a stale copy cannot advance any
RenewedAtUtc.
Freshness check. The heartbeat does not gate saves. It is read only by:
- enable and reset, which need a live worker for the protocol being stamped;
- the stamp advance;
- the fold sweep, which emits fold.worker_heartbeat_age for the worker-down alert.
Those reads use a projected find that returns the control's fields plus foldWorkerLive, and is
never cached, because a cached projection would freeze $$NOW (S-R4-8). The save path keeps
today's ResolveVerifiedActiveReviewerTrackingModeAsync read unchanged:
{ foldWorkerLive: { $gt: [ "$FoldWorkerHeartbeats.<P>.RenewedAtUtc", { $subtract: [ "$$NOW", 120000 ] } ] } }
This compares server time with server time.
Effect of a worker outage. Saves keep appending and readers stay exact through the overlay. After
10 minutes of age, 128 Studies or 1,024 entries, readers fall back to the authoritative calculation.
The worker-down alert fires on fold.worker_heartbeat_age > 2 min. When a worker returns, it folds
everything still pending; anything that overflowed meanwhile is Stale and is rebuilt by the repair.
No save fails at any point.
Per-pod fleet membership¶
pmProjectStatisticsFleetMember holds one document per pod of the api and project-management
deployments only. Other hosts, such as Quartz, have no allowlist role and never write a record. A
record with any other Deployment value is ignored (S-R4-2). Each record holds:
_id: the pod name;Deployment:apiorproject-management;BinaryVersion, the protocols the pod writes and folds, andFoldWorker;Allowlist: the pod's sortedProjectStatistics:ProjectAllowlistitself, not a digest. It is bounded to 1,000 ids; a longer allowlist refuses to start fold mode with a typed configuration error. Staging lists one id and production lists none, so a per-project question needs no extra storage;RenewedAtUtc: set by$currentDateevery 30 seconds.
Lifecycle. A pod deletes its own record on graceful shutdown (IHostApplicationLifetime.ApplicationStopping).
A TTL index removes a stale record after 10 minutes, but only for garbage collection.
Liveness is server time, not TTL existence. A member is live when
RenewedAtUtc > $$NOW - 2 minutes, evaluated server-side with $expr in the query. A SIGKILLed pod
therefore stops counting within 2 minutes. Readers cache the result per process for at most 10
seconds, so a member's state reaches a reader at most 2 minutes 10 seconds late.
It is written outside every save. It serves three purposes:
- Enable, reset and the protocol-stamp advance require every live member of both deployments to declare the protocol being stamped, and none to declare a lower one (S11).
- Guard consistency (NM2), per project. A reader, the repair and the fold worker refuse project
P only while live members disagree on whether P is in their allowlist. The reader's answer is
PendingOverlayUnavailable/AllowlistInconsistent, falling back to authoritative. - An allowlist edit that adds or removes another project affects only that project, so no other fold project falls back.
- This makes the guard's own rolling rollout fail closed between fold-aware binaries for the projects the guard changes.
- Pre-fold replicas cannot be detected this way, because they write no membership. They are covered by explicit rollout-complete gates in the enable, stamp-advance and rollback runbooks (S-R4-6, runbook) and by the narrow gate, which every binary honours.
Batch selection and size¶
- Per transaction: the whole pending sets of up to 32 Studies, capped at 32 entries, oldest
OldestAtUtcfirst. - Why 32: every applied entry becomes one delta record at the batch's single revision, and delta
maintenance (
SelectCompleteRevisions) refuses a sibling group of more than 32. - Per Study: each Study gets exactly one update.
- Minimum age: the background fold skips any Study whose newest entry is less than 500 ms old. This reduces collisions with the immediately following presence and reservation writes on that Study. A rebuild drain ignores the minimum age.
- Livelock bound (S12). One hot Study can conflict a whole batch:
- after a write-conflict abort, the worker halves the batch and retries, down to a batch of one;
- a Study whose single-Study batch aborts 5 times in a row is skipped for 2 seconds (in memory, per worker), so the rest of the project keeps folding;
-
after 20 consecutive single-Study aborts, the worker switches that Study to the two-step discard, which does not need a multi-document transaction to touch the Study:
- A transaction that does not write the Study. It marks Stale every family that the Study's entries touch, and every family an overflow marker implies. It inserts quarantine records for:
- the entries (reason
Livelocked, with digests); - every
Overflow.Droppedrecord (Overflowed); - the
UnrecordedDropsrange, if any (S-R4-3).
It reads the Study in its snapshot but writes only statistics documents, so a concurrent Study write cannot conflict with it. 2. A single-document, non-transactional pipeline update. It: - removes from
Entriesexactly the(OperationNamespace, OperationId)pairs step 1 quarantined; - removes fromOverflow.Droppedexactly the pairs step 1 quarantined; - clears the unrecorded range only ifUnrecordedToRevisionstill equals the value step 1 quarantined, because a newer unrecorded drop is kept for the next pass; -$unsetsOverflowwhen nothing remains in it, recomputesOldestAtUtc(or$unsets the whole set when it empties), and$incsAudit.VersionandStatisticsFoldSequence.It has no version filter, so it always applies. A single-document update cannot be aborted by a write conflict; it waits on the document lock.
Drops and entries appended after step 1 have other keys and survive for the next pass. A hot overflowed Study therefore cannot keep the bundle falling back, or rebuilds refusing, for ever.
A crash between the two steps leaves entries that are both pending and quarantined. The next fold treats a quarantine record for the same id as poison (
QuarantineExists) and consumes it without inserts. The families are already Stale, so readers fall back throughout, and the bound is two commands.
A rebuild drain uses the same escalation, so it cannot be held off indefinitely by one Study. - Per message: a holder runs at most 16 batches, then republishes itself.
Fold transaction¶
The transaction is opened with ProjectStatisticsTransaction.SnapshotOptions, committed through
CommitWithUnknownResultRetryAsync, and wrapped in RetryDefinitelyAbortedAsync. A rerun is safe
because it re-reads everything.
1. Read in the snapshot¶
Read:
- the global control, the project control and the Project (for context digests);
- the guards of the fold families;
- the lease;
- the batch, through the partial index, projecting
_id,Audit.Version,StatisticsFoldSequenceandPendingStatistics; - with one
$infindeach, existing receipts and existing delta records for the batch's ids; - for each family and scope the batch's entries touch, the visible row through
ProjectStatisticsVisibleRowSelector; - the ledger admission state through
ProjectStatisticsRevisionLedgerAdmission.IsAtCapacityAsync.
2. Classify each entry¶
Each entry, and each family it touches, takes exactly one disposition:
- Poison, handled per entry, is set out below.
- Apply (per family) when every one of these holds:
- the global and project
Modeare Enabled; - the entry's
Contextdigest for that family group equals the snapshot's; CatalogueVersionandSourceVersionequal the writer constants recorded on the row;AssumedTrackingModeequals the durable mode;- the family guard is Fresh with no active fence, and no project source-visibility token is held;
- the target row passes
EvaluateRow; - the ledger is not at capacity.
A derived move for a scope with no row is never dropped. It marks that scope Stale (see
families and scopes with no row).
- Discard (per family) in any other case: the entry's moves for that family reach no row. The
family is marked Stale in this transaction (MarkFamilyStale, which also advances
FamilyWriteEpoch) unless it is already non-Fresh. A later rebuild re-derives it from the Studies
(subtraction rule).
- Invalidation intents are applied as MarkStale (scope) or MarkFamilyStale (family), exactly
as the coordinator's dependent writes do today.
- Overflow on any Study in the batch: every fold family for which that Study could have produced
moves is discarded (marked Stale). Two things move into the quarantine collection, so they stay
detectable after the marker is cleared with the entries:
- the Overflow.Dropped records, with reason Overflowed and their digests;
- one UnrecordedDrops { StudyId, FromRevision, ToRevision } record when UnrecordedCount > 0
(S8).
- Protocol or kind mismatch: an entry whose own ProtocolVersion is outside the worker's accepted
window (its protocol and, under N-1, the previous one), or with an unknown Kind, is poison.
3. Apply¶
- Run
TryApplyMovesover the batch's derived moves per row. A capacity refusal (50,000 keys or 8 MiB) switches that family to discard with reasonCapacity. - No negativity check runs on the stored row alone. Stored + batch may be transiently negative because a transactional writer applied a later move directly (invariant 1). The full overlay is what must be non-negative, and the reader checks it.
control.AdvanceClocks(sourceChanged: true, projectionCommitted: anyApplied)once per batch.SourceInvalidationRevisionandClientInvalidationRevisionalways advance;CommittedProjectionRevisionadvances when anything was applied.- For each moved row,
RecordChange(CommittedProjectionRevision, epoch(global, project, guard.FamilyWriteEpoch), Fresh, SourceVersion, CatalogueVersion, ConfigurationDigest). These are the same values the coordinator stamps today; there is no row marker. guard.RecordPointMaterialization(CommittedProjectionRevision)for each moved family.- One
Deltarevision record per applied entry (stats.delta.point,RecordId = OperationId), all at the batch revision (one sibling group, at most 32 records), carrying the derived moves. This is the same shape history and replay read today. - One receipt per consumed, non-poison entry.
- Coalesce the outbox slot of every family that was moved, staled or discarded, to the new
ClientInvalidationRevision. Only the worker writes these slots.
4. Remove entries and finish¶
- Duplicate-key handling. A
DuplicateKey(E11000) on a receipt, delta or quarantine insert means a racing writer committed after the snapshot. It is classified exactly like a transient transaction error: abort and retry from a new snapshot, where the poison check then sees the row. It never surfaces as a non-retryable fault.ProjectStatisticsUnchangedDayRecorderalready classifies it this way. - Update each Study with filter
{ _id, "Audit.Version": <read> }and{ $unset: { PendingStatistics: "" }, $inc: { "Audit.Version": 1, StatisticsFoldSequence: 1 } }. Whole sets are always taken, andmatchedCountmust equal 1. - Renew the lease (the fencing write), CAS the control and guards, and commit.
- After a confirmed commit:
- record
source_operation_receipts.committedper receipt,fold.batches,fold.entries,fold.discarded{family,reason}and thefold.laghistogram; - invalidate this pod's Study repository cache for the folded ids.
Poison entries never halt a project¶
An entry is poison when any of these holds:
- a receipt already exists for its id, whether it matches or conflicts;
- a delta record already exists for its id;
- its revision is at or below the Study's
RetiredThroughSourceRevision(RejectRetiredSourceAsync); - its protocol version is outside the worker's accepted window, or any transition
Kindis unknown; - a quarantine record already exists for its id (
QuarantineExists, after a crash between the two steps of the livelock discard); - its transition cannot be deserialized.
Each of these can only follow an invariant breach or a legacy-path race. A poison entry is consumed in the same fold transaction without inserting anything keyed by its id:
- Remove it from the Study with the rest of that Study's set.
- Write no receipt and no delta.
- Mark every family the entry could touch Stale (
MarkFamilyStale). - Insert a bounded quarantine record into
pmProjectStatisticsFoldQuarantine: - key
(ProjectId, StudyId, OperationNamespace, OperationId, QuarantinedAtUtc); - a typed
Reason:DuplicateReceipt,ConflictingReceipt,DuplicateDelta,RetiredSourceRevision,UnknownProtocol,UnknownKind,Undeserializable,QuarantineExists,Overflowed,UnrecordedDropsorLivelocked; - the
SubmissionDigestwhen it can be read. A copy of the entry is kept only while the project has fewer than 1,000 records, and records beyond that keep the id and digest only, about 150 bytes; - an index on
(ProjectId, OperationNamespace, OperationId)for the save's duplicate check. - Retention (S7), one rule. A TTL index on
QuarantinedAtUtcremoves records after 30 days, and nothing else ever deletes a record. Duplicate detection therefore always covers at least 30 days, far beyond any retry or redelivery window. - Growth bound, fail-closed (S-R4-4). Before inserting, the fold counts the project's records
with an indexed
countDocumentslimited to 100,001. It is index-only and runs only on batches that quarantine something. Overflow is not a breach (it happens whenever the worker is down), so the bound must never cost detection:- At 90,000 records (soft cap), the fold transaction that crosses it sets the project's
FoldModetoDisablingand marks every fold family Stale. Writers stop appending within the 30-second pre-write deadline. The remaining 10,000 records of headroom are far more than the pending sets in flight can need: at most 32 entries plus 64 recorded drops per Study, and the sets drain. - At 100,000 records (hard cap), which needs the headroom itself exhausted, the fold stops
inserting and writes a durable per-project marker,
control.FoldQuarantineIncompleteSinceUtc, in the same transaction. While that marker is set, every unknown write result and every duplicate check that finds nothing on that project answersStatisticsOperationUnresolved, never Saved or a retry. Ids consumed without a record therefore remain fail-closed.fold.quarantine_capacity_exhaustedis raised. The marker is cleared only by an operator, after the TTL has brought the count below the soft cap and aresethas run.
- At 90,000 records (soft cap), the fold transaction that crosses it sets the project's
- Record
fold.quarantined{reason}.
Poison detection happens before any insert, from the snapshot's $in reads. A conflicting insert
from a racing legacy writer that lands after the snapshot raises a write conflict or a duplicate key
and aborts the transaction. The retry sees the new receipt and quarantines. No entry can make every later batch
abort, and the project keeps folding.
Conflicts with concurrent writers¶
- Different Studies: never conflict with each other or with the fold. This is the property the benchmark says is missing today.
- A same-Study save commits first: the fold's Study update raises a write conflict, and the fold retries from a new snapshot that includes the new entry. Only the worker retries.
- The fold commits first: the save's compare-and-set fails. The reload shows a fold-only change, which is a free retry.
- Version-guarded upsert writers that use shared cached instances:
- Generic
SaveAsyncupserts on{ _id, Version }, so a stale version raises E11000 rather than a clean miss. The affected writers areNotificationHub,MarkSessionIdleConsumer,RemoveIdleSessionConsumerandCheckConnectionLivenessConsumer. - These already lose to concurrent saves today; the fold adds a second bump shortly after each
save. Slice 1 gives each of them reload-and-retry on E11000 (bounded to 3), and treats a
fold-only reload as free, only while
MaterializedProjectStatisticsFoldis on. With the flag off the behaviour is byte-identical. - The 500 ms minimum-age rule keeps most folds out of the window just after a save.
- The same-Study benchmark cells, including the reviewer-correction pattern (save, then a correction within about 1 s), quantify what remains.
Crash and restart¶
- Atomicity: the fold is one transaction. A crash before commit leaves every entry pending.
- Unknown commit: re-commit on the same session. If that is exhausted, re-read: consumed entries are either gone (committed) or still present (not committed).
- Takeover: a takeover never sees a half-applied batch.
Families and scopes with no row¶
- Missing scope rows. Take a derived move for a scope with no visible row, for example a member who joined after the last backfill, or an investigator's first session in a stage. The fold inserts a published Stale placeholder row for that scope (S4):
- it sits at the guard's visible generation and token, with
IsPublished = true,State = Stale, zero counters andLastFailureReason = FoldMissingBaseline; - the move itself is not applied;
- published matters. Publication generations are per family. An unpublished row stops occupying
the visible key as soon as a sibling scope is republished. A published Stale row stays the
"greatest published at or below" selection, and
EvaluateRowrefuses it, until the scope's own rebuild replaces it; - if the exact visible key is already occupied by an unpublished candidate, the fold does not
insert. It runs
MarkFamilyStalefor that family instead; DuplicateKeyon the placeholder insert is a retry from a new snapshot, like every other fold insert.- Why this and not "start from zero": a legacy transactional writer, whether a flag-off pod in
fold mode or an older binary, would otherwise start that row from zero over the missing move
(
IsServableBaselinelets non-RequireExistingFreshfamilies do so) and stamp a Fresh undercount. With a Stale row present, the legacy writer sees a non-servable baseline and routes source-only, which also stales the scope. The undercount can never be served. - The reader falls back for that scope. The scheduled repair or a backfill republishes it with
authoritative(B) - pending(B). - No start from zero, while it matters (S5). The fold never starts a row from zero. The new
coordinator refuses to start from zero for a fold family only while
FoldMode != Disabledor the project's pending set is non-empty; in that case it routes source-only and inserts the same placeholder. AfterDisabledwith nothing pending, starting from zero is exact again, as it is today. Even in fold mode, a later fold onto a row created from zero is exact. Only a folded move that found no row needs refusing, and the placeholder provides that. - Undeclared scopes (F2). A MembershipStageAnnotation scope for a session investigator who is not
a member is never enumerated by the family's backfill, so its placeholder is never republished. Such
a placeholder carries
LastFailureReason = FoldMissingBaselineUndeclared. The scheduled repair backs it off until the membership context digest changes, as it already does for its other answers, instead of retrying it every run. It is never read, because readers only request member scopes. - Missing guards. Fold-mode enable records every fold family's control lifecycle as Stale, whether or not a guard exists, and creates a Stale guard for each family that lacks one. No Missing-guard family can then be point-materialized from zero by a legacy writer on a skewed pod.
History and client-invalidation timing¶
- Clocks advance per fold batch. Checkpoint identity rules are in checkpoints.
- Clients are invalidated after the fold. Reads are exact before that, so an early refetch is still correct. The added delay for other viewers is the fold lag, typically under one second, on top of today's one-second dispatcher tick.
- No early invalidation on save. It would need a hot shared write or a second per-save publish. The saving reviewer's own client refreshes as part of its flow, and every page polls every 30 seconds.
- If the pilot measures a fold lag above 2 seconds p95, a coalesced early publish from the API coalescer can be added without any protocol change.
4. The reader overlay¶
Algorithm¶
This runs in ProjectStatisticsBundleReader.ReadScopesAsync, in the same pinned snapshot
transaction (ReadConcern.Snapshot, primary) that already reads the controls, guards, rows, fences
and authorization context. An unpinned snapshot still refuses. The overlay is consulted for every
requested fold-family scope whenever control.FoldEverEnabled is true, including after Disabled,
so a late entry can never be skipped. When the project has never been in fold mode it costs nothing.
- The project gates and scope gates are unchanged, except that versions are compared through
ProjectStatisticsFleetVersions.Matches(5). - Read the pending set:
pmStudy.find({ ProjectId, "PendingStatistics.OldestAtUtc": { $exists: true } }, { PendingStatistics: 1 }).limit(MaxPendingStudies + 1), using the partial index. - Fall back authoritatively for the whole bundle (
PendingOverlayUnavailable, with a sub-reason tag) if any of these holds: - more than 128 Studies or 1,024 entries are pending;
- any overflow marker exists;
- the oldest entry is more than 10 minutes old;
- any entry whose own
ProtocolVersionis outside the binary's accepted window (its protocol and, under N-1, the previous one), or with an unknown transitionKind; - live fleet members disagree on whether the project is allowlisted
(
AllowlistInconsistent; see fleet membership). - any entry whose context digest, for a requested family's group, differs from the snapshot Project's;
- any entry carrying an invalidation intent that covers a requested scope;
- any entry whose
AssumedTrackingModediffers from the durable mode. - For each requested scope, clone the selected row and apply
PendingEntryMoves.Derive(entry, context, [scope])for every entry with the sameTryApplyMovesthe fold uses. TryApplyMoveskeeps keys that reach zero and does not decrementDynamicKeyCount, exactly as the point path does today. Consumers and parity already normalise zero cells.- A capacity refusal or any negative counter in the overlaid clone is an invariant violation:
fall back and record
overlay.invariant_violation. - The bundle carries the overlaid clone in place of the stored row. Every consumer that reads
row.Counterssees the overlaid values, never stored-only ones. This includesProjectStatisticsDerivedSummaries,ProjectScreeningCounters.ToStatsandStageOverviewStatisticsQuery.
The coherent-bundle rule is unchanged: one failing scope sends the whole bundle to fallback.
Read identity¶
The bundle's CheckpointId still names the stored baseline
(SourceInvalidationRevision, CommittedProjectionRevision, ModeEpoch). The result gains a
PendingFingerprint: entry count, overflow flag, and SHA-256 of the sorted entry ids. Parity, drift
and checkpoint code use it (6).
The wire DTOs are unchanged. The coherent web pages reject only a lower client revision, so an equal revision with a different overlay is accepted.
Consistency with the serving gate¶
EvaluateRowis unchanged and rows carry the same identity as today.LastChangedRevision <= CommittedProjectionRevisionkeeps its meaning. Folds, rebuilds and the remaining transactional writers each stamp the revision they commit.- Invariant kept: a stale or incompatible scope is never served because its values look plausible. The overlay only adds refusals.
Cost¶
| Case | Added per read of fold families |
|---|---|
| Steady state (10 reviewers, about 2 writes per second, fold lag under 1 s) | One indexed find returning 0-3 small documents, about one round trip (0.3-1 ms locally). The Fresh read p95 is about 12 ms (read benchmark). Derivation is microseconds |
| Worker lagging, at the caps | 128 documents and 1,024 entries, roughly 1 MB projected in the worst case. Reviewer-scope derivation runs only for requested scopes. Beyond the caps the read falls back |
| Never in fold mode | Nothing |
5. Fold mode, the rollback tripwire and version identity¶
Durable fold mode¶
New fields on ProjectStatisticsControl (unknown fields are preserved by ExtraElements):
FoldMode:Disabled(default),Enabled, orDisablingwithFoldDisablingSinceUtc;FoldProtocolVersion: the stamped protocol;FoldEverEnabled;FoldEnabledAtUtc.
Enable runbook gate (S-R4-6). Before the first enable on a deployment, and after every image
change, the operator confirms that both the api and project-management rollouts have completed (ArgoCD
Healthy, and kubectl rollout status read-only showing no old ReplicaSet pod alive). Pre-fold pods
write no fleet membership, so the code check below cannot see them.
Enable is POST api/admin/project-statistics/{id}/fold/enable. It refuses with:
PendingIndexMissinguntil the index exists;FoldFleetUnsupportedunless every live member of both deployments (see fleet membership) declares the protocol being stamped, none declares a lower one, and the fold-worker heartbeat for that protocol is live;FoldCoverageIncompleteunless one of these holds:- the binary's compile-time
PendingEntryCoverage.Completeis true (slice 6 onward); - all of the following hold (owner decision g):
- the deployment sets the allow setting
ProjectStatistics:Fold:AllowPartialCoverage = true; - the API host's
IHostEnvironment.IsStaging()is true. This is a positive check: staging setsASPNETCORE_ENVIRONMENT=Staging, while production and preview resolve toProduction(S13). A test pins it; - the configured database is not the production database (
syrftest); - the binary covers at least ProjectScreening (slice 2 onward).
- the deployment sets the allow setting
Neither condition can be satisfied by a runtime flag override.
Enable then runs one small transaction. It:
- sets
FoldMode = Enabled,FoldProtocolVersion,FoldEverEnabled = trueand the tripwire (below); - runs
MarkFamilyStalefor every fold family, and creates any missing guard as Stale; - advances
SourceInvalidationRevisionandClientInvalidationRevision, and coalesces every fold family's slot.
Ordinary non-forced backfills, the scheduled repair or the fleet runner then republish every fold
family with the subtraction rule. Enable needs no forced rebuild, so RecordVersions and sibling
identity replacement never run because of it.
Under partial coverage (the staging pilot), families whose writers are not yet migrated keep their legacy transactional writers. Those writers move rows directly, which invariant 1 allows. The migrated screening saves carry those families as invalidation intents, so they stay Stale and are served authoritatively.
Disable is POST …/fold/disable. It sets Disabling, and from then on saves take the
transactional path. The worker finalizes Disabled in a fold transaction only when:
- at least
FoldWriterGrace(5 minutes, longer than the 30-second pre-write deadline) has elapsed; - the pending set in its snapshot is empty.
The same transaction clears the tripwire. The reader and the worker keep handling any entry that
somehow appears later, because both run while FoldEverEnabled is true.
Reset¶
POST …/fold/reset is the guard-exit and recovery operation. Its preconditions and effect depend
on the prior mode (S6):
| Prior state | Preconditions | Effect, in one transaction |
|---|---|---|
FoldMode is Enabled or Disabling (including a cleared or other-protocol tripwire) |
Every enable refusal applies: the index, the fleet check and coverage with its staging-only rule. Only the allowlist is waived, so a reset can run inside a guard window, and it cannot bypass owner decision (g) | Re-stamps the tripwire at the binary's protocol, keeps FoldMode as it was, runs MarkFamilyStale for every materialized family of the project, SearchPopulation included (NM1), advances the source and client clocks and coalesces every slot |
FoldMode is Disabled |
None beyond administrator authorization | Keeps Disabled, leaves no tripwire, and stales every materialized family and advances the clocks, as above |
Staling every family, not only the fold families, is required. A non-allowlisted project gets no
statistics bookkeeping at all: no fences, no stales and no clocks. That covers the staged
import/removal and bulk fences (ProjectStatisticsStagedOperationFenceService answers NotRequested),
the source-visibility and definition-rewrite fences, and every writer. Any family, including
SearchPopulation after a search import during the window, may therefore have drifted.
reset is the recovery for each of these:
- an old binary's forced rebuild having cleared the tripwire. That row is
authoritative(B)with nothing subtracted. Marking the family Stale means the fold discards against it rather than adding to it, and the rebuild republishesauthoritative(B') - pending(B'); - the end of a guard window (below);
- a protocol jump that the N-1 rules do not cover (staging only, between MVP slices).
The fold protocol version and N-1 compatibility¶
FoldProtocolVersion is a compile-time constant. It is bumped by every change that adds meaning
to an entry: a new transition Kind, a new payload field the derivers read, or a new intent shape.
Durable identity.
- The control: the tripwire value is
global.StorageVersion + FoldStorageOffset + control.FoldProtocolVersion, withFoldStorageOffset = 1000. - The fleet: the heartbeats and fleet membership carry the protocols each pod writes and folds.
- Each entry: carries its own
ProtocolVersion.
Owner decision (h): previous-protocol (N-1) support. A binary at protocol N:
- accepts a control stamped
NorN-1(Matches, below); - reads, overlays, folds and rebuilds entries at
NandN-1, deriving each under its own protocol's rules; - writes entries at the control's stamped protocol. That is
N-1while the stamp isN-1; it never writesNbefore the stamp advances; - quarantines entries at
N-2or older (UnknownProtocol), and the reader falls back on them; - renews fold-worker heartbeats for both
NandN-1.
Allowed changes between N-1 and N (enforced by review and by the compatibility test matrix):
- additive only: new transition kinds, and new optional payload fields with defined absent-value semantics;
- no removed or re-typed fields, and no change to the meaning or derivation of an existing kind at
N-1; - a change to a
Contextdigest's inputs requires theNbinary to compute both digest forms, and to compare each entry against the form for that entry's protocol; - while the stamp is
N-1, anNbinary must express its new behaviour withN-1constructs, usually an invalidation intent, because it cannot writeNentries yet.
Limits of N-1 (F-R4-1).
- A derivation-changing release is a breaking bump. This covers a classifier bug fix, or any
change to how an existing kind derives. It cannot use the N-1 window, because
N-1entries would derive differently on the two binaries. It is handled as follows: - ship protocol
Ndeclaring the affected kinds as breaking; Nbinaries treat pendingN-1entries of those kinds as discards (stale the affected families), never deriving them;- after the stamp advance, run
resetfor each enabled project, followed by backfills of the affected families.
Only that cost is incurred. Additive bumps keep the no-reset path.
- "An upgrade needs no reset or rebuild" holds only when nothing older is pending. Entries still
pending at N-1 when the next bump (N → N+1) advances would be N-2 to an N+1 binary. The stamp
advance's pending-protocol precondition therefore waits for the worker to drain them. A project
whose worker is down simply advances later.
The rolling deploy from N-1 to N:
- During the rollout the stamp stays
N-1. N-1binaries acceptN-1(their own protocol).Nbinaries acceptN-1as their previous. Both writeN-1entries and both fold them.- The heartbeat for
N-1stays live throughout, because every pod renews it. - A rollback during this window is seamless. No
Nentry exists yet, so nothing needs a reset. - Stamp advance (S-R4-7). The owner is the project-management consumer
ProjectStatisticsFoldStampAdvanceConsumer, triggered by the fold sweep's one-minute schedule, withConcurrentMessageLimit = 1. It also has an administrator route that runs the same code. - Fleet precondition. In one read with server time it checks that every live member of
both deployments (liveness) declares
N, that none declares lower, and thatheartbeat[N]is live. - The runbook's rollout-complete gate must also have been recorded for the deployment, by
setting
ProjectStatistics:Fold:StampAdvanceAllowed = truein the values after both rollouts report complete, because pre-fold pods are invisible to the fleet check. - Per project, CAS. Filter:
control.FoldProtocolVersion == N-1, theN-1tripwire value and the control'sAudit.Version. Set:FoldProtocolVersion = Nand theNtripwire. - Pending-protocol precondition. The project has no pending entry below
N-1. Such an entry is impossible after a completedN-2 → N-1advance, but the precondition is cheap and is checked anyway. - It needs no stale, reset or rebuild, because pending
N-1entries remain valid forNbinaries. From then on, writers writeN.
Enable and reset during a mixed fleet stamp the lowest protocol declared by any live member
that the binary accepts: N-1 while any N-1 member is live. They no longer refuse until the
rollout completes.
3. After the advance, an N-1 binary sees the stamp N, which is outside its window
(N-2..N-1), and fails closed exactly like a pre-fold binary. Rolling back past an advanced stamp
therefore uses the rollback runbook.
When N-1 support starts. It is required from the first protocol bump after production
activation. The protocol that slice 6 ships with is the production baseline P0, and every later
bump must ship with N-1 support for P0 onward.
Between MVP slices 2-5 (protocols 1-4), each slice may change kinds freely, and the staging pilot uses
reset (with its rebuild) after each slice deploys. Recommendation, adopted: the slice sequence
is still shaping the entry vocabulary. Carrying N-1 writers and derivers for intermediate shapes would
roughly double each slice's test matrix for a single pilot project, where a rebuild costs seconds.
Production carries no fold project until P0.
Test matrix (from P0 onward). Every deriver, and the fold, overlay and rebuild paths, run at both
N-1 and N:
- each with entries of both protocols in one pending set;
- an
N-2entry is quarantined; - a mixed
N-1/Nfleet simulation keeps the stamp atN-1and stays exact; - the stamp advance requires full membership;
- an
N-1binary refuses a project stampedN.
The storage-version tripwire¶
What pre-fold binaries (and fold binaries outside their protocol window) do with the tripwire set:
| Old code path | Behaviour |
|---|---|
ProjectStatisticsServingGate.EvaluateProjectGates, used by every current reader, ProjectStatisticsCurrentRowCopier (211) and the drift check (584, 872) |
Incompatible for every family of the project, so it serves, copies and drift-checks nothing |
ProjectScreeningBackfillService (301, 346, RefuseOnControlVersionMismatch), inherited by every family backfill; EnsureControlAsync (adoption needs StorageVersion == 0) |
Refuses |
ProjectStatisticsStaleRepairRunner (397) |
FleetVersionMismatch, skip |
ProjectStatisticsCheckpointBuildService and ProjectStatisticsUnchangedDayRecorder |
Do not compare StorageVersion. In a pre-fold binary they are kept away from fold projects by the allowlist guard and the enable rollout gate, not by the tripwire |
| Old point writers (coordinator) | Do not compare control versions. They move a Fresh row directly, which is exact under invariant 1, and stale dependants correctly. They cannot start a row from zero over a fold-dropped move, because of the published placeholder. They can never make a row servable to readers that check the tripwire |
The narrow-gate re-enable (POST …/narrow-gate {enabled:true} → RebuildForReEnableAsync) |
Checks neither the allowlist nor StorageVersion. An old binary would publish authoritative(B) without subtraction. The runbook forbids narrow-gate transitions on old binaries during a window, and the mandatory reset stales its result anyway (S10) |
The old forced rebuild (/rebuild, RecordVersions) |
Would clear the tripwire. It is refused by the allowlist guard (Admission answers NotAllowlisted first); if it is run without the guard, reset recovers |
History writers.
- New history writers (fold-aware binaries) check
Matches, and they never record a value that ignores pending entries. Copies need an empty pending set. An authoritative build with entries pending first commits an observation barrier. An unchanged day needs unchanged clocks and an empty pending set (checkpoints). A fold-worker outage therefore cannot corrupt history written by new binaries. - Old history writers run only in a pre-fold project-management binary. They meet a fold project in exactly two situations, and each has a gate:
- a rollback. The allowlist guard is applied and fully rolled out before any image change,
and every existing history writer is reached only through allowlist-gated entry points:
ProjectStatisticsDailyObservationRunner(42-43),ProjectStatisticsDailyObservationProducer(66, 147) andProjectStatisticsBackfillService(146); - an enable while an old project-management pod is still alive. The enable runbook's rollout-complete gate excludes it, and fold-aware members of a lower protocol are refused by the fleet check.
- A runbook violation (a project-management rollback without the guard, while fold projects are enabled) can let an old recorder record an unchanged day over unfolded entries. This is bounded to the daily history of those fold projects for the length of the violation. Current values stay correct, because old readers fail closed on the tripwire and new readers overlay. It is listed in the failure modes.
Version comparison rule at every site¶
control.SourceVersion, control.CatalogueVersion, the global control and every row's
SourceVersion and CatalogueVersion stay at today's values. Rows never carry a fold marker.
New code routes every control-versus-global storage comparison through one helper,
ProjectStatisticsFleetVersions.Matches(global, control). It accepts:
control.StorageVersion == global.StorageVersionwhenFoldMode == Disabled;== global.StorageVersion + FoldStorageOffset + control.FoldProtocolVersionwhenFoldMode != Disabled, andcontrol.FoldProtocolVersionis the binary's own protocol or its previous one.
Anything else is Incompatible: a tripwire whose mode, protocol or offset disagrees, or a cleared
tripwire while FoldMode != Disabled.
| Site | Fold-mode rule |
|---|---|
ProjectStatisticsServingGate.EvaluateProjectGates (global vs control) |
Matches |
ProjectStatisticsServingGate.EvaluateRow (row vs control) |
Unchanged; rows carry plain versions |
ProjectScreeningBackfillService 301, 346 and RefuseOnControlVersionMismatch 564-590, and the per-family subclasses |
Matches; writer constants unchanged |
ProjectStatisticsStaleRepairRunner 397 |
Matches |
ProjectStatisticsRebuildService.PublishAsync 1202-1208 (RecordVersions when forced) |
Preserves the tripwire while FoldMode != Disabled |
ProjectStatisticsModeAdministrationService 103, 384, 392 (including the narrow-gate re-enable); ProjectStatisticsWriteEpochLifecycleService 98, 122, 158 |
Preserve the tripwire; the narrow-gate re-enable also checks Matches and the allowlist |
ProjectStatisticsBundleReader watermarks (567) |
Report the effective storage version (offset and protocol stripped) |
ProjectStatisticsCheckpointBuildService (446, 627-628, 857) and ProjectStatisticsUnchangedDayRecorder (124-136) |
Add a Matches check; neither compares storage today. History roots record the effective versions |
ProjectStatisticsCurrentRowCopier (211 gate, 260 row versions) |
Through the gate, so Matches; row versions unchanged |
ProjectStatisticsDerivedSummaries 167 (row.SourceVersion != 1) |
Unchanged; rows stay at 1 |
Rollback guard and runbook¶
Owner decision (f) requires two things: the runbook order, and a GitOps guard that existing binaries already honour.
The guard. The cluster-gitops values of both api and project-management remove every fold-mode
project from ProjectStatistics:ProjectAllowlist (SYRF__ProjectStatistics__ProjectAllowlist). Every
binary shipped so far reads that key once per process, and for a non-allowlisted project it:
- serves nothing materialized. The reader's project gate and every consumer's pre-snapshot gate check the allowlist;
- writes no statistics. Saves are plain source writes, with no fences, stales or clocks;
- refuses every backfill and admin rebuild, forced included (
AdmissionanswersNotAllowlisted); - skips daily observation, checkpoint builds, repair, drift, maintenance and the fleet runner.
The one exception is the narrow-gate re-enable route (S10), which the runbook forbids during the window. Nothing new needs to ship for the guard.
The guard is ordinary pod configuration, so its own rollout mixes guarded and unguarded pods. The
runbook therefore puts the durable, every-binary narrow gate around it (NM2). The narrow gate is
control.Mode (POST …/narrow-gate). Every binary's reader refuses a project whose control Mode
is not Enabled, and every binary's writer routes source-only, staling and advancing the clocks.
Rollback order:
- Disable fold mode. Run
POST …/fold/disablefor every fold project, and wait forFoldMode = Disabled(the grace has elapsed and pending is empty). - Close the narrow gate for each project (
POST …/narrow-gate {enabled:false}). From now on, every binary serves those projects authoritatively, whatever its allowlist. - Apply the guard alone in its own cluster-gitops PR. Wait until both deployments' rollouts
have completed: ArgoCD Healthy, and
kubectl rollout status(read-only) shows every pod on the new ReplicaSet. New-code readers additionally refuse while fleet members disagree (fleet membership). - Change the images in a separate PR. API rolls back no later than project-management. Make no narrow-gate transitions and run no forced rebuild while the window is open.
Roll-forward order:
- Deploy the new images with the guard still in place, and wait for both rollouts to complete.
POST …/fold/resetfor each project. It is accepted for a non-allowlisted project, and stales every materialized family.- Remove the guard. Wait until both rollouts have completed. During this rollout, still-guarded pods make plain source saves that move no clock. Nothing is served meanwhile, because the narrow gate is still closed, and the next step stales anything published in between.
POST …/fold/resetagain, now that every pod writes statistics. KeepProjectStatisticsRepair:Enabledoff from step 1 until this point, or accept that any publication it made is staled here.- Optional pre-warm. Running backfills while the narrow gate is closed is redundant, because step 6 rebuilds anyway, and a backfill may refuse while the project mode is not Enabled. If it refuses, skip this step (F-R4-2).
- Re-open the narrow gate on the new binary. Its re-enable runs the existing rebuild proof
through
ProjectStatisticsRebuildService(ProjectStatisticsModeAdministrationService220-226), which applies the subtraction rule. - If fold mode is wanted again,
POST …/fold/enable.
History gap (F1). The window's days have no daily observation for those projects, because every history writer skips a non-allowlisted project. The runbook records that as an expected gap, and the history API already reports missing days truthfully.
If the runbook is not followed, the tripwire still keeps pre-fold readers, backfill and repair fail-closed, and new readers keep overlaying exactly. Daily history for the window may be wrong (see history writers). Rows can still be wrong during an unguarded rollout, which is exactly why the narrow gate comes first. Recovery is the same: reset twice around guard removal, then backfill. A rebuild is always required after a window.
6. Interaction with existing mechanisms¶
Rebuild, backfill and the subtraction rule¶
A rebuild or backfill of a fold-family scope while FoldEverEnabled:
- Leases. It takes the scope lease and the project fold lease, using the yield request if the fold holds it.
- Drain under the lease. It marks the visible row
Rebuilding(today's step). It then runs fold transactions for the project. The row is non-servable, so entries for that family are discarded, while other families' entries apply normally. It repeats until no entry that existed at the start remains. Each discarded entry's Study write committed before the snapshot taken next. - Calculate in one snapshot
B. The family's scope calculator already reads the Project, the control and the source in one pinned read-only snapshot. It now also reads every pending entry of the project in that snapshot. The read is paged by_idthrough the partial index, with no cap; it is bounded only by the snapshot lifetime, and a timeout returns a retryablePendingNotDrained. It derives the scope's moves and returnsauthoritative(B),pending(B)and thePendingFingerprint. - Publish
authoritative(B) - pending(B). Entries created after the drain but beforeBare in both terms, so they cancel. Entries afterBstay pending. Rows keep plain versions. - Refuse with retryable
PendingNotDrainedifpending(B)contains an overflow marker, poison, an unknown protocol or kind, or an entry whose context digest or tracking mode differs fromB's for this family. Such entries can only come from a write racing a configuration or mode change. The worker discards them while the row is non-servable, and the next attempt publishes. - Keep today's publication CAS. It still checks clocks, the lease, the candidate token, fences, tokens, the mode epoch and the digest. Fold-path writes do not advance the control clocks, so a write during a rebuild no longer makes it lose. A legacy transactional writer still advances the clocks and still makes it lose, as today.
This closes the saves-during-a-backfill gap only for the point-maintained families. In the final MVP those are:
- ProjectScreening;
- MembershipScreening and ReviewerScreening, for saves with before-state evidence and at most 100 memberships;
- StageAnnotation, MembershipStageAnnotation and DomainReconciliation.
It does not close the gap for the intent-only families:
- ReviewerAnnotation, which today's
DependentFamiliesstales on every annotation save that moves allocation or inclusion, and otherwise for the affected reviewers; - QuestionAnswers, staled on every content change;
- MembershipScreening and ReviewerScreening for saves without evidence.
A rebuild of an intent-only family publishes Fresh, and the next fold of an entry carrying its intent stales it again. That is exactly today's behaviour. Making those families point-maintained is a separate catalogue decision (deferred item 6).
The helpers change to match:
IsAlreadyCurrentAsynccompares the authoritative value with stored + pending in one pinned snapshot;VerifyCompletedScopesAsyncverifies through the overlay;- the fleet runner, scheduled repair, sibling rebuilds and the admin routes all go through
ProjectStatisticsRebuildService, so they inherit this protocol with no change of their own.
No double counting. Every entry falls into exactly one of these cases:
- discarded before
B: counted only inauthoritative(B); - pending at
B: in both terms, so it cancels; - created after
B: counted only when folded.
Fences¶
| Fence | Behaviour |
|---|---|
| Staged operation fences: search import parse, complete and fail (#3842); bulk Study update; search removal; question-tally refresh | Admission and release are unchanged. Fold-path writes keep appending; they read no guard. The fold sees ActiveFenceCount > 0 and discards for the fenced families. The release leaves them Stale, and the next rebuild uses the subtraction rule. Bulk writers under a fence may preserve, rewrite or drop PendingStatistics on the Studies they touch, because nothing in a fenced family is served and its rebuild re-derives from the Studies |
Source-visibility fences: InclusionRecalculationToken (#3847), DefinitionRewriteToken |
Reads answer the typed 503 before the overlay is consulted. The fold discards for the affected families while the token is held. Entries classified under the old thresholds carry the old context digest. If one is still pending at the post-fence rebuild, that rebuild refuses PendingNotDrained until the worker discards it |
InvalidateDefinitionChangeInTransactionAsync (stage SessionCountTarget) |
Unchanged. It is a Project-save transaction that stales ReviewerAnnotation. The annotation context digest includes SessionCountTarget, so entries classified before the change no longer match afterwards: the reader falls back and the fold discards |
Stale repair, drift check and parity audit¶
- Scheduled repair is unchanged apart from
Matches. Its settle check sees the clocks moved by folds. A lease refusal is an ordinaryBusy, retried on the next run. - Drift check compares the overlay with the authoritative calculation. Both reads must match
on the existing identity triple plus
PendingFingerprint; a moved fingerprint is inconclusive, never drift. A real mismatch goes through the existingMarkDriftedAsync. - Parity audit (
ProjectScreeningParityAuditServiceand its family siblings) reads through the audit reader with the overlay, and extends its identity withPendingFingerprintin the same way. AstoredOnlydiagnostic reports the pending count and the stored-only difference; it is never a verdict.
Delta ledger and maintenance¶
- Ledger capacity. The fold runs
ProjectStatisticsRevisionLedgerAdmission.IsAtCapacityAsync(100,000 rows or 90 days) in its snapshot. At capacity it discards withDeltaLedgerCapacityExceeded, which marks the families Stale. This is the same fail-closed behaviour the coordinator has today. - Delta maintenance (
ProjectStatisticsDeltaMaintenanceService) treats a fold lease as non-blocking, because it captures no revision. Its one transaction CASes the same control and guards as the fold, so the two serialize through ordinary write conflicts and the loser retries. A live rebuild lease still makes maintenance refuse. - Compaction (
ProjectStatisticsDeltaCompactionService) ignores fold leases when pinning its floor. They have noCapturedProjectionRevision, and the separate scope key means a rebuild takeover never inherits one.
Checkpoints, daily observations and history¶
History must stay free of pending sets, and each identity must name exactly one value.
- Copied checkpoints.
ProjectStatisticsCurrentRowCopiercopies a row only if the project's pending set is empty in its pinned snapshot. Otherwise that scope is calculated authoritatively. - Authoritative builds: daily observation, the checkpoint build service and the backfill's
bootstrap checkpoint (
RecordBootstrapCheckpointAsync). They hold the project fold lease from capture to publication. If the pending set is non-empty at capture, the build first commits an observation barrier: a one-document control transaction that advancesSourceInvalidationRevision. The authoritative values, which already include every pending Study's state, are then labelled with aCheckpointIdthat no other build can share. Two builds can therefore never publish different values under one identity, even while the worker is down. The bootstrap checkpoint follows the same rule. - Unchanged days.
ProjectStatisticsUnchangedDayRecorderrecords an unchanged day only when the clocks are unchanged and the pending set is empty. - No history format change. Roots, pages, observations,
DeltaWatermarkRevision, retention and seals are unchanged. Fold batches produce ordinaryDeltasibling groups.
Project deletion, Study deletion and bulk updates¶
- Project deletion answers the typed 503 today, and
DeleteProjectAsynchas no caller, so nothing changes. The future deletion scheduler must set the tombstone first. The reader already refuses a deleted project, and the fold must discard for one. - Study deletion happens only in fenced paths: search removal, import failure and parse cancellation.
- Bulk updates are fenced. The unfenced
UpdateManywritersBackfillWorkloadShareBucketsAsyncandMarkBulkPdfDeliveredAsyncnever touchPendingStatistics.
Multiple pods and flag skew¶
- API pods need no coordination.
- PM pods compete for fold commands. The lease and the fencing write give one writer per project.
- A pod with the fold flag off in a fold-mode project takes today's transactional path. The new coordinator also checks the loaded Study's pending set, and the quarantine record, for its envelope id, so the duplicate rules hold across paths. It moves Fresh rows directly, which is exact under invariant 1. It never starts a fold-family row from zero (missing rows).
- The durable
FoldMode, the protocol-stamped tripwire and the per-pod fleet membership decide what readers, writers and the worker do. Process flags do not. The server-time heartbeat gates enable and the stamp advance, and drives the worker-down alert, but never a save.
What does not change¶
- The authoritative calculations and their fallback delegates.
- Authorization and the durable-mode 503.
- The control-digest authority (
ForThresholds). - The global control's write epochs and mode epoch.
- The history formats and retention.
- The outbox dispatcher,
ProjectStatisticsChangedand the web client's dedupe and refetch. - The capacity and reservation guards.
- The inclusion operation-id bytes (
screening.v1). - SearchPopulation, which has no point writer.
7. Invariants¶
These replace, for fold families in fold mode, the point-mutation bullets of the Lifecycle and the point-save items of the Transaction and idempotency rules.
- Exactly one place. Every committed point operation's effect on a fold family is either:
- in that family's stored rows, applied once by a fold or by a transactional writer;
- in exactly one pending entry on its Study;
- or the affected family or scope is non-servable until a rebuild re-derives it. That covers overflow, discard, poison, a missing row (Stale placeholder) and an invalidation intent.
No move is ever silently dropped while a Fresh row could serve without it.
2. Served value. A fold-family value served from storage is the overlaid clone: stored row plus
every pending entry's derived moves, from one pinned snapshot, whenever FoldEverEnabled.
3. Conservative refusal. The overlay refuses on overflow, context or tracking-mode mismatch, an
invalidation intent for the scope, an unknown protocol or transition kind, the count or age caps, a
negative sum or a capacity refusal.
4. Hot-path isolation. A fold-path write writes one Study document plus whatever non-statistics
documents the same writer already writes today (eligibility Project token, presence rows), and
no pmProjectStatistics* document and no Project statistics-admission token.
5. Row writers. Stored rows are written only by:
- the fold holder (lease plus fencing write);
- a rebuild holding the scope lease and the fold lease;
- transactional coordinator writers (flag-off pods and the administrative writers listed in 2).
Concurrent writers serialize through write conflicts on the row, and stored + pending stays exact.
6. No resurrection. Removing pending entries always bumps Audit.Version and
StatisticsFoldSequence, so no stale full-document replace can restore a folded entry.
7. Unique ids, digest-checked. A pending entry id is never appended if the same id is pending,
overflow-dropped, received or quarantined, or if its envelope revision lies in an unrecorded-drop
range. The unrecorded-drop case is always answered StatisticsOperationUnresolved, and so is any
miss while the project's quarantine-incomplete marker is set. A matching digest answers Saved and writes nothing; any
other digest fails closed. The same rule decides an unknown write result. A changed decision under
a reused id is never reported Saved.
8. Poison is consumed. An entry that cannot be applied or discarded normally is removed without
inserting anything keyed by its id, its families are staled, and it is quarantined. No entry can
stop a project's fold.
9. Rebuild subtraction. A fold-mode publication equals authoritative(B) - pending(B), where
pending(B) is the complete pending set at B with no overflow, poison or mismatched entry,
published under both leases.
10. Clocks. Fold-path writes advance no statistics clock. A fold batch advances the source and
client clocks once, and the committed clock once if anything applied. It writes at most 32 delta
records at one revision.
11. Receipts. Every consumed non-poison entry gets exactly one receipt, in the transaction that
removes it.
12. Old and out-of-window code fails closed. While FoldMode != Disabled, the control carries the
protocol-stamped storage tripwire. Every pre-fold reader, backfill, repair, copier and drift check,
and every fold binary whose window (its protocol and the previous one) excludes the stamp, treats it
as Incompatible. History writers that do not check it are covered by invariant 16 and the guard.
Rows and history never carry a fold marker.
13. Mode-transition grace. Any owner of a durable reviewer-mode transition must wait
FoldPreWriteDeadline + FoldWriterGrace after stage one before treating the fleet as
transitioned, and a fold-path write never issues its Study write more than
FoldPreWriteDeadline after its pre-write reads.
14. Clean callbacks. No bus publish, SignalR send or other external effect happens inside a write
or inside a retryable fold callback.
15. Clean history. Every published checkpoint names exactly one value. Copies need an empty
pending set, and authoritative builds with pending entries first advance the source clock.
16. Worker outages never fail saves. Fold-path writes do not depend on the fold worker's
liveness. During an outage, readers stay exact through the overlay up to the age and size
thresholds, then fall back. History writers never record a value that ignores pending entries.
Writers write only the stamped protocol.
17. Rollback discipline. A rollback window is bracketed by the closed narrow gate. The allowlist
guard is applied and fully rolled out before images change, and is removed only between two
resets that stale every materialized family. New-code readers, repair and the worker refuse a
project while live fleet members disagree on its allowlisting.
18. Whole pending set. Every write to PendingStatistics maintains the whole sub-document,
OldestAtUtc included. A bare $push is forbidden.
19. The claim stays atomic. The fold-path claim is today's single non-transactional
FindOneAndUpdate with an unchanged filter and one prepended, non-failing pipeline stage. Its
outcomes are identical to today's. Its entry derives against the real before and after profiles,
with the seeded tally and its ReviewerAnnotation intent derived identically by the fold and the
reader.
20. N-1. From P0 onward, a binary at N reads, overlays, folds and rebuilds N and N-1
entries, writes only the stamped protocol, quarantines N-2 and older, and advances the stamp
only when every live member declares N.
Explicitly not changed:
- every compatibility invariant in
.claude/rules/materialized-stats.mdfor non-fold families; - the rule that a Stale row becomes Fresh only through a rebuild (a fold never makes a non-Fresh row Fresh);
- the coherent-bundle fallback;
- source-visibility fence semantics;
- the durable-mode 503;
- the telemetry rules: one meter, no identifiers in tags.
8. Failure modes¶
| Failure | Effect | Why it is safe | Recovery |
|---|---|---|---|
| Lease lost mid-fold | The fencing write matches nothing, so the transaction aborts | Nothing commits without the lease | The new holder folds from a fresh snapshot |
| Entry from an obsolete configuration, membership or mode | Context or mode mismatch | The reader falls back; the fold discards and stales; the rebuild refuses until it is drained | Ordinary Stale path |
| Overflow | The write succeeds and the marker is set | The reader falls back. The fold discards and stales the touched families, and moves the dropped ids into quarantine with their digests | Stale path; pending_entries.overflowed |
| Duplicate or conflicting receipt or delta, retired revision, unknown protocol or kind, undeserializable | Poison | Consumed without keyed inserts; families staled; quarantined | Stale path; investigate fold.quarantined |
| Hot Study conflicts every fold batch | Livelock risk | The batch is halved down to one, the Study is skipped for 2 s, and after 20 aborts it goes through the two-step discard (entries and overflow) | Stale path |
| Racing receipt or delta insert after the fold snapshot | DuplicateKey |
Classified as retry from a new snapshot, then poison | None needed |
| Move for a scope with no row | — | A Stale placeholder row; legacy writers then route source-only and cannot start from zero | Repair or backfill |
| Rollback within the N-1 window (stamp not yet advanced) | Mixed N-1/N fleet |
Both accept the stamp and write and fold N-1 |
None needed |
Rollback past an advanced stamp, or to a pre-P0 slice binary |
Protocol outside the window | Fails closed exactly like a pre-fold binary; its worker does not touch the project | Rollback runbook; reset |
| Fold worker down (PM outage or consumer stopped) while API is new | Entries accumulate; the heartbeat ages | Saves keep appending and never fail. Readers stay exact through the overlay, then fall back after 10 minutes, 128 Studies or 1,024 entries. Overflowing Studies stale their families. New history writers wait for an empty pending set or commit a barrier | Restore the worker; fold.worker_heartbeat_age alerts after 2 minutes; the repair rebuilds anything staled |
| Project-management rolled back without the guard (runbook violation) | An old daily recorder may meet unfolded entries | Current values stay correct (the tripwire and the overlay). Only that project's daily history for the violation window may record an unchanged day over unfolded changes | Apply the guard; treat that window's history as untrusted; the roll-forward runbook as usual |
| The guard's own rollout (mixed guarded and unguarded pods) | Some saves write no statistics | The narrow gate is already closed, so nothing is served. New-code readers refuse on allowlist disagreement. Anything published is staled by the second reset |
Runbook order |
| Old forced rebuild cleared the tripwire (guard not applied) | Matches Incompatible |
New code serves nothing current for the project | POST …/fold/reset, then backfills |
| Two reviewers claim one Study | Today: both wait on the document lock and the filter decides | The fold-path claim is the same single atomic update | None |
| Clock skew between pods | CreatedAtUtc is off by the skew |
Pod clocks feed only the age threshold and ordering. The heartbeat, fleet membership and claim timestamps use server time ($currentDate, $$NOW) |
Early or late fallback only |
| Mongo primary failover | Retryable-write retry or unknown result on writes; folds abort | w: majority writes. Unknown results resolve from pending, overflow, receipt or quarantine records with the digest rule, regardless of version |
Saved, DigestMismatch, StatisticsOperationUnresolved or reclassify; the fold retries |
| Redelivery or duplicate request | A duplicate id | Saved for a matching digest, fail closed for a different one; the lease admits one folder | None needed |
| Fold races a same-Study write | VersionStale or E11000 for the writer |
A fold-only reload is a free retry; upsert writers reload and retry | Measured in same-Study cells |
| Negative overlay sum | Detected by the reader | Falls back | overlay.invariant_violation, then a rebuild |
| Rolling or complete rollback | Old pods see the closed narrow gate, the tripwire and the allowlist guard | Old paths serve, publish, repair and record history for nothing in fold projects | Runbook order; reset twice around guard removal; a rebuild is always required |
| Rebuild starved by folds | — | Yield request on the fold lease | — |
| Very large pending set at rebuild | The paged read exceeds the snapshot lifetime | PendingNotDrained |
The worker drains; retry |
9. MVP boundary and acceptance¶
Scope¶
The MVP is the fold path for every point-saving writer of ProjectScreening, MembershipScreening, ReviewerScreening, StageAnnotation, MembershipStageAnnotation, ReviewerAnnotation, DomainReconciliation and QuestionAnswers. Owner decision (a). The families differ in what the fold does for them:
| Family | In the MVP |
|---|---|
| ProjectScreening, StageAnnotation, MembershipStageAnnotation, DomainReconciliation | Point-maintained. Moves are folded, and rows are served Fresh with the overlay |
| MembershipScreening, ReviewerScreening | Point-maintained for saves with before-state evidence and at most 100 memberships; otherwise Stale-on-save through an intent |
| ReviewerAnnotation, QuestionAnswers | Stale-on-save only. Every save that affects them carries an invalidation intent, so after a save they are Stale and are served from the authoritative calculation until a rebuild. This is unchanged from today. Point maintenance for them is deferred item 6 |
What point saves touch today (matrix §2):
- Screening saves: move ProjectScreening. They move MembershipScreening and ReviewerScreening with before-state evidence and at most 100 memberships, and otherwise stale them. They stale StageAnnotation, DomainReconciliation, MembershipStageAnnotation and ReviewerAnnotation on an include/exclude crossing, or move them through the annotation half.
- Annotation session save, complete and delete: move StageAnnotation, MembershipStageAnnotation and DomainReconciliation. They stale ReviewerAnnotation (the family or the affected reviewers) and stale QuestionAnswers at project scope.
- Reservations: move StageAnnotation and DomainReconciliation, and stale ReviewerAnnotation.
Moves become transitions, and stales become invalidation intents.
Flag¶
- Name:
materializedProjectStatisticsFold. BackendFeatureFlags.MaterializedProjectStatisticsFold, environment variableSYRF__FeatureFlags__MaterializedProjectStatisticsFold, defaultfalse. - Declarations:
src/charts/syrf-common/env-mapping.yaml, sectionmaterializedProjectStatisticsFlags, withservices: [api, project-management]and noweb:block;RuntimeFeatureFlagCatalogandRuntimeFeatureFlagKeys(catalogue count test 62 to 63);ProjectStatisticsFlagMap.EffectiveFlags(smoke count 17 to 18).- Pinned to the deployed value by
ProjectStatisticsRuntimeActivation, like Writes and the family keys. - Consumer manifest: every typed read gets an entry in
docs/planning/feature-flag-overhaul/consumer-manifest.json, andpnpm -w run validate:generatedmust pass. - Effective only with Writes, the writer's family flags and the allowlist, on a project whose
FoldModeis Enabled.
Activation (owner decision g).
- Staging: after slice 2 the staging pilot may enable ProjectScreening fold mode behind
ProjectStatistics:Fold:AllowPartialCoverage = true. That setting is set only in the staging values
and honoured only when IsStaging() holds and the database is not syrftest.
- Production: waits for the full MVP, every family through slice 5 plus the slice-6 gate (b)
pass, and its own activation decision.
Out of scope: SearchPopulation (no point writer), early client invalidation, change streams, the eligibility Project-token redesign (deferred item 2), point maintenance for the intent-only families (deferred item 6), production activation, and legacy-path removal.
Acceptance (gate b)¶
- Hot path. Red-first real-Mongo command-count tests per save shape assert the budgets in
2, zero
commitTransactioncaused by statistics, and no write to anypmProjectStatistics*collection or toProject.StatisticsAdmissionTokenfrom a statistics path. - Zero statistics-caused failures. The benchmark cells are 1, 2, 5 and 10 reviewers, same-Study and different-Study. A cell passes when:
- fold-arm exhausted submissions equal the source-only arm's in the same cell;
- fold-arm saved submissions equal submissions minus source-level exhaustion;
- statistics-caused conflicts are zero. They are attributed per attempt: a
VersionStalewhose reload showsΔAudit.Version == ΔStatisticsFoldSequenceis fold-caused, but is a free retry and never exhausts; any statistics write rejection counts as statistics-caused. - The same-Study cells will still show source-level conflicts between reviewers saving the same Study. These exist identically in the source-only arm and are reported in their own column.
- The cells run with
ReviewEligibilityPolicyboth off and on, in both arms, so the eligibility Project token is shown to be source-level. - Exact reads. After every benchmark batch and in dedicated tests, for every fold family:
- the overlay before draining equals the authoritative value;
- stored rows after draining equal it;
- receipts equal consumed entries;
- the pending set is empty after the drain;
- parity against the family's legacy calculation section.
- Latency. p95 overhead against source-only is under 10% or at most +2 ms in every cell, with idle-host hygiene. The fold adds one concurrent read round trip and a slightly larger document write to uncontended baselines of 5.8-7.8 ms p95. That is likely 5-30% relative but well under +2 ms absolute, and under 10% is likely in the contended different-Study cells.
- Reads. Fresh reads with 0 and with 32 pending Studies stay within the read benchmark's existing p95 gate.
- Safety.
- Old and other-protocol binaries. Rollback simulation with an older binary's unmodified gate, backfill, repair, copier and drift code shows fail-closed, and the same holds for a binary at another fold protocol.
- History writers. The unmodified old checkpoint build and unchanged-day recorder are shown to be stopped by the allowlist guard. With the worker stopped, saves still append and succeed, and new history writers never record past pending entries.
- Entries. Entries of an unknown protocol or kind are quarantined, never applied or ignored.
- Regression tests. The poison, duplicate-key, livelock, duplicate-id (including unknown result with a different digest), overflow-after-fold, missing-row placeholder, multi-stage crossing and late-save tests are green.
- Revision 4 tests. These also pass:
- the two-step livelock discard is bounded;
- unrecorded-drop ranges answer
StatisticsOperationUnresolved; - a published placeholder survives a sibling republish;
resetstales SearchPopulation;- readers refuse on allowlist disagreement;
- a heartbeat renewal after a protocol rollback is not wedged;
- an older replace cannot advance a heartbeat;
- the claim is equivalent to today's for two concurrent claimers (zero statistics-caused nulls);
- a claim on a Study with no pending set is visible to the overlay;
- the staging check is positive.
10. Implementation plan¶
Each slice is one PR, red-first against MongoDbReplicaSetTestFixture. With the flag off, every slice
leaves main byte-identical in behaviour, and the existing suites prove it. Every behaviour change in
the plan, including the cached-upsert E11000 retry, is gated on the fold flag.
Enable is refused:
- everywhere until slice 2;
- on staging only with the partial-coverage allow setting, from slice 2 to slice 5;
- everywhere else until slice 6 sets PendingEntryCoverage.Complete.
Slices 2 to 5 each bump FoldProtocolVersion, without N-1 support (owner decision h applies from
P0). A staging project enabled at an earlier protocol needs a reset, followed by backfills, after
the next slice deploys.
| # | Slice | Main files | Red-first tests |
|---|---|---|---|
| 0 | Operational prerequisite: the pending index. Admin-triggered build endpoint; production-sized timing on a restored copy; fleet runbook entry; the enable refusal PendingIndexMissing |
ProjectStatisticsAdminController, StudyRepository (index definition only, not at startup), runbook |
Endpoint idempotency; listIndexes gate; startup does not create the index |
| 1 | Framework, dark. Study pending schema with ProtocolVersion, typed transition Kinds and class maps; StatisticsFoldSequence; control FoldMode and FoldProtocolVersion; the protocol-stamped tripwire and ProjectStatisticsFleetVersions.Matches at every site in 5, including the added checks in the checkpoint build and the unchanged-day recorder, and Matches plus the allowlist check on the narrow-gate re-enable; per-pod fleet membership for api and project-management only, storing the allowlist itself, with server-time liveness, deletion on shutdown and the per-project allowlist-consistency refusal; per-protocol, server-time fold-worker heartbeats with no Version bump, read through a projected, uncached find; reset with its preconditions, staling every materialized family; the whole-pending-set write rule; the flag (catalogue, env mapping, manifest, pins, counts) and the staging-only allow setting; PendingEntryMoves.Derive; the fold lease, consumer skeleton, batch halving and livelock bound, and DuplicateKey as retry; poison, quarantine (digest, index, 30-day TTL, a soft cap at 90,000 that disables fold mode, and a hard cap at 100,000 with the FoldQuarantineIncompleteSinceUtc marker) including UnrecordedDrops and QuarantineExists; the two-step livelock discard covering entries and overflow; published Stale placeholders, with MarkFamilyStale when the key is occupied; overlay plumbing (clone replacement); receipt maintenance refusing Studies with a pending set; the flag-gated E11000 reload-and-retry in the four cached upsert writers |
Study.cs, StudyPendingStatistics*.cs, StudyRepository.CreateMappings, ProjectStatisticsControl.cs, ProjectStatisticsGlobalControl.cs, ProjectStatisticsFleetVersions.cs, every site in the table, FeatureFlags.cs, env-mapping.yaml, the runtime catalogue, the flag map, consumer-manifest.json, new Fold/* (Core lifecycle), PM consumer and schedule, NotificationHub and the three idle and liveness consumers |
Old-shape map round-trips PendingStatistics. The tripwire makes the unmodified old gate, backfill, repair, copier and drift code (copied as a test double at the current commit) refuse, and an other-protocol binary refuse. The Matches truth table, including a cleared tripwire and reset. Each poison type, including UnknownProtocol and UnknownKind, is consumed and quarantined while the project keeps folding. A DuplicateKey race goes to retry and then poison. A hot Study cannot livelock the project. Lease fencing and takeover. Saves keep appending with the worker stopped; a heartbeat that is absent or at another protocol only refuses enable and the stamp advance. Flag count tests |
| 2 | ProjectScreening end to end (protocol 1). Fold-path routing for the three screening save shapes: isolated reads with the envelope minted from them, the duplicate checks (pending, overflow, receipt and quarantine, all digest-checked), the deadline with CSOT, unknown-result resolution with the digest rule, and the free fold-only retry. Fold apply, discard and invalidation intents; all other families arrive as intents, so they stay safe and Stale. The reader overlay for ProjectScreening. Rebuild drain (ignoring minimum age), paged uncapped pending(B), subtraction and PendingNotDrained. Enable, disable and staging partial-coverage enable. Copier, barrier and unchanged-day rules. Drift and parity fingerprint |
ReviewController.cs, ProjectScreeningStatisticsWriter.cs, StudyRepository.cs, StudyRepository.ActivityReviewWrites.cs, ProjectStatisticsBundleReader.cs, ProjectStatisticsRebuildService.cs, ProjectScreeningScopeCalculator.cs, ProjectScreeningBackfillService.cs, the checkpoint, copier and recorder classes, the drift and parity services |
Command budget per shape. The M2 race (two digests on one id → fail closed before the fold, after the fold, and through an unknown result). Overflowed id with an unknown result after the fold → resolved from quarantine. First attempt on a newer isolated load runs the receipt check. Unknown result after a fold bump → Saved. Crash between fold steps. Fold versus save in both orders. Resurrection. Rebuild with writes running publishes exact with no PublicationRaceLost. Late write after Disabled. Guarded rollback round trip with reset. Barrier uniqueness with the worker down |
| 3 | Reviewer families (protocol 2): MembershipScreening and ReviewerScreening transitions on screening saves (evidence from the isolated load, membership context digest, the 100-membership intent fallback), with derivers, overlay and rebuild calculators | ReviewerScreeningPointClassifier (deriver), the reviewer scope calculators and backfills, ReviewController |
Exact overlay and drain for membership scopes; a membership added between write and fold → mismatch → fallback and Stale; more than 100 members → intent |
| 4 | Annotation group, writers 1 (protocol 3): StageAnnotation, MembershipStageAnnotation and DomainReconciliation transitions (AnnotationStage, and AnnotationMultiStage on a class change, with the 8 KiB intent fallback); ReviewerAnnotation and QuestionAnswers intents; on annotation session save, complete and delete, and the screening save's annotation half (one entry per namespace) |
ProjectAnnotationStatisticsWriter.cs, AnnotationStatisticsClassifier (deriver), SubmitAnnotationSessionService.cs, ReviewController.RemoveSession, StudyRepository (SaveWithAnnotationStatisticsAsync, TrySaveActivitySessionDeletionAsync), the annotation calculators and backfills |
Per-family overlay parity. An include/exclude crossing on a Study with tallies in two stages gives overlay = authoritative for both stages. A SessionCountTarget change makes earlier entries mismatch; a member add does not invalidate annotation entries. Question-content intents stale QuestionAnswers. Presence transaction unchanged. A two-namespace save resolves duplicates on both |
| 5 | Annotation group, writers 2: reservations (protocol 4). The claim stays today's single atomic FindOneAndUpdate, with one prepended, non-failing pipeline stage that $sets the whole pending set. That stage copies the claimed stage's pre-image tally (or null), session profile fields, reservations and InclusionInfo, plus sessionCountTarget, into an AnnotationReservationClaim entry. The deriver mirrors the seeded-tally pipeline and derives the ReviewerAnnotation intent. The fold-path overload takes the caller's Project. The eligibility admission keeps its transaction and $sets the whole pending set. Releases |
StudyRepository (TryAtomicAssignStudyCoreAsync, TryAdmitActivityReviewAsync, TrySaveReservationChangeAsync), StageReviewService, ReservationChangeSave |
Every mandatory shape in the claim path, the first claim with no tally first: the real images' classification and the authoritative value both equal derive(entry) and the overlay. Two concurrent claimers give exactly today's outcomes, with zero statistics-caused nulls. A claim on a Study with no pending set is visible to the overlay and the index. The cap and overflow are maintained in the pipeline. Release after a fold bump |
| 6 | Acceptance and P0. PendingEntryCoverage.Complete = true, enabling production eligibility. The shipped protocol becomes the production baseline P0. Adds the N-1 machinery for later bumps: the accept window, per-entry protocol derivation dispatch, the stamp-advance consumer and CAS with its fleet and pending-protocol preconditions, breaking-kind discards and the N-1 test harness. The benchmark gains an arm enum {SourceOnly, Transactional, Fold}; the fold arm drains through the worker outside the timed window, reports source-caused and fold-caused conflicts separately, runs ReviewEligibilityPolicy off and on, includes the reviewer-correction pattern, and asserts acceptance 3 per batch. The read benchmark gets cases with 0 and 32 pending. Idle-host rerun: 100 iterations, plus a counting run |
ProjectScreeningWriteBenchmark.cs, ProjectScreeningWriteRoundTripBudgetTests.cs, the read benchmark, screening-write-benchmark.md |
Gate (b) evaluated per cell |
| 7 | Docs and rules. Statistics reference as built, runbook (evidence timing, overflow and quarantine counters, rollback order, index operation), .claude/rules/materialized-stats.md, STATUS |
docs | docs/scripts/validate-docs.sh |
Dependencies:
- slice 0 is independent and can start at once;
- slice 2 needs slice 1;
- slices 3, 4 and 5 need slice 2 and can run in parallel;
- slice 6 needs 0 and 2 to 5;
- slice 7 follows slice 6.
The MVP is done after slice 6 passes gate (b).
11. Deferred items¶
- Inclusion recalculation lost update.
UpdateStudyInclusionInfoForProjectAsync'sUpdateManydoes not bumpAudit.Version, so a concurrent save can overwrite itsInclusionInfo. This predates the fold and is fenced. Its fix is a separate source-correctness change. - Eligibility Project token.
TrySaveActivityReviewAsyncwritesProject.StatisticsAdmissionTokenon every save to serialize eligibility against settings. This is source-level and present with statistics off. Moving it, for example to a per-stage settings revision checked in the Study filter, is an eligibility design change, outside FEAT-024 statistics. - Coalescing the two control reads into one
$unionWithaggregate, or movingFoldModeonto the Project. This is an optimisation, to be taken only if slice 6 misses +2 ms. - Early client invalidation, only if pilot fold lag exceeds 2 s p95.
- A pending-aware drift check that runs without a matching fingerprint. Not needed while inconclusive results retry weekly.
- Point maintenance for intent-only families. ReviewerAnnotation and QuestionAnswers are staled
by intents, as today, so the saves-during-a-backfill gap stays open for them. Making them
point-maintained needs catalogue work on
SufficientlyAllocatedand on question content. That is independent of the fold protocol. - N-1 support between MVP slices. Deliberately not built (owner decision h, recommendation adopted). The staging pilot resets after each slice instead.
- Detecting pre-fold replicas automatically. Pre-fold pods write no fleet membership, so enable's "no lower member" check cannot see them. The runbook's rollout-complete gate covers them. Adding the ReplicaSet hash to membership via the downward API would allow a code check, but needs a chart change; it is deferred until a rollout incident shows the need.
- The receipt-maintenance probe (F4). The probe for a pending set is a snapshot read, so an append that lands after it can still be retired and become poison. That outcome is fail-closed: the entry is quarantined and its families are staled. Closing the gap needs a write on the Study, which is not worth it.
12. Open questions for the owner¶
None at revision 5. The revision 3 question (the cost of protocol upgrades) is answered by owner decision (h).
Revision history¶
- Revision 1 (30 September 2026). ProjectScreening-only MVP using a row-level
SourceVersion = 2marker. - Revision 2 (1 October 2026). Records owner decisions a-e, including the per-stage reviewer-tracking follow-up (#3876). It also:
- replaces the row marker with the control storage-version tripwire, with a per-site comparison rule (review M1);
- keeps envelope ids, with pending, receipt and overflow duplicate checks and unknown-result resolution that ignores the version (M2);
- adds the rollback tripwire, fleet declaration and runbook (M3);
- adds the poison-entry quarantine path (M4);
- widens the scope to every screening- and annotation-dependent family, with entries that store transitions and invalidation intents;
- closes the review's should-fix and follow-up items: Project admission, command budget, late writes, ledger and maintenance, history identity, the uncapped rebuild read, guard-less families, the writer inventory, cached upserts, isolated reads, lease fields, the index operation and write concern.
- Revision 3 (1 October 2026). Records owner decisions f (allowlist GitOps guard plus runbook) and g (staging-only enable after slice 2). It also closes the second review:
- the fold protocol version in the tripwire, heartbeats and entries, with typed transition kinds (N1);
- full multi-stage annotation profiles on a class change (N2);
- Stale placeholder rows instead of dropped moves (N3);
- the claim path keeps its snapshot transaction (N4);
- the corrected tripwire table, the fold-worker heartbeat and the allowlist guard for history writers (N5);
- the digest rule for unknown results (N6);
- receipt and quarantine checks whenever an attempt does not compare-and-set on the envelope's revision;
- quarantine of overflowed ids with their digests;
resetas recovery from a cleared tripwire;- the livelock bound, duplicate key treated as retry, receipt maintenance refusing Studies with a pending set, the flag-gated upsert retry, per-namespace entries, membership dropped from the annotation digest, and the intent-only family list.
- Revision 4 (1 October 2026). Records owner decision h: N-1 support from production
P0, not between MVP slices. It also closes the third review: resetstales every materialized family, SearchPopulation included (NM1);- a rollback runbook bracketed by the narrow gate, with the guard rolled out alone, two
resets around guard removal, and allowlist-consistency refusals from per-pod fleet membership (NM2); - the claim is today's single atomic update with a prepended pending-set stage, and every writer maintains the whole pending set (NM3);
UnrecordedDropsranges (S8), and one quarantine retention rule with a 30-day TTL and a capacity stop (S7);- a two-step livelock discard that does not need a transaction on the Study (S12);
- per-protocol, server-time heartbeats with no
Versionbump (S1-S3); - published placeholders (S4), the start-from-zero refusal scoped to active fold mode (S5), and
resetpreconditions (S6); - per-entry protocol (S9), narrow-gate coverage and an allowlist-only fold worker (S10), membership
checks on enable (S11) and the positive
IsStaging()(S13); - follow-ups F1-F4.
- Revision 5 (1 October 2026). Closes the fourth review.
- Claim derivation (R4-M1). The claim derives against the real before and after profiles,
including the tally that
BuildAssignmentPipelineseeds. The payload carries the stage's session fields andsessionCountTarget. The ReviewerAnnotation intent is derived identically by the fold and the reader. The claim test shapes are mandatory, the Project is in the claim budget, and MongoDB 4.4 is the stated minimum. - Should-fixes. The prepended stage cannot fail (S-R4-1). Fleet membership stores the allowlist itself, uses server-time liveness, comes from api and project-management only, and refuses per project (S-R4-2). The two-step discard covers overflow (S-R4-3). The quarantine has a soft cap and a fail-closed hard-cap marker (S-R4-4). The livelock threshold is 20 throughout (S-R4-5). There are enable and stamp-advance rollout gates (S-R4-6), a stamp advance specified as a CAS with an owner (S-R4-7), and a projected, uncached heartbeat read (S-R4-8).
- Follow-ups. The limits of N-1 for breaking changes (F-R4-1), the roll-forward pre-warm (F-R4-2) and the claim overflow range (F-R4-3).