ADR-020: All-or-nothing Bulk Study Update with study locks¶
Status¶
Accepted by Chris on 2026-10-02, as written, with every recommendation in
decisions D1–D9. (The frontmatter says Approved because that is the docs
validator's value for an accepted decision.) Chris chose the direction the same day: Option C
(validate and compute every change before any write, record each study's before-image, and restore
automatically on failure, denial or cancellation), plus his own requirement that the affected
studies are locked for the duration of the job, so that nothing else can change them and the
rollback is exact. A locked study stays readable; writes are refused at save time.
This ADR records the design. Implementation follows in separate PRs, in the slices below.
Production behaviour does not change until a new default-off flag is turned on (rollout). The guarded restoration that is live in production and the 117 historical jobs it holds are untouched by this design.
Context¶
Code facts below were read on main at ace997841 (2026-10-02).
What a bulk study update does today¶
An administrator uploads a CSV/TSV/XLSX file keyed by study ID. Each row can set customId,
pdfRelativePath and, when screening import settings are present, add screening decisions for one
stage (feature notes). The flow is API signature
(StudyController.getSignatureForBulkStudyUpdate, policy BulkUpdateStudies) → S3 → notifier →
SubmitJob<IStartBulkStudyUpdateJobCommand> → StartBulkStudyUpdateJobConsumer (project-management,
5-minute job timeout, 4 incremental retries, 5 concurrent jobs per instance) →
ProjectManagementService.ParseBulkStudyUpdateJobFile → StudyReferenceFileParser.ParseBulkStudyUpdateAsync.
The parser validates the whole file once, then reads it again and writes in batches of 400 rows:
- Simple rows go to
StudyRepository.ApplySimpleUpdates: one unorderedBulkWriteof$setoperations with no transaction, noAudit.Versionpredicate and no version bump. - Screening rows, with
ReviewEligibilityPolicyon, go toTrySaveAdministrativeScreeningBatchAsyncin chunks of 50. Each chunk is one Mongo transaction that re-reads the project and studies, refuses rows (not an active member, existing decision, stage not found), replaces accepted studies with anAudit.Versionpredicate, and sets the stage's monotonicHasEverBeenReviewedmarker. With the policy off, the legacy path loads studies, callsStudy.AddScreeningandSaveManyAsyncs them. - With materialized statistics on, the whole job runs inside the FEAT-024 bulk screening fence
(
AdmitBulkScreeningOperationAsync/ReleaseBulkScreeningOperationSafelyAsync, released withsourceChanged: truewhether the job succeeds or fails).
A failure, timeout or cancellation after the first batch therefore leaves a partial commit: some
batches written, others not, and refused screening rows reported but not undone. The
parser-stall runbook is explicit that the current
cancellation is cooperative, not a rollback, and that partial commits need reconciliation. The
incident in September 2026 left 28 jobs Parsing whose committed writes are still unknown.
Work this design builds on¶
- ADR-017 recovery design (PR #2900, open,
In-Review). It fences parser batches with an attempt-scoped Mongo guard document: each batch is one
transaction that conditionally writes the guard at an expected generation together with every study
mutation and the progress update, so a recovery fence acquired mid-job aborts the in-flight batch.
It keeps legacy attempts quarantined and makes retry depend on an immutable S3 input manifest
(
VersionIdplus checksum) and a per-unit write ledger. This ADR adopts the guard, the manifest and the ledger rather than inventing parallel ones. (PR #2900's file is numbered ADR-017, whichmainhas since used for the bulk-PDF notifier decision; it will need renumbering when it merges. This ADR refers to it as "the recovery ADR".) - Restoration guard (live in production since 2026-09-09).
BulkStudyUpdateJob.ExecutionVersionis stamped1by the API only; withBulkStudyUpdateRequireCurrentExecutionVersionon, the consumer rejects any other version before touching PM state. The 117 historical jobs have no version and stay held. - Application authority M5a (queued-work ledger).
DelegatedWorkGaterechecks the actor's authority (BulkUpdateStudies) at every execution and retry, behindDelegatedWorkAdmissionEnforced(default off). The P9 proposal says a bulk update must stop when the actor loses authority, and asks how legacy jobs are classified. - Reviewer sessions and claims (active reviewer tracking).
A reviewer's presence is an
ActiveReviewSession/ activity claim inStudy.SlotReservations. It exists while the study is open and is removed by an intentional leave, a 5-minute idle mark plus the stage'sIdleSessionTimeoutMinutes(default 120), or a 2-hour suspended-session grace period. Timeouts use MassTransit scheduled messages with a cancel/reschedule token. This design reuses that scheduling pattern for lock leases. - Optimistic concurrency. Almost every whole-study save filters on
Audit.Version(MongoExtensions.GetFilter,TrySaveExistingAsync, the capacity and activity-review writes). This gives the lock a cheap enforcement seam: bumping the version when the lock is taken makes every copy loaded before the lock fail its save.
Decision¶
A Bulk Study Update created under the new flag runs as a durable operation with five phases. Nothing is written to a study until every row has been validated and every change computed. Before the first change, every target study is locked; while locked, every other write path refuses to change it. Each change records the study's before-image in the same transaction. The operation then either commits all of it, or restores every before-image, and only then releases the locks.
Plan (no writes) ─► Lock ─► Apply (batch txns + before-images) ─► Commit ─► Release ─► Completed
│ │ │ ▲
│ │ └──── failure / denial / cancel / ceiling ───► Roll back ─► Release ─► Failed / Cancelled / Denied
│ └── busy study, conflict, or failure ─► Release ─► Refused
└── any invalid row ─► Refused (nothing written, no locks)
Two invariants carry the guarantee:
- Exclusive locks. From the moment a study is locked until its lock is released, the only process that may change its content is the operation named in the lock, through the one attempt that holds the current generation on the operation record. The study-side lock names only the operation; the generation lives on the operation record, and every transaction that writes a locked study also writes the guard at the expected generation, which fences stale attempts.
- Commit is one write. The operation reaches
Committedwith one conditional write on its operation record. Before that write, every interruption leads to rollback. After it, every interruption leads forward to release. There is no third outcome except an alerted quarantine when an invariant has been broken from outside.
Durable records¶
| Record | Where | Purpose |
|---|---|---|
| Operation record (the guard) | new collection pmBulkStudyUpdateOperation, one document per job |
Phase, lease, generation, plan digest, counters, outcome. This is the recovery ADR's attempt-scoped guard document: every batch transaction writes it conditionally on the expected generation. |
| Plan chunks | new collection pmBulkStudyUpdatePlan, keyed (operationId, chunk) |
The validated, computed change for each study: target study ID, Audit.Version at plan time, the fields to set or unset, and a hash of the intended post-image. Chunked because a plan can exceed 16 MB. Written once, immutable. |
| Before-images | new collection pmBulkStudyUpdateBeforeImage, keyed (operationId, studyId) |
The touched fields of the study as they were under the lock, the content hash of the whole study before and after the change, and its restore state (Applied, Restored). |
| Study lock | new field Study.BulkUpdateLock ({ operationId, lockedAtUtc }) |
Makes the lock visible on the study itself to every reader and writer. Absent on all existing documents, so it costs nothing until used. It deliberately carries no generation: a takeover changes the generation on the operation record only, so it never has to re-stamp locks. |
| Job status | the existing embedded BulkStudyUpdateJob |
Keeps the UI's status and progress. New statuses: Locking, Applying, RollingBack, RolledBack, RefusedBusy. |
The operation record carries the M5a admission key (the file ID), the ExecutionVersion (2) and the
S3 input manifest (VersionId, length, checksum) from the recovery ADR, so the plan is bound to
exact bytes and a retry re-reads the same bytes.
Phase 1: Plan (no study writes)¶
- Check M5a admission and authority. Check the execution version.
- Read the exact S3 object version named by the manifest; verify its length and checksum before parsing. A mismatch refuses the job.
- Parse every row. Resolve every study ID in the project. Apply today's row rules, including the
administrative screening refusals (
NotActiveMember,ExistingDecision,StageNotFound) and the blank-cell rules in the feature notes. Any error or refusal in any row refuses the whole job (see decision D1). The user sees the first rows with problems, as today, and nothing has been written. - Compute each study's change in memory with the same domain methods the current path uses
(
AdministrativeScreeningPolicy.ApplyorStudy.AddScreening, and the identifier setters), and reduce it to a field-level change set. Rows naming the same study fold into one change in file order. A row that changes nothing is dropped from the plan. - Refuse files above the size ceiling (size).
- Write the plan chunks and move the operation to
Plannedwith the plan digest.
Failure: the job ends Refused (validation) or Failed (infrastructure) with no study written and
no lock taken. Crash: a new attempt finds Planning or Planned, discards incomplete plan chunks,
and re-plans from the same manifest bytes. The plan is deterministic, so a re-plan yields the same
digest.
Phase 2: Lock¶
Locks are taken in batches of up to 100 studies, each batch in one transaction that:
- writes the guard at the expected generation;
- for each study, matches
_id,projectId, the plannedAudit.Version, noBulkUpdateLock, and no entry inSlotReservations(no reviewer session or claim of any kind); - sets
BulkUpdateLockand bumpsAudit.Version(so any copy loaded before the lock now fails its version check and reloads into a refusal); - advances the lock counter on the operation record.
A study that fails the match is classified in the same transaction: busy (a reservation exists), changed (its version moved since planning), locked by another operation, or missing.
- Busy: the whole job is refused before any change, with the list of busy studies (busy studies).
- Changed: the study is re-validated under its new lock; if its row is still valid and its change set is recomputed identically or validly, the plan chunk is superseded and locking continues; otherwise the job is refused (decision D3). Re-validation happens under the lock, so it cannot race.
- Locked by another operation: refused, naming the other job.
A refusal releases every lock already taken (Phase 5) before the job reports. Crash: the new
attempt bumps the generation (fencing the old attempt), counts the locks it holds via
BulkUpdateLock.operationId, and either continues locking or releases, according to the phase the
operation record shows. Lock acquisition is idempotent because it matches on the operation ID.
The project is also checked: locking is refused while the project is CalculatingInclusionInfo or
another project-wide operation in the write-path table
holds its own fence, and those operations are refused while this one holds locks.
Phase 3: Apply¶
Each batch of up to 100 studies is one Mongo transaction that:
- writes the guard at the expected generation (the recovery ADR's shared guard; an external stop, such as a recovery fence or an operator cancellation, or a takeover bumps the generation, so this write conflicts and the whole batch aborts);
- re-reads each study and requires
BulkUpdateLock.operationIdto match (the guard write in step 1 has already proved this attempt holds the current generation); - inserts the before-image (touched fields and the content hash before the change) unless one
already exists for
(operationId, studyId); - applies the planned field changes and bumps
Audit.Version; - records the post-change content hash on the before-image and checks it against the plan's intended hash;
- advances the applied counter and progress on the operation record.
The first-review marker on the stage (Stage.HasEverBeenReviewed) is not set here; it is
deferred to commit, because it is monotonic and a rollback could not unset it.
Between batches the executor renews its lease, checks the cancellation token, and rechecks authority
(decision D5). Failure of a batch (conflict that survives the bounded retries, a hash mismatch,
denial, cancellation, lease ceiling, or any exception) moves the operation to RollingBack. Crash:
the takeover attempt bumps the generation and rolls back (decision D2). Apply is also idempotent if
D2 chooses roll-forward: a study whose before-image exists and whose current content hash equals its
recorded post-change hash is already applied.
Phase 4a: Commit¶
When every planned study has a before-image in Applied state, one conditional write moves the
operation from Applying to Committed (matching the generation and the applied count). That write
is the point of no return. After it, the executor:
- sets the deferred stage marker (
RecordSavedReview) in one project transaction, if any screening decision was applied; - releases the FEAT-024 bulk screening fence with
sourceChanged: true(unchanged contract); - marks the embedded job
CompletedParsingwith the result counts; - continues to Phase 5.
Each of these is idempotent and is replayed by a takeover attempt that finds Committed.
Phase 4b: Roll back¶
Rollback walks the before-images (reverse batch order, 100 per transaction). Each transaction:
- writes the guard at the expected generation;
- re-reads the study and requires the lock to match and its content hash to equal the recorded post-change hash (the lock guarantees this);
- restores the touched fields to their before-image values (
$setthe old values,$unsetfields that were absent) and bumpsAudit.Versionforward; - checks the restored content hash equals the before-change hash;
- marks the before-image
Restored.
A study whose content hash does not match means something wrote it despite the lock. It is not
overwritten. The operation moves to RollbackQuarantined, keeps every lock, and raises an alert
naming the operation (not the study data). Release then requires an engineer, not elapsed time,
consistent with the recovery ADR's fail-closed rule.
When every before-image is Restored, the operation moves to RolledBack, the FEAT-024 fence is
released with sourceChanged: true (the population changed twice), and the embedded job is marked
with its outcome: Failed, Cancelled or Denied, with a message saying nothing was changed.
Rollback is idempotent: restored before-images are skipped, and a takeover after a crash finds
RollingBack and continues.
Phase 5: Release¶
Locks are removed in batches, each in one transaction that writes the guard at the expected
generation and unsets BulkUpdateLock where operationId matches. Release bumps Audit.Version.
Without the bump, a copy of the study loaded while it was locked (reads are never refused) would
carry the lock field and a version that still matches after release; its whole-document replace
would pass both predicates and write the lock back with no owner. With the bump, that save fails
its version check, reloads the unlocked study and proceeds normally. Release is idempotent and runs
on every exit path (Refused, Committed, RolledBack). Only after the lock count reaches zero does
the operation become terminal (Completed, Failed, Cancelled, Denied, Refused).
Crash and takeover at every step¶
Every executing attempt holds a lease on the operation record (leaseOwner, leaseExpiresAtUtc,
generation). A new MassTransit attempt, or the lease watchdog, takes over by a conditional write
that requires an expired lease and increments the generation. The previous attempt's next guard write
then fails, so it cannot write again. This is the same fencing the recovery ADR requires. Because
study locks name only the operation, the takeover finds and drives the same locks without touching
them first.
| Crash point | Durable state found by the takeover | Takeover action |
|---|---|---|
| During planning | Planning, partial plan chunks |
Delete the chunks, re-plan from the same manifest |
| After planning, before the first lock | Planned, no locks |
Lock |
| During locking | Locking, some locks |
Finish locking (or release on refusal) |
| During apply, mid-transaction | Batch aborted by Mongo: nothing of that batch is visible | Roll back (or roll forward under D2) |
| During apply, between batches | Applying, n before-images Applied |
Roll back (or roll forward under D2) |
| After the commit write, before deferred effects | Committed |
Replay deferred effects, release |
| During rollback | RollingBack, some before-images Restored |
Continue rollback |
| During release | Terminal-pending, some locks left | Finish release |
| Rollback hash mismatch | RollbackQuarantined |
Nothing automatic; alert, locks held |
The lease watchdog reuses the review-session pattern: when the operation starts, it schedules a
CheckBulkStudyUpdateLease message through the existing MassTransit scheduler; each lease renewal
cancels and reschedules it. If it fires with the lease expired and no live attempt has taken over,
it takes over itself and runs the recovery action above. A takeover always runs as the system
obligation in authority, never as the original actor.
Lock semantics¶
Scope¶
Only studies named in the file and changed by the plan are locked. A row that changes nothing is not locked. The rest of the project is unaffected. Project-wide operations are the one exception because they touch every study (table).
Enforcement¶
The lock is a field on the study, so enforcement is a predicate on the study write:
- Whole-document saves that already filter on
Audit.Versionget the lock for free on any copy loaded before the lock. A copy loaded after the lock carries the field, so the shared save seam addsBulkUpdateLock: { $exists: false }to the filter forStudy(a per-aggregate write filter inMongoUnitOfWorkBase/MongoExtensions), except for the lock holder's own writes. - Direct updates (
UpdateOne,UpdateMany,BulkWrite, aggregation-pipeline updates) add the same predicate explicitly. - A write that fails only because of the lock is reported as
StudyLockedByBulkUpdate, not as a generic concurrency conflict, so callers do not retry it in a loop. - An architecture test fails the build if any write to the study collection omits the predicate. The table below is the starting allow-list; a new write path must be added to it.
Write paths that must honour the lock¶
From main at ace997841. "Refuse" means the request fails with the lock message and writes nothing.
| Area | Path (entry point → repository method) | Behaviour while locked |
|---|---|---|
| Screening decisions | ReviewController → SaveScreeningWithCapacityGuardAsync, TrySaveScreeningWithStatisticsAsync, TrySaveActivityReviewAsync |
Refuse |
| Screening, fold path | ReviewScreeningFoldTarget → SaveScreeningOnFoldPathAsync, TrySaveActivityReviewAsync |
Refuse |
| Annotation saves and submit | ReviewController (SaveAsync(study), SaveWithAnnotationStatisticsAsync), SubmitAnnotationSessionService (SaveWithCapacityGuardAsync, SaveWithAnnotationStatisticsAsync, TrySaveActivityReviewAsync, SaveAsync) |
Refuse |
| Session deletion | ReviewController → TrySaveActivitySessionDeletionAsync |
Refuse |
| Study assignment and admission (screening, annotation, allocated, reconciliation) | NotificationHub and StageReviewService → TryAtomicAssignStudyAsync, TryAtomicAssignScreeningStudyAsync, TryAtomicAssignAllocatedStudyAsync, TryAdmitActivityReviewAsync; NextRandomStudyAvailableAndUnstartedForReconciliationAsync |
Skip the locked study: pool sampling excludes locked studies so the reviewer gets another one; a direct open of a locked study is refused |
| Session lifecycle | NotificationHub (SaveAsync(study), SetSlotReservationIdleScheduleTokenAsync), ReservationChangeSave → TrySaveReservationChangeAsync |
Refuse (unreachable: a locked study has no reservation) |
| Session maintenance consumers | MarkSessionIdleConsumer, RemoveIdleSessionConsumer, RemoveSuspendedSessionConsumer, CheckConnectionLivenessConsumer |
Skip (unreachable for the same reason; asserted by a test) |
| Stage settings claim release | StageReviewSettingsController (direct study-collection writes) |
Refuse the settings change |
| Workload share backfill | StageWorkloadSharesController → BackfillWorkloadShareBucketsAsync |
Refuse |
| Risk of bias, single | StudyController.addRob / addRobs → SaveAsync(study) |
Refuse |
| Risk of bias, batch job | RobProcessingService → SaveManyAsync |
Fail the RoB batch with the lock reason (its existing failure path) |
| PDF attach | BulkPdfUploadFinalizeConsumer → MarkBulkPdfDeliveredAsync |
Leave the delivery pending and retry after release (its finalize is already idempotent) |
| PDF corrections | StudyController correction submit/approve (StudyPdfCorrection, a separate collection) |
Refuse approval (it changes the effective PDF a bulk update may set); submission may proceed |
| Other bulk study updates | Lock acquisition; the legacy v1 path (ApplySimpleUpdates, TrySaveAdministrativeScreeningBatchAsync, SaveManyAsync) |
New operation refused at lock time; a v1 job refuses the locked rows and fails |
| Reference import | ProjectManagementService.ParseReferenceFile → BatchedSaveManyAsync |
Not affected in practice (creates new study IDs); the predicate still applies |
| Search deletion | ProjectManagementService → DeleteStudiesWithSearchAsync |
Refuse the deletion |
| Question deletion | ProjectManagementService → RemoveAnnotationsFromStudiesInProjectForQuestionAsync (project-wide UpdateMany) |
Refuse while any lock exists in the project |
| Screening inclusion recalculation | UpdateStudyScreeningStatsConsumer / ProjectManagementService → UpdateStudyInclusionInfoForProjectAsync (project-wide) |
The screening-settings save that triggers it is refused while any lock exists; lock acquisition is refused while the project is CalculatingInclusionInfo |
| Statistics fold bookkeeping | ProjectStatisticsFoldStudyStore.RemovePendingSetAsync (unsets PendingStatistics, increments Audit.Version and StatisticsFoldSequence) and DiscardAsync (pipeline update of PendingStatistics, no version filter) |
Allowed, and listed in the architecture test's allow-list with this reason: both touch only fold metadata, which is excluded from the content hash and never restored. Apply, rollback and release match locks by operation, not by version, so a fold drain on a locked study does not disturb them. A drain between planning and locking does move Audit.Version, so the study is classified changed; re-validation (D3) recomputes the same change set and continues |
| Preview seeding | DatabaseSeeder |
Exempt (preview databases only) |
The UploadSignature endpoint for a single study's PDF writes only to S3 and is not a study write.
The refusal message¶
For a reviewer or editor who hits a locked study (HTTP 409, problem type study-locked-by-bulk-update,
with a Retry-After from the lease):
This study is being changed by a bulk study update that started at 14:05. It will be available again when the update finishes, usually within a few minutes. Nothing you entered has been saved.
The second sentence matters for annotation forms: the form keeps the reviewer's unsaved input, and the existing save-retry path can resubmit once the lock is gone.
For a project-wide operation (settings save, question deletion, search deletion):
A bulk study update is in progress in this project. Try again when it finishes.
The study list shows a "Bulk update in progress" badge on locked studies, from BulkUpdateLock.
Lease duration and ceiling¶
- Lease: 2 minutes, renewed every 30 seconds and between batches. A lease expiry triggers takeover, not unlocking.
- Ceiling: 30 minutes from the first lock, configurable per environment, with a hard maximum of 2 hours. At the ceiling the executor stops applying and rolls back. The ceiling bounds how long reviewers can be kept off a study; the size limit keeps normal jobs far below it.
- Expired lease does not free studies. A study that holds an applied change must not become
writable until it is either committed or restored; otherwise the rollback is no longer exact.
Studies are released only by Phase 5, or by an engineer after a
RollbackQuarantinedalert. - Job timeout: today's 5-minute
JobTimeoutwould kill most large jobs mid-apply. Version 2 jobs need their own job type or options with a timeout at the ceiling (decision D6).
A review is in progress on a target study¶
Decision (D4): refuse to start. If any target study has a reservation in SlotReservations
(an active, idle or suspended session, or an activity claim), the job is refused during Phase 2
before any change, and every lock taken is released. The administrator sees:
12 studies in this file are open in a review session, so the update was not started and nothing was changed. Ask the reviewers to finish or leave these studies, or remove them from the file, then upload it again.
Study ID Custom ID Stage Session state … … Screening Active
The list shows at most 50 studies and a count of the rest. It names the stage and session state, not the reviewer (the administrator can already see who is reviewing in the project's activity view, under its own authorization).
The alternative, waiting, would lock the free studies, then wait for each busy session to end. It was rejected as the default because:
- a session can legitimately last up to 2 hours idle plus a 2-hour suspended grace period, which is far beyond the lock ceiling, so a wait would usually end in a timeout and rollback anyway;
- while waiting, the job holds locks on the free studies, keeping other reviewers off them for no benefit;
- a waiting job must be woken by session removal events, which adds a new coupling between session maintenance and the bulk executor.
Waiting could be offered later as an explicit option ("start when these studies are free, within N minutes") without changing the lock model (decision D4).
Saved but unsubmitted annotation drafts are not sessions and do not block the start. They cannot be saved again while the study is locked; the reviewer gets the lock message.
Rollback guarantees and limits¶
What rollback guarantees¶
For every locked study, the content after rollback is byte-identical to its content when the lock was taken, checked by hash. Content means the whole study document except the metadata listed below. Because no other writer can change a locked study, the before-image cannot be stale, and the restore cannot overwrite anyone's work.
What rollback does not restore (by design)¶
| Item | Why it is not restored |
|---|---|
Audit.Version, Audit.LastModified, Audit.LastModifiedBy |
Versions must only move forward. Restoring an old version would let a save from a copy loaded before the lock succeed against the restored document. Rollback writes a new, higher version. |
StatisticsFoldSequence, PendingStatistics |
Owned by the asynchronous statistics fold, which must see the apply and the restore as two changes. |
The stage's HasEverBeenReviewed marker |
Not written until commit, so there is nothing to undo. |
| Change-stream events | Clients watching the study list see the change and then the restore. No reviewer is on a locked study, so no review form is affected. |
| Logs and metrics | The attempt is recorded; nothing is retracted. |
Follow-on effects¶
- Statistics (FEAT-024). The bulk screening fence is raised before Phase 3 (as today) and stays
raised until commit or rollback finishes. It is released with
sourceChanged: trueon both paths, so derived statistics are rebuilt either way. A rollback never leaves aFreshstatistics row describing the rolled-back population. - Screening inclusion recalculation. The bulk update does not trigger it today, and v2 does not either. Project-wide recalculation and locking exclude each other (see the table).
- Stage marker. Deferred to commit (above).
- Notifications and events.
Studyraises no domain events and the bulk path publishes no study messages today; v2 publishes none during apply either. Progress goes to the project stream, as today, and the final message states the outcome. No email is sent. - Other jobs. A RoB batch or PDF finalize that meets a locked study fails or defers through its own existing path; neither has partial state that depends on the bulk outcome.
Before-image retention and audit¶
Before-images contain study identifiers, PDF paths and screening decisions (who decided what). They
are project data, not participant data. They are kept for 30 days after the operation becomes
terminal, then deleted by a TTL index; the operation record keeps a permanent summary (counts,
outcome, reason code, plan digest, hashes) without field values. A RollbackQuarantined operation
keeps its before-images until an engineer resolves it. Logs carry the operation ID and counts only,
never field values. Retention length is decision D7.
Authority (M5 and P9)¶
- Denial mid-job means rollback. With
DelegatedWorkAdmissionEnforcedon, the executor rechecks the actor's authority at each attempt (M5a, as today) and also between apply batches (decision D5, mirroring the batch-RoB proposal in P9). The executor checks between batches, so a denial leaves no batch in flight: it moves the operation toRollingBackand rolls back. Restoring access later does not resume it. - Rollback must run without the actor. Once authority is lost, the rollback cannot be authorised
by that actor. It runs as a system obligation, proposed contract
bulk-study-update-rollback.v1: it may only restore before-images and release locks for an operation that was admitted, and never applies anything. This is a new P9 item and needs Chris's approval like the screening and bulk-PDF obligations (decision D8). Without it, a denied job would have to leave its studies locked, which is worse than either outcome. - "Stop" means a clean rollback. P9 item 3 asks to confirm that a bulk update stops when the actor loses authority. Under this design, stopping never leaves partial state: a stopped v2 job is either fully rolled back or, if it had already committed, fully applied.
- Legacy jobs. P9 item 4 (classify legacy jobs from
StartedByInvestigatorIdor quarantine them) is unaffected. v1 and legacy jobs never take locks.
Relation to the recovery ADR¶
- v2 jobs use its guard document (the operation record), its S3 manifest and its idea of a write ledger (the plan plus before-images is the ledger: every intended and committed change is recorded atomically with the change).
- For a v2 job, the recovery coordinator does not need a cancel-only quarantine: a recovery fence bumps the generation, the in-flight batch aborts, and the operation rolls back automatically.
- Legacy and v1 attempts remain under the recovery ADR's quarantine and gates. This design does not make them recoverable and does not change the production guard.
Size and performance¶
- File size ceiling: 10,000 changed studies per job by default (configurable, decision D9). Historical production jobs were between about 100 and 1,500 rows. Larger files are refused at plan time with a message asking the user to split the file.
- Transaction size: 100 studies per lock, apply, rollback or release transaction. The before-image holds only touched fields (identifiers, the stage's screening subtree), so a batch stays well under Mongo's 16 MB document limit and the default 60-second transaction lifetime on Atlas. Batches are sequential, as in the parser-stall fix; write conflicts retry the batch a bounded number of times, then fail into rollback.
- Time: for 10,000 studies, roughly 100 lock transactions, 100 apply transactions and 100 release transactions. At the current local rate (about 120 rows in under 30 seconds end to end, including S3 and the job service), apply is expected in single-digit minutes; this must be measured on the replica-set fixture before the ceiling is fixed.
- Plan memory: planning holds one row's parsed values plus the compact change set per study, streamed to plan chunks in groups of 1,000, so memory is bounded by the chunk size, not the file.
- Lock contention: locks last only for the apply window. Pool sampling skips locked studies, so reviewers keep working on the rest of the project. The main contention is a study opened by a reviewer between planning and locking: that study fails the lock predicate and the job is refused as busy; the user retries once the session ends. Two v2 jobs on overlapping studies cannot run together; the second is refused at lock time, naming the first.
- Index: a sparse index on
BulkUpdateLock.operationIdsupports release, takeover counting and the badge query; its cost is nil for unlocked studies.
Rollout¶
Flag¶
New flag bulkStudyUpdateAtomicApply, default off everywhere, defined in
src/charts/syrf-common/env-mapping.yaml for the API and project-management services, not
runtime-overridable (it changes job semantics mid-flight). Flag decision: flagged, because it
changes observable semantics on a production data path, spans two services, depends on the lock
honouring being deployed everywhere first, and needs a kill switch for new jobs.
Compatibility and order¶
- Lock honouring, dark. Add the
BulkUpdateLockfield, the write predicate on every path in the table, the 409 handling and the architecture test, in API and project-management. With no locks in the database, every predicate matches and behaviour is unchanged. Deploy to production first. - v2 executor, dark. Add the operation, plan and before-image collections, the executor, the
lease watchdog and the new statuses in project-management. Admission accepts execution versions
{1, 2};BulkStudyUpdateRequireCurrentExecutionVersioncontinues to reject version 0 (the historical jobs) exactly as now. v1 jobs keep the current path. - API stamps version 2 only when the flag is on. Turning the flag off stops new v2 jobs; v2 jobs already admitted always run to commit or rollback, because project-management keeps the executor regardless of the flag.
- Enable in previews, then staging, then production with explicit approval.
Never roll project-management back to a version without the executor while any v2 operation is non-terminal, and never roll API or project-management back below step 1 while any lock exists. In-flight legacy and v1 jobs are not migrated and never take locks; nothing here touches the 117 held historical jobs.
Test strategy¶
All core tests run against a real Mongo replica set (transactions and write conflicts do not exist on a standalone or in-memory store) and the job path against real RabbitMQ and the Quartz-backed scheduler used in E2E, not in-memory transports.
- Plan: invalid rows, refused screening rows, duplicate rows for one study, unchanged rows, manifest mismatch, oversize file; each proves zero study writes and zero locks.
- Lock: busy study (active, idle, suspended session, claim), changed study (still valid and no
longer valid), study locked by another job, project in
CalculatingInclusionInfo; each proves all locks released and the busy list returned. - Every write path in the table against a locked study: refused, skipped or deferred as stated,
and nothing written. Also: load a study while it is locked, release the lock, then save the loaded
copy; the save must fail its version check and must never write
BulkUpdateLockback. The architecture test fails when a new unguarded study write is added. - Takeover: a takeover at generation g+1 applies, rolls back and releases locks taken at g, while the old attempt's guard writes at g abort.
- Crash at each phase: kill the executor (process exit, not exception) at every row of the crash table, including mid-transaction in apply and rollback; prove the takeover reaches the stated outcome, the old attempt's late write is fenced by the generation, and the final content hash of every study equals either all-before or all-after.
- Rollback triggers: batch failure, cancellation, M5 denial between batches, lease ceiling,
recovery-fence generation bump; each ends
RolledBackwith exact content. - Rollback quarantine: an out-of-band write to a locked study (bypassing the predicate in the test) is detected by hash and leaves locks held with an alert.
- Statistics: with FEAT-024 writes and the fold on, the fence stays raised through rollback and derived statistics match the restored population.
- Concurrency: two v2 jobs on overlapping studies; reviewers sampling a pool that contains locked studies; v1 and v2 jobs in one project.
- E2E: extend
bulk-update-recoverywith a v2 success, a forced mid-apply failure that restores every study, and a busy-study refusal.
Implemented crash suite: BulkStudyUpdateAtomicBrokerTests
(#3916). It submits a version 2 job exactly as the notifier
does, over a real replica set, real RabbitMQ, the MassTransit job service and Quartz. The production start and
run consumers handle it.
- The first attempt is killed at every row of the crash table: at the start of planning, after planning, locking, between lock batches, apply, between apply batches, after commit, release, between release batches, rollback, and between rollback batches.
- The job's own retry takes over at the next generation, and the suite proves the stated outcome. Every study ends all-before or all-after by content hash, and no lock remains.
- With the restoration guard off, the version 1 parse refuses the version 2 job before writing, and the job still runs atomically to completion. It never falls through to version 1.
- The old attempt's late write being fenced, and quarantine, are proven at the executor level
(
BulkStudyUpdateExecutorTests).
Implementation slices¶
- Lock field, write predicates on every path, 409 contract and architecture test (dark).
Implemented in #3909; the write-path table above is
enforced by
StudyWriteLockArchitectureTests, and.claude/rules/bulk-study-locks.mdrecords the invariant for new write paths. - Operation record, guard, lease, takeover and watchdog; plan phase with no writes. Partly implemented in #3911, dark: the operation record with its lease, generation and guard write, the planner, the plan-chunk and before-image stores. Nothing calls them yet; the executor that drives them came with slice 3, and the lease watchdog is still a follow-up. See implementation notes.
- Lock and release phases; busy-study refusal and UI list.
- Apply with before-images; commit and deferred effects.
- Rollback and quarantine; crash-at-each-phase suite.
Slices 3 to 5 (except the UI list) are implemented in #3913,
dark: BulkStudyUpdateExecutor and MongoBulkStudyUpdateStudyWriter, proven on a replica set (success, every
refusal, D3 both ways, rollback on denial, cancellation and the ceiling, a crash taken over at the next
generation with the old attempt fenced, quarantine, and two overlapping operations). The job wiring, the flag
and the authority recheck's M5 binding shipped in #3914. The crash-at-each-phase suite over a real broker
shipped in #3916; see the test strategy. The lease watchdog
is still a follow-up.
6. M5 recheck between batches and the rollback obligation (after P9 approval).
Implemented in #3914: DelegatedWorkBulkStudyUpdateAuthority
(dark with DelegatedWorkAdmissionEnforced; unknown authority stops and rolls back) and the approved obligation
bulk-study-update-rollback.v1.
7. Flag wiring, UI statuses and copy, user-guide update, staging rehearsal.
Reviewer copy and the user guide are implemented in #3915:
the review-save 409s (StudyLockedByBulkUpdate, with the start time in the reviewer's own time zone, and
BulkStudyUpdateInProgress), the hub's LockedByBulkUpdate denial dialog, and the "All-or-nothing updates"
section of Managing Studies. Dedicated job statuses, the study-list badge and the staging rehearsal remain.
Flag wiring is implemented in #3914: the generated
bulkStudyUpdateAtomicApply flag (API and project-management), version 2 stamping at job creation, the
version check widened to {1, 2} as in rollout step 2, the handoff to IRunBulkStudyUpdateOperationCommand
(30-minute job timeout, retries after the lease) and the host registration. Job retries take over a crashed
attempt; a scheduled lease watchdog for a job that exhausts its retries is still to come.
Implementation notes¶
Recorded as the slices land; each note says where the code departs from, or sharpens, the text above.
- Input manifest (slice 2). The S3 file service exposes no object version, so the manifest is the SHA-256
and length of the bytes the plan actually consumed, recorded on the operation record when planning starts. A
re-plan that reads different bytes is refused (
InputChanged), which is the recovery ADR's fail-closed rule without relying on bucket versioning. - Plan determinism (slice 2). The set of studies, the identifier changes and the fields touched are fixed by
the rows and the stored documents, but a screening decision records when it was made. A re-plan after a crash
in planning therefore yields a new digest. Nothing reads a plan before the operation reaches
Planned, and the digest binds the plan that is applied, so this is harmless. - Change set (slice 2). A study's change is the new value of each top-level field its rows touch:
CustomId,PdfRelativePathand, for screening rows,ScreeningInfo. The planner's in-memory$set/$unsetappends new fields in lexicographic order, as MongoDB 5.0+ does; a replica-set test proves a server-applied plan hashes to the planned post-change hash. - Indexes (slice 2). The three collections and their indexes (including the 30-day TTL on before-images) are created on first use, so nothing exists in an environment until a version 2 operation runs there.
- Routing with the guard off (rollout step 3). With
BulkStudyUpdateRequireCurrentExecutionVersionon (production), the start consumer's admission read already knows the job's version and hands every version 2 job to the executor, whatever the flag says. With the guard off, the consumer makes no extra read: the version 1 parse refuses a version 2 job in its first write, before saving anything (BulkStudyUpdateRequiresAtomicExecutionException), and the consumer hands it off. A version 2 job never runs the version 1 path, and a version 1 job makes exactly the calls it made before. - Failures before any lock (slice 6). The whole file is held in memory to bind the plan to its bytes, so
the read stops at
MaxInputBytes(50 MiB) and the file is refused withInputTooLarge. Any other failure while planning or preparing to lock ends the operation as Failed at once, as version 1 fails the job on a parse error; nothing is locked yet, so there is nothing to retry for. Failures after the first lock keep the lease-and-retry path. The embedded job is projected only after re-reading the record at this attempt's generation, so an attempt that was taken over never reports a stale phase.
Decisions¶
Chris accepted every recommendation below on 2026-10-02. The decision for D8 is approval of the
proposed P9 obligation bulk-study-update-rollback.v1.
| # | Question | Decision (accepted 2026-10-02) |
|---|---|---|
| D1 | Does one invalid or refused row fail the whole job? | Yes. All-or-nothing applies to validation too; the user fixes the file and resubmits. |
| D2 | After a crash mid-apply, roll back or roll forward? | Roll back. One recovery path, matches "stopped means clean rollback"; roll-forward is possible later because apply is idempotent. |
| D3 | A study changed between planning and locking: re-validate under the lock, or refuse? | Re-validate under the lock, and refuse only if its row is no longer valid. |
| D4 | Review in progress on a target study: refuse or wait? | Refuse with the busy list. Waiting can be added later as an explicit, bounded option. |
| D5 | Recheck authority between apply batches, or only at each attempt? | Between batches. A denial is seen before the next batch starts, so no batch is in flight and the operation rolls back. |
| D6 | Lease, ceiling and job timeout | 2-minute lease, 30-minute ceiling (2-hour hard maximum), v2 job timeout equal to the ceiling. |
| D7 | Before-image retention | 30 days, then TTL deletion; a permanent value-free summary on the operation record. |
| D8 | Approve rollback as a P9 system obligation (bulk-study-update-rollback.v1) |
Approve, restricted to restore-and-release for an admitted operation. |
| D9 | File size ceiling | 10,000 changed studies, configurable; refuse larger files with a "split the file" message. |
Consequences¶
Positive¶
- A Bulk Study Update either changes every study in the file or none, including after crashes, cancellations, timeouts and authority denials.
- The before-images and plan are the attempt-scoped ledger the recovery ADR requires, so v2 jobs are recoverable without operator reconciliation.
- Reviewers can never have their work overwritten by a rollback, because they cannot write a locked study.
Negative¶
- Every study write path gains a predicate and a refusal case, and future write paths must be registered in the architecture test.
- Reviewers are kept off locked studies for the apply window, and a busy study blocks the whole job at start.
- A file with one bad row no longer applies its good rows (a deliberate behaviour change, D1).
- Three new collections, a new study field and a scheduled watchdog add operational state to
monitor; a
RollbackQuarantinedoperation needs an engineer.
Alternatives considered¶
- One transaction for the whole job. Rejected: Atlas transaction lifetime and size limits make it fail for ordinary file sizes, and a long transaction conflicts with every reviewer save in the project.
- Validate everything first, then apply without locks or before-images (Option B). Rejected: it removes invalid-row partial commits but not crash, timeout or denial partial commits.
- Before-images and compensation without locks. Rejected: a reviewer's save between apply and rollback would be overwritten or would make the restore inexact. Chris required locks for this reason.
- Lock records in a separate collection only. Rejected: every write path would need an extra read; a field on the study is enforced by the write itself and visible to every reader.
- Wait for busy studies. Kept as a later option (D4).