Skip to content

Phase 6A bounded family fleet operations

Usable boundary and acceptance

This first operational slice lets an application administrator create an explicit bounded batch of project/family targets, process one target per request, inspect durable progress, pause and resume, and stop new fleet admissions globally. It composes the registered family backfill, rebuild and bootstrap coordinators. It introduces no scheduler, project discovery, wildcard admission or deployment activation. Families without a registered runtime backfill are rejected at creation. Registered family adapters include project, membership and reviewer screening; stage, membership-stage and reviewer annotation; question answers; search population; and domain reconciliation. All nine physical families use the same bounded dispatch protocol. Derived summaries do not need a physical backfill.

Acceptance is durable progress across process restart, one valid global admission lease, a server-owned minimum interval between admissions, denial of stale-owner progress, cancellation of a worker that loses renewal, and replay after publication without changing already-current screening rows or duplicating the bootstrap checkpoint. Source parity remains each family's existing authoritative calculation.

Configuration and authorization

Every action uses BatchAdminProjectsPolicy, the existing application administrator-only activity. Project membership alone cannot operate the fleet. The runner is disabled by default.

Server configuration under ProjectStatistics:Fleet:

Setting Default Enforced bound
Enabled false Required for create and advance
MinimumIntervalSeconds 30 Clamped to 10–3600 seconds
LeaseSeconds 300 Clamped to 60–1800 seconds

These are ordinary server IConfiguration settings, including the standard .NET environment-variable form ProjectStatistics__Fleet__Enabled; no runtime UI override is registered. Configuration is read at host startup. There are no client-supplied interval or lease settings. Creation and each dispatch also require the reviewed project allowlist, materialized writes and target family flags. The existing backfill coordinator rechecks its gates and owns durable epoch/fence publication admission.

Status, pause/resume and global stop controls remain available when the runner is disabled. Global stop blocks creation of new runs and new project admissions. An identical create retry can retrieve its existing run while stopped. Resuming a run does not clear the independent global stop.

Administrative protocol

All paths below are relative to /api/admin/project-statistics/fleet.

  1. POST runs with { "runId": "<stable-guid>", "targets": [{ "projectId": "<guid>", "family": "ProjectScreening" }], "repair": false } creates or retrieves a run (202). Reuse the same ID on an ambiguous network result. Identity binds the administrator, repair mode and canonical sorted project/family targets; changed payload under that ID returns RunIdentityConflict (409).
  2. GET runs/<runId> returns the durable cursor, status, attempts and latest result per target. The run ID is the resume token; callers never submit a cursor to advance past work.
  3. POST runs/<runId>/advance executes at most one project/family target and returns current progress (200). Poll GET control to inspect NextAdmissionAtUtc. A premature admission returns FleetRateLimited; another valid lease returns FleetBusy (409). Transaction contention returns FleetContended; the server bounds a transaction attempt to ten seconds (FleetOperationTimedOut). Inspect status and retry after a refusal. There is no queued retry loop.
  4. POST runs/<runId>/pause stops further admissions for that run. An admitted project finishes. POST runs/<runId>/resume permits retry of its current cursor, including a recorded failure.
  5. PUT stop with JSON true stops new fleet admissions. JSON false clears that stop. GET control reports stop and current lease state.

A refused publication, absent project or incomplete bootstrap pauses the cursor with a typed bounded result. Each result carries the backfill's lifecycleStatus and failureReason enum names, so an operator can tell retryable contention from a fence, digest mismatch or capacity refusal; free-text detail is not persisted. Resolve the cause and resume; no project is silently skipped. Unexpected faults or cancellation record Failed or Interrupted when the owner still holds its lease and then propagate the failure. A lease lost mid-dispatch returns FleetLeaseLost (409) rather than an unhandled cancellation. Pause and resume declare the same 409 refusals (FleetContended, FleetControlChanged, FleetOperationTimedOut) as the other mutating actions. An expired lease can be reclaimed at the same cursor without an explicit failure record.

The API returns the durable run's canonical order, which may differ from submission order. There are at most 32 distinct project/family targets per run and 64 retained run documents across the fleet. Results replace the previous result for a retried cursor rather than growing a history list. The 65th new run returns FleetRunCapacity; existing status, resume and stop controls continue to work. This bounded pilot does not silently delete progress. Administrators can reclaim completed records using the explicit pruning protocol below.

Completed-run pruning

POST runs/prune accepts { "runIds": ["<completed-run-id>"] }, with 1–32 distinct non-empty IDs. Only application administrators can call it; the server records their authenticated identity. The entire request fails with RunNotCompleted (409) if any existing run is unfinished or currently claimed. Validation and deletion share the control transaction with creation and claims, so a refused batch deletes nothing. Missing IDs are harmless on retry. No projection, checkpoint, receipt, source record or paused run is deleted. Pruning remains available while fleet dispatch is disabled or stopped.

The control retains the latest 32 bounded pruning records, each containing a monotonic sequence, server observation time, administrator, requested IDs, deleted run and target counts, and a count of the deleted results per typed LifecycleStatus/FailureReason pair (results written before those fields existed are counted under null names rather than dropped). Export these records before they rotate if longer audit retention is required; this is a finite operational audit, not permanent history. At most 1,024 requested IDs are retained across all audit records.

Archive required completed-run evidence before pruning. Deleted run IDs no longer provide status or create-request deduplication: callers must retire them and use new IDs for future runs. A delayed create retry using a pruned ID can create another run; normal backfill remains idempotent, while forced repair is already at least once. Pruning is explicit because it ends this operational evidence and retry window.

Single-family screening runs retain their original fleet.v1 request digest. Canonical ordering adds family as a secondary key, so a project can request multiple distinct families without duplicate targets. Each dispatch rechecks its own family flag and calls that family’s backfill; a disabled family pauses the cursor instead of falling through to screening.

Concurrency and restart semantics

Admission, pause/stop, renewal and progress updates serialize through a single Mongo control document in snapshot transactions with majority commit. The default _id indexes give unique run and control identities. The control stores the next admission time durably, so process restart cannot reset rate limits. Transactions also serialize the retained-run count and insert. Fleet-control reads tolerate additive BSON fields, and control writes update only fields this version owns; older replicas therefore preserve newer prune-audit fields instead of rejecting or erasing them. Only a prune rewrites the sequence and bounded audit list; ordinary claims, renewals and completions leave them untouched, and audit records tolerate unknown fields on read. Re-typing existing fields still requires an explicit compatibility plan.

An advance receives a unique token and increasing lease generation. Renewal runs every third of the lease period. Failed renewal cancels the token passed to the backfill and no subsequent project is dispatched. Completion requires the same token/generation, an unexpired lease and the expected cursor. A replacement worker cannot have its progress overwritten by the late owner.

This bounds valid admissions, not the number of physically executing processes during a process pause or network partition. Cancellation is cooperative; an already-running atomic publication may finish. Existing per-scope leases, epochs, source fences and publication guards remain the authority for its correctness. Global stop lets the currently admitted publisher renew and finish safely, then blocks the next project. It is not an instruction to roll back committed source or projection data.

After a crash between publication and progress, restart repeats the same project. Ordinary backfill recognizes already-current rows and its existing bootstrap point; no generation or history churn is needed. Forced repair is deliberately at least once: repeating a completed forced publication can create a further valid generation. Fleet progress is not an exactly-once source-operation receipt.

Evidence and required continuation

The replica-set regression exercises competing store instances, durable rate limits, expired-owner rejection, global stop, pause/resume, finite storage admission and replay through the real screening calculator/rebuild/bootstrap pipeline. Replay compares published counts to the production authoritative aggregation and checks unchanged publication generation and a single bootstrap checkpoint. Service regressions cover default-off storage isolation, invalid family/batch denial and cancellation on failed renewal. HTTP authorization uses the real application policy and deployed permission catalogue.

These remain required ordered work before full Phase 6A completion:

  1. Establish the remaining approved family backfills and their parity proof. Fleet dispatch now uses registered family services, so it cannot invent a missing calculator or grant an unregistered family admission. A mixed-family run retains separate durable progress for each target.
  2. Hard bounds for projection/history/receipt retention and cleanup work, including avoiding accumulation of every root after paged reads in existing cleanup. Completed fleet-run pruning is implemented separately with finite audit retention. This slice bounds its own progress documents; it does not prove bounded checkpoint/delta/receipt growth across the fleet.
  3. Scheduled bounded dispatch and deployment wiring only when operational activation is authorized; the explicit advance protocol already provides a usable administrator-driven pilot.

Phase 6B's actual seven-day staging soak, representative read/mutation volume, exact parity, fallback, no unresolved rebuild backlog and bounded storage-growth evidence remain live activation/evidence gates. Local tests and existing checkpoint retention regressions do not substitute for that soak. Credentials, allowlist approval and production activation are separate; this change enables none of them.

Combined programme integration

The fleet branch includes the operational family baselines, daily observations and guarded consumer stack through #3572. Registration tests resolve the fleet with all nine actual family adapters; the all-family dispatch test verifies independent routing and preserves the original screening-only run identity. Real Mongo tests cover multi-family publication, lease recovery, pause/resume, stop and rolling-reader field preservation. This integration does not enable the runner, any consumer flag or a deployed scheduler; sustained-load and operational acceptance remain separate.

Explicit bounded checkpoint history maintenance

POST /api/admin/project-statistics/{projectId}/history-maintenance runs one per-project maintenance transaction. It requires the application administrator policy, the existing explicit project allowlist, and server configuration ProjectStatistics:HistoryMaintenance:Enabled=true (default false). The only timer is the separately default-off scheduled maintenance job; there is no deployment change or implied permission to execute maintenance in staging or production. The operation deliberately remains available with serving flags off so rollback does not require re-enabling application reads to maintain retained data.

Send {} for the initial observation page. The response reports roots scanned/removed/deleted, reference pages deleted, observations scanned/deleted, nextObservationId, markersReclaimed, unchangedDaysDeleted and verificationsDeleted. Send {"afterObservationId":"<returned nextObservationId>"} to continue the observation sweep; null ends that sweep. Start a later sweep at null so observations referenced during an earlier sweep can be re-evaluated after their final referring root is removed. Retention/root cleanup independently resumes from durable root state on every request. Retrying after interruption is safe.

One invocation reads at most 1,025 root headers and refuses with HistoryScanCapacity, without changes, if more than 1,024 roots exist. It marks at most 32 roots removed, deletes at most 32 reference pages and 32 removed roots, and examines at most 32 observations. This explicit refusal avoids an unbounded historical scan; projects already beyond that admission ceiling need a separately reviewed migration. The surviving-root selection retains the existing daily/weekly/monthly policy and 256-root cap. Reader grace starts when a root is removed, and cleanup waits the configured lifecycle grace interval.

Checkpoint admission is compare-and-set in the same snapshot transaction as cleanup. An occupied build slot refuses maintenance (CheckpointBuildBusyOrUninitialized); a racing admission conflicts rather than publishing a reference to an observation being deleted. An admission attempted inside an open pass fails its own transaction and retries after the pass commits; one that committed after the pass read the admission row aborts the pass, which returns 409 CheckpointAdmissionChanged without changes. Any other transient transaction abort (for example a primary stepdown or network error) that leaves the admission row unchanged returns 409 MaintenanceTransactionConflict, also without changes; both mean "retry the request". Only these typed refusals are 409s: any other failure is an invariant break and returns 500. Every remaining reference page, including one belonging to another removed root awaiting cleanup, protects its observations. A bounded global observation keyset scan prevents deleted creator roots or a first page of referenced observations from losing orphan discovery. Cancellation before commit aborts the complete pass; the next request restarts from committed state. The older unbounded internal retention loops (ApplyRetentionAsync, RunCleanupAsync) had no production caller and were removed in #3674; RunCleanupAsync could leave a retention-failed root with some pages deleted that readers still treated as retained.

Checkpoint build marker reclamation (#3674). Checkpoint admission refuses a new build (CheckpointCapacityExceeded) once a project holds more than 512 terminal (Published or Cleaned) build markers. Without reclamation, every retention-removed root left a Cleaned marker behind, so a project building one checkpoint a day stopped admitting after roughly 512 days and daily history stopped. Each pass now also deletes at most 32 terminal markers whose build token has no root of any retention state, inside the same admission-CAS snapshot transaction as the rest of the pass. A marker is reclaimed only once it provably owns no reference page and no creator observation; observation collection is keyed on the creator marker's Cleaning/Cleaned state, so the marker of a removed root whose observation another root still references stays Cleaning until that observation is collected. A retained root's Published marker is never touched, and Building, Abandoned and Cleaning markers are never candidates (an unpublished build also holds the slot, which refuses the whole pass). Steady state is therefore about one marker per retained root (at most the 256-root cap) plus markers awaiting their final cleanup. A racing admission conflicts with the pass exactly as above: an admission inside an open pass cannot commit and retries afterwards; one committed after the pass read the admission row aborts the pass, which reclaims nothing and returns 409 CheckpointAdmissionChanged. Repeat the call. Real Mongo tests run 552 daily builds with weekly maintenance and keep admitting (the same test is refused at day 513 without reclamation), keep retained and still-referenced markers, reclaim a cleaned abandoned build, and cover both race orderings.

Unchanged-day reclamation (#3636). The daily runner records an inactive project's day in pmProjectStatisticsUnchangedDay instead of building a checkpoint (see the historical charts rollout). Each pass also deletes, inside the same admission-CAS transaction and oldest first, at most 32 of those records that are older than the daily retention window (90 days), that reference a root which is no longer retained (the reader already treats them as gaps), or that are the oldest beyond 128 per project. It reads at most 160 records to decide. The scheduled maintenance run counts these deletions as progress and reports them as the unchanged_days reclaimed item. Real Mongo tests prove 200 records converge to the retained window in passes of at most 32 and that the cap applies inside the window.

Verification-record reclamation (#3636 part 2). When copied daily snapshots are enabled (ProjectStatisticsDailyObservations:CopyFreshMaterializedRows), each daily build writes one pmProjectStatisticsCheckpointVerification record before publishing. The same pass deletes, at most 32 per pass, records whose root no longer exists (deleted by retention, or a build that never published). The admission slot is free during the pass, so no build is between writing its record and publishing. A live root's record is never deleted. The run reports these as the checkpoint_verifications reclaimed item.

Real Mongo tests cover shared-observation survival, reader grace, final-reference collection, active build refusal, both admission race orderings, a transient abort without an admission change, interrupted-transaction rollback/retry, observation cursors, 32-root mutation bounds, the 32-page bound with a removed root whose pages span two passes (the root stays Compacted until its last page goes), capacity refusal before changes and disabled/unlisted refusal without writes. A history read through the production bundle reader before, during and after a pass serves every retained root completely (including one sharing the removed root's observation) and reports the removed root as Compacted until its cleanup deletes it. HTTP tests cover anonymous/ordinary-user denial and default-off admin refusal. This adds executable checkpoint maintenance; delta/receipt storage-growth proof and the live soak acceptance gates remain outstanding.

Enabling copied daily snapshots

ProjectStatisticsDailyObservations:CopyFreshMaterializedRows (environment variable ProjectStatisticsDailyObservations__CopyFreshMaterializedRows, default false) makes the daily runner copy fresh current rows instead of recalculating them (see the historical charts rollout). It is a per-environment deployment change on the Project Management host. Enable it only in this order:

  1. Every replica runs the build first. Confirm that every API replica and every Project Management replica runs a build containing #3725 and #3727, and that no older replica set is still serving (the rolling update has completed). An older API replica ignores verification records and shows copied points as ordinary, unmarked history; it also saves a stage's session-count target without invalidating Reviewer annotation, so a copy could pair old counters with the new stage definition.
  2. Daily observations already run. ProjectStatisticsDailyObservations:Enabled=true has run cleanly for the target projects with the option off, and history maintenance is enabled so verification records of deleted roots are reclaimed.
  3. Pilot size is measured. The copy reads every scope's row and fences inside one snapshot transaction (read batching by family is a follow-up). Pilot on projects whose membership × stage scope count is known, and watch for copy faults: a fault releases the build at once and the next tick retries, but a project that exceeds the snapshot lifetime fails the same way each time.
  4. Set the option on the Project Management host and watch daily_observations.scopes by provenance; copied should dominate for active projects and calculated should cover search population plus any non-servable scope.

Rollback. Set the option back to false: new daily builds calculate every scope again. Existing Unconfirmed roots keep their state until the drift check decides them. Before rolling back to a binary older than #3725/#3727, set the option to false first.

The build marker and root of a copied snapshot carry IsAuthoritativeBuild = true, because they were published through the authoritative protocol. That flag is not provenance: only the root's pmProjectStatisticsCheckpointVerification record says which scopes were copied. Do not read the flag as evidence that a snapshot was calculated.

Explicit bounded delta maintenance

POST /api/admin/project-statistics/{projectId}/delta-maintenance performs one ledger-compaction pass. It requires application administrator authorization, the project allowlist, and independently default-off server setting ProjectStatistics:DeltaMaintenance:Enabled=true. No live maintenance is enabled by this change; the optional scheduled job is default-off. The older unbounded CompactAsync loop had no production caller and was removed in #3674; its read-only watermark calculation remains.

The pass uses a single snapshot transaction. It serializes with checkpoint admission and compare-and-sets the project control with the floor update. Existing family guards are also compare-and-set to serialize with source-fence/candidate admission. Any active checkpoint build, live rebuild lease or nonterminal publication coordinator or family fence/candidate guard refuses the pass; a truncated list of owners is never interpreted as a complete pin set. Retained roots, outstanding reconciliation, pending notification slots and the maximum redelivery window still restrict eligibility. Root inspection has the same 1,024-header admission ceiling as history maintenance. Seal reads cap at 65 to detect the 64-seal ceiling without silently omitting older evidence.

At most 33 exact rows are inspected and 32 are deleted. The extra row proves that the last selected projection revision is complete. Adjacent seal time bounds use the minimum/maximum observed times, so clock skew cannot move an audit boundary backwards. A sibling group exceeding the batch, a gap in revisions, or any young sibling stops selection; nothing skips over such a boundary. The result reports Compacted or NoEligibleCompleteRevision, the count deleted, the durable replay floor and the safe watermark. Repeat the call to continue from the committed floor. A no-op is a safety result, not evidence that the ledger is empty or storage growth is bounded. Oversized sibling groups require a separately reviewed larger-work protocol; this endpoint never splits them to force progress.

The seal insert or adjacent-seal merge, exact-row deletion, and replay/identity-floor advancement commit atomically. Interrupted work leaves all of them unchanged. The identity floor moves to the newest compacted observation but always stays strictly below the oldest observation of any exact row left above the new revision floor, so a clock-skewed surviving row is never classified too old. That proof reads at most 129 remaining rows in replay order; when more remain, the pass advances only the revision floor and leaves the identity floor where it was. A source write, checkpoint admission, guard transition or rebuild that commits a row the pass read after its snapshot aborts the whole pass with a Mongo write conflict; the service aborts and refuses with the typed ConcurrentStatisticsWrite reason (HTTP 409), and the caller simply repeats the call. A commit whose outcome stays unknown after the bounded commit retries refuses with CommitOutcomeUnknown (HTTP 409): the seal, deletions and floor moves either all committed or none did, and a repeat pass re-reads the floor, so repeating is safe. Every refusal and every write conflict is therefore a typed 409; only a cancelled request (the caller went away) and an unexpected server fault are not. Adjacent batches normally retain one merged seal; a full nonadjacent seal set refuses additional compaction. Cancellation is checked between bounded steps. Source-operation receipts and their idempotency floor are never deleted or advanced here: receipt reservation, pin/reachability and redelivery reclamation need their own operational contract.

Real Mongo regressions cover bounded continuation, shared-revision page edges, young siblings, retained checkpoint pins, active-owner refusal, rollback after an injected deletion interruption, the ledger-gap stop, the full nonadjacent seal set, the guard/root/notification scan ceilings and an unresolved commit outcome. Against real coordinator writes they also prove the visible projection and control clocks are identical before and after compaction, a rebuild from the authoritative source agrees with the compacted projection, a redelivered operation whose delta was compacted is still refused by its receipt (source applied once, counters unchanged), and a source write, checkpoint admission or rebuild racing an open pass either loses its own conflict or aborts the pass typed, with no delta committed during or after the pass ever deleted by it. Redelivery idempotency rests on source-operation receipts, which this pass never touches; production callers test only for a receipt. A duplicate whose delta was compacted reports its replayed routing as SourceOnlyFallback with projection revision 0, which no caller consumes today. HTTP tests prove anonymous/non-administrator denial, the default-off administrator result and production container resolution. Live fallback/rollback, storage-growth and sustained-load acceptance remain unproven until captured in the approved environment; local compaction tests do not replace that evidence.

Explicit bounded source receipt maintenance

POST /api/admin/project-statistics/{projectId}/receipt-maintenance requires application administrator access, project allowlisting and independently default-off server setting ProjectStatistics:ReceiptMaintenance:Enabled=true. This deployment-only setting also gates the two new receipt BSON identity fields: while disabled, receipts physically omit SourceAggregateId and HasTrustedSourceIdentity, so older strict readers can consume newly written receipts during a mixed-version rollout. Before enabling it, deploy the retirement-aware screening and annotation writers to every API and project-management mutation host and initialize the new indexes. Set it coherently on both hosts only after all old strict readers have exited. Keep it disabled during mixed-version deployment. No live cleanup is performed by this change.

Disabling the maintenance flag safely stops further reclamation and new identity-field emission, but does not make a binary rollback to old strict readers safe while enabled-mode receipts remain in the collection, or to retirement-unaware source writers safe after any receipt has been deleted. Every mutation host must retain the retirement-guard protocol thereafter. An older writer binary requires coordinated restoration of the database/receipt backup and reconciliation before resuming writes; ordinary application rollback must not discard the durable retry protection while keeping the compacted database.

The minimum usable slice reclaims only enabled-mode point receipts whose server factory explicitly bound a Study identity in stats.source.screening or stats.source.annotation. Legacy receipts without that provenance, child receipts, pinned receipts and other namespaces remain retained. No operation-ID parsing or fabricated creation age establishes eligibility. The old internal unbounded reclamation helper (AdvanceIdempotencyFloorAsync) had no production caller and was removed in #3674.

Each request reads at most 32 receipts using an explicitly hinted index on project, pin state, trusted identity, creation time and ID. It checks observed age within that page rather than scanning an arbitrary prefix to find 32 eligible rows. Both creation and observation must satisfy the larger audit/redelivery window. The time floor uses millisecond precision and deletion requires a strictly older stored creation time, preserving envelopes whose sub-millisecond fraction was lost in BSON. The earliest pinned creation time caps the floor, independently of commit observation order. A young observed receipt can temporarily occupy a page; the response counts scanned and deleted separately. Repeat requests continue after committed deletions, including when the time floor has not advanced.

Deleting a receipt also advances a monotonic source-revision retirement row in pmProjectStatisticsReceiptRetirement, keyed uniquely by project, Study and the fixed source namespace. There are at most two constant-size rows per Study and a hard 100,000-row project ceiling. Existing rows may still advance at the ceiling; new keys refuse the whole transaction before any deletion. The guard uses explicit $max updates, preserving unknown fields. Rows are durable retirement evidence and are not themselves pruned. Active publication coordinators and family fences/candidates refuse cleanup. Receipt deletion, retirement watermark updates and the project control CAS share one snapshot transaction; cancellation rolls the complete pass back. Source writers CAS that same control, preventing a writer which read an absent retirement row from committing after a concurrent cleanup.

Receipt-first replay preserves committed retries, including legacy receipts. With no matching receipt, a trusted envelope at or below its source retirement revision fails TerminalTooOld, even when a new server timestamp was minted for the old identity. Once a project has retirement evidence, a missing receipt in either supported namespace also requires explicit source metadata. A legitimately new source revision above the watermark remains admissible. Existing reviewer retry loops only handle duplicate or Mongo concurrency outcomes; they do not recapture and reapply TerminalTooOld intent. That rejection requires refresh/reconciliation rather than silently repeating a decision.

The guard adds indexed retirement reads to receipt-miss preflight and transactional admission. Command budget tests measure 23 Mongo commands per uncontended screening save (previously 21), including exactly two retirement lookups; they do not establish the programme's wall-time performance target. Real Mongo tests cover reminted identity rejection, normal new writes, retained committed retries, legacy/pin protection, bounded continuation, source snapshot conflicts and interruption rollback. Live storage-growth, fallback/rollback and elapsed soak acceptance remain unproven.

A source write that commits, or holds an uncommitted control write, after the pass took its snapshot aborts the pass with a Mongo write conflict. The service aborts and returns the typed ConcurrentStatisticsWrite refusal (409), never a raw 500; retirement rows, deletions and the floor roll back together, the source write commits exactly once, and the administrator simply repeats the call. Real replica-set tests drive the real transaction coordinator after a pass: the original envelope is refused by the floor, the same identity re-minted above the floor and any never-seen identity at or below the Study's retired revision are refused by retirement, an envelope without source metadata is refused, the source and counters never move twice, and new revisions of the same or another Study still commit. A missing IX_StatsReceipt_TrustedMaintenance index refuses with the typed ReceiptMaintenanceIndexMissing (409) rather than a server error; run index initialisation, then repeat. A pinned receipt already stranded below an existing floor refuses with PinnedReceiptBelowExistingFloor and writes nothing. Every 409 reports only the refusal's reason name. The 100,000-row ceiling is exercised with real rows: a new Study key fails closed with ReceiptRetirementCapacity and deletes nothing, while an existing key still advances.

Child/inverse and other source namespaces need a separate retention protocol that proves their durable retry authority before deletion. They are explicitly retained by this MVP; do not turn their cleanup on by widening the namespace predicate or deriving identities from operation-ID text.

Scheduled bounded maintenance

The Project Management host can call the three bounded services above on a Quartz schedule, so an allowlisted project that builds a checkpoint every day keeps its roots, build markers and delta ledger bounded without an administrator repeating the POSTs. The job adds no storage protocol: each pass is one ordinary call to the history, delta or receipt service, with the same gates, transaction and typed refusals as the administrative endpoint. Nothing is enabled by this change, and no environment configures it.

Configuration

All keys are read by the Project Management host only (environment variables carry the SYRF__ prefix and use __ for :).

Key Default Meaning
ProjectStatisticsMaintenance:Enabled false Registers the schedule and lets the consumer run. Off, the consumer acknowledges a tick without work even if an earlier Quartz trigger still delivers.
ProjectStatisticsMaintenance:CronExpression 0 30 2 * * ? Quartz cron, evaluated in UTC (daily 02:30, clear of the 00:05 daily observation tick). A change takes effect on the next start: the same schedule identity is re-registered with the new trigger.
ProjectStatisticsMaintenance:ProjectsPerRun 16 Allowlisted projects visited per run (1–256).
ProjectStatisticsMaintenance:PassesPerProject 4 Passes of each service per visited project per run (1–32); a pass that makes no progress ends that service's passes early.
ProjectStatisticsMaintenance:ReceiptMaintenanceEnabled false Separate opt-in for receipt reclamation (see rollback below).
ProjectStatistics:HistoryMaintenance:Enabled false The history service's own gate, which the schedule does not bypass.
ProjectStatistics:DeltaMaintenance:Enabled false The delta service's own gate.
ProjectStatistics:ReceiptMaintenance:Enabled false The receipt service's own gate. Receipt reclamation runs only when this and ReceiptMaintenanceEnabled are true.
ProjectStatistics:ProjectAllowlist empty The reviewed allowlist, the only source of projects the job visits.

With Enabled=true, the host refuses to start unless the cron is a valid Quartz expression and both budgets are in range. A service whose own gate is off is skipped and counted as skipped with its *_maintenance_disabled reason; it is not called.

Bounds and continuation

A run never enumerates projects: it sorts the configured allowlist (minus any project the host's flag source does not also admit) and visits at most ProjectsPerRun of them, starting after the last project the previous run visited. A budget smaller than the allowlist therefore reaches every project within ceil(allowlisted / ProjectsPerRun) runs. History maintenance's observation cursor is chained within a run and persisted between runs, so a long observation sweep resumes instead of restarting; a finished sweep clears it. Both hints live in one singleton document, pmProjectStatisticsMaintenanceCursor, bounded by the allowlist size. They are only resume hints: each service re-validates everything it touches, so a lost cursor or two overlapping runs cost re-scanning, never correctness. One run executes at a time per process (the consumer's concurrency limit is 1), but two Project Management replicas can still run concurrently, for example a misfire-sent tick and the next regular tick landing on different pods. That is safe: every service serializes through its own compare-and-swap transaction, so the loser refuses typed (refused, contention) and the cursor is only a hint. If such contention shows up in the refused counts, a lease on the cursor document is the follow-up. Unlike the daily observation schedule, startup publishes no catch-up run; a missed tick is sent once by the misfire policy.

Per project and service, the per-pass service bounds above still apply (32 roots, pages, observations or markers per history pass; 32 revisions per delta pass; 32 receipts per receipt pass). At the default budgets one run does at most 16 × 3 × 4 = 192 bounded transactions.

Refusals and failures

A refusal writes nothing and ends that service's passes for that project only; other services and projects continue. Typed refusals fall into two groups, and they need different responses:

  • Contention: clears by itself. CheckpointBuildBusyOrUninitialized (a daily build holds the slot), CheckpointAdmissionChanged, ConcurrentStatisticsWrite, the other *Changed reasons, PublicationGuardBusy, RebuildActive, PublicationOperationActive, MaintenanceTransactionConflict (history) and CommitOutcomeUnknown (delta). Logged at warning level with the project ID, counted with outcome refused, and normally absent on the next run.
  • Capacity: recurs on every run until an operator acts. HistoryScanCapacity (more than 1,024 checkpoint roots, refused before anything is retired), PublicationGuardCapacity, NotificationScanCapacity, SealCapacity, ReceiptRetirementCapacity and ReceiptMaintenanceIndexMissing (the receipt index must be created). The schedule cannot recover such a project. These are logged at error level and counted with the distinct outcome capacity_refused so they can be alerted on (outcome="capacity_refused", grouped by reason). Each needs the separately reviewed migration or protocol named in its service section above.

An unexpected exception is logged at error level and counted with outcome failed, reason unexpected; it does not stop the other services or projects. The consumer never faults its message, so a failure cannot poison or dead-letter the queue; only shutdown cancellation propagates.

Telemetry

syrf.project_statistics.maintenance.passes counts passes by operation, outcome (completed, refused, capacity_refused, failed, skipped) and reason. A skipped pass names the switch that was off: history_maintenance_disabled, delta_maintenance_disabled or receipt_maintenance_disabled for a service's own gate, and scheduled_receipt_maintenance_disabled when the schedule's ReceiptMaintenanceEnabled is off, so "switched off" is distinguishable from "never reached"; syrf.project_statistics.maintenance.reclaimed counts items by operation and item (roots retired or deleted, reference pages, observations, build markers, delta revisions, receipts). Neither carries a project identifier; use the structured logs for that. See the runbook Telemetry section.

Enablement order

  1. Confirm the target projects are in ProjectStatistics:ProjectAllowlist and that the explicit administrative passes behave as expected for them.
  2. Enable the services' own gates on the Project Management host: ProjectStatistics:HistoryMaintenance:Enabled=true and ProjectStatistics:DeltaMaintenance:Enabled=true. Leave ProjectStatistics:ReceiptMaintenance:Enabled false.
  3. Set ProjectStatisticsMaintenance:Enabled=true, optionally with a cron and budgets. Watch maintenance.passes by outcome and reason for at least one run. refused (contention) passes should clear on the next run; any capacity_refused pass needs an operator before it will.
  4. Only after every source mutation host runs retirement-aware writers (see receipt maintenance above), and only if receipt reclamation is wanted, set ProjectStatistics:ReceiptMaintenance:Enabled=true and ProjectStatisticsMaintenance:ReceiptMaintenanceEnabled=true.

Rollback

Set ProjectStatisticsMaintenance:Enabled=false (or turn off any single service's gate). The consumer then acknowledges ticks without work, so while this binary is deployed no Quartz clean-up is needed; the stale trigger is harmless.

A binary rollback to a Project Management build without this consumer is different: the previously registered Quartz trigger keeps publishing IRunProjectStatisticsMaintenanceCommand, and the broker queue bound to it collects one unconsumed message per tick. After such a rollback, unschedule the recurring ProjectStatisticsMaintenanceSchedule trigger in the Quartz store and delete the maintenance consumer's queue (or disable the job and let one tick drain before rolling back). The daily observation schedule has the same property. History and delta maintenance leave nothing that blocks rolling back the binary. Receipt maintenance does: once one receipt has been reclaimed, stopping the job is safe but rolling back to a retirement-unaware writer binary is not (see receipt maintenance). That is why it has its own default-off switch.

Evidence

Real replica-set tests over the production services show that a disabled schedule writes no statistics row; only allowlisted projects are visited, within the per-run budget, and the rotation resumes on the next run; a busy admission slot refuses that project typed while another project proceeds, and the next run retries and succeeds; receipt maintenance never runs with its switch off; a service whose own gate is off is skipped; and 420 simulated days of one daily checkpoint followed by one scheduled run keep every project root (retained or awaiting deletion) under the 256-root cap and the 1,024-header scan ceiling, terminal markers under the admission ceiling, and every build admitted. Live soak evidence remains outstanding.

Periodic drift check

The Project Management host can recalculate every allowlisted project from source on a slow schedule and compare it with the statistics the reader serves (#3636 part 3). It is what verifies the daily snapshots that copy fresh materialised rows (ProjectStatisticsDailyObservations:CopyFreshMaterializedRows), and it is the only check that sees drift the activity clocks cannot: a write path that changed data without advancing them, or a period with statistics writes switched off. Nothing is enabled by this change, and no environment configures it. What a pass and a failure do to history is described in the historical charts rollout.

Configuration

All keys are read by the Project Management host only (environment variables carry the SYRF__ prefix and use __ for :).

Key Default Meaning
ProjectStatisticsDriftCheck:Enabled false Registers the schedule and lets the consumer run. Off, the consumer acknowledges a tick without work even if an earlier Quartz trigger still delivers.
ProjectStatisticsDriftCheck:CronExpression 0 30 3 ? * SUN Quartz cron, evaluated in UTC: weekly, Sunday 03:30, clear of the 00:05 daily observation and 02:30 maintenance ticks. It must fire less often than daily; a cron with any gap of 24 hours or less between fire times refuses host start.
ProjectStatisticsDriftCheck:ProjectsPerRun 16 Allowlisted projects visited per run (1–256), shared by persisted retry and normal rotations. Pending projects receive up to half the budget; a one-project budget alternates retry and normal visits.
ProjectStatisticsDriftCheck:ScopesPerProject 256 Scopes recalculated per visited project per run (1–4096). A larger project continues its check on the next run from a persisted cursor. Size this to include all copied scopes in one visit if their snapshots should become Confirmed.
ProjectStatistics:ProjectAllowlist empty The reviewed allowlist, the only source of projects the job visits.

With Enabled=true, the host refuses to start unless the cron is valid and less often than daily and both budgets are in range. The check compares only what the reader would serve, so it needs the same serving state as the daily copy: the fleet and project statistics modes enabled, the serving flag and the per-family serving flags on, and the project allowlisted by the host's flag source. A project whose gates are closed is counted refused and retried next run; nothing is written for it. Continuing, inconclusive, conflicting, refused and errored projects remain in a durable retry set; completed visits leave it. The normal rotation still advances in each run, so repeated refusals or a large check cannot indefinitely block other allowlisted projects.

What one visit does

  1. Enumerates the enabled families' scopes (the same calculators the daily build uses) in (family, scope) order and resumes after the cursor of a check already in progress.
  2. For each scope, reads the served row and the project identity in one pinned snapshot, then calculates the scope from source. The two are compared only when the calculation captured the same source and projection revisions, the fleet mode epoch has not moved, and the calculation pinned its own snapshot (a calculation that fell back to unpinned reads could have read its revisions and its counters at different times). Otherwise the comparison is retried once and, if it still cannot be made, the visit ends inconclusive and the next run retries that scope. A concurrent ordinary write is therefore never reported as drift. A null calculation carries no identity of its own: a family that can prove absence in a pinned snapshot at the served identity must also recheck that absence in the Stale transaction. Without that proof the visit stays inconclusive.
  3. On a mismatch, one transaction re-checks the identity and the selected row, marks the row Stale (LastFailureReason = DriftDetected), advances the project's source clock, coalesces the family's notification slot and records the failure on the check's state. If the identity or the row moved, nothing is recorded and the scope is compared again. Reads fall back to the authoritative path immediately.
  4. When every scope has been visited it reaches a verdict:
  5. Passed (at least one scope compared, none drifted): confirms bracketed Unconfirmed snapshots whose every copied scope the final visit found in parity, and records the new last passing boundary. A confirmation walk resumes across runs at its root budget and rechecks scopes on each visit; an admitted checkpoint inside the bracket must finish before the boundary advances and invalidates the earlier parity evidence. If a root is first requested at the pass identity after the pass committed, the next check includes that exact-boundary root only when it was absent at the previous pass; its copied scopes still have to verify. Scopes it could not compare do not stop the pass but are not verified.
  6. Unverified (no scope could be compared, for example every family fenced or not served): confirms nothing and leaves the last passing boundary where it was.
  7. Failed: marks Unconfirmed snapshots since the last pass Suspect only if they copied a drifted scope at or before that scope's own pre-Stale checkpoint (resumable if interrupted; nothing is confirmed until it finishes). A snapshot after one scope's repair can remain safe while copying other scopes, but is still suspect if another copied scope is later found drifted. The failure stays pending while an admitted build captured within the failure boundary can still publish; a changed admission generation restarts the bounded walk, and its final state claim fences checkpoint publication.

The cursor and check tally persist across visits, but parity evidence does not. At the start of a resumed visit, the check discards earlier verified scopes: a source mutation invisible to the activity clocks could have changed one between runs. A multi-visit check can still pass based on the scopes it compared, but a snapshot that copied scopes from different visits remains Unconfirmed. The next scheduled check visits those scopes again and can detect any drift. This keeps the per-visit calculation budget intact and avoids confirming history from stale evidence.

A passing worker first claims the versioned project check state as confirmation-pending. Only after that compare-and-set wins may it begin confirming verification records. If another replica has already recorded a failure, the would-be pass reports a conflict without confirming anything. If confirmation is interrupted, the pending state retains its boundary and bounded root cursor; the next run rechecks scope parity before it advances the confirmation walk. This also applies after the walk waits for an in-boundary admitted build, whose publication may have copied a source change invisible to the activity clocks.

Bracketing is strong evidence, not proof per snapshot. If a copied row drifted and was republished from source (by a backfill or rebuild) before the next passing check, that check passes and confirms the snapshot that copied the drifted value. A verdict walks at most 1,024 checkpoint roots per visit: a Suspect walk resumes from where it stopped on the next run, while a confirm walk that reaches the cap leaves the remaining roots Unconfirmed (conservative). A root published late at the previous pass's exact boundary counts toward this cap and the persisted cursor prevents a resumed visit from walking it again. Failure-set overflow suspects only roots that actually copied at least one scope.

A drifted row is not rebuilt by the check itself. It serves authoritatively until the family's automatic backfill republishes it from source: the scheduled stale-statistics repair does that on its own when enabled; otherwise an administrator runs POST api/admin/project-statistics/{id}/backfill for project screening, /{id}/{family}/backfill for the others, or a fleet pilot run. An automatic backfill first needs the family's backfill-observed bootstrap point: if that root no longer exists and a daily root already occupies the current identity, the backfill reports the bootstrap unconfirmed and publishes nothing; the family's administrative /rebuild then republishes the scope.

Cost

One visit costs at most ScopesPerProject authoritative scope calculations (each the family's ordinary source aggregation, as a daily build with the copy off would run), plus two small pinned reads per scope and a bounded walk of at most 1,024 checkpoint roots in pages of 32 when a verdict moves verification state. At the defaults one weekly run recalculates at most 16 × 256 = 4,096 scopes. For a project with N scopes that is N calculations a week instead of the 7N a week the uncopied daily build costs; the daily copy itself costs no aggregation for copied scopes. Per-project state is one small document in pmProjectStatisticsDriftCheck (its unverified-scope list is capped at 1,024 entries), and the rotation is one document in pmProjectStatisticsMaintenanceCursor (id drift-check).

Refusals and failures

unverified (nothing could be compared; not a failure, but no pass either), refused (a closed gate or no enabled scope), conflict (another run wrote the project's state first), inconclusive (writes kept landing) and error (an unexpected exception, logged with the project ID at error level) each affect one project and are retried by the next run. The consumer never faults its message and runs one run at a time per process.

Telemetry

  • syrf.project_statistics.drift_check.projects counts visits by outcome (passed, failed, unverified, continuing, inconclusive, refused, conflict, error).
  • syrf.project_statistics.drift_check.comparisons counts each scope's final comparison answer once (family, outcome = the parity outcomes, reader_reason). It is deliberately separate from parity.audits so the screening parity-audit soak evidence stays clean and a retried comparison is not counted twice.
  • syrf.project_statistics.parity.invalidations counts drifted rows made Stale, by family.
  • syrf.project_statistics.drift_check.verifications counts verification records moved, by verification_state (confirmed, suspect).

None of them carries a project identifier; the structured logs name the project for a failing check. See the runbook Telemetry section.

Enablement order

  1. Confirm the target projects are allowlisted and serving materialised statistics, and that the daily observation schedule (and, if wanted, CopyFreshMaterializedRows) is running for them.
  2. Set ProjectStatisticsDriftCheck:Enabled=true on the Project Management host, optionally with a cron and budgets. Size ProjectsPerRun so ceil(allowlisted / ProjectsPerRun) runs fit inside the cadence you want every project checked at.
  3. Watch drift_check.projects by outcome for the first runs. drift_check.comparisons with outcome=mismatch, together with parity.invalidations, is real drift: find the project in the logs, republish the Stale rows (see above), and investigate the write path that bypassed the clocks.

Rollback

Set ProjectStatisticsDriftCheck:Enabled=false. The consumer then acknowledges ticks without work, so no Quartz clean-up is needed. Nothing the check wrote blocks rolling back the binary: Stale rows rebuild through the ordinary backfill, an advanced source clock only makes the next daily build a real one, and verification states are plain values an older binary reads (it treats a missing record as Confirmed, and the Confirmed and Suspect values already exist). The drift-check state documents can be left in place or dropped.

Evidence

Real replica-set tests over the production services show that a first pass records its boundary and confirms nothing, and a second pass confirms exactly the snapshot bracketed by the two; that a failure after three Unconfirmed days marks those three Suspect, leaves the Confirmed one and every root, page and observation untouched, makes the served row Stale and advances only the source clock; that the drifted project then serves authoritatively, its next daily build is a real (calculated) one rather than an unchanged day, an administrative rebuild restores materialised serving, and the next check passes; that a disabled schedule writes nothing; that a one-scope budget spreads a two-scope check across two runs through the persisted cursor; that an unchanged-day project is still visited and caught; and that a real screening write landing between the two halves of a comparison is retried and is not a failure. A calculation reporting no snapshot boundary is never compared, so drifted counters it returns end the visit inconclusive without recording a failure; a scope first enumerated in the final run of a multi-run check is not taken as verified, so the snapshot that copied it stays Unconfirmed; a snapshot taken after a mode-epoch-only advance during a check is not confirmed by it; and a check that compares nothing is unverified and leaves the last passing boundary alone. Live cadence, cost and drift numbers remain to be measured on staging.

Scheduled stale-statistics repair

Many correctness paths leave a live row Stale on purpose: the drift check above, a question-definition rewrite (the fence completes without a rebuild), shared-allocation and other dependent-family invalidations, fence completion and lost publication races. A Stale row serves correct values through the slower authoritative fallback, but before this job nothing republished it except an administrative backfill, a forced rebuild or a manual fleet run. The Project Management host can now look for such families on a schedule and republish them through each family's ordinary, non-forced backfill (FEAT-024 C1). Nothing is enabled by this change, and no environment configures it.

Configuration

All keys are read by the Project Management host only (environment variables carry the SYRF__ prefix and use __ for :).

Key Default Meaning
ProjectStatisticsRepair:Enabled false Registers the schedule and lets the consumer run. Off, the consumer acknowledges a tick without work even if an earlier Quartz trigger still delivers.
ProjectStatisticsRepair:CronExpression 0 15 * * * ? Quartz cron, evaluated in UTC: hourly at quarter past. It must never fire twice within 15 minutes; a tighter cron refuses host start.
ProjectStatisticsRepair:ProjectsPerRun 16 Allowlisted projects inspected per run (1–256), in a persisted rotation.
ProjectStatisticsRepair:RepairsPerRun 8 Family backfills started per run across all projects (1–64). Families found after the budget is spent are reported deferred.
ProjectStatisticsRepair:ScopesPerProject 1024 Scope inspections (served-row reads) per inspected project per run (1–4096). It bounds detection only; a backfill always recalculates every declared scope of its family.
ProjectStatisticsRepair:RequireSettledSource true Settle check: backfill a family only when the project's source clock has not moved since the previous run observed it; otherwise report skipped with reason=unsettled. The first visit to a project only records the observation, so with the default a repair lands one run after a row goes Stale at the earliest.
ProjectStatistics:ProjectAllowlist empty The reviewed allowlist, the only source of projects the job visits.

With Enabled=true, the host refuses to start unless the cron is valid and fires no more often than every 15 minutes and every budget is in range. The job needs materialized writes (materializedProjectStatisticsWrites) and each family's own flag; a family whose flag is off is neither inspected nor reported. It does not use the fleet runner, so ProjectStatistics:Fleet:Enabled does not gate it (see "Why not the fleet runner" below).

Why hourly. A stale row is a performance problem, not a correctness one, so the job only needs to bound how long a page stays on the fallback, not race writers. An inspection that finds nothing costs indexed reads only. Each repair, though, is one full authoritative recalculation of one family for one project, and on a project that is still being written a family invalidated by every write would only get a short Fresh window from each attempt, so a much shorter cadence would mostly repeat expensive work. Quarter past keeps the job off the 00:05 daily observation tick; overlapping with the 02:30 maintenance or 03:30 drift-check ticks is harmless because each service serializes through its own leases and compare-and-swap writes.

What one run does

  1. It sorts the configured allowlist (minus any project the host's flag source does not also admit) and inspects at most ProjectsPerRun projects from a persisted rotation start. The start moves every run: to the first project that had a family deferred, otherwise one project on when the whole allowlist fits in one run, otherwise past the last project visited. So no project always spends the repair budget first, and a budget smaller than the allowlist reaches every project within ceil(allowlisted / ProjectsPerRun) runs.
  2. A project is skipped (with a typed reason, retried next run) when the fleet mode is not Enabled, materialized writes are off, it has no statistics control or is deleted, its own mode is not Enabled (including the Disabling drain quarantine), a definition-rewrite or inclusion-recalculation fence is raised, a staged import is active, the durable reviewer mode disagrees with the flag, or its declared versions differ from the fleet's. The project gate, family guards, authoritative scope enumeration, served rows and operation fences are read in one pinned source snapshot. If the host cannot open that snapshot, it reports snapshot_unavailable and repairs nothing. The snapshot is released before any backfill begins; each backfill then rechecks its own admission against current state.
  3. For each family with a registered backfill and its flag on, it reads the publication guard and then each declared scope through ProjectStatisticsServingGate, the reader's own predicate. A family whose guard is absent, Stale or Missing, or with any served row Stale, Missing or published under an older write epoch, needs repair. A fence, a staged operation or a live rebuild anywhere in the family skips it for this run. A row the reader refuses for another reason (for example an incompatible row) is reported not_repairable and left alone. A never-published family that declares no scope for the project (for example search population without searches) has nothing to publish and is current; it is never backfilled. If it previously published a scope but now declares none, the job reports skipped/not_repairable. A family with scopes but no guard at all has never been backfilled for that project: it is missing, and the same ordinary backfill bootstraps it. Turning a family flag on for an allowlisted project therefore establishes its baseline on the next run; keep the flag off to keep a family unbootstrapped. The starting family rotates per project across visits, so even ScopesPerProject=1 eventually inspects a later family when an earlier current family spends the whole scope budget.
  4. Before any backfill: a family with an active backoff (below) is reported backed_off; with the settle check on, a project whose source clock moved since the previous run is skipped/unsettled; and once RepairsPerRun is spent the family is deferred. None of them spends the budget.
  5. A family that needs repair is handed to its ordinary backfill with force: false, the same call the administrative backfill route and the fleet runner make, with owner scheduled-stale-repair. The backfill re-checks the allowlist and flags, records the family's backfill-observed bootstrap point first, takes its own per-scope rebuild leases, publishes through the control compare-and-swap and verifies the final bundle. A scope that is already current is skipped by the backfill itself.

The job never forces: it never adopts a changed configuration, never re-stamps a control and never publishes rows without their bootstrap point. If a family previously published scopes but now declares none, it reports skipped/not_repairable: ordinary backfill also enumerates only current source scopes and cannot republish those rows. A single published scope that vanishes while sibling source scopes remain is not yet classified by the detector; #3792 tracks bounded inspection of those non-Fresh published scopes. The reader's fallback and drift quarantine still apply. An operator must reconcile the source and then use the appropriate family backfill or forced rebuild; the schedule does not invent a zero row.

Outcomes and what an operator does

Each family's answer is one of these (also the outcome telemetry tag):

Outcome Meaning Operator action
current Every declared scope is served; nothing was called. None.
repaired The backfill republished the family; its scopes serve materialised again. None.
deferred Needs repair, but the run's RepairsPerRun budget was spent. None; the next run starts at this project. Raise the budget if it persists.
skipped A fence, staged operation, quarantine, live rebuild, closed gate, unsettled source, scope budget or a row the backfill cannot repair (reason says which). None; the next run looks again.
backed_off An earlier configuration_refused or bootstrap_identity_occupied answer still holds, because nothing that could change it has moved. The backfill is not called; the report repeats the original reason and failure_reason. As for the original answer.
refused A typed, retryable backfill refusal (failure_reason, e.g. publication_race_lost, statistics_rebuild_busy, family_fenced). None unless it persists across many runs.
configuration_refused The authoritative configuration or versions differ from the project's established ones (digest_mismatch, control_version_mismatch, fleet_version_mismatch). The rows stay on the fallback. Confirm the configuration change is intended, then run the family's forced rebuild (POST api/admin/project-statistics/{id}/rebuild, or /{id}/{family}/rebuild). The schedule never does this for you.
bootstrap_not_confirmed The backfill could not record the family's bootstrap point, so it published nothing. reason=bootstrap_identity_occupied is the known limitation below; any other reason is a transient checkpoint refusal retried next run. See below.
failed An unexpected exception, logged at error level with the project ID. Investigate the log if it repeats; other families and projects were not affected.

Backoff. The two answers an automatic retry cannot change are remembered per (project, family) in pmProjectStatisticsStaleRepair, with a key read right after the refusal: for configuration_refused the control's configuration digest, the family guard's version and the project's source revision; for bootstrap_identity_occupied the checkpoint identity (source revision, projection revision, fleet mode epoch). A fleet_version_mismatch is retried on the next run rather than persisted: deploying a compatible writer can resolve it without moving any stored project identity. While a persisted key is unchanged the family is backed_off and costs three small reads. Any write to the project, a forced rebuild or a configuration change moves the key, and the family is attempted again on the next (settled) run. A family that becomes current or repaired drops its entry; entries are capped at 32 per project.

The occupied-bootstrap limitation. An ordinary backfill must record exactly one backfill-observed point for the family before it publishes. If that root no longer exists (history maintenance retired it, or the family was seeded by rebuilds and never had one) and a daily observation root already occupies the project's current (source revision, projection revision, mode epoch) identity, the checkpoint admission answers "already published" and the backfill stops with nothing published. The job reports bootstrap_not_confirmed with reason=bootstrap_identity_occupied, logs a warning with the project ID, and does not force; later runs report it backed_off until the identity moves. Two ways out:

  • Wait. Any later write to the project moves its identity, and the first settled run after it records a new bootstrap point and republishes the family.
  • Repair now. An administrator runs the family's forced rebuild (POST api/admin/project-statistics/{id}/rebuild, or /{id}/{family}/rebuild). It republishes the Stale scopes and moves the projection revision; the next scheduled run then finds the family current.

Refusals and failures

Every refusal and exception is isolated to one family of one project and retried by the next run; one project's failure never stops the others. The consumer never faults its message, runs one run at a time per process and publishes no startup catch-up run (a missed tick is sent once by the misfire policy). Only shutdown cancellation propagates.

Cost

Inspection costs at most ScopesPerProject served-row selections (one or two indexed reads plus a fence lookup each) and one scope enumeration per enabled family, per inspected project, plus one small state document read and write. Repair costs at most RepairsPerRun ordinary backfills per run. A backfill is not bounded by ScopesPerProject: it enumerates and handles every declared scope of the family, and each automatic scope costs about two authoritative calculations (the already-current check, then the rebuild's own calculation), plus the bootstrap point's calculation the first time. A configuration_refused attempt pays the same scope pass (each scope then refuses in its compare-and-swap), and a bootstrap_identity_occupied attempt pays a scope enumeration and a walk of the project's history; the backoff makes each of those a one-off until something relevant moves. At the defaults one hourly run inspects at most 16 projects and runs at most 8 family backfills. The rotation is one document in pmProjectStatisticsMaintenanceCursor (id stale-repair); per-project state is one document per visited project in pmProjectStatisticsStaleRepair (the last observed source revision, last starting family, at most 32 backoff entries and a per-family last-inspected scope cursor).

A family that ordinary writes invalidate as a whole (for example the reviewer families after a screening decision without point-maintenance evidence, which mark the family Stale) is Stale again after every such write. The settle check (on by default) repairs it only after a run interval with no write to the project, so a busy project's budget is not spent on a Fresh window that closes minutes later; its readers stay on the fallback while it is busy and are repaired within about an hour of it going quiet.

The scope budget is a bound on served-row inspections, not on the calculator's authoritative scope enumeration. Each visited family remembers its last inspected canonical scope in the per-project state and resumes after it, wrapping at the end. A stale scope beyond the first run's inspection budget is therefore reached on a later run. A family-wide invalidation marks the guard and is found without reading any scope. Scope enumeration still materializes the authoritative family set on each visit; keep this default-off job on reviewed allowlisted projects until source-side keyset paging is available. Track source-side paging in #3790.

Why not the fleet runner

The fleet runner can already call the non-forced backfill (a run with repair: false), but it cannot serve a schedule: a run pauses at its first failed target and needs an administrator to resume it, so one refusing project would stall every other; at most 64 runs are retained and only an administrator can prune them, so an hourly job would exhaust the capacity within three days and then block operators' own runs; one global lease admits one target per 30 seconds and would contend with operator runs; and its targets are whole project families, so detection is needed either way. The job therefore calls the same backfill service directly, with the same leases, bounds and typed outcomes; fleet runs remain the operator's tool.

Telemetry

  • syrf.project_statistics.stale_repair.projects counts project visits by outcome (inspected or skipped) and, when skipped, reason (fleet_not_enabled, writes_disabled, project_unavailable, project_not_enabled, source_visibility_fence, staged_operation, durable_mode_disagreement, incompatible, unexpected).
  • backed_off families carry the original reason and failure_reason; skipped families include reason=unsettled for the settle check.
  • syrf.project_statistics.stale_repair.families counts family answers by family, outcome (the table above), reason (the trigger stale, missing or epoch_mismatch, or the skip reason, including bootstrap_identity_occupied) and, for a typed backfill refusal, failure_reason.
  • Each backfill the job starts is also counted by the existing rebuilds.requested / rebuilds.completed / rebuilds.failed instruments with operation=backfill.

None of them carries a project identifier; the structured logs name the project for refusals that need an operator. See the runbook Telemetry section.

Enablement order

  1. Confirm the target projects are in ProjectStatistics:ProjectAllowlist and serving materialised statistics, and that exactly the families you want maintained have their flags on (a flagged family with no baseline is bootstrapped by the first run).
  2. Enable the repair before, or together with, the drift check: the drift check and the daily copy only ever mark rows Stale, so without the repair a drifted row stays on the fallback until someone acts. Order on the Project Management host:
  3. the daily observation schedule (and, if wanted, ProjectStatisticsDailyObservations:CopyFreshMaterializedRows);
  4. ProjectStatisticsRepair:Enabled=true (optionally with cron and budgets);
  5. ProjectStatisticsDriftCheck:Enabled=true.
  6. Watch stale_repair.families by outcome for the first runs. repaired should follow every parity.invalidations within about two runs once the project is quiet; configuration_refused and bootstrap_not_confirmed need the operator actions above; persistent deferred means RepairsPerRun is too small.

For a staging configuration PR, the keys go in syrf/environments/staging/project-management/values.yaml under env: in cluster-gitops, for example:

  SYRF__ProjectStatisticsRepair__Enabled: "true"
  # Optional; the defaults are shown.
  SYRF__ProjectStatisticsRepair__CronExpression: "0 15 * * * ?"
  SYRF__ProjectStatisticsRepair__ProjectsPerRun: "16"
  SYRF__ProjectStatisticsRepair__RepairsPerRun: "8"
  SYRF__ProjectStatisticsRepair__ScopesPerProject: "1024"
  SYRF__ProjectStatisticsRepair__RequireSettledSource: "true"

Rollback

Set ProjectStatisticsRepair:Enabled=false. The consumer then acknowledges ticks without work, so no Quartz clean-up is needed. Nothing the job wrote blocks rolling back the binary: every publication it caused is an ordinary backfill publication (and bootstrap point) that an older binary already reads, and its rotation and state documents can be left in place or dropped.

Evidence

Real replica-set tests over the production services show that a drift-check-marked scope and a definition-rewrite-marked scope each become Fresh and are served materialised after one run (and the next run finds them current without another backfill); that a source-visibility fence and a bulk operation fence are skipped with nothing backfilled and the family is repaired once the fence is gone; that a foreign configuration is refused typed (digest_mismatch) with the control's digest unchanged and the rows still on the fallback, and then backed_off without another backfill until a write moves the key; that a disabled schedule writes nothing; that the settle check waits for a run interval without writes; that ProjectsPerRun=1 rotates through the allowlist and the run after a RepairsPerRun=1 deferral starts at the deferring project; that the start moves on by one when every project fits in one run; that a first project whose family keeps refusing does not stop a later project being repaired; that one family's exception does not stop another family's repair; that an occupied bootstrap identity is reported bootstrap_not_confirmed / bootstrap_identity_occupied without forcing, then backed_off, after which the documented forced rebuild restores serving and the backoff is dropped; that a flagged family with no scopes is current and never backfilled; that a flagged family with no baseline is bootstrapped as missing; that an older-write-epoch row is repaired; and that a live rebuild, a row refused for another reason, the scope budget, a durable reviewer-mode disagreement and a fleet version mismatch are all skipped without a backfill. Every recorded backfill call used force: false. Live cadence and cost numbers remain to be measured on staging.