Phase 6A bounded family fleet operations¶
Usable boundary and acceptance¶
This first operational slice lets an application administrator create an explicit bounded batch of project/family targets, process one target per request, inspect durable progress, pause and resume, and stop new fleet admissions globally. It composes the registered family backfill, rebuild and bootstrap coordinators. It introduces no scheduler, project discovery, wildcard admission or deployment activation. Families without a registered runtime backfill are rejected at creation. Registered family adapters include project, membership and reviewer screening; stage, membership-stage and reviewer annotation; question answers; search population; and domain reconciliation. All nine physical families use the same bounded dispatch protocol. Derived summaries do not need a physical backfill.
Acceptance is durable progress across process restart, one valid global admission lease, a server-owned minimum interval between admissions, denial of stale-owner progress, cancellation of a worker that loses renewal, and replay after publication without changing already-current screening rows or duplicating the bootstrap checkpoint. Source parity remains each family's existing authoritative calculation.
Configuration and authorization¶
Every action uses BatchAdminProjectsPolicy, the existing application administrator-only activity.
Project membership alone cannot operate the fleet. The runner is disabled by default.
Server configuration under ProjectStatistics:Fleet:
| Setting | Default | Enforced bound |
|---|---|---|
Enabled |
false |
Required for create and advance |
MinimumIntervalSeconds |
30 |
Clamped to 10–3600 seconds |
LeaseSeconds |
300 |
Clamped to 60–1800 seconds |
These are ordinary server IConfiguration settings, including the standard .NET environment-variable
form ProjectStatistics__Fleet__Enabled; no runtime UI override is registered. Configuration is read at
host startup. There are no client-supplied interval or lease settings. Creation and each dispatch also
require the reviewed project allowlist, materialized writes and target family flags. The existing
backfill coordinator rechecks its gates and owns durable epoch/fence publication admission.
Status, pause/resume and global stop controls remain available when the runner is disabled. Global stop blocks creation of new runs and new project admissions. An identical create retry can retrieve its existing run while stopped. Resuming a run does not clear the independent global stop.
Administrative protocol¶
All paths below are relative to /api/admin/project-statistics/fleet.
POST runswith{ "runId": "<stable-guid>", "targets": [{ "projectId": "<guid>", "family": "ProjectScreening" }], "repair": false }creates or retrieves a run (202). Reuse the same ID on an ambiguous network result. Identity binds the administrator, repair mode and canonical sorted project/family targets; changed payload under that ID returnsRunIdentityConflict(409).GET runs/<runId>returns the durable cursor, status, attempts and latest result per target. The run ID is the resume token; callers never submit a cursor to advance past work.POST runs/<runId>/advanceexecutes at most one project/family target and returns current progress (200). PollGET controlto inspectNextAdmissionAtUtc. A premature admission returnsFleetRateLimited; another valid lease returnsFleetBusy(409). Transaction contention returnsFleetContended; the server bounds a transaction attempt to ten seconds (FleetOperationTimedOut). Inspect status and retry after a refusal. There is no queued retry loop.POST runs/<runId>/pausestops further admissions for that run. An admitted project finishes.POST runs/<runId>/resumepermits retry of its current cursor, including a recorded failure.PUT stopwith JSONtruestops new fleet admissions. JSONfalseclears that stop.GET controlreports stop and current lease state.
A refused publication, absent project or incomplete bootstrap pauses the cursor with a typed bounded
result. Each result carries the backfill's lifecycleStatus and failureReason enum names, so an operator
can tell retryable contention from a fence, digest mismatch or capacity refusal; free-text detail is not
persisted. Resolve the cause and resume; no project is silently skipped. Unexpected faults or cancellation
record Failed or Interrupted when the owner still holds its lease and then propagate the failure.
A lease lost mid-dispatch returns FleetLeaseLost (409) rather than an unhandled cancellation.
Pause and resume declare the same 409 refusals (FleetContended, FleetControlChanged,
FleetOperationTimedOut) as the other mutating actions.
An expired lease can be reclaimed at the same cursor without an explicit failure record.
The API returns the durable run's canonical order, which may differ from submission order. There are
at most 32 distinct project/family targets per run and 64 retained run documents across the fleet. Results replace the
previous result for a retried cursor rather than growing a history list. The 65th new run returns
FleetRunCapacity; existing status, resume and stop controls continue to work. This bounded pilot does
not silently delete progress. Administrators can reclaim completed records using the explicit pruning
protocol below.
Completed-run pruning¶
POST runs/prune accepts { "runIds": ["<completed-run-id>"] }, with 1–32 distinct non-empty IDs.
Only application administrators can call it; the server records their authenticated identity.
The entire request fails with RunNotCompleted (409) if any existing run is unfinished or currently
claimed. Validation and deletion share the control transaction with creation and claims, so a refused
batch deletes nothing. Missing IDs are harmless on retry. No projection, checkpoint, receipt, source
record or paused run is deleted. Pruning remains available while fleet dispatch is disabled or stopped.
The control retains the latest 32 bounded pruning records, each containing a monotonic sequence,
server observation time, administrator, requested IDs, deleted run and target counts, and a count of
the deleted results per typed LifecycleStatus/FailureReason pair (results written before those
fields existed are counted under null names rather than dropped). Export these records before
they rotate if longer audit retention is required; this is a finite operational audit, not permanent
history. At most 1,024 requested IDs are retained across all audit records.
Archive required completed-run evidence before pruning. Deleted run IDs no longer provide status or create-request deduplication: callers must retire them and use new IDs for future runs. A delayed create retry using a pruned ID can create another run; normal backfill remains idempotent, while forced repair is already at least once. Pruning is explicit because it ends this operational evidence and retry window.
Single-family screening runs retain their original fleet.v1 request digest. Canonical ordering adds
family as a secondary key, so a project can request multiple distinct families without duplicate targets.
Each dispatch rechecks its own family flag and calls that family’s backfill; a disabled family pauses
the cursor instead of falling through to screening.
Concurrency and restart semantics¶
Admission, pause/stop, renewal and progress updates serialize through a single Mongo control document
in snapshot transactions with majority commit. The default _id indexes give unique run and control
identities. The control stores the next admission time durably, so process restart cannot reset rate
limits. Transactions also serialize the retained-run count and insert. Fleet-control reads tolerate
additive BSON fields, and control writes update only fields this version owns; older replicas therefore
preserve newer prune-audit fields instead of rejecting or erasing them. Only a prune rewrites the
sequence and bounded audit list; ordinary claims, renewals and completions leave them untouched, and
audit records tolerate unknown fields on read. Re-typing existing fields still
requires an explicit compatibility plan.
An advance receives a unique token and increasing lease generation. Renewal runs every third of the lease period. Failed renewal cancels the token passed to the backfill and no subsequent project is dispatched. Completion requires the same token/generation, an unexpired lease and the expected cursor. A replacement worker cannot have its progress overwritten by the late owner.
This bounds valid admissions, not the number of physically executing processes during a process pause or network partition. Cancellation is cooperative; an already-running atomic publication may finish. Existing per-scope leases, epochs, source fences and publication guards remain the authority for its correctness. Global stop lets the currently admitted publisher renew and finish safely, then blocks the next project. It is not an instruction to roll back committed source or projection data.
After a crash between publication and progress, restart repeats the same project. Ordinary backfill
recognizes already-current rows and its existing bootstrap point; no generation or history churn is
needed. Forced repair is deliberately at least once: repeating a completed forced publication can
create a further valid generation. Fleet progress is not an exactly-once source-operation receipt.
Evidence and required continuation¶
The replica-set regression exercises competing store instances, durable rate limits, expired-owner rejection, global stop, pause/resume, finite storage admission and replay through the real screening calculator/rebuild/bootstrap pipeline. Replay compares published counts to the production authoritative aggregation and checks unchanged publication generation and a single bootstrap checkpoint. Service regressions cover default-off storage isolation, invalid family/batch denial and cancellation on failed renewal. HTTP authorization uses the real application policy and deployed permission catalogue.
These remain required ordered work before full Phase 6A completion:
- Establish the remaining approved family backfills and their parity proof. Fleet dispatch now uses registered family services, so it cannot invent a missing calculator or grant an unregistered family admission. A mixed-family run retains separate durable progress for each target.
- Hard bounds for projection/history/receipt retention and cleanup work, including avoiding accumulation of every root after paged reads in existing cleanup. Completed fleet-run pruning is implemented separately with finite audit retention. This slice bounds its own progress documents; it does not prove bounded checkpoint/delta/receipt growth across the fleet.
- Scheduled bounded dispatch and deployment wiring only when operational activation is authorized; the explicit advance protocol already provides a usable administrator-driven pilot.
Phase 6B's actual seven-day staging soak, representative read/mutation volume, exact parity, fallback, no unresolved rebuild backlog and bounded storage-growth evidence remain live activation/evidence gates. Local tests and existing checkpoint retention regressions do not substitute for that soak. Credentials, allowlist approval and production activation are separate; this change enables none of them.
Combined programme integration¶
The fleet branch includes the operational family baselines, daily observations and guarded consumer stack through #3572. Registration tests resolve the fleet with all nine actual family adapters; the all-family dispatch test verifies independent routing and preserves the original screening-only run identity. Real Mongo tests cover multi-family publication, lease recovery, pause/resume, stop and rolling-reader field preservation. This integration does not enable the runner, any consumer flag or a deployed scheduler; sustained-load and operational acceptance remain separate.
Explicit bounded checkpoint history maintenance¶
POST /api/admin/project-statistics/{projectId}/history-maintenance runs one per-project maintenance
transaction. It requires the application administrator policy, the existing explicit project allowlist,
and server configuration ProjectStatistics:HistoryMaintenance:Enabled=true (default false).
The only timer is the separately default-off scheduled maintenance
job; there is no deployment change or implied permission to execute maintenance in staging or production.
The operation deliberately remains available with serving flags off so rollback does not require
re-enabling application reads to maintain retained data.
Send {} for the initial observation page. The response reports roots scanned/removed/deleted,
reference pages deleted, observations scanned/deleted, nextObservationId, markersReclaimed,
unchangedDaysDeleted and verificationsDeleted. Send
{"afterObservationId":"<returned nextObservationId>"} to continue the observation sweep; null ends
that sweep. Start a later sweep at null so observations referenced during an earlier sweep can be
re-evaluated after their final referring root is removed. Retention/root cleanup independently resumes
from durable root state on every request. Retrying after interruption is safe.
One invocation reads at most 1,025 root headers and refuses with HistoryScanCapacity, without changes,
if more than 1,024 roots exist. It marks at most 32 roots removed, deletes at most 32 reference pages
and 32 removed roots, and examines at most 32 observations. This explicit refusal avoids an unbounded
historical scan; projects already beyond that admission ceiling need a separately reviewed migration.
The surviving-root selection retains the existing daily/weekly/monthly policy and 256-root cap. Reader
grace starts when a root is removed, and cleanup waits the configured lifecycle grace interval.
Checkpoint admission is compare-and-set in the same snapshot transaction as cleanup. An occupied
build slot refuses maintenance (CheckpointBuildBusyOrUninitialized); a racing admission conflicts
rather than publishing a reference to an observation being deleted. An admission attempted inside an
open pass fails its own transaction and retries after the pass commits; one that committed after the
pass read the admission row aborts the pass, which returns 409 CheckpointAdmissionChanged without
changes. Any other transient transaction abort (for example a primary stepdown or network error) that
leaves the admission row unchanged returns 409 MaintenanceTransactionConflict, also without changes;
both mean "retry the request". Only these typed refusals are 409s: any other failure is an invariant
break and returns 500. Every remaining reference page,
including one belonging to another removed root awaiting cleanup, protects its observations. A bounded
global observation keyset scan prevents deleted creator roots or a first page of referenced observations
from losing orphan discovery. Cancellation before commit aborts the complete pass; the next request
restarts from committed state. The older unbounded internal retention loops (ApplyRetentionAsync,
RunCleanupAsync) had no production caller and were removed in
#3674; RunCleanupAsync could leave a retention-failed
root with some pages deleted that readers still treated as retained.
Checkpoint build marker reclamation (#3674). Checkpoint admission refuses a new build
(CheckpointCapacityExceeded) once a project holds more than 512 terminal (Published or Cleaned)
build markers. Without reclamation, every retention-removed root left a Cleaned marker behind, so a
project building one checkpoint a day stopped admitting after roughly 512 days and daily history
stopped. Each pass now also deletes at most 32 terminal markers whose build token has no root of any
retention state, inside the same admission-CAS snapshot transaction as the rest of the pass. A marker is
reclaimed only once it provably owns no reference page and no creator observation; observation collection
is keyed on the creator marker's Cleaning/Cleaned state, so the marker of a removed root whose
observation another root still references stays Cleaning until that observation is collected. A
retained root's Published marker is never touched, and Building, Abandoned and Cleaning markers
are never candidates (an unpublished build also holds the slot, which refuses the whole pass). Steady
state is therefore about one marker per retained root (at most the 256-root cap) plus markers awaiting
their final cleanup. A racing admission conflicts with the pass exactly as above: an admission inside an
open pass cannot commit and retries afterwards; one committed after the pass read the admission row
aborts the pass, which reclaims nothing and returns 409 CheckpointAdmissionChanged. Repeat the call.
Real Mongo tests run 552 daily builds with weekly maintenance and keep admitting (the same test is refused
at day 513 without reclamation), keep retained and still-referenced markers, reclaim a cleaned abandoned
build, and cover both race orderings.
Unchanged-day reclamation (#3636). The daily runner records an inactive project's day in
pmProjectStatisticsUnchangedDay instead of building a checkpoint (see the
historical charts rollout). Each pass also deletes, inside the same
admission-CAS transaction and oldest first, at most 32 of those records that are older than the daily
retention window (90 days), that reference a root which is no longer retained (the reader already treats
them as gaps), or that are the oldest beyond 128 per project. It reads at most 160 records to decide.
The scheduled maintenance run counts these deletions as progress and reports them as the
unchanged_days reclaimed item. Real Mongo tests prove 200 records converge to the retained window in
passes of at most 32 and that the cap applies inside the window.
Verification-record reclamation (#3636 part 2). When copied daily snapshots are enabled
(ProjectStatisticsDailyObservations:CopyFreshMaterializedRows), each daily build writes one
pmProjectStatisticsCheckpointVerification record before publishing. The same pass deletes, at most
32 per pass, records whose root no longer exists (deleted by retention, or a build that never
published). The admission slot is free during the pass, so no build is between writing its record
and publishing. A live root's record is never deleted. The run reports these as the
checkpoint_verifications reclaimed item.
Real Mongo tests cover shared-observation survival, reader grace, final-reference collection, active
build refusal, both admission race orderings, a transient abort without an admission change,
interrupted-transaction rollback/retry, observation cursors, 32-root mutation bounds, the 32-page bound
with a removed root whose pages span two passes (the root stays Compacted until its last page goes), capacity refusal before changes and disabled/unlisted refusal without
writes. A history read through the production bundle reader before, during and after a pass serves
every retained root completely (including one sharing the removed root's observation) and reports the
removed root as Compacted until its cleanup deletes it. HTTP tests cover anonymous/ordinary-user denial and default-off admin
refusal. This adds executable checkpoint maintenance; delta/receipt storage-growth proof and the live
soak acceptance gates remain outstanding.
Enabling copied daily snapshots¶
ProjectStatisticsDailyObservations:CopyFreshMaterializedRows (environment variable
ProjectStatisticsDailyObservations__CopyFreshMaterializedRows, default false) makes the daily
runner copy fresh current rows instead of recalculating them (see the
historical charts rollout). It is a
per-environment deployment change on the Project Management host. Enable it only in this order:
- Every replica runs the build first. Confirm that every API replica and every Project Management replica runs a build containing #3725 and #3727, and that no older replica set is still serving (the rolling update has completed). An older API replica ignores verification records and shows copied points as ordinary, unmarked history; it also saves a stage's session-count target without invalidating Reviewer annotation, so a copy could pair old counters with the new stage definition.
- Daily observations already run.
ProjectStatisticsDailyObservations:Enabled=truehas run cleanly for the target projects with the option off, and history maintenance is enabled so verification records of deleted roots are reclaimed. - Pilot size is measured. The copy reads every scope's row and fences inside one snapshot transaction (read batching by family is a follow-up). Pilot on projects whose membership × stage scope count is known, and watch for copy faults: a fault releases the build at once and the next tick retries, but a project that exceeds the snapshot lifetime fails the same way each time.
- Set the option on the Project Management host and watch
daily_observations.scopesbyprovenance;copiedshould dominate for active projects andcalculatedshould cover search population plus any non-servable scope.
Rollback. Set the option back to false: new daily builds calculate every scope again. Existing
Unconfirmed roots keep their state until the drift check decides them. Before rolling back to a binary
older than #3725/#3727, set the option to false first.
The build marker and root of a copied snapshot carry IsAuthoritativeBuild = true, because they were
published through the authoritative protocol. That flag is not provenance: only the root's
pmProjectStatisticsCheckpointVerification record says which scopes were copied. Do not read the
flag as evidence that a snapshot was calculated.
Explicit bounded delta maintenance¶
POST /api/admin/project-statistics/{projectId}/delta-maintenance performs one ledger-compaction pass.
It requires application administrator authorization, the project allowlist, and independently
default-off server setting ProjectStatistics:DeltaMaintenance:Enabled=true. No live maintenance is
enabled by this change; the optional scheduled job is default-off. The older unbounded CompactAsync loop had no production caller
and was removed in #3674; its read-only watermark
calculation remains.
The pass uses a single snapshot transaction. It serializes with checkpoint admission and compare-and-sets the project control with the floor update. Existing family guards are also compare-and-set to serialize with source-fence/candidate admission. Any active checkpoint build, live rebuild lease or nonterminal publication coordinator or family fence/candidate guard refuses the pass; a truncated list of owners is never interpreted as a complete pin set. Retained roots, outstanding reconciliation, pending notification slots and the maximum redelivery window still restrict eligibility. Root inspection has the same 1,024-header admission ceiling as history maintenance. Seal reads cap at 65 to detect the 64-seal ceiling without silently omitting older evidence.
At most 33 exact rows are inspected and 32 are deleted. The extra row proves that the last selected
projection revision is complete. Adjacent seal time bounds use the minimum/maximum observed times,
so clock skew cannot move an audit boundary backwards. A sibling group exceeding the batch, a gap in revisions, or any young
sibling stops selection; nothing skips over such a boundary. The result reports Compacted or
NoEligibleCompleteRevision, the count deleted, the durable replay floor and the safe watermark. Repeat
the call to continue from the committed floor. A no-op is a safety result, not evidence that the ledger is
empty or storage growth is bounded. Oversized sibling groups require a separately reviewed larger-work
protocol; this endpoint never splits them to force progress.
The seal insert or adjacent-seal merge, exact-row deletion, and replay/identity-floor advancement commit
atomically. Interrupted work leaves all of them unchanged. The identity floor moves to the newest
compacted observation but always stays strictly below the oldest observation of any exact row left above
the new revision floor, so a clock-skewed surviving row is never classified too old. That proof reads at
most 129 remaining rows in replay order; when more remain, the pass advances only the revision floor and
leaves the identity floor where it was. A source write, checkpoint admission, guard
transition or rebuild that commits a row the pass read after its snapshot aborts the whole pass with a
Mongo write conflict; the service aborts and refuses with the typed ConcurrentStatisticsWrite reason
(HTTP 409), and the caller simply repeats the call. A commit whose outcome stays unknown after the bounded
commit retries refuses with CommitOutcomeUnknown (HTTP 409): the seal, deletions and floor moves either all
committed or none did, and a repeat pass re-reads the floor, so repeating is safe. Every refusal and every
write conflict is therefore a typed 409; only a cancelled request (the caller went away) and an unexpected
server fault are not. Adjacent batches normally retain one merged
seal; a full nonadjacent seal set refuses additional compaction. Cancellation is checked between bounded
steps. Source-operation receipts and their idempotency floor are never deleted or advanced here:
receipt reservation, pin/reachability and redelivery reclamation need their own operational contract.
Real Mongo regressions cover bounded continuation, shared-revision page edges, young siblings, retained
checkpoint pins, active-owner refusal, rollback after an injected deletion interruption, the ledger-gap
stop, the full nonadjacent seal set, the guard/root/notification scan ceilings and an unresolved commit outcome. Against real
coordinator writes they also prove the visible projection and control clocks are identical before and after
compaction, a rebuild from the authoritative source agrees with the compacted projection, a redelivered
operation whose delta was compacted is still refused by its receipt (source applied once, counters
unchanged), and a source write, checkpoint admission or rebuild racing an open pass either loses its own
conflict or aborts the pass typed, with no delta committed during or after the pass ever deleted by it.
Redelivery idempotency rests on source-operation receipts, which this pass never touches; production callers
test only for a receipt. A duplicate whose delta was compacted reports its replayed routing as
SourceOnlyFallback with projection revision 0, which no caller consumes today. HTTP tests
prove anonymous/non-administrator denial, the default-off administrator result and production container
resolution. Live fallback/rollback, storage-growth and sustained-load acceptance remain unproven until
captured in the approved environment; local compaction tests do not replace that evidence.
Explicit bounded source receipt maintenance¶
POST /api/admin/project-statistics/{projectId}/receipt-maintenance requires application administrator
access, project allowlisting and independently default-off server setting
ProjectStatistics:ReceiptMaintenance:Enabled=true. This deployment-only setting also gates the two
new receipt BSON identity fields: while disabled, receipts physically omit SourceAggregateId and
HasTrustedSourceIdentity, so older strict readers can consume newly written receipts during a
mixed-version rollout. Before enabling it, deploy the retirement-aware screening and annotation
writers to every API and project-management mutation host and initialize the new indexes. Set it
coherently on both hosts only after all old strict readers have exited. Keep it disabled during
mixed-version deployment. No live cleanup is performed by this change.
Disabling the maintenance flag safely stops further reclamation and new identity-field emission, but does not make a binary rollback to old strict readers safe while enabled-mode receipts remain in the collection, or to retirement-unaware source writers safe after any receipt has been deleted. Every mutation host must retain the retirement-guard protocol thereafter. An older writer binary requires coordinated restoration of the database/receipt backup and reconciliation before resuming writes; ordinary application rollback must not discard the durable retry protection while keeping the compacted database.
The minimum usable slice reclaims only enabled-mode point receipts whose server factory explicitly bound a
Study identity in stats.source.screening or stats.source.annotation. Legacy receipts without that
provenance, child receipts, pinned receipts and other namespaces remain retained. No operation-ID
parsing or fabricated creation age establishes eligibility. The old internal unbounded reclamation
helper (AdvanceIdempotencyFloorAsync) had no production caller and was removed in
#3674.
Each request reads at most 32 receipts using an explicitly hinted index on project, pin state, trusted identity, creation time and ID. It checks observed age within that page rather than scanning an arbitrary prefix to find 32 eligible rows. Both creation and observation must satisfy the larger audit/redelivery window. The time floor uses millisecond precision and deletion requires a strictly older stored creation time, preserving envelopes whose sub-millisecond fraction was lost in BSON. The earliest pinned creation time caps the floor, independently of commit observation order. A young observed receipt can temporarily occupy a page; the response counts scanned and deleted separately. Repeat requests continue after committed deletions, including when the time floor has not advanced.
Deleting a receipt also advances a monotonic source-revision retirement row in
pmProjectStatisticsReceiptRetirement, keyed uniquely by project, Study and the fixed source namespace.
There are at most two constant-size rows per Study and a hard 100,000-row project ceiling. Existing rows
may still advance at the ceiling; new keys refuse the whole transaction before any deletion. The guard
uses explicit $max updates, preserving unknown fields. Rows are durable retirement evidence and are
not themselves pruned. Active publication coordinators and family fences/candidates refuse cleanup.
Receipt deletion, retirement watermark updates and the project control CAS share one snapshot
transaction; cancellation rolls the complete pass back. Source writers CAS that same control, preventing
a writer which read an absent retirement row from committing after a concurrent cleanup.
Receipt-first replay preserves committed retries, including legacy receipts. With no matching receipt,
a trusted envelope at or below its source retirement revision fails TerminalTooOld, even when a new
server timestamp was minted for the old identity. Once a project has retirement evidence, a missing
receipt in either supported namespace also requires explicit source metadata. A legitimately new source
revision above the watermark remains admissible. Existing reviewer retry loops only handle duplicate
or Mongo concurrency outcomes; they do not recapture and reapply TerminalTooOld intent. That rejection
requires refresh/reconciliation rather than silently repeating a decision.
The guard adds indexed retirement reads to receipt-miss preflight and transactional admission. Command budget tests measure 23 Mongo commands per uncontended screening save (previously 21), including exactly two retirement lookups; they do not establish the programme's wall-time performance target. Real Mongo tests cover reminted identity rejection, normal new writes, retained committed retries, legacy/pin protection, bounded continuation, source snapshot conflicts and interruption rollback. Live storage-growth, fallback/rollback and elapsed soak acceptance remain unproven.
A source write that commits, or holds an uncommitted control write, after the pass took its snapshot
aborts the pass with a Mongo write conflict. The service aborts and returns the typed
ConcurrentStatisticsWrite refusal (409), never a raw 500; retirement rows, deletions and the floor roll
back together, the source write commits exactly once, and the administrator simply repeats the call.
Real replica-set tests drive the real transaction coordinator after a pass: the original envelope is
refused by the floor, the same identity re-minted above the floor and any never-seen identity at or below
the Study's retired revision are refused by retirement, an envelope without source metadata is refused,
the source and counters never move twice, and new revisions of the same or another Study still commit.
A missing IX_StatsReceipt_TrustedMaintenance index refuses with the typed ReceiptMaintenanceIndexMissing
(409) rather than a server error; run index initialisation, then repeat. A pinned receipt already stranded
below an existing floor refuses with PinnedReceiptBelowExistingFloor and writes nothing. Every 409
reports only the refusal's reason name.
The 100,000-row ceiling is exercised with real rows: a new Study key fails closed with
ReceiptRetirementCapacity and deletes nothing, while an existing key still advances.
Child/inverse and other source namespaces need a separate retention protocol that proves their durable retry authority before deletion. They are explicitly retained by this MVP; do not turn their cleanup on by widening the namespace predicate or deriving identities from operation-ID text.
Scheduled bounded maintenance¶
The Project Management host can call the three bounded services above on a Quartz schedule, so an allowlisted project that builds a checkpoint every day keeps its roots, build markers and delta ledger bounded without an administrator repeating the POSTs. The job adds no storage protocol: each pass is one ordinary call to the history, delta or receipt service, with the same gates, transaction and typed refusals as the administrative endpoint. Nothing is enabled by this change, and no environment configures it.
Configuration¶
All keys are read by the Project Management host only (environment variables carry the SYRF__ prefix
and use __ for :).
| Key | Default | Meaning |
|---|---|---|
ProjectStatisticsMaintenance:Enabled |
false |
Registers the schedule and lets the consumer run. Off, the consumer acknowledges a tick without work even if an earlier Quartz trigger still delivers. |
ProjectStatisticsMaintenance:CronExpression |
0 30 2 * * ? |
Quartz cron, evaluated in UTC (daily 02:30, clear of the 00:05 daily observation tick). A change takes effect on the next start: the same schedule identity is re-registered with the new trigger. |
ProjectStatisticsMaintenance:ProjectsPerRun |
16 |
Allowlisted projects visited per run (1–256). |
ProjectStatisticsMaintenance:PassesPerProject |
4 |
Passes of each service per visited project per run (1–32); a pass that makes no progress ends that service's passes early. |
ProjectStatisticsMaintenance:ReceiptMaintenanceEnabled |
false |
Separate opt-in for receipt reclamation (see rollback below). |
ProjectStatistics:HistoryMaintenance:Enabled |
false |
The history service's own gate, which the schedule does not bypass. |
ProjectStatistics:DeltaMaintenance:Enabled |
false |
The delta service's own gate. |
ProjectStatistics:ReceiptMaintenance:Enabled |
false |
The receipt service's own gate. Receipt reclamation runs only when this and ReceiptMaintenanceEnabled are true. |
ProjectStatistics:ProjectAllowlist |
empty | The reviewed allowlist, the only source of projects the job visits. |
With Enabled=true, the host refuses to start unless the cron is a valid Quartz expression and both
budgets are in range. A service whose own gate is off is skipped and counted as skipped with its
*_maintenance_disabled reason; it is not called.
Bounds and continuation¶
A run never enumerates projects: it sorts the configured allowlist (minus any project the host's flag
source does not also admit) and visits at most ProjectsPerRun of them, starting after the last project
the previous run visited. A budget smaller than the allowlist therefore reaches every project within
ceil(allowlisted / ProjectsPerRun) runs. History maintenance's observation cursor is chained within a
run and persisted between runs, so a long observation sweep resumes instead of restarting; a finished
sweep clears it. Both hints live in one singleton document, pmProjectStatisticsMaintenanceCursor,
bounded by the allowlist size. They are only resume hints: each service re-validates everything it
touches, so a lost cursor or two overlapping runs cost re-scanning, never correctness. One run executes at
a time per process (the consumer's concurrency limit is 1), but two Project Management replicas can
still run concurrently, for example a misfire-sent tick and the next regular tick landing on different
pods. That is safe: every service serializes through its own compare-and-swap transaction, so the loser
refuses typed (refused, contention) and the cursor is only a hint. If such contention shows up in the
refused counts, a lease on the cursor document is the follow-up. Unlike the daily observation schedule,
startup publishes no catch-up run; a missed tick is sent once by the misfire policy.
Per project and service, the per-pass service bounds above still apply (32 roots, pages, observations or markers per history pass; 32 revisions per delta pass; 32 receipts per receipt pass). At the default budgets one run does at most 16 × 3 × 4 = 192 bounded transactions.
Refusals and failures¶
A refusal writes nothing and ends that service's passes for that project only; other services and projects continue. Typed refusals fall into two groups, and they need different responses:
- Contention: clears by itself.
CheckpointBuildBusyOrUninitialized(a daily build holds the slot),CheckpointAdmissionChanged,ConcurrentStatisticsWrite, the other*Changedreasons,PublicationGuardBusy,RebuildActive,PublicationOperationActive,MaintenanceTransactionConflict(history) andCommitOutcomeUnknown(delta). Logged at warning level with the project ID, counted with outcomerefused, and normally absent on the next run. - Capacity: recurs on every run until an operator acts.
HistoryScanCapacity(more than 1,024 checkpoint roots, refused before anything is retired),PublicationGuardCapacity,NotificationScanCapacity,SealCapacity,ReceiptRetirementCapacityandReceiptMaintenanceIndexMissing(the receipt index must be created). The schedule cannot recover such a project. These are logged at error level and counted with the distinct outcomecapacity_refusedso they can be alerted on (outcome="capacity_refused", grouped byreason). Each needs the separately reviewed migration or protocol named in its service section above.
An unexpected exception is logged at error level and counted with outcome failed, reason
unexpected; it does not stop the other services or projects. The consumer never faults its message,
so a failure cannot poison or dead-letter the queue; only shutdown cancellation propagates.
Telemetry¶
syrf.project_statistics.maintenance.passes counts passes by operation, outcome (completed,
refused, capacity_refused, failed, skipped) and reason. A skipped pass names the switch that
was off: history_maintenance_disabled, delta_maintenance_disabled or receipt_maintenance_disabled
for a service's own gate, and scheduled_receipt_maintenance_disabled when the schedule's
ReceiptMaintenanceEnabled is off, so "switched off" is distinguishable from "never reached"; syrf.project_statistics.maintenance.reclaimed counts
items by operation and item (roots retired or deleted, reference pages, observations, build markers,
delta revisions, receipts). Neither carries a project identifier; use the structured logs for that. See
the runbook Telemetry section.
Enablement order¶
- Confirm the target projects are in
ProjectStatistics:ProjectAllowlistand that the explicit administrative passes behave as expected for them. - Enable the services' own gates on the Project Management host:
ProjectStatistics:HistoryMaintenance:Enabled=trueandProjectStatistics:DeltaMaintenance:Enabled=true. LeaveProjectStatistics:ReceiptMaintenance:Enabledfalse. - Set
ProjectStatisticsMaintenance:Enabled=true, optionally with a cron and budgets. Watchmaintenance.passesby outcome and reason for at least one run.refused(contention) passes should clear on the next run; anycapacity_refusedpass needs an operator before it will. - Only after every source mutation host runs retirement-aware writers (see receipt maintenance above),
and only if receipt reclamation is wanted, set
ProjectStatistics:ReceiptMaintenance:Enabled=trueandProjectStatisticsMaintenance:ReceiptMaintenanceEnabled=true.
Rollback¶
Set ProjectStatisticsMaintenance:Enabled=false (or turn off any single service's gate). The consumer
then acknowledges ticks without work, so while this binary is deployed no Quartz clean-up is needed; the
stale trigger is harmless.
A binary rollback to a Project Management build without this consumer is different: the previously
registered Quartz trigger keeps publishing IRunProjectStatisticsMaintenanceCommand, and the broker queue
bound to it collects one unconsumed message per tick. After such a rollback, unschedule the recurring
ProjectStatisticsMaintenanceSchedule trigger in the Quartz store and delete the maintenance consumer's
queue (or disable the job and let one tick drain before rolling back). The daily observation schedule has
the same property.
History and delta maintenance leave nothing that blocks rolling back the binary. Receipt maintenance
does: once one receipt has been reclaimed, stopping the job is safe but rolling back to a
retirement-unaware writer binary is not (see receipt maintenance).
That is why it has its own default-off switch.
Evidence¶
Real replica-set tests over the production services show that a disabled schedule writes no statistics row; only allowlisted projects are visited, within the per-run budget, and the rotation resumes on the next run; a busy admission slot refuses that project typed while another project proceeds, and the next run retries and succeeds; receipt maintenance never runs with its switch off; a service whose own gate is off is skipped; and 420 simulated days of one daily checkpoint followed by one scheduled run keep every project root (retained or awaiting deletion) under the 256-root cap and the 1,024-header scan ceiling, terminal markers under the admission ceiling, and every build admitted. Live soak evidence remains outstanding.
Periodic drift check¶
The Project Management host can recalculate every allowlisted project from source on a slow schedule and
compare it with the statistics the reader serves (#3636 part 3). It is what verifies the daily snapshots
that copy fresh materialised rows (ProjectStatisticsDailyObservations:CopyFreshMaterializedRows), and
it is the only check that sees drift the activity clocks cannot: a write path that changed data without
advancing them, or a period with statistics writes switched off. Nothing is enabled by this change,
and no environment configures it. What a pass and a failure do to history is described in the
historical charts rollout.
Configuration¶
All keys are read by the Project Management host only (environment variables carry the SYRF__ prefix
and use __ for :).
| Key | Default | Meaning |
|---|---|---|
ProjectStatisticsDriftCheck:Enabled |
false |
Registers the schedule and lets the consumer run. Off, the consumer acknowledges a tick without work even if an earlier Quartz trigger still delivers. |
ProjectStatisticsDriftCheck:CronExpression |
0 30 3 ? * SUN |
Quartz cron, evaluated in UTC: weekly, Sunday 03:30, clear of the 00:05 daily observation and 02:30 maintenance ticks. It must fire less often than daily; a cron with any gap of 24 hours or less between fire times refuses host start. |
ProjectStatisticsDriftCheck:ProjectsPerRun |
16 |
Allowlisted projects visited per run (1–256), shared by persisted retry and normal rotations. Pending projects receive up to half the budget; a one-project budget alternates retry and normal visits. |
ProjectStatisticsDriftCheck:ScopesPerProject |
256 |
Scopes recalculated per visited project per run (1–4096). A larger project continues its check on the next run from a persisted cursor. Size this to include all copied scopes in one visit if their snapshots should become Confirmed. |
ProjectStatistics:ProjectAllowlist |
empty | The reviewed allowlist, the only source of projects the job visits. |
With Enabled=true, the host refuses to start unless the cron is valid and less often than daily and
both budgets are in range. The check compares only what the reader would serve, so it needs the same
serving state as the daily copy: the fleet and project statistics modes enabled, the serving flag and the
per-family serving flags on, and the project allowlisted by the host's flag source. A project whose gates
are closed is counted refused and retried next run; nothing is written for it.
Continuing, inconclusive, conflicting, refused and errored projects remain in a durable retry set;
completed visits leave it. The normal rotation still advances in each run, so repeated refusals or a
large check cannot indefinitely block other allowlisted projects.
What one visit does¶
- Enumerates the enabled families' scopes (the same calculators the daily build uses) in (family, scope) order and resumes after the cursor of a check already in progress.
- For each scope, reads the served row and the project identity in one pinned snapshot, then
calculates the scope from source. The two are compared only when the calculation captured the same
source and projection revisions, the fleet mode epoch has not moved, and the calculation pinned its
own snapshot (a calculation that fell back to unpinned reads could have read its revisions and its
counters at different times). Otherwise the comparison is retried once and, if it still cannot be
made, the visit ends
inconclusiveand the next run retries that scope. A concurrent ordinary write is therefore never reported as drift. A null calculation carries no identity of its own: a family that can prove absence in a pinned snapshot at the served identity must also recheck that absence in the Stale transaction. Without that proof the visit stays inconclusive. - On a mismatch, one transaction re-checks the identity and the selected row, marks the row Stale
(
LastFailureReason = DriftDetected), advances the project's source clock, coalesces the family's notification slot and records the failure on the check's state. If the identity or the row moved, nothing is recorded and the scope is compared again. Reads fall back to the authoritative path immediately. - When every scope has been visited it reaches a verdict:
- Passed (at least one scope compared, none drifted): confirms bracketed Unconfirmed snapshots whose every copied scope the final visit found in parity, and records the new last passing boundary. A confirmation walk resumes across runs at its root budget and rechecks scopes on each visit; an admitted checkpoint inside the bracket must finish before the boundary advances and invalidates the earlier parity evidence. If a root is first requested at the pass identity after the pass committed, the next check includes that exact-boundary root only when it was absent at the previous pass; its copied scopes still have to verify. Scopes it could not compare do not stop the pass but are not verified.
- Unverified (no scope could be compared, for example every family fenced or not served): confirms nothing and leaves the last passing boundary where it was.
- Failed: marks Unconfirmed snapshots since the last pass Suspect only if they copied a drifted scope at or before that scope's own pre-Stale checkpoint (resumable if interrupted; nothing is confirmed until it finishes). A snapshot after one scope's repair can remain safe while copying other scopes, but is still suspect if another copied scope is later found drifted. The failure stays pending while an admitted build captured within the failure boundary can still publish; a changed admission generation restarts the bounded walk, and its final state claim fences checkpoint publication.
The cursor and check tally persist across visits, but parity evidence does not. At the start of a resumed visit, the check discards earlier verified scopes: a source mutation invisible to the activity clocks could have changed one between runs. A multi-visit check can still pass based on the scopes it compared, but a snapshot that copied scopes from different visits remains Unconfirmed. The next scheduled check visits those scopes again and can detect any drift. This keeps the per-visit calculation budget intact and avoids confirming history from stale evidence.
A passing worker first claims the versioned project check state as confirmation-pending. Only after that compare-and-set wins may it begin confirming verification records. If another replica has already recorded a failure, the would-be pass reports a conflict without confirming anything. If confirmation is interrupted, the pending state retains its boundary and bounded root cursor; the next run rechecks scope parity before it advances the confirmation walk. This also applies after the walk waits for an in-boundary admitted build, whose publication may have copied a source change invisible to the activity clocks.
Bracketing is strong evidence, not proof per snapshot. If a copied row drifted and was republished from source (by a backfill or rebuild) before the next passing check, that check passes and confirms the snapshot that copied the drifted value. A verdict walks at most 1,024 checkpoint roots per visit: a Suspect walk resumes from where it stopped on the next run, while a confirm walk that reaches the cap leaves the remaining roots Unconfirmed (conservative). A root published late at the previous pass's exact boundary counts toward this cap and the persisted cursor prevents a resumed visit from walking it again. Failure-set overflow suspects only roots that actually copied at least one scope.
A drifted row is not rebuilt by the check itself. It serves authoritatively until the family's
automatic backfill republishes it from source: the scheduled stale-statistics repair
does that on its own when enabled; otherwise an administrator runs
POST api/admin/project-statistics/{id}/backfill for project screening, /{id}/{family}/backfill for
the others, or a fleet pilot run. An automatic backfill first needs the family's backfill-observed bootstrap point: if that root
no longer exists and a daily root already occupies the current identity, the backfill reports the
bootstrap unconfirmed and publishes nothing; the family's administrative /rebuild then republishes the
scope.
Cost¶
One visit costs at most ScopesPerProject authoritative scope calculations (each the family's
ordinary source aggregation, as a daily build with the copy off would run), plus two small pinned reads
per scope and a bounded walk of at most 1,024 checkpoint roots in pages of 32 when a verdict moves
verification state. At the defaults one weekly run recalculates at most 16 × 256 = 4,096 scopes. For a
project with N scopes that is N calculations a week instead of the 7N a week the uncopied daily build
costs; the daily copy itself costs no aggregation for copied scopes. Per-project state is one small
document in pmProjectStatisticsDriftCheck (its unverified-scope list is capped at 1,024 entries), and the
rotation is one document in pmProjectStatisticsMaintenanceCursor (id drift-check).
Refusals and failures¶
unverified (nothing could be compared; not a failure, but no pass either), refused (a closed gate or
no enabled scope), conflict (another run wrote the project's state first),
inconclusive (writes kept landing) and error (an unexpected exception, logged with the project ID at
error level) each affect one project and are retried by the next run. The consumer never faults its
message and runs one run at a time per process.
Telemetry¶
syrf.project_statistics.drift_check.projectscounts visits byoutcome(passed,failed,unverified,continuing,inconclusive,refused,conflict,error).syrf.project_statistics.drift_check.comparisonscounts each scope's final comparison answer once (family,outcome= the parity outcomes,reader_reason). It is deliberately separate fromparity.auditsso the screening parity-audit soak evidence stays clean and a retried comparison is not counted twice.syrf.project_statistics.parity.invalidationscounts drifted rows made Stale, byfamily.syrf.project_statistics.drift_check.verificationscounts verification records moved, byverification_state(confirmed,suspect).
None of them carries a project identifier; the structured logs name the project for a failing check. See the runbook Telemetry section.
Enablement order¶
- Confirm the target projects are allowlisted and serving materialised statistics, and that the daily
observation schedule (and, if wanted,
CopyFreshMaterializedRows) is running for them. - Set
ProjectStatisticsDriftCheck:Enabled=trueon the Project Management host, optionally with a cron and budgets. SizeProjectsPerRunsoceil(allowlisted / ProjectsPerRun)runs fit inside the cadence you want every project checked at. - Watch
drift_check.projectsby outcome for the first runs.drift_check.comparisonswithoutcome=mismatch, together withparity.invalidations, is real drift: find the project in the logs, republish the Stale rows (see above), and investigate the write path that bypassed the clocks.
Rollback¶
Set ProjectStatisticsDriftCheck:Enabled=false. The consumer then acknowledges ticks without work, so
no Quartz clean-up is needed. Nothing the check wrote blocks rolling back the binary: Stale rows rebuild
through the ordinary backfill, an advanced source clock only makes the next daily build a real one, and
verification states are plain values an older binary reads (it treats a missing record as Confirmed,
and the Confirmed and Suspect values already exist). The drift-check state documents can be left in
place or dropped.
Evidence¶
Real replica-set tests over the production services show that a first pass records its boundary and
confirms nothing, and a second pass confirms exactly the snapshot bracketed by the two; that a failure
after three Unconfirmed days marks those three Suspect, leaves the Confirmed one and every root, page and
observation untouched, makes the served row Stale and advances only the source clock; that the drifted
project then serves authoritatively, its next daily build is a real (calculated) one rather than an
unchanged day, an administrative rebuild restores materialised serving, and the next check passes; that
a disabled schedule writes nothing; that a one-scope budget spreads a two-scope check across two runs
through the persisted cursor; that an unchanged-day project is still visited and caught; and that a real
screening write landing between the two halves of a comparison is retried and is not a failure. A
calculation reporting no snapshot boundary is never compared, so drifted counters it returns end the visit
inconclusive without recording a failure; a scope first enumerated in the final run of a multi-run
check is not taken as verified, so the snapshot that copied it stays Unconfirmed; a snapshot taken after
a mode-epoch-only advance during a check is not confirmed by it; and a check that compares nothing is
unverified and leaves the last passing boundary alone. Live
cadence, cost and drift numbers remain to be measured on staging.
Scheduled stale-statistics repair¶
Many correctness paths leave a live row Stale on purpose: the drift check above, a question-definition rewrite (the fence completes without a rebuild), shared-allocation and other dependent-family invalidations, fence completion and lost publication races. A Stale row serves correct values through the slower authoritative fallback, but before this job nothing republished it except an administrative backfill, a forced rebuild or a manual fleet run. The Project Management host can now look for such families on a schedule and republish them through each family's ordinary, non-forced backfill (FEAT-024 C1). Nothing is enabled by this change, and no environment configures it.
Configuration¶
All keys are read by the Project Management host only (environment variables carry the SYRF__ prefix
and use __ for :).
| Key | Default | Meaning |
|---|---|---|
ProjectStatisticsRepair:Enabled |
false |
Registers the schedule and lets the consumer run. Off, the consumer acknowledges a tick without work even if an earlier Quartz trigger still delivers. |
ProjectStatisticsRepair:CronExpression |
0 15 * * * ? |
Quartz cron, evaluated in UTC: hourly at quarter past. It must never fire twice within 15 minutes; a tighter cron refuses host start. |
ProjectStatisticsRepair:ProjectsPerRun |
16 |
Allowlisted projects inspected per run (1–256), in a persisted rotation. |
ProjectStatisticsRepair:RepairsPerRun |
8 |
Family backfills started per run across all projects (1–64). Families found after the budget is spent are reported deferred. |
ProjectStatisticsRepair:ScopesPerProject |
1024 |
Scope inspections (served-row reads) per inspected project per run (1–4096). It bounds detection only; a backfill always recalculates every declared scope of its family. |
ProjectStatisticsRepair:RequireSettledSource |
true |
Settle check: backfill a family only when the project's source clock has not moved since the previous run observed it; otherwise report skipped with reason=unsettled. The first visit to a project only records the observation, so with the default a repair lands one run after a row goes Stale at the earliest. |
ProjectStatistics:ProjectAllowlist |
empty | The reviewed allowlist, the only source of projects the job visits. |
With Enabled=true, the host refuses to start unless the cron is valid and fires no more often than
every 15 minutes and every budget is in range. The job needs materialized writes
(materializedProjectStatisticsWrites) and each family's own flag; a family whose flag is off is neither
inspected nor reported. It does not use the fleet runner, so ProjectStatistics:Fleet:Enabled does
not gate it (see "Why not the fleet runner" below).
Why hourly. A stale row is a performance problem, not a correctness one, so the job only needs to bound how long a page stays on the fallback, not race writers. An inspection that finds nothing costs indexed reads only. Each repair, though, is one full authoritative recalculation of one family for one project, and on a project that is still being written a family invalidated by every write would only get a short Fresh window from each attempt, so a much shorter cadence would mostly repeat expensive work. Quarter past keeps the job off the 00:05 daily observation tick; overlapping with the 02:30 maintenance or 03:30 drift-check ticks is harmless because each service serializes through its own leases and compare-and-swap writes.
What one run does¶
- It sorts the configured allowlist (minus any project the host's flag source does not also admit) and
inspects at most
ProjectsPerRunprojects from a persisted rotation start. The start moves every run: to the first project that had a familydeferred, otherwise one project on when the whole allowlist fits in one run, otherwise past the last project visited. So no project always spends the repair budget first, and a budget smaller than the allowlist reaches every project withinceil(allowlisted / ProjectsPerRun)runs. - A project is skipped (with a typed reason, retried next run) when the fleet mode is not Enabled,
materialized writes are off, it has no statistics control or is deleted, its own mode is not Enabled
(including the Disabling drain quarantine), a definition-rewrite or inclusion-recalculation fence is
raised, a staged import is active, the durable reviewer mode disagrees with the flag, or its declared
versions differ from the fleet's.
The project gate, family guards, authoritative scope enumeration, served rows and operation fences
are read in one pinned source snapshot. If the host cannot open that snapshot, it reports
snapshot_unavailableand repairs nothing. The snapshot is released before any backfill begins; each backfill then rechecks its own admission against current state. - For each family with a registered backfill and its flag on, it reads the publication guard and then
each declared scope through
ProjectStatisticsServingGate, the reader's own predicate. A family whose guard is absent, Stale or Missing, or with any served row Stale, Missing or published under an older write epoch, needs repair. A fence, a staged operation or a live rebuild anywhere in the family skips it for this run. A row the reader refuses for another reason (for example an incompatible row) is reportednot_repairableand left alone. A never-published family that declares no scope for the project (for example search population without searches) has nothing to publish and iscurrent; it is never backfilled. If it previously published a scope but now declares none, the job reportsskipped/not_repairable. A family with scopes but no guard at all has never been backfilled for that project: it ismissing, and the same ordinary backfill bootstraps it. Turning a family flag on for an allowlisted project therefore establishes its baseline on the next run; keep the flag off to keep a family unbootstrapped. The starting family rotates per project across visits, so evenScopesPerProject=1eventually inspects a later family when an earlier current family spends the whole scope budget. - Before any backfill: a family with an active backoff (below) is reported
backed_off; with the settle check on, a project whose source clock moved since the previous run isskipped/unsettled; and onceRepairsPerRunis spent the family isdeferred. None of them spends the budget. - A family that needs repair is handed to its ordinary backfill with
force: false, the same call the administrative backfill route and the fleet runner make, with ownerscheduled-stale-repair. The backfill re-checks the allowlist and flags, records the family'sbackfill-observedbootstrap point first, takes its own per-scope rebuild leases, publishes through the control compare-and-swap and verifies the final bundle. A scope that is already current is skipped by the backfill itself.
The job never forces: it never adopts a changed configuration, never re-stamps a control and never
publishes rows without their bootstrap point. If a family previously published scopes but now declares
none, it reports skipped/not_repairable: ordinary backfill also enumerates only current source scopes
and cannot republish those rows. A single published scope that vanishes while sibling source scopes
remain is not yet classified by the detector; #3792
tracks bounded inspection of those non-Fresh published scopes. The reader's fallback and drift quarantine
still apply. An operator must reconcile the source and then use the appropriate family backfill or
forced rebuild; the schedule does not invent a zero row.
Outcomes and what an operator does¶
Each family's answer is one of these (also the outcome telemetry tag):
| Outcome | Meaning | Operator action |
|---|---|---|
current |
Every declared scope is served; nothing was called. | None. |
repaired |
The backfill republished the family; its scopes serve materialised again. | None. |
deferred |
Needs repair, but the run's RepairsPerRun budget was spent. |
None; the next run starts at this project. Raise the budget if it persists. |
skipped |
A fence, staged operation, quarantine, live rebuild, closed gate, unsettled source, scope budget or a row the backfill cannot repair (reason says which). |
None; the next run looks again. |
backed_off |
An earlier configuration_refused or bootstrap_identity_occupied answer still holds, because nothing that could change it has moved. The backfill is not called; the report repeats the original reason and failure_reason. |
As for the original answer. |
refused |
A typed, retryable backfill refusal (failure_reason, e.g. publication_race_lost, statistics_rebuild_busy, family_fenced). |
None unless it persists across many runs. |
configuration_refused |
The authoritative configuration or versions differ from the project's established ones (digest_mismatch, control_version_mismatch, fleet_version_mismatch). The rows stay on the fallback. |
Confirm the configuration change is intended, then run the family's forced rebuild (POST api/admin/project-statistics/{id}/rebuild, or /{id}/{family}/rebuild). The schedule never does this for you. |
bootstrap_not_confirmed |
The backfill could not record the family's bootstrap point, so it published nothing. reason=bootstrap_identity_occupied is the known limitation below; any other reason is a transient checkpoint refusal retried next run. |
See below. |
failed |
An unexpected exception, logged at error level with the project ID. | Investigate the log if it repeats; other families and projects were not affected. |
Backoff. The two answers an automatic retry cannot change are remembered per (project, family) in
pmProjectStatisticsStaleRepair, with a key read right after the refusal: for configuration_refused
the control's configuration digest, the family guard's version and the project's source revision; for
bootstrap_identity_occupied the checkpoint identity (source revision, projection revision, fleet mode
epoch). A fleet_version_mismatch is retried on the next run rather than persisted: deploying a
compatible writer can resolve it without moving any stored project identity. While a persisted key is
unchanged the family is backed_off and costs three small reads. Any write to
the project, a forced rebuild or a configuration change moves the key, and the family is attempted again
on the next (settled) run. A family that becomes current or repaired drops its entry; entries are
capped at 32 per project.
The occupied-bootstrap limitation. An ordinary backfill must record exactly one backfill-observed
point for the family before it publishes. If that root no longer exists (history maintenance retired it,
or the family was seeded by rebuilds and never had one) and a daily observation root already occupies the
project's current (source revision, projection revision, mode epoch) identity, the checkpoint admission
answers "already published" and the backfill stops with nothing published. The job reports
bootstrap_not_confirmed with reason=bootstrap_identity_occupied, logs a warning with the project ID,
and does not force; later runs report it backed_off until the identity moves. Two ways out:
- Wait. Any later write to the project moves its identity, and the first settled run after it records a new bootstrap point and republishes the family.
- Repair now. An administrator runs the family's forced rebuild
(
POST api/admin/project-statistics/{id}/rebuild, or/{id}/{family}/rebuild). It republishes the Stale scopes and moves the projection revision; the next scheduled run then finds the familycurrent.
Refusals and failures¶
Every refusal and exception is isolated to one family of one project and retried by the next run; one project's failure never stops the others. The consumer never faults its message, runs one run at a time per process and publishes no startup catch-up run (a missed tick is sent once by the misfire policy). Only shutdown cancellation propagates.
Cost¶
Inspection costs at most ScopesPerProject served-row selections (one or two indexed reads plus a fence
lookup each) and one scope enumeration per enabled family, per inspected project, plus one small state
document read and write. Repair costs at most RepairsPerRun ordinary backfills per run. A backfill is
not bounded by ScopesPerProject: it enumerates and handles every declared scope of the family, and each
automatic scope costs about two authoritative calculations (the already-current check, then the rebuild's
own calculation), plus the bootstrap point's calculation the first time. A configuration_refused
attempt pays the same scope pass (each scope then refuses in its compare-and-swap), and a
bootstrap_identity_occupied attempt pays a scope enumeration and a walk of the project's history; the
backoff makes each of those a one-off until something relevant moves. At the defaults one hourly run
inspects at most 16 projects and runs at most 8 family backfills. The rotation is one document in
pmProjectStatisticsMaintenanceCursor (id stale-repair); per-project state is one document per
visited project in pmProjectStatisticsStaleRepair (the last observed source revision, last starting
family, at most 32 backoff entries and a per-family last-inspected scope cursor).
A family that ordinary writes invalidate as a whole (for example the reviewer families after a screening decision without point-maintenance evidence, which mark the family Stale) is Stale again after every such write. The settle check (on by default) repairs it only after a run interval with no write to the project, so a busy project's budget is not spent on a Fresh window that closes minutes later; its readers stay on the fallback while it is busy and are repaired within about an hour of it going quiet.
The scope budget is a bound on served-row inspections, not on the calculator's authoritative scope enumeration. Each visited family remembers its last inspected canonical scope in the per-project state and resumes after it, wrapping at the end. A stale scope beyond the first run's inspection budget is therefore reached on a later run. A family-wide invalidation marks the guard and is found without reading any scope. Scope enumeration still materializes the authoritative family set on each visit; keep this default-off job on reviewed allowlisted projects until source-side keyset paging is available. Track source-side paging in #3790.
Why not the fleet runner¶
The fleet runner can already call the non-forced backfill (a run with repair: false), but it cannot
serve a schedule: a run pauses at its first failed target and needs an administrator to resume it, so one
refusing project would stall every other; at most 64 runs are retained and only an administrator can
prune them, so an hourly job would exhaust the capacity within three days and then block operators' own
runs; one global lease admits one target per 30 seconds and would contend with operator runs; and its
targets are whole project families, so detection is needed either way. The job therefore calls the same
backfill service directly, with the same leases, bounds and typed outcomes; fleet runs remain the
operator's tool.
Telemetry¶
syrf.project_statistics.stale_repair.projectscounts project visits byoutcome(inspectedorskipped) and, when skipped,reason(fleet_not_enabled,writes_disabled,project_unavailable,project_not_enabled,source_visibility_fence,staged_operation,durable_mode_disagreement,incompatible,unexpected).backed_offfamilies carry the originalreasonandfailure_reason;skippedfamilies includereason=unsettledfor the settle check.syrf.project_statistics.stale_repair.familiescounts family answers byfamily,outcome(the table above),reason(the triggerstale,missingorepoch_mismatch, or the skip reason, includingbootstrap_identity_occupied) and, for a typed backfill refusal,failure_reason.- Each backfill the job starts is also counted by the existing
rebuilds.requested/rebuilds.completed/rebuilds.failedinstruments withoperation=backfill.
None of them carries a project identifier; the structured logs name the project for refusals that need an operator. See the runbook Telemetry section.
Enablement order¶
- Confirm the target projects are in
ProjectStatistics:ProjectAllowlistand serving materialised statistics, and that exactly the families you want maintained have their flags on (a flagged family with no baseline is bootstrapped by the first run). - Enable the repair before, or together with, the drift check: the drift check and the daily copy only ever mark rows Stale, so without the repair a drifted row stays on the fallback until someone acts. Order on the Project Management host:
- the daily observation schedule (and, if wanted,
ProjectStatisticsDailyObservations:CopyFreshMaterializedRows); ProjectStatisticsRepair:Enabled=true(optionally with cron and budgets);ProjectStatisticsDriftCheck:Enabled=true.- Watch
stale_repair.familiesby outcome for the first runs.repairedshould follow everyparity.invalidationswithin about two runs once the project is quiet;configuration_refusedandbootstrap_not_confirmedneed the operator actions above; persistentdeferredmeansRepairsPerRunis too small.
For a staging configuration PR, the keys go in syrf/environments/staging/project-management/values.yaml
under env: in cluster-gitops, for example:
SYRF__ProjectStatisticsRepair__Enabled: "true"
# Optional; the defaults are shown.
SYRF__ProjectStatisticsRepair__CronExpression: "0 15 * * * ?"
SYRF__ProjectStatisticsRepair__ProjectsPerRun: "16"
SYRF__ProjectStatisticsRepair__RepairsPerRun: "8"
SYRF__ProjectStatisticsRepair__ScopesPerProject: "1024"
SYRF__ProjectStatisticsRepair__RequireSettledSource: "true"
Rollback¶
Set ProjectStatisticsRepair:Enabled=false. The consumer then acknowledges ticks without work, so no
Quartz clean-up is needed. Nothing the job wrote blocks rolling back the binary: every publication it
caused is an ordinary backfill publication (and bootstrap point) that an older binary already reads, and
its rotation and state documents can be left in place or dropped.
Evidence¶
Real replica-set tests over the production services show that a drift-check-marked scope and a
definition-rewrite-marked scope each become Fresh and are served materialised after one run (and the
next run finds them current without another backfill); that a source-visibility fence and a bulk
operation fence are skipped with nothing backfilled and the family is repaired once the fence is gone;
that a foreign configuration is refused typed (digest_mismatch) with the control's digest unchanged and
the rows still on the fallback, and then backed_off without another backfill until a write moves the key; that a disabled
schedule writes nothing; that the settle check waits for a run interval without writes; that
ProjectsPerRun=1 rotates through the allowlist and the run after a RepairsPerRun=1 deferral starts at
the deferring project; that the start moves on by one when every project fits in one run; that a first
project whose family keeps refusing does not stop a later project being repaired; that one family's
exception does not stop another family's repair; that an occupied bootstrap identity is reported
bootstrap_not_confirmed / bootstrap_identity_occupied without forcing, then backed_off, after which
the documented forced rebuild restores serving and the backoff is dropped; that a flagged family with no
scopes is current and never backfilled; that a flagged family with no baseline is bootstrapped as
missing; that an older-write-epoch row is repaired; and that a live rebuild, a row refused for another
reason, the scope budget, a durable reviewer-mode disagreement and a fleet version mismatch are all
skipped without a backfill. Every recorded backfill call used force: false.
Live cadence and cost numbers remain to be measured on staging.