Materialized statistics programme status¶
For what each statistic counts, how it is rebuilt and what makes it stale, see the statistics reference. The 30 September as-built audit of every mutation family is in the mutation ownership matrix; it opened #3839, #3840 and #3841; #3839 and #3841 are since fixed (below), #3840 remains open.
Delivery reconciliation, 28 September 2026. The daily observation producer #3563, four observed-history consumers #3568/#3573/#3574/#3578, current-statistics consumers #3570/#3572/#3580/#3582/#3583/#3586/#3587, fleet operations #3253/#3263, checkpoint retention #3581, delta maintenance #3584, source-receipt maintenance #3588, review/reconciliation route loading #3590, and transient inclusion-fence recovery #3764 have merged. Their code delivery does not establish staging activation, an operating schedule, parity, live browser delivery, soak or production acceptance.
Annotation reservation/baseline correctness #3763 has also merged; older stage and membership-stage annotation baselines must be rebuilt before serving. #3838 (merged) routes the screened-reservation release through the one reservation save with a per-attempt operation id and adds a bounded retry of transient statistics rejections on claims, admissions and review-settings saves (#3740 items 3 and 4). Still open are rollback refresh #3591, settings-only refresh #3766, their dependent frontend/backend slices, and independent two-API local proof #3765. An earlier stronger local page/family-gated proof was blocked by a reproducible staged-fence release write conflict; #3767 has a pushed but unmerged fix. A subsequent bounded local #3765 run passed with Pages on, initial Stage Overview materialized reads, parity and two-API delivery before the first poll. This is local proof, not deployed cross-API-pod browser or Angular automatic-reconnect acceptance.
A read-only staging observation on 28 September found API, project-management and web images at
sha.9d6b6bf (one API, one writer, two web replicas). This includes #3764/#3588 code, but receipt
maintenance remains off and no supported staging fence-recovery mutation or annotation activation
is established. The 22 September deployment/parity table below remains a dated baseline, not a
statement of current images. No performance, soak, cross-API-pod browser or production acceptance
is established.
Search import parse-phase fence (#3839). A search import now admits its statistics fence before its first Study write, so the stored screening and annotation rows fall back to the live calculation for the whole parse instead of being served Fresh over a population that already includes the imported Studies. Imports into an allowlisted project are refused with a parse error while another import or bulk Study update holds the fence. Code delivery only; the families still end Stale and need a backfill after each import.
Agreement-setting change fence (#3841). For an
allowlisted project the screening agreement-mode save now raises the inclusion-recalculation fence in the
same transaction, so rows counted under the old threshold are never served Fresh while the recalculation
command is in flight. The administrative update-study-inclusion routes refuse while the setting change's own
recalculation job is open, instead of clearing its token mid-pass; the fence diagnostic's recovery text now
says so. Code delivery only; families still end Stale and need a backfill after each threshold change.
Project Overview own-progress slice #3781 (merged). The page's reviewer screening and annotation progress now has a route-provided SignalStore that acquires its project-scoped authorized read on direct Overview entry, shares it among mounted consumers, and releases it on route exit. A same-scope transient failure may show the last accepted response as stale; permission loss, identity change, incompatible payload and definitive refusal clear it. The existing flag-off and legacy-shape restoration remain. This is code delivery only; it is not deployed or activation evidence.
Stage Review personal-progress slice #3784 (merged). Review and reconciliation routes now demand one stage-scoped SignalStore for the signed-in reviewer's screening and annotation response. The Stage Review header, progress dialog and completion page select its signals; save and existing refresh actions are handed to that route owner, while flag-off keeps the global effect and its live refresh path. Form drafts remain in their existing owners. The stage response has no server revision or captured definition ID, so the store cancels on scope or current-stage-setting changes and computes only from the accepted response plus the stage settings captured with that request. After a transient refresh failure, retained personal counts are visibly labelled as from the last successful update on Stage Review, the completion page, and either progress dialog; the label clears on recovery. Parent-consumer activation, parity and rollback proof remain, followed by the separate legacy-row retirement gate. This is not deployed or activation evidence.
Historical 22 September 2026 baseline at main commit 9ddb8ec659bd63fa3b7b9fb65a05bad205471790.
At that point the screening pilot was deployed and configured on staging, several families still
lacked operational baseline routes, and most consumers used authoritative aggregates. Those code
gaps have since changed as described above; fleet scheduling and retention acceptance, performance,
soak, production and retirement gates remain open.
There is no meaningful completion percentage across these different obligations.
The phase 0–3 requirement audit and historical chart rollout record the newly requested audit and exact chart dependencies.
This is the current delivery record for epic #1831. The feature brief and technical plan remain the design and acceptance contract. The September recovery ledger and local Claude/Codex handovers are dated historical evidence, not current readiness declarations. This document supersedes their delivery summaries; it neither changes the approved design nor authorizes activation.
Evidence and interpretation¶
“Implemented” means code exists at the pinned main commit; “merged” does not mean activated or accepted. “Configured” means observed deployment settings or an effective API flag; it does not establish freshness, writer health, browser receipt or parity. “Observed” describes one dated result, not sustained acceptance. An open issue may contain delivered work; a closed issue may still leave an operational obligation.
The sanitized observation snapshot records GitHub states, deployment images, selected statistics settings and the explicitly user-reported parity result. Application tests were not rerun for this documentation audit. Test results below refer to the linked implementation PRs and committed proof reports. No authenticated staging browser or database mutation was performed. Ownership below names the responsible role; named delivery assignees and dates remain unassigned.
Staging preflight (#3802, extended by
#3850). scripts/capture-statistics-preflight.py and the
Statistics Staging Preflight workflow record live /health/live versions, the API's anonymous runtime-flag
snapshot and pinned cluster-gitops desired state, then fail on version or deployment-managed flag disagreement.
Since #3850 the main-only job also reads the running syrf-staging Deployments and pods as the read-only
syrf-stats-preflight identity (#3806): pod image digests,
rollout state and the writer host's deployed FEAT-024 flags, allowlist, SharedReaderMode and lifecycle
switches, compared with GitOps (Secret-sourced values are reported as "from secret", never resolved). A second
job captures one parity report per allowlisted project with the read-only evidence credential and uploads it in
the validator's parity JSONL format, so each run is one soak parity observation; a six-hourly schedule exists
behind the STATISTICS_PREFLIGHT_SCHEDULE repository variable, off by default. This is tooling only, with no
runtime behaviour and no flag. It does not read runtime overrides on the writer host, start a soak or supply
read or receipt evidence. See the runbook's preflight section.
Dated deployment and flag baseline (22 September)¶
| Environment / surface | Observation on 22 September | Limit |
|---|---|---|
| Staging API | 9.117.0-sha.9ddb8ec, one ready replica |
Contains the current fixes; one API replica cannot establish cross-API-pod browser delivery |
| Staging project management | 11.113.0-sha.9ddb8ec, one ready replica |
Writer deployment settings agree with API for the pilot |
| Staging web | 7.147.4-sha.9ddb8ec, two ready replicas |
Two web replicas are not two SignalR API hosts |
| API and writer deployment flags | Writes, Serving and Screening on; allowlist contains 00000000-0000-0000-0000-000000000102 (Ready for Annotation) |
MembershipScreening, Annotation, MembershipAnnotation, QuestionAnswers, SearchPopulation and DerivedSummaries off |
| Effective staging API flags, revision 69 | Pages, ProjectOverview, SignalR and Exports on; StageOverview off | Runtime consumer overrides differ from static defaults. Exports on does not prove an aggregate export consumer exists |
| Production API / writer | 9.44.1-restore.104eca9 / 11.43.1-restore.104eca9 |
Older code; no materialized-statistics deployment keys observed. Do not describe production as current code deployed dark |
| Production web | 7.39.0, three ready replicas |
No production rollout or activation established by this audit |
Update, 30 September. The staging consumer flags that differed from GitOps only through runtime
overrides are now recorded in GitOps: cluster-gitops
#1391 (merge 9f37cb7b) sets
materializedProjectStatisticsProjectOverview, …Pages, …SignalR, …Exports and the new …History to
true in both the staging API and project-management values files. The matching runtime overrides
(revision 72) are to be cleared by an administrator; this record does not claim that has happened. Family flags are unchanged: only project screening is on in staging.
The project name “Ready for Annotation” does not mean its annotation-statistics family is enabled. Likewise, enabling screening and annotation activities on a stage is distinct from enabling materialized statistics for those activities. Today's API observation cannot independently verify every writer or durable project control.
The latest supplied screening audit was 21 September, 15:32:11 UTC, checkpoint 9.11.0:
17 metrics matched exactly, the published scope was Fresh/Available, materialized reads were used,
and configuration matched. There were 32 studies, 31 sufficiently screened, 28 sufficiently included,
3 sufficiently excluded and 2 overscreened. The earlier 16 September audit at 0.2.0 also matched.
These are useful point observations, not a current audit or seven-day soak proof.
Phase reconciliation¶
| Approved phase | Delivered | Still required |
|---|---|---|
| 0 — inventory and baselines | Catalogue and method-level ownership (#3070), benchmark definitions (#3076), recovered-design reconciliation | Keep inventory aligned as remaining consumers migrate; measured acceptance is separate |
| 1 — shared foundation | Versioned bounded projections, transactions, receipts, fences, checkpoint/history contracts, compatibility maps, compaction and fallback seams | Operational scheduling/retention and rollout proof; not a new foundational redesign |
| 2 — screening | Classifier and writes (#3149/#3159), transaction/capacity admission (#3232), current read/parity adapter (#3196), pilot administration (#3523/#3531) | Sustained staging proof and performance gates |
| 3 — annotation, questions, reconciliation | Annotation moves, question/definition fences and post-commit fixes; all nine physical-family baseline routes merged; #3763 corrected reservation and zero-candidate paths | Rebuild old annotation baselines, prove lifecycle/concurrent writes and parity before family activation |
| 4 — remaining families / summaries | Reviewer baseline (#3239) and transactional maintenance (#3297); search population (#3240); derived summaries (#3244/#3256); source-row export compatibility (#3241) | Remaining adapters/consumers and measured acceptance; phase is partially delivered, not unstarted |
| 5 — consumer migration | Project/stage overview, reviewer progress, search, question counts, observed-history consumers and review/reconciliation route cutover #3590 have merged in guarded slices | Open rollback and settings refresh slices; transport, activation and eventual legacy-removal acceptance |
| 6 — fleet / soak / retirement | Durable fleet/project controls, bounded family-routed fleet operations #3253, completed-run pruning #3263, checkpoint/delta maintenance | Open scheduled maintenance/drift/snapshot stack, family follow-up #3254, controlled soak, production approvals and retirement |
Numbers in the following tables refer to the PR and issue evidence inventory. Phase labels describe the approved programme, not independent claims that all acceptance gates have passed.
Every metric family¶
The source calculator registry and administrative backfill router now support all nine physical families listed below. This establishes code routes, not deployed baselines, Fresh rows, parity or permission to enable a family. DerivedSummary is a virtual composition of coherent bundles, not a tenth physical backfill adapter.
| Family | Main implementation and operational baseline | Consumers / exact remaining work |
|---|---|---|
| ProjectScreening | Transactional point changes, lifecycle fences, baseline/rebuild, parity and Fresh/fallback reader implemented | Project Overview screening shipped; broad legacy reads and rollout/performance acceptance remain |
| MembershipScreening | Baseline calculator plus transactional before/after maintenance in the screening writer; #3297 | Guarded member consumers merged; authorization/parity and activation proof remain; currently off |
| ReviewerScreening | Baseline and transactional point maintenance implemented | Guarded own-reviewer/project progress consumers merged; other legacy reads and activation proof remain |
| StageAnnotation | Classifier/writer moves; runtime calculator/backfill added #3505 | Dedicated Stage Overview pie shipped but off. Validate lifecycle changes, enable only after family proof; other annotation graphs remain legacy |
| MembershipStageAnnotation | Classifier moves and scope-calculator code exist; runtime calculator registration and admin baseline/rebuild route merged #3561 | Guarded consumers exist, but family flag is off; rebuild old zero-candidate baselines and prove Fresh maintenance before activation |
| ReviewerAnnotation | Stage-scoped calculator, registered runtime router entry and guarded admin baseline/rebuild routes merged #3567 | Guarded own-reviewer progress consumer merged; family flag remains off and permission/parity proof plus baseline rebuild remain |
| QuestionAnswers | Authoritative calculator, tally/definition refresh and fences, current-question binding fixes; operational baseline/rebuild route merged #3562 | Guarded designer counts and assignment locks merged; family flag off by default; prove update/rebuild lifecycle before activation |
| DomainReconciliation | Authoritative calculator and annotation-classifier moves exist; operational baseline route merged #3564. Existing stage/membership annotation responses carry reconciliation fields; the frontend has no standalone public materialized reconciliation read | Family flag remains off by default; lifecycle, parity and activation proof remain. Any future direct family consumer needs a separate API contract and cutover |
| SearchPopulation | Calculator, baseline, import/unlink fences and authoritative search ownership merged in #3240 | Guarded project/search consumers merged in #3587; activation proof and #3474 follow-ups remain |
| DerivedSummary | Coherent screening and stage/member annotation composition with retained configuration (#3244/#3256); frontend percentages and chart segments derive from their authorized parent responses, with no standalone public DerivedSummary read | The default-off MaterializedProjectStatisticsDerivedSummaries flag is currently unused by production current/history queries; toggling it cannot activate a derived consumer. Verify parent-family configuration/version and consumer rollout under their actual gates; no standalone family activation proof exists |
Source anchors: screening writer, annotation writer, question calculator, derived contract, reviewer maintenance contract. The implementation of a signed move or calculator is not proof that an operator can build and serve that family today.
Every consumer surface¶
| Surface / path | Current behavior | Remaining implementation or acceptance |
|---|---|---|
| Project Overview screening totals | Dedicated screening endpoint and adapter, authoritative fallback; enabled on staging | Read-path performance, parity and soak acceptance |
Project Overview other panels / broad ProjectController.GetFullStats |
Still authoritative broad aggregation | Migrate approved sections independently; remove broad fetch only when all callers have replacements |
ReviewController.GetFullStats |
Runs authoritative aggregate, optionally substitutes an equal screening section; Pages guard is present | Replace mixed FullStats splice; this path is not a performance cutover |
| Stage Overview annotation pie | Dedicated stage family adapter, separate consumer gates | Family/consumer activation and acceptance; currently off |
| Stage Overview screening/member leaderboards, annotation member tables and historical areas | Coherent current-statistics and observed-history consumers merged behind separate gates; legacy fallback remains | Activation, parity, browser history and retained-observation acceptance |
| Stage Review header, open progress dialog and Review Completed | Shared live/recovery refresh and guarded current reviewer consumers are merged; #3590 avoids unused broad loads on review/reconciliation routes | Review-specific data still uses authorized endpoints; #3590 activation and parity proof remain |
| Screening Overview leaderboard | Coherent Screening Information consumer merged behind a gate; legacy fallback and visibility rules remain | Activation and permission/parity proof; #3514 name-visibility fix merged |
| Question answer counts / design and assignment UI | Guarded live designer counts and assignment locks merged; legacy path remains for rollback | Activation and parity proof; remove manual tally mechanism only after replacement |
| Search lists / systematic searches / project overview study count | Guarded imported search counts merged; authoritative shared-search links remain | Activation, permission and import-visibility proof |
| Existing bibliographic, screening, annotation and outcome exports | Source-row exports intentionally remain authoritative; compatibility #3241 and authorization #3243 merged | Preserve compatibility; they are not unfinished aggregate-statistics cutovers |
| Aggregate-statistics reports / exports | No concrete materialized aggregate export endpoint established in current source | Product owner must identify an approved consumer before adding one; flag alone creates no report |
| Historical API / UI | Bounded observed-history APIs and the four existing chart/dialog consumers merged behind separate default-off gates | Verify deployed observation capture, authorization, real browser rendering, gap handling and retention before activation |
| SignalR full-statistics payload | Existing legacy path remains alongside new revision invalidations | Migrate all dependent consumers and verify replacement before deleting legacy push/calculation |
| Derived summaries | Reader can compose current or retained coherent bundles; current frontend displays derive from their parent route owners | Verify each parent consumer and retained-definition contract before activation; no blanket page migration claim |
Frontend reconciliation and derived-family accounting¶
The generated Web client exposes DomainReconciliation administrative backfill/rebuild methods and the
family enum, but no public read dedicated to DomainReconciliation or DerivedSummary. The existing
StageAnnotationStats and MembershipStageStats DTOs carry reconciliation counters and candidate
availability. Their named reconciliation fields have no runtime display reader today; the candidate
session cards on the reconcile route use session/form state, not a statistics request. This is a
frontend-consumer inventory, not a change to the backend family calculator or its maintenance routes.
The Stage Reviewer Progress SignalStore
owns the existing authorized getReviewerStatsForStage read across review and reconcile routes. It
includes DomainReconciliation in its invalidation families and computes progress segments from one
accepted response and the stage definition captured for that request. Stage Reconcile
acquires and releases that owner; Stage Review, its progress dialog and Review Completed select its
signals when the existing Pages and ReviewerProgress flags are on. The router regression
covers direct review/reconcile entry, one shared read, coalesced refresh, route release, permission
loss, terminal refusal and flag-off restoration.
The Stage Overview presentation derives screening/annotation chart series from one accepted bundle; its standalone annotation pie has a separate route SignalStore for the existing stage response. Project Reviewer Progress derives its displayed screening and annotation segments from its accepted project response. Global selectors in Project Overview, Stage Review, the progress dialog and Review Completed are explicit flag-off legacy adapters. They remain until rollback and legacy-retirement gates are met; they do not imply another DerivedSummary endpoint or a new polling loop.
The backend ProjectStatisticsDerivedSummaries decoder validates parent-family rows. Stage Overview,
Stage Reviewer Progress and Project Reviewer Progress queries call it while gating on their physical
source families and consumer flags. No production current/history query checks
IsFamilyServingRequested(DerivedSummary); the default-off DerivedSummaries flag currently changes
effective-flag reporting, not these consumer reads. Parent-family configuration/version and rollout
proof belongs to those actual query gates. Enabling this unused flag would prove no UI cutover.
This accounting does not activate a family or consumer, prove deployed parity or performance, remove legacy loading, or complete Phase ⅚. A future standalone family display needs an independently reviewed public contract and consumer cutover.
The original epic milestone #1834 and #1845 require a basic historical view, while the older technical-plan phase 5.6 labels visible history UI optional. The current instruction to implement phases 0–3 followed by 4–6 includes the existing basic area-chart surfaces described in the history rollout. Backend history alone does not complete that UI outcome. Richer #1849 visualization remains a later enhancement.
Consumer and operations source trail¶
These links pin the inspected code, so later changes do not rewrite the evidence:
- ProjectController: screening-only and broad FullStats endpoints.
- ReviewController: reviewer/FullStats reads and screening-write sufficiency checks.
- StageStatisticsController: dedicated annotation consumer.
- ProjectStatisticsAdminController: supported baseline/rebuild families and awaited execution.
- StageOverviewComponent: refresh subscription and remaining legacy selectors.
- Stage progress refresh: authenticated flag-gated scope, invalidation and recovery stream.
- Notification implementation: dispatch, broker consumer, subscription filtering and delivery checks.
- History implementation: backend history and daily observation contracts.
Live updates and screening eligibility¶
#3534 implemented leased outbox draining, broker fan-out, authorized project subscriptions, permission checks at delivery and client revision invalidation. #3547 added stage progress refetches, including the open dialog. #3551 subscribed Stage Overview and its annotation pie, coalesced bursts, retained a trailing read, and added signal-backed state and recovery refreshes. It also added computed screening eligibility and server rejection of new decisions after sufficiency, while retaining annotation and existing decision corrections/reconciliation.
#3551 is merged and its commit is on staging. Its browser acceptance run passed four tests including a remote overview change, editable annotation after screening completion, HTTP 409 for an extra screening, and exactly two persisted decisions. Focused controller/Angular regressions and current-head PR checks passed. That browser evidence exercises authorized HTTP/store/signals and polling recovery, not staging SignalR transport or cross-pod delivery. An outbox observed drained on 21 September similarly proves delivery progress, not receipt/rendering in every browser. The release encountered the separately tracked PDF test timeout #3539; do not summarize every release job as green.
Remaining staging acceptance for #3511: Use independent test accounts and automated proof, never a personal account or interactive sign-in. The observed staging and production Auth0 tenant/client/audience settings are shared, so granting global administrator metadata is not a staging-isolated workaround; no such grant is recorded. Local #3765 proof and open #3767 are separate from deployed cross-API-pod browser acceptance.
- Two authenticated viewers see the same committed change without manual reload; record event arrival and render.
- Connect viewers to distinct API pods and prove broker fan-out; current one-API-pod staging topology is insufficient.
- Disconnect/reconnect one viewer, prove missed-event recovery, duplicate/out-of-order handling and burst/in-flight behavior.
- Verify revoked permissions cannot receive data, flag-off rollback works, and active annotation drafts survive refresh.
- Repeat for each activated family and its lifecycle/configuration changes, not only screening point decisions.
#3552 merged as eligibility-policy documentation, not a replacement for #3551. #3553 merged its separate feature-flag UX work; atomic group controls do not prove family readiness, parity or cross-host maintenance safety.
Operations, fleet and history¶
| Obligation | Delivered | Remaining / dependency |
|---|---|---|
| Single-project controls | Durable fleet/project epochs, prepared stop/drain modes (#3371); prepare/build/maintain/compare/use controls (#3523/#3531) | Operator evidence for restart, quarantine/drain and rollback; current flag state alone is insufficient |
| Backfill execution | Bounded services and nine-family admin routing merged. A publication aborted by a transient Mongo write conflict (the outbox dispatcher claiming the family slot) is now re-run a bounded number of times and, if still contended, refused as a typed retryable 409 RetryExhausted instead of HTTP 500 (#3826) |
Admin requests still await work; returning 202 is not an asynchronous durable job/status runner. No live annotation baseline or parity accepted here |
| Automated parity evidence | Read-only statistics-parity evidence credential (#3807): Identity client syrf-statistics-evidence with the single scope statistics:parity:read, accepted by the API on the parity read only (allowlist still applies), plus scripts/fetch-statistics-parity.py for the runbook's parity JSONL. Off by default: not seeded without its secret |
Infra must create the syrf-statistics-evidence Secret, set statisticsParityEvidence.enabled for Identity and add SYRF_STATISTICS_EVIDENCE_CLIENT_SECRET as an environment secret of the main-only statistics-evidence GitHub environment (since #3908; not a repository secret) before any automated capture; the preflight workflow's parity job (main only) uses it |
| Statistics operator identity | Least-privilege automation identity (#3908, off by default): Identity client syrf-statistics-operator, sole scope statistics:operate, accepted by the API on the pending-index check/build, fold status/enable/disable/reset, the per-family backfill/rebuild routes and the parity read only (ProjectStatisticsOperatorPolicy: operator or an unchanged administrator); recorded as service:syrf-statistics-operator. Statistics Operator workflow dispatches one operation on staging. See the runbook |
Staging enablement in cluster-gitops (ExternalSecret + statisticsOperator.enabled) only after this ships to staging, then SYRF_STATISTICS_OPERATOR_CLIENT_SECRET as an environment secret of the main-only statistics-operator GitHub environment (not a repository secret); no production client |
| Fleet reconciliation | Bounded family-routed fleet operations merged in #3253 | Deployed schedule, restart/resume, failure and throughput proof before fleet readiness |
| Fleet retention | Completed-run pruning #3263 and bounded checkpoint/delta maintenance merged | #3254 follow-up, operating retention and storage-growth proof remain; run pruning is not all history retention |
| Daily history | Quartz daily/startup capture and bounded same-day retry merged in #3563 | Prove actual deployed scheduling, missed-day handling, retention and storage growth; registration alone is not operating evidence |
| Observability | Shared Meter, bounded instrument tags, startup records (#3482), JSON logs (#3515) | Verify collector/export destination, dashboards and durable soak records; no sustained collector evidence established here |
| Multi-family maintenance | Shared transaction/configuration fixes and committed local two-family proof | The 22 September staging observation enabled only screening; second-family lifecycle and configuration proof still needed |
| Cross-host flags | Deployment-managed maintenance guards and matched API/writer settings | #3360 runtime propagation architecture remains separate; do not bypass guards with API-only overrides |
| Async fold pending index | Slice 0 of the async point-fold design (ADR-019): administrator-only POST/GET api/admin/project-statistics/fold/pending-index builds and checks the pmStudy partial index IX_Study_PendingStatistics with commit quorum votingMembers; never built at start-up (runbook) |
Not built in any environment. Run on staging, then in an approved production window; local estimate 25–35 s of scan, Atlas under five minutes, unmeasured on a production-sized copy |
| Configuration repair | Forced rebuild and coherence checks | #3370 sibling-scope repair: a forced rebuild that replaces the control identity now rebuilds that family's stranded sibling scopes (#3700); other families still need their own backfill after a replacement; #3506 family-version coupling needs disposition |
#3343 is closed, but that does not establish async job delivery or an operational collector. The executable source and deployment evidence above take precedence over its state.
Performance and rollout gates¶
| Gate | Available evidence | Remaining acceptance |
|---|---|---|
| Read performance | Read benchmark: PS-DS02 Fresh repository adapter p95 approximately 12 ms versus 1,077 ms, no authoritative aggregate for Fresh reads. The local HTTP harness (runbook: Read performance gate — local harness) has passed only a smoke run on a loaded host (29 September), which is not acceptance evidence | Run the harness in acceptance mode on the 5,000-study fixture on an idle host: at least 20% p95 improvement in both phases and 80% fewer authoritative aggregations. Staging needs an operations identity, approval for runtime arm control, and a way to count legacy aggregations |
| Write overhead | Write benchmark, #3299; #3475 reduced measured command count from 30 to 21 on a loaded host; idle-host run 2026-09-30 (main at 65ae1099b, load per CPU at most 0.31) is acceptance-grade |
FAIL on an idle host: all eight cells exceed the 10% p95 gate (46% to 1,283%; the single-reviewer cells fail by 575-689%, so cost is per-save structure, not only contention). 25 commands per materialized save against 2 source-only (+15 find, +4 update, +2 insert, +1 aggregate, +1 commitTransaction); different-study cells exhaust up to 40% of submissions. Enabled reviewer maintenance is not an arm of this harness and remains unmeasured. Reduce round trips and hot-row contention (#3255), or obtain an explicit revised gate |
| Capacity and repeatability | Fixed benchmark contracts exist | PS-DS03/04/05 and CI harness #3194, storage growth and retention evidence |
| Correctness | User-supplied screening parity snapshots; local unit/integration and two-family proof | Exact integer parity across activated families, no stale-as-Fresh reads, failure injection and lifecycle/configuration coverage |
| Staging soak | Pilot configured; two point audits | At least seven actual days and controlled 10,000 reads / 1,000 mutations, recorded outcomes, restart/resume, fallback, rollback and bounded growth. Calendar elapsed time is not proof |
| Production pilot | No current-code production activation established | Accept preceding evidence and separately authorize deployment/configuration and pilot |
| Wider rollout | Not started/authorized by this record | Pilot acceptance, fleet operations and separate rollout decision |
| Legacy retirement | Existing aggregates, FullStats, old push and source-row exports still used | Prove caller replacement and rollback boundary; separately approve calculation removal, then collection/index cleanup where applicable |
Write-path redesign (building, dark). After the 30 September idle-host run failed every
write-overhead cell, the owner approved a
single-document async point-fold design, behind a new default-off
materializedProjectStatisticsFold flag. On 1 October the owner widened the MVP to every screening-
and annotation-dependent family and set the write gate (under 10% or at most +2 ms p95, and zero
statistics-caused failures at ½/5/10 reviewers). A staging-only pilot may enable ProjectScreening
fold mode after the second implementation slice; production waits for the full MVP: saves append a pending entry to the
Study, a leased worker folds entries, and reads serve stored + pending. Slice 1, the dark framework, is being built in
#3881: the Study pending-statistics schema, the
protocol-stamped control tripwire with one version comparison at every site, the flag, receipt
maintenance that never retires a Study with a pending set, and the flag-gated reload-and-retry in the
cached upsert writers. Part 2 (#3883) adds the
control plane: per-pod fleet membership (pmProjectStatisticsFleetMember, keyed {deployment}/{pod}, api and
project-management only, server-time liveness, deleted on shutdown) with the per-project allowlist-consistency refusal in
the scheduled repair, the per-protocol server-time fold-worker heartbeat on the global control, the
project fold lease with its fencing write, the fold quarantine (pmProjectStatisticsFoldQuarantine,
30-day TTL, soft and hard caps; the fold collections' indexes are created at host start-up, never inside a
fold transaction), and published Stale placeholders instead of starting a fold family's
scope from zero. Part 3 (#3885) adds the worker
skeleton in project-management (consumer, one-minute sweep, heartbeat renewal, batch halving, the
two-step livelock discard covering entries and overflow, poison and overflow quarantine, the soft and
hard quarantine caps, Disabled finalization), the shared deriver (protocol 0 derives no transition
kind, so every transition is poison), the reader overlay's plumbing (a project with fold history falls
back while anything is pending), and the enable, disable and reset service (enable is refused while
coverage is incomplete, which it is throughout slice 1; no HTTP route yet). The livelock counters are
durable (pmProjectStatisticsFoldAbort, one record per aborting Study, cleared on commit or discard,
one-hour TTL), so the 20-abort bound holds when the fold lease moves between pods. The sweep runs its
unbounded pending query only while the slice-0 pending index
(#3880) is Ready, hinted to that index so it can never
scan pmStudy, and enable's PendingIndexMissing check reads the same operation. The sweep also
visits every project Disabling for longer than the five-minute writer grace, so a drained project
finalizes to Disabled without a new save. Slices 0 and 1 are merged
(#3880, #3890,
3881, #3883, #3885).¶
Slice 2, ProjectScreening end to end at fold protocol 1 (#3895,
merged; the audit overlay #3902 merged). With the fold flag on, an allowlisted project in fold mode saves a screening decision (all
three shapes: capacity guarded, plain and the eligibility transaction) in one Study write that appends
a ProjectScreeningProfile entry, with the duplicate checks, the free fold-only retry, the 30-second
pre-write deadline and unknown-result resolution by digest; an unprovable operation answers a typed
409 (StatisticsOperationUnresolved). Every other family the save touches travels as a family-wide
invalidation intent, so it stays Stale and is served authoritatively. The worker applies the moves
(row increments, one delta and one receipt per entry, published Stale placeholders) and discards and
stales a family whose moves cannot apply; the reader overlay serves stored + pending exactly with
every fallback threshold; a ProjectScreening rebuild publishes authoritative - pending under the fold
lease and refuses the retryable PendingNotDrained; the copier, the unchanged-day recorder and the
drift check refuse while entries are pending. Enable, disable, reset and status have administrator
routes; enable is allowed only on staging (positive IsStaging(), a database other than syrftest,
ProjectStatistics:Fold:AllowPartialCoverage) and refused everywhere else until slice 6. The staging
pilot procedure is in the runbook.
Also fixed from the slice-1 review: young or skipped Studies no longer hide eligible ones from a fold
run (which now answers Deferred), and a drained holder re-queries once after releasing its lease.
With the flag off behaviour is unchanged.
As-built deviations from the design in slice 2, each fail-closed:
- No rebuild drain. The rebuild subtracts every derivable pending entry in its snapshot and refuses
PendingNotDrainedon anything it cannot subtract, leaving the worker to discard it, instead of running fold batches itself first. Exactness is the same; a rebuild on a project with poison waits one sweep. - Family-wide intents. The reviewer and annotation families a screening save touches are staled family-wide rather than per scope, which keeps every entry small (no 100-member fan-out).
- Receipt duplicates by digest. The save's receipt check compares the submission digest, the same
rule as the pending, overflow and quarantine checks, rather than
MatchesEnvelope's creation time. - No checkpoint observation barrier yet. An authoritative build at an unchanged identity with entries pending computes the authoritative value; a second build at that identity with a different value is already refused by the root's identity rule, so history never holds two values for one identity. The cost is a missing day while the worker is down. Deferred to a later slice.
- Parity audit through the overlay (#3902). Slice 2 alone built the audit's reader with no pending
overlay, so every project with fold history reported
Unavailable(PendingOverlayUnavailable) permanently. #3902 gives the host's audit reader the overlay and puts the pending-set fingerprint in the audit's race bracket: a fold project is audited exactly (stored row plus pending entries against the authoritative calculation), and is inconclusive only when entries change between the audit's two reads. It is the pre-enable gate for the staging pilot. Not done: the design'sstoredOnlydiagnostic (pending count and stored-only difference in the report); the report's verdict does not need it. - One more command on the plain and capacity-guarded saves. The durable reviewer mode is re-read immediately before the Study write (today's saves re-check it inside their transaction), so a mode flip can never commit an ordinary screening without its capacity guard. These shapes cost five commands on a first attempt, not four.
- Unrecorded-drop ranges and the quarantine-incomplete marker are consulted only when an earlier
attempt of the request had an unknown result. They name a revision range, not an operation, so an
attempt known never to have written would otherwise answer
Unresolvedfor another screener's drop. - Reset needs fold history.
POST …/fold/resetrefusesInvalidStatefor a project that has never been in fold mode, because the route is live in every environment. - No PM Study-cache eviction after a fold. Every whole-Study write is version-guarded and the cache holds an instance for two seconds, so a stale cached Study can only lose its write and reload.
Slice 3, the reviewer families at fold protocol 2 (#3922,
in review). A fold-path screening save now carries MembershipScreening and ReviewerScreening as a
ReviewerScreeningPoint transition instead of family-wide intents: the before and after classifier
inputs (each screener's decision counts, the profile key and the persisted sufficient and insufficient
flags), with a membership context digest (SHA-256 of the sorted canonical investigator ids). The
before-state comes from the attempt's own private load: the save reads the whole Study document once
and takes the persisted inclusion statuses from it, so the separate before-state read of the
transactional path is not needed. The fold, the reader overlay and both reviewer calculators' rebuilds
derive the moves for every member with today's ReviewerScreeningPointClassifier, only when the entry's
screening and membership digests equal the snapshot Project's; a member added or removed since the save
is a context mismatch (the reader falls back, the fold discards and stales both families, a rebuild
refuses PendingNotDrained until the worker has discarded it). A project with more than 100 canonical
members, a Project with no threshold, or a load without the persisted statuses keeps today's
family-wide intents, counted by fold.reviewer_fallbacks{reason}. Both reviewer rebuilds publish
authoritative - pending in one pinned snapshot. The protocol bump means the staging pilot needs a
reset and backfills after this slice deploys (runbook).
With the flag off behaviour is unchanged.
As-built deviations from the design in slice 3:
- The canonical member set excludes an empty investigator id, which the transactional writer's candidate list includes. Such a scope is never enumerated by the backfill or read, so no served value changes; the fold simply never creates a placeholder for it.
- The transition is written whatever the reviewer family flags say. It records what the Study changed; a family that is not maintained is discarded and staled at the fold, never moved.
- A re-save of an unchanged decision can append a reviewer entry. When a Study's persisted status differs from the recomputed one (older documents), the re-save re-persists the recomputed status and moves the reviewer families; slice 2 appended nothing and left them wrong until a rebuild.
- No reviewer-family parity audit. The audit exists for ProjectScreening only. On the pilot, the reviewer families are verified by comparing the served value with the authoritative one while nothing is pending (runbook).
- The reviewer family flag stays off on staging.
materializedProjectStatisticsMembershipScreeninggates both reviewer families and is off, so on the pilot the fold discards the reviewer transitions (FamilyNotServable) and the families are served by the legacy calculation. The runbook lists what must hold before it is turned on: the flag's fleet-wide scope, the defined manual exactness check, and no stored-versus-recalculated inclusion status drift (writers other than the fold-path screening save re-persist the recalculated status without a reviewer move).
The production MongoDB server version (the design needs 4.4 or later) is to be confirmed by the operator before slice 1 deploys.
#3255 records the decision to pilot the existing write path before redesigning it. This defers speculative optimization until evidence; it does not waive production performance acceptance. Do not reopen already merged reviewer maintenance as a prerequisite to that experiment.
Ordered remaining work¶
These are reviewable slices. Owners are roles, with individual assignment pending; none of the dates below is a promised delivery deadline. Complete the usable single-project screening proof before broad activation.
| Order | Work / owner | Dependency and completion evidence |
|---|---|---|
| 1 | Screening pilot proof — operations + QA | Existing deployed code; record real SignalR/multiple-viewer evidence, controlled performance and soak. Address demonstrated correctness failures before expanding |
| 2 | Fleet operations — backend + operations | #3253/#3263 are merged; integrate the open scheduled maintenance, drift and snapshot stack, then exercise resumability, bounds, cancellation/retention and failure recovery; complete #3254 only for demonstrated family gaps |
| 3 | Physical-family operational gaps — backend | MembershipStageAnnotation #3561, QuestionAnswers #3562, DomainReconciliation #3564 and ReviewerAnnotation #3567 baseline/rebuild routes are now merged. Remaining: authorized rebuild/parity proof on real projects and supported source-write/lifecycle tests before any consumer or flag activation |
| 4 | Consumer slices — API + frontend | Migrate reviewer/member screening, remaining stage/member/reviewer annotation, question counters, search population and derived summaries independently after their baseline contracts are usable. Keep fallback and permission boundaries |
| 5 | Basic history acceptance — product + backend/frontend | The four default-off history consumers #3568/#3573/#3574/#3578 have merged. Resolve #1834/#1845 scope, prove deployed observations, retention, authorized browser history and rollback before flag activation |
| 6 | Production and retirement — maintainer + operations | Gate evidence accepted; separate production pilot, wider rollout and legacy-removal approvals. No deployment authorized by this docs PR |
Operations work and family implementation can proceed independently, but activation depends on proven family maintenance and readers. The family and consumer tables above define the exact scope of rows 3 and 4; “finish annotation” or “finish Phase 4” alone is not an actionable work item.
Tracking and recent changes¶
The observation snapshot contains each fetched title, state and URL. All following states were checked on 22 September; refresh before acting on them.
| Tracking group | Reconciliation |
|---|---|
| #1831; #1832–#1835 milestones | Open. Preserve outcome tracking; do not close based on a single pilot or merged foundation |
| #1836/#1837, #1840/#1841 | Open historical implementation tickets with substantial merged work; reconcile against current separate-family architecture, not obsolete embedded-schema wording |
| #1838/#1842 | Open consumer migration obligations; the consumer matrix explains the remaining paths |
| #1839/#1843/#1850 | Open reconciliation/fleet obligations; #3371 controls and merged #3253 fleet operations still need scheduled-operation and acceptance proof |
| #1844/#1847/#1848 | Open historical storage/snapshot/retention outcomes, partly supplied by foundation; operating schedule and retention evidence still required |
| #1845/#1849 | Basic history API/UI versus richer visualization; exact existing area-chart gaps and phased dependencies are in the chart rollout plan |
| #1846 | Historical optional storage refactor; current family/block architecture supersedes the old embedded-to-single-document proposal. Reconcile issue wording rather than implement the obsolete design |
| #3145, #3185, #3085 | Closed; rolling-reader class maps (#3512), transactional admission (#3232), legacy project authorization (#3340) are delivered |
| #3361 | Open, but ReviewController now checks Pages before optional materialized substitution. Reconcile/close the delivered guard after review; do not list it as missing code |
| #3197 | Open follow-up; critical post-commit dispatch and question binding are delivered by #3234. Remaining hygiene is not those original defects |
| #3190 | Open; committed local two-family proof exists, broader staging proof remains |
| #3364 / #3474 / #3506 / #3370 | Open reviewer/search/version/configuration follow-ups; distinguish correctness blockers for the affected activation from optional cleanup |
| #3141/#3143/#3147, #3187/#3188 | Open foundation/test/retention follow-ups; do not make optional test hygiene block the screening MVP without a demonstrated correctness or acceptance gap |
| #3510 / #3511 | Open staging and live-update evidence. Bodies predate pilot activation and notification implementation; remaining acceptance is described above |
| #2504 / PR #2469 | Open chart refactoring; separate presentation work, not proof of materialized reads |
| #3083 / PR #3514 | #3514 merged; preserve the name-visibility permission boundary during activation |
| #3189 / #3539 | Open CI/test reliability evidence; record failures honestly without reclassifying them as missing statistics families |
Merged milestones commonly missed by older records include #3240 search population, #3297 transactional reviewer maintenance, #3243 export authorization, #3505 stage-annotation baseline, #3534 notifications, #3547 stage progress and #3551 overview/eligibility fixes. Closed-unmerged candidate PRs #2985, #2534, #3296, #3298 and #3300 are superseded evidence, not five additional deliveries still to implement.
Broad runtime flag architecture (#3527), flag UX (#3553), allocation redesign and screening-policy proposal (#3552) remain separate workstreams. Source-row exports, speculative new reports, richer historical charts and optional cleanup should not expand the first usable screening release. Basic history, approved family consumers, correctness and explicit acceptance gates must not be silently removed from the full programme.
Implementation resumed — 22 September 2026¶
The owner requested implementation of phases 0–3 followed by 4–6, using parallel independent work. The dated main/deployment inventory above remains a baseline, not a claim these new branches are deployed.
| Slice | Current implementation evidence | Remaining boundary |
|---|---|---|
| Membership-stage annotation baseline #3561 | Runtime calculator/backfill routes; focused Mongo, API and PM-host tests pass | MERGED 2026-09-23 00:26 UTC (a80f65528). Family flag stays off; consumer migration and activation separate |
| Question-answer baseline #3562 | Shared authoritative tally pipeline, pinned runtime calculator, administrative baseline/rebuild; focused database/API/host tests pass | MERGED 2026-09-23 01:03 UTC (c30ec96e6). Family flag stays off; consumer migration and activation separate |
| Daily observation recovery #3563 | Capture/publication executor, default-off Quartz daily/startup scheduling and bounded same-day retry implemented; integrated real-Mongo test publishes all nine physical families | Deployed schedule, restart/midnight behavior and retention acceptance remain unproven |
| Domain reconciliation baseline #3564 | Stage-grain pinned calculator/backfill, 3 Mongo and 76 API tests pass | MERGED 2026-09-23 02:12 UTC (a44db615f). Family flag stays off; consumer migration and activation separate |
| Reviewer annotation baseline #3567 | Stage-scoped reviewer calculator/rebuild implemented with distinct reviewer target/tracking semantics; combined 721 Core tests pass | MERGED 2026-09-23 18:38 UTC (6e72eb37b515523207006c2787bbe2391af38f95). Family flag (materializedProjectStatisticsMembershipAnnotation) stays off; consumer migration and activation separate |
| Observed screening history #3568 | Bounded graph-authorized endpoint and existing Overview chart use actual checkpoint times and explicit gaps; 10 endpoint and 10 Angular tests pass | Daily capture deployment and browser acceptance remain unproven |
| Acceptance evidence validator #3569 | 13 offline tests cover receipts, elapsed span and conclusive parity; rebuild revisions cannot count as mutations | Does not supply missing elapsed soak, storage or performance evidence |
| Own reviewer current progress #3570 | Independent default-off consumer, coherent current/source reads and own-scope authorization; 723 Core, 7 Mongo, 79 API, 11 host and 17 frontend tests pass | Staging activation and sustained materialized-read performance separate |
| Question designer counts #3572 | Snapshot-covered counts and component-local signals preserve drafts; 52 Core, 4 Mongo, 49 API and 36 Angular tests pass | Non-pilot editing and unavailable-count locking are covered; activation remains separate |
| Stage observed history #3573 | Existing screening location reuses #3568; annotation history uses retained definitions and explicit gaps; 13 endpoint and 16 Angular tests pass | Merged; deployment and browser acceptance remain separate |
| Reviewer screening observed history #3574 | Existing reviewer dialog implemented with membership scope, preserved graph/decision permissions and signed availability | 144 API, 48 Core and 18 Angular checks passed; activation remains separate |
| Reviewer annotation observed history #3578 | Paired retained stage/member observations, graph/peer authorization and signed availability; combined 65 API / 31 Angular plus 49 flag tests pass | Activation and deployed evidence remain separate |
| Fleet operations #3253 / retention #3263 | Multi-family dispatch and bounded run retention updated; rolling-schema unknown fields preserved and tested | Deployment, fleet operation and storage acceptance separate |
Combined consumer regression through #3572 passes 746 Core, 636 Mongo, 189 API and 154 frontend tests, plus the final host-registration suites (20 API / 12 worker). Eight opt-in benchmarks were skipped; the deterministic screening-write command budget was included. These checks exercise combined flags, history contracts and consumer state rather than only each isolated endpoint. The reviewer-progress runtime flag propagation gap found during integration is fixed. Fleet operations are integrated above the consumer stack, with all nine physical families covered by dispatch tests (5 Core, 18 Mongo, 92 API and 13 worker-registration tests pass). Retention integration passes 22 Mongo and 95 API tests, including terminal-only pruning and preservation of newer schema fields. Generated contracts are rebuilt from the integrated API.
Combined baseline regression on the stack through #3564 passed 720 Core and 654 Mongo tests (with 8 opt-in benchmarks skipped). Later #3567 fixes the stage-scope fixture and passes 721 Core tests. These are local regression checks, not an assertion that stacked PR CI, staging soak or write-overhead acceptance has passed.
A fresh targeted rebuild ran 246 existing screening/annotation/question Core tests successfully against
PR #3551; statistics Core source/tests matched main a89d1f136 at that check. This is regression evidence,
not the missing write-overhead acceptance, sustained soak or cross-pod SignalR proof. Fleet continuation
reuses existing #3253/#3263 instead of creating a duplicate. The implementation stack contains reviewable slices; further consumer and maintenance work is recorded below.
No merge, production activation or destructive retirement is implied by this status update.
Compatibility FullStats endpoints retain their authoritative paths; source-row exports retain their existing contract. The seven-day
soak, read/mutation evidence targets, write-overhead budget, deployed scheduler and cross-pod browser
checks, storage/rollback evidence and retirement gates remain open. These are not satisfied by the
local regression suites above.
Further consumer and maintenance implementation¶
The table mixes merged code and open follow-up PRs. None of its local test evidence establishes deployment, activation or the separate acceptance gates above. Check each linked PR state before use.
| Slice | Local implementation evidence | Boundary |
|---|---|---|
| Coherent Stage Overview #3580 | One pinned graph/roster/definition snapshot or whole-response guarded fallback; component-local signals and route guards prevent broad eager loads. 12 Mongo, 751 Core, 102 API, 13 host and 40 Angular tests passed; subsequent flag/startup coverage passed | Independent default-off consumer; permissions and legacy compatibility retained |
| Screening Information #3586 | Project-wide screening and member charts use the same snapshot contract without requiring annotation or a stage. 15 Mongo, 122 API, 6 Core, 13 host and 49 Angular tests passed | Independent default-off consumer; SignalR, reconnect, polling, cancellation and failure clearing covered locally |
| Own project reviewer progress #3583 | One coherent screening/all-stage response with caller and stage-definition metadata; Project Overview local signals handle live updates. 17 Mongo, 25 controller, 76 flag/runtime, 13 host, 6 Core and 7 Angular tests passed | Integrated with the consumer stack; ordinary and inactive stages retain legacy response shape |
| Search population consumers #3587 | Existing project-details and search-list reads preserve imported-file population semantics and authoritative shared-search links. 7 Mongo, 150 API, 13 host and 4 Angular tests passed | This replaces metadata-count composition; it is not a claim of avoiding a previously expensive Study aggregation |
| Live question assignment locks #3582 | Component-local question counts update assignment locks without replacing drafts; 10 Angular tests and focused lint pass | Complements designer migration #3572; integrated consumer regression passes |
| Bounded checkpoint-history maintenance #3581 | Transactional admission fences, bounded retained-root scans, reader grace, reference-safe page/observation deletion and cursor progression; 24 Mongo, 99 API and 751 Core checks passed | Default off and allowlisted; no live deletion performed |
| Checkpoint build marker reclamation #3674 | History maintenance deletes terminal markers whose root is gone and which own no page or creator observation, inside the same admission-CAS pass, so daily builds keep admitting past the 512-marker ceiling (552 daily builds proven on real Mongo; red at day 513 without it). Unused unbounded retention, delta-compaction and receipt-floor loops removed | Default off and allowlisted with history maintenance; no live deletion performed |
| Bounded delta-ledger maintenance #3584 | Atomic complete-revision compaction, bounded pin inspection, source/publication guards and replay-floor advancement; 12 dedicated Mongo, 14 existing reclamation, 103 API and 31 Core tests passed | Source receipts remain a separate protocol and are not deleted by this operation |
| Source-receipt maintenance #3588 | Merged and deployed dark: a durable per-Study/namespace retirement watermark guards eligible receipt cleanup; rolling-reader BSON compatibility is gated by the default-off maintenance control | No receipt cleanup or activation evidenced; trusted point receipts only. Child/other-namespace cleanup is a separate protocol |
| Scheduled bounded maintenance #3717 | Default-off PM Quartz job calls the history, delta and (behind its own switch) receipt services for allowlisted projects only, within per-run project and per-project pass budgets, with a persisted rotation and history cursor; typed refusals and failures are isolated per project and retried next run. 420 simulated daily builds with daily scheduled runs stay under the 256-root cap and keep admitting on real Mongo | Default off (ProjectStatisticsMaintenance:Enabled); each service keeps its own gate; no environment configured |
| Periodic drift check #3636 part 3 (#3726) | Default-off PM Quartz job recalculates every allowlisted project from source regardless of the activity clocks, within per-run project and per-project scope budgets and a persisted cursor. It compares each scope with the reader's served row at one captured identity; a pass confirms Unconfirmed snapshots bracketed by two passing checks, a failure makes the drifted row Stale (advancing the source clock, which ends any unchanged-day streak) and marks the Unconfirmed snapshots since the last pass Suspect; Stale rows are republished by an administrative backfill/rebuild, not automatically. Real-Mongo tests cover bracketing, suspect marking, stale fallback and rebuild, the budget cursor, a racing write, unpinned calculations, late-enumerated scopes, epoch-only advances and vacuous checks | Default off (ProjectStatisticsDriftCheck:Enabled); weekly default cadence; no environment configured; history is never rewritten |
| Scheduled stale-statistics repair (C1, #3254, #3731) | Default-off PM Quartz job (hourly default) inspects each allowlisted project inside one pinned source snapshot through the reader's serving gate (skipping unpinned reads) and republishes families whose rows are Stale, Missing or epoch-mismatched through each family's ordinary non-forced backfill, within per-run project, repair and scope budgets and a persisted rotation. Fences, staged operations, quarantine and closed gates are skipped; configuration mismatches are refused, never forced; an occupied bootstrap identity is reported typed with an operator remedy; both are then backed off per (project, family) until their identity moves, and a default-on settle check repairs only after a run interval without writes. Real-Mongo tests cover drift-check- and definition-rewrite-marked repair to served, fences, digest refusal and backoff, settle, disabled, rotation/deferral/starvation, empty and missing families, failure isolation, a concurrent source advance, unpinned refusal and the bootstrap limitation (25 focused real-Mongo tests) | Default off (ProjectStatisticsRepair:Enabled); no environment configured; calls the backfill directly rather than the fleet runner (reasons in the fleet runbook) |
| Review/reconciliation route cutover #3590 | Merged: guarded review/current/completed and reconciliation routes skip unused broad FullStats; nested navigation and rollback are covered | Routes requiring FullStats and flag-off loading retain legacy behavior; deployment/flag activation and acceptance remain open |
| Historical chart page limit #3600 | Four checkpoint-backed charts stop cumulative requests at the API's 50-point cap; 41 Angular tests across eight files, lint and generated-contract validation pass at combined tip fc922d698 |
Prevents the third “Show more” click from replacing a chart with HTTP 400; wider cursor pagination remains a separate enhancement |
| History continuation UI #3774 | In review: four charts request one 20-checkpoint continuation page per “Show more” action, retain loaded history across first-page refreshes and reset on scope change | API limit remains 50 per request; deployed browser acceptance remains open |
Source inspection found no active frontend invocation of SubscribeToProjectFullStats or
SubscribeToFullStatsForProjectAndInvestigator; the old listener and server methods remain for
compatibility. The active SubscribeToProject stream still carries entity metadata, including searches.
Removing those compatibility methods is deferred retirement, not a prerequisite for the new view-local
statistics consumers. No aggregate reporting product is invented from the Exports flag.
The additional consumer stack passes 187 API, 39 Mongo, 6 Core startup, 13 worker and 54 Angular checks, plus 10 assignment-lock checks and 39 narrow route checks. Checkpoint maintenance integration passes 236 API, 38 consumer/maintenance Mongo and 95 index-contract checks. Delta-maintenance integration passes 240 API and 127 Mongo checks. Generated contracts are regenerated from the combined API rather than edited by hand. Final receipt-maintenance integration passes 751 Core, 925 Mongo and 521 API tests, with eight explicitly opt-in benchmark tests skipped. Generated contracts and documentation validation pass; skipped benchmarks remain acceptance work.
The receipt protocol adds two indexed reads to the screening write command budget (23 rather than 21); this does not establish the wall-time overhead target. After any receipt reclamation, all mutation hosts must retain the retirement-aware writer protocol. Disabling cleanup is safe; reverting to older writer binaries without coordinated receipt/database restoration is not a supported rollback. No such cleanup or deployment was performed during implementation.
Two final rollback corrections preserve the usability of the independently reversible consumers: #3591 remains open and requests fresh legacy FullStats or own-reviewer statistics once instead of restoring an old cached success, without repeating broad reads during disabled-response retries. Open #3766 adds a separate settings-only refresh path after stage settings saves. #3594 preserves conservative question edit locks after live count evidence is disabled; it clears displayed counts, respects authenticated project identity and leaves drafts untouched. Initial legacy behavior remains unchanged. The question change passes 58 designer/assignment/state tests. The combined frontend tip passes 377 tests across 47 files, including project and stage overviews, screening, question management, route guards and statistics invalidation. Generated-contract validation passes at the combined tip.
The notification implementation has a two-queue RabbitMQ integration test for separate API-pod consumers and multiple viewer sockets; Angular tests cover resubscription, revision deduplication and bounded recovery polling. That is local transport evidence. The 22 and 28 September staging observations each had one ready API replica. A multi-pod deployment, browser reconnect/revocation and full operational acceptance have not been verified in this record.
The daily checkpoint producer and Quartz schedule/retry merged in #3563. Actual deployed scheduling and captured daily observations still need staging evidence.
Ordered follow-up boundaries remain explicit:
- Review and merge the remaining open slices, deploy them dark, and verify writer protocol/index coordination before requesting activation. Earlier slices have merged; this record does not establish their deployment or activation.
- Complete the staging live-update and eligibility proof tracked by #3510/#3511, including browser, scheduler, parity, rollback and cross-API-host checks. Then collect the seven-day soak, read/mutation, write-latency and storage evidence before a production decision.
- Under the remaining #3254 retention work, specify safe child/other-namespace receipt retirement before extending cleanup beyond the delivered trusted point-operation slice. Legacy unproven receipts remain retained; their deletion is not justified by a time-only floor.
- Retire legacy endpoints, push methods, manual tally mechanisms and indexes only in separate PRs after all consumers have passed acceptance and the rollback window has closed.