Skip to content

Materialized Project Statistics

SyRF will provide fast, authorized current and historical project statistics without repeatedly running the complete authoritative MongoDB aggregate on ordinary reads. Materialized values remain disposable, versioned read projections; Project and Study remain the source of truth and every current consumer can fall back safely to the approved live calculation.

Start here: the statistics reference explains, in plain English, every materialized statistic as built: what it counts, how it is rebuilt, what keeps it current, what makes it stale, and the known gaps.

The current programme status records implemented families, consumer migrations, remaining work and dated deployment/acceptance evidence. The phase 0–3 audit and historical chart rollout make the current implementation gaps and chart dependencies explicit. The September 2026 recovery ledger is historical and must not be used as the current readiness record.

The detailed architecture and phased delivery gates are defined in the technical plan. The recovered-design reconciliation records every material difference from the March-April 2026 plans and M008-M011, with an explicit disposition. The async point-fold design (planned, 30 September 2026) replaces the same-transaction point save for every screening- and annotation-dependent family with a single-document pending entry and a leased fold.

User outcomes

  • Project users can load approved screening and annotation statistics without the latency of a full project-wide recalculation on every ordinary read.
  • Users can understand meaningful distributions of Study workflow state, rather than seeing only simple totals.
  • Authorized reviewers and project managers can request equivalent membership/reviewer breakdowns without exposing peer data they cannot currently access.
  • Users can request durable historical checkpoints with clear calculation time, source provenance and schema version.
  • A stale, rebuilding, missing or incompatible current projection falls back to the current authoritative calculation; historical APIs never fabricate an old point from current data.
  • Operators can disable writes or serving globally, per family, per consumer or per pilot project without deleting projection data.

State-profile requirement

The primary product statistic is a bounded, catalogue-defined distribution over meaningful current Study states. It is not an arbitrary analytics cross-product.

Approved profile families will cover:

  • project screening combinations such as reviewer include/exclude decisions, screening-decision number/status, sufficiency, started and overscreening classifications;
  • authorized membership/reviewer screening profiles;
  • project/stage annotation-session combinations such as no session, incomplete/in-progress and completed; and
  • authorized membership/reviewer-stage annotation profiles, including availability and reconciliation state where the current live calculation defines them.

A Study may validly contain include and exclude decisions from different reviewers. Project profiles must preserve and tally these conflicting reviewer-decision combinations. A profile is excluded as impossible only when an enforced application invariant proves it cannot occur. Every profile key, formula, permitted overlap, ordering rule and authorization class must mirror an approved current live calculation and be versioned in the metric catalogue.

MVP boundary

The minimum independently useful runtime delivery is:

  1. a dark shared projection/history foundation with transactional signed deltas, immutable operation history and fail-safe current fallback;
  2. project screening state-profile materialization with parity and performance evidence; and
  3. one screening-only API/Project Overview consumer behind an independently reversible flag.

The first consumer must not request broad FullStats, because that bundle also requires annotation and membership families and would correctly fall back until they exist.

Page consumers and their flags

Each page consumer is independently reversible and additionally requires the materializedProjectStatisticsPages kill switch. All default to false.

Consumer Dedicated flag Families required Contract
Project Overview screening totals materializedProjectStatisticsProjectOverview project screening screening-overview-cutover.md
Stage Overview annotation pie materializedProjectStatisticsStageOverview stage annotation stage-overview-cutover.md

Both families now have an administrative maintenance half. The stage-annotation baseline is established through the family-routed trigger, POST api/admin/project-statistics/{id}/stage-annotation/backfill and /rebuild, on the same BatchAdminProjects policy and the same gates as every other family; the runbook's step 4b is the sequence. Every other physical family now has the same /backfill and /rebuild pair; see the statistics reference for the full list.

Requirements

Correctness and safety

  • Materialized values never become domain truth and can always be rebuilt from Project and Study.
  • A current response is evaluated against one project committed revision or wholly from the authoritative calculation. An unchanged Fresh scope may retain an older last-changed revision.
  • Supported same-MongoDB point mutations update authoritative source, affected materialized scopes and an immutable unique delta/history record in one transaction.
  • Family guards select a visible generation; each scope selects its newest eligible published row at/below that generation without compatibility filtering, and evaluates compatibility only after selection, so an incompatible newest selected row forces whole-bundle authoritative fallback and is never skipped in favour of a superseded older row. Bounded bulk candidates remain isolated while unchanged older scopes stay visible. A constant-size guard flip assigns a distinct public revision and replacement provenance.
  • Bulk, cross-system, non-additive and unsupported mutations fence affected scopes until a rebuild or bounded bulk-delta publication completes.
  • Large imports remain wholly hidden from authoritative queries behind one Project-level visibility token, while a separate admission lock blocks overlap until every affected family publishes or a verified terminal abort releases the gates. A publication-capacity abort marks the affected families Stale and deliberately reveals already committed source only through authoritative fallback; a source- loading abort instead proves and removes the hidden prefix before release. Bounded Study cleanup never becomes the visibility mechanism.
  • Historical checkpoint-set roots publish a complete compatible bundle through BSON-bounded reference pages of immutable real scope observations. Unchanged scopes may reuse an earlier observation reference; missing or incompatible history is explicitly unavailable and is never reconstructed from current state.
  • Current and historical reads apply current authorization, including own-row and peer-row rules.
  • The screening and stage annotation statistics and history endpoints, the project-statistics parity audit, the reconciliation submission re-check, and the client permission report evaluate the caller's effective application groups: the syrf_groups claim plus the application roles on their investigator record, the same union the authorization handlers use. A role-only grant is therefore neither refused by these consumers after the route policy admitted it nor hidden by the permission report. Other statistics-adjacent consumers, including the stage review settings Design checks and evidence reader, still use claim groups only and remain part of #3642.
  • Source capture, retries, redelivery, concurrent writers, rebuild leases and publication are idempotent and fail closed.

Screening leaderboard visibility

The per-reviewer screening leaderboard is shaped for the caller who asked for it (#3083). Every legacy FullStats producer goes through one shaper, and the materialized read applies the matching metric gates. The two permissions the screening visibility dialog already configures decide the tier:

ViewScreeningProgressGraph ViewScreeningProgressGraphDecisions Legacy FullStats (web leaderboard) Materialized peer rows
no either No membership screening rows at all, and no screening graph. None.
yes no Every row's counts, with no reviewer id, under a displayLabel of "Reviewer A", "Reviewer B" and so on assigned per request in a random order. The caller's own row carries isSelf, so a client can mark it without identifying anyone. None (own row only).
yes yes Rows exactly as computed, reviewer ids included. Served only if the caller also holds ViewMemberships; otherwise own row only.

The materialized filter's extra ViewMemberships requirement for peer rows is deliberate and not mirrored in the legacy shaper. The default catalogue grants the decisions permission to every project member but ViewMemberships only to project administrators and the owner, so requiring it on the legacy path would take the named leaderboard away from ordinary members of every default project. The two rules differ only in that the materialized path withholds more; neither discloses anything the other withholds (#3642).

Counts are identical in every tier that returns rows: shaping changes who a row is attributed to, never what it counts. The anonymised label is neither stable across requests nor derivable from any identifier, and a peer row's membership id is replaced by a per-request surrogate, so two responses cannot be correlated to re-identify a reviewer. The caller's own row keeps the caller's real membership id (they already hold it), so the web client can still show them their own progress.

The anonymisation is weak in very small projects. A member knows who else is in the project, and their own row is marked, so in a project with two reviewers the other row is the other reviewer, and with three the counts may narrow it down further. The leaderboard's membershipIds is the project's membership list, which the same response already carries in its (unrestricted) annotation section, so withholding it would not change this. Where reviewer identity matters in a small project, withhold the graph permission rather than relying on the letters.

The legacy SignalR push captures the subscriber's application groups when it subscribes, re-reads the project on every emission, and treats the caller's tier as part of what makes an emission new. A change to the project's screening-graph permissions therefore reaches the subscriber on the next project notification even when no statistic changed; a change to the caller's application groups or investigator roles takes effect on resubscription. Both paths evaluate the same group set as the authorization handlers: the syrf_groups claim plus the investigator's application roles.

The materialized path has no anonymous form of an identified scope, so it withholds a peer's reviewer-grained screening row from a caller who lacks the decisions permission rather than serving it anonymously.

History and provenance

  • Checkpoints record project, metric/scope keys, catalogue/schema/source versions, source watermark, observed/calculated time, trigger and operation/event provenance.
  • Current/delta publication and its durable revision/scope notification outbox row commit together; a fixed-size checkpoint root names one BuildToken and atomically publishes BSON-bounded immutable reference pages/observations only after epoch, lifecycle and fence guards pass.
  • Retention is bounded so compaction cannot remove delta provenance required by surviving checkpoints, retries, rebuilds or audits.
  • Exact delta identities and moves have measured replay/idempotency floors and a hard ceiling; safe old ranges compact into bounded seals, while unsafe pressure disables materialized writes and falls back.
  • Bootstrap records one truthful observation at enablement and never invents earlier history.

Rollout and evidence

  • All serving starts disabled and is controlled by durable epoch-aware global/project safety gates plus family, consumer and project-pilot flags; cached process flags are not correctness authorities.
  • Shadow parity compares materialized and authoritative calculations before any consumer cutover.
  • Each family proves exact integer parity, safe fallback, mutation/invalidation coverage, authorization, history behavior and bounded storage.
  • Each consumer demonstrates a material read-p95 improvement and authoritative aggregation reduction against named reproducible datasets before activation.

Explicit exclusions

  • Unit-level statistics/materialization: an annotation unit is scoped to one Study and reviewer annotation; no useful cross-study aggregate or performance need has been demonstrated.
  • Outcome-level statistics/materialization: no stable authoritative cross-study aggregate, authorized consumer or measured performance need has been demonstrated, so outcomes are not represented in the programme's scope key or metric catalogue.
  • Arbitrary profile dimensions: clients cannot construct unbounded cross-products or invent formulas.
  • Broad agreement/kappa: PR #2534's cross-stage calculation is invalid; any future measure requires a separately approved same-stage cohort and denominator.
  • Operational progress/presence: import/export job state, Bulk PDF progress and live connections remain in their operational models.
  • Runtime implementation or rollout: this planning PR creates no product code, migration, staging or production activation.

Unit or outcome materialization may return only through a future separately approved feature with a concrete useful aggregate, authoritative definition, consumer and measured performance need.

Delivery and approval gates

This brief and its technical plan received architecture-review approval, and all phases were approved for implementation on 2026-09-02. Phase 0 completed with the exhaustive metric/consumer catalogue, method-level mutation ownership, profile formulas, fixed benchmark datasets and executable commands (#3070, #3076).

The shared foundation and multiple family/consumer slices are merged. The authorized single-project screening pilot is configured on staging; this programme is no longer wholly dark. See the current status for the exact family, consumer and deployment boundaries, including the remaining operational adapters and legacy consumers. Default-off flags are defaults, not a claim about current staging configuration.

Merging a slice does not satisfy its activation gate. The staging-proof runbook requires at least seven actual days, 10,000 reads and 1,000 mutations before any production pilot. Performance, production pilot, wider rollout and legacy retirement retain their separate technical-plan gates.

Telemetry exists independently of activation. Both hosts register the SyRF.ProjectManagement.ProjectStatistics meter and record reads by source and fallback reason, confirmed write commits by path, rebuild/backfill outcomes and parity results. Startup logging reports configuration and allowlist size; JSON logging is also implemented. These facilities do not establish an operating collector, dashboard or retained soak evidence. See the runbook's Telemetry section for instruments and the current status audit for the remaining operational verification.

Pilot activation's #3185 dependency is resolved. #3232 closed #3185: capacity-sensitive source writes (assignment, screening and annotation capacity saves) now read the durable reviewer-tracking mode and commit their statistics delta in the same snapshot transaction as the source write, so a Project.AgreementThreshold write can no longer commit between a rebuild's pinned snapshot and its publication without advancing the control's source revision. The remaining pilot gates are the Phase 2C staging proof above.

Success criteria

  • Approved profile distributions exactly match authoritative live calculations, including valid mixed include/exclude reviewer decisions.
  • No stale, rebuilding, incompatible or pending-event scope is served as Fresh.
  • Historical unavailability is explicit and never substituted with current data.
  • The first screening-only consumer executes at least 80% fewer authoritative statistics aggregations and improves response p95 by at least 20% on the approved benchmark.
  • Supported source-mutation p95 regresses by less than 10%; realistic ½/5/10-reviewer benchmarks show acceptable retry and write-conflict rates. A small shared summary is allowed unless measurements justify splitting or striped counters.
  • Rollback to authoritative reads is immediate and does not require data deletion.
  • Epic #1831 — pre-calculated statistics programme
  • PR #2534 — original broad design evidence; stale implementation is not resumed
  • PR #2985 — screening-only implementation evidence to reconcile after shared contracts are approved
  • FEAT-006 — domain reconciliation compatibility input
  • FEAT-009 — screening annotations compatibility input
  • FEAT-013 — export/report consumers enabled by the shared query contract

Compatibility invariants

ProjectStatisticsConfigurationDigest.ForThresholds owns the versioned project-wide projection configuration identity. Every metric family sharing a control row must use that authority; do not introduce family-specific control digests. Extending its inputs or format requires an explicit compatibility/reconciliation plan for established controls, current rows and retained checkpoints before serving is re-enabled. An automatic rebuild refuses an established incompatible control digest and returns the scope to Stale; it has no authority to decide the control is the stale side. The administrative forced rebuild (POST api/admin/project-statistics/{id}/rebuild) is the sanctioned reconciliation for the control and its current rows: it adopts the authoritative identity inside the publication's own control compare-and-swap, so a concurrent source write defeats it rather than being overwritten. That reconciliation is deliberately partial. Retained history checkpoints keep the identity they were captured under and are read through their own captured settings; they are never rewritten. So a change to the digest's inputs or format still requires the explicit migration plan above — the forced rebuild is one step in it, not a substitute for it.

Durable source-operation identities are a separate protocol. Inclusion-recalculation operation IDs retain the original screening.v1 threshold representation so an in-flight fence can resume across deployments. Never derive those job tokens from the evolving shared projection digest or change their bytes without an explicit resume migration.

Every host that Lamar-scans SyRF.ProjectManagement.Core registers the statistics seams through ProjectStatisticsRegistry and ProjectStatisticsLifecycleRegistry; Lamar resolves the last registration per type, so overrides must be registered after them — ProjectStatisticsProductionRegistry first, then ProjectStatisticsFamilyCalculatorRegistry, which replaces the fail-closed placeholder authoritative calculator with a router over the real per-family scope calculators.