Skip to content

Bulk PDF durable capture: merge acceptance and remaining activation gates

PR #3094 adds the disabled-by-default capture and ordinary terminal-cleanup slice on top of the authority contract in #3042. Its usable boundary is an immutable, correctly bound object: capture survives duplicate events, publication consults Project Management, terminal cleanup retains the exact version until its retention deadline, and acknowledgement retires the hold through a durable outbox. Publication pause permits cleanup to continue.

Acceptance for this slice:

  • An exact-binding duplicate preserves the capture; conflicting binding data poisons it.
  • An expired lease can be recovered without replacing a newer durable state from a stale index read.
  • A cancel or abandon that lands after capture but before the hand-off (job Aborting) is handed to Project Management's cleanup-only object-ready branch even while publication is paused; it terminalises the job with durable cleanup-only authority in one transaction and sends no agent work, so the job cannot deadlock in Aborting (staging upload 5ba76ac4, 2026-09-25). Covered by the reconciler unit tests and the Bulk PDF lane's cancel-while-paused E2E case.
  • A terminal object is never published to the agent. A retriable claim remains held for as long as the job's retry window is open (BulkPdfUploadJob.RetryAuthorityWindow, 30 days from the terminal instant); once it lapses the claim is gone and cleanup proceeds.
  • Cleanup uses the exact bucket/key/version and preserves the first retention deadline.
  • Capture retirement and hold-retirement outbox creation are one conditional transaction; the outbox is deleted only after its acknowledgement is recorded.
  • Each inventory invocation reads one bounded version page and checkpoints only after all its captures succeed. Overlapping invocations cannot overwrite a newer checkpoint. Due cleanup and outbox work run before inventory, so an inventory outage cannot starve them.
  • Missing/malformed metadata and metadata naming another environment remain non-publishable under the configured environment's reserved-object boundary.
  • The reconciler requests the scopes issued by Identity, and its role permits the configured failure-queue send. Existing chart defaults render no new durable resources.

Validation includes notifier unit and LocalStack integration tests, chart unit tests, chart-change detection, and the PR's current-commit CI/reviews. The PR description records the final counts.

Ordered follow-ups owned by #3168

  1. Complete rejected-object cleanup and quarantine adoption. Done (PR #3558). Mechanism: the authenticated resolver now also discloses the server-owned release for a single unambiguous match, because a rejected binding derives its release id from the object coordinates alone and matches nothing in Project Management. Zero matches delete the exact version after the snapshotted invalid-object retention and retire the row by a conditional delete covering every immutable binding field, gated on a completed inventory checkpoint later than that deletion plus an exact-version absence read; no acknowledgement, outbox or tombstone is fabricated. One match pins the disclosed release on the row before any authority is requested for it, then registers cleanup-only authority and joins the ordinary retention path without ever publishing; a pinned row resumes from its release rather than resolving again. Many matches, or resolver evidence disagreeing with the digest recorded on the row, register a durable quarantine fence under a capture-key-derived fence-set id before custody changes; the fence route names no release and Project Management re-resolves the candidates, adopting only unowned ones and failing closed on any held elsewhere. An unavailable resolver, a lost response or a crash re-drives from durable state, and a competing worker loses the conditional write. Covered by notifier unit tests, two LocalStack retirement tests and the Project Management/API contract tests.
  2. Complete tombstone replay/queue/inventory retirement. Done (PR #3575). Tombstones still have no TTL; they now leave the active partition of the due index for a TOMBSTONE one and become due at exactly their snapshotted replay deadline, so the existing StateBucket/DueAtEpoch index serves the purge pass and no new index was added. Three facts are each recorded on the row by a row-version-conditioned write before an exact conditional DeleteItem names every one of them alongside every immutable binding field: the snapshotted deadline (read, never recomputed); the failure queue, proved by two observations of the visible, in-flight and delayed counters, all zero, at least FailureQueueQuietPeriod (PT15M) apart and both after the deadline — GetQueueAttributes has no ApproximateAgeOfOldestMessage, which is a CloudWatch metric, so the first empty observation is the watermark and a delayed redrive landing between the samples is seen and withdraws the whole proof, inventory evidence included; and a completed inventory checkpoint newer than both the deadline and the queue proof, plus an exact-version read that still reports the object gone. An unconfigured or unavailable queue, a queue that omits a counter, a still-unretired hold outbox, a truncated or older checkpoint, a version that is present again, and a lost conditional write all retain the row and retry; no age or TTL path can delete one. A QuarantineTransferTombstone is declared and refused by name — ADR-015 (lines 534-538) adds destination cleanup acknowledgement, fence-set clearance and shared-debit release on top of this fence, and follow-up 4 owns them. Covered by the notifier purge unit tests and three LocalStack tests (exact conditional purge, due visibility, real SQS attribute reads).
  3. Schedule and validate receipt-drain reconciliation. Done (#3556). BulkPdfReceiptDrain now carries a bounded producer beside the consumer: BulkPdfReceiptDrainBackgroundService runs one BulkPdfReceiptDrainProducer pass per BulkPdfReceiptDrain:Interval (default 10 minutes, BatchSize 100, both validated on start), and each pass discovers retained receipts through IBulkPdfUploadReleaseRecordRepository.GetDrainCandidatesAsync — CleanupReceipt records whose hold retirement is acknowledged and whose drain is either unproved or proved and past ReceiptExpiresAt — oldest first, then sends one IReconcileBulkPdfReceiptDrainCommand per record over the existing command queue. The producer decides nothing: it never fabricates an outbox key and never records an acknowledgement, so exact-key outbox absence, history pruning and receipt expiry are still proved together inside the consumer before a purge. A receipt whose outbox row still exists is left untouched and rediscovered on the next pass, and a pass runs to completion before the interval delay, so the retry cadence is the interval. One faulting candidate does not abandon the rest of its page. Proved by BulkPdfReceiptDrainProducerTests (consumer and producer wired over the MassTransit harness) and the GetDrainCandidatesAsync repository tests. Still feature-gated: BulkPdfReceiptDrain:Enabled is off by default and a disabled pass queries and sends nothing. Chart wiring for this and for the receipt-retention settings on BulkPdfNotifierAuthorityOptions landed in #3560 (bulkPdfReceiptDrain.* and bulkPdfNotifierAuthority.receiptPolicyVersion / receiptRetentionDays.*), so cluster-gitops can now set them; no activation-time environment values changed as part of that PR.
  4. Complete quarantine transfer, preview provisioning and the separately approved rollout evidence in ADR-015. These are later slices, not properties established by merging #3094.

Activation gates owned outside this chart

  • Setup-role IAM grant (blocking). durableCapture.reconciler.setupRoleAuthorized is false by default and gates the reconciler's lambda add-permission (Sync hook) and lambda update-function-configuration (PostSync hook). Both hook Jobs run as the external syrf-ack-setup-job role, whose Lambda mutation scope is the exact ingress function ARNs (technical plan, syrf-ack-setup-job). Infrastructure must add the exact reconciler function ARN to that role before the flag is flipped; until then the reconciler stays unconfigured and its EventBridge rule stays DISABLED (even with durableCapture.scheduleEnabled: true, see below), and the chart issues no call that would fail the sync. Once authorized, both Jobs first poll lambda get-function-configuration (the only read the setup role holds on the reconciler) until the reconciler reports Active with no update in progress, bounded by setupJob.timeouts.lambdaActive. On the first sync ACK may not have created the function yet, and an unguarded call would fail the hook with ResourceNotFoundException. Covered by .chart/tests/reconciler_setup_wait_test.yaml.
  • Schedule switch. durableCapture.scheduleEnabled (default false) is the only way the EventBridge rule renders ENABLED, and only when durableCapture.enabled and reconciler.setupRoleAuthorized are also true. Enable it in a separate cluster-gitops change after the setup Jobs have configured the reconciler. Covered by .chart/tests/durable_capture_schedule_test.yaml (all eight gate combinations).
  • Hold-table point-in-time recovery is a per-environment values decision: durableCapture.pointInTimeRecovery (default false) is always rendered as an explicit spec.continuousBackups.pointInTimeRecoveryEnabled boolean on the ACK Table. An omitted field would be unmanaged by ACK, so recovery could be turned on but never back off.
  • Ingress role needs s3:GetObjectVersion. Since syrf c08e82e0d (2026-08-20, first released in s3-notifier-v1.19.0) the S3 event handler reads every event's metadata with a version-pinned HeadObject, which S3 authorizes only with s3:GetObjectVersion. The chart's S3ReadAccess granted just s3:GetObject, so on versioned buckets (staging, previews) every upload notification would fail its first metadata read. This went unnoticed because staging had 0 uploads in the week before durable-capture activation. The chart now grants both actions on the environment bucket's objects (SyRFS3NotifierLambdaBoundary already allows both). Terraform still declares a same-named inline S3ReadAccess for the bootstrapped staging/production/preview roles, and it must be updated in step or a terraform apply reverts the grant. Production runs pre-c08e82e0d code and is not affected until its code moves to v1.19.0 or later.
  • Held-version absence is proved by listing, never by HEAD. On staging on 2026-09-25, the reconciler deleted the held version for Proof A (CAPTURE#fa9edf5b…) at 20:48:23Z. An admin head-object then returned 404, but every later reconciler pass logged Reconciliation failed … Forbidden, so the row never retired. The absence read was a versioned HeadObject that treated only 404/405 as "gone". The reconciler role has prefix-conditioned s3:ListBucketVersions but no s3:ListBucket, and without s3:ListBucket S3 answers a HEAD on a missing key with 403. LocalStack enforces no IAM, so no test could see it. HeldVersionExistsAsync now answers every presence/absence decision with this procedure:
  • list ListObjectVersions with Prefix set to the exact key, which satisfies the role's s3:prefix condition, and page to the end;
  • treat the version as present if its exact id appears for the exact key as a version or a delete marker.

The same helper serves four decisions: pre-cleanup presence, unowned-rejected deletion, rejected-row retirement and the tombstone purge's "still gone" check. Every error, and a truncated page that does not advance, fails closed. No IAM changed. The only remaining HEAD is the inventory sweep's metadata read of a version it has just listed. A source-contract test pins both rules. Unit tests use an IAM-shaped fake S3 (403 on a missing-key HEAD, prefix-conditioned listing). LocalStack cannot catch this class of bug: the E2E and Testcontainers images are Community editions without IAM enforcement, and the E2E stack does not run the reconciler. - Rendered YAML must survive YAML 1.1 parsing. The first staging sync (cluster-gitops#1098, chart s3-notifier-v1.29.3) was rejected because the hold Table rendered attributeType: N unquoted: Helm reads it as a string, but ArgoCD/kubectl decode it with YAML 1.1 rules as boolean false. The chart now quotes every attributeType, and the PR chart lane runs .github/scripts/assert-chart-renders-have-no-yaml11-booleans.sh, which fails on any plain y/n/yes/no/on/off scalar in a rendered chart. helm lint and helm unittest cannot see this class of defect. - ACK controllers for the durable kinds. The Table, Queue and Rule are ACK dynamodb, sqs and eventbridge resources. Their controllers and AWS grants must exist in the cluster before any environment sets durableCapture.enabled, or the sync fails on unknown kinds. - Secret provisioning for durableCapture.reconciler.clientSecretName, and the Identity client whose scopes the reconciler requests, are separate prerequisites of the same activation window.

Where the durable workflow is documented

The cross-service workflow this slice introduces — durable S3 capture, the DynamoDB capture ledger, scheduled reconciliation, retention cleanup and the hold-retirement outbox — is documented for future agents in .claude/rules/pdf-agent.md (Bulk PDF / notifier section) and .claude/rules/s3-notifier.md. The root CLAUDE.md is deliberately an index and points at those rule files; do not restate the workflow there.

Environment status

Merging this code does not establish ARRNC, storage, broker or cloud readiness. The schedule remains disabled by default (durableCapture.scheduleEnabled: false); production capture is rejected. The activation receipt required by #3042 must be recorded by an operator during a fresh, complete zero-inventory proof window; the chart values, proof counters and authority-before-upload ordering are in Activate the Bulk PDF notifier authority. Infrastructure apply, secret provisioning, staging cutover, preview admission and production promotion retain their existing explicit gates. Do not remove expiry or activate durable takeover based on this merge alone.