Bulk PDF durable capture: merge acceptance and remaining activation gates¶
PR #3094 adds the disabled-by-default capture and ordinary terminal-cleanup slice on top of the authority contract in #3042. Its usable boundary is an immutable, correctly bound object: capture survives duplicate events, publication consults Project Management, terminal cleanup retains the exact version until its retention deadline, and acknowledgement retires the hold through a durable outbox. Publication pause permits cleanup to continue.
Acceptance for this slice:
- An exact-binding duplicate preserves the capture; conflicting binding data poisons it.
- An expired lease can be recovered without replacing a newer durable state from a stale index read.
- A cancel or abandon that lands after capture but before the hand-off (job
Aborting) is handed to Project Management's cleanup-only object-ready branch even while publication is paused; it terminalises the job with durable cleanup-only authority in one transaction and sends no agent work, so the job cannot deadlock inAborting(staging upload5ba76ac4, 2026-09-25). Covered by the reconciler unit tests and the Bulk PDF lane's cancel-while-paused E2E case. - A terminal object is never published to the agent. A retriable claim remains held for as long
as the job's retry window is open (
BulkPdfUploadJob.RetryAuthorityWindow, 30 days from the terminal instant); once it lapses the claim is gone and cleanup proceeds. - Cleanup uses the exact bucket/key/version and preserves the first retention deadline.
- Capture retirement and hold-retirement outbox creation are one conditional transaction; the outbox is deleted only after its acknowledgement is recorded.
- Each inventory invocation reads one bounded version page and checkpoints only after all its captures succeed. Overlapping invocations cannot overwrite a newer checkpoint. Due cleanup and outbox work run before inventory, so an inventory outage cannot starve them.
- Missing/malformed metadata and metadata naming another environment remain non-publishable under the configured environment's reserved-object boundary.
- The reconciler requests the scopes issued by Identity, and its role permits the configured failure-queue send. Existing chart defaults render no new durable resources.
Validation includes notifier unit and LocalStack integration tests, chart unit tests, chart-change detection, and the PR's current-commit CI/reviews. The PR description records the final counts.
Ordered follow-ups owned by #3168¶
Complete rejected-object cleanup and quarantine adoption.Done (PR #3558). Mechanism: the authenticated resolver now also discloses the server-owned release for a single unambiguous match, because a rejected binding derives its release id from the object coordinates alone and matches nothing in Project Management. Zero matches delete the exact version after the snapshotted invalid-object retention and retire the row by a conditional delete covering every immutable binding field, gated on a completed inventory checkpoint later than that deletion plus an exact-version absence read; no acknowledgement, outbox or tombstone is fabricated. One match pins the disclosed release on the row before any authority is requested for it, then registers cleanup-only authority and joins the ordinary retention path without ever publishing; a pinned row resumes from its release rather than resolving again. Many matches, or resolver evidence disagreeing with the digest recorded on the row, register a durable quarantine fence under a capture-key-derived fence-set id before custody changes; the fence route names no release and Project Management re-resolves the candidates, adopting only unowned ones and failing closed on any held elsewhere. An unavailable resolver, a lost response or a crash re-drives from durable state, and a competing worker loses the conditional write. Covered by notifier unit tests, two LocalStack retirement tests and the Project Management/API contract tests.Complete tombstone replay/queue/inventory retirement.Done (PR #3575). Tombstones still have no TTL; they now leave the active partition of the due index for aTOMBSTONEone and become due at exactly their snapshotted replay deadline, so the existingStateBucket/DueAtEpochindex serves the purge pass and no new index was added. Three facts are each recorded on the row by a row-version-conditioned write before an exact conditionalDeleteItemnames every one of them alongside every immutable binding field: the snapshotted deadline (read, never recomputed); the failure queue, proved by two observations of the visible, in-flight and delayed counters, all zero, at leastFailureQueueQuietPeriod(PT15M) apart and both after the deadline —GetQueueAttributeshas noApproximateAgeOfOldestMessage, which is a CloudWatch metric, so the first empty observation is the watermark and a delayed redrive landing between the samples is seen and withdraws the whole proof, inventory evidence included; and a completed inventory checkpoint newer than both the deadline and the queue proof, plus an exact-version read that still reports the object gone. An unconfigured or unavailable queue, a queue that omits a counter, a still-unretired hold outbox, a truncated or older checkpoint, a version that is present again, and a lost conditional write all retain the row and retry; no age or TTL path can delete one. AQuarantineTransferTombstoneis declared and refused by name — ADR-015 (lines 534-538) adds destination cleanup acknowledgement, fence-set clearance and shared-debit release on top of this fence, and follow-up 4 owns them. Covered by the notifier purge unit tests and three LocalStack tests (exact conditional purge, due visibility, real SQS attribute reads).Schedule and validate receipt-drain reconciliation.Done (#3556).BulkPdfReceiptDrainnow carries a bounded producer beside the consumer:BulkPdfReceiptDrainBackgroundServiceruns oneBulkPdfReceiptDrainProducerpass perBulkPdfReceiptDrain:Interval(default 10 minutes,BatchSize100, both validated on start), and each pass discovers retained receipts throughIBulkPdfUploadReleaseRecordRepository.GetDrainCandidatesAsync—CleanupReceiptrecords whose hold retirement is acknowledged and whose drain is either unproved or proved and pastReceiptExpiresAt— oldest first, then sends oneIReconcileBulkPdfReceiptDrainCommandper record over the existing command queue. The producer decides nothing: it never fabricates an outbox key and never records an acknowledgement, so exact-key outbox absence, history pruning and receipt expiry are still proved together inside the consumer before a purge. A receipt whose outbox row still exists is left untouched and rediscovered on the next pass, and a pass runs to completion before the interval delay, so the retry cadence is the interval. One faulting candidate does not abandon the rest of its page. Proved byBulkPdfReceiptDrainProducerTests(consumer and producer wired over the MassTransit harness) and theGetDrainCandidatesAsyncrepository tests. Still feature-gated:BulkPdfReceiptDrain:Enabledis off by default and a disabled pass queries and sends nothing. Chart wiring for this and for the receipt-retention settings onBulkPdfNotifierAuthorityOptionslanded in #3560 (bulkPdfReceiptDrain.*andbulkPdfNotifierAuthority.receiptPolicyVersion/receiptRetentionDays.*), so cluster-gitops can now set them; no activation-time environment values changed as part of that PR.- Complete quarantine transfer, preview provisioning and the separately approved rollout evidence in ADR-015. These are later slices, not properties established by merging #3094.
Activation gates owned outside this chart¶
- Setup-role IAM grant (blocking).
durableCapture.reconciler.setupRoleAuthorizedisfalseby default and gates the reconciler'slambda add-permission(Sync hook) andlambda update-function-configuration(PostSync hook). Both hook Jobs run as the externalsyrf-ack-setup-jobrole, whose Lambda mutation scope is the exact ingress function ARNs (technical plan,syrf-ack-setup-job). Infrastructure must add the exact reconciler function ARN to that role before the flag is flipped; until then the reconciler stays unconfigured and its EventBridge rule staysDISABLED(even withdurableCapture.scheduleEnabled: true, see below), and the chart issues no call that would fail the sync. Once authorized, both Jobs first polllambda get-function-configuration(the only read the setup role holds on the reconciler) until the reconciler reportsActivewith no update in progress, bounded bysetupJob.timeouts.lambdaActive. On the first sync ACK may not have created the function yet, and an unguarded call would fail the hook withResourceNotFoundException. Covered by.chart/tests/reconciler_setup_wait_test.yaml. - Schedule switch.
durableCapture.scheduleEnabled(defaultfalse) is the only way the EventBridge rule rendersENABLED, and only whendurableCapture.enabledandreconciler.setupRoleAuthorizedare also true. Enable it in a separate cluster-gitops change after the setup Jobs have configured the reconciler. Covered by.chart/tests/durable_capture_schedule_test.yaml(all eight gate combinations). - Hold-table point-in-time recovery is a per-environment values decision:
durableCapture.pointInTimeRecovery(defaultfalse) is always rendered as an explicitspec.continuousBackups.pointInTimeRecoveryEnabledboolean on the ACKTable. An omitted field would be unmanaged by ACK, so recovery could be turned on but never back off. - Ingress role needs
s3:GetObjectVersion. Since syrfc08e82e0d(2026-08-20, first released ins3-notifier-v1.19.0) the S3 event handler reads every event's metadata with a version-pinned HeadObject, which S3 authorizes only withs3:GetObjectVersion. The chart'sS3ReadAccessgranted justs3:GetObject, so on versioned buckets (staging, previews) every upload notification would fail its first metadata read. This went unnoticed because staging had 0 uploads in the week before durable-capture activation. The chart now grants both actions on the environment bucket's objects (SyRFS3NotifierLambdaBoundaryalready allows both). Terraform still declares a same-named inlineS3ReadAccessfor the bootstrapped staging/production/preview roles, and it must be updated in step or a terraform apply reverts the grant. Production runs pre-c08e82e0dcode and is not affected until its code moves to v1.19.0 or later. - Held-version absence is proved by listing, never by HEAD. On staging on 2026-09-25, the
reconciler deleted the held version for Proof A (
CAPTURE#fa9edf5b…) at 20:48:23Z. An adminhead-objectthen returned 404, but every later reconciler pass loggedReconciliation failed … Forbidden, so the row never retired. The absence read was a versioned HeadObject that treated only 404/405 as "gone". The reconciler role has prefix-conditioneds3:ListBucketVersionsbut nos3:ListBucket, and withouts3:ListBucketS3 answers a HEAD on a missing key with 403. LocalStack enforces no IAM, so no test could see it.HeldVersionExistsAsyncnow answers every presence/absence decision with this procedure: - list
ListObjectVersionswithPrefixset to the exact key, which satisfies the role'ss3:prefixcondition, and page to the end; - treat the version as present if its exact id appears for the exact key as a version or a delete marker.
The same helper serves four decisions: pre-cleanup presence, unowned-rejected deletion,
rejected-row retirement and the tombstone purge's "still gone" check. Every error, and a
truncated page that does not advance, fails closed. No IAM changed. The only remaining HEAD is
the inventory sweep's metadata read of a version it has just listed. A source-contract test pins
both rules. Unit tests use an IAM-shaped fake S3 (403 on a missing-key HEAD, prefix-conditioned
listing). LocalStack cannot catch this class of bug: the E2E and Testcontainers images are
Community editions without IAM enforcement, and the E2E stack does not run the reconciler.
- Rendered YAML must survive YAML 1.1 parsing. The first staging sync (cluster-gitops#1098,
chart s3-notifier-v1.29.3) was rejected because the hold Table rendered attributeType: N
unquoted: Helm reads it as a string, but ArgoCD/kubectl decode it with YAML 1.1 rules as
boolean false. The chart now quotes every attributeType, and the PR chart lane runs
.github/scripts/assert-chart-renders-have-no-yaml11-booleans.sh, which fails on any plain
y/n/yes/no/on/off scalar in a rendered chart. helm lint and helm unittest
cannot see this class of defect.
- ACK controllers for the durable kinds. The Table, Queue and Rule are ACK dynamodb,
sqs and eventbridge resources. Their controllers and AWS grants must exist in the cluster
before any environment sets durableCapture.enabled, or the sync fails on unknown kinds.
- Secret provisioning for durableCapture.reconciler.clientSecretName, and the Identity client whose
scopes the reconciler requests, are separate prerequisites of the same activation window.
Where the durable workflow is documented¶
The cross-service workflow this slice introduces — durable S3 capture, the DynamoDB capture ledger,
scheduled reconciliation, retention cleanup and the hold-retirement outbox — is documented for future
agents in .claude/rules/pdf-agent.md (Bulk PDF / notifier
section) and .claude/rules/s3-notifier.md. The root
CLAUDE.md is deliberately an index and points at those rule files; do not restate the workflow
there.
Environment status¶
Merging this code does not establish ARRNC, storage, broker or cloud readiness. The schedule remains
disabled by default (durableCapture.scheduleEnabled: false); production capture is rejected. The activation receipt required by #3042
must be recorded by an operator during a fresh, complete zero-inventory proof window; the chart
values, proof counters and authority-before-upload ordering are in
Activate the Bulk PDF notifier authority. Infrastructure
apply, secret provisioning, staging cutover, preview admission and production promotion retain their
existing explicit gates. Do not remove expiry or activate durable takeover based on this merge alone.