Bulk PDF Upload¶
Project admins upload a folder of full-text PDFs for a systematic search directly from the SyRF UI, replacing the manual email-the-team workflow (epic #2223, SOP discussion #1930).
How it works (user view)¶
- On the Systematic Searches page, choose Bulk upload PDFs for a search (requires the
BulkPdfUploadpermission; feature-flagged viabulkPdfUpload). - Drag in (or pick) a folder containing only PDFs. The app validates the contents and shows a preview: which files match which studies (by relative path), which files have no matching study, and which studies still lack a PDF.
- Confirm. The app zips the folder in the browser and uploads it directly to S3 with a progress bar.
- After receipt,
Uploadeddisplays Received, waiting to be processed until processing begins; it does not imply scanning has started.Scanningdisplays Checking files. The pipeline scans every file for viruses, validates it is a real PDF, and copies it to the PDF web server. Live status streams into the UI. - On completion: a result summary (copied / skipped / replaced / unmatched / missing / invalid counts — or Infected, in which case nothing is delivered) and a downloadable CSV report with one row per file.
A study's Open PDF link activates only once its file is verifiably in place — zero broken links.
The Searches operations card offers Open upload, which expands and focuses the owning
search row without starting another session. Commands remain in that permission-gated row.
The expander retains Material's full-size icon-button target so its chevron is not clipped; it
carries an accessible name (Show/Hide PDF uploads for <search>), aria-expanded and a tooltip.
The Searches page also honours a deep link, /projects/{projectId}/searches?openUpload={searchId}
(OPEN_UPLOAD_QUERY_PARAM in search-operation.presentation.ts). Once the upload gate is open and
the search is in the table, it expands, scrolls to and focuses that row, then removes the
parameter (replacing the history entry) so a reload does not reopen a row the reader closed. A
link is honoured once: a one-shot guard stops the gate or the search list re-emitting before the
router clears the parameter from opening the row and navigating again; a rejected or cancelled
clearing navigation releases the guard. If the row's expander is
not rendered yet, nothing polls; the parameter stays and the next change to the gate or the
search list tries again (#3812). The
Processing row detail for a Bulk PDF upload links there as Open upload — a link, not a
command, so Processing stays read-only. It is shown only when the reader could open Searches and
the page would render the expander (the bulkPdfUpload flag and capability). No new feature
flag: the link rides the existing bulkPdfUpload and studyManagementProcessing gates.
Upload and Processing timestamps use the viewer's local time with an explicit UTC offset.
The operations card's Updated field uses only a supplied update timestamp, never creation
or completion as a substitute. For Bulk PDF that is BulkPdfUploadJob.UpdatedAt (#3812), exposed
as BulkPdfUploadJobDto.updatedAt: the instant of the most recent accepted mutation, seeded with the
creation time. It moves
exactly when StateRevision moves (one MarkStateChanged step in the aggregate; $max on the
repository's direct partial-update path), so status transitions, count-only progress, lease
renewals, part acknowledgements and cleanup markers all advance it, while idempotent
redeliveries do not. It never moves backwards across replicas with skewed clocks. Progress is
stamped with the PM consumer's receipt time, not the agent's clock. The card and Processing read
it through the same formatter, so they agree.
Rows persisted before the field existed have no value and still show Not available. They are
not given a substitute such as CompletedAt, UploadCompletedAt or CreatedAt: a legacy row's
later progress and cleanup writes left no time behind, so any of those could understate how
recently it changed, and the created and completed times are already shown separately. A legacy
row gains an UpdatedAt from its next accepted mutation. Unfinished jobs have no completion time.
No new feature flag: this is an additive, read-only field on the existing bulkPdfUpload surface.
Indeterminate progress animations make no measurable-progress claim and need not be synchronised
between tabs.
Documents¶
S3ReadAccess conformance guard¶
The chart's iam_role_s3_read_test.yaml pins the full rendered S3ReadAccess document
for staging, production and a per-PR preview role. Coordinate intentional changes with
camarades-infrastructure/terraform/lambda/main.tf
and its mocked s3_notifier_ingress_read_policies contract in
tests/s3_notifier_boundary_contract.tftest.hcl (#3670). Both CI lanes fail on a one-sided
policy edit; failures name the companion owner. There is no live cross-repository fetch.
Both owners pin s3:GetObject and s3:GetObjectVersion. Resource patterns are intentionally
not identical for previews: Terraform's legacy shared preview role covers production and
preview bucket patterns, whereas an ACK per-PR role covers only its exact bucket. The guards
pin those existing scopes separately; they do not broaden IAM or perform an infrastructure apply.
Design references¶
- Design: 2026-08-11-bulk-pdf-upload-v2-design.md
- Implementation plan: 2026-08-11-bulk-pdf-upload-v2-plan.md
- Decision record: ADR-012
- Hosting correction: ADR-015
- ARRNC migration plan: 2026-08-30-bulk-pdf-agent-arrnc-hosting-plan.md
Prior attempt PR #2373 is superseded by these documents.