RabbitMQ broker security contract (B0)¶
This is B0 of the implementation plan: the broker topology, target design and rollback contract for the design's P10 decision. B0 changes nothing live. It creates no users, permissions, listeners, certificates, secrets or workload configuration. Each B1/B2 live step needs its own approval. The steps marked (P) in B1 and B2 touch production, or the broker that production shares, and need Chris's explicit go for that step.
Contents¶
- Evidence and markers
- 1. Current topology
- 2. Target design
- 3. MassTransit endpoints and permission regexes
- B1 overlapping rollout
- B2 retirement
- Risks
- Decisions needed before B1
Evidence and markers¶
| Marker | Meaning |
|---|---|
| V | Verified on 2026-09-30 by read-only kubectl get on context gke_camarades-net_europe-west2-a_camaradesuk, or from camaradesuk/cluster-gitops origin/main 23aaef71 / camaradesuk/server-config origin/main. |
| D | Dated snapshot of 2026-09-28 01:01 UTC from the verification ledger. It was not re-run: the listing needs kubectl exec, which B0 may not use. The MCP broker tool was also not used, because its user listing returns password hashes and its connection listing returns more than the name/user/ssl columns. |
| S | Derived from SyRF source at 65ae1099b (MassTransit 8.4.0). |
| L | Observed on a local synthetic E2E broker (syrf-e2e-rabbitmq, about 3 weeks old, so it predates newer consumers). Read-only rabbitmqctl list_queues/list_exchanges/list_bindings. Holds no production data. |
| U | Unknown. It blocks the step that depends on it and is never taken to mean "absent". |
No Secret values, password hashes, message bodies or client properties were read.
1. Current topology¶
Broker deployment¶
| Item | Observed | Marker |
|---|---|---|
| Owner | cluster-gitops plugin plugins/helm/rabbitmq: Bitnami chart 14.6.6, image bitnamilegacy/rabbitmq:3.13.6-debian-12-r1, namespace rabbitmq |
V |
| Replicas | StatefulSet rabbitmq, 1/1 (rabbitmq-0). One broker serves production, staging, previews and local. A broker restart is a production outage. |
V |
| Plugins | rabbitmq_delayed_message_exchange (community plugin, downloaded at pod start) |
V |
| Definitions | rabbitmq-load-definition Secret committed in git. It holds no credentials: users: [], vhosts /, syrf-production, syrf-staging, local, and rabbit has .*/.*/.* on all four. Loaded at boot only. |
V |
| Bootstrap user | Chart auth.username: rabbit; the password and the Erlang cookie come from Secret rabbit-mq (ESO) |
V |
| Credential distribution | ClusterExternalSecret rabbit-mq (in argocd/local/argocd-secrets) copies both rabbitmq-password and rabbitmq-erlang-cookie into syrf-staging, syrf-production, rabbitmq and every namespace labelled syrf.org.uk/environment: preview (8 at present) |
V |
| Users | Only rabbit, tagged administrator |
D |
| Permissions | rabbit has .* configure/write/read on syrf-production and syrf-staging. There are no topic permissions. |
D |
| Vhosts | Production, staging and //local come from the definitions. Previews get pr-<N> at runtime (see below). The number of leftover pr-* vhosts is U. |
V/D/U |
| Listeners | AMQP 5672, management 15672, metrics 9419, clustering 25672, epmd 4369. No AMQPS 5671. | D; the container ports were re-verified V |
| Connections | 30 in total, all rabbit and all ssl=false: 3 production, 4 staging, 23 elsewhere |
D |
| Cluster Service | rabbitmq ClusterIP: 5672, 4369, 25672, 15672, 9419 |
V |
| Public AMQP | ingress-nginx tcp: 5672: rabbitmq/rabbitmq:5672 on LoadBalancer 34.13.63.21, i.e. rabbitmq.camarades.net:5672. Plaintext AMQP is reachable from the Internet. It was added so the Lambda notifier can connect. |
V |
| Public management | Ingress rabbitmq.camarades.net on 443 (Let's Encrypt letsencrypt-prod, cert-manager Certificate rabbitmq-tls, notAfter 2026-10-30) exposes the management UI/API for the shared administrator. |
V |
| NetworkPolicy | rabbitmq allows ingress from any source to 4369, 5672, 5671, 25672, 15672 and 9419. Egress is open. |
V |
Security consequences, in order of severity:
- The Erlang cookie is in every SyRF and preview namespace, and epmd/distribution ports are open to any
pod. A pod that can read the
rabbit-mqSecret in its own namespace and reach 4369/25672 can join the broker's Erlang distribution. That gives it full node control, beyond AMQP ACLs. Only therabbitmqnamespace needs the cookie. No SyRF chart or cluster-gitops file other than the broker's own values referencesrabbitmq-erlang-cookie(V,git grep). - Every workload, preview and external host authenticates as the same administrator. The same credential opens the public management UI.
- AMQP is plaintext, and the Lambda and the ARRNC agent send the administrator password over the Internet.
Components that talk to the broker¶
| Component | Where it runs | Broker endpoint | Credential source today | Marker |
|---|---|---|---|---|
| API | GKE syrf-{production,staging}, pr-N |
rabbitmq.rabbitmq.svc.cluster.local:5672, vhost per env |
rabbit + rabbit-mq Secret through syrf-common.env.rabbitmq |
V/S |
| Project Management | same | same | same | V/S |
| Quartz (MassTransit scheduler, job sagas; SQL Server) | same | same | same | V/S |
API preview PreSync hook presync-rabbitmq-vhost.yaml |
each pr-N |
management HTTP 15672 | rabbit admin: creates vhost pr-N and grants rabbit .* on it. No teardown deletes the vhost; the preview orphan sweep has no broker handling. |
S |
s3-notifier Lambda (syrfAppUploadS3Notifier, per-PR preview Lambdas) and reconciler |
AWS eu-west-1 |
amqp://rabbitmq.camarades.net:5672 (public) |
The env-vars-job copies the rabbit-mq password into plaintext Lambda environment variables. For uploads, the vhost comes from S3 object metadata virtualhost, which the API sets. |
V/S |
| PDF agent (production and staging slots) | ARRNC host, outside GKE (ADR-015) | rabbitmq.camarades.net:5672, amqp (public) |
Staging: server-config age scope syrf-pdf-agent-staging. Production: the arrnc-api-deploy workflow secret. The only broker user is rabbit (D), so any working credential is the administrator's. |
V/D |
local and / vhosts |
U (probably developer or historic) | public 5672 | rabbit |
U |
Consumer of IStartSearchUpdateCommand (API request client) and of ILivingSearch*Event |
U: none in this repository | S/U |
2. Target design¶
Principals: one per workload per environment¶
| Env | Vhost | Principals |
|---|---|---|
| Production | syrf-production |
syrf-production-api, -pm, -quartz, -s3notifier, -pdfagent |
| Staging | syrf-staging |
syrf-staging-api, -pm, -quartz, -s3notifier, -pdfagent |
Rehearsal (when a syrf-rehearsal namespace exists; V: none today) |
syrf-rehearsal |
syrf-rehearsal-<workload>, provisioned like staging |
Preview pr-N |
pr-N |
syrf-pr-N, shared by that preview's in-cluster workloads and its Lambda, with permissions only on pr-N. A preview PDF agent gets syrf-pr-N-pdfagent, as ADR-015 requires. |
| Operations | none by default | syrf-broker-observer (tag monitoring, no vhost permissions) for read-only verification without exec. The break-glass administrator is described in B2. |
Every principal is tagged none (no management UI or API), except the observer (monitoring) and the
break-glass administrator. No principal has permissions in more than one vhost. Vhost isolation is kept as
a second boundary and is never the only one.
Transport: AMQPS with a cert-manager certificate¶
- Enable Bitnami
auth.tlswith a 5671 listener.failIfNoPeerCert: falseandsslOptionsVerify: false: this is server-authenticated TLS, and clients authenticate with passwords. mTLS is a possible later option. - Certificate: a dedicated cert-manager
Certificaterabbitmq-amqps-tlsinrabbitmqforrabbitmq.camarades.net, issued by the existingletsencrypt-prodClusterIssuer (HTTP-01 already works for this host through the management Ingress). Keeping it separate from the Ingress'srabbitmq-tlsmeans AMQPS does not depend on the Ingress lifecycle. - Clients validate the public name. External clients (Lambda, ARRNC) connect to
amqps://rabbitmq.camarades.net:5671. They use public CA trust, so no CA bundle has to be distributed. In-cluster clients keep the hostrabbitmq.rabbitmq.svc.cluster.localand set the TLS server name torabbitmq.camarades.net. This avoids hairpinning through the public LB, and it keeps MassTransit addresses stable for persisted Quartz schedules (see Risks). It needs one small SyRF change:RabbitMqConfiggetsTlsServerName(plusPort5671 andSchemeNameamqps), andConfigureSyrfMassTransitcallshost.UseSsl(s => s.ServerName = …)with TLS 1.2 or later. There is no certificate-validation bypass setting. - Strict validation (B1, implemented): MassTransit 8.4 enables TLS for port 5671 but by default
accepts
RemoteCertificateChainErrors, so a self-signed certificate with the right name would pass. Every SyRF bus (sharedConfigureSyrfMassTransit, the s3-notifier and its reconciler) now declares its host throughRabbitMqTls.ConfigureSyrfRabbitMqHost. For anamqps/rabbitmqsaddress or port 5671 it enforces chain, name-mismatch and missing-certificate errors and allows TLS 1.2/1.3 only. The expected certificate name is the address host, unlessMessageBusConfig:RabbitMqConfig:TlsServerNameis set. That is the in-cluster server-name override above: dialrabbitmq.rabbitmq.svc.cluster.localand expectrabbitmq.camarades.net. The override changes only the expected name, never the validation. Its chart and env-mapping plumbing ships with the first in-cluster AMQPS switch. Plaintext addresses are unchanged. Local proof on RabbitMQ 3.13.6 (port 5671): the MassTransit default connected to an untrusted chain; the strict client rejected it and connected once the chain was trusted. The s3-notifier and the ARRNC agent must run code containing this change before they switch to AMQPS. - Alternative (decision D2): a private CA through a cert-manager CA Issuer, with SANs for the service DNS and the public name. This avoids the server-name override, but the CA must reach the Lambda and ARRNC.
- Rotation: the Let's Encrypt certificate renews every 60–90 days. B1 must prove in the fixture that the
broker serves the renewed certificate without a manual
exec, either through Erlang's PEM-cache refresh or through a declarative, scheduled rolling restart. Otherwise the rotation mechanism blocks B2. - External route: add
5671: rabbitmq/rabbitmq:5671to ingress-nginxtcp:(L4 passthrough, so TLS terminates at RabbitMQ). B2 then removes5672. - NetworkPolicy: restrict 4369/25672 to the RabbitMQ pods, 15672 to ingress-nginx and the provisioner, and 5671/5672 to SyRF namespaces and ingress-nginx.
Credentials: ESO Password generators, following the Valkey pattern¶
Reuse plugins/local/valkey-nonprod/resources/acl-credentials.yaml:
- A new
ClusterGeneratorrabbitmq-principal-password(Password, alphanumeric only, 48 characters, because the value lands in a URI or connection settings) inargocd/local/argocd-secrets. - In the
rabbitmqnamespace: oneExternalSecretper principal withrefreshPolicy: CreatedOnce, and a ServiceAccount + Role + RoleBinding that maygetonly that Secret. - One
ClusterSecretStore(Kubernetes provider,remoteNamespace: rabbitmq) per principal, conditioned on the consuming namespace only (syrf-staging,syrf-production, or the preview label forsyrf-pr-N-style stores). Workloads read their own principal's password through anExternalSecret. - Broker-side principals and permissions are declared by an owner that reads the same generated Secret.
Recommended (decision D1): the RabbitMQ Messaging Topology Operator. Its
UserCR imports its credentials from the ESO Secret,PermissionCRs carry the regexes, andVhostCRs handle previews. It reaches the Bitnami broker through a connection Secret held only inrabbitmq. RabbitMQ requires theadministratortag to manage users and vhosts, so that Secret belongs to a dedicated user,syrf-topology-operator(tagadministrator, no vhost permissions). This is a standing admin identity: it is never shared with workloads, and B2 lists it explicitly. Its password comes from the same PasswordClusterGeneratorthrough aCreatedOnceExternalSecret in therabbitmqnamespace only. This is a new operator, so it needs explicit approval (SyRF CLAUDE.md: no established declarative owner → stop for approval). The fallback is definitions only: an ESO-templatedload_definition.jsonwith users and permissions. Every change then needs a broker restart, which is a production outage on this single-replica broker, and it cannot serve preview churn. Whether definitions acceptpasswordrather thanpassword_hashis U and must be tested locally. - External consumers: the Lambda chart's
env-vars-jobreads the per-environment notifier Secret, notrabbit-mq. ARRNC slots keep their server-config age scope andarrnc-api-deployworkflow secret (ADR-015 ownership), and receive the-pdfagentprincipal's value by a manual out-of-band copy that is recorded without the value. An automated sync is a follow-up. - The Erlang cookie moves into a separate ExternalSecret that only the
rabbitmqnamespace receives, and is then rotated (B1.5b), because the current value is already exposed.
Rotation after B2: create a second principal (for example -api-2), switch, then delete the old one. This
reuses the B1 overlap mechanics and never edits a live password in place.
3. MassTransit endpoints and permission regexes¶
RabbitMQ semantics used¶
Declaring a queue or exchange needs configure, which also allows deleting it. basic.publish needs
write on the exchange. queue.bind needs write on the queue and read on the exchange. An
exchange-to-exchange bind needs write on the destination and read on the source. Consuming and purging
need read on the queue (RabbitMQ access control).
Whether a passive declare avoids configure must be confirmed in the fixture (U).
How SyRF names topology (S, L)¶
- Message-type exchanges:
<Namespace>:<Type>, for exampleSyRF.ProjectManagement.Messages.Commands:IMarkSessionIdleCommand, with exchange-to-exchange bindings up the interface hierarchy (for exampleSyRF.SharedKernel.Interfaces:ICommand). - Command queues:
MapCommandQueues()maps everyICommandtoqueue:<Assembly minus .Messages>.<Type>, for exampleSyRF.ProjectManagement.IMarkSessionIdleCommand. The exceptions are the three bulk-PDF writers, which map to the sharedSyRF.ProjectManagement.BulkPdfUpload. Aqueue:send endpoint declares the destination queue, its same-named exchange and the binding from the sender's connection. - Event/saga queues:
EventConsumerQueueName()=<Assembly minus .Endpoint>.<Type>, for exampleSyRF.ProjectManagement.SearchImportJobState. - Default-formatted endpoints (no
EndpointName):RunProjectStatisticsMaintenance,RunProjectStatisticsDriftCheck,RunProjectStatisticsRepair,StartDailyProjectStatistics,ObserveDailyProjectStatistics. These names are inferred from MassTransit's default formatter (S, confirm in the fixture). They have noSyRF.prefix. - MassTransit infrastructure:
quartz;Job,JobType,JobAttempt;MassTransit.Scheduling:*;MassTransit.Contracts.JobService:*, includingSubmitJob--<type>--;MassTransit:FaultandMassTransit:Fault--<type>--; lazily created<queue>_errorand<queue>_skipped. - Temporary queues: every bus has a bus endpoint
<machine>_<process>_bus_<id>(auto-delete), which request/response replies target. The API adds two per-pod fan-out endpoints:api-project-statistics-invalidationsandapi-activity-claim-revocationswith a GUIDInstanceId,Temporary = true. The exact runtime name patterns are captured in the fixture before any regex is final. SyrfAutoConfigureEndpointsinMassTransitHelpersis private and unused. All services useConfigureEndpointswith consumer definitions.
Per-workload resource matrix (S; L where the local broker showed it)¶
| Workload | Declares/consumes (configure + read) | Publishes/sends (write) |
|---|---|---|
| API | its bus endpoint; api-project-statistics-invalidations*, api-activity-claim-revocations*; the exchanges for SyRF.API.SignalR.Statistics:ProjectStatisticsChanged and SyRF.API.SignalR.ClaimRevocations:ActivityClaimRevoked |
publish SyRF.API.Messages.Events:{ISearchUploadStartedEvent,ILivingSearchEnabledEvent,ILivingSearchDisabledEvent}, the two SignalR fan-out types, MassTransit.Scheduling:{ScheduleMessage,CancelScheduledMessage} (IMarkSessionIdle, IRemoveSuspendedSession, ICheckConnectionLiveness), MassTransit.Contracts.JobService:SubmitJob--…IStartRobProcessingJobCommand--; send SyRF.ProjectManagement.IUpdateStudyScreeningStatsCommand; request the seven SyRF.ProjectManagement.I*BulkPdf*Command notifier-authority queues and SyRF.ProjectManagement.IStartSearchUpdateCommand (consumer U) |
| PM | every SyRF.ProjectManagement.* queue and endpoint exchange (command queues, BulkPdfUpload, SearchImportJobState, SearchImportJobErrorConsumer, SearchUploadSavedFaultConsumer, _error/_skipped), the five default-formatted statistics endpoints, its bus endpoint; binds to the SyRF.API.Messages.Events, SyRF.S3FileSavedNotifier.Messages.Events, SyRF.ProjectManagement.Messages.*, SyRF.ProjectManagement.Endpoint.Sagas and MassTransit:Fault--… exchanges it consumes, and to MassTransit.Contracts.JobService:SubmitJob--… for its three job consumers |
publish SyRF.ProjectManagement.Messages.Events:{IReferenceFileParsingCompletedEvent,IReferenceFileParsingFaultedEvent,ISearchImportJobErrorEvent}, IObserveDailyProjectStatisticsCommand, IStartDailyProjectStatisticsCommand, scheduling (ScheduleMessage, ScheduleRecurringMessage, cancel/pause/resume) and job-service contracts; send IStartParsingReferenceFileCommand and SyRF.ProjectManagement.IProcessBulkPdfUploadCommand (the PDF agent's queue); respond to the bus endpoints of the API, the PDF agent and itself; MassTransit:Fault* |
| Quartz | quartz, Job, JobType, JobAttempt (+ _error), its bus endpoint; binds MassTransit.Scheduling:* and MassTransit.Contracts.JobService:* |
every scheduled destination: the scheduled-command exchanges (IMarkSessionIdle, IRemoveIdleSession, IRemoveSuspendedSession, ICheckConnectionLiveness, IObserveDailyProjectStatistics, IStartDailyProjectStatistics, IRunProjectStatistics{Maintenance,DriftCheck,Repair}, ISearchUploadTimeoutExpiredEvent) and the input queues of every consumer using UseScheduledRedelivery/Redeliver (the PM bulk-PDF endpoints, the review-session queues, IProcessBulkPdfUploadCommand); job-service contracts and the PM job queues |
| s3-notifier (Lambda, reconciler) | its bus endpoint only | publish SyRF.S3FileSavedNotifier.Messages.Events:ISearchUploadSavedToS3Event; SubmitJob--…IStartBulkStudyUpdateJobCommand--; send SyRF.ProjectManagement.IBulkPdfUploadObjectReadyCommand |
| PDF agent (ARRNC) | SyRF.ProjectManagement.IProcessBulkPdfUploadCommand (+ _error/_skipped) only when ConsumerState=Enabled; its bus endpoint |
request IClaimBulkPdfUploadProcessingCommand and send IReportBulkPdfUploadProgressCommand/IFinalizeBulkPdfUploadCommand, all to SyRF.ProjectManagement.BulkPdfUpload; MassTransit.Scheduling:ScheduleMessage for durable redelivery; MassTransit:Fault* |
Regex construction¶
The target is one regex per permission type per principal. It is built from anchored alternations of the names above, with the environment isolated by vhost. A worked draft for staging PM:
configure: ^(SyRF\.ProjectManagement\.[A-Za-z]+(_error|_skipped)?|(RunProjectStatistics(Maintenance|DriftCheck|Repair)|(Start|Observe)DailyProjectStatistics)(_error|_skipped)?|SyRF\.(ProjectManagement\.(Messages\.(Commands|Events)|Endpoint\.Sagas)|API\.Messages\.Events|S3FileSavedNotifier\.Messages\.Events|SharedKernel\.Interfaces):[A-Za-z]+|MassTransit(:Fault(--.+--)?|\.Scheduling:[A-Za-z]+|\.Contracts\.JobService:.+)|<pm-bus-endpoint-pattern>)$
write: <configure alternatives> | <api/pdfagent/pm bus-endpoint patterns for responses>
read: <PM queues> | <every exchange PM binds from, including the API and notifier event exchanges>
This is a draft, not a proposed value. B1's local fixture turns each row into generated regexes with
positive tests (every declare/publish/consume path the service uses) and negative tests. The negative
tests cover the notifier consuming a PM queue, the PDF agent publishing IBulkPdfUploadObjectReadyCommand,
the API deleting SyRF.ProjectManagement.*, any principal using another vhost, and preview → staging.
Commit the regexes only after both sets pass.
Topology-declaration needs¶
MassTransit declares on start and on first send, so pure least privilege is not reachable without changing SyRF:
- Senders declare destination queues, so the API, notifier and PDF agent need configure on some PM queues. Configure also allows delete.
- A fault on any consumer declares
_error/_skippedandMassTransit:Fault--…exchanges. - A publisher declares the whole interface hierarchy of each message type.
The recommended order is:
- The owning consumer declares. PM declares every
SyRF.ProjectManagement.*queue. Quartz declaresquartz/Job*. - Senders get a narrow configure grant on the exact destination names they send to, never a
SyRF.ProjectManagement.*wildcard. The remaining risk is that a sender with configure on a destination could delete that queue. - Fixture question for B1: can MassTransit 8.4 send without declaring (a passive or skip-topology send address), so senders need only write? If yes, remove senders' configure grants. If no, record the remaining risk under D6.
- Name the five default-formatted statistics endpoints under
SyRF.ProjectManagement.only in a separate SyRF PR that drains and retires the old queues. B1 does not require it.
B1 overlapping rollout¶
Mechanism: new principals and AMQPS run alongside rabbit/5672. Every workload switch is a GitOps
values change (credential Secret name, scheme, port, server name) in its own commit, so rollback is a
revert. Topology names and vhosts do not change, so queues, saga state (Mongo, SQL) and Quartz triggers
survive both directions.
Verification for every switch:
- Argo Healthy.
- Readiness, including the PDF agent's broker readiness.
kubectl logsfor the broker and the workload show noACCESS_REFUSEDor TLS errors.- Queue consumer counts and backlog return to baseline. Read these through the
syrf-broker-observermanagement API: names, users,ssland counts only. - One full operational cycle of the affected scheduled work (idle/liveness schedules, daily statistics).
- No live test messages. Only organic traffic is observed.
Stop and roll back on any auth failure, any stuck scheduler, a queue left without consumers, or an unknown consumer.
Order: the IDs below group the steps by kind; they are not a sequence. B1.5 (Erlang cookie and NetworkPolicy) depends on none of B1.0–B1.4 and should run first after approval. Its NetworkPolicy part immediately stops other pods reaching epmd/distribution, with no broker restart, SyRF change or new operator. Moving the Secret prevents new copies, but the current cookie value has already been copied into every SyRF and preview namespace, so it stays compromised until B1.5b rotates it. Rotation restarts the broker.
Interim exposure of public 5672: no network-level interim mitigation is feasible on the current route.
The Lambda notifier runs outside a VPC with no fixed egress address, and loadBalancerSourceRanges on the
shared ingress-nginx Service would also restrict the public web and API. The interim reduction is
therefore the staged switch itself: each workload that moves to a scoped principal over AMQPS stops
sending the admin credential in plaintext. A dedicated LoadBalancer Service for AMQP, which could later
allowlist ARRNC's egress, is an optional follow-up. Lambda VPC egress is out of B0–B2 scope.
| # | Step | Owner repo | Touches production? |
|---|---|---|---|
| B1.0 | SyRF: RabbitMqConfig TLS (amqps, TlsServerName, TLS 1.2+), chart value plumbing, s3-notifier/PDF agent TLS; unit tests; defaults unchanged (amqp). Flag decision: no runtime flag, because the setting is inert until values change it per environment. |
syrf | No |
| B1.1 | Local fixture (E2E broker): generated regexes, positive/negative ACL tests, TLS hostname/CA failure, reconnect, and persisted-schedule address compatibility (amqp→amqps). The resulting regexes are committed. | syrf | No |
| B1.2 | Decision D1: install the Topology Operator (or approve the definitions fallback) | cluster-gitops | (P): it holds admin to the production broker |
| B1.3 | rabbitmq-amqps-tls Certificate + Bitnami auth.tls (adds 5671). Restarts the shared broker. |
cluster-gitops | (P) |
| B1.4 | ingress-nginx tcp: 5671 passthrough (additive; controller rollout on the production LB) |
cluster-gitops | (P) |
| B1.5 | Split the Erlang cookie out of rabbit-mq; rabbit-mq stops carrying the cookie in syrf-*/preview namespaces; tighten the NetworkPolicy |
cluster-gitops | (P): changes the syrf-production Secret and the broker policy |
| B1.5b | Rotate the Erlang cookie: a new ESO-generated value, rabbitmq namespace only, then a broker restart. It can be combined with B1.3's restart. |
cluster-gitops | (P): restarts the shared broker |
| B1.6 | syrf-broker-observer principal (monitoring, read-only) |
cluster-gitops | (P): user on the shared broker, no production vhost permissions |
| B1.7 | Staging principals + ESO/ClusterSecretStores + Permission CRs | cluster-gitops | No (staging vhost only) |
| B1.8 | Staging switches, one per commit, each verified before the next: (a) s3-notifier + reconciler → amqps://rabbitmq.camarades.net:5671 and syrf-staging-s3notifier; (b) PDF agent staging (ARRNC, Paused) via server-config/arrnc-api-deploy; © Quartz; (d) PM; (e) API |
cluster-gitops, server-config | No |
| B1.9 | Previews: per-pr-N Vhost/User/Permission from the preview chart/ApplicationSet; the API PreSync hook stops using rabbit; preview teardown and the orphan sweep delete the vhost and user; the preview Lambda gets syrf-pr-N |
syrf, cluster-gitops | No, but the rabbit-mq distribution to previews changes in B2 |
| B1.10 | Production principals + stores | cluster-gitops | (P) |
| B1.11 | Production switches in B1.8 order: s3-notifier/reconciler, PDF agent (production slot), Quartz, PM, API. Each is its own gate. | cluster-gitops, arrnc-api-deploy | (P) each |
Rollback during B1: revert the workload's values commit, or re-run its Lambda env-vars-job or ARRNC
deploy with the previous values. The previous rabbit/5672 route stays intact until B2. Reverting B1.3 is
another broker restart (P).
B2 retirement¶
Entry proof, collected by the observer: for a proposed 7 days, including a daily-statistics run,
every connection in syrf-production and syrf-staging is a new principal with ssl=true. The same holds
for every live preview. There are no broker ACCESS_REFUSED entries, every queue has its expected
consumers, and backlog drains. The owners of local//, IStartSearchUpdateCommand and the living-search
events are identified, or explicitly retired by Chris. An unknown consumer blocks B2.
| # | Step | Touches production? |
|---|---|---|
| B2.1 | Remove 5672 from the ingress-nginx tcp: map (no public plaintext) |
(P) |
| B2.2 | Remove rabbit permissions on syrf-production, syrf-staging, local and / from definitions/CRs; stop distributing rabbit-mq to syrf-* and preview namespaces |
(P) |
| B2.3 | Disable the plaintext listener (listeners.tcp = none); restarts the broker |
(P) |
| B2.4 | Replace the chart bootstrap user with a break-glass administrator (below); retire rabbit |
(P) |
| B2.5 | Restrict the public management Ingress (source allowlist or authentication in front) | (P) |
Break-glass administrator: syrf-broker-breakglass, tag administrator, no standing vhost
permissions. Its password lives only in GCP Secret Manager and a rabbitmq-namespace-only
ExternalSecret. It is never in a workload namespace, the Lambda or ARRNC. Use requires Chris's approval
and a recorded reason, and the password is rotated afterwards. B2 end state: exactly two administrator
identities remain, syrf-broker-breakglass (dormant, approval-gated) and syrf-topology-operator
(standing, used only by the operator). Both credentials exist only in the rabbitmq namespace and GCP
Secret Manager. No workload, Lambda or ARRNC host holds an administrator credential.
Recovery after B2: a locked-out workload is fixed by re-reconciling its ESO Secret and User or Permission CR, or by using the rotation overlap. It is not fixed by restoring the shared admin to workloads, as the plan requires. If the broker itself is lost, the definitions and CRs rebuild vhosts, users and permissions declaratively. Durable queue contents are lost if the PVC is lost; that is today's exposure too. Reverting a B2 step needs its own (P) go.
Risks¶
| Risk | Consequence | Mitigation / contract |
|---|---|---|
| MassTransit auto-declaration | Senders need configure (therefore delete) on destination queues. Faults create _error/_skipped and fault exchanges lazily, so an ACL gap appears only when something fails. A denied operation raises ACCESS_REFUSED, which closes the channel, and the consumer retries. |
Fixture exercises the fault and redelivery paths explicitly; owner-declares ordering; the passive-send question; negative tests. Endpoints use UseInMemoryOutbox, so a denied publish after DB side effects causes redelivery. Consumers must stay idempotent, as they already must. |
| Quartz scheduler / saga | Quartz writes wherever a schedule points. Anyone with write on MassTransit.Scheduling:ScheduleMessage can have Quartz deliver arbitrary messages to anything in Quartz's write regex. Quartz is a confused deputy. Persisted one-shot triggers (session idle/suspended/liveness, bulk-PDF redelivery, the search-upload timeout) store destination addresses that include host and scheme. |
The Quartz write regex is an explicit allowlist of scheduled destinations, never .*. Keep in-cluster hosts unchanged (server-name override); the fixture proves that rabbitmq:// addresses persisted before the switch still deliver after amqps. Recurring schedules are re-registered at PM start. Job sagas (Job*) persist in SQL and are unaffected by credential changes. Switch Quartz after its destinations are in the allowlist. |
| ARRNC PDF agent outside the cluster | It reaches the broker over the Internet; its credential lives in server-config/arrnc-api-deploy, outside ESO; it needs ScheduleMessage (the confused-deputy path above). |
A dedicated -pdfagent principal whose write is limited to the BulkPdfUpload sends, the claim request, ScheduleMessage and its own bus endpoint. AMQPS with public-CA validation. A future option routes redelivery through a narrower path so the agent does not need the shared scheduler. ADR-015's per-slot least-privilege requirement is met by this principal. |
| Preview churn | 8 previews exist now. Vhosts are created by an admin hook and never deleted. Per-preview principals add create/delete churn on a shared broker. | Operator-managed Vhost/User/Permission CRs whose lifecycle is tied to the preview app; the orphan sweep also removes broker users/vhosts for closed PRs; a cap on preview principals and vhost limits (max-connections, max-queues), as ADR-015 already asks. |
| Single shared broker | TLS enablement, definitions changes and B2.3 restart production too. | Each is a (P) step in a low-traffic window; MassTransit reconnect behaviour is proven in the fixture first. |
| Lambda credential in plaintext environment variables | Anyone who can read the function configuration sees the broker password. | Per-env notifier principals limit the blast radius. Moving to Secrets Manager/SSM is a follow-up, outside B0–B2. |
| Vhost chosen from S3 metadata | Under the shared admin, a notifier could publish into any vhost that the object metadata names. | Per-environment notifier principals make a cross-environment vhost value fail with an auth error. |
| Certificate rotation | An expired server certificate breaks every client at once. | Rotation proof is a B1 exit criterion (see Transport). |
PDF notifier transport: HTTP versus MassTransit. This question is open, and the plan defers it "until bulk-PDF work next touches the notifier". B0 does not decide it. It only notes the effect:
- Notifier cleanup authority already uses HTTP (
/api/notifier/bulk-pdf/*with OpenIddict), and object readiness is still a MassTransit send. - If the notifier moves fully to HTTP, the
-s3notifierprincipal loses itsIBulkPdfUploadObjectReadywrite and keeps only the reference-import publish and study-update job submission. - If it stays on MassTransit, the matrix above stands.
Either outcome changes a regex, not the principal model.
Decisions needed before B1¶
| ID | Decision | Recommendation |
|---|---|---|
| D1 | Declarative owner of broker users, permissions and preview vhosts | Topology Operator (new cluster-gitops plugin); fallback is definitions-only, which needs restarts and has no preview support |
| D2 | TLS server identity | Let's Encrypt rabbitmq.camarades.net + in-cluster server-name override; alternative is a private CA |
| D3 | Public AMQP route | Keep the public LB for 5671 only (Lambda has no fixed egress; ARRNC is off-cluster); private connectivity is a follow-up |
| D4 | Fate of the local and / vhosts and their clients |
Identify them in the B1 observer listing; retire in B2 unless an owner claims them |
| D5 | Break-glass custody and procedure | As described in B2 |
| D6 | Accept the remaining configure-implies-delete risk if MassTransit cannot send passively | Accept for B1 with exact-name grants; revisit after the fixture |
Related: design P10, implementation plan B0–B2, service/client inventory, ADR-015.