E2E Testing Principles¶
Distilled principles for SyRF end-to-end testing, derived from industry research (see e2e-test-raw-research.md) and shaped for our monorepo, Playwright stack, and GitOps deployment model.
These principles should be referenced whenever new E2E tests are added or the testing infrastructure is modified.
1. Test Pyramid Discipline¶
E2E tests sit at the top of the testing pyramid. They are the most expensive to write, run, and maintain -- but provide the highest confidence that the full system works together.
Rules:
- E2E tests should be 5-10% of the total test suite. The bulk of coverage belongs in unit and integration tests.
- Only test critical user journeys end-to-end: authentication, project creation, screening workflows, annotation, data export. If a scenario can be validated by a unit or integration test, it should be.
- Resist the "inverted pyramid" anti-pattern (many E2E tests, few integration tests). This creates slow, brittle pipelines.
- When a new feature is built, ask: "Can this be covered by a unit or integration test?" If yes, write that test. Reserve E2E for the full-stack user flow.
2. Test Data Strategy¶
Test data management is the hardest part of E2E testing. There is no single perfect approach -- use the right tool for each situation.
Rules:
- Dynamic over static: Each test should create its own data and not depend on pre-existing database state. This prevents cross-test pollution and allows tests to run in any order.
- Seed via the domain model, not raw database manipulation: Use the
E2ESetupControlleron the PM service, which creates entities throughProjectBuilderandStudyBuilder— the same code path as production. Never insert documents directly into MongoDB from test code. Raw inserts bypass schema validation, computed properties, enum encoding, and entity relationships, causing silent failures that only surface at runtime. Thefixtures/db.fixture.tsshould be used for read-only verification (checking that screening decisions persisted), not for data creation. See e2e-testing-principles.md. - Clean up after yourself: Tests that create data should tear it down. For database-seeded data, use
afterAll/afterEachcleanup. For ephemeral preview environments, the environment teardown handles this. - Use unique identifiers: Generate unique project names, emails, and identifiers per test run (timestamps or UUIDs) to guarantee isolation when running in parallel.
- Static data is acceptable for read-only reference data: If a test only reads data that never changes (e.g., predefined annotation schemas), manually managed fixtures are fine. Don't over-engineer what doesn't need automation.
Choosing an approach:
| Situation | Approach |
|---|---|
| Test needs a project with studies, stages, questions | E2E setup API (POST /api/e2e-setup/project on PM service) — uses domain model builders |
| Test validates a UI workflow creates something correctly | API seed for prerequisites, then test the UI creation |
| Test needs a user with specific auth state | Auth fixture (cached session via Playwright storageState) |
| Data is read-only and stable | Static fixture file in fixtures/data/ |
| Test needs to verify data persisted correctly | Read-only MongoDB query via fixtures/db.fixture.ts |
3. Environment Parity¶
Tests are only as trustworthy as the environment they run in.
Rules:
- Test environments must mirror production: Same database engine (MongoDB), same service configurations, same auth flow (OpenIddict via mock OIDC server locally). Our
docker-compose.ymlandappsettings.e2etest.jsonachieve this. - Ephemeral preview environments for PRs: Each PR gets its own isolated namespace with real services. E2E tests run against these, not a shared staging environment. This prevents contention and "it works on staging" surprises.
- Mock external services at the boundary, not internally: Use sandbox/mock versions of Auth0/OpenIddict, S3 (LocalStack), and email. Do not mock internal API calls between SyRF services -- those interactions are exactly what E2E tests exist to validate.
- Never test against production databases.
4. Test Reliability and Flakiness¶
A flaky test is worse than no test -- it erodes trust in the entire suite and trains the team to ignore failures.
Rules:
- Never use hard-coded waits (
sleep,setTimeout). Use Playwright's built-in auto-waiting,waitForSelector,waitForResponse, orexpect().toBeVisible(). - Use stable selectors: Prefer
data-testidattributes, ARIA roles, or semantic selectors over CSS classes or DOM structure that changes with styling. - Isolate test state: Each test file should be runnable independently. No test should depend on another test's side effects.
- Retry on CI, investigate locally: CI retries (we use
retries: 2) catch transient infrastructure issues. But if a test fails repeatedly, fix the root cause -- don't just increase retries. - Quarantine, don't delete: If a test becomes persistently flaky, tag it and remove it from the critical path while investigating. Don't delete it -- the scenario it covers still matters.
5. Test Design Patterns¶
Rules:
- Page Object Model: All page interactions go through page objects (
pages/*.page.ts). Tests read like user stories; page objects encapsulate selectors and actions. This is already our pattern -- maintain it. - One assertion concern per test: A test named "screening workflow" should test the screening workflow. Don't bundle unrelated assertions to "save time" -- when it fails, you want to know exactly what broke.
- Test the login flow once, reuse auth state: Authenticate in the
setupproject, savestorageState, and reuse it across test projects. This is already implemented insetup/auth.setup.ts. - Prefer API-level validation alongside UI checks: After a UI action (e.g., creating a project), verify the result via API call or database query, not just by checking what the UI renders. This catches cases where the UI shows cached/optimistic state that doesn't match reality.
6. CI/CD Integration¶
Rules:
- Smoke tests on every PR: A fast subset tagged
@smokeruns on every pull request. These cover the critical path (auth, project creation, basic screening). - Full suite on merge to main or nightly: The comprehensive suite runs when changes land on main or on a scheduled basis. This catches integration issues without blocking every PR.
- Fail the build on E2E failure: E2E failures should be treated with the same urgency as unit test failures. A broken E2E test means a broken user journey.
- Capture diagnostics on failure: Screenshots, videos, and traces are captured automatically (
trace: 'retain-on-failure',screenshot: 'only-on-failure',video: 'retain-on-failure'). These are non-negotiable for debugging CI failures. - Parallelise where safe: Tests that don't share state should run in parallel (
fullyParallel: true). Tests that share a project or database state should be grouped in the same file and run sequentially within that file.
7. Test Maintenance¶
Rules:
- Treat test code as production code: Tests get code review, follow naming conventions, and avoid duplication. Helper functions belong in
helpers/, fixtures infixtures/, page objects inpages/. - Update tests when you change the feature: If a PR changes a user flow, the E2E test for that flow must be updated in the same PR. Don't leave broken tests for someone else.
- Audit quarterly: Review the test suite periodically. Remove tests for deleted features. Consolidate overlapping tests. Check that the suite still reflects real user journeys.
- Don't test copy or translations via E2E: If a label changes, the E2E test shouldn't break. Assert on structural elements (
data-testid, element roles) rather than display text where possible. When text assertions are necessary, reference constants or translation keys, not hardcoded strings.
8. Testing Transactional Emails¶
When user flows involve email (password reset, invitations), test them without relying on real email providers.
Rules:
- Use a test email service (e.g., Mailosaur, MailSlurp, or a local SMTP trap) that provides API access to received messages.
- Generate unique email addresses per test run to isolate assertions.
- Assert on email subject, body content, and link targets -- verify the email would actually enable the user to complete their flow.
- Never test the email provider itself -- only that your application sent the right content to the right address.
9. Reporting and Observability¶
Rules:
- HTML reports with traces: Playwright's HTML reporter is configured. Every CI run should produce a downloadable report artifact.
- Track trends over time: Monitor test duration and failure rates across runs. A test that gradually gets slower indicates a performance regression in the app.
- Log context on failure: When a test fails, the report should make it possible to diagnose the issue without re-running. This means: screenshot of failure state, video replay, network trace, and console logs.
10. When NOT to Write an E2E Test¶
Not every scenario needs E2E coverage. Skip E2E when:
- The scenario is purely computational (use a unit test)
- The scenario tests API contract compliance between two services (use an integration test or contract test)
- The scenario tests a UI component in isolation (use a component test)
- The scenario duplicates coverage that a lower-level test already provides
- The scenario is for an edge case that doesn't represent a real user path
The goal is maximum confidence with minimum E2E tests. Every E2E test should justify its execution time by covering something that no lower-level test can.
Quick Reference: Our Stack¶
| Concern | Tool / Pattern |
|---|---|
| Test framework | Playwright |
| Browser | Chromium (headless in CI) |
| Auth | Mock OIDC server (local), cached storageState |
| Database seeding | Domain-model API via E2ESetupController (PM service) + MongoDB read-only verification via fixtures/db.fixture.ts |
| Page interactions | Page Object Model (pages/*.page.ts) |
| Test data isolation | Unique IDs per run, cleanup in afterAll |
| CI smoke tests | @smoke tag, runs on every PR |
| CI full suite | Runs on merge to main |
| Diagnostics | Screenshots, video, traces on failure |
| Infrastructure | Docker Compose (MongoDB, RabbitMQ, mock OIDC, LocalStack) |
| Preview environments | Ephemeral per-PR via ArgoCD + Kubernetes namespaces |
Derived from research compiled in e2e-test-raw-research.md. Last updated: 2026-03-21.