The testing pyramid
Lots of fast, focused tests at the bottom; fewer, slower, broader ones at the top. Every layer runs on every push and pull request, and a red layer blocks the merge.
| Layer | Tools | What it proves | Runs |
|---|---|---|---|
| Unit | xUnit v3, Shouldly, NSubstitute, FakeTimeProvider | Business rules: tax, FX locks, idempotency, validation, reports | Every push · ~5 s |
| Architecture | NetArchTest | The layers stay clean; Avalara stays inside its adapter | With unit tests |
| Integration | WebApplicationFactory, Testcontainers Postgres, WireMock | Real HTTP + SQL + migrations + locks; the AvaTax SDK against a stub | Every push · ~45 s |
| Contract | Snapshot (approval) tests | Messages and routes don't break other teams | Every push |
| End-to-end | Playwright | A real browser can use the playground and VPW; the API works over the wire | Every push · post-deploy |
| Accessibility | axe-core in Playwright | WCAG 2.1 AA: labels, contrast, ARIA | With E2E |
| Performance | Budgets, k6, Server-Timing, OpenTelemetry, GitHub schedule | No slow-downs ship; the demo stays up and fast | Every push · on demand · every 6 h |
Unit tests 163
Small, fast tests of one rule at a time with no network and no real database. They are where we pin down the money logic: rounding, which rate a committed line gets, what happens when Avalara is down.
What's covered
- Tax documents: create, replace, commit, credit, void, idempotent replays, PendingTax fallback
- FX: daily rate table, same-day lock sharing, weekend/holiday fallback, rounding
- Every request validator (field-level messages, camelCase paths)
- The simulated tax engine and rate source the demo relies on
- Sales-tax, by-jurisdiction and vendor-tax reports
- The Bank of Canada client at the HTTP boundary (stubbed handler)
- The job scheduler with a fake clock: no real waiting, no flakiness
Example
[Fact]
public async Task Replaying_the_same_request_is_idempotent()
{
await _host.SeedEntityAsync();
var (first, _) = await _host.Get<TaxService>().CreateAsync(Invoice(), Ct);
var (second, outcome) = await _host.Get<TaxService>().CreateAsync(Invoice(), Ct);
outcome.ShouldBe(TaxDocumentOutcome.Unchanged);
second.Id.ShouldBe(first.Id);
_host.Provider.Requests.Count.ShouldBe(1); // Avalara called once
}
dotnet test --project tests/Integrations.UnitTestsHow to add one
- Find the class under test in
tests/Integrations.UnitTests/<Area>(or add a file named<Thing>Tests.cs). - Name the test as a sentence about behaviour:
Committed_documents_cannot_be_replaced. - Arrange with
TaxTestHost(in-memory DB, fake providers, fake clock), act once, assert with Shouldly. - Pass
TestContext.Current.CancellationTokento async calls (the analyzer enforces it).
Architecture tests 11
The design rules, written as tests. If someone adds the "quick" reference that breaks the layering, the build fails with the name of the offending type.
- Domain depends on nothing: no EF Core, ASP.NET, MassTransit or Avalara
- Contracts are self-contained (other teams consume them as a package)
- Application reaches the outside world only through its own ports
- The Avalara SDK is only used inside
Providers.AvaTax, so a second provider is a new adapter, not a rewrite - Adapters stay
internal; validators end inValidator; jobs end inJob - A control test proves the checker sees real dependencies (no rule can pass by accident)
[Fact]
public void The_avalara_sdk_is_only_used_inside_the_avatax_adapter()
{
foreach (var assembly in new[] { Domain, Contracts, Application, Infrastructure, Api })
ShouldPass(Types.InAssembly(assembly).ShouldNot().HaveDependencyOn("Avalara"));
}
Integration tests 45
The whole service in-process, with a real PostgreSQL in a throwaway container and a fake AvaTax server. This is where mocks stop lying: real SQL, real migrations, real advisory locks, real JSON over HTTP.
Building blocks
- WebApplicationFactory hosts the API in memory; tests call it over HTTP like a client
- Testcontainers starts
postgres:17once per run; each test class gets its own database - WireMock.Net plays AvaTax, with responses in the documented v2 shapes, so the real SDK, retries and error mapping run
- A test auth handler issues scoped callers, so permissions are tested too
Highlights
- Two requests racing for the same FX lock get the same rate (Postgres advisory lock)
- An Avalara outage saves the bill as PendingTax, and the retry recovers it under the same document code
- Vendor sync sends only code, name, address and email, never proprietary data
- Redelivered messages don't create duplicate documents (inbox/outbox)
Example: the AvaTax adapter against a stub
Stub("PUT", "/api/v2/companies/1001/vendors/V-100", 404, NotFound);
Stub("POST", "/api/v2/companies/1001/vendors", 200, CreatedVendor);
var response = await client.PutAsJsonAsync(
"/api/v1/entities/PVC/vendors/V-100",
new UpsertPartyRequest("ABC Plumbing", Columbus, "ap@abc.example"));
(await response.ReadAsync<PartyDto>()).SyncStatus
.ShouldBe(PartySyncStatus.Synced);
dotnet test --project tests/Integrations.IntegrationTests (needs Docker)Contract tests
PVC Connect's legacy app and the Modernization services depend on our messages and routes. Snapshot ("approval") tests make every change to that surface a deliberate, reviewed decision.
Message contracts
All 38 event, command and DTO shapes (property names, types, nullability, enum values) live in
message-contracts.approved.txt. A rename or removal fails the build.
HTTP surface
Every route, its query parameters and response codes, read from our own OpenAPI document, live in
api-surface.approved.txt.
Auth sweep
Every documented endpoint is called anonymously and must answer 401. A forgotten
RequireAuthorization can't ship.
When a snapshot test fails
- Open the
.received.txtfile next to the approved one and diff them. - Additive (a new optional field or route): copy it over the approved file and commit both.
- Breaking (rename, removal, type change): add a new message or route version instead, and retire the old one after consumers move.
End-to-end tests with Playwright 34
A real Chromium drives the deployed pages and calls the API over the wire, like a user would. Slowest layer, so it covers the journeys that matter, not every rule (those live lower in the pyramid).
Suites (tests/e2e)
- playground: key gate (locked, wrong key, unlock, remembered), FX rate, tax estimate, record + credit, validation errors, reports, the guys
- api: health, 401 without a key, Server-Timing, estimates, idempotent documents, problem details; safe as a post-deploy smoke test
- vpw: the Vendor Portal Wizard summons three vendors and sends them back
- accessibility and responsive (Pixel 7): see below
No retries: a flaky E2E test gets fixed, never retried, because retries hide exactly the bugs these tests exist to catch. Failures keep a trace, screenshot and HTML report.
Example
test("records a document, then credits it", async ({ page }) => {
await unlock(page);
await page.locator("#document-form .send").click();
await expectStatus(page, "document", 201);
await expect(page.locator("#lastDoc")).toContainText("Committed");
await page.locator('[data-doc-action="credit"]').click();
await expectStatus(page, "document", 201);
});
cd tests/e2e && npm ci && npx playwright install chromium && npx playwright testE2E_BASE_URL=https://avatax.bowers.cloud E2E_API_KEY=… npx playwright test apiAccessibility & layout
axe-core scans each page (locked, unlocked, VPW and this guide) for WCAG 2.1 A/AA problems and fails on anything serious. Phone-sized runs check nothing scrolls sideways.
It earned its keep on day one: it caught low-contrast POST/PUT badges and an invalid ARIA label on the portal,
both fixed. The same day, the API suite caught validation errors coming back as CounterpartyCode
instead of the counterpartyCode the client sent; that's fixed too.
Performance testing & monitoring
Four layers, from "did this commit make it slower?" to "is the demo up right now?".
1 · Budgets in CI
200 requests at 8 concurrent through the full stack (auth, validation, EF Core, Postgres). The build fails over budget.
| Scenario | p50 | p95 | Budget |
|---|---|---|---|
| Tax estimate | 8 ms | 30 ms | 250 ms |
| Record a document | 43 ms | 121 ms | 400 ms |
Measured on a dev container; CI prints its own numbers in the job summary.
2 · Load test (k6)
tests/load/smoke.js ramps to 10 virtual users for a minute against any deployment, checking
rates and estimates. Thresholds: under 1% errors, p95 under 800 ms / 1.2 s. Run it from the
load-test GitHub workflow or locally:
k6 run -e BASE_URL=… -e API_KEY=… tests/load/smoke.js
3 · In the service
Every response carries Server-Timing: app;dur=… (see it in dev tools). Requests over 1 s
are logged as warnings. OpenTelemetry records request-duration histograms, provider call counts and
deferred-tax counts, exported to any OTLP backend (Grafana, Datadog, Azure Monitor).
4 · Synthetic monitoring
The monitor workflow wakes avatax.bowers.cloud every 6 hours, checks /health/ready and
times 20 warm requests. Down, unhealthy or a p95 over 1.5 s fails the run and notifies the repo.
Quality gates & CI
Coverage gate (a ratchet)
| Project | Now | Floor |
|---|---|---|
| Domain | 99% | 95% |
| Application | 95% | 90% |
| Contracts | 98% | 90% |
| Api | 91% | 85% |
| Infrastructure | 87% | 80% |
| Providers.AvaTax | 80% | 75% |
| Total | 92% | 88% |
Floors only go up. build/coverage-gate.py merges unit and integration coverage and writes this table into every CI run's summary.
Always on
- Warnings are errors; .NET analyzers at
latest-recommended - Vulnerable NuGet packages (including transitive) fail the build
- Dependabot opens weekly update PRs (NuGet, npm, Actions, Docker), and every one runs this whole pipeline
- Deterministic tests: fake clocks, per-class databases, no sleeps, no retries
- Health checks:
/health/live,/health/ready(database),/health/providers
What's next
Good practices we'd add as the service grows. None of them block today.
- Avalara sandbox contract suite (opt-in): the same tax, vendor and certificate scenarios against
sandbox-rest.avatax.comonce credentials arrive, run nightly. - Mutation testing (Stryker.NET) on the tax and FX core, to prove the tests would notice a flipped
>or a dropped rounding call. - Static security analysis (CodeQL) if GitHub Advanced Security is enabled for the repo.
- Alerting: point OpenTelemetry at Grafana or Azure Monitor, with alerts on p95 latency, error rate and PendingTax backlog.
- Load test on staging at month-end volumes (about 2,500 vendor transactions) before go-live.
- Resilience drills: inject Avalara latency and outages (Polly chaos strategies) to rehearse the PendingTax fallback.