← 9. The API · Contents · Next → 11. Personal data
10. Security¶
An access control system is a security product twice over: it protects a site, and it is itself a target. This chapter gathers everything that protects the system, and ends with an honest list of what does not yet.
This chapter is what the product does. What a deployment must add on top of it — TLS, HSTS, network placement, secret custody — is 18. The security agreement, which is platform-independent and ends in a checklist to sign.
Who can be, and what they can do¶
| Identity | Proves itself with | Sees |
|---|---|---|
| An operator | An OIDC access token | Every workspace |
| A machine | A workspace API key | Exactly one workspace |
| A controller | An X.509 client certificate | Its own MQTT subtree, nothing else |
| The backend | An X.509 client certificate | The whole fleet tree |
Operators¶
OpenID Connect, authorization code flow with PKCE, against any standards-compliant provider. Keycloak is the default, but Auth0, Okta, Azure AD, or a self-hosted alternative work the same way. wardn holds no password and no client secret: the dashboard is a public client, and the code exchange is bound to a verifier the browser tab generated.
Neither the backend nor the dashboard is told the provider's endpoints by name. Both read them from the one document every conformant provider serves at a fixed path, {issuer}/.well-known/openid-configuration: where to sign in, where to exchange a code, where to fetch signing keys, where to sign out.
Configuration is an issuer URL and a client id. Nothing provider-specific, and a signing-key rotation is picked up without a restart because wardn never holds one.
wardn asks two things of whatever token comes back: a stable sub claim, which becomes the operator's identity, and a roles claim (or Keycloak's realm_access.roles, read the same way) that carries wardn-user for every operator and wardn-admin for the one gesture held back: pushing firmware to a controller.
The backend enforces exactly one role check, wardn-admin on the OTA endpoint, and audits everything else.
Why so few roles? A fine-grained permission matrix is a document nobody maintains and everybody works around with an exception. Everything here is open to an operator and traced, except the one gesture whose mistake is most expensive to undo: pushing firmware to a fleet.
Configuring an identity provider¶
Two places read the provider's URL, and neither needs the other:
| Where | Variable | What it is |
|---|---|---|
| Backend | OIDC_ISSUER_URL | The issuer exactly as it appears in tokens |
| Backend | OIDC_JWKS_URL | Optional. Set only when this process cannot reach the issuer the way a browser does. See 12. Operating |
| Backend | OIDC_AUDIENCE | wardn-spa by default. See below |
| Dashboard | WARDN_OIDC_AUTHORITY | The same issuer, reached from a browser |
| Dashboard | WARDN_OIDC_CLIENT_ID | A public client: PKCE, no secret |
On the provider's side, whatever it is, the client needs:
- A public (SPA) client type. No client secret, ever.
- The authorization code grant with PKCE, challenge method
S256. - The dashboard's origin as a redirect URI and a post-logout redirect URI.
- The
wardn-adminrole reaching the access token asroles,realm_access.rolesorcognito:groups. Most providers call this a role mapping or a custom claim, configured once per client. Cognito puts a user's groups there on its own. - An audience mapper adding the client's own id to the access token's
audclaim, matchingOIDC_AUDIENCE. A Cognito access token has noaudand names the client inclient_idinstead, which is read in its place, and only whenaudis absent.
Keycloak does all five out of the box for the realm this repository imports, loaded by make up. infra/aws/ does the same for a Cognito user pool. Pointing wardn at Auth0, Okta or Azure AD instead means the same five boxes ticked in that provider's console. Nothing in wardn changes.
The operator's name is name, then preferred_username, then Cognito's username, whichever the token carries first. A Cognito access token carries no email, so the audit trail attributes an action to the username alone.
OIDC_AUDIENCE defaults to wardn-spa, checked against a token's aud claim on every request. This closes a gap the realm itself used to have: wardn-backend, a second, confidential client, already shares this realm (for its own service-account use, unrelated to operator sign-in). So a token minted for that client, or for any future one, no longer passes as a wardn operator just by sharing the issuer. Pointing OIDC_AUDIENCE at a provider whose client carries no matching audience mapper rejects every token. Confirm with a decoded access token (aud present and correct) before relying on a new one.
Machines¶
An API key is wapi_live_ plus 24 random bytes. Only a bcrypt hash is stored, computed inside PostgreSQL by pgcrypto, so the plaintext never travels further than the response that issued it.
The first sixteen characters are stored in the clear and indexed: enough to find the row in one lookup, never enough to be the key. The hash still decides. See 9. The API for why that matters.
A key can carry an expiry, and revoking it takes effect on the next request.
Bcrypt rather than a memory-hard hash. The seed is plain SQL and has to produce valid hashes without a runtime, and
crypt()/gen_salt('bf', 10)is what PostgreSQL offers for that. The trade is deliberate and worth revisiting: a random 192-bit key is not a password, so the slow-hash property is defence in depth rather than the thing standing between an attacker and the key.
Confinement between customers¶
This is the guarantee customers care about most, and it is enforced in four places rather than one.
flowchart TD
K["A workspace API key"] --> R["REST: every query carries the caller's scope"]
K --> W["Live socket: each message filtered as it is sent"]
K --> A["Alerts: only those whose controller resolves to this workspace"]
K --> F["Fleet-wide actions: refused outright"]
R --> N["A resource of another workspace answers 404, never 403"] - The filter is inside the query, not applied after it, so a caller can never be handed a row it should not have seen.
- 404, never 403. A status that distinguishes "does not exist" from "not yours" lets a client map what it cannot read.
- A message carrying no workspace reaches nobody who is confined. An unattributable payload is not a reason to widen what an API client can see.
- The debug console is closed to API keys. Raw reader frames carry no workspace and cannot be confined to one. Confining what cannot be confined is worse than saying no.
make api # asserts that Acme's key sees nothing of Northwind's
The link between a controller and the server¶
A controller sits on someone else's network and publishes to a broker on the public internet. A shared password would be one leak away from a fleet. A certificate is per device, revocable on its own, and cannot be read off a screen over somebody's shoulder.
make mtls # both ends on the mutually authenticated listener: asserted
The identities¶
tls-init issues, once, into a shared volume:
ca.crt / ca.key the development authority
emqx.crt CN=emqx, what the broker presents
backend.crt CN=wardn-backend, what the backend presents
edge-01.crt CN=edge-01, what that controller presents
flowchart LR
CA["Development CA"] -->|"signs"| E["edge-01 certificate<br/>CN=edge-01"]
CA -->|"signs"| BE["backend certificate<br/>CN=wardn-backend"]
CA -->|"signs"| MQ["emqx certificate<br/>CN=emqx"]
E -->|"connects, presents CN"| MQ
MQ -->|"username = CN"| RULES["Access rules:<br/>wardn/devices/edge-01/# only"] It is idempotent: restarting the stack does not hand every party a new identity. On a subsequent start it re-reads what is there and prints anything expired or inside the warning window. An expired authority reused in silence would make every handshake fail with no explanation of why.
The authority outlives what it signs: TLS_CA_DAYS (3650) against TLS_DAYS (825). Expiring together would mean that the day the controllers' certificates die, the key that could have replaced them is dead too. Renewal would become a recall rather than a rotation.
The two listeners¶
| Port | What it asks for |
|---|---|
1883 | Nothing. Kept for development so a laptop needs no keys |
8883 | A certificate signed by the CA, from both sides, or the handshake fails |
Moving a party across is a URL and three paths:
MQTT_URL=mqtts://emqx:8883
MQTT_TLS_CA=/tls/ca.crt
MQTT_TLS_CERT=/tls/edge-01.crt
MQTT_TLS_KEY=/tls/edge-01.key
Both ends refuse to start half-configured. A controller with two of the three variables set fails loudly rather than falling back to the plaintext port. That is precisely what whoever set them was trying to avoid.
The certificate decides which controller is talking¶
A valid certificate is not the same thing as this controller. On the mutually authenticated listener, EMQX replaces the MQTT username with the common name of the client certificate (peer_cert_as_username = cn) before consulting its access rules, so a client cannot announce somebody else's name.
The rules then confine each party to its own subtree:
%% The server sees the whole fleet, because it is the only party that has to.
{allow, {username, "wardn-backend"}, all, ["wardn/#"]}.
%% A controller gets the subtree its certificate names.
{allow, all, all, ["wardn/devices/${username}/#"]}.
%% Everything else under the fleet tree belongs to somebody else.
{deny, all, all, ["wardn/#"]}.
The placeholder is not rendered when there is no username, so a client that announces nothing matches nothing.
make mtls asserts all four claims:
- Both ends connect and carry a real passage.
- A client with no certificate is turned away at the handshake.
- One with a valid certificate is let through.
- A controller cannot reach another controller's topics even when it announces that controller's name.
On the plain
1883listener the username is self-asserted. That port is a convenience for development, not a control, and that is the whole extent of it.
Certificates fail on a date, not for a reason¶
Every certificate in a fleet is minted in the same pass, so they lapse on the same day. A lapsed certificate does not fail immediately: an established session survives its own credential, so a controller keeps working until something makes it reconnect, and then never comes back.
Three things stand between here and that morning:
- The authority is issued for longer than what it signs.
- Each controller reports the instant its certificate expires, not the days left, because counting days needs a clock and that is the one thing a controller cannot vouch for. The server does the subtraction with its own clock and raises a
certificate_expiringalertTLS_EXPIRY_WARNING_DAYS(30) ahead. - For the broker's and the backend's own certificates, which no controller reports: each backend instance reads its own certificate file directly and raises the same
certificate_expiringalert on the same schedule. The broker and the backend are minted in the same pass and share an expiry date by design, so one file's check stands in for both.
make tls-check remains available as a manual, on-demand cross-check against every certificate at once:
make tls-check # exits non-zero inside the warning window
What is stored, and how¶
On the controller¶
The local database is encrypted with SQLCipher, because the box is physically reachable: someone can open the enclosure and take the SD card.
Even decrypted, it holds no credential values, only scrypt(canonicalize(type, value)) — badge numbers and plates are too small a keyspace for a fast digest to resist brute force once the file itself is what got taken, so the lookup key both this database and the backend derive is memory-hard instead.
A stolen controller does not hand over a list of everyone's plates. The raw value exists in exactly one place: the outbound event.
It also holds no history, no names and no emails. It is a cache, rebuildable from the server, except for the queue of undelivered passages.
The encryption key comes from SQLITE_ENCRYPTION_KEY_SOURCE:
| Value | Behaviour |
|---|---|
env | Read from SQLITE_ENCRYPTION_KEY. What the local stack uses |
none | Unencrypted. Explicit, never a silent fallback |
secure_element | ⚠️ Not implemented: the controller refuses to start and says so |
There is deliberately no default key. A silent fallback would produce a database that looks encrypted and is not.
In the server¶
Three database roles, and the separation is the point:
| Role | Log tables | Business tables |
|---|---|---|
wardn_owner | owns them | owns them |
wardn_app | SELECT, INSERT. UPDATE and DELETE revoked | full |
wardn_maintenance | full, for retention and erasure only | full |
The application connects as wardn_app. An application bug cannot rewrite history, because the permission does not exist. It is not because a code path avoids it. make verify asserts the grants.
Retention drops whole monthly partitions as the owner, so the append-only guarantee costs nothing operationally. See 11. Personal data.
Secrets that pass through¶
Firmware links. The backend signs a GET valid for OTA_LINK_VALIDITY_S (900 s) and sends it to one controller. The artifact never travels through the API, and the link dies on its own rather than being a credential somebody has to remember to revoke. The audit trail records the version and the digest, not the link.
The live socket. A browser cannot set headers on a WebSocket handshake, so a credential travels in the query string one way or another: the one place it is unavoidable.
The dashboard does not put its real bearer token there: it first spends an authenticated POST /realtime/tickets for a single-use ticket, good for a few seconds and one connection attempt, and puts that in the URL instead. Whatever a proxy or load balancer logs about the handshake is not something a second attempt can replay.
A script talking to the socket directly (?token=/?apiKey=, 09. The API) still takes the same shortcut a browser is forced into, deliberately.
The audit trail and error reporting share one redaction function rather than two that drifted independently. The divergence itself was the finding: the audit trail's own copy missed email, fullName and other personal-data keys entirely, and never looked inside a string for one embedded in free text. Sentry is wired into the backend and the dashboard, and is off unless a DSN is set. A controller has no Sentry at all: it masks addresses in its own journal entries and sends them over MQTT, so a cabinet never talks to anything but its broker. Applied to the whole event/request body rather than to the places trouble is expected:
- Keys named like credentials:
authorization,cookie,x-api-key,password,token,apikey, … - Keys naming a person:
email,value,rawValue,scannedValue,rawFrame,sampleValue,plate,credentials, … The last two are the same badge or plate one step earlier: the frame a reader sent before a strategy decoded it, and the sample an operator pastes into a strategy preview to find out why a real credential is being refused. - Any email address appearing inside a string.
- Any
token=,access_token=,api_key=,ticket=in a URL. That is how a live-socket credential travels, and the kind a key-name filter would have missed.
It is the one copy of our data we do not administer.
Ceilings and origins¶
Rate limiting applies before authentication, because verifying a key or a token costs real work and an unauthenticated caller should not be able to spend it freely. Three ceilings, each keyed differently:
RATE_LIMIT_MAX(3000) perRATE_LIMIT_WINDOW_MS, the general ceiling.WS_HANDSHAKE_LIMIT(60 per address per window), the live socket's own lower ceiling, applied before the credential is even looked at.API_KEY_INVALID_ATTEMPT_LIMIT(20 per address per window), counting only failed API-key lookups, so guessing keys hits a ceiling far below 3000/minute while a valid key at any rate never does.
All three ceilings live in process memory, not a shared store. Fine for one instance. Behind a load balancer fronting N of them, the effective ceiling for a single attacker spread across instances is N times the number above, since each instance counts only what it personally saw. Not urgent for a mono-instance topology, which is this deployment's own, but the fix if that ever changes is a shared store (Redis) behind the same two limiters, not a redesign.
All three are keyed on the caller's address, so something has to decide whose word that is. TRUST_PROXY is it, and false by default: with nothing in front, the socket address is the caller, and believing a header instead would let anyone claim any address and walk past the ceilings above.
Behind a proxy the socket address is the proxy's, identical for every caller, and those ceilings collapse into one shared bucket while the audit trail records the hop rather than the origin. A deployment behind a proxy must state it — true when the proxy is the only route in, a hop count or the proxy's own addresses when it is not.
CORS origins are listed, never reflected. A wildcard would let any page an operator visits call this API with their token.
Nothing connects into a site. The controller dials out and keeps the connection open. There is no inbound port to open on a customer's router, and no fixed address to negotiate with their IT department.
/docs and /docs-json are off unless ENABLE_API_DOCS=true. Swagger mounts outside the application's normal request pipeline, so the usual authentication check never sees these routes: "reachable" has always meant "reachable by anyone with network access, unauthenticated." The local dev stack turns it back on. Production should not.
What is not there yet¶
Every known limit, in one place.
⚠️ OSDP Secure Channel is not implemented. Frames on the RS-485 bus are in the clear. The secure block is parsed and skipped so a secure peripheral does not confuse the decoder, but nothing is encrypted or authenticated: that needs an SCBK provisioned per reader and can only be trusted once it has been proved against real hardware.
⚠️
SQLITE_ENCRYPTION_KEY_SOURCE=secure_elementis not implemented. The controller refuses to start rather than pretending. Useenvwith a key from your secret store, ornoneto run unencrypted on purpose.⚠️ The simulated reader bus (
SIMULATION_LISTEN, unset by default) accepts any peer that reaches it, with no authentication. Evaluated and deliberately not restricted to "local" peers: on Docker Desktop for Mac, a connection from the host machine through the published port arrives at the container carrying a peer address (observed:172.66.152.246) that is neither loopback, RFC1918-private nor link-local by any standard classification. That is the exact traffic the local development scripts and CI's badge-simulation checks depend on. A same-host check narrow enough to reject a genuine outsider would also reject that traffic, on this platform if not on Linux's own Docker networking. This bus stands in for a physical RS-485 line and is unreachable in the shape production actually runs: real hardware speaks OSDP over a serial port, never TCP, and opens the TCP bus only whenSIMULATION_LISTENnames an address. It opens on top of the real wires rather than instead of them, so a bench can keep its readers. Left on by mistake, it would go unnoticed, since nothing stops working. So the controller logs a warning when it opens the port, and reports the address in every heartbeat, which the controller's page shows in red. Every passage that came in on it is marked simulated in the activity log and its CSV export, so a door opened that way is told apart from one opened by somebody standing at it. Bind it to one interface (192.168.1.20:9000, or127.0.0.1:9000for a test over SSH) rather than to all of them where the LAN allows.⚠️ In the local stack, private keys are files on a shared volume. That is fine for a laptop and nowhere else. In production the authority is the manufacturing one and a controller's private key never leaves its secure element.
⚠️ A self-hosted identity provider's accounts are not backed up by wardn. The default, Keycloak, has its realm imported from this repository, so what a backup would add is whatever was created afterwards. Pointing wardn at a managed provider (Auth0, Okta, Azure AD) instead removes this gap entirely: that provider backs up its own accounts. See 12. Operating.
⚠️ A transitive
devDependencystill carries a known CVE.@compodoc/compodoc(used by Storybook to generate the Design reference) pulls a vulnerableundicithroughcheerio, with no fix available upstream. It never reaches the bundle served to a browser: absent from the runtime dependency graph, and compodoc is only ever invoked for its local TypeScript-AST JSON export, never the HTTP/cheerio path that would actually exerciseundici. Exposure is bounded to a developer's machine or CI, covered by an automated gate that ignores exactly this chain's advisories by CVE — seefrontend/pnpm-workspace.yaml'sauditConfigfor the current list — not by skipping the check, so anything new here still fails the build. Tracked until compodoc ships a fix, or removing the integration becomes worth the loss of generated docs.⚠️ A person's
nameis no longer redacted by key from the audit trail or Sentry. It used to be, under the keyfullName. Renaming the column tonameput it in direct collision with the genericnameevery other entity uses — zones, doors, workspaces, controllers, API keys — and redacting that key would have blanked all of those out of the audit trail on every create or rename. Key-name redaction was dropped for this field rather than accepting either failure mode; the email-address scan across string values still catches an email wherever it appears.
The system is defended. Now: what it holds about people.
Next → 11. Personal data