← 2. Getting started · Contents · Next → 4. A decision
3. Architecture¶
The whole picture¶
Two worlds, joined by a single channel.
flowchart TB
subgraph SITE["Physical site"]
RD["Badge reader"] -->|"OSDP / RS-485"| EDGE
CAM["ANPR camera"] --> EDGE
QR["QR reader"] --> EDGE
EDGE["Controller<br/>(Rust)"] --> DB[("Encrypted<br/>database")]
EDGE -->|"dry contact"| BAR["Door"]
EDGE -.->|"/metrics"| MON["Site monitoring"]
end
subgraph SERVER["Server"]
MQ["MQTT broker<br/>(EMQX locally)"]
BE["Backend"]
PG[("PostgreSQL")]
KC["Identity provider"]
S3["Object storage<br/>(S3-compatible)"]
UI["Dashboard"]
MQ <--> BE
BE <--> PG
BE <--> KC
BE <--> S3
UI -->|"REST"| BE
BE -->|"WebSocket"| UI
end
EDGE ==>|"MQTT over mTLS, outbound only"| MQ
EDGE -.->|"presigned HTTPS, firmware only"| S3 Two arrows leave the site, and both are outbound. Nothing ever connects into a customer's network.
The seven components¶
The controller¶
The box on site. It reads credentials, decides, drives the relays and journals. Written in Rust, split into two layers:
- The decision logic. No I/O, no clock, no database. A pure function from the door's mode, the candidate rights, and the instant, to a decision.
- Everything that touches the world. The encrypted store, the reader bus, the relay, the MQTT forwarder, the heartbeat, OTA, the watchdog.
Why Rust? This program runs unattended, in an enclosure, for years, on modest hardware. You want no unpredictable garbage collection, a small memory footprint, and a compiler that rejects the class of bugs that take a service down at three in the morning. A crash here is a door that no longer opens.
Its local store is an encrypted database (SQLCipher). Encrypted because the box is physically reachable: someone can open the enclosure and take the SD card, which would otherwise hold badge numbers and plates.
It holds hashes rather than values, so even decrypted it does not hand over a list of everyone's plates.
It keeps only active rights (the current, in-force revision of each one, as opposed to a revision a passage superseded or a right that was revoked), never history. Its database is a cache, rebuildable from the server.
The one exception is the queue of pending passages. That is the one piece of data the controller is the sole holder of, for the duration of an outage.
The broker¶
The post office between controllers and the server. An MQTT broker, chosen for the guarantees the protocol itself makes rather than for one vendor's implementation of it.
Why MQTT rather than HTTP? Three reasons. The controller connects outwards and keeps the connection open, so no inbound port has to be opened on the customer's router. The protocol handles reconnection and quality of service natively. And QoS 1 guarantees a message is delivered at least once, which is exactly what replaying a backlog of passages needs.
Any broker that speaks MQTT with shared subscriptions can hold this role. Locally, that is EMQX. In production, it can be the same EMQX, or a managed service such as AWS IoT Core. No application code knows which one is running underneath.
The backend¶
The brain on the server side. It serves the REST API, consumes MQTT messages, writes to PostgreSQL, pushes the live feed to the dashboard, keeps the audit trails, and runs the scheduled maintenance.
It is stateless and leaderless. No instance is "the leader": start two, and MQTT shared subscriptions spread the messages between them, with each event handled exactly once.
make scale # starts a second instance and proves it
Scheduled work runs in every instance: the offline sweep, the firmware convergence, the retention purge, and the two reconciliation passes. Each takes a lease in sweep_leases before it starts. Whoever takes the lease does the work. The others move on.
The retention purge takes its lease itself rather than being handed one by the schedule, because it can also be asked for over the API. A purge run by hand competes for the same lease as the scheduled one, and hears that it lost rather than running beside it.
The lease lasts as long as the interval between runs of that sweep, and it is claimed and released in single statements. Nothing sits inside an open transaction while a sweep runs — an open one holds back the snapshot horizon, and autovacuum would stay off the very tables a sweep churns for as long as it took. An instance that dies mid-sweep blocks the next one no longer than the run it was already going to miss.
The rights queue is drained differently, and deliberately so: each instance claims its own batch of tasks under a lease of its own, so several instances drain the same queue at the same time, safely, with no leader and no need to take turns.
Why leaderless? A leader election adds a mechanism that can get it wrong, and a moment when nobody is leader. Shared subscriptions and leases move that problem into components whose job it is.
The database (PostgreSQL)¶
The memory. Rights, passages, audit trails.
The three journal tables are partitioned by month. Deleting a month of data then becomes instant: it is just dropping a partition, instead of a DELETE over millions of rows. That is what makes retention actually enforceable.
They are also immutable at the database level. Three roles exist:
| Role | Used by | Privileges |
|---|---|---|
wardn_owner | Migrations and the seed | Owns everything |
wardn_app | The backend's ordinary pool | SELECT, INSERT on the log tables. Full access on the business tables |
wardn_maintenance | Retention and erasure only | Adds the ability to drop partitions and to rewrite a log row |
The application connects as wardn_app, which has UPDATE and DELETE revoked on activity_logs, system_audit_logs and control_plane_audit_logs.
Why at the database level? A rule enforced by code is a rule a future developer will work around without meaning to. A permission refused by PostgreSQL produces an immediate, visible error.
make verifyasserts it.
Identities (OIDC)¶
Humans sign in to the dashboard through OpenID Connect, authorization code flow with PKCE, against any standards-compliant provider. Keycloak is the default. Two roles travel in the token:
wardn-user. The default role, held by every operator.wardn-admin. Adds one thing: pushing firmware to a controller.
The backend accepts any token the configured provider signed, and enforces exactly one role check: wardn-admin on the OTA endpoint. Everything else is open to any authenticated operator, and recorded in the audit trail instead. Configuring a provider other than Keycloak is covered in 10. Security.
Why so few roles? A fine-grained permission matrix is a document nobody maintains and everybody works around with an exception. Here everything is open and traced, except the one gesture whose mistake is most expensive to undo: pushing firmware to a fleet.
Machines use API keys, and a key is confined to a single workspace. See 9. The API and 10. Security.
Object storage¶
Firmware artefacts and technical documents. Anything that speaks the S3 API can hold this role: MinIO locally, Amazon S3 or a compatible service in production.
The backend never carries a firmware: it signs a temporary link and the controller downloads it itself, directly. A multi-megabyte binary has no business passing through an API's memory.
The backend speaks only the little of the S3 protocol it needs: read an object, list a prefix, presign a GET. It does this with a hand-written signature rather than a vendor SDK, which is what keeps it portable across any S3-compatible store.
The dashboard¶
A single-page application. It calls the API over REST and receives the live feed over a WebSocket. It is configured at container start, not at build: the image writes a config.js from its environment, holding the addresses a browser dials, so the same built image runs everywhere. See 13. Developing for how it is put together, and 8. The dashboard for the screens themselves.
The message channels¶
Every topic lives under one tree: wardn/devices/{deviceId}/{channel}.
flowchart LR
subgraph UP["Controller → Server"]
E1["events<br/>passages, QoS 1"]
E2["heartbeat<br/>state + presence, QoS 0"]
E3["debug<br/>raw frames on demand, QoS 0"]
E4["ack<br/>command and delta receipts, QoS 1"]
E5["rights/request<br/>send me everything, QoS 1"]
end
subgraph DOWN["Server → Controller"]
D1["commands<br/>open, restart, debug, OTA…, QoS 1"]
D2["rights/sync<br/>rights deltas, QoS 1"]
end | Topic | Direction | QoS | Payload |
|---|---|---|---|
wardn/devices/{id}/events | controller → server | 1 | One passage. Detailed in chapter 6 |
wardn/devices/{id}/heartbeat | controller → server | 0 | The controller's state. Chapter 7 |
wardn/devices/{id}/debug | controller → server | 0 | One raw diagnostic signal. Never stored |
wardn/devices/{id}/ack | controller → server | 1 | {refId, status, revision?, error?, timestamp} |
wardn/devices/{id}/rights/request | controller → server | 1 | Empty. "Send me my whole set" |
wardn/devices/{id}/commands | server → controller | 1 | {commandId, action, parameters}. Chapter 7 |
wardn/devices/{id}/rights/sync | server → controller | 1 | {syncId, revision, fullSync, set?, operations[]}. A full set too large for one message arrives in several, each carrying set. Chapter 5 |
wardn/internal/broadcast | backend ↔ backend | 0 | What one instance tells the others to show live |
Why those quality levels. Passages, commands and deltas are QoS 1: the broker holds them for a controller that is away, and losing one is not recoverable. The heartbeat and debug are QoS 0: a heartbeat held by the broker and delivered in a burst on reconnect would describe a controller as it was minutes ago, which is worse than the gap it is trying to fill.
The backend subscribes to the controller-to-server topics through a shared subscription ($share/<group>/wardn/devices/+/…), so each message reaches exactly one instance. It subscribes to wardn/internal/broadcast outside the group, because there every instance must receive it. That is how a passage recorded by instance A reaches a dashboard connected to instance B.
When the backend acknowledges. QoS 1 is only worth its name if the acknowledgement means the work is done, so the backend answers the broker once a message is handled rather than when it arrives. A passage that could not be written leaves its message unacknowledged and the broker keeps it — which is the only thing standing between a database that is briefly away and a passage nobody ever sees, since the controller's own copy is gone the moment the broker takes it. Holding it across a restart needs the persistent session a stable INSTANCE_ID gives (12. Operating).
Two things follow. Ingestion is serial per instance: a backend that cannot write is one that stops reading, which is precisely what the shared subscription lets the other instances absorb. And a payload nothing will ever parse is acknowledged all the same — redelivering it reaches the same refusal for ever, while filling the window the broker delivers through.
Idempotence, end to end¶
Every passage carries a eventId generated by the controller. Replaying it creates no duplicate: activity_logs has a unique index on (event_id, occurred_at) and the insert says ON CONFLICT DO NOTHING.
That is what makes MQTT's "at least once" acceptable. Rather than chasing an "exactly once" nobody knows how to guarantee simply, the duplicate is made harmless.
The same idea runs through the rest:
- A rights delta carries a
revision. A controller ignores anything not ahead of what it already holds, so a redelivery is inert. - Provisioning a right is idempotent on
(zoneId, externalId). make up, the migrations and the seed are idempotent, and CI asserts it.
The principles that keep coming back¶
The server never decides an opening in real time. The subject of chapter 1, and the reason for the rest.
Communication is outbound only. Nothing connects into a site. No port to open, no fixed address to negotiate with a customer's IT department.
Access rights are immutable. A right is never edited: a new one is created that supersedes it, and the old one is deactivated. History stays readable, and "why could this person get in that day?" always has an answer.
Instants are stored in UTC. The timezone is separate data: each zone carries its own. A passage is stored in UTC and displayed at the site's local time. Support looking for "the badge at 8am" sees the same hour as the person on the phone.
Hardware sits behind an interface, and each device says which one. A reader declares its wiring: an address on the OSDP bus, two pins on the controller's own header, or the controller itself. A door declares its own: a relay terminal, or a node on that same bus. The decision engine reads neither — it is handed a frame and it opens a door, and what carried either is the adapter's business.
That is also what makes a controller and a reader the same box when a site wants them to be: a controller opens every wire its topology puts a door or a reader on, and several of them run side by side.
Everything is an environment variable with a documented default. No constant is hardcoded where an operator might need to change it. The complete list is in 12. Operating.
Local versus production¶
| Local | Production | What changes |
|---|---|---|
| EMQX | AWS IoT Core, or any MQTT broker with shared subscriptions | Nothing: same protocol, same QoS, same shared subscriptions |
| MinIO | S3, or any compatible object store | Nothing: same API |
| PostgreSQL in a container | Managed PostgreSQL | Nothing |
| Keycloak in a container | Keycloak or any standards-compliant OIDC provider | Nothing |
| Development CA on a shared volume | Manufacturing PKI | A controller's private key never leaves its secure element |
| Plaintext MQTT on 1883 | mTLS on 8883 | A URL and three file paths |
| A door's relay reported in the log | The contact on GPIO_CHIP, or the door's own OSDP node | Nothing in the door: its wiring already said which. An empty GPIO_CHIP is what makes a laptop report instead of drive |
| Simulated reader (text over TCP) | OSDP readers, readers on the header, or both | SIMULATION_LISTEN unset: no TCP port, and each wire opens when the topology uses it. The business logic is identical |
| The controller's SQLite key from the environment | Key from a secure element | ⚠️ Not implemented. See 10. Security |
You know the pieces. Now let us follow a badge, from the read to the door.
Next → 4. A decision