← 6. Going offline · Contents · Next → 8. The dashboard
7. The fleet¶
Controllers are on customers' sites, sometimes in another country, behind their router. Everything you can do to them, you do from a screen.
The heartbeat is the report¶
Every ten seconds, each controller publishes what it knows about itself. There is no separate ping.
Why one signal? Because "it is alive" and "here is how it is doing" can then never disagree. A controller that answers a ping while its queue is full and its readers are dead is a controller that looks healthy when it is anything but.
What a controller reports¶
{
"firmwareVersion": "v0.2.0",
"osVersion": "Debian GNU/Linux 12 (bookworm)",
"hardwareModel": "UniPi Patron",
"ipAddress": "10.42.0.17",
"simulationListen": null,
"freeMemoryKb": 381244,
"diskUsagePercent": 43.18,
"certificateExpiresAt": "2028-11-02T09:31:00Z",
"pendingEvents": 0,
"queueSaturated": false,
"droppedEvents": 0,
"droppedTelemetry": 0,
"clockTrusted": true,
"readersDown": [],
"readerBusStopped": false,
"readerStates": [{ "address": 1, "up": true }, { "address": 2, "up": true }],
"accessRightsHash": "618c4da9…",
"accessRightsCount": 128,
"appliedRevision": 42,
"logLevel": "info",
"debugUntil": null,
"timestamp": "2026-08-10T07:14:20.001Z"
}
| Field | Why it is there |
|---|---|
firmwareVersion | What it is actually running: the release it installed, not the one it shipped with |
osVersion, hardwareModel | The distribution name and what the device tree says the box is |
ipAddress | Discovered by asking the routing table which interface carries outbound traffic. No packet is sent |
simulationListen | Where the controller also accepts credentials as text over TCP, from anyone who reaches it (SIMULATION_LISTEN). Null on a cabinet. The controller's page shows it in red so a bench setting is not left on |
freeMemoryKb, diskUsagePercent | The partition holding the queue, which is the one whose fullness actually matters |
certificateExpiresAt | The instant, not the days left. Counting days needs a clock, and that is the one thing the controller cannot vouch for |
pendingEvents | How many passages are still waiting for the server |
queueSaturated | The controller's own verdict: it knows its own capacity |
droppedEvents | Cumulative since commissioning. Non-zero means the record has a hole somewhere |
droppedTelemetry | Telemetry samples lost to the same bound. Non-zero means a gap in the history charts that will never fill |
clockTrusted | Whether wall time is consistent with what it has already lived through |
readersDown | Which bus addresses stopped answering. The controller is the only party that knows what is actually wired to it |
readerBusStopped | The poll loop itself is over. Distinct from the above: a stopped bus reports no reader down, and would otherwise look like a site nobody happens to be badging at |
readerStates | Every address polled at least once and whether it last answered — kept on file, so a reader's or a peripheral door's own screen can show it without waiting for an alert |
accessRightsHash, accessRightsCount | A fingerprint of what it holds, so drift is visible without shipping the whole set back for comparison |
appliedRevision | The last rights update it confirmed applying |
logLevel, debugUntil | What it is currently doing about diagnostics |
The server keeps the latest report per controller, marks the controller online, and broadcasts it so any open dashboard updates live.
A report from a controller the server does not know about is dropped: a box that has not been declared first must not be able to create its own row just by talking to the server.
Going offline¶
A sweep runs every few seconds, checking every controller marked online or updating. Any of them that has missed three heartbeats in a row — about thirty seconds of silence — gets flipped to offline, with a note in the technical journal and an alert. The note carries the last heartbeat the server saw, which is when the outage actually began rather than when the sweep noticed it. The next heartbeat writes the matching note that it is reporting again, and the two together are one outage: what the controller's own screen draws as its availability, read straight from the journal rather than from a table of its own.
Why three missed heartbeats and not one? A single lost heartbeat on a 4G link says nothing on its own, and a dashboard that flickers between online and offline teaches operators to stop trusting it.
The sweep only believes silence it has witnessed. A server that was asleep, or just restarted, finds every heartbeat stale when it comes back, and none of that staleness is the controllers' doing. So the sweep keeps the instant it last ran: after a gap it notes in the technical journal that nobody was watching (sweep_resumed, with how long), then watches for a full thirty seconds before marking anybody. A controller that is alive has reported by then. One that is not is marked thirty seconds later than it otherwise would have been, with its outage still dated from its real last heartbeat.
Silence is the only evidence available. A controller that lost power, lost its uplink, or crashed all look exactly the same from here — and all three are things an operator should be told about, rather than left staring at a status that says "online" when it no longer means it.
The statuses¶
| Status | Meaning |
|---|---|
provisioning | Declared, not yet reporting |
online | Reporting on time |
offline | Silent past the threshold |
updating | A firmware update was pushed. It is expected to go quiet for a while |
maintenance | Deliberately taken out of service |
updating exists so the silence of a firmware install does not read like a fault. A controller that never comes back from one is still swept to offline like any other, once enough time has passed.
Alerts¶
The system raises nine named conditions, each written to the technical journal and, where that has been set up, forwarded to whoever is on call — because the people who get paged at 3am don't read the database directly.
| Kind | Raised when |
|---|---|
controller_silent | A controller stopped reporting |
reader_error | Readers stopped answering, or the bus stopped polling altogether |
queue_saturated | The local queue crossed its warning level |
events_dropped | The queue was full and passages were lost |
telemetry_dropped | The telemetry queue was full and samples were lost |
clock_untrusted | Wall time is behind an instant the controller already lived through |
sync_failed | A rights change was refused, or never delivered after several tries |
rights_drift | The fingerprint still disagrees after repeated full syncs |
certificate_expiring | A controller's certificate lapses within the warning window, or already has |
firmware_update_failed | A controller is still not running the release it was assigned, after every attempt |
topology_drift | A controller holds a different number of doors or readers than it was configured with |
commissioning_incomplete | Doors declared with nothing wired to them, devices answering on the bus that nothing owns, or a peripheral that moved onto an address already taken |
The last two are not the same kind of problem, and the split is deliberate. topology_drift is a fault: something was sent and did not land, and nothing else would have noticed, because a controller acknowledges a topology when it applies it rather than when it applies all of it. commissioning_incomplete is work to finish — the closest thing to confirmation a dry contact can produce, which is to say the absence of it. Both are computed from the inventory each controller reports alongside its heartbeat: how many doors and readers it is actually holding, how many of those doors are wired to nothing, and which bus addresses answered that it has no reader for. commissioning_incomplete has one source outside the heartbeat: a peripheral confirming a move onto an address something else on that bus already holds.
That last list is the only discovery channel there is. Nothing sweeps the bus outside commissioning, so a device nobody declared makes itself known by being used: its frame is refused, its address is remembered, and the dashboard offers it as something to assign rather than only complaining about it.
The same condition is raised once, then stays quiet for about an hour. A controller reports every ten seconds, so without that pause a saturated queue alone would raise hundreds of alerts an hour and drown in its own noise. Long enough to stop the flood, short enough that a condition nobody has fixed yet gets mentioned again. Two conditions are exceptions, because the journal records their end: a silent controller's outage ends with the heartbeat that brings it back, and a clock the controller stopped vouching for ends with the report in which it vouches again. A second outage inside the hour is then a second alert rather than a repeat of the first, and an alert about a clock NTP fixed a minute after boot says so instead of reading like one still wrong an hour later.
Alerts are read-only. One is not acknowledged by hand — it simply stops being raised once whatever caused it goes away. Anything else would leave two different answers to "is this still a problem" floating around at once.
An alert whose controller cannot be resolved to a customer is shown to signed-in operators, never to a key scoped to a single customer. Withholding an alert from someone is a nuisance. Showing one customer's controller trouble to another customer is a breach.
Telemetry history¶
The heartbeat above is a liveness signal: it says "alive, and here is how it's doing" every ten seconds, and nothing before it is kept. Telemetry is a different signal with a different job — a system-metrics sample taken every TELEMETRY_INTERVAL_S (60s by default) and kept for TELEMETRY_RETENTION_DAYS (5 by default), which is what the fleet's history charts are drawn from.
{
"sampleId": "9f2b7c1e-4a10-4d3b-8e55-0c1f2a3b4c5d",
"sampledAt": "2026-08-10T07:14:00.000Z",
"clockTrusted": true,
"cpuLoad1m": 0.12,
"uptimeSeconds": 384712,
"freeMemoryKb": 381244,
"diskUsagePercent": 43.18,
"networkType": null,
"networkSignal": null,
"ambientTemperatureC": 21.6,
"readersDown": []
}
Where the heartbeat is fire-and-forget (QoS 0, never queued — a controller with nothing to say still has to say it, and a stale one delivered late would be worse than a gap), a telemetry sample goes through the same Store & Forward queue as a passage: bounded, kept until acknowledged, and delivered at QoS 1 with the instant it was actually taken.
Why bother queuing it? Because it is the one way to tell "the controller was up but had no internet" from "the controller was down", after the fact. A gap the network caused gets filled in once the link comes back, stamped with when it really happened; a gap a dead process caused never fills, because nothing was ever running to queue a sample in it. The server computes
receivedAt - sampledAton arrival rather than trusting a flag the controller sets — the same reasoning asisOfflineBatchon passages (chapter 6), applied to a chart instead of a log.
sampleId is minted once, when the sample is taken, and a redelivery repeats it verbatim. It is what the server deduplicates on: QoS 1 is at-least-once, and the instant cannot serve as the key because the server rewrites it whenever the controller admits its clock is wrong — which would leave exactly the degraded controllers this stream exists to watch inserting a fresh row per redelivery.
sampledAt is read off the controller's own clock, which is exactly what it cannot always vouch for — the same problem the heartbeat's clockTrusted already covers. A sample rides in with the same verdict, and the server does not trust a sampledAt from a clock it was told is wrong: a controller that rebooted with a dead battery would otherwise backfill a chart at whatever date it woke up believing.
networkType and networkSignal are placeholders today: reading a cellular modem's signal quality needs an integration this firmware does not have yet, so both report null rather than a guess. ambientTemperatureC is the cabinet's own probe, not the CPU's — also null on hardware with none wired.
Fetched per controller with GET /devices/{id}/telemetry?from=&to= (chapter 9).
Commands¶
A command is published to the controller, and acknowledged by it once applied. Nothing waits for that acknowledgement: a controller on a 4G link can be a minute away, and asking someone in the dashboard to wait that long is not acceptable.
sequenceDiagram
participant O as Operator
participant B as Backend
participant M as Broker
participant E as Controller
O->>B: Send a command from the dashboard
B->>B: Recorded: who asked
B->>M: Command published
B-->>O: Confirmed, tracked by its own id
M->>E: Delivered (held while offline if needed)
E->>E: Applied
E->>M: Acknowledged
M->>B: Recorded: what became of it The commands¶
| Action | What it does |
|---|---|
| Change the log level | debug · info · warn · error, applied without restarting |
| Open or close live diagnostics | Opens or closes the live diagnostic stream, always time-boxed |
| Ask for a full resync | Makes the controller ask for its whole rights set again |
| Blink its LED | The box blinks its own LED for ten seconds, so the one in front of you is the one on screen. Refused where there is no LED |
| Open now | Pulses a door's relay immediately |
| Change a door's mode | normal · unlocked · free_access · locked |
| Restart the service | Stops the process. The supervisor brings it back |
| Reboot the machine | A full reboot, where the hardware has an OS to reboot |
| Install an update | Fetches, verifies and installs a firmware release |
An action a given controller's firmware doesn't implement is rejected, not retried — replaying it would only fail the exact same way again.
A door is always addressed by the door, and routed automatically to whichever controller actually drives it: an operator on the phone with a driver knows which lane they're standing at, not which controller happens to be behind it.
Restarting versus rebooting¶
They are separate actions because they are separate costs: a few seconds, against a site with no controller at all for a minute. Collapsing them into one button would make the cheap remedy feel as risky as the expensive one.
- Restarting the service stops the process and lets it come back on its own, exactly as it would after a crash. The local rights survive, because they live in the controller's own database, not in memory that a restart would wipe.
- Rebooting the machine is refused wherever there is no actual operating system to reboot — inside a container, for instance. Quietly turning that into a service restart instead would let an operator believe a box came back clean when nothing was actually rebooted.
In both cases, the acknowledgement is sent and given a moment to actually reach the broker before the controller does what it was told — an operator must never see a controller go quiet with no way of knowing whether their command got through first.
The controller can also restart itself, entirely unasked: a topology update (see chapter 5) that adds or removes an address on the reader bus changes what that bus should be listening for, and that list is only re-read when the software starts. The same short grace period applies — an operator watching the fleet screen sees the change acknowledged, then a brief offline, never a controller that just went silent for no reason.
The door mode lives on both sides¶
The server keeps a copy so the dashboard can show it. The controller keeps its own copy so an uplink outage cannot change how a door behaves — a locked lane stays locked even with the network gone.
Changing it from the dashboard updates both: the server's own record, and the controller itself, which is where it actually takes effect.
Reaching a device directly¶
Every command above is addressed to a door, or to the controller as a whole. Some hardware needs a third target: one reader on its own, or a door that is itself directly addressable on the bus, reached without touching anything else sharing the same line.
| Output | What value means |
|---|---|
| Relay | Milliseconds to pulse for, or "on" / "off" to hold it |
| LED | A colour: red, green, amber, or off, shown for ten seconds before the device's own colour comes back |
| Buzzer | How many times to beep |
| Text | A message of up to four lines for the device's own screen, where it has one. It replaces the whole screen |
| Reset | Nothing: the device goes back to rest, its own colour, a silent buzzer and its resting screen |
Nothing sent from here sticks. A colour that stayed would leave a door showing green while it is locked, long after whoever set it has moved on. A message stays until the next one, or until Reset puts the resting screen back. On the wire a message is one TEXT per line, from the top: a line on the first row starts a new message and clears the rows below.
Sent, and only actually applied once the device's own turn comes up. Unlike opening a door, which the controller answers immediately, this queues behind whatever the shared bus is already doing. The acknowledgement can lag a little behind the request as a result.
A door's own version of this is refused if that door has no address of its own on the bus: without one, it is only reachable through its relay, which opening it directly already covers.
Asking a peripheral about itself¶
The commands above tell a peripheral to do something. Three more ask it something, and are answered with a report rather than an acknowledgement. They are what commissioning a cabinet actually needs.
| Question | What comes back |
|---|---|
| What is it? | Vendor, model, version, serial number and firmware, as the unit itself reports them |
| What can it do? | Every function it claims, and how many of each: outputs, LEDs, readers. For the receive buffer, its size in bytes |
| Move to another address | The address and baud rate it actually took |
The first two settle arguments an installer otherwise has no way to settle: whether the reader at address 3 is the model the drawing says, and whether a unit really has the screen the Text button is driving. The answers are kept, so a bus that is down later still shows what it answered when it was up.
Moving a peripheral is the one that changes the system. A controller polls the addresses its configuration names, so a reader that moved and a record that did not is a reader gone silent on a bus it is perfectly well connected to. The order here is deliberate: the command goes out, the peripheral confirms the address it took, and only then does the record follow and the controller start polling it there. A command that never lands leaves nothing out of step.
The peripheral named in that confirmation has to be one the answering controller actually owns. A controller identifies itself by the topic it publishes on, which it cannot forge; the device it names sits in the message body, which it can. Naming somebody else's reader is refused and logged, not followed.
The address it names has to be free on that controller's bus. One address answers for one device there, door node or reader alike, and a confirmation landing on an address something else already holds is a collision on the wire itself, two units answering the same poll. Nothing here can undo a move the hardware has already made, so the record is left as it was and a commissioning_incomplete alert says which device took which address. Somebody at the cabinet decides which of the two moves.
Why no key installation yet? OSDP also defines a command that installs a Secure Channel key, and wardn's controller does not speak Secure Channel (the frames are clear OSDP). A peripheral given a key may then refuse every unencrypted command — which would make it unreachable from the very controller that keyed it. The command is implemented and tested at the protocol layer; it is deliberately not offered on a screen until there is a channel to use it with.
⚠️ Not yet proved against real hardware. These commands are built from the specification, not yet confirmed against a physical peripheral that has actually accepted one. 16. The lab suitcase is what closes that gap.
The live debug console¶
Reached from a controller's own page in the dashboard, the console streams raw reader activity, relay pulses, sync events and broker traffic as they happen — and writes none of it down anywhere: not to a table, not to a file, not to a log.
It is deliberately exhaustive while it is on. Anything the controller does that an installer or an operator might be chasing shows up under one of these types, filterable from the console's own chips:
| Type | What it carries |
|---|---|
osdp_frame, osdp | Every line off the reader bus, and every reply the bus decoded |
gpio_input, gpio_output | Each decision and each relay pulse, including remote ones |
mqtt | Connects, subscriptions, disconnects, and every packet published or received |
events, telemetry | Each queue drain: how many went out, how many were replays, how many came back unconfirmed |
heartbeat | The whole ten-second reading, exactly as it is published |
sync, error | Rights deltas applied or rejected, and readers a delta named no door for |
flowchart LR
E["Controller emits a signal"] --> D{"Is debug on<br/>right now?"}
D -->|"no"| N["Nothing happens"]
D -->|"yes"| Q["Held briefly in memory"]
Q --> P["Sent out, never stored"]
P --> B["Backend passes it on"]
B --> W["The live console in your browser"] Three things about it are deliberate:
It is always time-boxed, capped both by the server and by the controller itself — fifteen minutes by default. It is the controller paying for the data over its own connection, and it forgets the session on its own. Nobody has to remember to turn it off.
It never gets in the way of a decision. If the stream can't keep up, signals are simply dropped rather than held — waiting for a browser somewhere to catch up would put a door's own decision behind it, which is never acceptable.
It is closed to anything but a signed-in operator. Raw frames carry no information about which customer they belong to and cannot be safely confined to one, so rather than pretend to confine what can't be confined, this is refused outright to anything scoped to a single customer.
Why store nothing? A debug stream that accumulated over time would quietly become a second copy of everything a door sees, sitting outside every retention rule and outside the audit trail entirely.
Updating over the air¶
Publishing a release¶
A release is a tag. Pushing edge-v0.3.0 is what builds it: the release pipeline compiles the controller once per architecture, signs each build, and uploads them to object storage with a manifest describing the lot.
Nothing is published by hand, from the dashboard or anywhere else — an artifact somebody had on their disk is not traceable to the commit it came from, and a release that cannot be traced is one nobody can reason about later.
The firmware has its own version sequence for the same reason: the edge-v* tags are separate from anything else the repository releases, so a controller's version means one thing only.
That signature is produced by the pipeline itself, never by the server, and proves the release actually came from that pipeline — a digest alone only proves a file wasn't corrupted in transit, not who produced it (see 10. Security).
The private half of the signing key lives on the release environment alone; the job that holds it is the only one in the pipeline that can produce a signature a controller will trust.
Storage itself is the catalogue: a release exists because its manifest is sitting there, not because some separate row says so. The builds are uploaded first and the manifest last, so a run that dies half way leaves no release rather than half of one.
Any operator can see what's available straight from the dashboard — seeing that a controller is two versions behind is part of watching a fleet.
One release, several architectures¶
A controller might be a Raspberry Pi, a UniPi, or an industrial box in a rack, and a binary built for one will not run on another. So a release carries one build per target, and every controller reports which target it was compiled for as part of its heartbeat.
Binaries are linked statically against musl: they carry their own libc, so the only thing that varies between builds is the architecture, never how old the controller's distribution happens to be.
Nobody picks a build by hand. An operator chooses a version; the server resolves it to the build matching what that controller reported, and refuses outright if the release has none — pushing the wrong architecture would clear every digest check and simply fail to start.
Assigning it¶
Installing a release is the one action in the whole system reserved to an administrator; everything else is open to any operator and simply audited instead.
A release is assigned, not delivered. The server records the version a controller is meant to be running, and the controller converges on it: one that was offline when the decision was made installs the release when it comes back, and one that rolled a release back is simply seen to be drifted again.
What an operator reads on the dashboard is the gap between the two — running v0.2.0, assigned v0.3.0 — rather than whether a message happened to be delivered.
That convergence is bounded. A release is pushed to a controller a few times, spaced far enough apart to let a real install finish, and then stopped: a firmware that will not stay installed is a fault to report, not something to retry forever. The failure is called once the last attempt has had that same time to come up and the controller still reports the old release, never while that attempt may yet succeed. It raises firmware_update_failed once, and the dashboard shows the release as failed rather than pending. A new assignment starts the count again.
The controller never downloads the file through the API itself. It is handed a link straight to storage, valid for a short window, and fetches the release on its own schedule from there. The record kept afterwards is the version and its digest, not that link — a link is a credential that has to expire; a version number is just a fact.
What the controller does¶
flowchart TD
C["An update arrives"] --> SV{"Does its signature<br/>check out?"}
SV -->|"no"| SR["Rejected.<br/>Nothing is downloaded at all"]
SV -->|"yes"| D["Downloaded and checked<br/>as it arrives"]
D --> V{"Does it match<br/>what was announced?"}
V -->|"no"| R["Discarded.<br/>The running firmware is untouched"]
V -->|"yes"| S["Swapped in as one atomic step"]
S --> A["Acknowledged, then the controller restarts"]
A --> N["The new firmware starts up"]
N --> OK{"Does it come up cleanly<br/>and start serving doors?"}
OK -->|"yes"| CF["Confirmed. This is now<br/>the running release"]
OK -->|"no"| RB["Rolled back automatically<br/>on the next start"] A few guarantees make this safe to leave unattended:
- The signature is checked before a single byte is downloaded. It answers a different question than the digest does: not "was this corrupted on the way", but "did this actually come from the release pipeline at all". Anything that fails this check is refused outright.
- Nothing already running is touched until the new file has fully checked out. A failed update simply leaves the controller running exactly what it had before.
- The switch itself cannot land half-done. There is no moment where the file the controller would start from is only partially written.
- Confirmation is the boot itself. A release only counts as confirmed once the new firmware has actually come up, found its rights, and declared its doors ready to work. One that crashes before that point is rolled back automatically, with no one needing to notice and intervene by hand.
Only one update can be in progress on a given controller at a time; a second one arriving mid-update is simply refused.
⚠️ A release that changes the controller's local store refuses to start on the file the previous one left behind. That store has no migrations: it holds what the cloud already owns — the wiring, the rights, the queue — so the firmware says so in one line and stops, the rollback above puts the previous release back, and the controller keeps deciding. Moving forward means removing the store and letting the next sync refill it, which also drops whatever passages were still queued in it.
What else a controller exposes¶
Its own metrics endpoint¶
Each controller also serves a small set of live numbers over its local network — door openings so far, how many passages are still queued, and whether each reader on the bus is currently answering. Meant to be read on-site, by whatever already monitors the network there, rather than over the same connection the regular reports already use.
A controller that cannot even count its own queue still answers rather than going silent — silence there would read as the whole box being down, a much louder problem than the one it's actually having.
Links and documents¶
A controller's own page in the dashboard can carry a short list of useful links — a 4G router's admin page, a monitoring board — and documents such as a wiring diagram or a network configuration sheet. The address of the site's router at two in the morning is exactly what someone troubleshooting a controller actually needs in that moment, and remembering it by heart was never going to be their job.
The fleet is watchable. Let us look at the screen it is watched from.
Next → 8. The dashboard