← 11. Personal data · Contents · Next → 13. Developing
12. Operating¶
Everything needed to keep wardn running: deploying it, watching it, backing it up, restoring it, and the complete list of knobs.
Deploying¶
The local stack is Docker Compose. Production replaces four things with managed equivalents and changes nothing in the code:
| Component | Local | Production |
|---|---|---|
| Broker | EMQX container | AWS IoT Core, or any MQTT broker with shared subscriptions |
| Object storage | MinIO | S3 |
| Database | PostgreSQL container | Managed PostgreSQL |
| Identity | Keycloak container | Keycloak, or any standards-compliant OIDC provider |
A worked example of the right-hand column, on AWS with a Helm chart and Terraform, is 14. Deploying to production. A cheaper one, on free and near-free tiers, is 15. A near-free deployment.
The database socle¶
db-init is the single source of truth for the schema, the roles, the grants, the partitions and the seed. It is idempotent and runs on every up.
make up # migrations, partitions, seed
make verify # assert the schema, the grants and the seed
make seed # re-run migrations and seed
make reseed # re-run the seed even if already applied
Set SEED_ENABLED=false for an empty database. PARTITION_MONTHS_BACK and PARTITION_MONTHS_AHEAD control how many monthly partitions are created up front. The retention sweep keeps two months ahead from then on.
Running more than one backend¶
The backend is stateless and leaderless.
make scale # a second instance on port 3001, and the proof that it works
Two rules:
- Each instance needs its own
INSTANCE_ID, stable across restarts. Two MQTT clients sharing an id kick each other off the broker. In Kubernetes this is the StatefulSet ordinal. - Stable matters as much as unique. With
INSTANCE_IDset, the broker keeps a persistent session and queues events for an instance that is restarting. Without it, the backend takes a clean session, says so in its log, and accepts that a restart misses whatever arrives while it is down.
Everything else coordinates by itself:
- MQTT shared subscriptions spread events.
- The rights queue is drained concurrently with leases.
- The scheduled sweeps take a lease in
sweep_leases.
Certificates¶
tls-init issues a development CA and one identity per party on first start. In production the authority is the manufacturing one, and a controller's key never leaves its secure element. See 10. Security.
make tls-check # how long every certificate in the stack has left
| Exit code | Meaning |
|---|---|
0 | Nothing lapses within TLS_EXPIRY_WARNING_DAYS |
1 | Something lapses inside the window, or already has |
2 | The check could not be made (no openssl) |
The 2 exists precisely so that an inability to check does not masquerade as a certificate alarm.
Watching day to day¶
| Signal | Where to read it |
|---|---|
| A controller is offline | Fleet screen, or the controller_silent alert |
| A queue is rising | pendingEvents in the heartbeat, queue_saturated alert |
| Passages were lost | droppedEvents non-zero, events_dropped alert |
| History samples were lost | droppedTelemetry non-zero, telemetry_dropped alert |
| A clock is wrong | clockTrusted: false, clock_untrusted alert |
| A reader stopped answering | readersDown, readerBusStopped, reader_error alert |
| Rights are not landing | sync_failed, rights_drift alerts. appliedRevision against rightsRevision |
| A certificate is dying | certificate_expiring alert, and make tls-check |
| Application errors | Backend and dashboard: Sentry, when SENTRY_DSN is set. A controller: controller_panicked and controller_task_stopped in the technical journal |
| A specific request | The X-Trace-Id on its response |
| A controller, from the site's own monitoring | http://<controller>:9090/metrics |
| The backend is up | GET /health, and which instance answered |
| The database answers | GET /health/ready, or 503 when it does not |
Alerts are read through the API:
curl -H "Authorization: Bearer $TOKEN" 'http://localhost:3000/api/v1/alerts?limit=50'
The eight kinds and what raises them are in 7. The fleet.
Backups¶
make backup # database, artefacts, erasure ledger
make backup-verify BACKUP=../backups/<stamp> # restore it elsewhere to prove it
make restore BACKUP=../backups/<stamp> CONFIRM=yes
flowchart TD
B["make backup"] --> D["pg_dump, custom format"]
B --> L["Erasure ledger<br/>written BESIDE the dump"]
B --> O["Mirror the object storage<br/>(firmware, documents)"]
D --> M["MANIFEST: what this can<br/>and cannot restore"]
L --> M
O --> M
M --> ENC{"Off-site<br/>configured?"}
ENC -->|"yes, with a passphrase"| E["Encrypt, upload,<br/>prune the remote"]
ENC -->|"yes, no passphrase"| R["REFUSE: nothing<br/>leaves in the clear"]
ENC -->|"no"| K["Stays on this machine"] What is backed up¶
| Why | |
|---|---|
| PostgreSQL | Rights, passages, the audit trail. The only copy |
| Object storage | Firmware artefacts and documents. Their digests are in the database, so a database restored without them points at files that do not exist |
| The erasure ledger | The list of people erased, written beside the dump. It is what makes this more than a pg_dump |
What is not backed up¶
- The controllers' local databases: derived from the server, no history. They rebuild themselves on the next connect.
- EMQX state: sessions and retained messages rebuild themselves.
- The development certificate authority: production uses the manufacturing PKI.
- ⚠️ A self-hosted identity provider's accounts. The default, Keycloak, has its realm imported from this repository, so what a backup would add is whatever was created at runtime afterwards. That is a real gap, stated rather than hidden, and one a managed provider (Auth0, Okta, Azure AD) does not have, since it backs up its own accounts. See 10. Security.
The recovery point¶
One pg_dump a day. Up to 24 hours of passages can disappear.
That is not theoretical. The controller is built never to lose a passage locally. But once an event is acknowledged and cleared from its queue, last night's dump holds the only copy.
It is the accepted cost of not operating a write-ahead log archive. If that cost becomes unacceptable, point-in-time recovery is the way out, and it is an operations project rather than a script change.
A backup that has never been restored is not a backup¶
make backup-verify BACKUP=../backups/<stamp>
This is the step everyone skips, and the reason backups are found to be worthless on the one day they are needed.
It restores into a throwaway database beside the real one, never over it, then asserts three things a corrupt or truncated dump cannot satisfy:
- The schema comes back whole, checked by the same
verify.sqlthe socle uses. - The tables carrying the business are not empty.
- The audit trail came back with them, since a restore that loses the trail restores the data without the accountability.
Off site¶
A backup on the same machine as the database is not a backup: the fire, the theft and the failed disk take both.
BACKUP_REMOTE_ENDPOINT=https://s3.eu-west-3.amazonaws.com
BACKUP_REMOTE_BUCKET=wardn-backups
BACKUP_REMOTE_ACCESS_KEY=…
BACKUP_REMOTE_SECRET_KEY=…
BACKUP_ENCRYPTION_PASSPHRASE=…
Encrypted before it leaves, never after. With a destination configured and no passphrase, make backup refuses, rather than sending names, credentials and the history of every passage through a site to somebody else's disk in the clear because a variable was forgotten.
⚠️ A lost passphrase makes the backup as useless as its absence. It belongs neither in this repository nor in the remote store.
The SHA-256 of the content before encryption stays on site, where the remote store never saw it. Comparing a copy that comes back against that value proves what returns is what left. The cipher alone does not prove that.
The same run prunes the remote store beyond BACKUP_KEEP_DAYS. Retention off site is not the local find: without this, a 30-day commitment would hold everywhere except the one place the data had been copied to.
Backups compose with retention rather than coinciding with it. The newest backup already holds logs 30 days old, so by the time it is deleted at 30 days it has been holding data up to 60 days old. That belongs in a register of processing activities.
Restoring¶
make restore BACKUP=../backups/<stamp> CONFIRM=yes
Destructive, and deliberately hard to run by accident: without CONFIRM=yes it refuses and prints the command it wanted.
It runs in order:
- Stops the backend.
- Restores the dump.
- Re-runs
db-init(which is the source of truth for the schema, the roles and the append-only grants, and is idempotent). - Starts the backend again.
The two steps that are not automatic¶
The script prints them instead of doing them, because neither is safe to automate.
flowchart TD
R["Restore finished"] --> S1["1. Replay the erasures<br/>recorded after this backup"]
R --> S2["2. Run the retention sweep"]
S1 --> A["Collect the ids from every<br/>LATER backup's ledger"]
A --> B["Re-issue each erasure"]
B --> C{"Response?"}
C -->|"404"| D["Already erased,<br/>nothing to do"]
C -->|"2xx"| E["That person HAD come back.<br/>Count them: this number<br/>cannot be recovered later"]
S2 --> F["Data purged after the backup<br/>is back, and out of policy<br/>until the next sweep"] 1. Replay the erasures. Restoring rolls back later erasures and the record that they happened. The ledgers written beside the following backups are the only things that know who must stay erased. The script prints the list of ledger files with the real paths filled in:
cut -d, -f2 ../backups/<later>/erasures.csv … | sort -u
Then, for each id, re-issue the erasure and read the answer:
404: already erased, nothing to do.2xx: that person had come back, and has just been erased again.
That count is what goes into the incident record and to the data protection officer. After this step it is no longer recoverable from anywhere.
If no later ledger is reachable, stop: nothing in the world knows who had to be erased.
2. Run the retention sweep. A backup taken before a purge restores the data the purge removed, putting the system back outside its own announced window until the next sweep. A 409 here means the scheduled sweep is doing that very work already — wait for it rather than forcing anything.
Recovering a backup that went off site¶
The scenario off-site backup exists for is the one where nothing local is available, so the command cannot live only in the script that disappeared with the machine. It is copied into the manifest inside the archive:
mc alias set offsite "$BACKUP_REMOTE_ENDPOINT" "$BACKUP_REMOTE_ACCESS_KEY" "$BACKUP_REMOTE_SECRET_KEY"
mc cp "offsite/$BACKUP_REMOTE_BUCKET/<stamp>.tar.enc" .
openssl enc -d -aes-256-cbc -pbkdf2 -iter 600000 \
-in <stamp>.tar.enc -out <stamp>.tar \
-pass env:BACKUP_ENCRYPTION_PASSPHRASE
mkdir -p ../backups/<stamp> && tar -C ../backups/<stamp> -xf <stamp>.tar
make backup-verify BACKUP=../backups/<stamp>
A wrong passphrase fails loudly (bad decrypt). It never returns a half-correct file.
Commissioning a controller¶
- Create the zone, then the controller, its doors and their readers, through the dashboard or the API (
POST /devices,POST /doors,POST /identification-devices). Bulk provisioning a whole fleet at once is still better done through the seed or by SQL directly. - Give each door its relay GPIO (
relayGpio), or anosdpAddressif it has its own node on the bus. A door with neither is held by the controller and refuses every read at it withDOOR_NOT_COMMISSIONED— declared, not wired. - Issue the controller an identity whose common name is its
device_identifier. The broker turns that name into the topic subtree it may reach. - Prepare its
edge.envfrom the controller's page (Prepare its edge.env) and copy it to/etc/wardn/edge.env.DEVICE_ID, the broker and the OTA key are already filled in, and the page says what is still missing. Generate the database key on the machine withopenssl rand -hex 32, and put the certificate at the paths the file names. - Start it. On its first connect it asks for its whole rights set, and its zone, doors, readers and formats arrive alongside it.
Nothing else goes on the machine. BOOTSTRAP_RIGHTS_PATH still exists and is optional: it holds rights only, for a site that has to keep deciding through an uplink that has not come up yet. It is read once, when the local store holds no credential, and the server replaces what it loaded on the first full sync.
What the controller cannot check for you is the terminal itself. A dry contact has no return path, so relayGpio is taken at its word until somebody opens the door and watches it move. An OSDP address is the opposite: the bus reports whether the node answers, and the dashboard shows it. A node keeps its output as it was last told, so the controller closes every door on its bus when it starts. A door held open from the dashboard (Hold open) stays open until Close, or until the next restart of the controller.
Check it landed:
curl -H "Authorization: Bearer $TOKEN" http://localhost:3000/api/v1/devices/fleet
status: online, an appliedRevision matching rightsRevision, and an accessRightsHash the server agrees with.
Common situations¶
| Symptom | What to look at |
|---|---|
| A controller is grey on the fleet screen | It has missed three heartbeats. Power, uplink or a crash all look the same from here. controller_silent was raised, and its own screen shows how long it has been away |
appliedRevision lags rightsRevision | Deltas queued or in flight. sync_failed says if delivery gave up. POST /devices/{id}/actions/full-resync forces the remedy |
rights_drift keeps being raised | The controller has been sent its whole set three times and still disagrees. That is far more likely a fault in the fingerprint comparison than a controller losing the same rights every time. The nightly pass still covers it meanwhile |
pendingEvents climbing | The uplink is down, or the broker is refusing. Passages are safe until the queue reaches PENDING_EVENTS_MAX |
droppedEvents non-zero | The record for that controller has a hole. The number is cumulative since commissioning |
droppedTelemetry non-zero | The telemetry queue reached PENDING_TELEMETRY_MAX and samples were lost. The charts have a gap there, and it is a hole in the record rather than a controller that was down |
clockTrusted: false | The controller's wall clock is behind an instant it already lived through. Its decisions and its timestamps cannot be trusted until real time catches up |
| A gap in the telemetry history that never fills | The controller was actually down for that stretch, not merely cut off. A gap the network caused instead fills in once the link returns, with deliveryLagSeconds showing how late (chapter 7) |
| A door behaves oddly | Check its mode. A lane left in unlocked after roadworks opens for everybody, and still logs every read |
reboot_os is refused | The controller runs in a container. It refuses rather than degrading into a service restart |
| The dashboard cannot reach the API | CORS_ORIGINS and WARDN_API_URL must both name the address the browser dials |
| Retention never runs | MAINTENANCE_DATABASE_URL is unset. The backend says so once at startup and keeps serving |
Configuration reference¶
Everything is an environment variable with a documented default. make env creates dev/.env from dev/.env.example. This list is kept against the code that actually reads each one, not against memory of what it used to be.
Backend¶
Every number below is read once at startup and checked: it has to be a whole one, and at least one. A word, a fraction, a negative or a zero stops the boot and names the variable, rather than reaching the thing it bounds — a pool allowed to hold no connection never lends one out, a page size of -1 empties every list, and neither says so. The few where zero means something accept it and refuse anything below: the three pg timeouts, the hour of the nightly full sync, its jitter, the expiry grace and the certificate warning. An empty value counts as unset and takes the default.
| Variable | Type | Default | Effect |
|---|---|---|---|
PORT | number | 3000 | HTTP port |
DATABASE_URL | string | required | The application role's connection |
DATABASE_POOL_MAX | number | 10 | Connections per instance |
DATABASE_CONNECTION_TIMEOUT_MS | number (ms) | 5000 | How long a request waits for a free connection |
DATABASE_IDLE_TIMEOUT_MS | number (ms) | 10000 | How long an unused connection is held |
DATABASE_STATEMENT_TIMEOUT_MS | number (ms) | 15000 | Enforced by PostgreSQL, so it releases the connection even when the client stopped waiting |
MAINTENANCE_DATABASE_URL | string | (empty) | Retention and erasure only. Empty means those two are unavailable |
MQTT_URL | string | required | mqtt:// or mqtts:// |
DEVICE_MQTT_URL | string | MQTT_URL | The broker as a controller reaches it, used to fill in the edge.env the dashboard prepares. Set it when the backend's own URL only works inside its network |
MQTT_SHARED_GROUP | string | wardn-backend | Every instance joins the same group |
MQTT_USERNAME | string | wardn-backend | Must match the certificate's common name on the mTLS listener |
MQTT_PASSWORD | string (secret) | (empty) | Checked by the plain listener's authenticator. Ignored once MQTT_URL is mqtts:// |
MQTT_TLS_CA · _CERT · _KEY | string (paths) | (empty) | Read only when MQTT_URL is mqtts://. Missing one is fatal |
INSTANCE_ID | string | wardn-backend-<hostname> | Stable and unique per instance. Setting it deliberately enables a persistent session |
OIDC_ISSUER_URL | string (URL) | required | Exactly as it appears in tokens |
OIDC_JWKS_URL | string (URL) | (empty) | Discovered from the issuer when empty. Set only when this process cannot reach it the way a browser does. See 10. Security |
OIDC_AUDIENCE | string | wardn-spa | Checked against a token's aud claim, or client_id on a token that carries no aud. See 10. Security |
ENABLE_API_DOCS | boolean | false | Registers /docs and /docs-json, unauthenticated. See 10. Security |
CORS_ORIGINS | list (comma-separated) | http://localhost:4200 | Listed, never reflected |
TRUST_PROXY | true · false · hop count · list of addresses/CIDRs | false | Whose word to take for a caller's address. Must be set behind a proxy, or every per-address ceiling shares one bucket. See 10. Security |
LOG_LEVEL | enum: debug · info · warn · error | info | |
API_DEFAULT_PAGE_SIZE | number | 50 | |
API_MAX_PAGE_SIZE | number | 500 | Ceiling on limit |
DEFAULT_ZONE_TIMEZONE | string (IANA zone) | Europe/Brussels | The timezone a zone created without one gets |
RATE_LIMIT_WINDOW_MS | number (ms) | 60000 | |
RATE_LIMIT_MAX | number | 3000 | Requests per window, applied before authentication |
WS_HANDSHAKE_LIMIT | number | 60 | Live-socket handshakes per address per window |
WS_TICKET_TTL_MS | number (ms) | 15000 | How long a live-socket ticket survives unredeemed. See 09. The API |
API_KEY_INVALID_ATTEMPT_LIMIT | number | 20 | Failed API-key lookups per address per window |
OWN_CERTIFICATE_CHECK_MS | number (ms) | 86400000 (24h) | How often this instance checks its own certificate's expiry. Runs only when MQTT_TLS_CERT is set |
REPLAY_THRESHOLD_MS | number (ms) | 30000 | Lateness past which a passage is recorded as held |
DEVICE_HEARTBEAT_INTERVAL_S | number (s) | 10 | The cadence the server measures silence against |
DEVICE_OFFLINE_AFTER_MISSED | number | 3 | Missed heartbeats before a controller is called offline |
DEVICE_OFFLINE_SWEEP_MS | number (ms) | 5000 | How often the sweep looks |
DEVICE_DEBUG_MAX_DURATION_S | number (s) | 900 | Ceiling on a debug stream, enforced here and again on the controller |
ALERT_REPEAT_MINUTES | number (minutes) | 60 | Silence after a condition has been raised once. A controller reporting again ends the silence for controller_silent early |
TLS_EXPIRY_WARNING_DAYS | number (days) | 30 | Notice before a controller's certificate lapses |
SYNC_TICK_MS | number (ms) | 1000 | How often the rights queue is drained |
SYNC_TASK_BATCH | number | 20 | Deltas claimed per tick |
SYNC_TASK_LEASE_S | number (s) | 30 | How long a claim is held before another instance may take it |
SYNC_TASK_MAX_ATTEMPTS | number | 10 | Attempts before giving up and alerting |
SYNC_OPERATIONS_PER_MESSAGE | number | 200 | Rights one message carries, which is what splits a full sync into several. Sized for IoT Core's 128 KB ceiling, topology included |
RIGHTS_FULL_SYNC_LOCAL_HOUR | number (0–23) | 3 | Nightly full sync, in each zone's own timezone |
RIGHTS_FULL_SYNC_JITTER_MINUTES | number (minutes) | 60 | Window the fleet is spread over |
RIGHTS_FULL_SYNC_TICK_MS | number (ms) | 60000 | How often the nightly schedule is examined |
RIGHTS_VERIFY_INTERVAL_S | number (s) | 7200 | How often fingerprints are compared. 0 disables the check |
RIGHTS_MAX_FINGERPRINT_MISMATCHES | number | 3 | Disagreements before alerting instead of resynchronising |
EXPIRED_RIGHT_GRACE_DAYS | number (days) | 7 | How long a lapsed right stays in the set a controller is expected to hold. Must match the value the controllers use, or every controller reads as drifted |
RETENTION_SWEEP_MS | number (ms) | 21600000 | Six hours |
ACTIVITY_LOG_RETENTION_DAYS | number (days) | 30 | |
SYSTEM_LOG_RETENTION_DAYS | number (days) | 30 | |
AUDIT_LOG_RETENTION_DAYS | number (days) | 365 | |
ACCESS_RIGHT_RETENTION_DAYS | number (days) | 30 | How long a revoked right is kept |
SYNC_TASK_RETENTION_DAYS | number (days) | 7 | How long an acknowledged sync task is kept |
TELEMETRY_RETENTION_DAYS | number (days) | 5 | How long a telemetry sample is kept. Not partitioned — a plain DELETE on a small table, unlike the log tables above |
S3_ENDPOINT | string (URL) | http://minio:9000 | Signed into the links handed to controllers, so it must be an address a controller can dial |
S3_PUBLIC_ENDPOINT | string (URL) | (empty) | Where a browser can reach the same store, when that differs from S3_ENDPOINT. Empty reuses it |
S3_REGION | string | us-east-1 | |
S3_ACCESS_KEY · S3_SECRET_KEY | string (secret) | (empty) | Scoped credentials. Never the root account |
S3_OTA_BUCKET | string | wardn-ota | |
OTA_LINK_VALIDITY_S | number (s) | 900 | How long a presigned link handed to a controller, to fetch a release, lives |
OTA_MAX_ATTEMPTS | number | 3 | How many times a release is pushed to a controller before it is called a failure and raises firmware_update_failed |
OTA_RETRY_AFTER_S | number (s) | 900 | How long to leave a controller alone between two attempts. Long enough for a real download, install and reboot |
FIRMWARE_CONVERGE_MS | number (ms) | 60000 | How often the fleet is checked for controllers not running the release they were assigned |
SENTRY_DSN | string (URL, secret) | (empty) | Empty means no error reporting |
SENTRY_ENVIRONMENT | string | development | |
OTEL_EXPORTER_OTLP_ENDPOINT | string (URL) | (empty) | Empty means spans are created but not exported |
OTEL_SERVICE_NAME | string | wardn-backend | The service name traces are reported under |
OTEL_TRACES_CONSOLE | boolean | false | Print spans where the logs are |
Controller¶
| Variable | Type | Default | Effect |
|---|---|---|---|
DEVICE_ID | string | required | The name it publishes under, and the common name of its certificate |
SQLITE_PATH | string (path) | /data/edge.db | |
SQLITE_ENCRYPTION_KEY_SOURCE | enum: env · none · secure_element | secure_element | ⚠️ secure_element not implemented: the process refuses to start |
SQLITE_ENCRYPTION_KEY | string (secret) | (none) | Required when the source is env. No default, ever |
BOOTSTRAP_RIGHTS_PATH | string (path) | (none) | Fallback rights only, read once when the store holds no credential. The wiring comes from the cloud |
SIMULATION_LISTEN | string (host:port) | (none) | Development and bench only. Where credentials are also accepted as text over TCP, from anyone who reaches it. It opens on top of the real wires, and stands in for the ones this box has none of (an empty GPIO_CHIP or OSDP_SERIAL_PORT). Reported in the heartbeat, and every passage it carries is marked simulated in the activity log. Whatever this says, the controller opens the RS-485 line when a door or a reader has a bus address, and its own header when a reader is wired to it. A wire that cannot be opened is reported and stops nothing else |
GPIO_CHIP | string (path) | /dev/gpiochip0 | The character device this controller's own pins live on. Empty is a box saying it has none: doors on a terminal are then reported rather than opened |
OSDP_SERIAL_PORT | string (path) | /dev/ttyV0 | Empty is a box saying it has no RS-485 adapter. The addresses polled on it are not set here. They come from the cloud, one per door and reader the topology wires to the bus |
OSDP_BAUD_RATE | number | 9600 | |
OSDP_POLL_INTERVAL_MS | number (ms) | 100 | |
SIMULATED_SILENT_ADDRESSES | list (comma-separated) | (none) | Development only. A simulated line answers for every address the topology declares; these do not, which is how a reader that is down is reproduced without hardware |
PIN_ENTRY_DIGITS | number | 0 | How many digits complete a PIN on their own. 0 waits for the enter key |
PIN_ENTRY_IDLE_S | number (s) | 10 | How long a half-typed PIN survives between two keystrokes |
METRICS_LISTEN | string (host:port) | 0.0.0.0:9090 | Prometheus endpoint, on the local network |
LOG_LEVEL | enum: debug · info · warn · error | info | Changeable at runtime from the server |
DECISION_TIMEOUT_MS | number (ms) | 300 | The budget a decision is measured against |
PENDING_EVENTS_MAX | number | 50000 | Store & Forward capacity. A real maximum |
PENDING_EVENTS_WARN_RATIO | number (0–1) | 0.8 | Depth at which the controller reports itself saturated |
CLOCK_TOLERANCE_S | number (s) | 60 | How far wall time may run backwards before the clock is called wrong |
EXPIRED_RIGHT_GRACE_DAYS | number (days) | 7 | How long a lapsed right stays distinguishable from an unknown one. Set the same value on the backend: it is the window both sides apply to the rights set they compare by digest |
EXPIRED_PURGE_INTERVAL_S | number (s) | 3600 | How often lapsed rights are swept out past that window |
MQTT_HOST · MQTT_PORT | string · number | emqx · 1883 | |
MQTT_PASSWORD | string (secret) | (empty) | Checked by the plain listener's authenticator. Ignored once MQTT_TLS_* is set |
MQTT_TLS_CA · _CERT · _KEY | string (paths) | (empty) | All three or none. A half-configured identity fails at startup |
MQTT_KEEP_ALIVE_S | number (s) | 30 | |
EVENT_BATCH_SIZE | number | 100 | Passages published per drain |
EVENT_PUBLISH_INTERVAL_MS | number (ms) | 1000 | How often the queue is drained |
HEARTBEAT_INTERVAL_S | number (s) | 10 | Heartbeat and state report |
TELEMETRY_INTERVAL_S | number (s) | 60 | How often a system-metrics sample is taken and queued |
PENDING_TELEMETRY_MAX | number | 7200 | Telemetry queue capacity. Sized to five days at the default interval — bufferising more than the server keeps buys nothing |
AMBIENT_TEMPERATURE_PATH | string (path) | (empty) | A 1-Wire probe's sysfs w1_slave file. Empty means no probe wired, reported as null rather than a guessed temperature |
DEVICE_DEBUG_MAX_DURATION_S | number (s) | 900 | The controller's own ceiling on a debug stream |
HARDWARE_MODEL | string | read from the device tree | What the box is |
REBOOT_COMMAND | string | /sbin/reboot | Refused outright inside a container |
LED_PATH | string (path) | /sys/class/leds/ACT | The box's own LED, blinked on request. The local stack points it at an empty directory standing in for one. Absent, the command is refused |
WATCHDOG_PATH | string (path) | /dev/watchdog | Absent in a container. The controller says so |
WATCHDOG_INTERVAL_S | number (s) | 10 | |
FIRMWARE_PATH | string (path) | /opt/wardn/bin/wardn-edge | What the supervisor starts, and what an update replaces |
OTA_PUBLIC_KEY | string (hex, 32 bytes) | (required) | Ed25519 public key. A release with no valid matching signature is refused before anything is downloaded |
A controller has no Sentry setting. Its starts, its stops, its panics and the tasks that die under it are written to a journal folder beside its local store (journal/journal.jsonl, next to SQLITE_PATH) and sent over MQTT on wardn/devices/<id>/journal once it is connected. They land in the technical journal, dated when they happened, as controller_started, controller_stopped, controller_panicked and controller_task_stopped. A stop is journaled whenever it was asked for (SIGTERM from systemd or a reboot, a restart command, an update, a wiring change), so a controller that went offline with no controller_stopped crashed or lost power. An entry is deleted once the broker has acknowledged it, and at most 256 KiB wait unsent: past that, new entries are dropped rather than filling the disk.
Dashboard¶
Read at container start and written into config.js. These are the addresses a browser dials, never the service names the containers use.
| Variable | Type | Default | Effect |
|---|---|---|---|
WARDN_API_URL | string (URL) | http://localhost:3000 | |
WARDN_OIDC_AUTHORITY | string (URL) | http://localhost:8080/realms/wardn | |
WARDN_OIDC_CLIENT_ID | string | wardn-spa | |
WARDN_OIDC_ENDPOINT_ORIGIN | string (origin) | (empty) | The origin the provider serves its token endpoint from, when it is not the issuer's. Cognito's hosted domain. Added to connect-src |
WARDN_SENTRY_DSN | string (URL, secret) | (empty) | |
WARDN_ENVIRONMENT | string | development | |
WARDN_HELP_URL | string (URL or path) | /help | Where the sidebar's help links point to |
Backups¶
| Variable | Type | Default | Effect |
|---|---|---|---|
BACKUP_DIR | string (path) | ./backups | |
BACKUP_KEEP_DAYS | number (days) | 30 | Applied locally and in the remote store |
BACKUP_REMOTE_ENDPOINT | string (URL) | (empty) | Empty keeps backups on this machine |
BACKUP_REMOTE_BUCKET · _ACCESS_KEY · _SECRET_KEY | string · string (secret) · string (secret) | (empty) | |
BACKUP_ENCRYPTION_PASSPHRASE | string (secret) | (empty) | Mandatory once an endpoint is set. make backup refuses without it |
Local infrastructure¶
Used by Compose and by db-init, not by the applications.
| Variable | Type | Default |
|---|---|---|
POSTGRES_DB · POSTGRES_USER · POSTGRES_PASSWORD | string · string · string (secret) | wardn · wardn_owner · wardn_owner_dev |
APP_DB_USER · APP_DB_PASSWORD | string · string (secret) | wardn_app · wardn_app_dev |
MAINTENANCE_DB_USER · MAINTENANCE_DB_PASSWORD | string · string (secret) | wardn_maintenance · wardn_maintenance_dev |
PARTITION_MONTHS_BACK · PARTITION_MONTHS_AHEAD | number · number | 3 · 3 |
SEED_ENABLED · SEED_FORCE | boolean · boolean | true · false |
SEED_DIR | colon-separated paths | /db/seed |
SEED_ACTIVITY_DAYS · SEED_ACTIVITY_EVENTS_PER_DAY | number · number | 30 · 12 |
MQTT_PORT · MQTT_TLS_PORT | number · number | 1883 · 8883 |
EMQX_DASHBOARD_PORT · EMQX_DASHBOARD_PASSWORD | number · string (secret) | 18083 · wardn_dev_dashboard |
IDP_PORT · IDP_ADMIN · IDP_ADMIN_PASSWORD · IDP_REALM | number · string · string (secret) · string | 8080 · admin · admin · wardn |
MINIO_PORT · MINIO_CONSOLE_PORT · MINIO_ROOT_USER · MINIO_ROOT_PASSWORD | number · number · string · string (secret) | 9000 · 9001 · wardn · wardn_dev_secret |
MINIO_DOCS_BUCKET · MINIO_OTA_BUCKET | string · string | wardn-docs · wardn-ota |
MINIO_APP_ACCESS_KEY · MINIO_APP_SECRET_KEY | string · string (secret) | wardn-backend · wardn_backend_dev_secret |
TLS_DAYS · TLS_CA_DAYS | number (days) · number (days) | 825 · 3650 |
BACKEND_PORT · BACKEND_2_PORT · FRONTEND_PORT | number · number · number | 3000 · 3001 · 4200 |
EDGE_READER_PORT · EDGE_METRICS_PORT | number · number | 9100 · 9101 |
wardn is operable. Last chapter: contributing to it.
Next → 13. Developing