Skip to content

← 11. Personal data · Contents · Next → 13. Developing

12. Operating

Everything needed to keep wardn running: deploying it, watching it, backing it up, restoring it, and the complete list of knobs.

Deploying

The local stack is Docker Compose. Production replaces four things with managed equivalents and changes nothing in the code:

Component Local Production
Broker EMQX container AWS IoT Core, or any MQTT broker with shared subscriptions
Object storage MinIO S3
Database PostgreSQL container Managed PostgreSQL
Identity Keycloak container Keycloak, or any standards-compliant OIDC provider

A worked example of the right-hand column, on AWS with a Helm chart and Terraform, is 14. Deploying to production. A cheaper one, on free and near-free tiers, is 15. A near-free deployment.

The database socle

db-init is the single source of truth for the schema, the roles, the grants, the partitions and the seed. It is idempotent and runs on every up.

make up          # migrations, partitions, seed
make verify      # assert the schema, the grants and the seed
make seed        # re-run migrations and seed
make reseed      # re-run the seed even if already applied

Set SEED_ENABLED=false for an empty database. PARTITION_MONTHS_BACK and PARTITION_MONTHS_AHEAD control how many monthly partitions are created up front. The retention sweep keeps two months ahead from then on.

Running more than one backend

The backend is stateless and leaderless.

make scale   # a second instance on port 3001, and the proof that it works

Two rules:

  • Each instance needs its own INSTANCE_ID, stable across restarts. Two MQTT clients sharing an id kick each other off the broker. In Kubernetes this is the StatefulSet ordinal.
  • Stable matters as much as unique. With INSTANCE_ID set, the broker keeps a persistent session and queues events for an instance that is restarting. Without it, the backend takes a clean session, says so in its log, and accepts that a restart misses whatever arrives while it is down.

Everything else coordinates by itself:

  • MQTT shared subscriptions spread events.
  • The rights queue is drained concurrently with leases.
  • The scheduled sweeps take a lease in sweep_leases.

Certificates

tls-init issues a development CA and one identity per party on first start. In production the authority is the manufacturing one, and a controller's key never leaves its secure element. See 10. Security.

make tls-check   # how long every certificate in the stack has left
Exit code Meaning
0 Nothing lapses within TLS_EXPIRY_WARNING_DAYS
1 Something lapses inside the window, or already has
2 The check could not be made (no openssl)

The 2 exists precisely so that an inability to check does not masquerade as a certificate alarm.

Watching day to day

Signal Where to read it
A controller is offline Fleet screen, or the controller_silent alert
A queue is rising pendingEvents in the heartbeat, queue_saturated alert
Passages were lost droppedEvents non-zero, events_dropped alert
History samples were lost droppedTelemetry non-zero, telemetry_dropped alert
A clock is wrong clockTrusted: false, clock_untrusted alert
A reader stopped answering readersDown, readerBusStopped, reader_error alert
Rights are not landing sync_failed, rights_drift alerts. appliedRevision against rightsRevision
A certificate is dying certificate_expiring alert, and make tls-check
Application errors Backend and dashboard: Sentry, when SENTRY_DSN is set. A controller: controller_panicked and controller_task_stopped in the technical journal
A specific request The X-Trace-Id on its response
A controller, from the site's own monitoring http://<controller>:9090/metrics
The backend is up GET /health, and which instance answered
The database answers GET /health/ready, or 503 when it does not

Alerts are read through the API:

curl -H "Authorization: Bearer $TOKEN" 'http://localhost:3000/api/v1/alerts?limit=50'

The eight kinds and what raises them are in 7. The fleet.

Backups

make backup                                     # database, artefacts, erasure ledger
make backup-verify BACKUP=../backups/<stamp>    # restore it elsewhere to prove it
make restore BACKUP=../backups/<stamp> CONFIRM=yes
flowchart TD
    B["make backup"] --> D["pg_dump, custom format"]
    B --> L["Erasure ledger<br/>written BESIDE the dump"]
    B --> O["Mirror the object storage<br/>(firmware, documents)"]
    D --> M["MANIFEST: what this can<br/>and cannot restore"]
    L --> M
    O --> M
    M --> ENC{"Off-site<br/>configured?"}
    ENC -->|"yes, with a passphrase"| E["Encrypt, upload,<br/>prune the remote"]
    ENC -->|"yes, no passphrase"| R["REFUSE: nothing<br/>leaves in the clear"]
    ENC -->|"no"| K["Stays on this machine"]

What is backed up

Why
PostgreSQL Rights, passages, the audit trail. The only copy
Object storage Firmware artefacts and documents. Their digests are in the database, so a database restored without them points at files that do not exist
The erasure ledger The list of people erased, written beside the dump. It is what makes this more than a pg_dump

What is not backed up

  • The controllers' local databases: derived from the server, no history. They rebuild themselves on the next connect.
  • EMQX state: sessions and retained messages rebuild themselves.
  • The development certificate authority: production uses the manufacturing PKI.
  • ⚠️ A self-hosted identity provider's accounts. The default, Keycloak, has its realm imported from this repository, so what a backup would add is whatever was created at runtime afterwards. That is a real gap, stated rather than hidden, and one a managed provider (Auth0, Okta, Azure AD) does not have, since it backs up its own accounts. See 10. Security.

The recovery point

One pg_dump a day. Up to 24 hours of passages can disappear.

That is not theoretical. The controller is built never to lose a passage locally. But once an event is acknowledged and cleared from its queue, last night's dump holds the only copy.

It is the accepted cost of not operating a write-ahead log archive. If that cost becomes unacceptable, point-in-time recovery is the way out, and it is an operations project rather than a script change.

A backup that has never been restored is not a backup

make backup-verify BACKUP=../backups/<stamp>

This is the step everyone skips, and the reason backups are found to be worthless on the one day they are needed.

It restores into a throwaway database beside the real one, never over it, then asserts three things a corrupt or truncated dump cannot satisfy:

  • The schema comes back whole, checked by the same verify.sql the socle uses.
  • The tables carrying the business are not empty.
  • The audit trail came back with them, since a restore that loses the trail restores the data without the accountability.

Off site

A backup on the same machine as the database is not a backup: the fire, the theft and the failed disk take both.

BACKUP_REMOTE_ENDPOINT=https://s3.eu-west-3.amazonaws.com
BACKUP_REMOTE_BUCKET=wardn-backups
BACKUP_REMOTE_ACCESS_KEY=…
BACKUP_REMOTE_SECRET_KEY=…
BACKUP_ENCRYPTION_PASSPHRASE=…

Encrypted before it leaves, never after. With a destination configured and no passphrase, make backup refuses, rather than sending names, credentials and the history of every passage through a site to somebody else's disk in the clear because a variable was forgotten.

⚠️ A lost passphrase makes the backup as useless as its absence. It belongs neither in this repository nor in the remote store.

The SHA-256 of the content before encryption stays on site, where the remote store never saw it. Comparing a copy that comes back against that value proves what returns is what left. The cipher alone does not prove that.

The same run prunes the remote store beyond BACKUP_KEEP_DAYS. Retention off site is not the local find: without this, a 30-day commitment would hold everywhere except the one place the data had been copied to.

Backups compose with retention rather than coinciding with it. The newest backup already holds logs 30 days old, so by the time it is deleted at 30 days it has been holding data up to 60 days old. That belongs in a register of processing activities.

Restoring

make restore BACKUP=../backups/<stamp> CONFIRM=yes

Destructive, and deliberately hard to run by accident: without CONFIRM=yes it refuses and prints the command it wanted.

It runs in order:

  1. Stops the backend.
  2. Restores the dump.
  3. Re-runs db-init (which is the source of truth for the schema, the roles and the append-only grants, and is idempotent).
  4. Starts the backend again.

The two steps that are not automatic

The script prints them instead of doing them, because neither is safe to automate.

flowchart TD
    R["Restore finished"] --> S1["1. Replay the erasures<br/>recorded after this backup"]
    R --> S2["2. Run the retention sweep"]
    S1 --> A["Collect the ids from every<br/>LATER backup's ledger"]
    A --> B["Re-issue each erasure"]
    B --> C{"Response?"}
    C -->|"404"| D["Already erased,<br/>nothing to do"]
    C -->|"2xx"| E["That person HAD come back.<br/>Count them: this number<br/>cannot be recovered later"]
    S2 --> F["Data purged after the backup<br/>is back, and out of policy<br/>until the next sweep"]

1. Replay the erasures. Restoring rolls back later erasures and the record that they happened. The ledgers written beside the following backups are the only things that know who must stay erased. The script prints the list of ledger files with the real paths filled in:

cut -d, -f2 ../backups/<later>/erasures.csv … | sort -u

Then, for each id, re-issue the erasure and read the answer:

  • 404: already erased, nothing to do.
  • 2xx: that person had come back, and has just been erased again.

That count is what goes into the incident record and to the data protection officer. After this step it is no longer recoverable from anywhere.

If no later ledger is reachable, stop: nothing in the world knows who had to be erased.

2. Run the retention sweep. A backup taken before a purge restores the data the purge removed, putting the system back outside its own announced window until the next sweep. A 409 here means the scheduled sweep is doing that very work already — wait for it rather than forcing anything.

Recovering a backup that went off site

The scenario off-site backup exists for is the one where nothing local is available, so the command cannot live only in the script that disappeared with the machine. It is copied into the manifest inside the archive:

mc alias set offsite "$BACKUP_REMOTE_ENDPOINT" "$BACKUP_REMOTE_ACCESS_KEY" "$BACKUP_REMOTE_SECRET_KEY"
mc cp "offsite/$BACKUP_REMOTE_BUCKET/<stamp>.tar.enc" .

openssl enc -d -aes-256-cbc -pbkdf2 -iter 600000 \
  -in <stamp>.tar.enc -out <stamp>.tar \
  -pass env:BACKUP_ENCRYPTION_PASSPHRASE

mkdir -p ../backups/<stamp> && tar -C ../backups/<stamp> -xf <stamp>.tar
make backup-verify BACKUP=../backups/<stamp>

A wrong passphrase fails loudly (bad decrypt). It never returns a half-correct file.

Commissioning a controller

  1. Create the zone, then the controller, its doors and their readers, through the dashboard or the API (POST /devices, POST /doors, POST /identification-devices). Bulk provisioning a whole fleet at once is still better done through the seed or by SQL directly.
  2. Give each door its relay GPIO (relayGpio), or an osdpAddress if it has its own node on the bus. A door with neither is held by the controller and refuses every read at it with DOOR_NOT_COMMISSIONED — declared, not wired.
  3. Issue the controller an identity whose common name is its device_identifier. The broker turns that name into the topic subtree it may reach.
  4. Prepare its edge.env from the controller's page (Prepare its edge.env) and copy it to /etc/wardn/edge.env. DEVICE_ID, the broker and the OTA key are already filled in, and the page says what is still missing. Generate the database key on the machine with openssl rand -hex 32, and put the certificate at the paths the file names.
  5. Start it. On its first connect it asks for its whole rights set, and its zone, doors, readers and formats arrive alongside it.

Nothing else goes on the machine. BOOTSTRAP_RIGHTS_PATH still exists and is optional: it holds rights only, for a site that has to keep deciding through an uplink that has not come up yet. It is read once, when the local store holds no credential, and the server replaces what it loaded on the first full sync.

What the controller cannot check for you is the terminal itself. A dry contact has no return path, so relayGpio is taken at its word until somebody opens the door and watches it move. An OSDP address is the opposite: the bus reports whether the node answers, and the dashboard shows it. A node keeps its output as it was last told, so the controller closes every door on its bus when it starts. A door held open from the dashboard (Hold open) stays open until Close, or until the next restart of the controller.

Check it landed:

curl -H "Authorization: Bearer $TOKEN" http://localhost:3000/api/v1/devices/fleet

status: online, an appliedRevision matching rightsRevision, and an accessRightsHash the server agrees with.

Common situations

Symptom What to look at
A controller is grey on the fleet screen It has missed three heartbeats. Power, uplink or a crash all look the same from here. controller_silent was raised, and its own screen shows how long it has been away
appliedRevision lags rightsRevision Deltas queued or in flight. sync_failed says if delivery gave up. POST /devices/{id}/actions/full-resync forces the remedy
rights_drift keeps being raised The controller has been sent its whole set three times and still disagrees. That is far more likely a fault in the fingerprint comparison than a controller losing the same rights every time. The nightly pass still covers it meanwhile
pendingEvents climbing The uplink is down, or the broker is refusing. Passages are safe until the queue reaches PENDING_EVENTS_MAX
droppedEvents non-zero The record for that controller has a hole. The number is cumulative since commissioning
droppedTelemetry non-zero The telemetry queue reached PENDING_TELEMETRY_MAX and samples were lost. The charts have a gap there, and it is a hole in the record rather than a controller that was down
clockTrusted: false The controller's wall clock is behind an instant it already lived through. Its decisions and its timestamps cannot be trusted until real time catches up
A gap in the telemetry history that never fills The controller was actually down for that stretch, not merely cut off. A gap the network caused instead fills in once the link returns, with deliveryLagSeconds showing how late (chapter 7)
A door behaves oddly Check its mode. A lane left in unlocked after roadworks opens for everybody, and still logs every read
reboot_os is refused The controller runs in a container. It refuses rather than degrading into a service restart
The dashboard cannot reach the API CORS_ORIGINS and WARDN_API_URL must both name the address the browser dials
Retention never runs MAINTENANCE_DATABASE_URL is unset. The backend says so once at startup and keeps serving

Configuration reference

Everything is an environment variable with a documented default. make env creates dev/.env from dev/.env.example. This list is kept against the code that actually reads each one, not against memory of what it used to be.

Backend

Every number below is read once at startup and checked: it has to be a whole one, and at least one. A word, a fraction, a negative or a zero stops the boot and names the variable, rather than reaching the thing it bounds — a pool allowed to hold no connection never lends one out, a page size of -1 empties every list, and neither says so. The few where zero means something accept it and refuse anything below: the three pg timeouts, the hour of the nightly full sync, its jitter, the expiry grace and the certificate warning. An empty value counts as unset and takes the default.

Variable Type Default Effect
PORT number 3000 HTTP port
DATABASE_URL string required The application role's connection
DATABASE_POOL_MAX number 10 Connections per instance
DATABASE_CONNECTION_TIMEOUT_MS number (ms) 5000 How long a request waits for a free connection
DATABASE_IDLE_TIMEOUT_MS number (ms) 10000 How long an unused connection is held
DATABASE_STATEMENT_TIMEOUT_MS number (ms) 15000 Enforced by PostgreSQL, so it releases the connection even when the client stopped waiting
MAINTENANCE_DATABASE_URL string (empty) Retention and erasure only. Empty means those two are unavailable
MQTT_URL string required mqtt:// or mqtts://
DEVICE_MQTT_URL string MQTT_URL The broker as a controller reaches it, used to fill in the edge.env the dashboard prepares. Set it when the backend's own URL only works inside its network
MQTT_SHARED_GROUP string wardn-backend Every instance joins the same group
MQTT_USERNAME string wardn-backend Must match the certificate's common name on the mTLS listener
MQTT_PASSWORD string (secret) (empty) Checked by the plain listener's authenticator. Ignored once MQTT_URL is mqtts://
MQTT_TLS_CA · _CERT · _KEY string (paths) (empty) Read only when MQTT_URL is mqtts://. Missing one is fatal
INSTANCE_ID string wardn-backend-<hostname> Stable and unique per instance. Setting it deliberately enables a persistent session
OIDC_ISSUER_URL string (URL) required Exactly as it appears in tokens
OIDC_JWKS_URL string (URL) (empty) Discovered from the issuer when empty. Set only when this process cannot reach it the way a browser does. See 10. Security
OIDC_AUDIENCE string wardn-spa Checked against a token's aud claim, or client_id on a token that carries no aud. See 10. Security
ENABLE_API_DOCS boolean false Registers /docs and /docs-json, unauthenticated. See 10. Security
CORS_ORIGINS list (comma-separated) http://localhost:4200 Listed, never reflected
TRUST_PROXY true · false · hop count · list of addresses/CIDRs false Whose word to take for a caller's address. Must be set behind a proxy, or every per-address ceiling shares one bucket. See 10. Security
LOG_LEVEL enum: debug · info · warn · error info
API_DEFAULT_PAGE_SIZE number 50
API_MAX_PAGE_SIZE number 500 Ceiling on limit
DEFAULT_ZONE_TIMEZONE string (IANA zone) Europe/Brussels The timezone a zone created without one gets
RATE_LIMIT_WINDOW_MS number (ms) 60000
RATE_LIMIT_MAX number 3000 Requests per window, applied before authentication
WS_HANDSHAKE_LIMIT number 60 Live-socket handshakes per address per window
WS_TICKET_TTL_MS number (ms) 15000 How long a live-socket ticket survives unredeemed. See 09. The API
API_KEY_INVALID_ATTEMPT_LIMIT number 20 Failed API-key lookups per address per window
OWN_CERTIFICATE_CHECK_MS number (ms) 86400000 (24h) How often this instance checks its own certificate's expiry. Runs only when MQTT_TLS_CERT is set
REPLAY_THRESHOLD_MS number (ms) 30000 Lateness past which a passage is recorded as held
DEVICE_HEARTBEAT_INTERVAL_S number (s) 10 The cadence the server measures silence against
DEVICE_OFFLINE_AFTER_MISSED number 3 Missed heartbeats before a controller is called offline
DEVICE_OFFLINE_SWEEP_MS number (ms) 5000 How often the sweep looks
DEVICE_DEBUG_MAX_DURATION_S number (s) 900 Ceiling on a debug stream, enforced here and again on the controller
ALERT_REPEAT_MINUTES number (minutes) 60 Silence after a condition has been raised once. A controller reporting again ends the silence for controller_silent early
TLS_EXPIRY_WARNING_DAYS number (days) 30 Notice before a controller's certificate lapses
SYNC_TICK_MS number (ms) 1000 How often the rights queue is drained
SYNC_TASK_BATCH number 20 Deltas claimed per tick
SYNC_TASK_LEASE_S number (s) 30 How long a claim is held before another instance may take it
SYNC_TASK_MAX_ATTEMPTS number 10 Attempts before giving up and alerting
SYNC_OPERATIONS_PER_MESSAGE number 200 Rights one message carries, which is what splits a full sync into several. Sized for IoT Core's 128 KB ceiling, topology included
RIGHTS_FULL_SYNC_LOCAL_HOUR number (0–23) 3 Nightly full sync, in each zone's own timezone
RIGHTS_FULL_SYNC_JITTER_MINUTES number (minutes) 60 Window the fleet is spread over
RIGHTS_FULL_SYNC_TICK_MS number (ms) 60000 How often the nightly schedule is examined
RIGHTS_VERIFY_INTERVAL_S number (s) 7200 How often fingerprints are compared. 0 disables the check
RIGHTS_MAX_FINGERPRINT_MISMATCHES number 3 Disagreements before alerting instead of resynchronising
EXPIRED_RIGHT_GRACE_DAYS number (days) 7 How long a lapsed right stays in the set a controller is expected to hold. Must match the value the controllers use, or every controller reads as drifted
RETENTION_SWEEP_MS number (ms) 21600000 Six hours
ACTIVITY_LOG_RETENTION_DAYS number (days) 30
SYSTEM_LOG_RETENTION_DAYS number (days) 30
AUDIT_LOG_RETENTION_DAYS number (days) 365
ACCESS_RIGHT_RETENTION_DAYS number (days) 30 How long a revoked right is kept
SYNC_TASK_RETENTION_DAYS number (days) 7 How long an acknowledged sync task is kept
TELEMETRY_RETENTION_DAYS number (days) 5 How long a telemetry sample is kept. Not partitioned — a plain DELETE on a small table, unlike the log tables above
S3_ENDPOINT string (URL) http://minio:9000 Signed into the links handed to controllers, so it must be an address a controller can dial
S3_PUBLIC_ENDPOINT string (URL) (empty) Where a browser can reach the same store, when that differs from S3_ENDPOINT. Empty reuses it
S3_REGION string us-east-1
S3_ACCESS_KEY · S3_SECRET_KEY string (secret) (empty) Scoped credentials. Never the root account
S3_OTA_BUCKET string wardn-ota
OTA_LINK_VALIDITY_S number (s) 900 How long a presigned link handed to a controller, to fetch a release, lives
OTA_MAX_ATTEMPTS number 3 How many times a release is pushed to a controller before it is called a failure and raises firmware_update_failed
OTA_RETRY_AFTER_S number (s) 900 How long to leave a controller alone between two attempts. Long enough for a real download, install and reboot
FIRMWARE_CONVERGE_MS number (ms) 60000 How often the fleet is checked for controllers not running the release they were assigned
SENTRY_DSN string (URL, secret) (empty) Empty means no error reporting
SENTRY_ENVIRONMENT string development
OTEL_EXPORTER_OTLP_ENDPOINT string (URL) (empty) Empty means spans are created but not exported
OTEL_SERVICE_NAME string wardn-backend The service name traces are reported under
OTEL_TRACES_CONSOLE boolean false Print spans where the logs are

Controller

Variable Type Default Effect
DEVICE_ID string required The name it publishes under, and the common name of its certificate
SQLITE_PATH string (path) /data/edge.db
SQLITE_ENCRYPTION_KEY_SOURCE enum: env · none · secure_element secure_element ⚠️ secure_element not implemented: the process refuses to start
SQLITE_ENCRYPTION_KEY string (secret) (none) Required when the source is env. No default, ever
BOOTSTRAP_RIGHTS_PATH string (path) (none) Fallback rights only, read once when the store holds no credential. The wiring comes from the cloud
SIMULATION_LISTEN string (host:port) (none) Development and bench only. Where credentials are also accepted as text over TCP, from anyone who reaches it. It opens on top of the real wires, and stands in for the ones this box has none of (an empty GPIO_CHIP or OSDP_SERIAL_PORT). Reported in the heartbeat, and every passage it carries is marked simulated in the activity log. Whatever this says, the controller opens the RS-485 line when a door or a reader has a bus address, and its own header when a reader is wired to it. A wire that cannot be opened is reported and stops nothing else
GPIO_CHIP string (path) /dev/gpiochip0 The character device this controller's own pins live on. Empty is a box saying it has none: doors on a terminal are then reported rather than opened
OSDP_SERIAL_PORT string (path) /dev/ttyV0 Empty is a box saying it has no RS-485 adapter. The addresses polled on it are not set here. They come from the cloud, one per door and reader the topology wires to the bus
OSDP_BAUD_RATE number 9600
OSDP_POLL_INTERVAL_MS number (ms) 100
SIMULATED_SILENT_ADDRESSES list (comma-separated) (none) Development only. A simulated line answers for every address the topology declares; these do not, which is how a reader that is down is reproduced without hardware
PIN_ENTRY_DIGITS number 0 How many digits complete a PIN on their own. 0 waits for the enter key
PIN_ENTRY_IDLE_S number (s) 10 How long a half-typed PIN survives between two keystrokes
METRICS_LISTEN string (host:port) 0.0.0.0:9090 Prometheus endpoint, on the local network
LOG_LEVEL enum: debug · info · warn · error info Changeable at runtime from the server
DECISION_TIMEOUT_MS number (ms) 300 The budget a decision is measured against
PENDING_EVENTS_MAX number 50000 Store & Forward capacity. A real maximum
PENDING_EVENTS_WARN_RATIO number (0–1) 0.8 Depth at which the controller reports itself saturated
CLOCK_TOLERANCE_S number (s) 60 How far wall time may run backwards before the clock is called wrong
EXPIRED_RIGHT_GRACE_DAYS number (days) 7 How long a lapsed right stays distinguishable from an unknown one. Set the same value on the backend: it is the window both sides apply to the rights set they compare by digest
EXPIRED_PURGE_INTERVAL_S number (s) 3600 How often lapsed rights are swept out past that window
MQTT_HOST · MQTT_PORT string · number emqx · 1883
MQTT_PASSWORD string (secret) (empty) Checked by the plain listener's authenticator. Ignored once MQTT_TLS_* is set
MQTT_TLS_CA · _CERT · _KEY string (paths) (empty) All three or none. A half-configured identity fails at startup
MQTT_KEEP_ALIVE_S number (s) 30
EVENT_BATCH_SIZE number 100 Passages published per drain
EVENT_PUBLISH_INTERVAL_MS number (ms) 1000 How often the queue is drained
HEARTBEAT_INTERVAL_S number (s) 10 Heartbeat and state report
TELEMETRY_INTERVAL_S number (s) 60 How often a system-metrics sample is taken and queued
PENDING_TELEMETRY_MAX number 7200 Telemetry queue capacity. Sized to five days at the default interval — bufferising more than the server keeps buys nothing
AMBIENT_TEMPERATURE_PATH string (path) (empty) A 1-Wire probe's sysfs w1_slave file. Empty means no probe wired, reported as null rather than a guessed temperature
DEVICE_DEBUG_MAX_DURATION_S number (s) 900 The controller's own ceiling on a debug stream
HARDWARE_MODEL string read from the device tree What the box is
REBOOT_COMMAND string /sbin/reboot Refused outright inside a container
LED_PATH string (path) /sys/class/leds/ACT The box's own LED, blinked on request. The local stack points it at an empty directory standing in for one. Absent, the command is refused
WATCHDOG_PATH string (path) /dev/watchdog Absent in a container. The controller says so
WATCHDOG_INTERVAL_S number (s) 10
FIRMWARE_PATH string (path) /opt/wardn/bin/wardn-edge What the supervisor starts, and what an update replaces
OTA_PUBLIC_KEY string (hex, 32 bytes) (required) Ed25519 public key. A release with no valid matching signature is refused before anything is downloaded

A controller has no Sentry setting. Its starts, its stops, its panics and the tasks that die under it are written to a journal folder beside its local store (journal/journal.jsonl, next to SQLITE_PATH) and sent over MQTT on wardn/devices/<id>/journal once it is connected. They land in the technical journal, dated when they happened, as controller_started, controller_stopped, controller_panicked and controller_task_stopped. A stop is journaled whenever it was asked for (SIGTERM from systemd or a reboot, a restart command, an update, a wiring change), so a controller that went offline with no controller_stopped crashed or lost power. An entry is deleted once the broker has acknowledged it, and at most 256 KiB wait unsent: past that, new entries are dropped rather than filling the disk.

Dashboard

Read at container start and written into config.js. These are the addresses a browser dials, never the service names the containers use.

Variable Type Default Effect
WARDN_API_URL string (URL) http://localhost:3000
WARDN_OIDC_AUTHORITY string (URL) http://localhost:8080/realms/wardn
WARDN_OIDC_CLIENT_ID string wardn-spa
WARDN_OIDC_ENDPOINT_ORIGIN string (origin) (empty) The origin the provider serves its token endpoint from, when it is not the issuer's. Cognito's hosted domain. Added to connect-src
WARDN_SENTRY_DSN string (URL, secret) (empty)
WARDN_ENVIRONMENT string development
WARDN_HELP_URL string (URL or path) /help Where the sidebar's help links point to

Backups

Variable Type Default Effect
BACKUP_DIR string (path) ./backups
BACKUP_KEEP_DAYS number (days) 30 Applied locally and in the remote store
BACKUP_REMOTE_ENDPOINT string (URL) (empty) Empty keeps backups on this machine
BACKUP_REMOTE_BUCKET · _ACCESS_KEY · _SECRET_KEY string · string (secret) · string (secret) (empty)
BACKUP_ENCRYPTION_PASSPHRASE string (secret) (empty) Mandatory once an endpoint is set. make backup refuses without it

Local infrastructure

Used by Compose and by db-init, not by the applications.

Variable Type Default
POSTGRES_DB · POSTGRES_USER · POSTGRES_PASSWORD string · string · string (secret) wardn · wardn_owner · wardn_owner_dev
APP_DB_USER · APP_DB_PASSWORD string · string (secret) wardn_app · wardn_app_dev
MAINTENANCE_DB_USER · MAINTENANCE_DB_PASSWORD string · string (secret) wardn_maintenance · wardn_maintenance_dev
PARTITION_MONTHS_BACK · PARTITION_MONTHS_AHEAD number · number 3 · 3
SEED_ENABLED · SEED_FORCE boolean · boolean true · false
SEED_DIR colon-separated paths /db/seed
SEED_ACTIVITY_DAYS · SEED_ACTIVITY_EVENTS_PER_DAY number · number 30 · 12
MQTT_PORT · MQTT_TLS_PORT number · number 1883 · 8883
EMQX_DASHBOARD_PORT · EMQX_DASHBOARD_PASSWORD number · string (secret) 18083 · wardn_dev_dashboard
IDP_PORT · IDP_ADMIN · IDP_ADMIN_PASSWORD · IDP_REALM number · string · string (secret) · string 8080 · admin · admin · wardn
MINIO_PORT · MINIO_CONSOLE_PORT · MINIO_ROOT_USER · MINIO_ROOT_PASSWORD number · number · string · string (secret) 9000 · 9001 · wardn · wardn_dev_secret
MINIO_DOCS_BUCKET · MINIO_OTA_BUCKET string · string wardn-docs · wardn-ota
MINIO_APP_ACCESS_KEY · MINIO_APP_SECRET_KEY string · string (secret) wardn-backend · wardn_backend_dev_secret
TLS_DAYS · TLS_CA_DAYS number (days) · number (days) 825 · 3650
BACKEND_PORT · BACKEND_2_PORT · FRONTEND_PORT number · number · number 3000 · 3001 · 4200
EDGE_READER_PORT · EDGE_METRICS_PORT number · number 9100 · 9101

wardn is operable. Last chapter: contributing to it.

Next → 13. Developing