Skip to content

← 14. Deploying to production · Contents · Next → 16. The lab suitcase

15. A near-free deployment

14. Deploying to production is the high-availability shape: multi-AZ RDS, an EKS cluster, Greengrass on every controller. That is the right target once a site's uptime has a cost attached to it.

It is the wrong target for a pilot, a single-workspace install, or an evaluation: the actual requirement there is the smallest bill that still runs the whole system, not a text a manager wants to hear before the first customer exists.

The swap is the same one 12. Operating already describes: four managed services and a place to run the backend and the dashboard, with every one of the six chosen for its free or near-free tier instead of its uptime guarantee:

Component Local This chapter
Broker EMQX container AWS IoT Core
Object storage MinIO Cloudflare R2
Database PostgreSQL container Supabase
Identity Keycloak container Auth0
Backend docker compose, one process A container host billed by usage, kept warm (worked example: Cloud Run, min-instances=1)
Dashboard nginx container A static host behind a CDN (worked example: Cloudflare Pages)
flowchart LR
    subgraph Local["docker compose"]
        L1["EMQX"]
        L2["MinIO"]
        L3["PostgreSQL container"]
        L4["Keycloak container"]
    end
    subgraph Cheap["This chapter"]
        C1["AWS IoT Core"]
        C2["Cloudflare R2"]
        C3["Supabase"]
        C4["Auth0"]
        C5["Cloud Run<br/>min-instances=1"]
        C6["Cloudflare Pages"]
    end
    L1 -.->|"same MQTT protocol,<br/>same QoS 1"| C1
    L2 -.->|"same S3 API,<br/>zero egress fee"| C2
    L3 -.->|"same schema,<br/>same db-init"| C3
    L4 -.->|"same OIDC contract"| C4

Every figure below is a free-tier limit as these providers publish it today. Free tiers change without notice. Treat this chapter the way 14. Deploying to production asks to be treated: a worked shape, not a bill to sign.

What "serverless" cannot mean here

The backend is not a request handler that happens to touch a database. It holds an open MQTT session on the broker's shared subscription, and it runs its own clock: the offline sweep, the nightly rights resync, the retention sweep, all in-process timers (12. Operating).

None of that survives a platform that scales an instance to zero between HTTP requests: a cold controller reconnecting to an empty subscriber, a sweep that never fires because nothing woke the process up to run it.

That rules out classic FaaS: AWS Lambda, Vercel or Netlify functions, and any PaaS's free web-service tier that sleeps after a few idle minutes.

It does not rule out usage-billed hosting in general: a container platform that keeps one instance resident and bills for what that instance actually uses is still far cheaper than provisioning fixed capacity, and loses nothing.

12. Operating already prices the failure mode for when an instance does get replaced: without a stable INSTANCE_ID, "the backend takes a clean session, says so in its log, and accepts that a restart misses whatever arrives while it is down." That is an accepted, occasional cost on a machine you control.

On a platform that recycles instances on its own schedule, it stops being occasional, one more reason min-instances has to be at least 1, not 0.

The backend: a container that never scales to zero

Cloud Run is the worked example: a genuine perpetual free tier (2,000,000 requests and roughly 180,000 vCPU-seconds a month), and it happens to already agree with wardn's own contract. The backend reads its listen port from PORT (12. Operating), which is exactly the variable Cloud Run injects. Nothing about the image built in 14 changes.

gcloud artifacts repositories create wardn --repository-format=docker --location=europe-west1

gcloud auth configure-docker europe-west1-docker.pkg.dev

docker build -t europe-west1-docker.pkg.dev/$PROJECT/wardn/wardn-backend:$GIT_SHA ./backend
docker push europe-west1-docker.pkg.dev/$PROJECT/wardn/wardn-backend:$GIT_SHA

gcloud run deploy wardn-backend \
  --image europe-west1-docker.pkg.dev/$PROJECT/wardn/wardn-backend:$GIT_SHA \
  --region europe-west1 \
  --min-instances 1 --max-instances 1 \
  --allow-unauthenticated \
  --set-env-vars "MQTT_URL=mqtts://xxxxxxxxxxxxxx-ats.iot.eu-west-1.amazonaws.com:8883,OIDC_ISSUER_URL=https://wardn.eu.auth0.com/,CORS_ORIGINS=https://dashboard.wardn.example.com,S3_ENDPOINT=https://<account-id>.r2.cloudflarestorage.com,S3_OTA_BUCKET=wardn-ota" \
  --set-secrets "DATABASE_URL=wardn-database-url:latest,MAINTENANCE_DATABASE_URL=wardn-maintenance-url:latest,S3_ACCESS_KEY=wardn-r2-access-key:latest,S3_SECRET_KEY=wardn-r2-secret-key:latest,MQTT_TLS_CERT=wardn-mqtt-cert:latest,MQTT_TLS_KEY=wardn-mqtt-key:latest"

Two flags in that command are deliberate, not incidental:

  • --allow-unauthenticated is not a gap: wardn's own OIDC and API-key checks are the security boundary here, exactly as they are behind nginx locally. See 10. Security. Cloud Run's own IAM layer is not asked to do that job.
  • --max-instances 1 alongside --min-instances 1 is a deliberate choice, not just an economy: a second concurrent instance sharing the same INSTANCE_ID would fight the first one for the broker session. max 1 avoids that question entirely rather than pinning an id per revision. The INSTANCE_ID variable is left unset on purpose: Cloud Run replaces the instance on every deploy and occasionally on its own, and each replacement already pays the "clean session" cost described above. Nothing here makes that worse than it already is on any platform that does not hand you a stable pod identity. See 14 for the one that does, at the price of running it yourself.

Cost: inside the free tier for a single always-warm instance at low traffic. A few dollars a month once the idle time exceeds it, since a min-instances replica is billed for memory while idle even with no requests arriving.

Render, Railway or any other container platform with an "always on" setting is a drop-in substitute: the requirement is the setting, not the vendor.

The database: Supabase

A standard managed PostgreSQL instance, not a proprietary dialect: db-init runs against it exactly as it runs locally, creating wardn_owner, wardn_app and wardn_maintenance with the same grants (12. Operating).

That has not been run against a live Supabase project as part of writing this chapter. Confirm it with make verify against a real project before relying on it.

Supabase hands out two connection strings: a direct one and a pooled one (Supavisor, transaction mode). Use the pooled string for DATABASE_URL. The free tier's direct-connection ceiling is low, and a backend with DATABASE_POOL_MAX=10 will exhaust it fast on its own.

Transaction-mode pooling is exactly what the retention, reconciliation and fleet sweeps were built for: they take a lease row rather than a session-scoped advisory lock, so a sweep is a claim, then the work, then a release — each its own statement, none of them assuming the next one lands on the same server connection. That is the only shape a transaction pooler can serve across a piece of work that takes minutes.

⚠️ The free project pauses after seven days without a query. A fleet that heartbeats every ten seconds keeps it awake by construction. A project with no controllers talking to it yet, a demo between visits, a staging environment, will find it paused, and the first request after pays a cold-start delay rather than a clean answer.

Object storage: Cloudflare R2

Same S3 API MinIO already speaks locally, so S3_ENDPOINT, S3_ACCESS_KEY/S3_SECRET_KEY and S3_OTA_BUCKET keep their shape. Only the endpoint changes, to https://<account-id>.r2.cloudflarestorage.com, reached with an R2 API token scoped to the one bucket.

Presigned URLs work the same way they do against S3, so OTA_LINK_VALIDITY_S (12. Operating) needs nothing rewritten.

Why R2 specifically, and not the cheapest S3-compatible bucket in general: firmware artifacts are pulled by controllers in the field, on whatever uplink a site has, and S3's egress fee is a cost that scales with fleet size and rollout frequency. R2 charges nothing for egress, at any volume, on top of a free tier that covers 10 GB of storage and the request volumes a small fleet's OTA and document traffic will not come close to.

The broker: AWS IoT Core

Everything 14. Deploying to production already writes about IoT Core applies unchanged: same MQTT protocol, same IoT policy shape, the certificate's common name still confining a controller to its own topic subtree.

This chapter skips only Greengrass: the OTA provisioning layer is worth adopting once there is a fleet to provision, not before.

Cost, worked from wardn's own defaults: a controller publishes a heartbeat every HEARTBEAT_INTERVAL_S seconds, 10 by default (12. Operating).

  • That is 8,640 messages a day, about 259,000 a month, from heartbeats alone.
  • At IoT Core's per-message rate that is roughly $0.26 a month per controller.
  • A permanent MQTT connection adds a few thousandths of a dollar more in connection-minutes.
  • Passage events add to that total, but the heartbeat floor alone shows the order of magnitude: cents per controller, not dollars.
  • New AWS accounts also get 500,000 messages a month free for their first twelve months, which is close to covering one controller's heartbeats outright during that window.

Identity: Auth0

wardn asks nothing provider-specific of an OIDC issuer (10. Security), so the client-side configuration is exactly 10. Security's four boxes:

  • A public SPA client.
  • Authorization code with PKCE.
  • The dashboard's origin as the redirect and post-logout URI.
  • The wardn-admin role reaching the token.

That last box is the one Auth0 does not do by default. Unlike Keycloak, it does not put role assignments on the token without being told to. An Action on the Login flow does it:

exports.onExecutePostLogin = async (event, api) => {
  if (event.authorization) {
    api.accessToken.setCustomClaim('roles', event.authorization.roles);
  }
};

Assign the wardn-admin role to whichever operators may push firmware, everyone else gets wardn-user, same as the realm this repository imports for Keycloak.

The free plan covers up to 25,000 monthly active users, several orders of magnitude past what an operator headcount will ever be.

The dashboard: Cloudflare Pages

The repository already publishes to Cloudflare Pages: a workflow builds and deploys the documentation site and the developer reference on every push to main. The dashboard is the same pattern, one more cloudflare/pages-action step pointed at the dashboard's build output.

The one thing that step has to do differently: locally and in 14, config.js is written by an nginx entrypoint script that runs when the container starts. A static host has no container and no start, so generate the same file as a build step instead, before the upload:

cat > frontend/dashboard/dist/wardn/browser/config.js <<EOF
window.wardnConfig = {
  apiUrl: '${WARDN_API_URL}',
  oidcAuthority: '${WARDN_OIDC_AUTHORITY}',
  clientId: '${WARDN_OIDC_CLIENT_ID}',
  sentryDsn: '${WARDN_SENTRY_DSN:-}',
  environment: 'production',
  helpUrl: '${WARDN_HELP_URL}',
};
EOF

WARDN_HELP_URL needs its own value here, too: the nginx image mounts the built documentation at /help alongside the dashboard, which a bare static deploy of the dashboard build does not carry. Point it at the documentation site this same workflow already publishes, whether that is https://developer.wardn.xyz or wherever wardn-docs ends up, rather than a path that will 404.

What this costs, and where it stops being free

At the scale of a pilot or a single workspace, one Cloud Run instance, a few controllers, an operator headcount in the tens, four of these six pieces cost nothing, and the two that do not (the Cloud Run instance's idle time, IoT Core's messaging) come to a few dollars a month between them.

That is the entire pitch of this chapter over 14: no RDS instance, no EKS cluster, no Helm release to operate.

What is traded away is exactly what §14 buys:

  • No multi-AZ database.
  • One backend instance rather than three.
  • A Supabase project that pauses if nothing talks to it.
  • Free-tier ceilings (500 MB of database, 10 GB of object storage) that a growing fleet or a second workspace will reach.

None of that is a rewrite when it happens. The same argument 14 closes on holds here too: the backend, the controller and the dashboard read a URL and a credential, never a vendor's name. Moving from this chapter's services to §14's is swapping which secret a variable points at, not touching a line of code.


wardn runs the same way everywhere it runs, including the cheapest place it can.

Next → 16. The lab suitcase