Skip to content

← 13. Developing · Contents · Next → 15. A near-free deployment

14. Deploying to production

2. Getting started walks the same stack a production site runs, with every managed service played by a container: EMQX for the broker, MinIO for object storage, PostgreSQL in a container instead of a managed one. 12. Operating already says what changes:

Component Local Production
Broker EMQX container AWS IoT Core, or any MQTT broker with shared subscriptions
Object storage MinIO S3
Database PostgreSQL container Managed PostgreSQL
Identity Keycloak container Keycloak, or any OIDC provider

Nothing in the backend, the controller or the dashboard knows the difference: each of the four is reached through a URL and a credential, never through code that says "this is AWS".

This chapter is one concrete way to make that swap, on AWS, with Terraform for the server and a Helm chart for the backend. Treat the snippets as a worked shape, not a chart to copy-paste: sizing, networking and secrets handling belong to whoever runs the site.

flowchart LR
    subgraph Local["docker compose"]
        L1["EMQX"]
        L2["MinIO"]
        L3["PostgreSQL container"]
        L4["Keycloak container"]
    end
    subgraph AWS["This chapter"]
        A1["AWS IoT Core<br/>+ Greengrass on each controller"]
        A2["S3"]
        A3["RDS for PostgreSQL"]
        A4["Keycloak on EKS"]
        A5["Backend on EKS<br/>(Helm chart)"]
        A6["S3 + CloudFront<br/>(dashboard)"]
    end
    L1 -.->|"same MQTT protocol,<br/>same QoS 1, same<br/>shared subscriptions"| A1
    L2 -.->|"same S3 API"| A2
    L3 -.->|"same schema,<br/>same db-init"| A3
    L4 -.->|"same realm import"| A4

Building what gets deployed

The three images are the ones 13. Developing already builds in CI: wardn-backend, wardn-edge, wardn-frontend. Nothing about them is dev-specific. Tag them by commit, push to a registry (ECR here), and every environment runs the exact bytes CI tested.

aws ecr get-login-password --region eu-west-1 | \
  docker login --username AWS --password-stdin "$ECR_REGISTRY"

docker build -t "$ECR_REGISTRY/wardn-backend:$GIT_SHA" ./backend && docker push "$ECR_REGISTRY/wardn-backend:$GIT_SHA"

# The frontend image is built from the repo root, not ./frontend: its
# Dockerfile COPYs across the whole pnpm workspace (dashboard + the ui and
# tokens packages it depends on), which only resolves from there.
docker build -f frontend/Dockerfile -t "$ECR_REGISTRY/wardn-frontend:$GIT_SHA" . && docker push "$ECR_REGISTRY/wardn-frontend:$GIT_SHA"

The controller does not take an image in the field: it takes the wardn-edge binary, published to S3 and pushed through the OTA mechanism 7. The fleet already describes. Building it for a controller's actual hardware, rather than the container target docker build -t wardn-edge-device ./edge produces, is a cross-compilation concern outside this chapter.

The controllers: AWS IoT Core and Greengrass

The local broker's own comment states the intent directly: EMQX "stands in for AWS IoT Core: same MQTT protocol, same QoS 1 semantics, same shared-subscription support the backend relies on to scale out." Moving to IoT Core changes an endpoint and a certificate, not a line of the backend's MQTT handling.

A controller's certificate still decides which controller is talking, the same principle as 10. Security. An IoT policy expresses the same confinement EMQX's ACL does, with the same shape:

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": "iot:Connect",
      "Resource": "arn:aws:iot:eu-west-1:*:client/${iot:Connection.Thing.ThingName}",
      "Condition": { "Bool": { "iot:Connection.Thing.IsAttached": "true" } }
    },
    {
      "Effect": "Allow",
      "Action": ["iot:Publish", "iot:Receive"],
      "Resource": "arn:aws:iot:eu-west-1:*:topic/wardn/devices/${iot:Connection.Thing.ThingName}/*"
    },
    {
      "Effect": "Allow",
      "Action": "iot:Subscribe",
      "Resource": "arn:aws:iot:eu-west-1:*:topicfilter/wardn/devices/${iot:Connection.Thing.ThingName}/*"
    }
  ]
}

One IoT policy, attached to every controller's certificate: the ${iot:Connection.Thing.ThingName} variable is what confines each connection to its own subtree, exactly as ${username} does in 10. Security's EMQX rule. The IsAttached condition is what makes the variable mean something: a certificate no thing is attached to resolves it to nothing and connects as nobody.

The backend's own certificate gets a second policy scoped to wardn/#, matching wardn-backend's rule today, with its shared group named explicitly: on IoT Core a shared subscription is authorised on the $share/<group>/… filter itself, so the policy and MQTT_SHARED_GROUP have to agree.

This part of the chapter is not illustrative. infra/aws/ provisions the things, the certificates and both policies with Terraform, alongside the Cognito user pool and the S3 buckets below, and its docker-compose.yml runs the server half on a laptop against them. Its README walks one backend and one controller onto the broker, including the two limits IoT Core adds that EMQX does not have (sessions expiring after an hour away, a 30-second keep-alive floor). Its third, 128 KB per message, a full rights sync now splits itself across as many messages as it needs (chapter 5).

Greengrass runs the binary. It does not replace wardn's own OTA safety. Registering each controller as a Greengrass core device buys AWS-managed provisioning, remote log shipping and a local secret store for the device certificate, on top of IoT Core connectivity.

wardn-edge deploys as a Greengrass component, but the signature check, the digest check, the staged install and the boot-confirms-the-release rollback 7. The fleet describes stay exactly as written: that logic lives in the binary, not in whatever supervises it, on a container, a bare board or under Greengrass alike.

Generate a production keypair for OTA_SIGNING_KEY/OTA_PUBLIC_KEY before the first real release. Never reuse the dev pair .env.example ships. The private half stays with the release pipeline alone (a CI secret, never deployed to a controller or the backend). The public half is the OTA_PUBLIC_KEY every controller's Greengrass component config sets.

A controller with the wrong public key rejects every release, signed or not, which is the fail-closed behaviour to expect from a fleet that was provisioned with the dev key by mistake.

The server side, in Terraform

Four resources, matching the table at the top of this chapter. Illustrative, trimmed of the networking (VPC, subnets, security groups) every real deployment already has its own conventions for.

# ── Managed PostgreSQL ────────────────────────────────────────────────
resource "aws_db_instance" "wardn" {
  identifier          = "wardn"
  engine              = "postgres"
  engine_version      = "16"
  instance_class      = "db.r6g.large"
  allocated_storage   = 100
  multi_az            = true
  storage_encrypted   = true
  db_subnet_group_name   = aws_db_subnet_group.wardn.name
  vpc_security_group_ids = [aws_security_group.wardn_db.id]

  # db-init still owns the schema, the roles and the grants. This
  # provisions the instance. db-init is run once against its endpoint,
  # exactly as it runs against the local container today.
  username = "wardn_owner"
  password = data.aws_secretsmanager_secret_version.wardn_owner.secret_string
}

# ── Object storage: firmware, documents, backups ──────────────────────
# The two application buckets and the backend's key are real: infra/aws/.
resource "aws_s3_bucket" "ota"     { bucket = "wardn-ota" }
resource "aws_s3_bucket" "docs"    { bucket = "wardn-docs" }
resource "aws_s3_bucket" "backups" { bucket = "wardn-backups" }

resource "aws_iam_user" "backend" { name = "wardn-backend" }

# The scope S3_ACCESS_KEY / S3_SECRET_KEY carry: object read/write on the
# two application buckets, nothing on the account. The same confinement
# dev/infra/minio/entrypoint.sh gives the backend's service account locally.
resource "aws_iam_user_policy" "backend" {
  user = aws_iam_user.backend.name
  policy = jsonencode({
    Version = "2012-10-17"
    Statement = [{
      Effect   = "Allow"
      Action   = ["s3:GetObject", "s3:PutObject", "s3:DeleteObject", "s3:ListBucket"]
      Resource = [
        aws_s3_bucket.ota.arn, "${aws_s3_bucket.ota.arn}/*",
        aws_s3_bucket.docs.arn, "${aws_s3_bucket.docs.arn}/*",
      ]
    }]
  })
}

# ── The fleet's broker ─────────────────────────────────────────────────
# Real, not illustrative: infra/aws/ is the module, see above.

# ── The dashboard: a static build behind a CDN ─────────────────────────
resource "aws_s3_bucket" "dashboard" { bucket = "wardn-dashboard" }

resource "aws_cloudfront_distribution" "dashboard" {
  enabled             = true
  default_root_object = "index.html"

  origin {
    domain_name              = aws_s3_bucket.dashboard.bucket_regional_domain_name
    origin_id                = "dashboard"
    origin_access_control_id = aws_cloudfront_origin_access_control.dashboard.id
  }

  default_cache_behavior {
    target_origin_id       = "dashboard"
    viewer_protocol_policy = "redirect-to-https"
    allowed_methods        = ["GET", "HEAD"]
    cached_methods          = ["GET", "HEAD"]
    forwarded_values { query_string = false; cookies { forward = "none" } }
  }

  # A single-page-app route that is not a file: fall back to index.html
  # and let the client-side router take it from there.
  custom_error_response {
    error_code         = 404
    response_code       = 200
    response_page_path = "/index.html"
  }

  restrictions { geo_restriction { restriction_type = "none" } }
  viewer_certificate { cloudfront_default_certificate = true }
}

WARDN_API_URL in the dashboard's build still names the address a browser dials, per 12. Operating. Here that is the backend's own domain, not CloudFront's.

Keycloak, and the backend, on EKS

Keycloak stays exactly what it is locally: a realm imported from this repository, running as a normal workload rather than a managed AWS service, because 12. Operating already scopes identity as "Keycloak, or any OIDC provider" rather than promising a managed one. Running it on the same EKS cluster as the backend is the simplest option, not the only one.

The Helm chart's shape

wardn/
├── Chart.yaml
├── values.yaml
└── templates/
    ├── backend-statefulset.yaml
    ├── backend-service.yaml
    ├── backend-configmap.yaml
    ├── db-init-job.yaml
    └── ingress.yaml

The backend is a StatefulSet, not a Deployment. 12. Operating already says why: each instance needs a stable INSTANCE_ID, and "in Kubernetes this is the StatefulSet ordinal." A pod named wardn-backend-0 keeps that name across restarts. A Deployment's pods do not.

# values.yaml (excerpt)
backend:
  image: "111111111111.dkr.ecr.eu-west-1.amazonaws.com/wardn-backend"
  tag: "GIT_SHA"
  replicas: 3
  env:
    MQTT_URL: "mqtts://xxxxxxxxxxxxxx-ats.iot.eu-west-1.amazonaws.com:8883"
    OIDC_ISSUER_URL: "https://auth.wardn.example.com/realms/wardn"
    S3_ENDPOINT: "https://s3.eu-west-1.amazonaws.com"
    S3_OTA_BUCKET: "wardn-ota"
    CORS_ORIGINS: "https://dashboard.wardn.example.com"
  secretName: wardn-backend-secrets   # DATABASE_URL, S3 keys, MQTT client cert
# templates/backend-statefulset.yaml (excerpt)
apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: {{ .Release.Name }}-backend
spec:
  serviceName: {{ .Release.Name }}-backend
  replicas: {{ .Values.backend.replicas }}
  selector:
    matchLabels: { app: wardn-backend }
  template:
    metadata:
      labels: { app: wardn-backend }
    spec:
      containers:
        - name: backend
          image: "{{ .Values.backend.image }}:{{ .Values.backend.tag }}"
          # The pod's own name is `<statefulset>-<ordinal>`, e.g.
          # wardn-backend-0, wardn-backend-1: the same shape docker-compose
          # gives backend / backend-2 today.
          env:
            - name: INSTANCE_ID
              valueFrom: { fieldRef: { fieldPath: metadata.name } }
          envFrom:
            - configMapRef: { name: {{ .Release.Name }}-backend-config }
            - secretRef: { name: {{ .Values.backend.secretName }} }
          ports: [{ containerPort: 3000 }]
          readinessProbe: { httpGet: { path: /health/ready, port: 3000 } }
          livenessProbe: { httpGet: { path: /health, port: 3000 } }

db-init runs unchanged: still the single source of truth for the schema, the roles and the grants, still idempotent, now as a Helm pre-install/pre-upgrade hook Job against the RDS endpoint instead of the local container.

# templates/db-init-job.yaml (excerpt)
apiVersion: batch/v1
kind: Job
metadata:
  name: {{ .Release.Name }}-db-init
  annotations:
    "helm.sh/hook": pre-install,pre-upgrade
spec:
  template:
    spec:
      restartPolicy: Never
      containers:
        - name: db-init
          image: postgres:16-alpine
          envFrom: [{ secretRef: { name: {{ .Values.backend.secretName }} } }]
          command: ["/bin/sh", "/db/init/entrypoint.sh"]

Every other environment variable in this chart is the same one documented in 12. Operating: production changes where the value points, never what the variable means.

Using a managed OIDC provider instead

Running Keycloak on EKS is the simplest option, not the only one. Nothing above needs it: the backend and the dashboard read an issuer URL and a client id, never a brand. Pointing wardn at Auth0, Okta, Azure AD or Cognito instead removes A4 and its realm-import job from the picture entirely. infra/aws/ takes the Cognito route: a user pool, a public client with the dashboard's origin as its callback, the two roles as groups, and the operators invited by email.

# ── Auth0, as one example provider ─────────────────────────────────────
# Illustrative: this chapter otherwise uses only the aws provider, and
# adding another one is specific to whichever provider you land on.
resource "auth0_client" "dashboard" {
  name                 = "wardn-dashboard"
  app_type             = "spa" # a public client: Auth0 never issues it a secret
  callbacks            = ["https://dashboard.wardn.example.com/"]
  allowed_logout_urls  = ["https://dashboard.wardn.example.com/"]
  grant_types          = ["authorization_code"]
}

output "oidc_issuer_url" {
  value = "https://${var.auth0_domain}/"
}

The values in values.yaml and config.js change. Nothing in the chart or the images does:

Variable Keycloak on EKS A managed provider
OIDC_ISSUER_URL (backend) https://auth.wardn.example.com/realms/wardn https://${auth0_domain}/
OIDC_JWKS_URL (backend) (unset, discovered) (unset, discovered)
WARDN_OIDC_AUTHORITY (dashboard) https://auth.wardn.example.com/realms/wardn https://${auth0_domain}/
WARDN_OIDC_CLIENT_ID (dashboard) wardn-spa the provider's client id

The one thing to configure by hand, whatever the provider: the wardn-admin role reaching the access token as roles or realm_access.roles, done with an Auth0 Action, an Okta claims policy, or an Azure AD app role, depending on where you land. See 10. Security for the full contract.

Hardening checklist

This chapter has deliberately left network and firewall rules to whatever conventions a deployment already has. Three things are not a matter of convention, though. They are absent from every environment until someone adds them, and are worth naming explicitly rather than leaving as an exercise.

A WAF and rate limiting in front of the API. The application's own RATE_LIMIT_MAX (§10 Security) bounds cost after a request already reached a backend pod. Nothing upstream of that exists yet. Attach AWS WAFv2 to the ingress ALB, with at minimum a rate-based rule and the managed common-threats rule group:

resource "aws_wafv2_web_acl" "api" {
  name  = "wardn-api"
  scope = "REGIONAL"

  default_action { allow {} }

  rule {
    name     = "rate-limit"
    priority = 1
    action { block {} }
    statement {
      rate_based_statement {
        limit              = 6000 # requests per 5-minute window, per IP
        aggregate_key_type = "IP"
      }
    }
    visibility_config {
      cloudwatch_metrics_enabled = true
      metric_name                = "wardn-api-rate-limit"
      sampled_requests_enabled   = true
    }
  }

  rule {
    name     = "aws-common"
    priority = 2
    override_action { none {} }
    statement {
      managed_rule_group_statement {
        name        = "AWSManagedRulesCommonRuleSet"
        vendor_name = "AWS"
      }
    }
    visibility_config {
      cloudwatch_metrics_enabled = true
      metric_name                = "wardn-api-common-threats"
      sampled_requests_enabled   = true
    }
  }

  visibility_config {
    cloudwatch_metrics_enabled = true
    metric_name                = "wardn-api"
    sampled_requests_enabled   = true
  }
}

resource "aws_wafv2_web_acl_association" "api" {
  resource_arn = aws_lb.api.arn # the ALB behind templates/ingress.yaml
  web_acl_arn  = aws_wafv2_web_acl.api.arn
}

A secret-rotation procedure, covering at minimum:

  • The database password (RDS supports rotation through Secrets Manager natively, the same secret values.yaml's DATABASE_URL already reads from).
  • The S3 access key issued to the backend's service account.
  • The backend's own mTLS certificate against IoT Core.
  • If a client secret is in play (a confidential OIDC client, or wardn-backend's Keycloak secret if that client is ever actually used), that too.

None of these need to be dramatic: a scheduled Lambda rotating the DB credential every 90 days, alongside a runbook for the two that are manual (mTLS cert, OIDC client secret). But "rotated on a schedule" has to be a decision someone made, not a gap nobody noticed.

Certificate-expiry monitoring for the broker and backend identities in production, mirroring certificate_expiring on the controller side (10. Security). IoT Core's device and backend certificates expire on the schedule this provisioning step gives them. Nothing watches that deadline unless something is told to.

A scheduled Lambda checking iot:DescribeCertificate against every certificate this stack issued, alerting through the same channel as the fleet's own alerts, closes the gap the local TLS-check script closes locally: that script has no production equivalent on its own, since it is wired to the dev stack's own TLS volume.

What stays exactly as documented

  • Certificates. The authority is the manufacturing CA, and a controller's key never leaves its secure element, per 10. Security. IoT Core's device certificates are provisioned through it the same way.
  • Backups. make backup already targets an S3-compatible endpoint through BACKUP_REMOTE_ENDPOINT. See 12. Operating. Point it at the wardn-backups bucket above and nothing else changes: same manifest, same encryption before it leaves, same make backup-verify.
  • Retention and erasure. Unaffected by where PostgreSQL runs. Both are the application talking to its own database, per 11. Personal data.

wardn runs the same way everywhere it runs. That was the whole argument of this documentation, and this chapter is where it gets tested against a real server.

This chapter is one platform's answer. The obligations that hold on every platform — and the checklist to sign before real doors depend on them — are 18. The security agreement.

Next → 15. A near-free deployment