Developer docs
System

Hosting & deploy

Compose locally, Kubernetes in production.

docs/DEPLOYMENT.md

Local is Docker Compose. Production is Kubernetes. Same two images either way.

Local — day-to-day

docker compose up -d postgres minio     # data services only
cp .env.example .env                    # then set JWT_SECRET
npm install
npm run db:push && npm run seed
npm run dev:api                         # terminal 1 → :4000
npm run dev                             # terminal 2 → :3000
ServicePortNotes
web3000Next dev server
api4000Hono, tsx watch
postgres5433offset so it never collides with another local stack
minio (S3)9100console on 9101

Local — production topology

docker compose --profile apps up -d --build

Four containers: postgres, minio, api, web. Use this before touching Kubernetes — if the images don't work here, they won't work there.

NEXT_PUBLIC_API_URL is a build arg, not a runtime env var: Next inlines NEXT_PUBLIC_* into the client bundle at build time. Changing the API address means rebuilding the web image.

Production — the droplet (what runs today)

A push to main runs .github/workflows/deploy.yml: verify (the PR's own checks, re-run against the merge commit) and the two image builds in parallel, then deploy — which waits on all three — ssh's deploy/roll-vm.sh onto the droplet.

The roll is ordered so that everything able to fail does so before anything that serves traffic is touched:

#StepOn failure
1record the running API_IMAGE / WEB_IMAGE pins— (they are the rollback target)
2stage compose + Caddyfile + .envaborts; nothing recreated
3docker compose pullaborts; prod untouched, still on the previous release
4docker compose run --rm migrateaborts; prod untouched
5docker compose up -d—
6poll /api/v1/health and / through Caddy, 180srestores the previous pins, re-rolls, exits non-zero
7prune to the 3 most recent images per repo—

Step 4 is the one that matters. migrate is also a compose dependency of api and worker (service_completed_successfully), so running it only as part of up meant a failed migration destroyed the live api container to recreate it and then blocked the replacement from starting — the site went down, while Caddy kept answering its static /healthz and the outage read as a tunnel fault. Running migrations as their own gated step makes a bad migration a no-op deploy.

What step 6 does not cover. /api/v1/health returns 200 whenever Postgres is up and reports a dead broker as queue: "down" in the body rather than a 503 — deliberately, so a queue outage does not take the API out of rotation. The worker has no healthcheck of its own. And only the api/web pins are restored on rollback, so a moving tag on any other service cannot be put back. The gate catches a broken release, not a degraded one.

Step 6 rolls back images, not the database. A migration applied in step 4 stays applied. Roll back only rescues the ordinary case where the new code is broken; a release whose migration is not backward-compatible with the previous image has to be rolled forward. That is the reason to keep migrations expand-then-contract.

A failed deploy opens (or comments on) an issue labelled deploy-failure, naming the stage that failed — which is what decides whether prod is serving the old release or is down.

Production — Kubernetes (F-081, not yet built)

WorkloadReplicasNotes
web Deployment2+stateless; readiness on /
api Deployment2+stateless; readiness on /api/v1/health (returns 503 when Postgres is down, so an unhealthy pod is removed from the Service rather than left serving errors)
worker Deployment1same image as api, different command. Scheduled jobs — publishing, metric collection, token refresh. Exactly one replica until jobs are made idempotent with leader election. Since F-159 it also hosts the 22 interval workers (TICK_WORKERS_HOST=worker), so scheduled publishing depends on it.
postgresmanagednot run in-cluster
object storagemanagedS3-compatible
Ingress—TLS; / → web, /api → api

Secrets (JWT_SECRET, DATABASE_URL, provider keys) come from a Secret, never an image layer.

Migration strategy

Run prisma migrate deploy as a Job in a pre-upgrade hook, not on container start. Two API replicas starting at once would otherwise race the migration lock.

Environment variables

Full list in .env.example. The rule:

PrefixRead byVisibility
(none)apps/api/src/core/config.ts onlyserver-side, may hold secrets
NEXT_PUBLIC_the browser bundlepublic — never a secret

Backup (F-084, not yet built)

Daily pg_dump to object storage, 30-day retention, plus a manifest recording the app version and the storage backend. Media is already durable in S3. The restore procedure is a documented command, not a runbook paragraph — an untested restore is not a backup.