Hosting & deploy
Compose locally, Kubernetes in production.
docs/DEPLOYMENT.md
Local is Docker Compose. Production is Kubernetes. Same two images either way.
Local — day-to-day
docker compose up -d postgres minio # data services only
cp .env.example .env # then set JWT_SECRET
npm install
npm run db:push && npm run seed
npm run dev:api # terminal 1 → :4000
npm run dev # terminal 2 → :3000
| Service | Port | Notes |
|---|---|---|
| web | 3000 | Next dev server |
| api | 4000 | Hono, tsx watch |
| postgres | 5433 | offset so it never collides with another local stack |
| minio (S3) | 9100 | console on 9101 |
Local — production topology
docker compose --profile apps up -d --build
Four containers: postgres, minio, api, web. Use this before touching Kubernetes — if the
images don't work here, they won't work there.
NEXT_PUBLIC_API_URL is a build arg, not a runtime env var: Next inlines NEXT_PUBLIC_* into
the client bundle at build time. Changing the API address means rebuilding the web image.
Production — the droplet (what runs today)
A push to main runs .github/workflows/deploy.yml: verify (the PR's own checks, re-run
against the merge commit) and the two image builds in parallel, then deploy — which waits on
all three — ssh's deploy/roll-vm.sh onto the droplet.
The roll is ordered so that everything able to fail does so before anything that serves traffic is touched:
| # | Step | On failure |
|---|---|---|
| 1 | record the running API_IMAGE / WEB_IMAGE pins | — (they are the rollback target) |
| 2 | stage compose + Caddyfile + .env | aborts; nothing recreated |
| 3 | docker compose pull | aborts; prod untouched, still on the previous release |
| 4 | docker compose run --rm migrate | aborts; prod untouched |
| 5 | docker compose up -d | — |
| 6 | poll /api/v1/health and / through Caddy, 180s | restores the previous pins, re-rolls, exits non-zero |
| 7 | prune to the 3 most recent images per repo | — |
Step 4 is the one that matters. migrate is also a compose dependency of api and worker
(service_completed_successfully), so running it only as part of up meant a failed migration
destroyed the live api container to recreate it and then blocked the replacement from starting —
the site went down, while Caddy kept answering its static /healthz and the outage read as a
tunnel fault. Running migrations as their own gated step makes a bad migration a no-op deploy.
What step 6 does not cover. /api/v1/health returns 200 whenever Postgres is up and reports a
dead broker as queue: "down" in the body rather than a 503 — deliberately, so a queue outage does
not take the API out of rotation. The worker has no healthcheck of its own. And only the api/web
pins are restored on rollback, so a moving tag on any other service cannot be put back. The gate
catches a broken release, not a degraded one.
Step 6 rolls back images, not the database. A migration applied in step 4 stays applied. Roll back only rescues the ordinary case where the new code is broken; a release whose migration is not backward-compatible with the previous image has to be rolled forward. That is the reason to keep migrations expand-then-contract.
A failed deploy opens (or comments on) an issue labelled deploy-failure, naming the stage that
failed — which is what decides whether prod is serving the old release or is down.
Production — Kubernetes (F-081, not yet built)
| Workload | Replicas | Notes |
|---|---|---|
web Deployment | 2+ | stateless; readiness on / |
api Deployment | 2+ | stateless; readiness on /api/v1/health (returns 503 when Postgres is down, so an unhealthy pod is removed from the Service rather than left serving errors) |
worker Deployment | 1 | same image as api, different command. Scheduled jobs — publishing, metric collection, token refresh. Exactly one replica until jobs are made idempotent with leader election. Since F-159 it also hosts the 22 interval workers (TICK_WORKERS_HOST=worker), so scheduled publishing depends on it. |
postgres | managed | not run in-cluster |
| object storage | managed | S3-compatible |
Ingress | — | TLS; / → web, /api → api |
Secrets (JWT_SECRET, DATABASE_URL, provider keys) come from a Secret, never an image layer.
Migration strategy
Run prisma migrate deploy as a Job in a pre-upgrade hook, not on container start. Two API
replicas starting at once would otherwise race the migration lock.
Environment variables
Full list in .env.example. The rule:
| Prefix | Read by | Visibility |
|---|---|---|
| (none) | apps/api/src/core/config.ts only | server-side, may hold secrets |
NEXT_PUBLIC_ | the browser bundle | public — never a secret |
Backup (F-084, not yet built)
Daily pg_dump to object storage, 30-day retention, plus a manifest recording the app version and
the storage backend. Media is already durable in S3. The restore procedure is a documented command,
not a runbook paragraph — an untested restore is not a backup.