Infrastructure
Postgres, RabbitMQ, trigger.dev, the WhatsApp bridge, and the scraping rules.
docs/INFRASTRUCTURE.md
What runs, why it runs, and what it would take to remove it. Local is Docker Compose; production is Kubernetes with the same images.
The pieces
| Service | Role | Local | Production |
|---|---|---|---|
postgres | System of record. Also the rate-limit store and the read-model for billing. | container, :5433 | managed Since F-160 the prod service runs with a tuned command (256 MB buffers, SSD planner cost) and a 256 MB /dev/shm; pool sizes are per process via DATABASE_POOL_SIZE. |
minio | S3-compatible object storage — media, ad assets, generated audio. | container, :9100 | managed S3-compatible |
rabbitmq | Work queue for short-lived, high-fan-out jobs. | container, :5673 | managed or in-cluster |
api | The HTTP surface. Stateless. | container / npm run dev:api | Deployment, 2+ replicas |
worker | Same image as api, different command. Consumes the queues and, since F-159, hosts the 21 interval workers (TICK_WORKERS_HOST=worker), so scheduled publishing runs here. Healthcheck: dist/worker-healthcheck.js reads the heartbeat mark (fresh within 45 s ⇒ healthy); the roll gate requires it. | container | Deployment |
web | The UI. No database driver. | container / npm run dev | Deployment, 2+ replicas |
whatsapp | Baileys bridge. Long-lived WhatsApp Web session. | container, :4100 | StatefulSet, 1 replica |
| trigger.dev | Durable multi-step workflow orchestration. | cloud (or self-hosted) | cloud |
The MinIO image (#308)
No public registry serves the MinIO image prod runs (quay.io/minio/minio:RELEASE.2025-09-07T16-13-09Z): Docker Hub answers 401 and quay.io refuses anonymous pulls. The prod droplet's cached copy was copied, layer for layer, by .github/workflows/mirror-minio.yml to:
ghcr.io/verjson/marketing-studio/minio@sha256:a1a8bd4ac40ad7881a245bab97323e18f971e4d4cba2c2007ec1bedd21cbaba2
The package is private and owned by this repository, like marketing-studio/api, so a droplet logged in to GHCR can pull it. docker-compose.prod.yml reads MINIO_IMAGE. A droplet with no cached copy (nonprod, a rebuilt prod) sets it to the reference above. Unset, it falls back to prod's existing image, so prod's MinIO is never recreated by this change.
The Google Cloud copy of each candidate (#381)
The org's release workflow will only promote a candidate that also exists in Google Artifact Registry (container-release.yml refuses a candidate without exactly one GAR destination). So every candidate is also copied, digest for digest, to:
us-central1-docker.pkg.dev/verjson-ci-640463/marketing-studio-candidates/marketing-studio-candidate-{api,web}
It lives in verjson-ci's Google Cloud project, next to verjson-ci's own container-canary repository, and is set up the same narrow way:
| Piece | Name | Scope |
|---|---|---|
| Docker repository | marketing-studio-candidates (us-central1) | Only ms-oci-publisher can write to it |
| Service account | ms-oci-publisher@verjson-ci-640463.iam.gserviceaccount.com | No project roles; artifactregistry.writer on that one repository |
| Trust provider | marketing-studio-main in pool github-oci-candidates | GitHub tokens from this repository (id 1316850610), on refs/heads/main, from the org's container-candidate-publish.yml (push) or container-release.yml (dispatch) at contract SHA a835699 |
The trust condition names the contract SHA. When the generated files move to a new Verjson/.github pin, the publish fails at Google sign-in until the condition names the new SHA too. Update it in the same change, with gcloud iam workload-identity-pools providers update-oidc marketing-studio-main --workload-identity-pool=github-oci-candidates --location=global --project=verjson-ci-640463 --attribute-condition=... (swap the SHA in the current condition, read with describe).
Two kinds of "background work", and why both exist
They are not competing choices — they solve different problems, and using one for the other's job goes badly in a predictable way.
| RabbitMQ | trigger.dev | |
|---|---|---|
| Unit | One short job | One long, multi-step run |
| Lifetime | Milliseconds to seconds | Minutes to days |
| Example | "Publish this post to LinkedIn", "pull yesterday's metrics for account X", "scrape this page" | "Research the competitor set → draft a plan → wait for approval → schedule 40 posts" |
| Needs | Fan-out, fair dispatch, per-queue rate limiting, dead-lettering | Durable state across steps, resumability, waiting on a human, retry-from-step |
| Failure mode if misused | A multi-day workflow in a queue means either a message with a huge TTL or a hand-rolled state machine in the payload | A per-post publish job in a workflow engine means paying orchestration overhead thousands of times an hour |
RabbitMQ because nearly every consumer here talks to a rate-limited third-party API. Per-queue prefetch, fair dispatch and a native dead-letter path are exactly the primitives that need — all of which a polling table has to reimplement, usually badly. See D-014.
trigger.dev for the agentic layer. An agent run is long, multi-step, and has to survive a deploy and a wait-for-human — which is precisely what a durable workflow engine is for and what a queue is not. See D-016.
Postgres keeps the rate limiter (a fixed-window counter, not a queue) and every read-model.
Queue topology
verjson.publish → publish a post to one channel (per-channel queues, prefetch 1)
verjson.metrics → pull metrics for one account/day
verjson.tokens → refresh OAuth tokens expiring soon
verjson.scrape → fetch and parse one URL (politeness delay per host)
verjson.whatsapp → send one approved template
verjson.dlq → anything that exhausted its retries
Each queue is bound to a direct exchange by routing key. Retries use a delayed retry queue rather than an in-process sleep, so a redeploy does not lose the backoff.
WhatsApp: Baileys
Baileys speaks the WhatsApp Web protocol from a persistent socket, so it runs as its own container with its own volume — a long-lived paired session cannot share a lifecycle with stateless API pods that get rescheduled. The volume holds the pairing; losing it means re-pairing by QR code.
The honest trade-off. Baileys is an unofficial, reverse-engineered client. Bulk marketing over it is against WhatsApp's terms, and the realistic downside is the number being banned — not a rate-limit, a ban. It is the pragmatic choice for an internal tool and for conversational replies; it is the wrong choice for a large cold broadcast.
The mitigation is structural, not procedural: the bridge sits behind the same ChannelProvider
interface as every other channel, so the official WhatsApp Business Cloud API is a transport
swap rather than a rewrite. If volume grows or a ban lands, changing one adapter moves the whole
feature across. See D-015.
Operationally: opt-in is enforced before anything is queued, sends are paced, and template content still passes the approval gate.
The bridge's surface
services/whatsapp — a standalone ESM Node 22 service, outside the npm workspace so it can
hold a stale Baileys pin without dragging the API's dependency tree along with it. Baileys is
pinned to 6.7.24: npm's latest tag is a 7.0.0 release candidate, and legacy is the last
stable line.
Every route but /health requires x-bridge-token, compared in constant time. /health is open
because it is the container's HEALTHCHECK and the StatefulSet's probe, and it reveals only that
the process is up.
| Method | Path | Returns | Notes |
|---|---|---|---|
GET | /health | { status: "ok", bridge: "up" } | No token. Liveness only |
GET | /status | { connectionState, phone, pushName, qrAvailable, lastDisconnectReason, connectedAt } | connectionState is Baileys' own open | connecting | close |
GET | /qr | { qr: "data:image/png;base64,…", connectionState, phone } | 409 once paired, or before the first QR is issued. A QR older than 60 s is not served — a stale code just wastes a scan |
POST | /send | 202 { id, to, status: "sent" } | Body { to, text?, mediaUrl?, mediaType?: "image" | "video" }. to takes a raw number, +-formatted number, or a full jid. 409 when the socket is not open, 502 when WhatsApp rejects it |
POST | /disconnect | { status: "disconnected", …status } | Logs the device out, wipes /app/auth, comes back up un-paired so the next /qr pairs a different number. Destructive and unrecoverable |
Outbound, the bridge POSTs to WEBHOOK_URL (default http://api:4000/api/v1/whatsapp/events)
with the same x-bridge-token. Payloads are { event, occurred_at, data } and snake_case,
because that is what the API validates inbound; the bridge's own REST responses follow the table
above instead.
event | Fires on | data |
|---|---|---|
messages.upsert | A real inbound message (type: "notify" only — history backfill is dropped) | { type, messages: [{ id, chat_jid, sender_jid, from_me, is_group, push_name, timestamp, type, text }] } |
message.status | Delivery / read receipts for messages we sent | { statuses: [{ id, chat_jid, from_me, status }] } |
connection.update | The socket opening, connecting or closing | { connection_state, logged_out?, reason?, phone?, is_new_login? } |
qr.updated | A new pairing QR was issued | { connection_state } — a nudge to re-fetch /qr |
Delivery is fire-and-forget. A webhook failure is logged and dropped: it must never take the socket down or fail a caller's send, because the API is the bridge's audience, not its dependency. Durability for anything that matters belongs on the API side of that hop.
A dropped socket reconnects with exponential backoff to 30 s. A close carrying loggedOut (the
number was un-linked from the phone) clears the auth state instead of looping on dead credentials,
so the bridge lands back at "show a QR".
Scrapers
Scraping is a queue consumer like any other (verjson.scrape), which means it inherits retries,
dead-lettering and rate limiting for free rather than growing its own scheduler.
Rules that keep it defensible:
- Respect
robots.txtand a per-host politeness delay. Enforced in the consumer, not left to each scraper. - Identify honestly in the user agent, with a contact URL.
- Cache aggressively. A page fetched today is not fetched again tomorrow unless something asks for a refresh.
- Public pages only. Nothing behind a login, and nothing a
robots.txtdisallows. - Personal data has a retention clock from the moment it lands — see the Apollo note in FEATURES.md (F-092).
Agent-Reach
Agent-Reach is a capability layer that gives agents unified read access across ~14 platforms (X, Reddit, LinkedIn, YouTube, GitHub, RSS, podcasts, web search) by routing to native CLI tools with ordered fallbacks, rather than wrapping them in a new abstraction.
Where it fits here: the research half of the growth-intelligence work (F-091, F-092, F-094) and the agentic layer (F-071). It is installable as an MCP server, which means our agents can reach it through the same MCP path they already use — no bespoke integration, and no scraper of our own for platforms it already covers.
Where it does not fit: anything that writes. It is a read and search layer; publishing stays with our own channel adapters, behind the approval gate.
Tracked as F-098.
Nonprod (#379, #383)
A second droplet, ms-nonprod (144.126.249.116, Pulumi stack nonprod), runs the same Compose stack at http://144.126.249.116.nip.io. It is deployed only by dispatching deploy nonprod (.github/workflows/deploy-nonprod.yml), which builds nothing: it rolls an existing candidate's signed digests through prod's own deploy/roll-vm.sh. Its GitHub environment nonprod holds its own SSH key and freshly generated secrets; none of prod's are shared. The database started empty; SEO providers run in mock mode. With no CLOUDFLARED_TOKEN, the cloudflared container restarts in a loop, which the health gate ignores; it stops once nonprod has a tunnel hostname.
Environment
Every value is listed in .env.example. The rule stands: server-side values are
read only in apps/api/src/core/config.ts, and anything prefixed NEXT_PUBLIC_ is public and must
never hold a secret.
| Added by this work | For |
|---|---|
AMQP_URL | RabbitMQ connection |
WHATSAPP_BRIDGE_URL · WHATSAPP_BRIDGE_TOKEN | The API↔Baileys hop (the bridge is never internet-facing) |
TRIGGER_SECRET_KEY · TRIGGER_API_URL | trigger.dev |
SEMRUSH_API_KEY · APOLLO_API_KEY | Growth intelligence |
SCRAPER_USER_AGENT · SCRAPER_CONTACT_URL | Identifying ourselves honestly when fetching |