Logs
Decision log
Decisions with their alternatives and the reason — including the ones that turned out wrong.
docs/DECISIONLOG.md
Architecture, product, and library decisions with rationale. Newest first.
Add a row when you pick between real options and the reason won't be obvious from the code.
| ID | Date | Decision | Rationale | Alternatives considered | Status |
|---|---|---|---|---|---|
| D-061 | 2026-10-06 | Nonprod is deployed only by dispatch, and only a candidate whose Cosign signature names that main commit and the org publisher; the droplet pulls with the job's own token (#383) | Nonprod is where a release soaks before prod (#387), so a merge must not replace what is soaking: no push trigger. Verifying the signature before the roll means nonprod runs only what the org publisher built from main, and the digest rolled is the digest verified, so a moved tag cannot change it. The job token (packages: read, the packages are linked to this repository) expires with the job, so nonprod holds no long-lived registry credential, unlike prod's GHCR_PAT. The roll itself is prod's deploy/roll-vm.sh, unchanged, so nonprod exercises the same migration and health gates prod will. | Auto-deploy every merge (nonprod then tracks main and nothing soaks); copy prod's GHCR_PAT (a long-lived credential on a second box); a separate nonprod roll script (two gates to keep in step) | ✅ Active |
| D-060 | 2026-10-06 | Candidates are also copied to Google Artifact Registry, in verjson-ci's Google Cloud project, through a trust provider that admits only this repository's main and the org's exact publish and release workflows (#381) | The org's release workflow will not promote a candidate without exactly one GAR destination, so releases (#387) need one. verjson-ci's project already has the GitHub trust pool, so marketing-studio adds its own repository, service account and provider there instead of standing up a new project. Isolation is per repository: the provider's condition pins this repository's id, main, and the contract SHA, and the service account can write only to marketing-studio-candidates. Cost: a new contract pin needs the condition updated (see docs/INFRASTRUCTURE.md). | A new Google Cloud project for marketing-studio (its own bill, but a project, billing account, pool and provider to stand up); skipping releases and deploying from candidates directly (gives up numbered, signed releases) | ✅ Active |
| D-059 | 2026-10-06 | Every merge to main builds api and web once, as signed release candidates, with the org's generated publisher (#381) | A release has to run the same bytes on nonprod and prod, so the image is built once per commit and promoted unchanged (#387), never rebuilt per environment. The publisher is generated from Verjson/.github at a835699, never hand-written. Its candidates go to their own GHCR repos (marketing-studio-candidate-api / -web; the unprefixed names belong to older org packages this repository cannot write), apart from deploy.yml's images, until the cutover (#384). It signs with keyless Cosign, which writes the repository name, commit, workflow and image digest to Sigstore's public Rekor log; the user accepted that on 2026-10-06. GitHub's private Artifact Attestations are not available on verJSON's plan, which is why the org switched (Verjson/.github #1690). Candidates take no build args, so the web image's API URL defaults to the same-origin /api/v1, and per-environment values move to runtime (#400). The API candidate likewise carries no SOURCE_COMMIT, so the release deploy sets it at runtime beside RELEASE_VERSION (#383), or /version would answer commit: null. | Keep building in deploy.yml per environment (nonprod would test different bytes than prod ships); pin the older attestation-based contract as self-publish does (needs a GitHub plan verJSON lacks); hand-write a publisher (the org contract forbids it). | ✅ Accepted |
| D-058 | 2026-10-06 | Nonprod is its own droplet, its own Pulumi stack, its own DO project and firewall, and its own SSH deploy key (#379) | Releases soak on nonprod before prod (#387), so it has to resemble prod without being able to touch it. A second Compose stack on the prod droplet would share a 4 GB box with prod's Postgres, RabbitMQ and MinIO. A separate stack keeps the two Pulumi states apart, and a separate deploy key means a leak of nonprod's credentials cannot reach prod. The program reads per-stack names from config with prod's existing values as defaults, so prod previews unchanged. | A second Compose project on the prod droplet (free, but it competes for memory and shares the failure domain); reusing prod's deploy key (simpler, but a nonprod leak would open prod). | ✅ Accepted |
| D-056 | 2026-10-06 | Social listening is a project feature under Performance, open to project members, with SocialCrawl spending limited to owners/admins and capped per organization per month (supersedes the platform-admin-only part of D-054) | Product decision to ship listening to customers. The data was project-scoped from the start, so opening it is an access change: reads use projectScope like every project feature (another tenant's project is a 404); creating, changing or running terms and SocialCrawl searches need an org owner or admin, like tracked SEO keywords. The deployment holds one SocialCrawl key for all tenants, so each organization gets a monthly credit budget (LISTENING_MONTHLY_CREDITS_PER_ORG, 402 when exceeded) read from listening_runs — explorer searches are recorded as runs so nothing escapes it. The explorer reads official APIs with the project's own connected accounts only (it previously used the newest account in the deployment). Also: a per-term run claim (run_started_at) so a term never runs twice at once, per-project term cap, rate limits, audit entries, skipping suspended orgs and archived projects, and a retention purge (LISTENING_RETENTION_DAYS). Open item: SocialCrawl collects from logged-out public pages; the legal/compliance review D-054 asked for has not been recorded. | Keep platform-admin only (no customer value); per-plan credit meters in billing (needs pricing decisions — the env budget is the interim); no budget (one tenant could drain the shared key) | Accepted |
| D-055 | 2026-10-05 | The mention explorer reads platform-specific fields from raw at display time; only an engagement total is added as a column | Each platform reports a different set (TikTok country and downloads, Reddit subreddit and upvote ratio, YouTube subscribers), so a column per field would be empty for most rows and need a migration per new field. Reading raw in one presentation module (listening-tracking/present.ts) shows whatever a platform sent, including fields it adds later. The one exception is sorting: "most engagement" must order in the database across pages, so engagement_total is derived when a mention is written and backfilled once from existing columns. | A column per platform field (sparse, migration per field); sorting engagement in the browser (wrong across pages); a raw-SQL sum in ORDER BY (duplicates the filter logic outside Prisma) | Accepted |
| D-054 | 2026-10-01 | Listening tracking is platform-admin only for now, but project-scoped from day one | The data is SocialCrawl's, collected from logged-out public pages — scraping as far as Meta, TikTok and X's terms go — so customers should not see it before a compliance sign-off. Keying terms, runs and mentions on the project now means opening it to members later is a permission change, not a migration. | Org-level tables (wrong unit: an agency tracks different brands per client); customer-visible now (compliance risk); a platform-only table with no project (needs a data migration to open up) | Accepted |
| D-053 | 2026-10-01 | A listening mention is stored once per (term, platform, external id) and refreshed on every sweep, with the provider's full item kept in raw | A post seen on five daily sweeps is one conversation, not five; storing it per run would multiply every count. Refreshing keeps the latest engagement, first_seen_at gives the new-per-day trend, and run-level totals live on listening_runs. raw lets later analytics read fields not yet promoted to columns without re-fetching — every re-fetch costs credits. | One row per run (inflates counts, needs de-dup in every query); only the latest snapshot with no first-seen (loses the trend); columns only, no raw (every new analytic needs a paid re-crawl) | Accepted |
| D-052 | 2026-09-25 | API test files run in parallel, one cloned Postgres database per Vitest worker, and each case resets with DELETE instead of TRUNCATE | The serial suite was ~38 of the 43 minutes of ci / build-test. A per-worker database keeps the per-case wipe from reaching another file's rows without changing a single test. CREATE DATABASE … TEMPLATE clones the already-migrated database in about a second, so the migrations still run once. TRUNCATE had to go once files ran in parallel: it gives every table a new file, and the checkpoint fsync of millions of them stalled resets past the hook timeout; DELETE is ~60× faster per reset and creates no files. Workers default to one per core capped at 4, since past that they queue on the one Postgres they share | Transaction-per-test rollback — breaks every test whose code opens its own transaction or runs work concurrently. One Postgres schema per worker — advisory locks are per database, not per schema, so workers would contend on each other's locks, and setup.ts looks constraints up by name across the whole database; it would also need a migrate per schema. prisma migrate deploy per worker database — correct but ~20 s each. Keep TRUNCATE and raise the hook timeout — hides the stall rather than removing it | Accepted |
| D-051 | 2026-09-24 | The deploy skips its CI re-run when the merged tree provably already passed CI | verify re-ran the whole suite (~45 min) on every merge, although a squash merge of an up-to-date PR is byte-identical to the tree the PR's CI passed — so merge-to-live took ~50 min instead of ~10. deploy/already-tested.sh answers yes only when the merge commit's tree equals the PR head's tree and that head's successful CI run contains a ci / build-test job whose Run npm test step succeeded and whose Report deferred CI step was skipped (the org workflow reports a deferred build-test as green), or when this exact commit already deployed successfully. Any API error answers no, and a failed check leaves verify running (!cancelled()), so an error costs time, never the gate. | Require branches up to date (every merge forces a rebase and a fresh 45-min run); GitHub merge queue (org-level ruleset change); drop verify (lets an untested commit roll to prod again). | Accepted |
| D-050 | 2026-09-21 | A project may register several apps for one platform, and each account remembers which app connected it | Asked for directly: one project serving several brands whose Meta apps must stay separate. The load-bearing fact is that an OAuth token is only valid against the app that issued it — so the choice is not "which app is the project's" but "which app is this ACCOUNT's". social_accounts_connected.channel_credential_id records it at connect, and every account-scoped path (publish, metrics, inbox, reply, resubscribe) resolves the adapter from it. NULL keeps the old meaning (the project's primary app, first by created_at), so no backfill and no behaviour change for existing accounts. Each app mints its own webhook token, so the delivery URL names the app whose secret verifies it. Deleting an app with accounts on it is refused (409) rather than silently re-pointing them at another app, which would fail at the platform with an error naming neither. | (a) One project per client — already possible and simpler, but the ask was one project. (b) Several apps with no per-account link, picking "the" app by project — breaks every account connected through the non-primary app. (c) Cascade-delete accounts with their app — destroys reporting history. (d) Allow it for Google Ads/Analytics too — they resolve "the project's app" with nothing to say which of several was meant, so they stay single. | ✅ |
| D-046 | 2026-09-11 | The backdrop is its own axis from the colours, and one named "Look" sets all four | The brief already carried background, foreground and accent, and the UI sent none of them — so every video ever made was the hardcoded midnight mesh, and the three fields were unreachable from the product. Wiring them up alone would not have been enough: the mesh was the ONLY backdrop, so a cream background came out as pastel pools rather than as paper, which is the ground most explainer footage actually wants. Hence backdrop — mesh (the default, so nothing existing changes) · solid · paper · grid — as a separate axis, because "warm cream" and "warm cream under a mesh of coloured pools" are different videos built from the same three hex values. Two consequences fell straight out. The vignette is mesh-only: darkened corners on a light ground read as a printing fault, not as depth. And every text shadow now follows the foreground — they were all rgba(0,0,0,…), which lifts white type off a dark mesh and turns near-black type on cream into a muddy blur; isLight uses relative luminance rather than a channel average, since the average calls a saturated green light and the eye does not. That last one is also the cheapest fix for the real legibility problem in the first rendered video: white captions over white app screenshots were almost invisible, and a light look makes them dark and readable without a caption pill. The UI exposes ONE control, not four inputs — a named look sets backdrop plus the three colours together, because the palettes that look wrong vastly outnumber the ones that do not, and nobody should be inventing one in a form. The contract still accepts any combination for a caller that wants one | Wiring the three colour fields up as pickers and stopping there (leaves every light palette rendering as pastel mesh — the actual complaint); a free theme string resolved server-side (moves the palette into the API and makes adding a look a deploy); keeping one dark-shadow constant and telling people to avoid light looks (the shadow is a detail, not a constraint worth designing a product around); a caption pill instead of the shadow polarity fix (the right robust answer for captions over arbitrary footage, and still open — but it does not address the backdrop, which is what was asked for) | Accepted; the caption pill remains the outstanding piece for a silent screen-recording explainer |
| D-045 | 2026-09-10 | A camera move is a CSS keyframe whose animation-delay equals the segment's own data-start | A screen recording zoomed at the thing being described is the whole difference between a product walkthrough and a clip of a screen, so zoom_in / zoom_out / pan_left / pan_right plus a focus point are now in the brief. Two findings sit under that. First, the note in composition.ts claiming CSS animations do not advance between captures was wrong — the HyperFrames runtime does not let animations play, it SEEKS them, walking element.getAnimations(), setting currentTime and pausing. Motion in CSS renders exactly. Second, and the part that actually costs you a day: the runtime does not rebase that clock per element. currentTime comes from the composition's time, not the element's data-start, so a four-second move on a segment starting at 3.6s has already run to its end the moment the segment appears — every frame identical and fully zoomed, which reads as a broken zoom rather than as an error. The offset therefore has to live in animation-delay, set to the element's own start, with animation-fill-mode: both holding the opening pose until then. A carousel gets one rule per PICTURE, not per segment, because each picture has its own start. The presets are a closed set checked against a known list before interpolation, since a brief is JSON on a row and the value lands inside a <style> element, where esc() — which escapes for markup, not for CSS — would not save us. focus is CLAMPED rather than rejected: by the time the composition is built the voiceover is bought and paid for, and a rounding error must not cost the render | GSAP, which is HyperFrames' documented route and works — but it loads from a CDN, a network fetch inside the render worker, which this composition has deliberately stayed free of (the grain is an inline SVG for the same reason); vendoring GSAP's dist as a workspace asset (kills the fetch, but puts a disk read inside buildComposition, which is pure and testable precisely so it is neither); start/end transforms in the contract instead of named presets (the caller is choosing an intent — "get closer to this" — and the numbers expressing it are a renderer detail that becomes a compatibility problem the first time the composition changes); exposing a zoom scale knob as well (a fifth control for the thing that matters least — where it zooms beats how far); animating the mesh background too (a look, not a default; now possible, deliberately not taken) | Accepted; focus is in the contract but not yet in the UI — picking a point needs a preview of the frame to click, and x/y number boxes against a picture you cannot see is not a control |
| D-049 | 2026-09-15 | A project brings its own provider account, not its own SMTP settings — and the platform transport stays the fallback, not a substitute | The ask was per-project email sending; the shape it takes decides the security posture. Accepting raw SMTP host/port/credentials per project would mean storing an arbitrary outbound relay a tenant controls, with no way to tell a legitimate provider from an exfiltration endpoint. Accepting a Resend account is narrower in exactly the useful way: the credential is scoped to one provider we can validate against a known API, the domains it may send as are enumerable (GET /domains), and the client already owns the deliverability relationship. It is implemented over Resend's SMTP endpoint rather than their HTTPS API so the entire existing send path — attachments, headers, the text/plain-plus-html discipline — stays byte-identical and only the credentials change; switching to the HTTPS API later is a transport swap, not a rewrite. Three rules follow. The key is validated BEFORE it is stored, because a saved-but-broken credential looks configured and fails hours later to whoever pressed Send, with no visible link back to the settings screen. Transports are cached per project AND per key fingerprint, because nodemailer pools connections and a rotation that lives in a stale pool keeps transmitting through the credential someone was trying to revoke. And an undecryptable key raises rather than falling back to the platform transport: falling back would silently send a client's broadcast from the platform's account under a domain it cannot DKIM-sign — a deliverability-destroying substitution at precisely the moment nobody is watching | Per-project raw SMTP credentials (stores an arbitrary relay controlled by a tenant; unvalidatable; a generic exfiltration primitive); Resend's HTTPS API instead of SMTP (cleaner long-term and gives per-message ids, but forks the send path and re-implements attachment handling for no gain today); one platform Resend account with per-project API keys inside it (Resend keys are account-scoped, so this isolates nothing — the domains and suppression list are still shared); provider-agnostic BYO from the start (every provider's quirks become ours, and "our mail is not arriving" becomes unanswerable); falling back to the platform transport when a project's key fails (silent, and the failure mode is invisible until a client asks why their domain is not signing their mail) | Accepted; a second provider slots in behind the same resolveTransportForProject seam |
| D-048 | 2026-09-15 | A project owns the envelope, never the transport — and an unverified From address cannot send | Every "email integration" screen people have seen asks for an SMTP host and a password, so that is what they arrive expecting; the absence of those fields reads as a half-built feature unless the reason is stated. The reason is that SMTP credentials authorize sending as the ENTIRE deployment. Handing them to a project would let one tenant send as every other, and no amount of scoping around them fixes that — the credential itself is the authority. A sender identity is the opposite: it asserts an address but authorizes nothing, so it is safe to delegate to a project user, self-service, with no admin in the loop. That asymmetry is what makes the split work, and it is why the upsert schema is strictObject rather than merely ignoring unknown keys — a 201 that silently drops smtp_pass reads to anyone testing it as "the API accepted my credentials", and a client could grow such a form with nobody noticing. The second half is verification. An address costs nothing to type, including somebody else's, and mail from an unproved From address is what puts a deployment's IP on a blocklist — damage shared by every tenant on the transport, caused by the one project that benefited from skipping the check. So the gate is server-side at dispatch (assertSenderUsable, checked once before any recipient row is written, and applied to the test send too), the token is hashed, single-use and expires in 24 hours, and changing the address clears verification because proving hello@acme.com says nothing about billing@acme.com | Per-project SMTP credentials (hands every tenant the deployment's sending authority, and makes one project's misconfiguration everyone's outage); a platform-admin approval queue instead of email verification (a human bottleneck on a check a mailbox can perform itself, and it proves intent rather than control); verifying the domain by DNS instead of the mailbox (stronger and the right end state for a shared IP, but it needs DKIM/SPF per tenant and a provider that supports it — a much larger change than this asked for); trusting the address and rate-limiting instead (the reputation damage is done on the first send, and it lands on tenants who did nothing); requiring Reply-To to be verified too (it says where an answer lands, not who sent the message, so the bar would block a legitimate shared-inbox setup for no gain) | Accepted; DNS-level domain verification is the natural follow-up when the transport moves to a per-tenant provider |
| D-047 | 2026-09-15 | Migrations run as their own gated step, and the roll verifies itself and rolls back | migrate being a compose dependency of api/worker reads as the safe design — nothing starts until the schema is current — but it conflates two things: should the new release start and should the old one stop. Compose answers the second first, destroying the running api to recreate it, so a migration failure is an outage rather than a declined deploy. Verified, not reasoned about: on a local stack with the same dependency shape the api ends up a new container in created, whereas gating on docker compose run --rm migrate leaves it running with the same container id. Health is probed through Caddy rather than per-container because Caddy's /healthz is a static 200 and container health says nothing about the chain in front of it. | Leave it in the compose graph and read the logs after — the status quo; it is what made the P3009 incident an outage. Blue/green or two droplets — correct, and the real answer once this is more than one box, but it is a topology change to fix an ordering bug. Roll back the database too — rejected as a promise that cannot be kept: prisma migrate deploy has no down-migrations, and a rollback that silently half-works is worse than one with a documented boundary. Expand-then-contract migrations are the discipline that makes the image-only rollback sufficient. | Accepted |
| D-046 | 2026-09-14 | Project notification delivery is keyed on the event, not the recipient — and personal integrations stay | emitNotification is per-recipient by design: an approval request notifies three approvers, so it runs three times. Hanging the project's Slack and webhooks off it was the obvious cheap move and would have posted the same card to the channel three times, fired every project webhook three times, and made the duplicate count scale with the size of the team — which reads as flakiness rather than a bug and gets worked around instead of reported. So emitProjectEvent takes a caller-supplied eventKey stable for one real-world occurrence (platform_post:<id>:published) and claims it by INSERT against a unique index before sending. The insert races; exactly one caller wins. That covers the three shapes a duplicate actually arrives in — a per-recipient loop, two workers on the same job, and a retry after partial failure — without any of them needing to know about the others. The second half is that the per-USER Slack and webhook rows were kept rather than migrated: they answer a different question ("where do MY notifications go", following the person across every project) from the project rows ("where does THIS PROJECT announce itself", set by a lead and outliving them), and folding them into one table would mean either a project's channel dying when its creator left, or a person's private routing becoming visible to everyone on the project | Adding projectId to the existing user tables and letting one row mean both things (a nullable-owner column where exactly one of two FKs is set, and every query re-asserting which kind it wanted); de-duplicating inside emitNotification on a time window (works until two events legitimately land in the same second); firing project delivery only from a worker so there is one caller by construction (moves the problem rather than solving it, and delays the announcement); auto-migrating each user's personal hook into every project they are on (a user on six projects suddenly has six copies firing, and dedupe becomes their problem) | Accepted |
| D-045 | 2026-09-14 | A post carries exactly ONE content bucket, and the column is nullable forever | The bucket exists to answer "what share of our content is product?", and a share is only computable if every post contributes to exactly one denominator. Allowing two buckets per post — which is how a planner will describe a post that both teaches and sells — makes the shares sum past 100%, and "we are over-indexed on Product" stops being a sentence the report can support. So the field is single-valued, and the composer's helper text pushes the decision onto purpose ("why is this post being published?") rather than content, because a post that does two things was still commissioned for one. The second half is the nullability. Making it required would have meant back-filling every existing post, and the only available back-fill is a guess; once stored, a guessed classification is indistinguishable from a real one and it poisons the very report the field was added for. Nullable means old posts read as Unassigned, the mix report keeps them in its denominator so a project that has classified three posts out of fifty cannot look perfectly balanced, and ?content_bucket=unassigned is a first-class filter so the gap is findable and fixable rather than merely admitted | A multi-select of buckets with a designated "primary" (the primary is then the only field the report can use, so the others are decoration that still has to be maintained); a required field with a product_solutions default (defaults become the largest bucket in every report, measuring the default rather than the content); a required field back-filled by an LLM pass over existing captions (plausible, unauditable, and wrong in exactly the cases a planner would care about); free-text tags reusing Post.tags (drifts into "Education", "education" and "Edu" within a week, and the report becomes three rows of the same bucket) | Accepted |
| D-044 | 2026-09-02 | A channel tab may say "customised" and "you have read it", but never "approved" | PLATFORM_POST_STATUSES contains approved, changes_requested and rejected per variant, so marking each channel with its own verdict looks like a five-minute job. Nothing writes them. The only writers of platformPost.status in the API are the generic post update, the publisher and the scheduler — the approvals module writes none, because a decision is recorded against the POST and never touches a channel. Labelling a tab "approved" would therefore invent a per-channel verdict the system never made, and the reviewer acting on it would be acting on a claim nothing stands behind. variantPublishState returns null for every approval-shaped status for exactly this reason, with a test named after it, so the map is the one place to extend when the backend does start writing them. What the tabs CAN say is carried instead: a dot for a channel carrying its own caption or media (real, shared, derived from the payload), a tick for "you have opened this" (private to the viewer, localStorage, never styled as a status because it is a note to self), and a chip for where the variant actually got to (Scheduled / Published / Failed — written by the publisher and the scheduler, so true for everyone). Since one verdict covers every channel, Approve is held until every customised channel has been opened — the gate is the only honest answer to a per-post decision over per-channel content | Marking tabs with the variant's own status enum (invents a verdict; the field exists but is never written by the approval flow); recording "seen" server-side so the team shares it (a backend change, and this was scoped frontend-only); a warning instead of a gate on Approve (a warning about content you have not read is read by exactly the people who are not reading); making approval per-channel (the honest fix, but a schema and domain change — raised separately rather than faked in the UI) | Accepted; the per-channel verdict is a backend follow-up, and variantPublishState is where it lands |
| D-043 | 2026-09-02 | The approval review is a dialog over the queue AND a route, sharing one body | Reviewing a queue meant navigating in and back out per item, with the list's counts changing behind you — the round trip was the whole cost. But the route could not simply become a modal: four surfaces link into it (the approvals list, the calendar, the ad-proposals page, and the post's own analytics tab, whose back-link reads "Back to the post"), and a modal cannot be somewhere you navigate back TO. So the list opens a dialog, deep links open the page, and both render the same four components — the arrangement D-042 settled for bulk import, in the other direction. The parts that differ between hosts are passed as flags rather than forked: flush drops the card chrome where the panel is already the surface, fill lets the comments own the column height, layout=\"bar\" moves the decision into the dialog's pinned footer. One state machine either way — a separate footer bar would have been a second copy of the arming logic, and the copy is the one that forgets that a note is required, or that Approve is gated. The dialog is h-[85vh], a fixed height rather than a maximum, so the decision sits in the same place on a short post and a long one; Modal gained one optional bodyClassName so a dialog can take back the scrolling its body normally owns and give each column its own | Deleting the route and making the modal the only way in (breaks four inbound links and the analytics back-link, which names this route as the post's home); keeping the page and adding inline decisions on the list rows instead (cheaper, and still worth doing — but it answers triage, not "I need to read this properly"); a second modal-shaped component wrapping copies of the four panels (two implementations of the decision, the stepper and the thread, and the copy always falls behind) | Accepted; the page is the canonical URL and the compatibility surface, the dialog is the primary way in |
| D-042 | 2026-09-02 | Bulk import is a modal on the calendar, and its route survives as a second host of the same body | Importing a sheet acts on the calendar you are already looking at, which is the test for whether something should be a page. Sending someone to /calendar/import discarded the month they had scrolled to, and returned them to a route with no way to show what had landed; a modal keeps the grid behind it and lets the mutation's existing cache invalidation repaint it in place. The route is NOT deleted, because it was a sidebar child until D-034 removed it and project-ia.test.ts pins it resolvable for the bookmarks that outlived that change — but it renders the shared body rather than a copy, so two entry points cannot drift into two behaviours. The dry-run step survives the move, which is the part most at risk when a page becomes a modal: the temptation is a single "Upload & import" and a smaller dialog, and it was explicitly considered and rejected. The import is all-or-nothing server-side, and a half-filled calendar cannot be told apart from a full one by looking at it, so the preview is the only place a bad row is cheap to fix. Making the dialog smaller is not worth moving that discovery to after the write. State lives in useBulkImport rather than in the body component because the modal pins its commit button in the footer, outside the scrolling area — the button and the preview it acts on are then in different components and must read one state. The file input's ref is the one thing deliberately NOT in that hook: React 19's react-hooks/refs rule treats a returned object containing a ref as ref-like throughout, so every bulk.file read in render became an error. Reset bumps a nonce used as the input's key, which clears it by remount and, as a side effect, fixes re-picking the same file — the reason the old code assigned value = "" | Deleting the route and making the modal the only way in (breaks bookmarks and the test that exists to protect them, for one file saved); a one-step "Upload & import" modal matching the smaller dialog the brief sketched (drops the all-or-nothing gate, and moves every bad row from a preview to a rejected write); leaving the button as a link and adding a "back to calendar" affordance on the page (still loses the month, still cannot show the result in context); keeping the state inside the body and rendering the commit button inline in the modal too (an action that scrolls away with the preview table above it is an action people stop finding) | Accepted; the modal is the primary entry, the route is a compatibility surface and should be removed once bookmarks are no longer a concern |
| D-041 | 2026-09-02 | A channel tab is something you have; connecting is an action, and they stop sharing a row | The Content page listed every organic platform as a tab, most of them empty and offering + Connect. Three separate problems in one control. A filter bar's job is to describe the data underneath it, and seven-eighths of this one described data that did not exist — every one of those tabs filtered the grid to nothing, so the page's most prominent row was mostly buttons that produce an empty screen. Connecting is also not the same kind of act as filtering: every other control in the row narrows the grid in place, while + Connect left the application entirely for an OAuth round trip and came back on a different URL. And a project's channel list is small and stable, so a row sized to the full platform catalogue is permanently mostly wrong. Tabs are now connectedPlatforms(accounts), and one dashed button at the end opens a Manage Connected Channels modal. The dashed border is doing real work: it is the row's only non-tab, and it must not read as an unselected filter. The membership test is has an account, NOT has a live token — D-026 keeps a disconnected account's row precisely so its published posts and their metric history survive, and hiding the tab would hide that back catalogue from this page the instant a token expired. The predicate is shared with the page rather than written twice, because the page derives its selected channel from the same list: two copies that drift leave a tab selected that is no longer rendered, or a grid filtered to a channel with no tab. That selection is now also re-derived when its channel disappears, which the manager made reachable without leaving the page. The modal renders the Organic hub's real OrganicHubTiles, not a second grid — a copy would be two implementations of connect, disconnect, the account popover and the status badges, and the copy is always the one that falls behind. It is addressed by ?channels=1 rather than local state so Back closes it and the OAuth callback can be pointed back into it. Inside the modal the tiles navigate nowhere (linkToHub={false}): no "Open hub" footer, no stretched link on the title, and the account rows become static text rather than links into the platform hub. A modal opened to disconnect an account should not be able to walk you out of itself, and the per-platform hubs are on their way out — the sidebar's Channels section and its sub-tabs are going, so a card whose main affordance points at one is building on a floor that is being removed. The four affordances that make a tile read as clickable come off together with the destination, because a card that lifts on hover and shows a pointer while doing nothing is worse than one that never promised | Keeping every platform as a tab and greying the empty ones (the row is still mostly inert, and grey-but-clickable is worse than absent — it still filters to nothing); an "All channels" tab to absorb the empty ones (rejected: see D-038, and it was explicitly not wanted); putting + Connect channel inside the tablist as a final tab (it is not a filter, and role="tab" would put it in the same arrow-key group as the channels while doing something entirely different on activation); a bespoke compact channel-picker in the modal instead of the hub grid (a second vocabulary for connection state, and the disconnect/popover behaviour re-implemented); local useState for the modal (Back would not close it, and the OAuth return would land on a closed modal with the params stranded in the URL) | Accepted; supersedes nothing — D-038's one-channel rule is unchanged, and the "no all tab" consequence it records still stands |
| D-040 | 2026-09-02 | Per-channel content is written by creating the post, then patching each variant — and there is still only one uploader | postCreateSchema carries a single caption, media list and link that seed EVERY child variant; the contract states the rule itself — default_platform_extra is "extras to seed on every child PlatformPost" and each variant is "individually overridden later via PATCH". So per-platform content cannot ride along with the create, and the composer creates first and applies overrides after, sequentially, naming the channel in any failure: a partial apply leaves a real post with some channels customised and some on the shared copy, and "it didn’t work" would send someone hunting through all of them. The uploader is parameterised, not duplicated — it writes to whichever bucket is open rather than existing once per tab. An instance per tab means several drop targets and several in-flight upload counters against several buckets, and a real unanswered question about what happens to the one you switch away from mid-transfer; one uploader is one queue and one answer, and the media still lands per channel because platform_media is per-variant. Accepted cost, written at the call site: the media cap is still the most permissive across every selected platform rather than the open tab’s, because maxMedia is derived well above the tab state it would need to read | Extending postCreateSchema to accept per-variant content (a contracts change, and the PATCH route already exists for exactly this); rendering an uploader per tab (the upload-in-flight question above, and three copies of a drop target, a queue and a cap); keeping media shared and only splitting the words (the first thing anyone wants to differ per channel is the image); one PATCH batched for all variants (no such endpoint, and a batch would lose which channel failed) | Accepted; per-tab media cap is a known follow-up |
| D-039 | 2026-09-02 | An item’s own description does not change with how you arrived at it | Raised when the content-type pill was added: with a Type filter active, every card repeats the same word, so should it hide? The rule that settled it — a card describes the post, not the route. But the reasoning that made "dim it" look right was inherited from an earlier layout and did not survive scrutiny: once the badge moved to absolute on the media well, removing it reflows nothing, and because the type filter is page state rather than a URL param, nobody can arrive with it already set — you filtered seconds ago. So the grid badge is dropped outright while filtered, and the list pill stays, because that one is in normal flow and its row has no media well saying "video" for it. Hiding via opacity-0 was rejected outright: it leaves the badge in the accessibility tree, announced on every card and visible on none. Where "you are filtered" belongs is the filter control, which carries the active tint — the same reason the Content search field stays open while a query is live | Muting the pill everywhere (keeps a redundancy that competes for attention, and rests on a reflow argument that is void for an out-of-flow element); visibility: hidden to preserve layout (nothing to preserve — the badge is absolutely positioned, so this is conditional rendering with extra steps, worth it only for a fade); removing it from the list row too (in flow, so it shifts the meta line, and the row has no thumbnail doing the explaining) | Accepted |
| D-038 | 2026-09-01 | The Content page shows exactly one channel, and says what that costs | A post's shape, limits and audience are all per-platform, so an "all channels" list is a list of things that cannot be compared — the same draft appears once per channel it targets, or once with a row of icons that answers nothing. One channel at a time makes every row comparable and lets the grid card carry platform-correct media. The default is the first platform with a connected account rather than a fixed one, so a project that only runs LinkedIn does not open on an empty Instagram tab. The cost is real and is written next to the component rather than discovered later: a post with no channel attached matches no tab, and most drafts are exactly that, because the composer deliberately lets you write first and pick where it goes later. Those posts stay reachable from the calendar and the composer. Accepting an unreachable-here case is the honest trade; the alternative hides it behind a tab that cannot render a post correctly | An "All" tab listing every post with channel icons (cannot render platform-correct media or limits, and the icons restate what the channel bar already says); showing channel-less drafts under every tab (the same draft in eight places, and a publish button on a tab it does not target); a separate "Unassigned" tab (a bin whose contents are defined by absence — it grows, nobody clears it, and it is not where anyone looks for a draft they just wrote) | Accepted; the channel-less gap is a known follow-up, not a defect |
| D-037 | 2026-09-01 | What a post card offers belongs to the post's status, not to the screen showing it | The Instagram hub and the Content page render the same object — a post at some point in its life — and had drifted into two vocabularies: "Edit" against "Continue editing", "Insights" against "View post", and a Content card with no publish, reschedule or review action at all. Two names for one action is a bug the reader has to resolve at every glance, and syncing two hand-written lists is the kind of maintenance nobody does twice. lib/post-actions.ts returns one ordered list per status, and the ordering lives there too: which action leads is a property of the status, not of the page. Capabilities are passed in, not derived, because they genuinely differ by surface — a card owning a single platform variant can publish it, while a Content row for a post with three channels, or none, cannot answer "publish what?". It offers nothing rather than a button that fails: absent-because-impossible is consistent, present-but-broken is not | Each card keeping its own list with a lint rule or test asserting they match (machinery to maintain a duplication we did not want); deriving capabilities inside the helper from the post alone (it cannot know whether the surface is scoped to one platform, so it would either over-offer publish or never offer it); one shared card component for both surfaces (the two grids differ in layout, density and what else they show — sharing the list is the part that was actually duplicated) | Accepted |
| D-036 | 2026-09-01 | On the calendar, the dataset is navigation; the period and the presentation are page state | The old switcher put month, week, content and ads in one flat strip, so the four were mutually exclusive when only two of them ever were. Choosing "Content" discarded the month you were looking at, and "this month, as a list" was unaskable. Splitting them by what kind of thing each is: scheduled posts and ad campaigns are different tables with different shapes, and moving between them is genuinely navigation — it belongs in the sidebar and in the URL, where it can be linked and bookmarked. A month or a week is a zoom level on whichever dataset is open, and grid-or-list is how that is drawn; putting either in the URL made them look like destinations. monthly and weekly therefore retire to content rather than 404ing, since that is what those URLs were already showing. This restores two sidebar children D-034 removed and does not contradict it: D-034 deleted children that RESTATED the page's own tab strip, and once the period and the layout stay in the toolbar, the dataset is no longer in that strip to restate. Bulk import moves the other way — out of the sidebar into the page header — because it is an action on the content calendar, and listed between two datasets it read as a third one. Its route is untouched | Keeping one flat switcher and adding a separate list toggle (the month-vs-content exclusivity, the actual bug, survives); putting all three axes in the URL as calendar/content/month/list (three segments of which two are view state, and every control writing to the router); leaving the period in the URL and moving only the dataset (monthly and weekly stay linkable as though they were places, which is the misreading that caused this) | Accepted |
| D-032 | 2026-08-31 | CAPI events post to the PIXEL node, not to act_{account}/events as specified | The brief for this work named https://graph.facebook.com/v20.0/act_{externalAccountId}/events. That endpoint does not exist. Meta's Conversions API is scoped to a pixel — POST /{PIXEL_ID}/events — and an act_ node answers "Unsupported post request… path /events does not exist". Building it as written would have shipped a forwarder that failed on every single call and did so silently, because this path swallows its errors by design so it cannot slow the tracker. The ad account is still where the credentials and the tenant scoping come from; only the URL is the pixel's. Flagged rather than followed, because the failure mode of following it is invisible until somebody asks why Meta's reporting shows nothing | Implementing it literally and letting it fail (silently wrong, and the swallowed-errors design means nobody finds out); implementing it literally and making errors loud (turns a spec bug into checkout-page noise); asking before building (the correct endpoint is unambiguous and documented, so the answer was not in doubt — the deviation is recorded here instead) | Accepted |
| D-031 | 2026-08-31 | The CAPI destination pixel is read from AdGroup.targeting.pixel_id, not from a new setting | The forwarder needs to know which pixel should receive a project's conversions. The obvious build is a "CAPI pixel" field on the ad account or the project. The ad sets already carry one — it is what the ad-set builder writes when a conversion goal is picked — and the two can never legitimately differ: the pixel a campaign optimises against is by definition the pixel whose events it needs to see. Reading it from there means the pairing is correct by construction rather than by somebody remembering to keep two fields in step; a separate setting starts correct and silently stops being the first time an ad set is repointed. Deduplicated across ad sets, because several routinely share a pixel and sending the conversion once per ad set is triple-counting in the client's reporting, not redundancy | A capi_pixel_id column on AdAccount (a second source of truth for one fact, and the one nobody updates); providerMeta.pixel_id (same problem, less discoverable); forwarding to every pixel on the account via the new getPixels call (an outbound Graph call per conversion just to decide where to send it, and it would forward to pixels no campaign here uses) | Accepted |
| D-033 | 2026-08-31 | CAPI sends a hashed external_id and no raw client IP, accepting worse match rates | Meta accepts client_ip_address raw and it materially improves event matching, and we hold the raw IP for the duration of the tracker request. Sending it would still be wrong here: MarketingTouchpoint deliberately stores ipHash and never an address, with a comment saying so, and a forwarder that quietly shipped the raw one to a third party would make that stated posture a fiction — the data would leave the system through the side door the privacy design was built to close. external_id is the field Meta defines for exactly this: a stable, first-party, hashed identifier, and the visitor id is already that. Worse matching is the honest cost of the position the codebase already took | Sending client_ip_address raw (better matching, contradicts the storage posture silently); sending a hashed IP in client_ip_address (Meta does not hash that field — it would simply never match, which is worse than omitting it because it looks like it is working); adding a per-project consent switch (a real feature, but a bigger one than this slice, and defaulting it either way is the same decision made less visibly) | Accepted |
| D-030 | 2026-08-31 | The email sanitiser lives in @verjson/contracts and is hand-written, not a library | It has to run in three places that must agree exactly: the browser preview (before anything has been saved), the API on write, and the renderer at dispatch. Three sanitisers means the preview shows one email and the recipient receives another, and the only person who would ever notice is the recipient. The contracts package is the one thing both apps import, and it is the same reasoning that put parseGoogleSheetUrl there for F-111.2. Hand-written because the package must stay dependency-free and DOM-free — it runs in Node, in the browser and in Vitest's node environment — and because email HTML is a narrow enough subset (an allow-list of ~45 tags, ~15 attributes, 4 URL schemes, ~40 CSS properties) that a tokeniser is tractable and, more importantly, reviewable. 67 tests hold the allow-list property rather than enumerating today's attacks | DOMPurify (needs a DOM — jsdom on the server, and a transitive dependency on the wire contract both apps then have to agree on); sanitise-html (Node-only, so the browser preview would need a second implementation); sanitising on read instead of write (the one read path somebody forgets is the one that mails the payload); trusting the composer's output because it is our own editor (a paste, a PATCH from a script, or a restored backup all bypass the composer) | Accepted |
| D-029 | 2026-08-31 | contenteditable + execCommand for the composer, not TipTap/ProseMirror/Slate | Every capability the toolbar needs — bold, italic, underline, strikethrough, colour, highlight, alignment, both list kinds, indent/outdent, blockquote, undo, redo, remove formatting, links, images — is a single execCommand that every browser implements. It is marked deprecated, has been for years, is what the web's mail composers still run on, has no removal plan, and nothing replaced it (the Editing API meant to succeed it never shipped). The alternative is a document model, a schema, a serialiser and a paste pipeline that would then have to be taught to emit exactly the subset of HTML sanitizeEmailHtml permits — otherwise it emits markup the server silently strips, which is the worst failure mode available here because it looks like it worked. Two places the built-in behaviour is not good enough are overridden explicitly: font size (the legacy <font size=1-7> is not in the allow-list, so it is converted to an inline font-size span in the same tick) and paste (run through the shared sanitiser, so what you see after pasting is what will be stored) | TipTap (~15 packages, a ProseMirror schema to keep in step with the allow-list by hand, SSR configuration in Next); Slate (same, plus a less stable API); a Markdown editor (people expect to paste formatted text into an email composer, and Markdown discards it); accepting execCommand's <font size> output and widening the sanitiser to match (deprecated presentational markup that several mail clients render inconsistently) | Accepted |
| D-028 | 2026-08-31 | The signature is inserted into the body by the composer, not appended at dispatch | Appending at send is less code and would guarantee every broadcast carried the sign-off. It would also make the preview lie. The preview's whole job is to answer "what will they get", and a signature the sender cannot see in the editor, cannot see in the preview, and cannot edit is a paragraph added to their email by something they never saw. Making the preview show it instead would mean two pieces of code assembling the same message — precisely what D-030 exists to prevent. So signatureBlock() (shared, in contracts) emits a marked block the composer appends to body_html, and stripSignatureBlock() removes exactly that block when the toggle goes off. The data-email-signature marker is not in the sanitiser's allow-list, so it never reaches a recipient; it exists only while the body is being edited, which is exactly when it is needed | Appending at dispatch (preview lies); appending at dispatch AND in the preview (two assemblers of one message); a per-broadcast include_signature column (a third state to keep in step with a body somebody can also edit by hand); a per-user signature (broadcasts leave from the deployment's configured from address, so a personal sign-off would sit under mail that visibly came from somewhere else — and an agency running six clients out of one org needs six sign-offs, which is the project, not the person) | Accepted |
| D-035 | 2026-08-31 | Project identity lives in a sidebar switcher, not a breadcrumb band | The breadcrumb spent a full horizontal strip restating what the sidebar already showed — the active section is marked there — and its only irreplaceable element was the link back to /projects, which survives as a pinned row inside the new dropdown. What it could never do is move between projects: from inside project A, reaching project B meant navigating out to the list and back in, which is the single most common cross-project action in a multi-client tool. Putting the name, the status and every sibling project in one control makes that a click, and gives the vertical space back to the content. It is a disclosure, not a role="menu": the items are ordinary links, so native tab order and Enter-to-follow already do the right thing, and a menu role would advertise arrow-key roving this doesn't implement — a promise a screen reader holds you to. The status badge is aria-hidden because the trigger's accessible name already carries the status; without that it is announced twice. Status→variant is a local map, not the shared statusVariant() helper, which maps order/payment vocabulary (paid, refunded, no_show) and has no case for active or paused — routed through it, three of the four project statuses collapse to the neutral default and an Active project renders identically to an archived one | Keeping the breadcrumb and adding a switcher beside it (two controls naming the same project, and the vertical strip stays); a switcher in the global top nav (that nav is org-level — Projects · Billing · Team · Settings — and a project control there would be lit on pages that have no project); leaving navigation alone and only reclaiming the space (loses the cross-project jump, which was the actual complaint) | Accepted; not yet covered by a Playwright spec — the truncation and clipping behaviour is manual-verified only (see VERIFICATION.md 2026-08-31) |
| D-034 | 2026-08-31 | Sidebar sub-nav that duplicates a page's own tab bar is deleted, not kept in sync | Content, Calendar and Approvals each listed children that resolved to a route the page already reaches through its own tab strip — several of them to the same URL with only a query string between them, and the Content status filters landed on a view that reads neither status nor unscheduled posts. A second control for the same destination is not redundancy that costs nothing: it is two places to change when a tab is added, and the sidebar is the copy nobody remembers. Only children pointing somewhere the tab strip cannot reach are kept — Media Library (/measure/media) and Bulk import (a static route with its own upload + preview state). The load-bearing fact checked before touching any of it: the tab bars are not derived from these children. Tab ids live in lib/section-tabs.ts, and iaLeafPaths() — the only function that turns children into paths — has no non-test caller. Had either been otherwise, deleting the children would have emptied the calendar's view switcher or 404'd its URLs. Two tests now pin that independence in both directions, specifically because the monthly/weekly slug mappings are left unreferenced by the IA and read as dead code to the next person tidying up | Keeping the children and adding a lint or test that asserts they match the tab bars (machinery to maintain a duplication we did not want); collapsing the tab bars into the sidebar instead (the tabs are the better home — they sit on the page they filter, and the sidebar would then be the only way to filter, which is worse on mobile where it stacks above the content); leaving it alone (the sidebar keeps growing, and this pattern is still present in Campaigns, the organic hubs and the performance hubs — untouched here, and the obvious follow-up) | Accepted |
| D-029 | 2026-08-31 | Project identity lives in a sidebar switcher, not a breadcrumb band | The breadcrumb spent a full horizontal strip restating what the sidebar already showed — the active section is marked there — and its only irreplaceable element was the link back to /projects, which survives as a pinned row inside the new dropdown. What it could never do is move between projects: from inside project A, reaching project B meant navigating out to the list and back in, which is the single most common cross-project action in a multi-client tool. Putting the name, the status and every sibling project in one control makes that a click, and gives the vertical space back to the content. It is a disclosure, not a role="menu": the items are ordinary links, so native tab order and Enter-to-follow already do the right thing, and a menu role would advertise arrow-key roving this doesn't implement — a promise a screen reader holds you to. The status badge is aria-hidden because the trigger's accessible name already carries the status; without that it is announced twice. Status→variant is a local map, not the shared statusVariant() helper, which maps order/payment vocabulary (paid, refunded, no_show) and has no case for active or paused — routed through it, three of the four project statuses collapse to the neutral default and an Active project renders identically to an archived one | Keeping the breadcrumb and adding a switcher beside it (two controls naming the same project, and the vertical strip stays); a switcher in the global top nav (that nav is org-level — Projects · Billing · Team · Settings — and a project control there would be lit on pages that have no project); leaving navigation alone and only reclaiming the space (loses the cross-project jump, which was the actual complaint) | Accepted; not yet covered by a Playwright spec — the truncation and clipping behaviour is manual-verified only (see VERIFICATION.md 2026-08-31) |
| D-028 | 2026-08-31 | Sidebar sub-nav that duplicates a page's own tab bar is deleted, not kept in sync | Content, Calendar and Approvals each listed children that resolved to a route the page already reaches through its own tab strip — several of them to the same URL with only a query string between them, and the Content status filters landed on a view that reads neither status nor unscheduled posts. A second control for the same destination is not redundancy that costs nothing: it is two places to change when a tab is added, and the sidebar is the copy nobody remembers. Only children pointing somewhere the tab strip cannot reach are kept — Media Library (/measure/media) and Bulk import (a static route with its own upload + preview state). The load-bearing fact checked before touching any of it: the tab bars are not derived from these children. Tab ids live in lib/section-tabs.ts, and iaLeafPaths() — the only function that turns children into paths — has no non-test caller. Had either been otherwise, deleting the children would have emptied the calendar's view switcher or 404'd its URLs. Two tests now pin that independence in both directions, specifically because the monthly/weekly slug mappings are left unreferenced by the IA and read as dead code to the next person tidying up | Keeping the children and adding a lint or test that asserts they match the tab bars (machinery to maintain a duplication we did not want); collapsing the tab bars into the sidebar instead (the tabs are the better home — they sit on the page they filter, and the sidebar would then be the only way to filter, which is worse on mobile where it stacks above the content); leaving it alone (the sidebar keeps growing, and this pattern is still present in Campaigns, the organic hubs and the performance hubs — untouched here, and the obvious follow-up) | Accepted |
| D-027 | 2026-08-13 | PR #124 squash-merged to dev without CI validation, sole-operator | Verjson's GitHub Actions billing has been suspended since 2026-08-11 (every workflow run dies in 3-7 seconds with "The job was not started because recent account payments have failed"), so ci.yml on this PR never actually ran. The alternative — waiting for the billing fix — would leave three important commits stranded on the deploy branch: the deploy path itself (Pulumi + docker-compose + Caddy + workflow), the cloud-init em-dash fix (without which pulumi up on a fresh droplet silently produces a bare Ubuntu box), and the tsup noExternal: [/^@verjson\//] fix (without which node dist/main.js crashes with ERR_MODULE_NOT_FOUND on @verjson/contracts). The validation for these changes is the manual first deploy we ran end-to-end on 2026-08-13: /healthz, /api/v1/health, and / all served 200 from http://134.199.250.245.nip.io. Landing on dev rather than main because main-protection (ruleset 18098028) enforces require_last_push_approval: true with no bypass_actors, and there is no second reviewer available; dev has no protection and no ruleset condition matches it. main remains a strict-ruleset branch and this decision does not weaken that; the cascade dev → main will go through the ordinary flow once a second reviewer or the billing fix arrives | Waiting for GitHub Actions billing to be restored (unknown ETA, leaves three known-good fixes off dev); PATCHing main-protection to drop require_last_push_approval, merging to main, restoring (weakens repo governance for a solo-operator escape hatch, exactly what the rule exists to prevent); asking a second Verjson member to approve (not reachable in the deploy window); building a self-hosted-runner fallback in the workflows (uses the gha-general-* DO fleet — a real long-term move, but doesn't help right now because GitHub still queues on billing state) | Accepted, until GitHub Actions billing lifts and dev's subsequent commits get real CI coverage; rollback: git revert the squash commit on dev — the prod droplet is unaffected since it's pinned to image tags api:e43f06e... and web:645e8e8..., not to a branch |
| D-026 | 2026-08-12 | Disconnecting a social account is a state change, not a row deletion | Every published PlatformPost hangs off SocialAccount with onDelete: Cascade, and every PostMetricSnapshot / PostMetricDaily hangs off those — so the old hard delete erased the entire reporting history for work that really did go out. That history is the thing a user is least able to reconstruct and most likely to still want: you stop posting to a channel long before you stop caring how it performed. Disconnect now drops the credentials, sets connection_status=disconnected + disconnected_at, and keeps everything else. The status is the one a worker already sets when a token dies, so every consumer that skipped dead accounts skips these too — no new state to teach the system. The composite unique on (project, platform, account_platform_id) means a reconnect revives the same row and the history reattaches by itself. Permanent deletion survives as an explicit ?purge=true, for a mistakenly-connected account or an erasure request | Keeping the cascade and telling users to export first (puts the burden on the person who least expects the loss, and only works if they read the confirm); copying metrics onto a detached archive table before deleting (a second schema of the same numbers, and every analytics query needs a UNION forever); soft-delete via a generic deleted_at filter on every query (one missed filter leaks a deleted account back into a publish path) | Accepted |
| D-024 | 2026-07-29 | The UI adopts the Verjson parent brand rather than a palette of its own | Tokens, type and devices are taken from verjson.com / verjson.ai: cream #faf8f3 on ink #0a0a0a, indigo #1700b8 for action, acid lime #e8ff47 as the signature, Inter + JetBrains Mono, and the _section underscore kicker. This is an internal Verjson tool used by Verjson staff and the agencies they hire; a house style that matches nothing else they use is a cost with no upside, and "generic SaaS blue + a serif" was the actual complaint. Lime is deliberately not --accent — at 88% luminance it fails contrast on cream at any size, so it lives in --highlight and is only ever a fill with near-black on top, which is how the brand uses it (the logo inverts to lime on dark) | Keeping the blue/green default (looked like every component library); a bespoke identity for the studio (a second brand to maintain, and no reason for one) | Accepted |
| D-025 | 2026-07-29 | Photography is downloaded into public/img/, not hotlinked from Unsplash | Serving it ourselves means the page has no runtime dependency on a third-party CDN, no viewer IP is disclosed to one, and the images cannot change or vanish underneath us. The cost is ~2.3 MB in the repo and a CREDITS.json to keep provenance. Images are used only where the product genuinely has images — the media library (F-033), the ad asset pipeline (F-055), the composer's creative — plus two atmospheric bands; decorative photography is alt="" so a screen reader is not read a stock photo | Hotlinking images.unsplash.com (a CSP hole, an external dependency and a tracking vector for a page that has no other one); illustration instead (nothing to illustrate — a marketing studio's subject matter is photographs) | Accepted |
| D-023 | 2026-07-29 | OpenSEO for off-page SEO and AI-visibility data; our own audit stays for on-page | MIT, self-hostable, MCP-native, and priced per DataForSEO call rather than as a subscription — so it replaces the planned SEMrush integration on both cost and lock-in. The larger reason is its ai-search feature (shareOfVoice, citedSources, promptExplorer), which implements the AEO measurement LAB-008 was open on: a prompt panel run against assistants, measuring who gets cited. We were about to approximate that with heuristics, and approximating it badly is worse than not measuring it. The split is on-page vs off-page — our audit stays because it costs nothing per run and emits stable rule ids that map to fixes we open as PRs, which an external tool cannot do without knowing our projects, gate and repo | SEMrush subscription (recurring cost, no self-host, no AI-visibility); building all of it (months, and the underlying data still has to be bought) | Accepted |
| D-022 | 2026-07-29 | The approval gate is a state machine, not a status column | Non-negotiable #3 is only as real as the rule that a run cannot reach applied without passing through awaiting_approval. A status column any writer can set is a state machine in name only — so the legal transitions live in one dependency-free module, transition() is the sole writer, and the test asserts the negative property (no status other than awaiting_approval can reach applied) rather than an example of it, because an example only proves the paths someone thought of. The write is a filtered updateMany, so two simultaneous approvals cannot both apply the same spend | A boolean approved flag (no history, no concept of an illegal transition); enforcing in the route handlers (one forgotten call away from an auto-apply) | Accepted |
| D-021 | 2026-07-29 | Webhook idempotency is claimed by INSERT before the work, not checked after it | A check-then-insert races: two concurrent deliveries of the same event both find no row, both apply, and then the loser's insert violates the unique constraint, 500s, and makes Stripe retry an event that already succeeded. Inserting first makes the primary key the lock — the second delivery fails to claim it and returns quietly. A P2002 is success, not an error | Advisory lock (works, another concept to hold); a transaction around check+apply (serialises every webhook, and the apply may call out to Stripe) | Accepted |
| D-020 | 2026-07-29 | The subscription's plan is derived from its Stripe price, never from metadata.plan | Metadata is written once at checkout and is never updated when a customer changes plan in the Stripe billing portal, so an organization that downgrades scale → starter would keep scale limits indefinitely. The price is the thing that actually changes. An unrecognised price degrades to free and logs loudly, because the alternative — defaulting to a paid plan — silently grants entitlements to a subscription we cannot account for | Trusting metadata and updating it on portal events (needs a webhook for every path finance can take); storing the price id and mapping in the UI (moves the same problem to the client) | Accepted |
| D-019 | 2026-07-29 | Role checks and the project scope filter are allow-lists, never deny-lists | The scope filter previously read "if agency or client, restrict to memberships", which fails open: any role it did not recognise — a typo, a role added later, a hand-inserted row — got org-wide read of every project, its budgets and its team. Enumerating the trusted set means an unknown role is restricted by default. User.role is now a Prisma enum as well, so the column cannot hold a value no check anticipates | Keep the deny-list and rely on the enum alone (the enum is a good second layer, but the failure direction of the code should be safe on its own) | Accepted |
| D-018 | 2026-07-29 | OAuth state is bound to the browser with a nonce cookie — signing alone is not enough | /start is unauthenticated, so an attacker can mint a genuinely-signed state whenever they like. They can complete consent with their own account, capture code+state, and get a victim to open the callback — the victim's browser then stores the attacker's session and starts connecting real ad accounts to the attacker's organization. Via /link-url it is worse: a state carrying linkUserId = attacker makes the victim's consent attach their provider identity and repo-scoped token to the attacker's account. So /start also sets an HttpOnly, SameSite=Lax, __Host- cookie holding a random nonce, the state carries only its hash, and the callback requires both and compares in constant time. Also: an explicit link against a provider account already bound to someone else is a 409, never a silent login as that person | Signed state alone (what we had — prevents nothing, since the attacker can mint one); PKCE (protects the code exchange, not session fixation); server-side session state (needs shared storage, defeats statelessness) | Accepted, supersedes D-012 |
| D-017 | 2026-07-29 | Agent-Reach as the agents' read layer for external platforms | It already covers ~14 platforms (X, Reddit, LinkedIn, YouTube, GitHub, RSS, podcasts, web search) by routing to native CLI tools with ordered fallbacks, and installs as an MCP server — which is the path our agents already speak. Adopting it means not writing and maintaining a scraper for every platform whose access path breaks quarterly. Strictly read-only: publishing stays with our own adapters behind the approval gate | Our own scraper per platform (permanent maintenance against hostile targets); a paid unified API (recurring cost, and still a wrapper we do not control) | Accepted |
| D-016 | 2026-07-29 | trigger.dev for agentic workflows, alongside RabbitMQ for jobs | An agent run is long, multi-step, and has to survive a deploy and a wait-for-human approval. That is durable-workflow work: resumability, retry-from-step, and state that outlives a process. Putting it in a queue means either a message with a multi-day TTL or a hand-rolled state machine in the payload. Conversely, per-post publish jobs in a workflow engine would pay orchestration overhead thousands of times an hour — so both exist, for different shapes of work | Everything in RabbitMQ (hand-rolled saga state); everything in trigger.dev (overhead on high-frequency jobs); Temporal (more capable, materially more to operate for our size) | Accepted |
| D-015 | 2026-07-29 | Baileys in its own container for WhatsApp, behind the shared channel interface | Baileys speaks the WhatsApp Web protocol from a persistent socket, so it needs its own lifecycle and a volume for the paired session — it cannot live inside stateless API pods. It is an unofficial client: bulk marketing over it breaches WhatsApp's terms and the realistic downside is the number being banned, not throttled. It is accepted for an internal tool and conversational replies, and the mitigation is structural: the bridge implements the same ChannelProvider interface as every other channel, so moving to the official Cloud API is a transport swap rather than a rewrite | Official WhatsApp Business Cloud API only (sanctioned, but needs Meta business verification and per-template approval before anything can be sent); a third-party BSP (removes the ban risk, adds per-message cost and a vendor) | Accepted |
| D-014 | 2026-07-29 | RabbitMQ for the work queue — superseding Postgres-as-queue | Nearly every consumer talks to a rate-limited third-party API, and per-queue prefetch, fair dispatch and native dead-lettering are exactly what that needs. A polling table has to reimplement all three, and the version that gets written under deadline handles the dead-letter case worst. The rate limiter stays on Postgres — a fixed-window counter is not a queue | Keeping Postgres as the queue (fewer services, but reimplementing broker primitives); Redis + BullMQ (good, but a second stateful service either way, and weaker routing) | Accepted, supersedes D-005 |
| D-013 | 2026-07-29 | Bundle the API with tsup rather than emitting with tsc | The codebase uses @/* path aliases and tsc does not rewrite them on emit — the compiled output carried literal @/core/db specifiers that Node cannot resolve, so npm run build was producing an unrunnable dist. tsup (esbuild) resolves them at build time, so aliases stay a development convenience with no runtime cost. Prisma stays external: its generated client ships platform-specific engine binaries | tsc-alias as a post-step (another tool, same outcome, slower); relative imports everywhere (churn, and worse to read) | Accepted |
| 2026-07-29 | Claimed that signing alone prevented session fixation. It does not: /start is unauthenticated, so an attacker can mint a valid signed state at will. The verified-email half of this decision stands and is restated in D-018 | — | Superseded by D-018 | ||
| D-011 | 2026-07-29 | Money is an integer count of minor units, everywhere — API, database, and the contract package | 0.1 + 0.2 !== 0.3, and a budget is not the place to discover that. Storing cents as Int means arithmetic is exact and the roll-up is reproducible. The formatter lives in @verjson/contracts so the API and the UI cannot disagree by a factor of 100 | Decimal in Postgres (exact, but JS has no native decimal so it becomes a string everyone re-parses); floats (wrong, quietly) | Accepted |
| D-010 | 2026-07-29 | The marketing mix is a database enum, and a budget may only exist for a declared channel | Budgets, campaigns and analytics all key off the channel. Free text would let a typo create a phantom channel with its own budget line that silently inflates the total. Requiring the channel to be declared first makes the mix a decision rather than an accident of data entry — and dropping a channel that still holds money is a 409, not a silent orphan | Free-text channel names (typos become data); a lookup table (flexible, but every consumer then needs a join and a fallback for an unknown row) | Accepted |
| D-009 | 2026-07-29 | Project membership is what scopes an agency user — enforced as a where clause in the data-access layer, not a check in the route | An agency contractor belongs to our organization but must only see the projects they were invited to. Building the scope into the query means an out-of-scope project is indistinguishable from one that does not exist: the caller gets 404, never 403, and cannot probe for real ids. There is deliberately no unscoped findById, so a scoping bug cannot be one forgotten clause away | A role check in the route (one forgotten call from a leak); Postgres row-level security (stronger, but needs a per-request session variable and hides the rule from the code that depends on it — revisit at scale) | Accepted |
| D-008 | 2026-07-29 | The dashboard reads JSON from data/dashboard/, not the database | The board must be accurate on day one, before there is a schema to hold it, and it must be diff-able and editable by an agent in a PR. A DB-backed board would need auth, migrations and an admin UI before it showed anything. | DB tables + admin CRUD (too much scaffolding for a build board); a static markdown page (not queryable, no computed "coming up") | Accepted |
| D-007 | 2026-07-29 | Docs are rendered inside the dashboard, sourced from docs/*.md at request time | A doc nobody opens is a doc nobody updates. Putting them one tab away from the feature checklist means the board and the prose stay in the same field of view. Files stay the source of truth, so git history and PR review still work. | A separate docs site (another deploy); leaving docs in the repo only (they rot) | Accepted |
| D-006 | 2026-07-29 | Groq for text generation, ElevenLabs for voice | Both keys are already provisioned. Groq's latency makes interactive drafting feel live rather than batch. Provider access is confined to the infra layer, so swapping either is a one-file change. | A single multimodal provider (one vendor for everything, worse latency on text); self-hosted models (ops cost we are not ready for) | Accepted |
| 2026-07-29 | One stateful dependency instead of three. At our volume — one org, a handful of projects — Redis would be operational overhead with no measured benefit. The interface is narrow enough to swap if it ever binds. | Redis + BullMQ (a second stateful service to run and back up); a managed queue (vendor lock-in for an internal tool) | Accepted | ||
| D-004 | 2026-07-29 | Agents get the same API and the same approval gate as humans, authenticated with scoped keys | An agent that can bypass approvals to spend money is a liability. Reusing one authorization path means there is exactly one place where "may this actor do this?" is answered — and one audit trail. | A privileged internal agent API (a second authorization path to keep correct); direct DB access for agents (no audit, no scope) | Accepted |
| D-003 | 2026-07-29 | Hono for the API service | Small, TypeScript-first, Web-standard Request/Response — so tests drive the real app in-process with app.request(), no port and no mock server. Runs on plain Node. | Express (callback-era types, needs more glue for zod); Fastify (heavier, schema story duplicates zod); NestJS (decorator/DI weight we do not need); Next API-only (defeats D-001) | Accepted |
| D-002 | 2026-07-29 | @verjson/contracts holds zod schemas, and both apps import it | The contract becomes executable. The API validates with the same object the web app's types are inferred from, so a breaking change fails typecheck in both workspaces instead of at runtime in staging. | Hand-maintained TS interfaces on the client (drift is a matter of time); OpenAPI + codegen (a generation step to keep in sync for a two-app repo) | Accepted |
| D-001 | 2026-07-29 | Two deployables — @verjson/api and @verjson/web — in one npm-workspaces monorepo | They scale and fail for different reasons: a publishing worker hammering LinkedIn has nothing to do with a chart render. One repo keeps the contract and the CI gate atomic; two containers keep the runtimes independent, which is what the Compose→Kubernetes path in requirements.txt needs. | One Next.js app with route handlers (simplest, but the API can never scale or be rewritten independently — the original template's D-001, now superseded); two separate repos (independent releases, but every contract change becomes a two-PR dance for a two-person team) | Accepted |