stpierre.ai · internal · September 3, 2026 · companion to the Builder Delivery Handbook

Delivery Operating Manual

How this company actually operates — reverse-engineered from 13 weeks of running it (1,752 logged working sessions, 45,000+ ledgered changes). The Handbook tells you what exists; this tells you how to run it and how to decide.

The whole system is one loop, run twice:

Offense — set weekly targets Monday, plan to hit them, commit daily at the huddle, get graded nightly at 7:50pm, review Friday.

Defense — machines watch every rail and file issues automatically; you triage the register daily, fix with proof, and turn every repeated failure into a gate so it can't happen again.

Your job is to run both lanes so well that Akash never discovers a problem before you report it.

Offense

The weekly loop

WhenRitualOutput
Monday AMSet the week: pick weekly targets with Akash, write them into the scorecard (scorecard_metrics — the ONE place targets live), and name the single biggest constraint in the business in one sentence.Weekly plan: 3–5 dominoes that attack the constraint.
Daily 08:30Huddle off the auto-built page (rooms.stpierre.ai/huddle). Numbers vs pace, blockers, today's commitment said out loud.One committed domino per person per day.
Daily 19:50The EOD job grades commitments automatically — anything still "planned" is marked missed, in writing. You post your EOD report.Graded day + tomorrow's plan.
FridayWeekly scorecard review (rooms.stpierre.ai/scorecard): this week vs last, red/amber/green per metric.What worked, what changes Monday.

The huddle board — what you're both looking at

The huddle is three panes, always the same three, 15 minutes total. One tracker per kind of thing — never a second list of any of them:

PaneTrackerThe 5-minute question
KPIsThe scorecard (huddle page heroes)On pace or behind? Every behind number gets an owner and a same-day action.
BugsThe issue register (issues table), sorted severity × ageAny criticals? What closed yesterday with proof? What's aging?
ProjectsCommand home — the week's one big project + next action per active projectDid the big project advance yesterday? What's today's move on it?

Then commitments out loud, and done. Tasks live as the day's commitments plus P0–P3 tasks in the brain — the huddle is where they're chosen, the EOD job is where they're graded. So the full system is exactly three trackers: issues = problems found · command home = containers of work · commitments/tasks = today's chosen moves.

Prioritization is mechanical, not a debate: name the constraint first; only work that removes it gets a slot; everything else is parked by default. Never present a ranked list without the constraint named on top. During a cash sprint, infrastructure and tooling lose to delivery unless they ARE the bottleneck.

Offense

Picking the week's one big project

Context: after 13 weeks, a minimum viable version of everything exists — every process has a first, rough draft. The work now is optimization, and capacity is one big project per week. The pick is mechanical:

StepRule
1Walk the money path left to right on the Friday scorecard: outbound demand → booked sales calls → closes → market launch → leads → qualified → delivered appointments → client retained → cash collected.
2The first stage below target is the constraint — that's the week's project. Never fix stage 5 while stage 2 is the first red: downstream reds are usually upstream's fault, and fixing them buys nothing until the upstream flows.
3One override: anything threatening a paying client's results beats acquisition work, even if it sits later in the path. Existing cash defends before new cash attacks.
4Tie-break with the smallest fix that moves the number — not the most interesting build. If two candidates truly tie, pick the one with the shorter undo.

Where optimization candidates come from when the funnel is green: the rule-of-three build tasks (patterns that clunked 3× in a week) and issue-register clusters. That queue is the backlog; the funnel walk decides whether this is a week for it.

The week's project laddered up: it must serve the current 6-week sprint outcome, which serves $100K by Dec 31. If a candidate serves neither, it's parked no matter how broken it feels.

Process

The builder's daily sweep — every surface, every day

The no-orphan rule made operational. Run in this order; the whole sweep is ~60–90 minutes bracketing the huddle:

#SweepSurfaceDone means
1System check (before huddle)Watchdog logs + overnight jobs + issue registerEvery job green or filed as an issue; criticals claimed.
2B2C conversationsGHL (the B2C client CRM) + B2C boardEvery open homeowner conversation has a next action; stalls nudged inside ruled limits; nothing waiting on us.
3B2C lead transferClient portal + delivery receiptsEvery qualified lead from yesterday is with its client, with proof (card, text receipt, portal row). Zero unexplained gaps.
4B2C adsB2C ScoreboardSpend delivering where intended; kill rules applied; starred-but-paused decisions surfaced to Akash as one line.
5B2B railssales.stpierre.ai + Instantly + cold-SMS queueEvery reply routed to Akash's call list with context; sends inside caps; every B2B lead has a next action and an owner.
6B2B adsAds-Live (B2B view)Same read as B2C: spend-first, kill rules, decisions as one line.
7Content & creative pipelineContent tracker + creative production queueToday's organic post is publishing; client creative in production is on schedule; the idea backlog captured, not lost.
8EOD#daily-reportsThe five-part report posted; decision log current; tomorrow's first domino named.

Org

Departments & ownership

The ruled org shape is seven departments. Two people cover all seven: Akash owns every human conversation with money on the other end; you own every machine and every process. Per-department status of SOPs/playbooks is mapped live at team.stpierre.ai/sops/ — the honest summary: a v1 exists for nearly everything; almost nothing is optimized. That's the design, not a failure — zero-to-one is done, one-to-ten is your job, one department per week.

DepartmentFunctionsYou ownAkash ownsScorecard / tracking
B2B DemandCold email, cold SMS/iMessage, B2B ads, list building, Loom/demo assetsThe machines: lists, sends, caps, deliverability, demo pagesMessage strategy sign-off; every reply-turned-conversationsales.stpierre.ai · Instantly · send registry
B2C DemandMeta campaigns, creative production, market launches, pages/datasetsAll of it, through the launch gateContact-sheet approvals; budget ceilingB2C Scoreboard (Live/Markets/Creatives)
SettingLead qualification, booking, speed-to-leadThe bot + rails that qualify and bookAny human touch a prospect getsHuddle heroes · response-time metric
ClosingSales calls, offers, paymentCall sheets, demo pages, prep100% — this is the 9–5Sales CRM · Stripe (close truth)
DeliveryLead transfer, appointment delivery, client CRM hygieneAll of it — this is your core seatNothing (first time ever)Client portal · delivery receipts · ≥5 appts/client/wk
RetentionClient health, results proof, escalation preventionMonitoring, diagnosis, the draft of every client messageSending every client messageOnboarding & retention board
FinanceCash tracking, Profit First split, invoicingKeeping the finance room honestEvery dollar that movesFinance room · 4DX cash scorecard

Organic content is the one split seat: Akash records (his face, his voice — not delegable); you own everything after the recording — production, editing, publishing cadence, and the idea pipeline. Target: 5 reels/week + 1 long-form, tracked on the scorecard like any other metric.

Org

The org: two humans + an agent workforce

The org chart is: Akash (CEO, sales) → you (ops lead) → AI agents. Every department's labor is agents — the Messenger bot, the follow-up rails, the watchdogs, the ~50 scheduled jobs, and the Claude/Codex working sessions that build and fix things. There are no other humans. That makes your actual craft managing an agent workforce: building agents, directing them, and quality-controlling their output.

Direction

Where we are → where we're going, per department

Zero-to-one is done everywhere; this is the one-to-ten map. Each row's gap is a candidate for a weekly big project — the funnel walk (above) decides the order.

DepartmentWhere we are (Sep 3)Where we're goingThe gap to close
B2B Demand~156 cold emails sent, a handful of replies, one burned hot lead; dialer not yet live; lists partly wrong-serviceDialer live + clean lists feeding 2+ real sales conversations/day to AkashKixie go-live · list rebuild · reply routing that never re-asks a question
B2C DemandCampaigns paused; 514 ads standardized under a locked launch contractEvery sold market live inside its CPL target, launched through the gate in <48h of a closeControlled reactivation, market by market, behind the recovered backlog
SettingBot captures a phone on ~half of Messenger leads; response <60s when rails are up; humanity fixes just shippedQualified-and-booked above 20% of leads, zero silent stallsProve the repaired rails on live traffic; instrument every stall
ClosingOffer iterating toward a $500 one-call trial; scripts in draftOne fixed offer card, one standard call, cash collected weeklyAkash's lane — you supply call sheets, demo pages, and booked slots
DeliveryBacklog mid-recovery; 2 named holes; receipts partially proven; 27 open issues100% of qualified leads delivered with a receipt, same day; ≥5 appts/client/wk; zero criticals openFinish the repair SOW workstreams; own the register to zero
RetentionAll 3 clients paused; results proof ad-hocClients live, seeing weekly proof of results, renewing without being chasedA weekly per-client results packet (you draft, Akash sends)
FinanceCash below the $1,500 floor; invoicing manual; rooms recently caught staleFinance room honest daily; invoices same-day; 50/10/10/30 split applied on every dollar inKeep the room honest; tee up every invoice the moment cash is owed

Defense

How problems get found

Four detection channels, in order of preference. The 13-week pattern is clear: every big incident that hurt was caught by channel 4. Your mandate is to make channels 1–3 catch everything first.

#ChannelHow it works
1Watchdogs (automated)Monitors watch every rail — silent-lead watch, lead-handoff watchdog, rail watchdog, launch QA, page-hijack detector, secrets scan, weekend check-ins — and auto-file findings into the issue register with a severity. A nightly fixer auto-repairs known patterns first; only what it can't fix escalates.
2Reconcilers (automated)Jobs that compare two sources of truth that should agree (e.g. "every qualified lead has a delivery receipt") and file the differences. Data drift is how leads get lost silently.
3Standing audits (scheduled, human-run)The recurring sweeps Akash has been running by hand: daily "is everything on and firing?" on ads + conversations; weekly end-to-end lead-journey audit (pick 5 real leads, walk every hop); weekly reply audit (read what actually went out and what humans replied). These become YOUR calendar.
4Someone notices (bad)Akash or a client sees a broken thing. Every one of these is a defense failure: fix the thing, then add the watchdog or reconciler that would have caught it. Target: zero.

Diagnosis rules (learned the hard way)

Defense

The bug tracker & lifecycle

The tracker is the issues table in the brain (Postgres) — one unified register that the watchdogs and anomaly sweeps sync into automatically (issues_sync.py). Each issue has a kind, severity, owner, age, and status. Manual issues (something you or Akash notices) are inserted by hand into the same table — one register, never a second list.

Severity ladder & response time

SeverityMeansYour SLA
CriticalLeads not flowing, a client can see it, or money at risk. (SaaS: SEV-1.)Drop everything. Fix or mitigate same day. Tell Akash immediately with the fix in motion.
WarningA rail degraded but a backstop is holding, or a risk building. (SEV-2.)Fix this week. Appears in your EOD.
InfoWorth knowing, not urgent. (SEV-3.)Batch weekly; close or promote.

Lifecycle

open → in_progress → resolved (or wont_fix with one line of why). Issues auto-resolve when their source alert clears; everything else you close by hand. Three rules make closing honest:

Day-one reality check: the register currently has 27 open issues, every one owned by "akash", with criticals aging since Aug 11 — because the owner was also the salesman, the builder, and the support desk. Your first defense act is taking ownership of the register and working it to zero criticals.

History

13 weeks of recurring problems — and what fixed them

Seven problem families account for nearly every incident in the record. Learn these and you'll recognize the next incident before it grows — new problems here are almost always an old family wearing a new shirt.

Recurring problemWhat kept happeningAction takenOutcome today
Silent lead lossHandoffs died quietly: transfers stopped (late Jul), a DND flag blocked a hot lead for days (mid Aug), 19 homeowners stuck behind three stacked failures (late Aug), the 158-lead never-sent scare (Sep)Message mirror, watchdog fleet, reconcilers, the issue register, the 7-workstream repair SOWDetection is automated; the "lost" pile shrank from unknown to 2 named leads. Repair SOW still mid-execution — your first inheritance.
Over-messaging & robot tells3am texts to a client's crew, 4 texts in 5 seconds, re-asking questions leads already answered, chasing complete profilesSuppression lists, one-message-per-turn, 45-day repeat block, mechanical STOP, template-fingerprint enforcement, chase stand-down rulingMechanically blocked at send time. The lesson: send discipline lives in code, never in memory.
Copy & offer driftTwo unapproved deploys in one day (Aug 4); retired offers (the $8K/ten-contracts package, the retired guarantee-window wording) still teaching on live pages weeks after being replacedDeploy gate wired into every deploy command, before/after contact sheets, offer-consistency sweeps, canonical context files with retired-terms graveyardsDeploys are gated 100%. Stale copy on old pages still surfaces — flag it, never repeat it.
Platform config gaps15 of 17 Facebook pages couldn't receive a lead form; missing datasets made leads vanish; per-page ToS assumed fleet-wideThe fail-closed launch-QA gate: every market launch checks pages, datasets, ToS, tracking, routing before a dollar spendsLaunches can't skip the checklist. The gate is the memory.
Dashboards lyingA "LIVE" badge over 6-day-dead data; false success receipts; green renders over skipped stepsHonest-badge rules, freshness monitors, heartbeats on every job, renders that exit "degraded" instead of pretendingStaleness now self-reports. Standing rule: a stale number behind a green badge is an ops failure, whoever's number it is.
Decision sprawlApproval queues silently filled with reversible decisions nobody executed (12 deep at one point); work waited on asks that weren't neededThe reversible-default ruling: reversible → execute, ledger, report; only the short irreversible list asks first; decision log reviewed after the factQueues retired. This is why your decision log exists — it's the replacement for asking.
Work dying with its sessionUsage limits and crashes killed multi-hour work mid-flight; one 20-agent fan-out burned a day's capacityIncremental reports written as work proceeds, the ledger as the recovery point, SOW-file handoffs, agent waves capped at 5Work survives process death. Habit: write down progress as you go — never hold results only in your head or one window.

The meta-pattern across all seven: every durable fix was a mechanism, not a reminder. Memory-based fixes recurred; gate-based fixes didn't. When you fix something, ask "what makes this impossible to repeat?" — that answer is the real fix.

Reference

How SaaS companies do this → our simplest version

Over 13 weeks this company independently converged on the standard SaaS reliability stack. Here's the mapping, so you can borrow industry practice without importing its ceremony:

SaaS practiceOur versionStatus
Monitoring & alerting (Datadog/PagerDuty)Watchdog fleet + escalator, filing into the issue registerLive
Issue tracker (Jira/Linear)issues table + P0/P1 tasks in the brain; surfaced at the huddleLive — needs an owner (you) and close-with-proof discipline
Incident severity levels + on-callcritical/warning/info; you are on-call during work hours, the fixer overnightLive
Deploy pipeline + audit log + rollbackDeploy gate (automatic on every deploy command) + change ledger with per-change rollback notesLive
Postmortems with action itemsRule of three → "rulings become gates": recurring failures get turned into mechanical checksLive, young — 5 ruling-checks exist; grow this
RunbooksThe references library + per-skill "Gotchas" logs (append-only failure notes on every tool)Live
Data-integrity jobsReconcilers + nightly fixerLive, mid-buildout (workstream of the repair SOW)

What we deliberately skip: sprint points, ticket ceremonies, standup theater, dashboards for their own sake. Two people don't need Jira; they need one honest register, one scorecard, and one ledger.

Judgment

Akash's decision frameworks

Six rules explain ~90% of the decisions in the record. Use them and your calls will match his.

#RuleIn practice
1Name the number first.Every proposal opens with which target it moves and by how much. If it moves nothing on the current board, the honest recommendation is "park it." No silent orphan work.
2Recommendation, then defense.Never bring options without a pick. Format: the options, the one you choose, the plain-English why, and the cost of waiting. Bring decisions, not homework.
3Reversible → do it. Irreversible → ask.If an action can be cleanly undone, execute it, ledger it, report it. The ask-first list is short and fixed: money out, anything a client or homeowner receives, anything published live, credential/permission changes, deletions.
4Prove the gap before building.Before any new build: does this already exist? Is the problem real in the data, not just plausible? And after: verify before calling it done — run it, open it, fire a test through it.
580/20, smallest thing.Surface the 20% driving 80% and explicitly kill the rest. New builds pass a 4-part contract first: goal, constraints, output format, failure conditions.
6Corrections become rules, same day.When Akash corrects you or a wrong assumption causes a bug, write the rule down where it will be enforced (a gate, a checklist, a gotcha note) that session. "I'll remember next time" is a violation.

Focus framework behind it all: urgent loses to important; the constraint decides priority; and the daily three (value, content, conversations) are Akash's — yours is the delivery mirror: leads answered, appointments delivered, systems green.

Protocol

Communication protocol — what Akash sees, and when

The shared surface is the scorecard system, not chat. You both look at the same four pages; chat is for exceptions.

SurfaceWhat it answersCadence
Huddle page (/huddle)Are we on pace today? What's blocked? What did each of us commit to?Daily, together
Issue register (defense board)What's broken, how bad, how old, who owns it?You: daily triage. Akash: glances, never works it.
Command home (projects)What projects exist, next action, % complete. One tracker — never a parallel status doc.You update as things ship
Weekly scorecard (/scorecard)Did the week hit its targets?Friday, together

Your decision log — the trust mechanism

You will make judgment calls all day without asking. That's wanted. The deal is: every meaningful call gets logged the moment you make it, in one line, four parts: the fork · what I picked · the number it serves · how to undo it. Akash reviews the log after the fact instead of approving in advance — that's what makes 9–5 sales possible. A decision that would embarrass you in that log is a decision to ask about first.

The daily EOD (Slack #daily-reports)

Five parts, same order every day: Shipped (with proof) · Numbers (per live client) · Broken/behind (with fix + ETA) · Decisions I made (the day's log lines) · Blocked (split: needs-your-word vs your-lane) · Tomorrow's first domino.

Interruption rule

Interrupt Akash mid-day only for: leads not flowing, client-visible breakage, or money at risk — and always with the fix already started. Everything else waits for the EOD or the huddle.

The boundary, stated once: you do everything except client-facing and prospect-facing communication. When a fix requires telling a client something, you draft the message and the receipts; Akash sends it.

Scorecard

Your scorecard — leading & lagging

Leading (your inputs, graded daily)

  • Morning system check done by huddle: every watchdog and overnight job green or filed
  • Critical issues closed same day (count + %)
  • Backlog leads advanced per day (each one to a next action or delivered)
  • Issues closed with proof this week vs opened
  • 100% of deploys through the gate with an approved contact sheet

Lagging (the outcomes, graded weekly)

  • Booked appointments per live client per week (≥5)
  • Median lead → first response (<60s)
  • Problems Akash or a client found before you did (target: 0)
  • Leads ending the week with no next action (target: 0)
  • Client credits/refunds caused by delivery (target: 0)

These go into the scorecard as your seat, so the same Friday review that grades the company grades the role. If a leading metric is green for weeks while a lagging one stays red, the leading metric is measuring the wrong input — say so and propose a swap.

Reference

How to find anything

QuestionWhere the answer lives
"What happened / why was this decided?"The brain's session_handoffs table — 1,752 summaries of every working session, searchable. This is the company's memory; check it before re-deriving anything.
"What changed on a live system, and how do I undo it?"change_ledger — every external change with before/after and rollback notes.
"Is this number right?"The brain is the source of truth; dashboards only render it. Check the table behind the dashboard before distrusting the metric.
"How does this tool/system behave? What bites?"The references library (.codex/references/) + each skill's append-only Gotchas log — the runbooks, written from real failures.
"What is this URL / dashboard?"The URL directory (url-directory.json) maps every live page to its source file and purpose.
"Has this been built before?"Deliverables pre-check: search the deliverables table before creating anything new. Duplicates are how truth forks.

Reference

Tech stack inventory

LayerTools
DemandMeta Ads (2 B2C accounts) · Facebook Pages per market · Typeform (forms) · the Messenger bot
Conversation & CRMGoHighLevel (one shared location) · mycrmsim iMessage bridge · LoopMessage (parked backup) · Instantly (cold email) · Kixie dialer (being set up) · Calendly
Delivery to clientsClient portal (client.stpierre.ai) · DripJobs (AlphaLift) · group text (Dan) · Zapier + n8n glue
Truth & memoryPostgres "brain" (you get full access) · change ledger · issue register · session handoffs
SurfacesCloudflare Workers/Pages: rooms.stpierre.ai, data.stpierre.ai, command home, B2C scoreboard, team.stpierre.ai docs
Reliability~50 scheduled jobs on the host Mac: watchdogs, reconcilers, nightly fixer, deploy gate, nightly git snapshot
Payments / booksStripe (B2B close truth) · payments land in St. Pierre Checking via ACH

Onboarding

The 4-week ramp

Buyback logic: Akash hands off the lowest-judgment, highest-hour work first, audits your outputs weekly, and keeps only what genuinely needs the owner.

WeekYou ownAkash still doesExit test
1Shadow defense: run the morning check, read the register, walk one lead end to end daily. Write the EOD.All fixes, all decisionsYour EOD matches reality without corrections.
2Own defense: the register transfers to you. Triage, fix, close with proof. First decision-log entries.Offense; reviews your closesZero open criticals; every close has proof.
3Own offense: draft Monday's weekly plan and targets; run the huddle; ship one gated deploy end to end.Approves the plan; sellsThe week's plan survives Friday review without a rewrite.
4The full loop, both lanes. Akash reviews only the decision log + Friday scorecard.Sales 9–5, client comms, approvalsAkash went a full week learning about every problem from you first.

Standing bar, forever: nothing breaks silently, nothing needs asking twice, every number traces to a source, and "I don't know" beats a guess.