# AutomationAudit — Brain v1.0

> Internal strategy doc. Working map of what can be automated today, what bites in production, what operators say versus what they need, and how the audit agent decides.
>
> Compiled May 2026. Every quantitative claim has an inline URL citation. Reliability scores carry a ~1.3× haircut vs. vendor benchmarks (see §6). Sections are tagged for decay.

## Contents

1. Strategic thesis · [decay: 12mo]
2. The audit job — end-to-end engagement · [decay: 6mo]
3. Process database (200+) · [decay: 6mo]
4. Wedge — top 10 + never-automated tier · [decay: 6mo]
5. Anti-list — don't automate this · [decay: 6mo]
6. Agent capability matrix · [decay: 3mo]
7. Vendor & platform map · [decay: 3mo]
8. Recommended stacks · [decay: 3mo]
9. Failure modes & cross-cutting heuristics · [decay: 6mo]
10. Per-role pain inventory · [decay: 12mo]
11. Operator ↔ vendor glossary · [decay: stable]
12. Discovery question bank · [decay: 12mo]
13. Readiness scoring rubric · [decay: 12mo]
14. Decision tree · [decay: 12mo]
15. Recommendation deliverable template · [decay: 12mo]
16. Red flags during intake · [decay: 12mo]
17. Pricing & business reality · [decay: 6mo]
18. What I can't verify yet · [decay: 1mo]

## 1. Strategic thesis [decay: 12mo]

- **The wedge is the audit itself.** A 10-minute, free, AI-run audit that returns a ranked, dollar-quantified, builder-ready list of automations for one business plus an explicit "don't automate" list. The build/maintain offer monetizes; the audit hooks.
- **Compete on specificity, not capability.** Defensible asset is the 200-row, citation-backed process database (§3) plus the 50-trap anti-list (§5). Both signal we know which automations actually ship.
- **Opening move is the never-automated tier.** 46 entries in the database flag `never_automated_yet` — high pain, low awareness, no competitive density. Examples: vendor-renewal alerter, scope-creep tracker, customer-health rollup from telemetry + tickets, partner-attribution stitcher, customer-by-customer profitability.
- **Offer for v1: Free Audit → $99/mo build & maintain.** Subscription is low enough for an ops director to swipe AmEx (no procurement). Per-build pricing ($1.5k–$8k) for users who refuse recurring. A $499 "Pro Audit" is a v2 SKU.
- **First customer profile: 25–150 person services businesses.** Agencies, consultancies, dev shops. Founder/COO buys. Stack is HubSpot/Pipedrive, Slack, Notion/Asana, Gmail, Stripe. Avoid v1: enterprise (procurement), regulated (PHI/financial advice), <5-person (won't pay).

## 2. The audit job [decay: 6mo]

### Six stages

1. **Intake (60–180s).** Company URL, role, size, primary tools. Mode: voice note (default), AI voice call, or typed. Context priming: 8–12 prompts spanning week-shape, repeated work, hated tasks, last fire-drill.
2. **Discovery (3–5 min).** Conditional branching question bank (§12). Function-aware. Each surfaced pain quantified (frequency × duration × loaded hourly cost) and tagged to a candidate from the process database. Hard caps: max 25 candidates, max 12 minutes total.
3. **Scoring (~30s, server-side).** Each candidate scored on the 7-axis readiness rubric (§13): frequency, time cost, ambiguity, stakes-of-error, data availability, tool-fit, change-mgmt cost. Decision tree (§14) routes each item to: full agent, deterministic workflow, RPA, fix-process-first, or leave-human. Sub-threshold items go to the anti-list explicitly, never silently dropped.
4. **Recommendation generation (~60s).** Top 5 ranked by `dollar_value × readiness × sales_appeal`. Each: plain-English explanation, concrete stack, estimated build time, confidence score, similar prior deployments. Anti-list section calls out non-recommendations. Headline numbers: total recoverable spend, hours/month.
5. **Deliverable (gated).** Email gate unlocks full report. Public preview shows top item. PDF + shareable URL. Each item has three CTAs: DIY (instructions) · We build for $X · We build + maintain for $99/mo.
6. **Handoff (sales).** Build/subscribe click fires a human-shaped Slack ping with audit + pre-filled proposal. 24-hour SLA. No-pressure path captures email for nurture.

### Pricing model — recommendation

Default for v1: **Free Audit + $99/mo subscription, with per-build escape hatch ($1.5k–$8k).** Subscription is the conversion engine; per-build is the relief valve for buyers who refuse recurring. Comparable: Zapier Experts marketplace per-build rates (https://zapier.com/experts); emerging "AI ops" subscription shops at $99–$499/mo (https://goodspeed.studio/blog/n8n-agency-pricing-what-it-costs-to-work-with-an-n8n-partner).

### Unit economics targets

- Time per audit: <5 min of human time post-MVP (review + send). 100% agent-run is the target.
- Cost per audit: <$1.50 in LLM/voice/infra at scale. Voice (Retell) ~$0.05/min × 8 min + Claude Sonnet 4.5 ~$0.05 per 50k tokens ≈ $0.65 raw. [UNCITED — modelled]
- Conversion targets: 4–8% free→$99/mo, 1–3% free→$1.5k+ one-shot.
- LTV target: $1,800 ($99/mo × 18 mo expected SMB tenure).
- CAC target: <$200 via SEO + practitioner presence on X.

## 3. Process database [decay: 6mo]

Stored canonically in `research/processes.json` (206 entries; 46 `never_automated_yet`; coverage across Ops/Finance/Sales/CS/Support/HR/Recruiting/Marketing/Legal/Eng/Exec plus 7 verticals: agencies, e-commerce, real estate, healthcare admin, logistics, professional services, local services).

Status tiers:

- `frequently_automated` — Zapier/Make territory. Sell on superior reliability or stack consolidation.
- `rarely_automated` — Some shops do it; friction stops most. Bread-and-butter.
- `never_automated_yet` — Wedge tier. No vendor or agency has packaged a clean solution. Cheap acquisition.

Headline filter examples:
- 46 entries with `status = never_automated_yet`
- 57 entries with `status = frequently_automated` (use for trust-building / parity claims)
- 103 entries with `status = rarely_automated` (the meat of recommendations)

Top categories by entry count: Ops (37), Sales (36), Finance (35), then Vertical (66 across 7 industries), then the smaller function buckets.

## 4. Wedge — top 10 + never-automated tier [decay: 6mo]

### Top 10 highest-ROI starting processes

Ranked by `cost_saved × sales_appeal × reliability_today` from `processes.json`. Use these as cold-outbound proof points and landing page case studies first.

1. Scope creep tracking on services contracts (Ops, agencies/services)
2. Customer health rollup with leading-indicator churn flags (CS)
3. Vendor renewal alerts before auto-renew (Ops)
4. Shadow IT discovery & cost rationalization (Ops)
5. QBR deck auto-generation from usage + CRM (CS)
6. AI PR review pre-pass (Eng)
7. AR follow-up cadence on overdue invoices (Finance)
8. Internal IT ticket triage + first-touch response (Ops)
9. Vendor bill OCR + 3-way match (Finance)
10. Voice-of-customer rollup from tickets + surveys + reviews (Support)

### Wedge tier examples (`never_automated_yet`)

- Vendor renewal alerter (avg 18% off renewal per Vendr; SMBs don't have CLMs)
- Customer health rollup from telemetry + tickets + payment events
- Scope creep tracker on services contracts
- Partner-attribution stitcher for BD/partnership pipelines
- Customer-by-customer profitability (unit economics with cost allocation)
- Expansion-signal detection from feature adoption + seat growth
- Capacity planning for services teams (8-week forward look)
- Test-flake quarantine with auto-issue + last-passing commit
- Conference / event ROI tracker through to closed-won
- PTO conflict detector (3 engineers out same week)
- Holiday/leave coverage planner
- Recurring-meeting agenda hygiene from Slack + Linear + Gong context
- Internal data hygiene drift detector (semantic, not just schema)
- Internal wiki staleness detection
- COI (certificate of insurance) collection from vendors

## 5. Anti-list [decay: 6mo]

Stored canonically in `research/anti-list.json` (60 entries across 6 failure-mode categories).

### Failure modes — top 10 cited incidents

1. **Air Canada chatbot** invented bereavement-fare refund window; BC CRT held Air Canada liable. Tribunal rejected "chatbot is separate entity" defense (https://www.cbc.ca/news/canada/british-columbia/air-canada-chatbot-lawsuit-1.7116416).
2. **Klarna** replaced 700 CS agents with AI, then reversed and is rehiring (https://www.entrepreneur.com/business-news/klarna-ceo-reverses-course-by-hiring-more-humans-not-ai/491396).
3. **Cursor support bot** fabricated a "one-device login" policy in April 2025; customers cancelled (https://www.theregister.com/2025/04/18/cursor_ai_support_bot_lies/).
4. **Workday class action** (Mobley v. Workday): nationwide ADEA collective certified May 2025; 1.1B applications rejected (https://www.lawandtheworkplace.com/2025/06/ai-bias-lawsuit-against-workday-reaches-next-stage-as-court-grants-conditional-certification-of-adea-claim/).
5. **iTutorGroup** paid $365K to settle EEOC's first AI-discrimination case (auto-rejecting older applicants) (https://www.eeoc.gov/newsroom/itutorgroup-pay-365000-settle-eeoc-discriminatory-hiring-suit).
6. **FCC TCPA AI-voice ruling** (Feb 2024) makes AI cold calls illegal without prior express consent; civil penalty up to $1,500 per call (https://www.fcc.gov/document/fcc-makes-ai-generated-voices-robocalls-illegal).
7. **DPD chatbot** swore at customers + wrote anti-DPD poem (https://time.com/6564726/ai-chatbot-dpd-curses-criticizes-company/).
8. **McDonald's-IBM drive-thru AI** ended after viral failures (https://www.cnbc.com/2024/06/17/mcdonalds-to-end-ibm-ai-drive-thru-test.html).
9. **NYC Local Law 144** (since Jul 2023): bias audits + candidate notice; $500–$1,500/violation/day (https://www.nyc.gov/site/dca/about/automated-employment-decision-tools.page).
10. **CVE-2025-59944** (Cursor case-sensitivity prompt-injection → RCE); $250k bank-assistant fraud (Jun 2025) (https://www.mayhemcode.com/2026/02/real-world-prompt-injection-attacks-10.html).

### Categories the anti-list covers

- Customer & comms (refunds, replies, voice cold outbound, moderation, dispute responses, crisis comms)
- HR & people (hiring decisions, performance reviews, PIPs/terminations, talent screening without feedback loop)
- Legal & compliance (redlines without attorney review, regulatory filings, IP/patent drafting)
- Finance (invoice approvals without HITL, expense fraud rules, ML-only forecasting, dispute responses)
- Sales (lead scoring without feedback loop, AI cold dials, autonomous pricing)
- Vertical-specific (healthcare clinical advice, real-estate tenant screening, legal medical translation)

## 6. Agent capability matrix [decay: 3mo]

**Headline reality check.** MIT State of AI in Business 2025: **95% of corporate GenAI pilots delivered zero measurable P&L impact** across 300 deployments (https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/). Gartner predicts **>40% of agentic AI projects will be canceled by end of 2027** (https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027). MIT same study: purchased/partnered AI succeeds ~2× the rate of internal builds.

Full matrix lives in `research/raw-capability-matrix.md` (13 capability rows + 21 vendor entries). Summary:

| Capability | Reliability | HITL | Notes |
|---|---|---|---|
| Browser automation | 3/5 | Yes (writes) | Stagehand v3 + Browserbase; injection via page text is real |
| Code generation | 3–4/5 bounded, 2/5 long | Yes (PR) | SWE-bench Pro top **46%** vs Verified 81% (https://www.morphllm.com/swe-bench-pro) |
| Email triage/draft | 3/5 | Yes (send) | Never auto-send |
| Calendar | 2/5 unsupervised | Yes | Phantom events, defragging gone wrong |
| Doc/sheet | 3–4/5 | Yes (finance) | Claude + Skills leads |
| CRM ops | 3/5 | Yes (outbound) | Garbage-in/out; HubSpot Breeze Customer Agent is most production-ready |
| Voice agents | 3/5 | Yes (escalate) | Retell HIPAA/SOC2; <800ms latency required |
| Support deflection | 3–4/5 mature KB, 2/5 day-one | Always tier 2+ | Fin ~51% avg resolution; learn from Klarna |
| RAG / research | 3/5 | Yes for decisions | Naive RAG fails retrieval ~40% (https://lushbinary.com/blog/rag-retrieval-augmented-generation-production-guide/) |
| Data extraction | 4/5 typed, 3/5 scanned | Yes >threshold | Claude 4.5 97–98% field acc |
| Workflow orchestration | 3–4/5 | Yes (branches) | LangGraph + Temporal beats bare SDKs |
| Long-horizon | 2/5 | Checkpoints | 95%/step × 50 steps = 7% success |
| Computer-use | 2–3/5 | Yes (writes) | OSWorld-Human shows efficiency lag (https://arxiv.org/html/2506.16042v1) |

## 7. Vendor & platform map [decay: 3mo]

Full breakdown in `research/raw-capability-matrix.md` Part 2. Decision summary:

- **Anthropic API + Claude Agent SDK** — default for coding/doc/tool-heavy. Cache aggressively.
- **OpenAI API + Agents SDK** — multimodal-heavy or one-vendor preference. Agents SDK lacks checkpointing.
- **LangGraph + LangSmith** — production durability layer. Use over CrewAI for any real system.
- **CrewAI** — demos only; manager-worker executes sequentially despite docs.
- **n8n (cloud / self-host)** — 80–90% cheaper than Zapier at high volume. AGPL gotcha for productized resale.
- **Zapier** — non-tech founder, low volume. Surprise-billing complaints recurring.
- **Make.com** — SMB visual canvas with branching.
- **Lindy.ai** — solo operator / agency-fast assistant agents.
- **Gumloop** — batch document workflows.
- **Retool + Retool Agents** — internal ops dashboards with agent assist.
- **Supabase (pgvector)** — default agent backend. Postgres + auth + RAG in one.
- **Stagehand + Browserbase** — any browser automation across changing sites.
- **Retell / ElevenLabs Agents / Vapi** — Retell for regulated, ElevenLabs for brand voice, Vapi for max control.
- **Nango / Merge / Paragon** — Nango for startup, Merge for embedded breadth, Paragon for enterprise sales motion.

## 8. Recommended stacks [decay: 3mo]

- **No-code MVP, non-technical founder** — Lindy (or Zapier Agents) + connectors + Claude.
- **SMB automation, agency-owner-no-devs** — Make.com or n8n cloud + Claude/OpenAI + Airtable/Supabase + Firecrawl/Apify.
- **Internal ops dashboard** — Retool + Supabase + Claude API.
- **Browser-heavy** — Stagehand v3 on Browserbase + Claude Sonnet 4.5 + Temporal + Supabase.
- **Document-heavy** — Gumloop (batch) OR Claude API + Skills + LangGraph + Postgres.
- **Sales/support automation** — Intercom Fin or Zendesk AI Agents + Retell voice + Clay/HubSpot + Merge unified API.
- **Productized service** — n8n self-hosted (watch AGPL) OR LangGraph + Anthropic/OpenAI + Supabase + Browserbase.

## 9. Failure modes & cross-cutting heuristics [decay: 6mo]

Top 15 failure modes (citations in `raw-capability-matrix.md`):

1. Eval gap — works in dev, fails in prod
2. Long-horizon compounding (95%/step × 50 = 7%)
3. pass^k collapse (τ-bench pass^8 <25%)
4. Prompt injection (CVE-2025-59944; $250k bank-assistant fraud)
5. Brittle browser DOMs
6. Hallucinated tool calls (CrewAI, Devin)
7. Context cost spiral
8. Permissioning hell (OAuth scope sprawl)
9. Vendor pricing surprises
10. HITL fatigue → rubber-stamping
11. Reward hacking / spec gaming
12. Statelessness / restart loss
13. Manager-worker fan-out that doesn't fan out
14. Voice hallucinations & latency
15. Benchmark contamination (OpenAI dropped SWE-bench Verified)

### Cross-cutting heuristics (the audit agent's rules)

```sysprompt
- Prefer buy over build by ~2:1.
- Quote a reliability haircut: divide vendor-claimed accuracy by ~1.3× for real-world distribution shift.
- Always quote pass^k or k-attempt reliability, not pass@1.
- Budget retries (30% retry rate at $0.50/run = $0.15 hidden cost per task).
- Mandatory HITL for send (email/SMS), spend (>$X), legal/compliance, customer-visible writes.
- Pick durability infra first. LangGraph + Temporal (or Mistral Workflows pattern). Statelessness kills more deploys than model quality.
- Default vector DB: pgvector if you have Postgres.
- Default doc extraction: Claude 4.5/Opus. GPT-4o only on degraded scans.
- Voice in regulated industries: Retell or ElevenLabs (SOC2/HIPAA/GDPR out of box).
- Browser: Stagehand+Browserbase over Playwright unless target is owned and stable.
```

## 10. Per-role pain inventory [decay: 12mo]

Stored in `research/role-pains.json` (12 curated roles). Full 20-role research in `research/raw-role-pains.md`. Most-likely buyers (own pain + budget + sign-off):

1. **Agency Owner** — Friday reporting panic; scope-creep margin loss; founder buys at $0–$3k/mo on AmEx.
2. **Director of Operations (SMB / mid)** — 9 tools that don't talk; KPI rollups; $0–$5k/mo direct.
3. **CFO at 20–100 person co** — month-end close, AR follow-up, board pack assembly; $5k–$50k/mo discretionary AND gatekeeper for everyone else.
4. **RevOps Lead** — CRM hygiene, forecasting, lead routing; $0–$3k/mo direct.
5. **Head of Customer Success** — renewals, QBRs, health scores; $1k–$10k/mo direct.
6. **E-commerce Operator** — reviews, returns, ad creative; founder full sign-off on $0–$5k/mo.
7. **Solo SaaS Founder** — inbox + support + content distribution; $0–$2k/mo.
8. **Practice Manager** — intake, insurance verification, denials, recall; $0–$2k/mo (physician/partner above).
9. **Local Services Owner** — speed-to-lead, dispatch updates, review requests; $0–$2k/mo.
10. **Engineering Manager** — PR review, on-call, codebase Q&A; $1k–$5k/mo direct.
11. **Recruiting Ops** — sourcing personalization, scheduling, ATS hygiene; $500–$2k/mo.
12. **Chief of Staff** — exec inbox, board packs, calendar prep; $0–$3k/mo on AmEx, founder sponsorship for larger.

## 11. Operator ↔ vendor glossary [decay: stable]

Stored in `research/glossary.json` (80 entries). Drawn from r/sales, r/operations, r/CustomerSuccess, r/accounting, r/sysadmin, podcast operator interviews, JD reviews, and public sales-call breakdowns.

The underlying pattern: operators describe pain in workflow terms ("Renewals sneak up on me"); we translate to capability ("Renewal forecasting + auto-prompted CSM workflow") and concrete stack ("Gainsight/Catalyst + renewal-90/60/30 agent in Slack").

## 12. Discovery question bank [decay: 12mo]

Stored in `research/raw-glossary-discovery.md` Part 2. Wrap in `<sysprompt>` in deployment. Eleven function blocks (Sales, CS, Support, Finance, Ops/RevOps, Marketing, HR/Recruiting, Legal/GRC, Eng/IT, Exec, Owner-led SMB). Each block: openers → conditional branches → quantification → tool-stack probes → red flags.

See the rendered `/research` page §12 for the complete tree.

## 13. Readiness scoring rubric [decay: 12mo]

```sysprompt
For each candidate process, score 1-5 on:

- Frequency (5: ≥5×/day per user, multiple users; 1: ≤1×/month, one user)
- Time cost (5: ≥$2,000/mo recoverable; 1: <$200/mo)
- Ambiguity (5: deterministic rules cover >90%; 1: each case is judgment)
- Stakes-of-error (5: low — internal, reversible; 1: high — legal, financial, customer-facing, irreversible)
- Data availability (5: all inputs in APIs / structured; 1: in heads or scattered docs)
- Tool-fit (5: native APIs, mature integration; 1: walled-garden, no API)
- Change-mgmt cost (5: sole owner who wants the help; 1: multi-team rollout, prior failed attempts, no champion)

Composite = sum, max 35.

Recommend automation if composite ≥ 22 AND no axis < 2.
Otherwise: route to anti-list or fix-process-first.

≥ 28: full agent build, top of report.
22–27: recommend with caveats; spec HITL gates.
16–21: deterministic workflow (Zapier/n8n) instead of agent.
11–15: fix-process-first; surface SOP or data hygiene need.
≤ 10: leave human; place on anti-list with reasoning.
```

## 14. Decision tree [decay: 12mo]

```sysprompt
1. Composite ≤ 10 or any axis = 1? → leave-human + anti-list.
2. Data availability ≤ 2? → fix-process-first (data consolidation / SOP extraction Phase 1).
3. Ambiguity ≥ 4 AND change-mgmt ≥ 4? → deterministic workflow (Zapier/n8n/Make).
4. Closed system, no API, but UI is stable? → browser-automation agent with HITL.
5. Document-in / structured-out (PDF/invoice/contract)? → Claude + Skills, or Gumloop for batch.
6. Conversational customer-facing (support, outbound)? → buy the deflection layer (Fin, Zendesk AI, Retell).
7. Multi-step, multi-tool, with state across steps? → LangGraph + Temporal. Anchor durability before agent logic.
8. Long-horizon (> tens of minutes per run)? → re-scope to smaller checkpoints.
9. Default: composite ≥ 22, no blockers → full agent with HITL on send/spend/legal/customer-visible writes.
```

## 15. Recommendation deliverable template [decay: 12mo]

````sysprompt
# {{process_name}}
**Composite readiness:** {{composite}}/35 · **Confidence:** {{confidence}}%
**Owner / role:** {{owner_role}}
**Status:** {{status}}
**Reasoning:** {{one_sentence}}

## What you're doing today
{{manual_workflow}}

## What you'll save
- Time: {{time_saved}} hrs/mo
- Money: ${{cost_saved}}/mo ({{cost_saved_assumption}})
- Annualized: ${{cost_saved_year}}

## How we'd build it
**Stack:** {{tools_needed}}
**Approach:** {{automation_approach}}
**Reliability today:** {{reliability_today}}/5 — {{reliability_rationale}}
**HITL:** {{hitl_mode}} — {{hitl_reason}}
**Estimated build:** {{build_hours}} hours · {{build_calendar_days}} calendar days

## Risks & failure modes
- {{primary_failure_mode}}
- {{mitigation}}

## What you'd own going forward
- {{maintenance_estimate}} per month · {{maintenance_owner}}

## Choose your path
- [ ] DIY — full build instructions inside
- [ ] We build for one shot: **${{one_shot_price}}**
- [ ] We build + maintain monthly: **${{monthly_price}}/mo** (recommended)
````

Report header block (above all top-5 items):

```
# Audit for {{company_name}}, {{role}}
**Total recoverable:** ${{total_savings}}/mo · {{total_hours}} hrs/mo back
**Top 5 automations · {{anti_count}} traps avoided · {{wedge_count}} wedge plays**
Generated {{date}} · Confidence-weighted · See methodology at /research
```

## 16. Red flags during intake [decay: 12mo]

```sysprompt
30 walk-or-scope signals. Walk-away signals override everything.

1. "We don't have one source of truth for X." → scope-down to data consolidation Phase 1, or walk.
2. "The last team / agency got fired." → political risk. Walk unless explicit explanation.
3. "We're in healthcare/PHI / payments / SOX." → scope down to advisory or BAA-eligible infra.
4. "We need it deployed by Friday." → scope-down or walk.
5. "We tried Zapier and it failed." → 20-min post-mortem before any pitch.
6. "Our team will adopt this — they're excited." → ask for proof (training plans, change champion).
7. "Our process is too unique to automate." → probe for the stable slice.
8. "We don't want HITL — we want full automation." → educate; if refuse, walk.
9. "Our IT team will help build it." → ask for named engineer with hours.
10. "Legal needs to approve everything." → map approval path with named contact.
11. No clear process owner. → insist on a single named owner with decision authority.
12. No data inventory. → scope a paid 1–2 week discovery.
13. No KPI for the process. → define one on the call.
14. Regulated industry + no compliance lead. → walk unless we have compliance in-house.
15. Multiple stakeholders openly disagree on the goal. → insist on a single decision-maker.
16. Recent reorg (<90 days). → scope down or wait.
17. "We just need an MVP" with no defined outcome. → force a measurable target.
18. "We're talking to 5 other agencies." → walk if lowest-price decision.
19. "Process is in someone's head." → Phase 1 = SOP extraction. Charge for it.
20. "We just got new leadership." → confirm new leader is sponsor.
21. "No budget approved." → real budget conversation before discovery.
22. "Our customers are very different — every account is special." → scope to internal ops.
23. "Can you guarantee X% accuracy?" → educate on evals + HITL.
24. "Can it replace [person]?" → reframe to augment + redirect.
25. "We want to own the IP / build in-house eventually." → price for transfer.
26. Founder/CEO alone, no operator counterpart. → require operator in discovery #2.
27. "We tried hiring and couldn't find anyone." → probe whether the role is judgment-heavy.
28. "We don't have observability / logs." → scope eval infra in Phase 1.
29. Procurement-led with no business owner present. → request business stakeholder.
30. "We don't have time for discovery — just build it." → walk, or non-refundable discovery fee.
```

## 17. Pricing & business reality [decay: 6mo]

Full data in `research/raw-business-pricing.md`. Anchors:

- **Build pricing for one automation:** $5k–$40k SMB sweet spot. $5k Goodspeed Accelerator is the cleanest template (https://goodspeed.studio/blog/n8n-agency-pricing-what-it-costs-to-work-with-an-n8n-partner).
- **Hourly rates:** US $100–$450; Philippines $25–$85; India $30–$60.
- **Monthly retainers:** $2k–$8k SMB sweet spot; $15k–$50k enterprise. Goodspeed sets $10k/mo as their lowest tier.
- **Token cost share:** 10–25% of monthly run for typical SMB agent; 40–60% for heavy-reasoning agents [UNCITED — composite estimate].
- **Sales cycle:** SMB 14–40 days; mid-market 30–90; enterprise 90–180+. Win rates 30–40%/20–30%/15–20%.
- **Free-audit → paid conversion benchmark:** 15–30% defensible, 80% best-case (Stoute Web case study).
- **Closest direct competitor:** Goodspeed Studio (same audit → build → retainer ladder, n8n-branded).
- **Brand-tier ceiling for SMB→mid:** Liam Ottley's $10k AI Audit Blueprint → $60k+ build engagements (https://www.scribd.com/document/890182441/How-to-Perform-Your-First-10-000-AI-Audit-as-an-AI-Agency).
- **Whitespace:** vendor-neutral audits. Most AAA-school shops do "audits" that are thinly disguised sales calls; Goodspeed's is real but n8n-branded.

## 18. What I can't verify yet · decay-watch [decay: 1mo]

### Holes in this draft

- **Unit economics** for the audit are modelled, not measured. Conversion rates are targets. 100 audits needed to validate.
- **Voice cost in prod** — we expect drift to ~$0.08–$0.10/min with multimodal LLM cost included.
- **Wedge-tier count** is 46 candidates; the true number is whatever survives sales contact.
- **Pricing for Pro Audit ($499) is untested.**
- **Customer profile** "25–150 person services businesses" needs 20+ paid validations.

### Decay watch — re-run these first

- **1 month:** capability matrix & vendor map (§6, §7).
- **3 months:** failure-mode list (§9).
- **6 months:** process database (§3), wedge tier (§4), anti-list (§5), pricing reality (§17).
- **12 months:** thesis (§1), audit-job spec (§2), role pains (§10), discovery bank (§12), scoring rubric (§13), decision tree (§14), template (§15), red flags (§16). Stable, but re-check after first 100 paid engagements.
- **Stable:** operator ↔ vendor glossary (§11).

## Bibliography

See `research/sources.md` for the full list of cited URLs.

Raw research files live in `research/raw-*.md` for traceability:
- `raw-capability-matrix.md` — capability matrix + vendor map + failure modes
- `raw-glossary-discovery.md` — glossary + discovery bank + red flags
- `raw-processes-ops-fin-sales.md` — 108 processes across Ops, Finance, Sales
- `raw-vertical-and-antilist.md` — 66 vertical processes + 60 anti-list entries
- `raw-role-pains.md` — 20 roles with salaries, headcount, pains, buying authority
- `raw-business-pricing.md` — 39 named shops + pricing benchmarks + sales cycle data

Generated artifacts:
- `processes.json` (206 entries)
- `anti-list.json` (60 entries)
- `role-pains.json` (12 roles)
- `glossary.json` (80 entries)
