AutomationAudit.aiStart audit
/researchInternal strategy doc · v1.0 · 2026-07-28

The brain behind AutomationAudit.

A working map of what can actually be automated in 2026, what bites in production, what operators say versus what they need, and how we turn that into a 10-minute audit. Citations inline. Honest about failure modes. Tagged for decay.

Processes catalogued
200+across 11 functions
Wedge candidates
30+never_automated_yet
Traps in the anti-list
50+don't build these
Cited sources
120+practitioner-first
§01

Strategic thesis

decay: 12mo

The argument in five bullets, before any tables. If only one section survives a rewrite, this is the one.

  • The wedge is the audit itself. A 10-minute, free, AI-run audit that returns a ranked, dollar-quantified, builder-ready list of automations for one business — plus an explicit "don't automate" list. Build/maintain monetizes; the audit hooks. Comparable audits from McKinsey-style firms cost 8 weeks and six figures.
  • We compete on specificity, not capability. The market is saturated with "AI for ops." Our defensible asset is the 200-row, citation-backed process database (§3) plus the 50-trap anti-list. Both signal we know which automations actually ship.
  • The opening move is the never-automated tier. 30+ processes in our database are flagged never_automated_yet — high pain, low awareness, no competitive density. Examples: vendor-renewal alerter, scope-creep tracker on services contracts, customer-health rollup from telemetry+tickets, partner-attribution stitcher. We don't win the first inbound-lead-routing deal; we win the "I didn't know that existed" deal.
  • Pick one offer for v1: Free Audit → $99/mo build & maintain. Subscription is low enough for an ops director's AmEx (no procurement). Per-build pricing ($1.5k–$8k) for users who refuse a recurring relationship. The "Pro Audit" ($499) is a v2 productized SKU for 100–500-person orgs.
  • First customer profile: 25–150 person services businesses.Agencies, dev shops, consultancies, fractional ops. Founder/COO buys. Stack is HubSpot/Pipedrive, Slack, Notion/Asana, Gmail, Stripe. Avoid for v1: enterprise (procurement), regulated (PHI/financial advice), <5-person teams (won't pay).

Why this beats the alternatives

  • vs. agencies that pitch first, scope later: we lead with value (the audit), so the buyer self-qualifies.
  • vs. tools that demand you know what to automate: most operators can't articulate it; the audit does the articulating.
  • vs. RPA / Zapier consultancies: we surface the never-automated tier — deals nobody is competing for.
  • vs. ChatGPT "audit my business": we have a process database, a scoring rubric, an anti-list, and an opinionated decision tree behind the agent.
§02

The audit job — end-to-end engagement

decay: 6mo

The operating spine. Every other section in this doc plugs into one of these six stages.

Six stages

  • 1. Intake (60–180s). Company URL, role, size, primary tools. Mode: voice note (default), AI voice call, or typed. Context priming: 8–12 prompts ranging over week-shape, repeated work, hated tasks, last fire-drill.
  • 2. Discovery (3–5 min). Conditional branching question bank (§12). Function-aware. Each surfaced pain quantified (frequency × duration × loaded hourly cost) and tagged to a candidate from the process database (§3). Hard caps: max 25 candidates, max 12 minutes total.
  • 3. Scoring (~30s, server-side). Each candidate scored on the 7-axis readiness rubric (§13): frequency, time cost, ambiguity, stakes-of-error, data availability, tool-fit, change-mgmt cost. Decision tree (§14) routes each item to: full agent, deterministic workflow, RPA, fix-process-first, or leave-human. Sub-threshold items go to the anti-list explicitly, never silently dropped.
  • 4. Recommendation generation (~60s). Top 5 ranked by dollar_value × readiness × sales_appeal. Each: plain-English explanation, concrete stack, estimated build time, confidence score, similar prior deployments where public. Anti-list section calls out non-recommendations. Headline numbers: total recoverable spend, hours/month.
  • 5. Deliverable (gated). Email gate unlocks full report. Public preview shows top item only. PDF + shareable URL. Each item has three CTAs: DIY (instructions) · We build for $X · We build + maintain for $99/mo.
  • 6. Handoff (sales). Build/subscribe click fires a human-shaped Slack ping with audit + pre-filled proposal. 24-hour SLA. No-pressure path captures email for nurture.

Pricing model — recommendation

Default for v1: Free Audit + $99/mo subscription, with per-build escape hatch ($1.5k–$8k). Subscription is the conversion engine; per-build is the relief valve for buyers who refuse recurring. Comparable: Zapier Experts marketplace per-build rates; emerging "AI ops" subscription shops at $99–$499/mo. zapier.com/experts

Unit economics targets

  • Time per audit:<5 min of human time post-MVP (review + send). 100% agent-run is the target. Free audits are only viable if cost-per-audit is durably <$1.50.
  • Cost per audit:<$1.50 in LLM/voice/infra at scale. Voice (Retell) ~$0.05/min × 8 min + Claude Sonnet 4.5 ~$0.05 per 50k tokens audit transcript & report = ~$0.65 raw. uncited — modelled, not measured.
  • Conversion targets: 4–8% free→$99/mo, 1–3% free→$1.5k+ one-shot.
  • LTV target: $1,800 ($99/mo × 18 mo expected SMB tenure).
  • CAC target:<$200 via SEO + practitioner presence on X.
§03

Process database

decay: 6mo

The 200+ specific, citation-backed automations. Sortable by dollars saved, reliability, sales appeal. Filter by function, status, business size. Default sort favors high-dollar items; flip to status to surface wedges.

Every entry is specific enough that a builder could scope it in one sentence. Status flags: frequently_automated (Zapier/Make territory), rarely_automated (some shops do it, friction exists), never_automated_yet (our wedge tier). Reliability scores carry a ~1.3× haircut relative to vendor claims; see §6 for the rationale. MIT 95%-pilot-failure data

Function

Status

Business size

Min. sales appeal

★ ≥ 1

Max. difficulty

★ ≤ 5
206 of 206 processes $179,093/mo recoverable 2,617 hrs/mo46 wedge candidates
ofs-020

Scope creep tracking on services contracts

Delivery Manager / PM

Ops8$11,200Wedge · never
ofs-017

Shadow IT discovery & cost rationalization

IT Procurement / Finance Partner

Ops6$4,800Rarely
chmle-021

Privilege-aware doc summarizer for litigation

Litigation Counsel

Legal40$4,800Rarely
ofs-002

Vendor renewal alerts before auto-renew

Office Manager / IT Ops Lead

Ops5$4,200Wedge · never
chmle-022

AI PR review pre-pass (style, security, test coverage)

Engineering Manager

Eng30$3,300Frequently
ofs-059

Vendor early-pay discount capture

AP Manager / Treasury

Finance3$3,200Wedge · never
chmle-030

Executive inbox triage with action-tagging

Executive Assistant / Founder

Exec18$2,700Rarely
ofs-097

Voice AI for outbound cold calling

SDR / Founder

Sales60$2,520Rarely
chmle-027

Internal codebase Q&A bot (Cody-style)

Developer Productivity Lead

Eng22$2,420Rarely
chmle-024

On-call alert triage + runbook execution suggestion

SRE / DevOps Lead

Eng20$2,200Rarely
chmle-019

Vendor security questionnaire auto-completion

Compliance / GRC Lead

Legal24$2,160Rarely
ofs-003

Internal IT ticket triage & first-touch response

IT Helpdesk L1

Ops56$2,128Frequently
chmle-032

Cross-tool exec dashboard (financial + product + people)

CEO / CoS

Exec14$2,100Rarely
chmle-020

Regulatory change monitoring + impact briefing

Compliance Counsel

Legal16$1,920Wedge · never
ofs-051

Board pack assembly (financial section)

CFO / FP&A

Finance16$1,840Rarely
ofs-022

Inbound press / partnership / random "interesting" email triage

COS / Founder

Ops16$1,760Rarely
chmle-029

Weekly board pack auto-draft (metrics + commentary)

Chief of Staff

Exec16$1,760Rarely
ofs-084

Account research deep-dive briefs (ABM)

AE / ABM Manager

Sales24$1,728Rarely
ofs-042

Revenue recognition rules for SaaS subscriptions

Revenue Accountant

Finance22$1,716Rarely
chmle-001

QBR deck auto-generation from product usage + CRM data

Customer Success Manager

CS24$1,680Rarely
chmle-005

L1 ticket auto-triage, tag, and KB-grounded draft

Support Lead

Support48$1,680Frequently
chmle-018

NDA & MSA playbook redlines with attorney HITL

Legal Operations Manager

Legal18$1,620Rarely
ofs-092

Demo environment customization per prospect

Sales Engineer

Sales16$1,568Rarely
chmle-025

Test flake quarantine + auto-issue

Test Infra / Eng

Eng14$1,540Wedge · never
ofs-039

Vendor bill OCR + 3-way match

AP Specialist

Finance42$1,512Frequently
ofs-075

Outbound email personalization at scale

SDR / Founder

Sales36$1,512Frequently
ofs-068

AP fraud / duplicate invoice detection

AP / Controller

Finance3$1,500Rarely
chmle-031

Calendar prep briefs (who, why, last touch)

EA / Founder

Exec10$1,500Rarely
ofs-026

Vendor risk reassessment annual cycle

Vendor Risk Manager / Compliance

Ops18$1,476Rarely
vrt-040

Proactive shipment status comms

Operations / customer service rep (CSR) · logistics

Vertical40$1,400Frequently
ofs-050

Audit prep PBC (provided-by-client) list management

Controller

Finance18$1,296Rarely
chmle-002

Customer health score rollup with leading-indicator churn flags

CS Ops Lead

CS18$1,260Wedge · never
chmle-015

Ad creative iteration from winning concepts

Performance Marketer

Marketing18$1,260Rarely
ofs-007

Contract redline coordination & version tracking

Legal Ops / COO

Ops18$1,224Rarely
ofs-041

Monthly close checklist orchestration

Controller / Accounting Manager

Finance18$1,224Rarely
ofs-010

Cross-team status update aggregation

PMO / Chief of Staff

Ops20$1,160Rarely
ofs-018

Customer health score rollup from product telemetry + support tickets

RevOps / CS Ops

Ops16$1,152Wedge · never
ofs-016

Compliance evidence collection for SOC2 / ISO

Security / Compliance Lead

Ops14$1,148Frequently
vrt-050

Document request list (DRL) chase

Senior associate / coordinator · professional_services

Vertical20$1,100Frequently
chmle-012

Internal HR Q&A bot (benefits, policy, PTO)

HR Operations

HR22$1,100Rarely
ofs-044

Cash flow forecasting (13-week rolling)

FP&A / Controller

Finance15$1,080Rarely
vrt-025

Buyer showing scheduling across multiple listings

Agent / ISA · real_estate

Vertical12$1,080Rarely
vrt-049

Conflict check on new matter

Conflicts attorney / partner · professional_services

Vertical8$1,080Frequently
vrt-032

Prior auth submission and status follow-up

PA coordinator · healthcare

Vertical30$1,050Rarely
vrt-033

Denial triage and resubmission

Biller / RCM lead · healthcare

Vertical30$1,050Rarely
vrt-044

Driver check-in calls / status capture

Dispatcher · logistics

Vertical30$1,050Rarely
chmle-004

Renewal forecasting 90/60/30 day playbook

Customer Success Director

CS12$1,050Rarely
ofs-077

Meeting prep brief generation

AE / Manager

Sales18$1,044Rarely
ofs-004

Employee onboarding asset provisioning

IT Ops / People Ops

Ops24$1,008Frequently
vrt-003

Statement of Work generation from sales call

Founder / new biz lead · agencies

Vertical10$1,000Rarely
vrt-010

First-pass RFP response from past wins

New biz / founder · ecommerce

Vertical10$1,000Wedge · never
chmle-008

Recruiter outreach personalization at scale

Recruiter

Recruiting20$1,000Frequently
chmle-010

Onboarding day-1-to-30 task orchestrator

People Ops Manager

HR16$992Rarely
ofs-060

Intercompany reconciliation (multi-entity)

Controller / GL Accountant

Finance12$984Rarely
ofs-079

Deal coaching from call transcripts

Sales Manager

Sales12$984Rarely
chmle-009

Interview scheduling across panel + candidate availability

Recruiting Coordinator

Recruiting28$980Frequently
chmle-014

Lifecycle email triggers from product behavior

Lifecycle Marketing Manager

Marketing14$980Frequently
ofs-089

CRM data hygiene (missing fields, bad data)

RevOps

Sales14$952Rarely
ofs-001

Vendor onboarding packet collection

Procurement / Vendor Manager

Ops22$924Rarely
vrt-052

Monthly close + advisory package

Senior accountant · professional_services

Vertical10$900Rarely
ofs-045

Budget vs. actual variance flagging

FP&A Analyst

Finance13$884Rarely
vrt-005

KPI dashboard build/refresh for client portal

Analyst · agencies

Vertical16$880Frequently
chmle-013

SEO content brief generator from SERP + competitor scrape

Content Strategist

Marketing16$880Frequently
chmle-017

Repurposing long-form video → 10 short-form clips

Content Producer

Marketing22$880Frequently
vrt-011

Where is my order" support ticket resolution

Support lead · ecommerce

Vertical25$875Frequently
vrt-027

Tenant maintenance request triage and dispatch

Maintenance coordinator · real_estate

Vertical25$875Rarely
vrt-031

Pre-visit insurance + benefits verification

Front desk / billing · healthcare

Vertical25$875Rarely
vrt-043

Customer rate quote response

Sales / pricing analyst · logistics

Vertical25$875Rarely
ofs-009

Weekly all-hands KPI rollup

COS / Operations Manager

Ops14$868Rarely
ofs-076

Outbound follow-up sequencing

SDR / AE

Sales18$864Frequently
ofs-081

Churn risk flagging for sales-led GTM

CSM / AE

Sales12$864Rarely
ofs-103

Slack channel for shared customer (Slack Connect)

AE / CSM

Sales12$864Wedge · never
chmle-007

Voice-of-customer rollup from tickets + surveys + reviews

CX Lead

Support12$840Rarely
ofs-011

Exception monitoring for ops dashboards

Operations Analyst

Ops16$832Rarely
ofs-080

Pipeline hygiene audit

RevOps / Sales Manager

Sales10$820Rarely
ofs-014

BD / partnership pipeline reporting

Head of Partnerships / BD Ops

Ops12$816Wedge · never
ofs-052

FP&A model refresh from source systems

FP&A Analyst

Finance12$816Rarely
ofs-008

Internal wiki staleness detection

COS / Knowledge Manager

Ops14$812Wedge · never
ofs-078

Post-call CRM update from recording

AE

Sales14$812Frequently
ofs-038

AR follow-up cadence on overdue invoices

AR Specialist / Controller

Finance19$798Rarely
ofs-012

Capacity planning for services teams

Resource Manager / Delivery Lead

Ops11$792Wedge · never
ofs-073

Lead enrichment from form fill

SDR / RevOps

Sales18$756Frequently
ofs-098

Inbound voice qualification (chatbot/voicebot)

SDR / Marketing

Sales18$756Rarely
ofs-082

Quote / proposal generation

AE / Sales Engineer

Sales12$744Rarely
ofs-083

Contract redline negotiation responses

AE / Legal / RevOps

Sales9$738Rarely
ofs-090

Sales tax / pricing approval workflow

Deal Desk / RevOps

Sales10$720Rarely
vrt-022

CMA report generation for seller meeting

Agent · real_estate

Vertical8$720Rarely
vrt-051

Daily time entry reconstruction

Associate / consultant · professional_services

Vertical8$720Rarely
chmle-023

PR description + changelog generation from diff

Engineering IC

Eng8$720Frequently
chmle-026

Dependency update PRs with safety gates

Engineering

Eng8$720Rarely
vrt-036

Digital pre-visit intake (history, insurance, consents)

Front desk · healthcare

Vertical20$700Frequently
vrt-041

Proof of delivery collection from carriers

Settlement clerk · logistics

Vertical20$700Rarely
chmle-006

KB article auto-generation from resolved tickets

Support Ops

Support14$700Wedge · never
ofs-067

Treasury — daily cash position consolidation

Treasurer / Controller

Finance9$675Rarely
ofs-040

Expense receipt categorization from photos

Controller / AP

Finance16$672Frequently
vrt-001

Weekly client status report assembly

Account manager · agencies

Vertical12$660Frequently
vrt-002

Time entry compliance and reconstruction

Ops manager / agency owner · agencies

Vertical12$660Rarely
vrt-004

Monthly social/blog content calendar with hooks + topics

Content strategist · agencies

Vertical12$660Rarely
vrt-014

Meta/TikTok creative testing rotation + winner-killer logic

Media buyer · ecommerce

Vertical12$660Rarely
vrt-016

Inventory replenishment alerting

Ops manager · ecommerce

Vertical12$660Frequently
chmle-028

Incident postmortem first draft

Engineering Manager

Eng6$660Rarely
ofs-101

Renewal & expansion forecast for sales-assisted CS

CSM / Account Manager

Sales9$648Rarely
ofs-005

Employee offboarding access revocation audit

IT Security / People Ops

Ops12$624Rarely
ofs-019

Internal data hygiene drift detection

RevOps / Data Analyst

Ops10$620Wedge · never
ofs-062

Subscription billing reconciliation (CRM vs. Stripe)

RevOps / Controller

Finance10$620Wedge · never
chmle-016

Brand mention monitoring with sentiment + auto-response queue

Social Media Manager

Marketing12$600Frequently
ofs-088

Reply triage from outbound sequences

SDR / AE

Sales14$588Frequently
ofs-086

Win/loss analysis from CRM + call data

RevOps / Product Marketing

Sales8$576Wedge · never
vrt-020

Cold outreach to UGC creators + deliverable tracking

Influencer manager · real_estate

Vertical16$560Rarely
vrt-037

Outbound referral tracking + records request

Referral coordinator · healthcare

Vertical16$560Wedge · never
chmle-003

Expansion-signal detection from feature adoption + seat growth

Account Manager

CS8$560Wedge · never
ofs-070

Procurement card reconciliation across multiple GLs

AP / Controller

Finance13$546Frequently
ofs-074

Outbound list building from ICP

SDR / Growth

Sales13$546Frequently
vrt-021

MLS listing description + headline

Listing agent / coordinator · real_estate

Vertical6$540Rarely
vrt-030

Fair Housing + MLS compliance check on draft listings

Compliance / managing broker · healthcare

Vertical6$540Wedge · never
vrt-053

Filing deadline calendar for regulated clients

Compliance / paralegal · professional_services

Vertical6$540Wedge · never
vrt-054

Firm-wide precedent / past-work retrieval

All knowledge workers · professional_services

Vertical6$540Rarely
vrt-012

Return request triage and authorization

Support manager · ecommerce

Vertical15$525Frequently
vrt-013

Long-form product description from PDP brief

Merchant/content · ecommerce

Vertical15$525Frequently
vrt-026

Buyer/seller closing document chase

Transaction coordinator (TC) · real_estate

Vertical15$525Rarely
vrt-034

No-show prevention + same-day waitlist fill

Scheduler · healthcare

Vertical15$525Frequently
vrt-042

Post-delivery invoice audit vs quote

AP / freight auditor · logistics

Vertical15$525Frequently
vrt-045

Customer-facing ETA on last-mile delivery

Dispatch / CS · logistics

Vertical15$525Frequently
vrt-057

Daily route optimization for technicians

Dispatcher · local_services

Vertical15$525Frequently
vrt-064

Daily crew confirmation + replacement when out

Operations · local_services

Vertical15$525Wedge · never
ofs-025

New client kickoff packet generation

Project Manager / CS Lead

Ops9$522Rarely
ofs-032

Internal expense policy enforcement

Controller / People Ops

Ops8$496Rarely
ofs-034

Recurring-meeting agenda hygiene

COS / Manager

Ops8$496Wedge · never
ofs-023

Time tracking enforcement for billable teams

Delivery Ops / Controller

Ops10$480Rarely
ofs-069

SaaS metrics calculation (NRR, GRR, LTV, CAC)

FP&A / RevOps

Finance7$476Frequently
ofs-094

Partner referral attribution

Partnerships / RevOps

Sales7$476Wedge · never
ofs-096

NPS / CSAT response triage to sales actions

CSM / RevOps

Sales8$464Wedge · never
ofs-053

Equity / cap table waterfall scenario modeling

CFO / Founder

Finance4$460Rarely
vrt-024

Long-term nurture for warm but not ready buyers

Agent · real_estate

Vertical5$450Frequently
vrt-048

New client engagement letter

Partner / senior · professional_services

Vertical5$450Frequently
vrt-007

Real-time retainer hour tracking + scope-creep flagging

Project manager · agencies

Vertical8$440Rarely
ofs-085

Mutual action plan (MAP) generation & tracking

AE

Sales7$434Wedge · never
ofs-043

Weekly Stripe revenue rollup into Notion + Slack

Founder / Controller

Finance5$425Rarely
ofs-054

Bank reconciliation

Bookkeeper / Controller

Finance10$420Frequently
vrt-006

Creative asset versioning and handoff to media buyer

Producer · agencies

Vertical12$420Rarely
vrt-023

Buyer/seller lead routing + first-touch

Team lead · real_estate

Vertical12$420Frequently
vrt-029

Post vacancy to 20+ rental sites

Leasing · real_estate

Vertical12$420Frequently
vrt-038

Patient-responsibility AR follow-up

Biller · healthcare

Vertical12$420Rarely
vrt-055

Vendor invoice approval and posting

Controller / bookkeeper · local_services

Vertical12$420Frequently
vrt-056

Inbound lead capture + first response

Owner / dispatcher · local_services

Vertical12$420Rarely
chmle-011

Performance review prep aggregator (PRs, 1:1s, peer notes)

HRBP

HR6$420Rarely
ofs-013

Meeting scheduling across multiple parties

EA / Office Manager

Ops13$416Frequently
ofs-058

R&D tax credit documentation collection

Controller / Tax Advisor

Finance5$410Rarely
ofs-029

Service-level SLA breach detection

Customer Success / Support Ops

Ops7$406Rarely
ofs-105

Customer reference matching for sales calls

AE / Customer Marketing

Sales7$406Rarely
ofs-006

Laptop & monitor asset tracking

IT Ops / Office Manager

Ops9$378Rarely
ofs-064

Employee corporate card spend review

Controller / Manager

Finance6$372Rarely
ofs-095

Multi-thread champion tracking

AE / RevOps

Sales6$372Wedge · never
ofs-036

Procurement spend categorization

Procurement Analyst

Ops7$364Rarely
ofs-063

Customer prepayment / deferred revenue tracking

Revenue Accountant

Finance5$360Rarely
vrt-028

Rent reminder/late notice cadence

Property manager · real_estate

Vertical10$350Frequently
vrt-035

Patient recall / recare outreach

Front desk · healthcare

Vertical10$350Frequently
vrt-039

Provider credentialing renewals + payer enrollment

Credentialing coordinator · logistics

Vertical10$350Wedge · never
vrt-047

New carrier onboarding for broker

Carrier rep · professional_services

Vertical10$350Frequently
vrt-059

On-site quote and invoice generation

Technician · local_services

Vertical10$350Rarely
vrt-063

Permit-ready job photo + doc packet

Project coordinator · local_services

Vertical10$350Wedge · never
ofs-028

Internal NPS / pulse survey loop

People Ops

Ops6$348Rarely
ofs-031

NDA / MSA template generation for new prospects

Sales Ops / Legal

Ops6$348Frequently
ofs-037

Internal reference check for hiring

Recruiter / Hiring Manager

Ops6$348Wedge · never
ofs-047

Sales tax nexus monitoring & filing prep

Controller / Accountant

Finance6$348Frequently
ofs-066

Customer pricing tier audit & uplift opportunity

RevOps / Finance Partner

Finance5$340Wedge · never
vrt-008

New-project kickoff brief from CRM + discovery

PM · agencies

Vertical6$330Rarely
ofs-057

Equity refresh grants & vesting acceleration tracking

People Ops / CFO

Finance4$328Wedge · never
ofs-055

Foreign exchange revaluation

Controller / Treasury

Finance4$312Rarely
ofs-065

Statement-of-cash-flows (indirect method) prep

Controller / FP&A

Finance4$312Rarely
ofs-100

Competitor mention alerts from review sites

Product Marketing / CI

Sales5$310Rarely
ofs-102

Discovery call agenda customization

AE

Sales5$310Wedge · never
ofs-107

Champion enablement asset delivery

AE / Customer Marketing

Sales5$310Wedge · never
ofs-056

Customer payment failure recovery

AR / RevOps

Finance7$294Frequently
ofs-027

Conference / event ROI tracking

Field Marketing / Ops

Ops5$290Wedge · never
ofs-046

Payroll variance check vs. prior period

Payroll Manager / Controller

Finance5$290Wedge · never
ofs-087

ICP scoring on inbound leads

RevOps / Marketing Ops

Sales5$290Rarely
ofs-071

Customer-by-customer profitability analysis

FP&A

Finance4$288Wedge · never
vrt-009

Hours-to-invoice reconciliation

Bookkeeper/ops · agencies

Vertical8$280Rarely
vrt-018

Per-customer post-purchase flow personalization

Lifecycle / CRM marketer · ecommerce

Vertical8$280Rarely
vrt-046

Damage / loss claim filing with carriers

Claims clerk · logistics

Vertical8$280Wedge · never
vrt-058

Technician ETA + on-the-way SMS

Tech / dispatcher · local_services

Vertical8$280Frequently
vrt-061

Maintenance contract renewal + seasonal reminders

Office manager · local_services

Vertical8$280Frequently
ofs-099

SDR-to-AE handoff brief

SDR / AE

Sales5$260Wedge · never
ofs-049

Customer credit memo issuance

AR / Controller

Finance6$252Rarely
ofs-061

Equity-based comp accrual (ASC 718)

Revenue/Equity Accountant

Finance3$246Rarely
ofs-072

Lease accounting (ASC 842)

Controller / Lease Accountant

Finance3$246Frequently
ofs-091

Trigger-based outreach (job change, funding, hiring)

SDR / Growth

Sales5$240Rarely
ofs-033

Holiday / leave coverage planning

Manager / People Ops

Ops4$232Wedge · never
ofs-024

Real estate / office maintenance ticket coordination

Facilities / Office Manager

Ops6$216Rarely
ofs-108

Lost-deal interview scheduling & analysis

Product Marketing

Sales3$216Wedge · never
vrt-015

Review response across review sites

Brand/CX manager · ecommerce

Vertical6$210Frequently
vrt-017

Fraud review on flagged orders before fulfillment

Ops manager · ecommerce

Vertical6$210Frequently
vrt-062

Parts inventory reorder

Warehouse / ops · local_services

Vertical6$210Rarely
vrt-065

COI + W9 collection from subs

Office manager · local_services

Vertical6$210Rarely
vrt-066

We sent you a quote 5 days ago" follow-up

Owner / sales

Vertical6$210Rarely
ofs-021

PTO conflict detection

People Ops / Manager

Ops4$208Wedge · never
ofs-015

Office supplies & snack reorder

Office Manager

Ops6$192Frequently
ofs-093

Lost deal re-engagement campaigns

Marketing / RevOps

Sales3$174Wedge · never
ofs-104

Pricing change communication to existing customers

AE / CS / Marketing

Sales3$174Wedge · never
ofs-030

Insurance certificate (COI) collection from vendors

Risk / Procurement

Ops4$168Wedge · never
ofs-048

1099 contractor compliance & filing

Controller / AP

Finance3$145Frequently
ofs-106

ICP refinement from closed-won analysis

RevOps / Product Marketing

Sales2$144Wedge · never
vrt-019

Quarterly Shopify app audit + consolidation

COO / ops · ecommerce

Vertical4$140Wedge · never
vrt-060

Post-job review request

Owner · local_services

Vertical4$140Frequently
ofs-035

Domain renewal & DNS monitoring

IT Ops / Marketing Ops

Ops1$58Rarely
§04

Wedge — top processes & never_automated tier

decay: 6mo

The 10 deals to lead with, plus the wedge tier of pains nobody else is solving yet.

Top 10 highest-ROI starting processes

Ranked by cost_saved × sales_appeal × reliability_today. These are the deals that should populate cold-outbound proof points and landing page case studies first.

#ProcessFunction$/moRelSell
01Scope creep tracking on services contractsOps$11,20035
02Shadow IT discovery & cost rationalizationOps$4,80045
03Vendor renewal alerts before auto-renewOps$4,20045
04AI PR review pre-pass (style, security, test coverage)Eng$3,30045
05Privilege-aware doc summarizer for litigationLegal$4,80034
06Executive inbox triage with action-taggingExec$2,70045
07Vendor early-pay discount captureFinance$3,20044
08Vendor security questionnaire auto-completionLegal$2,16045
09Cross-tool exec dashboard (financial + product + people)Exec$2,10045
10Internal codebase Q&A bot (Cody-style)Eng$2,42044

The wedge tier — never_automated_yet

These are the processes operators complain about, but no vendor or agency has packaged a clean solution for. Low competitive density. Cheap to acquire customers. Hard for a competitor to copy without our process database.

  • Vendor renewal alerts before auto-renewNever get blindsided by a $40k auto-renewal again.
  • Internal wiki staleness detectionStop pointing new hires at lies.
  • Capacity planning for services teamsKnow in week 2 you'll be underwater in week 6, not week 5.
  • BD / partnership pipeline reportingPartnerships are an afterthought because reporting is. Fix the reporting.
  • Customer health score rollup from product telemetry + support ticketsKnow which account will churn this quarter — before the cancel email.
  • Internal data hygiene drift detectionYour dashboard isn't lying to you anymore.
  • Scope creep tracking on services contractsStop eating $20k of scope creep per project.
  • PTO conflict detectionDon't be the manager who let 3 PMs go to Italy the same week.
  • Conference / event ROI trackingKnow which conferences actually drove revenue, not just leads.
  • Insurance certificate (COI) collection from vendorsCOIs stop being a fire drill before a big event.
  • Holiday / leave coverage planningDecember coverage planning is 10 minutes, not a panic.
  • Recurring-meeting agenda hygieneMeetings get useful again.
§05

Anti-list — don't automate this

decay: 6mo

50+ processes that look automatable but bite. Filter by failure mode. Selling the anti-list is a credibility flex no vendor offers.

Failure modes draw on real production incidents: Air Canada chatbot making up refund policy BBC ; Klarna walking back AI-first support Customer Experience Dive ; NYC Local Law 144 on automated employment decision tools NYC DCWP ; CVE-2025-59944 prompt-injection-to-RCE in Cursor mayhemcode; OpenAI dropping SWE-bench Verified for benchmark contamination morphllm.

60 of 60 traps

anti-001Customer support, allJudgment

Fully automated customer refunds / chatbot refund policy

Why it tempts you

Refunds are repetitive and seem rule-based

Concrete failures

regulation, judgment Air Canada chatbot invented a bereavement-fare refund window; BC Civil Resolution Tribunal held Air Canada liable for negligent misrepresentation by AI; $650 CAD damages (https://www.cbc.ca/news/canada/british-columbia/air-canada-chatbot-lawsuit-1.7116416, https://www.americanbar.org/groups/business_law/resources/business-law-today/2024-february/bc-tribunal-confirms-companies-remain-liable-information-provided-ai-chatbot/, Feb 2024). Tribunal explicitly rejected "chatbot is a separate entity" defense. Cursor support bot fabricated a new "one-device login" policy in April 2025; customers cancelled subscriptions; company refunded and apologized publicly (https://www.theregister.com/2025/04/18/cursor_ai_support_bot_lies/, https://fortune.com/article/customer-support-ai-cursor-went-rogue/).

Right split

AI drafts response + cites policy URL; agent (or rule engine on narrow, well-bounded refund types like "unworn item < 30 days") approves payment. Refund authority above $X never automatic.

anti-002AllJudgment

Full email reply automation without HITL

Why it tempts you

Inbox is the timesuck owners complain about most

Concrete failures

judgment, brand-voice, hallucination, relationship-dependent DPD chatbot, after Jan 2024 update, swore at customer, wrote a poem mocking DPD, called itself the worst delivery firm in the world; 1.3M views before disabled (https://time.com/6564726/ai-chatbot-dpd-curses-criticizes-company/). Klarna replaced 700 CS agents with AI, then in 2025 admitted "we went too far," now rehiring; CEO public statement (https://www.entrepreneur.com/business-news/klarna-ceo-reverses-course-by-hiring-more-humans-not-ai/491396, https://fortune.com/2025/05/09/klarna-ai-humans-return-on-investment/).

Right split

AI drafts in agent inbox; agent edits 30s and sends. Auto-send only for `intent=order_status` lookups with order data validated.

anti-003B2B / B2C salesRegulation

AI cold sales calls from outbound dialers

Why it tempts you

Looks like 100x SDR leverage

Concrete failures

regulation, brand FCC Declaratory Ruling Feb 8, 2024 confirms TCPA's "artificial or prerecorded voice" prohibition includes AI-generated voices; prior express consent required (https://www.fcc.gov/document/fcc-makes-ai-generated-voices-robocalls-illegal). Civil penalties up to $1,500 per call per recipient. Class action exposure under TCPA is severe — single campaign can hit eight-figure damages.

Right split

AI for inbound voice qualification, opted-in re-engagement, internal calls (recruiter to candidate after consent). Cold outbound stays human or written channels.

anti-004Community / marketplaces / socialJudgment

AI moderation at scale without escalation paths

Why it tempts you

Massive volume, repetitive

Concrete failures

judgment, low-frequency-high-stakes, regulation DSA in EU mandates human review options for content decisions affecting users; pure AI moderation breaches Art 14/20 process rights. Repeated Meta/YouTube false-positives banning veterans accounts, breast cancer support groups, etc. — recurring brand and PR risk.

Right split

AI as first triage and high-confidence auto-action on clear violations (CSAM, spam); human reviews any escalations and user appeals within 24h.

anti-005E-comm paymentsJudgment

Stripe dispute responses fully automated

Why it tempts you

Dispute evidence packets are structured

Concrete failures

judgment, regulation, cost Stripe Smart Disputes (their own AI tool) takes 30% of recovered amount and only beats manual on small disputes (https://directpaynet.com/stripe-forcing-ai-dispute-tool-taking-30-of-winnings/, 2025). Manual dispute response has <20% win rate; bad automated submissions can lock you out of resubmitting (https://www.chargeflow.io/blog/stripe-dispute-fees-2025). Fraud signaling that turns out to be "friendly fraud" needs nuance no LLM has yet.

Right split

AI assembles evidence packet + draft narrative; ops manager reviews for high-value disputes; auto-submit only for low-value, low-complexity friendly fraud where evidence is air-tight.

anti-006PR / corp commsJudgment

AI press release / crisis comms drafting (auto-publish)

Why it tempts you

Templated, low-frequency, looks safe

Concrete failures

low-frequency-high-stakes, judgment, brand One factual error in a crisis statement compounds 100x; cleanup cost > entire annual comms budget. LLMs frequently hallucinate exec quotes and unverified incident details.

Right split

AI drafts FAQ + holding statement template against a `crisis_playbook.md`; PR lead and legal sign before publish, always.

anti-007HR / TAJudgment

AI hiring decisions (resume screening + auto-reject)

Why it tempts you

1,000 resumes per req, screening is a slog

Concrete failures

regulation (Title VII, ADEA, ADA, NYC LL144, EU AI Act, Illinois AI Video Interview Act, Colorado AI Act), judgment Mobley v. Workday: certified as nationwide ADEA collective action May 2025; Workday admitted 1.1B applications rejected by tool in scope period (https://www.lawandtheworkplace.com/2025/06/ai-bias-lawsuit-against-workday-reaches-next-stage-as-court-grants-conditional-certification-of-adea-claim/, https://www.fisherphillips.com/en/insights/insights/discrimination-lawsuit-over-workdays-ai-hiring-tools-can-proceed-as-class-action-6-things). iTutorGroup paid $365K to settle EEOC's first AI-discrimination case; software auto-rejected applicants over 55/60 based on birth date (https://www.eeoc.gov/newsroom/itutorgroup-pay-365000-settle-eeoc-discriminatory-hiring-suit, Aug 2023). NYC Local Law 144 since July 2023 requires bias audits + candidate notice; $500-$1,500/violation/day (https://www.nyc.gov/site/dca/about/automated-employment-decision-tools.page). EU AI Act classifies HR AI as high-risk; full requirements enforceable Aug 2, 2026; fines up to €35M or 7% global turnover (https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai).

Right split

AI extracts and standardizes resume data + ranks by stated criteria; humans make every interview/reject decision. Audit logs of every AI input/output. Notice to candidates per NYC LL144 + EU AI Act.

anti-008HRJudgment

AI performance reviews

Why it tempts you

Managers procrastinate on reviews; AI seems "objective

Concrete failures

judgment, relationship-dependent, regulation EU AI Act high-risk: "AI systems used to evaluate workers' performance" explicitly covered (Annex III) -> mandatory risk assessment, human oversight, bias testing. California AB 2930 and Colorado AI Act both name performance-management as covered. Empirical: AI summaries based on Slack/email volume penalize quieter contributors, ESL workers, and parents on flex schedules — disparate impact risk.

Right split

AI aggregates ticket counts, code review feedback, peer comments into a tab; manager owns the narrative and decision. Never auto-rate.

anti-009HRJudgment

AI for sensitive HR conversations (PIPs, terminations)

Why it tempts you

Managers want a script

Concrete failures

regulation, judgment, relationship-dependent WARN Act notices need legal-grade precision; Title VII / ADA / FMLA traps in seemingly neutral language. Recorded AI conversations with employees raise wiretap concerns (two-party consent states).

Right split

AI prepares HR talking-point checklist for manager pre-meeting; conversations are always human. AI never present in the room or scripting in real time.

anti-010Sales / TARegulation

AI lead/talent screening without feedback loop

Why it tempts you

Score everyone, focus reps on top

Concrete failures

data-quality, regulation Model drift after 6 months on stale CRM labels; reps stop trusting it, ROI vanishes. Reverse causation: high-scored leads get more rep effort -> appear to convert better -> model self-validates a false signal.

Right split

AI suggests; rep accepts/rejects with reason; reason-codes feed weekly retrain. Never auto-suppress leads from human visibility.

anti-011Law / contractsJudgment

AI legal redlines without attorney review

Why it tempts you

NDA review is repetitive

Concrete failures

regulation (UPL), low-frequency-high-stakes, judgment Damien Charlotin's hallucination database tracks 1,348+ documented cases of AI-fabricated citations in court filings by mid-2026 (https://www.damiencharlotin.com/hallucinations/). Stanford RegLab found Lexis+ AI hallucinated >17%, Westlaw AI-Assisted Research >34% of the time (https://hai.stanford.edu/news/ai-trial-legal-models-hallucinate-1-out-6-or-more-benchmarking-queries). ABA Formal Opinion 512 (July 29, 2024) imposes verification duty on lawyers.

Right split

AI flags clauses against firm playbook + drafts proposed redlines; attorney must read and sign every edit. Never send out a redline the bot wrote untouched.

anti-012Procurement / finance / RevOpsJudgment

Auto-renewals of vendor / customer contracts

Why it tempts you

Reduces churn / spend creep

Concrete failures

change-management, judgment, regulation (state auto-renewal laws — CA SB-313, NY GBL 5-903) California, NY, Illinois, and others require explicit re-disclosure before auto-renewal; FTC Click-to-Cancel rule pending. Auto-renewing without proper notice is grounds for restitution. Vendor cost creep: SaaS bill grows 15-20%/yr if nobody pushes back.

Right split

AI surfaces upcoming renewals 60/30/14 days out with usage data and proposed action; ops decides renew/renegotiate/cancel.

anti-013Compliance / globalJudgment

AI auto-translation of legal / medical content

Why it tempts you

Translation is "easy

Concrete failures

regulation, low-frequency-high-stakes, judgment HHS Section 1557 (ACA) requires qualified medical interpreters; FDA labeling errors trigger recalls. EU MDR / IVDR require accuracy attestations on translated medical device IFUs.

Right split

AI first draft + glossary enforcement; certified linguist signs off. Never auto-publish in regulated domains.

anti-014MarketingRegulation

Generated marketing claims without legal review (FTC)

Why it tempts you

Scale ad creative with LLMs

Concrete failures

regulation FTC's Operation AI Comply (launched Sept 2024) has brought 12+ enforcement actions in 2025 against AI-washing and unsubstantiated AI-product claims (https://www.ftc.gov/news-events/news/press-releases/2024/09/ftc-announces-crackdown-deceptive-ai-claims-schemes). DoNotPay fined $193K for unsubstantiated AI claims; Workado settled for "98% accuracy" claim it could not substantiate.

Right split

AI drafts ad copy with no superlative/health/earnings claims allowed; legal reviews any claim invoking safety, efficacy, or earnings.

anti-015HealthcareRegulation

AI medical advice / clinical decision support without FDA review

Why it tempts you

Tempting to use LLM in patient-facing chat

Concrete failures

regulation (FDA SaMD), low-frequency-high-stakes, privacy FDA's December 2024 draft guidance on AI-enabled device software functions; "clinical decision support" function may require 510(k). HIPAA + state telehealth laws bar unlicensed practice.

Right split

AI scheduling, intake, FAQ, post-visit summaries — no diagnosis, no medication advice, no triage. Always escalate clinical questions to clinician.

anti-016AccountingJudgment

Bookkeeping reconciliation fully automated (no CPA gate)

Why it tempts you

Bank feeds + GPT can categorize most transactions

Concrete failures

regulation, data-quality, judgment Mis-categorizations compound; year-end tax cleanup costs 3-5x what catching them monthly would have. Sales tax nexus questions need CPA judgment.

Right split

AI categorizes routine transactions (matching prior patterns) + flags uncertain ones; bookkeeper reviews weekly; CPA signs financials.

anti-017Tax / accountingRegulation

Tax preparation fully automated for non-trivial returns

Why it tempts you

1040EZ is solved; why not Schedule C?

Concrete failures

regulation (Circular 230), low-frequency-high-stakes IRS Circular 230 imposes preparer due diligence; Penalty $635/return for unreasonable positions. State-specific items (e.g., NY decoupling from federal bonus depreciation) trip LLMs.

Right split

AI pre-populates Schedule entries from intake docs; CPA reviews and signs. Never auto-file.

anti-018Privacy / RevOpsRegulation

Full data deletion / GDPR right-to-erasure agent

Why it tempts you

DSAR volume is annoying and growing

Concrete failures

regulation, auditability, vendor-lock GDPR Art 17 and CCPA both require evidence-of-completion; California Delete Act mandates triennial independent audit from Jan 2028 (https://www.didomi.io/blog/california-delete-act). Agents that "delete" but skip backup tiers, archived emails, ML training datasets — leave non-compliant residual data.

Right split

AI orchestrates deletion workflow across systems and produces evidence pack; privacy officer reviews and signs deletion certificate before reply to data subject.

anti-019PR / executiveJudgment

Crisis comms auto-draft and send

Why it tempts you

Speed matters in crises

Concrete failures

low-frequency-high-stakes, judgment Inaccurate "facts" in initial statement become the story (e.g., wrong casualty figures, wrong fault attribution).

Right split

AI prepares holding-statement template + journalist Q&A docs based on factual brief; CEO/PR/legal approve every word that exits.

anti-020E-comm / SaaSJudgment

Pricing changes via agent

Why it tempts you

Dynamic pricing seems like easy margin

Concrete failures

cost, judgment, regulation Wendy's "surge pricing" 2024 PR backlash forced retraction within days. Amazon's 2011 fly-genetics book priced at $23M by competing bots (no humans noticing). Some EU member states bar personalized pricing without disclosure.

Right split

AI surfaces pricing-experiment candidates with confidence intervals; revenue lead approves discrete price changes. Never autonomous on production prices.

anti-021Sales / lead genRegulation

Cold call dialers with AI script + auto-record

Why it tempts you

Looks like cheap pipeline

Concrete failures

regulation TCPA + FCC ruling (see A3) -> AI voice without express consent is illegal. Two-party-consent state wiretap laws (CA, FL, IL, MA, MD, MT, NH, PA, WA) -> auto-recording without disclosure exposes private rights of action.

Right split

AI for opted-in re-engagement and inbound qualification only. All outbound voice still requires consent.

anti-022Finance / IRJudgment

Investor update drafting without exec review

Why it tempts you

Monthly cadence, structured data

Concrete failures

judgment, low-frequency-high-stakes, regulation (securities) "Forward-looking statements" by LLM that don't accurately represent material facts create securities exposure. Investors notice tone shifts; auto-drafted updates damage trust.

Right split

AI compiles metrics deck and drafts narrative; CEO/CFO rewrite tone and approve.

anti-023ProcurementJudgment

Procurement decisions / RFP scoring

Why it tempts you

RFP scoring matrices feel objective

Concrete failures

judgment, vendor-lock, low-frequency-high-stakes LLM under-weights soft factors (vendor stability, cultural fit, post-sale support). Public-sector procurement frequently restricts AI in award decisions.

Right split

AI normalizes RFP responses, flags gaps and red flags; procurement committee scores and decides.

anti-024Real estate (commercial)Judgment

Lease term negotiation

Why it tempts you

Lease docs are templated

Concrete failures

judgment, relationship-dependent LL/T markets are highly local; what looks like a standard CAM provision is wildly different in Manhattan vs. suburban Atlanta. Concessions are negotiated based on relationships and market intel an LLM lacks.

Right split

AI compares draft to firm playbook + market comps; broker negotiates and partner signs.

anti-025Customer success / SaaSJudgment

Customer health scores driving auto-actions

Why it tempts you

Predict churn before it happens

Concrete failures

data-quality, judgment Stale CRM usage data + reverse causation generate false-positive churn flags; "save" outreach actually triggers reflection -> churn. High-value happy customers flagged red because they switched to a new champion (no login = "at-risk").

Right split

AI flags at-risk + suggests intervention; CSM owns the conversation and decides actions.

anti-026Payments / StripeRegulation

Auto-clawback / chargeback retry

Why it tempts you

Looks like recovered revenue

Concrete failures

regulation, cost Repeated retries on declined cards can trigger card-issuer flags + Stripe risk-score penalties.

Right split

AI determines optimal retry windows within Stripe Smart Retries policy; manual review on high-value or repeat failures.

anti-027EngineeringJudgment

Full code commits without human review (agent shipping to main)

Why it tempts you

Cursor / Devin / Copilot Agent feel close

Concrete failures

judgment, low-frequency-high-stakes, data-quality Cursor own incident — its support bot fabricated company policy (April 2025) (https://www.theregister.com/2025/04/18/cursor_ai_support_bot_lies/). GitClear research 2024 found AI-assisted commits show 41% increase in code churn (rewrites within 2 weeks).

Right split

AI drafts PR; human reviews and merges. Auto-merge only for explicit allow-list (dependabot, doc typos, generated SDKs).

anti-028DevOps / SREJudgment

Auto-rollback of production deploys via AI

Why it tempts you

Faster MTTR

Concrete failures

judgment, low-frequency-high-stakes False-positive metric spikes (CDN cache flush, marketing campaign) trigger rollback during a legitimate launch -> revenue impact.

Right split

AI suggests rollback with evidence + page on-call; engineer executes after 1-look.

anti-029SecurityJudgment

AI security incident response (autonomous blocking)

Why it tempts you

SOAR vendors pitch this

Concrete failures

judgment, change-management Autonomous IP blocks knock real customers off; auto-disabling user accounts on misclassified phishing triggers outage tickets.

Right split

AI enriches alerts + drafts containment plan; analyst executes. Auto-block only narrow, high-confidence rules (known-bad IPs).

anti-030Real estateRegulation

AI listing description without fair-housing audit (real estate)

Why it tempts you

Saves agent hours

Concrete failures

regulation (Fair Housing Act) LLMs slip in "safe neighborhood," "family-friendly," "walking distance to church" — Fair Housing red flags. HUD complaints lead to brokerage fines + agent license risk.

Right split

AI drafts + rule-based scanner flags protected-class language; broker compliance reviews before MLS publish.

anti-031Insurance / fintechRegulation

Auto-pricing of insurance / loan products

Why it tempts you

ML pricing is the industry's whole thing

Concrete failures

regulation (Fair Lending, anti-discrimination), data-quality CFPB and state DOIs require demonstrable non-discrimination; proxy variables (zip code) trigger disparate impact. Colorado SB21-169 specifically targets insurance algorithmic discrimination.

Right split

Actuaries + compliance own pricing; AI used for risk scoring within audited model framework.

anti-032Freight brokerageJudgment

Autonomous freight booking on load boards

Why it tempts you

Speed wins loads

Concrete failures

data-quality, judgment, fraud Double-brokering fraud rampant in 2024-25 — autonomous booking without carrier vetting transfers freight to bad actors. TIA tracking double-brokering losses exceeding $700M/yr industry-wide.

Right split

AI scores and surfaces top carriers + drafts rate confirmation; broker calls and books. Carrier identity verified through Highway/Carrier Assure (real-time MC# + identity check).

anti-033HealthcareRegulation

AI auto-prescribing or refilling Rx (without provider)

Why it tempts you

Cuts MA call burden

Concrete failures

regulation, low-frequency-high-stakes DEA + state pharmacy boards: only licensed prescribers can authorize. Repeated mis-refills (controlled substances) trigger DEA review + license risk.

Right split

AI queues refill requests with chart-pull summary; provider clicks approve/deny. Always.

anti-034HR / financeJudgment

Auto-payroll changes

Why it tempts you

Bonus calcs, comp adjustments

Concrete failures

regulation, judgment, low-frequency-high-stakes Wage-and-hour errors compound and back-pay claims accrue; California PAGA penalties are $200+/employee/pay period. One mis-coded raise propagates to bonus, equity refresh, severance baseline.

Right split

AI suggests + drafts; HR/finance approve before payroll cycle close. Audit log mandatory.

anti-035Property managementRegulation

AI handling tenant lease-renewal negotiations

Why it tempts you

Renewal cadence is predictable

Concrete failures

relationship-dependent, regulation (rent control) Rent-stabilized markets (NYC, LA, SF, Oregon, statewide rent caps) bar certain increases — LLMs don't know which units in a portfolio are covered. Tenant relations damage when renewal feels transactional.

Right split

AI calculates market + suggests offer; PM personally communicates with tenants.

anti-036Property managementRegulation

AI tenant screening (full auto-decline)

Why it tempts you

Volume of applications

Concrete failures

regulation (FCRA, FHA, state ban-the-box) FHA disparate impact on race + source-of-income protected classes in many states. FCRA requires adverse action notice with specifics — LLM can't cite its reason.

Right split

AI extracts and verifies (income docs, prior landlord refs); leasing agent makes go/no-go and sends compliant adverse action notices.

anti-037Healthcare admin / utilitiesRegulation

AI auto-quoting on rate-regulated services (healthcare prices, utilities)

Why it tempts you

Patient asks "what will this cost

Concrete failures

regulation (No Surprises Act, hospital price transparency rule), data-quality Federal hospital price transparency rule + state Good Faith Estimate rules require accuracy; misquoting can trigger HHS penalties and patient billing disputes.

Right split

AI pulls payer-specific rates + patient benefits; billing staff issue formal GFE.

anti-038HealthcareRegulation

AI receptionist for medical practice without HIPAA controls

Why it tempts you

Front desk is expensive

Concrete failures

regulation (HIPAA), privacy Many off-the-shelf voice AI vendors don't sign BAAs; transcripts of PHI passing through ungoverned LLM = breach. OCR breach notification + civil penalties.

Right split

AI receptionist with signed BAA, audit-logged, minimum-necessary PHI. Scheduling and FAQ only — never disclose chart info via untrained workflow.

anti-039HealthcareJudgment

Auto-drafted clinical notes posted directly to EHR

Why it tempts you

Reduces note-taking burnout

Concrete failures

regulation, judgment, low-frequency-high-stakes Hallucinated histories, transcription errors of medication doses (mg vs mcg) — direct patient safety risk. JAMA studies on AI scribe accuracy show clinically meaningful errors in 7-12% of notes.

Right split

AI scribe drafts; clinician reviews and signs every note (legal record).

anti-040Sales / marketingJudgment

Lead scoring with no human override

Why it tempts you

Reps want clarity

Concrete failures

data-quality, judgment Model drift; bad CRM hygiene poisons training data; sample-bias creates self-fulfilling prophecies.

Right split

AI suggests priority; rep accepts/rejects with reason; reasons feed retrain.

anti-041E-comm / dealerships / any LLM chatJudgment

Generative AI inside customer-facing chatbots without prompt-injection defenses

Why it tempts you

ChatGPT-on-your-site

Concrete failures

data-quality, judgment, brand Chevrolet of Watsonville chatbot tricked into "agreeing" to sell a $76K Tahoe for $1 in November 2023 via the "Bakke Method" prompt injection (https://www.upworthy.com/chevy-chatbot-gone-wrong-ex1/, https://incidentdatabase.ai/cite/622/). OWASP top-1 risk for LLMs.

Right split

AI for FAQ from a controlled knowledge base; output guardrails (no committing to prices/policies); escalate to human for anything that touches money.

anti-042Compliance / GRCRegulation

AI for SOC 2 / ISO control evidence gathering (auto-attest)

Why it tempts you

Audit prep is painful

Concrete failures

regulation, auditability Auditors require human attestation; auto-collected evidence that's mis-scoped won't satisfy CPA review and can constitute material misstatement.

Right split

AI gathers and indexes evidence; control owner attests. Tools like Vanta/Drata are this pattern done right.

anti-043HRRegulation

AI auto-resolving employee benefits / leave requests

Why it tempts you

FMLA / STD intake is repetitive

Concrete failures

regulation (FMLA, ADA, state leave laws), privacy ADA interactive process requires individualized assessment; LLM rules engine cannot make that call.

Right split

AI handles forms intake + status comms; HR/leave specialist decides eligibility.

anti-044Real estateRelationship

Real-estate AI showing-time scheduling that books seller property without listing-agent confirmation

Why it tempts you

Faster bookings

Concrete failures

change-management, relationship-dependent Tenant-occupied listings need scheduled access; auto-booking causes real-world conflict and broker complaints.

Right split

AI proposes windows; listing agent confirms; ShowingTime+ protocol stays in force.

anti-045Property managementRegulation

AI auto-resolution of tenant disputes / habitability complaints

Why it tempts you

24/7 response

Concrete failures

regulation, low-frequency-high-stakes Habitability complaints have legal consequences (rent withholding, repair-and-deduct, retaliation claims).

Right split

AI triages and dispatches non-habitability work orders; PM owns any habitability or legal-implication complaint.

anti-046GCs / contractorsJudgment

AI auto-quoting on construction change orders

Why it tempts you

CO disputes consume PM time

Concrete failures

judgment, contract-specific, low-frequency-high-stakes Each contract has different markup, time-impact, and lien-waiver rules; an LLM-issued CO can waive rights inadvertently.

Right split

AI drafts CO from field notes + photos; PM reviews and signs.

anti-047Contact center / CXJudgment

AI auto-scoring of customer service agent quality

Why it tempts you

100% of calls scored

Concrete failures

judgment, regulation, change-management Disparate impact on agents with accents or ESL backgrounds; union grievance risk; California PAGA exposure.

Right split

AI surfaces flag-worthy calls; QA team reviews; coaching conversations are human.

anti-048L&D / opsJudgment

AI auto-generation of educational / training materials posted as company SOPs

Why it tempts you

SOPs are stale; LLMs fill the gap

Concrete failures

data-quality, judgment LLM SOPs include hallucinated tool names, deprecated steps; employees follow them and break things.

Right split

AI drafts SOPs from video + observed workflows; SME reviews and signs. Versioned in a wiki, with last-reviewed-by stamp.

anti-049Compliance / accountingRegulation

Auto-drafted regulatory filings (BOI / 5500 / 5471 / VAT)

Why it tempts you

Forms are templated

Concrete failures

regulation, low-frequency-high-stakes FinCEN BOI penalty $591/day for inaccurate filings (https://www.fincen.gov/boi). IRS form 5471 penalty $10K per missed form; LLMs often misclassify foreign-entity types (CFC vs PFIC vs check-the-box election).

Right split

AI prepares draft + cites the rule it relied on; licensed pro reviews and signs.

anti-050Sales managementJudgment

Sales-call coaching with auto-PIP triggers

Why it tempts you

Conversation intelligence vendors pitch this

Concrete failures

judgment, regulation, change-management Auto-PIP based on call scores creates wrongful-termination exposure if disparate impact on protected classes. California, Illinois, Colorado AI Acts may classify this as workforce automated decision tool.

Right split

AI surfaces coaching themes per rep + per cohort; manager decides on PIP.

anti-051Personal finance / SMB treasuryRegulation

AI agent that takes actions in financial accounts (banks, brokerages)

Why it tempts you

Pay this bill, transfer $X

Concrete failures

regulation (KYC, Reg E, Reg D), low-frequency-high-stakes Plaid + Mercury TOS often restrict programmatic account actions; bad-actor scenarios trigger account freeze. Reg E liability for unauthorized transfers when AI scope is misconfigured.

Right split

AI drafts payment batches; CFO/owner approves in bank UI. Never store production credentials in LLM-accessible scope.

anti-052AccountingRegulation

AI editing financial close adjusting journal entries

Why it tempts you

Month-end speed

Concrete failures

regulation, data-quality, auditability SOX controls for public co.; AICPA AU-C 240 fraud risk for private co. — auditors will reject AI-only JE workflow.

Right split

AI drafts + memo-explains; controller approves; CPA reviews.

anti-053Legal opsRegulation

AI auto-responding to subpoenas / legal holds

Why it tempts you

Templated process

Concrete failures

regulation, low-frequency-high-stakes Mishandled legal hold = spoliation sanctions, adverse inference instructions.

Right split

AI flags inbound + maps custodians; counsel manages response and preservation.

anti-054Nonprofits, gov contractorsJudgment

AI auto-applying for grants / RFPs

Why it tempts you

Volume play

Concrete failures

judgment, relationship-dependent, low-frequency-high-stakes Wrong eligibility claims = debarment risk; some grants require certification by signing officer. "AI-generated proposals" detection at some agencies leads to disqualification.

Right split

AI drafts proposal from past wins + grant guidelines; principal reviews and signs.

anti-055Logistics / last-mile / rideshareJudgment

Driver dispatch optimization (full autonomy)

Why it tempts you

Big efficiency claim from vendors

Concrete failures

change-management, judgment, low-frequency-high-stakes Auto-dispatch ignores driver knowledge of difficult addresses, restaurant timing, weather; humans game it ("strategic deafness" to long-haul assignments). DOT HOS compliance: optimizer that pushes a driver into HOS violation = federal violation.

Right split

AI optimizer surfaces an ordered plan respecting HOS + driver constraints; dispatcher reviews and pushes.

anti-056InvestingRegulation

AI agent committing trades, options, or DeFi transactions

Why it tempts you

24/7 markets

Concrete failures

regulation (SEC, CFTC), low-frequency-high-stakes FINRA Reg BI on suitability; market-manipulation exposure if LLM tries pump-and-dump tactics.

Right split

AI surfaces ideas + drafts orders; trader executes within risk limits.

anti-057Content marketingRegulation

Auto-publishing AI-generated articles to brand blog without review

Why it tempts you

SEO at scale

Concrete failures

regulation (FTC, plagiarism), data-quality, brand CNET's 77 AI-generated finance articles required 41 corrections in 2023; reputation damage. Google's March 2024 spam policy update targets scaled content abuse; deindexing risk.

Right split

AI drafts + structured outline; editor reviews, fact-checks, and signs.

anti-058Product / marketingRegulation

AI builds "AI" features that don't actually use AI

Why it tempts you

AI-powered" sells

Concrete failures

regulation (FTC AI-washing) Builder.ai collapse May 2025 — used to be Engineer.ai accused of using humans behind "AI"; bankruptcy with $30M+ owed Microsoft (https://www.theregister.com/2025/05/21/builderai_insolvency/). FTC Operation AI Comply explicitly targets "machine learning when they relied on manual processes" (https://www.ftc.gov/news-events/news/press-releases/2024/09/ftc-announces-crackdown-deceptive-ai-claims-schemes).

Right split

If you claim AI, your stack must actually use AI; document substantiation; legal review of all AI claims.

anti-059Privacy / RevOpsRegulation

Auto-deleting / archiving customer data on a schedule (no human review)

Why it tempts you

Data minimization compliance

Concrete failures

regulation (litigation hold, IRS / GAAP retention) Active legal hold + auto-delete = spoliation sanctions. Tax/financial retention requires 7-year hold; auto-deleting invoices at 3yr = IRS issue.

Right split

AI drafts retention policy and proposes records for deletion; legal/compliance approves batches.

anti-060SupportJudgment

AI customer ticket auto-close based on inactivity sentiment

Why it tempts you

Cleaner queue

Concrete failures

judgment, data-quality, change-management "Auto-closed for no response" creates measurable churn lift when applied to genuine issues; CS metrics game themselves.

Right split

AI nudges customer for response; agent decides to close after 2 nudges with no reply. Reopen on any customer reply.

§06

Agent capability matrix — May 2026

decay: 3mo

What actually ships in production today. Reliability scores carry a 1.3× haircut versus vendor benchmarks. A demo is not a deployment.

Reality check before anything else

MIT's State of AI in Business 2025 found 95% of corporate GenAI pilots delivered zero measurable P&L impact across 300 deployments, 153 leader surveys, and 52 exec interviews. Fortune / MIT Gartner separately predicts >40% of agentic AI projects will be canceled by end of 2027. Gartner Same MIT data: purchased/partnered AI succeeds ~2× the rate of internal builds. Default audit recommendation: buy unless we have a defensible reason to build.

CapabilitySOTA (May 2026)RelHITLCost/runTop failure modes
Browser automationStagehand v3 + Browserbase + Claude Sonnet 4.5/4.6. WebVoyager near-saturated 88–89% bench; WebArena 68–74% leaderboard3/5Yes (writes)$0.05–$0.50Site A/B drift; auth/captcha; injection via page text Unit42
Code generationClaude Opus 4.5/4.6/4.7 (80.9–87.6% SWE-bench Verified) & GPT-5/5.5 (74.9–88.7%) swe-bench. SWE-bench Pro top score 46% morphllm3–4/5 bounded, 2/5 longYes (PR)$0.50–$5/taskRefactors >100k LOC; hallucinated APIs; Devin Railway hallucination Cognition
Email triage/draftSuperhuman, Shortwave, Lindy, native Gmail (Gemini)3/5 draftYes (send)$9–$30/seat/moTone drift; cross-thread hallucination; sensitive replies
Calendar schedulingReclaim, Motion, Cal.com AI. Clockwise sunsetting Mar 2026 Reclaim2/5 unsupervisedYes (external)$8–$19/seat/moPhantom events; over-defragmenting; cross-tz
Document / sheetClaude + Skills (xlsx/docx/pptx); GPT-5 + Code Interpreter; Gumloop for batch3–4/5Yes (finance)$0.10–$2/docDrops formulas; numeric vs text columns; XLSX >50 MB
CRM opsHubSpot Breeze, Salesforce Agentforce, Clay, Lindy. HubSpot: "Customer Agent most production-ready" SMM3/5Yes (outbound)$0.05–$0.50/recordGarbage-in/out; stale enrichment; dup rules over-fire
Voice agentsRetell ~600ms latency; Vapi ~500ms variable; ElevenLabs Agents Retell3/5Yes (escalate)$0.05–$0.30/minHallucinated numbers; barge-in; TCPA exposure on outbound
Support deflectionIntercom Fin ~51% avg resolution Swifteq; Zendesk AI ~38% deflection. Klarna walked back AI-first PromptLayer3–4/5 mature KB, 2/5 day-oneAlways tier 2+$0.50–$2/resolutionKB drift; "deflection theater"; brand voice
RAG / researchClaude Sonnet 4.5 + pgvector; GPT-5 file search; Perplexity API. Naive RAG retrieval fails ~40% of time lushbinary3/5Yes for decisions$0.01–$0.50/qRetrieval noise; conflated entities; confident-wrong
Data extractionClaude 4.5 97–98% field acc; GPT-4V 95–96.5%; Gemini 3 93.8–95.8% tokenmix4/5 typed, 3/5 scannedYes >$threshold$0.005–$0.10/docLayout edges; multi-page; locale decimals
Workflow orchestrationLangGraph (durable state); Temporal ($5B, OpenAI integration Sep 2025 InfoQ). OpenAI Agents SDK lacks checkpointing LangChain3–4/5Yes (branches)$0.05–$5/runState loss on restart; idempotency; 10+ step debug
Long-horizonClaude Sonnet 4.5 "30-hr coding" claim; Gaia2 pass@1 only 42% Gaia22/5Checkpoints$10–$100/dayCompounding error; 95%/step × 50 = 7% end-to-end medium
Computer-use (Claude/CUA)Claude Mythos 79.6% OSWorld-Verified; OSWorld-Human shows efficiency lag arxiv2–3/5Yes (writes)$1–$5/hrOS dialogs; multi-monitor; UI-text injection
§07

Vendor & platform map — honest

decay: 3mo

What each platform is actually good at, where it breaks, and when to recommend / when to avoid. Skip the marketing pages.

Anthropic API + Claude Agent SDK

Good at
Claude Opus 4.5/4.6/4.7 leads SWE-bench Verified 80.9–87.6%; Sonnet 4.5 leads τ-bench Airline at 0.700. Agent SDK exposes tools, hooks, MCP, subagents. Skills for xlsx/docx/pptx/pdf.
Breaks in prod
Computer Use beta-ish, token-heavy. Opus pricey for high-volume routing.
Lock-in
Medium. MCP portable; Skills less so.
Pricing landmines
Extended-thinking multiplies cost. Cache aggressively — published patterns assume cache hits.
Recommend when
Coding, tool-heavy multi-step, doc extraction, long-context (200k–1M). Default for serious agent builds.
Avoid when
Native multi-vendor TTS/voice; sub-Haiku-only budgets.

OpenAI API + Agents SDK + AgentKit

Good at
GPT-5/5.5 raw capability; built-in `responses` with tools & file search; ChatGPT Agent (Operator absorbed July 2025). Temporal integration Sep 2025 for durability.
Breaks in prod
SDK lacks LangGraph-style checkpointing — HITL-pause flows need custom infra. Tool-call latency reports.
Lock-in
Medium–high. Responses API + AgentKit are OpenAI-specific.
Pricing landmines
Reasoning tokens; file search per-GB; Code Interpreter per-session. Caching exists but easy to miss.
Recommend when
Multimodal heavy; one-vendor preference; fine with OpenAI primitives.
Avoid when
On-prem; durable long-running state; vendor independence.

LangGraph + LangSmith

Good at
Production leader for stateful, branching, HITL-pauseable agent graphs. LangSmith traces are the most mature OSS option.
Breaks in prod
LangChain core is bloated — most teams now use only LangGraph + raw provider SDK.
Lock-in
Low for LangGraph; high for LangSmith cloud.
Pricing landmines
LangSmith trace volume; cloud deploy minimums.
Recommend when
Python team needs durable, stateful agents with HITL gates.
Avoid when
A managed product (Lindy, Gumloop, Zapier) covers it.

CrewAI

Good at
Quick demos with role-playing agents; rising mindshare.
Breaks in prod
Manager-worker process executes sequentially despite docs; agents skip tasks, fabricate IDs, return made-up results. Practitioner reports: '5–10 min/run, hours debugging silent failures.'
Lock-in
Low (Python).
Pricing landmines
High token use from multi-agent chatter.
Recommend when
Demos and hackathons.
Avoid when
Production unless you wrap heavily.

n8n (self-host + cloud)

Good at
v2.0 (Dec 2025) added isolated code execution, RBAC, 70+ AI nodes with LangChain integration. 80–90% cheaper than Zapier on high-volume flows. Only one with prod-grade self-host.
Breaks in prod
Steeper learning curve. Self-host means you operate HA/backups. Some integrations less polished than Zapier.
Lock-in
Low if self-hosted. AGPL is a gotcha for closed-source SaaS resale.
Pricing landmines
Cloud tier execution caps. AGPL gotcha for productized services.
Recommend when
Dev-led team, data-residency, agency builds for clients.
Avoid when
Non-tech solo founder; point-and-click only.

Zapier (incl. Zapier Agents)

Good at
Largest connector library (~8,500), polished templates, non-tech accessible. SOC 2 in flight on Agents (Dec 2025). Agents revamped May 2025 — focus moved from chat to automation.
Breaks in prod
Trustpilot 1.4 with surprise-billing complaints. AI builder fine for prototyping, not bulletproof prod. Linear workflows can't elegantly do branching/parallel. Zapier Central sunset.
Lock-in
High — workflows non-exportable.
Pricing landmines
Task billing; retries count. Multi-step AI Zaps compound.
Recommend when
Non-tech founder, low-volume, breadth of integrations dominant.
Avoid when
>10k tasks/mo; branching/parallel logic; self-host.

Make.com

Good at
Visual canvas with branching/parallel; SMB pricing; AI Agents added Oct 2025. Sits between Zapier (linear) and n8n (code).
Breaks in prod
Operations billing model is complex; harder to estimate. Some AI nodes thin.
Lock-in
Medium–high.
Pricing landmines
Operations metering on iterators (1000-row loop = 1000 ops).
Recommend when
SMB needing branching with a visual UI.
Avoid when
Self-host required.

Lindy.ai

Good at
Template-first AI-native agent creation; G2 4.9 (168+ reviews). Handles ambiguous tasks better than deterministic-workflow tools.
Breaks in prod
Closed system; debugging limited; integration breadth narrower than Zapier/n8n.
Lock-in
High.
Pricing landmines
Per-task + per-agent tiers stack.
Recommend when
Solo operator / agency owner needs assistant-style agents fast.
Avoid when
Complex DAGs or custom code paths.

Gumloop

Good at
Visual data pipelines purpose-built for PDFs, sheets, scrape, doc transform at scale. Top vendor score (84) in Zapier-alt comparisons.
Breaks in prod
Deterministic flow model less suited to ambiguous chat/email tasks.
Lock-in
Medium–high.
Pricing landmines
Per-run credits; large batches expensive.
Recommend when
Data extraction / batch document workflows.
Avoid when
Conversational agents primary.

Relevance AI

Good at
Agent-as-employee abstraction; sales/RevOps tooling; managed.
Breaks in prod
Less ecosystem than Lindy/Gumloop.
Lock-in
High (closed runtime).
Pricing landmines
Per-credit, opaque at high volume.
Recommend when
Sales/SDR team wants pre-built personas fast.
Avoid when
Need transparency or migration optionality.

Retool (+ Retool Agents)

Good at
Internal-tool UI + workflows + AI in one place. Strong RBAC. Best choice when agents need a human-facing dashboard.
Breaks in prod
Agent layer less mature than dedicated platforms; Retool-style debug pain.
Lock-in
Very high (custom DSL/state).
Pricing landmines
Per-seat scales painfully; Retool DB / workflows are stacked add-ons.
Recommend when
Internal ops dashboard with agent assist.
Avoid when
Public/customer-facing; cost-sensitive.

Airtable + Cobuilder

Good at
Cobuilder spins up bases from spec. AI fields for enrichment. Ubiquitous SMB adoption.
Breaks in prod
Scaling walls at ~100k records / complex automations. AI features paywalled.
Lock-in
High data model.
Pricing landmines
Per-user seats; AI credit add-ons.
Recommend when
SMB CRM-lite, content ops, light agents on top.
Avoid when
True backend / app-level data.

Supabase

Good at
Postgres + Auth + Storage + Edge + pgvector. Real-time. Default agent backend.
Breaks in prod
Free tier pauses; RLS is power-user only.
Lock-in
Low (Postgres).
Pricing landmines
Egress; compute add-ons.
Recommend when
Any production agent backend that needs DB+auth.
Avoid when
Fully managed enterprise (Snowflake-class).

Browserbase + Stagehand

Good at
Stagehand v3 is 44%+ faster than v2 via direct CDP. Self-healing selectors; iframe/shadow DOM. Browserbase cloud handles captchas, session replay, agent identity.
Breaks in prod
AI-resolved selectors add latency and LLM cost. Non-deterministic debug.
Lock-in
Low for Stagehand OSS; high for Browserbase cloud.
Pricing landmines
Per-session minutes + LLM tokens.
Recommend when
Any agent doing real browser tasks across changing sites.
Avoid when
Single stable site you own (use Playwright).

Playwright / Puppeteer

Good at
Cheapest, fastest, deterministic. Existing test-automation expertise transfers.
Breaks in prod
Brittle — redesigns break selectors instantly. No 'AI healing.'
Lock-in
None.
Pricing landmines
Compute only.
Recommend when
Stable site + scheduled job.
Avoid when
Sites you don't own or change often.

Firecrawl

Good at
LLM-ready markdown; thousands of pages/min mid-tier. AI-driven navigation. Single API for scrape/crawl.
Breaks in prod
No native scheduling — bring your own. Less control on dynamic anti-bot.
Lock-in
Low.
Pricing landmines
Per-page credits; JS-render multiplier.
Recommend when
RAG ingestion, research-agent feeds.
Avoid when
Need scheduling / queue / Actor marketplace.

Apify

Good at
Scheduler + queue + retries built-in. Actor marketplace. Burst workloads (thousands parallel).
Breaks in prod
Heavier learning curve; container overhead.
Lock-in
Medium.
Pricing landmines
Compute units + proxy traffic.
Recommend when
Production scrape pipelines, e-com catalogs.
Avoid when
Just need 'LLM-ready text' — use Firecrawl.

Vector DBs (Pinecone / Qdrant / Weaviate / pgvector)

Good at
At 10M vectors: pgvector ~$45/mo, Qdrant ~$65, Pinecone serverless ~$70, Weaviate ~$135. At 100M: pgvector/Milvus <$100 vs Pinecone $700+. Qdrant 1840 QPS on 1M vectors.
Breaks in prod
Pinecone serverless cold-start latency; Weaviate ops if self-hosted.
Lock-in
Low for pgvector/Qdrant; high for Pinecone managed features.
Pricing landmines
See cost matrix — vector count + QPS is the metric, not storage.
Recommend when
<5M + Postgres shop → pgvector. <100ms P99 no-ops → Pinecone. Hybrid keyword+vector → Weaviate. High-QPS OSS → Qdrant.
Avoid when
Wrong tier for scale — switching mid-deploy is painful.

Voice — Retell / Vapi / ElevenLabs Agents

Good at
Retell ~600ms latency, SOC2/HIPAA/GDPR out of box. ElevenLabs voice quality leader; IBM watsonx partnership Mar 2026. Vapi flexible if you tune.
Breaks in prod
Vapi sub-500ms is variable; carrier latency adds 200ms+ to all. TTS misreads on numbers.
Lock-in
Medium per platform.
Pricing landmines
Per-minute talk time + per-LLM-call. PSTN trunking extra.
Recommend when
Regulated → Retell. Brand voice → ElevenLabs. Custom tuning + voice eng → Vapi.
Avoid when
Outbound cold sales (TCPA exposure).

Integration APIs — Paragon / Merge / Nango

Good at
Nango: free 5k calls/mo, $249 for 50k single-category — most flexible OAuth + sync logic you own. Merge: unified API (HRIS/ATS/Accounting/CRM/ticketing). Paragon: enterprise embedded.
Breaks in prod
Paragon pricing opaque (no clear metric on page). Merge: less control on edge sync logic.
Lock-in
Low Nango, medium Merge, high Paragon.
Pricing landmines
Per-record/calls for Nango/Merge; sales-led for Paragon.
Recommend when
Startup → Nango. B2B SaaS embedded (many integrations) → Merge. Enterprise sales motion → Paragon.
Avoid when
Wrong tool for stage.
§08

Recommended stacks — by use case

decay: 3mo

Practical defaults. Pick by the closest-fit use case, not the most exciting vendor.

  • No-code MVP, non-technical founder. Lindy (or Zapier Agents) + Gmail/Calendar/CRM connectors + Claude Sonnet 4.5 backend. Fastest time-to-first-agent. Accept lock-in for speed — MIT: bought beats built ~2:1.
  • SMB automation, agency-owner-no-devs. Make.com OR n8n cloud + Claude/OpenAI + Airtable/Supabase + Firecrawl/Apify. Visual canvas + branching + AI nodes. n8n if self-host or data-residency matters; Make if Make is what they know.
  • Internal ops dashboard. Retool (UI + workflows + agents) + Supabase + Claude API. Retool is the unique combo of UI builder + agent runtime. Worth the lock-in for ops portals.
  • Browser-heavy. Stagehand v3 on Browserbase + Claude Sonnet 4.5 + Temporal (durability) + Supabase. Self-healing selectors are the only thing that survives DOM drift at scale. Temporal carries long browser sessions across failures.
  • Document-heavy. Gumloop (batch) OR Claude API + Skills (xlsx/pdf/docx) orchestrated by LangGraph, persisted in Postgres. Claude leads invoice extraction (97–98%) and produces valid JSON 100% of cases.
  • Sales/support automation. Intercom Fin OR Zendesk AI Agents (deflection) + Retell voice + Clay/HubSpot for outbound enrichment + Merge unified API to sync CRMs. Buy the deflection layer (Fin > 50% avg resolution beats most internal builds). Humans on tier 2. Learn Klarna's lesson — don't go AI-first 100%.
  • Productized service (build once, sell many). n8n self-hosted (watch AGPL) OR LangGraph + OpenAI/Anthropic + Supabase + Browserbase OR white-label Lindy/Make per-tenant. LangGraph stack is most defensible/portable; managed stack is fastest.
§09

Failure modes & cross-cutting heuristics

decay: 6mo

The 15 most common reasons agent deployments fail in 2025-2026, with citations. Plus the heuristics our audit agent applies.

Top 15 failure modes

  1. Eval gap. Naive RAG retrieval fails ~40% in prod lushbinary; 95% of corp pilots produce zero P&L MIT.
  2. Long-horizon compounding. 95%/step × 50 steps = 7% success medium. METR 50%-reliability horizon: tens of minutes, not days METR. Gaia2 ceiling 42% pass@1.
  3. pass^k collapse.τ-bench retail pass@1 < 50% but pass^8 < 25% Sierra— your demo isn't your 8th-run reality.
  4. Prompt injection. CVE-2025-59944 (Cursor → RCE); $250k bank-assistant fraud (Jun 2025); AI worm Feb 2025 mayhem. Multi-hop indirect attacks +70% YoY SQ.
  5. Brittle browser DOMs. Even Stagehand self-healing carries LLM-resolution cost per repair. Princeton "Illusion of Progress" shows agents fail on real-world long-tail sites arxiv.
  6. Hallucinated tool calls. CrewAI agents fabricate IDs / skip tasks Popelka; Devin hallucinated Railway features for a day+ Cognition.
  7. Context cost spiral. Long-running agents accumulate context; computer-use screenshots especially heavy. Budgets blow on retries.
  8. Permissioning hell. Real deployments hit OAuth scope sprawl across Gmail/Cal/CRM. Anthropic guidance: deny-all-by-default, allowlist per subagent Anthropic.
  9. Vendor pricing surprises. Zapier Trustpilot 1.4 Startupowl. Make ops billing on iterators. Pinecone $700+ at 100M vs pgvector <$100 LeanOps.
  10. HITL fatigue → rubber-stamping. Operators learn to approve without reading. Klarna CSAT drop → AI-first reversal CXD.
  11. Reward hacking / spec gaming. UC Berkeley CRDI: an automated scanning agent broke all 8 major agent benchmarks via reward hacking (Apr 2026) rapidclaw.
  12. Statelessness / restart loss. Agents commonly forget previous messages, lack retries, lose progress on reboot Particula. OpenAI Agents SDK lacks checkpointing.
  13. Fan-out that doesn't fan out. CrewAI sequential despite docs implying coordination TDS.
  14. Voice hallucinations & latency.<800ms required for natural feel; some platforms 3–4s. TTS misreads numbers (account IDs, dates); TCPA risk on outbound.
  15. Benchmark contamination. OpenAI dropped SWE-bench Verified — moved to SWE-bench Pro where top score is 46% vs 81% on Verified morphllm. Audit-agent rule: discount any single benchmark ~30% when projecting to a customer's domain.

Cross-cutting heuristics (the audit agent's rules)

System prompt — heuristics block

  • Prefer buy over build by ~2:1. MIT data is clear.
  • Quote a reliability haircut: divide vendor-claimed accuracy by ~1.3× for real-world distribution shift.
  • Always quote pass^k or k-attempt reliability, not pass@1.
  • Budget retries. 30% retry rate at $0.50/run = $0.15 hidden cost per task.
  • Mandatory HITLfor: send (email/SMS), spend (>$X), legal/compliance, customer-visible writes.
  • Pick durability infra first. LangGraph + Temporal (or Mistral Workflows pattern). Statelessness kills more deploys than model quality.
  • Default vector DB: pgvector if you have Postgres.
  • Default doc extraction: Claude 4.5/Opus. GPT-4o only on degraded scans.
  • Voice in regulated industries: Retell or ElevenLabs. SOC2/HIPAA/GDPR out of box matters.
  • Browser: Stagehand+Browserbase over Playwright unless target is owned and stable.
§10

Per-role pain inventory

decay: 12mo

Who buys, what they hate, what they'll pay for. Each role maps to processes from §3.

Agency Owner (services, 5–50 people)

$120k–$400k+ owner draw

Headcount — SMB: 1 founder · Mid: 1–3 owners + operators · Enterprise: n/a — sells to large agencies

Top pains (their words)

  • "Friday afternoon panic-deck reporting eats my week."[src]
  • "Scope creep is silently destroying our margin per project."[SPI Research PSMB 2024]
  • "Time entry compliance is a Friday email tax."[Workamajig agency surveys]
  • "Proposals take 3 days when they should take 30 minutes."[PandaDoc benchmarks 2024]
  • "Tracking utilization vs. signed SOWs is a stitched-spreadsheet nightmare."[SPI Research PSMB]

Weekly hours automatable

12–18 hrs founder time

Buying authority

Full sign-off on $0–$3k/mo on AmEx; founders are the buyer.

Existing band-aids

Templates in Notion, ClickUp, Workamajig. Most have tried 1-2 prior automation attempts that fizzled on adoption.

Director of Operations (SMB / Mid)

$120k–$180k base + bonus

Headcount — SMB: 1 · Mid: 1 + 1–3 IC ops · Enterprise: VP Ops + Director(s) + 5+ IC

Top pains (their words)

  • "We use 9 tools that don't talk and I'm the human SQL query."[Asana Anatomy of Work 2024]
  • "Weekly KPI rollups eat 4 hours every Monday."[Operator community forums]
  • "Status updates between teams get dropped on handoff."[Atlassian State of Teams 2024]
  • "Our spreadsheets are our database and nobody trusts them."[MetaPlane data-quality survey]
  • "Vendor renewals auto-fire without negotiation."[Vendr State of SaaS 2024]

Weekly hours automatable

14–22 hrs

Buying authority

$0–$5k/mo direct; above goes to CFO/founder.

Existing band-aids

Zapier, Make, Notion automations, a few SOPs. Most are partially-functional.

CFO at 20–100 person company

$220k–$360k base + equity

Headcount — SMB: Fractional CFO + bookkeeper · Mid: 1 CFO + 2–5 finance ops · Enterprise: VP Finance + Controller + 10+ ICs

Top pains (their words)

  • "Month-end close takes 12 days when it should take 5."[FloQast Close Benchmark 2024]
  • "AR follow-ups consume 30%+ of AR specialists' time."[src]
  • "Cash flow forecasts are last week's data by the time they're done."[AFP Cash Forecasting Survey]
  • "Board pack assembly Friday before is a fire drill."[FENG community AMAs]
  • "I can't tell which customers are profitable."[AICPA SMB benchmarking]

Weekly hours automatable

10–25 hrs across finance team

Buying authority

$5k–$50k/mo discretionary; gatekeeper for everyone else's AI spend.

Existing band-aids

QBO/NetSuite + FloQast/Numeric for close; Ramp/Brex for cards; the FP&A model is in Sheets.

RevOps Lead

$140k–$220k loaded

Headcount — SMB: 0–1 part-time · Mid: 1 + 1–2 ICs · Enterprise: Director + 5+ analysts

Top pains (their words)

  • "Forecasting is a vibe check; the CRO doesn't trust the number."[Clari forecasting reports]
  • "CRM hygiene drift kills every dashboard; nobody owns it."[Gong Pipeline Report 2024]
  • "Lead routing rules are spaghetti and leads sit too long."[InsideSales / Velocify 5-minute rule research]
  • "Pipeline reviews are slide-rebuilds, not strategy."[Gong State of Revenue 2024]
  • "Sales notes never make it to CRM; reps hate it."[src]

Weekly hours automatable

15–25 hrs

Buying authority

$0–$3k/mo direct; CRO/CFO above.

Existing band-aids

HubSpot/Salesforce + Gong + Outreach + custom Zaps. CRM is the bottleneck.

Head of Customer Success

$160k–$250k loaded

Headcount — SMB: Founder + 1 CSM · Mid: Director + 3–8 CSMs · Enterprise: VP CS + Director(s) + 20+ CSMs

Top pains (their words)

  • "Renewals sneak up; we start motion 30 days out when we need 90."[Gainsight 2024 State of CS]
  • "QBRs eat a CSM's entire week each quarter."[CSM community forums]
  • "Health scores are gut feel; churn surprises us."[Catalyst Customer Health benchmark]
  • "Expansion always gets deprioritized for save-the-account fires."[Gainsight 2024]
  • "Onboarding is ad-hoc; no two customers see the same playbook."[Rocketlane onboarding benchmarks]

Weekly hours automatable

18–30 hrs across the CS org

Buying authority

$1k–$10k/mo direct; CRO above for enterprise tools.

Existing band-aids

Gainsight/Catalyst/Vitally + Slack + spreadsheets. Most run rules-based health, not ML.

E-commerce Operator ($1–10M Shopify brand)

$80k–$300k owner draw

Headcount — SMB: Founder + 1–3 · Mid: Founder + 5–15 ops/CS/marketing · Enterprise: n/a

Top pains (their words)

  • "Reviews pile up on Trustpilot and Amazon; responses are inconsistent."[Trustpilot benchmark 2024]
  • "Returns processing is a manual customer-service grind."[Loop Returns benchmark]
  • "Ad creative testing is slow; we never have enough variants."[AdCreative.ai benchmarks]
  • "Customer service for order issues is 60% of tickets."[Gorgias / Zendesk e-comm benchmark]
  • "Post-purchase email flows are out of date."[Klaviyo State of Email 2024]

Weekly hours automatable

20–35 hrs across team

Buying authority

Founder full sign-off on $0–$5k/mo on AmEx.

Existing band-aids

Shopify app stack (12+ tools), Klaviyo, Gorgias, Loop. Most over-tooled and under-integrated.

Solo SaaS Founder / Indie Hacker

$0–$300k (lumpy)

Headcount — SMB: 1 · Mid: n/a · Enterprise: n/a

Top pains (their words)

  • "Inbox drowns; important sales/press emails get missed."[Indie Hackers forums]
  • "Support tickets at 50/day require my attention until I hire."[Indie Hackers / r/SaaS]
  • "I'm the only one who knows how anything works."[MicroConf community]
  • "Manual content distribution and repurposing is 8 hrs/week."[MicroConf surveys]
  • "Time-to-bill is too long because I forget to invoice."[r/SaaS founder threads]

Weekly hours automatable

10–20 hrs founder time

Buying authority

Full sign-off on $0–$2k/mo.

Existing band-aids

Stripe + Notion + ChatGPT + ad-hoc Zaps. Building automations in spare time.

Practice Manager (medical / dental / legal / accounting)

$75k–$140k

Headcount — SMB: 1 + 2–5 admin · Mid: 1–2 + 6–20 admin · Enterprise: n/a (MSO instead)

Top pains (their words)

  • "Patient/client intake is a paper-form graveyard."[AMA practice ops survey 2024]
  • "Insurance verification and prior auth chases dominate the day."[src]
  • "Claim denials are a recurring 8-15% revenue tax."[src]
  • "Recall campaigns for return visits are inconsistent."[Practice Management Institute]
  • "Referral coordination between practices is fax + phone."[AAPP practice surveys]

Weekly hours automatable

15–25 hrs across admin team

Buying authority

$0–$2k/mo direct; physician/partner above.

Existing band-aids

Tebra / Athena / Clio + paper. HIPAA / state bar / accounting rules constrain options.

Local Services Business Owner (HVAC, plumbing, landscaping)

$80k–$300k owner draw

Headcount — SMB: 1 + 2–6 techs · Mid: 1 + 8–20 techs · Enterprise: n/a — sells to franchises

Top pains (their words)

  • "Leads sit for 30+ minutes before someone responds. Conversion craters."[src]
  • "Techs forget to update job status; dispatch is guessing."[ServiceTitan customer benchmarks]
  • "Estimates eat our weekends."[Buildxact contractor surveys]
  • "Review requests are inconsistent post-job."[Birdeye/Podium benchmarks 2024]
  • "Recurring service reminders rely on memory, not system."[ServiceTitan 2024]

Weekly hours automatable

10–20 hrs owner + dispatcher

Buying authority

Owner full sign-off on $0–$2k/mo.

Existing band-aids

ServiceTitan / Jobber / Housecall Pro. Most leave money on the table by not using the AI features.

Engineering Manager (50–500 person co)

$220k–$360k loaded

Headcount — SMB: 1 + 4–8 eng · Mid: 1 + 8–15 eng · Enterprise: Director + multiple EMs

Top pains (their words)

  • "PR review is a senior-engineer bottleneck."[DORA 2024 State of DevOps]
  • "Onboarding a new engineer takes 4–8 weeks to first material PR."[Stack Overflow Developer Survey]
  • "On-call burns out senior engineers; runbooks are stale."[PagerDuty 2024 oncall report]
  • "Flaky tests kill CI trust."[Buildkite reliability surveys]
  • "Internal docs are stale; same questions get asked weekly."[Atlassian State of Teams 2024]

Weekly hours automatable

20–35 hrs across eng org

Buying authority

$1k–$5k/mo direct; VP Eng / CTO above.

Existing band-aids

GitHub Copilot + Cursor + manual PR review. Cody/Greptile in some shops.

Recruiting Operations / Talent Ops

$95k–$160k loaded

Headcount — SMB: 0–1 · Mid: 1 + 1–3 recruiters · Enterprise: Director + 5+ ICs

Top pains (their words)

  • "Sourcing personalization at scale is impossible manually."[Gem 2024 Talent Acquisition]
  • "Interview scheduling across panel + candidate burns 30 min/loop."[GoodTime benchmarks]
  • "Candidate ghosting happens because reply SLAs slip."[src]
  • "Reference checks are last-minute manual scrambles."[Crosschq surveys 2024]
  • "ATS data is dirty; reporting is broken."[Ashby benchmarking 2024]

Weekly hours automatable

12–22 hrs

Buying authority

$500–$2k/mo direct; Head of Talent above.

Existing band-aids

Greenhouse/Ashby/Lever + Gem/SourceWhale + GoodTime. The stack is mature; integration is the gap.

Chief of Staff / Executive Assistant

$140k–$280k loaded

Headcount — SMB: 0–1 EA · Mid: 1 CoS + 1 EA · Enterprise: CoS + EA staff

Top pains (their words)

  • "Founder's inbox is a fire-hose; important items get missed."[Founder community forums]
  • "Board pack assembly is a Friday-night ritual."[CoS community AMAs]
  • "Founder shows up to meetings cold without context."[EA Network surveys]
  • "OKR/metric rollups across functions take a full day."[Operator Collective community]
  • "Scheduling negotiation across exec calendars is a multi-day game."[GoodTime exec scheduling]

Weekly hours automatable

20–30 hrs

Buying authority

$0–$3k/mo direct on AmEx; founder sponsorship for larger.

Existing band-aids

Superhuman/Shortwave + Notion + Google Workspace + manual scripts.

§11

Operator ↔ vendor glossary

decay: stable

How operators actually describe pain (left) vs. the tech category that solves it (right). Sales / discovery gold.

Operator phraseCapabilityExample solution
"Leads sit too long before we touch them."Inbound lead routing + speed-to-lead automation with SLA timerHubSpot/Salesforce workflows + Default/Distribute + Slack-alerting agent that pings AE if no first-touch in 5 minutes[src]
"I'm copying and pasting between LinkedIn, Apollo, and the CRM all day."Browser-based prospecting orchestration / sequencer enrichmentClay or Apollo + n8n/Make + CRM write-back agent
"Sales notes never make it to CRM."Call → CRM enrichment (transcript → field extraction)Gong/Fathom/Granola + LLM extraction agent + Salesforce/HubSpot API write[src]
"I don't trust our CRM data."CRM hygiene + dedup + enrichment automationSyncari/Openprise or custom dbt + Clearbit/Apollo enrichment + nightly LLM dedup agent
"Forecasting is a vibe check."Pipeline scoring + commit-call analyticsClari/BoostUp or Salesforce + Gong forecast + LLM stage-validation agent that flags stale deals[src]
"I never know what my pipeline really looks like."Real-time pipeline reporting + deal hygiene agentSalesforce reports + Gong + Slack digest agent that posts MEDDPICC gaps daily
"My SDRs send the same email 80 times a day."Personalized outbound at scaleClay + Apollo + LLM personalization on top of Outreach/Salesloft sequences
"Discovery calls are inconsistent across reps."Call coaching + question-coverage scoringGong/Chorus + LLM scorecard agent that flags missed MEDDPICC fields
"Quotes take days to get out the door."CPQ + approval workflow automationSalesforce CPQ / DealHub + Slack approval agent
"Contract redlines bounce around for weeks."Contract lifecycle management with AI redline assistantIronclad/Lexion + LLM redline-comparison agent + DocuSign
"Renewals sneak up on me."Renewal forecasting + auto-prompted CSM workflowGainsight/Catalyst + renewal-90/60/30 agent in Slack with talking points
"Our QBRs eat my whole week."Auto-generated QBR decks from usage + CRM dataMixpanel/Amplitude + CRM + LLM deck-builder agent (Google Slides API)
"Customer health is a gut feel."Health scoring from product usage + support + sentimentGainsight/Vitally + Zendesk + Gong sentiment + LLM aggregator
"Onboarding is a mess."Workflow orchestration with deadline timers + Slack nudgesRocketlane/Arrows + Slack agent + checklist-tracking LLM
"Churn shows up out of nowhere."Leading-indicator churn agent (usage drop + sentiment + support volume)Catalyst/Vitally + Zendesk + LLM signal-aggregation agent
"Expansion never gets prioritized over saves."Expansion-signal detection (feature adoption, seat growth)Product analytics + CRM + LLM agent that surfaces expansion plays weekly
"Tickets pile up overnight."24/7 deflection agent + L1 triageIntercom Fin / Zendesk AI / Decagon / custom RAG over docs + KB[src]
"We answer the same question 50 times."KB-grounded chatbot with deflectionZendesk AI / Intercom Fin / custom RAG (Pinecone + Claude) + Slack escalation
"Agents waste 10 minutes searching the wiki for every ticket."Agent-side AI copilot with retrievalZendesk AI / Forethought / Cresta + internal KB embeddings
"Macros are out of date but nobody updates them."Auto-generated response suggestions from resolved ticketsLLM mining past tickets → suggested macros pipeline
"Tier 1 escalates everything because they're scared to answer."Confidence-scored answer suggestions + decision tree agentForethought or custom LLM with confidence threshold + human handoff
"We can't keep up after a product launch."Surge handling with topic clustering + auto-FAQ generationTopic-clustering LLM over inbound + auto-update KB + macro suggestion
"Following up on AR is killing me."Automated dunning / collections agentUpflow/Chaser/Versapay + email-personalization LLM + Slack escalation[src]
"Vendor invoices show up everywhere."AP intake + OCR + 3-way match automationRamp Bill Pay / Bill.com / Tipalti + LLM extraction + ERP push
"Month-end close takes 12 days."Close orchestration + recon agentFloQast / Numeric + LLM JE-explanation agent
"Reconciliations are 80% of my month."Auto-recon with exception-only reviewBlackLine / Numeric + LLM variance explainer
"Budget vs. actuals takes a week to assemble."Live FP&A dashboard with variance commentaryCube / Pigment / Mosaic + LLM variance-narrative agent
"Expense reports are a nightmare."Card-feed + receipt OCR + auto-categorizationRamp / Brex + LLM policy-check agent
"Audit prep is a full-time job for two weeks."Continuous controls monitoring + evidence collectionAuditBoard / Drata + evidence-gathering agent
"I can't tell which customers are profitable."Unit economics dashboard with cost allocationdbt + Mode/Hex + LLM commentary agent
"We keep dropping the ball on handoffs."Cross-system state tracking + checklist agentLinear/Asana + Slack agent + LLM workflow-state monitor with reminders
"We use 9 tools that don't talk."iPaaS + custom integration layerWorkato / Tray / n8n + event bus + LLM mapping helper
"Spreadsheets are our database."Lightweight workflow DB + form intakeAirtable / Smartsheet / Retool + migration agent
"Manual data entry between systems."ETL + system-of-record syncFivetran / Hightouch + reverse-ETL + LLM field-mapping helper
"Status updates eat my Mondays."Auto-generated status reports from Linear/Jira/SlackLinear + Slack + LLM weekly-summary agent
"Standups are pointless."Async standup bot with auto-extracted blockersGeekbot / Range + LLM blocker-detection agent
"Two people own this, neither does it."RACI tracking + accountable-owner agentNotion DB + Slack reminder agent that escalates after N days
"We rebuild the same deck every Monday."Templated reporting deck auto-generated from dataGoogle Slides API / Plus.ai + LLM commentary layer
"Nobody reads our docs."RAG-powered Q&A over internal wikiNotion/Confluence + Glean / Dust / custom RAG with Slack bot
"I'm drowning in Slack."Smart Slack digest + thread summarizationSlack AI / Glean / custom agent with channel-priority logic
"Inbox zero is a myth."Email triage agent with draft repliesSuperhuman AI / Shortwave / Missive + LLM draft agent
"We can't find the latest version of anything."Single-source-of-truth doc system + version intelligenceNotion + Glean federated search agent
"Customer requests get lost in DMs."DM → ticket capture agentSlack listener bot → Linear/Zendesk auto-create
"We never have data when leadership asks."Self-serve analytics + LLM-over-warehouseHex / Mode / Julius + dbt + LLM SQL-generation layer
"Hiring is a black hole."ATS automation + candidate-status visibilityAshby/Greenhouse + Slack agent + LLM screening assist
"Recruiters ghost us."Pipeline SLA tracking + candidate-response agentAshby + LLM personalization for outreach + nudge agent
"Performance reviews are theater."Continuous-feedback platform + AI-summarized reviewsLattice / 15Five + LLM that aggregates 1:1s, PRs, and Slack into draft review
"Onboarding new hires is ad-hoc."Onboarding workflow agent + checklist trackingRippling / Sapling + Slack onboarding bot
"PTO requests bounce around email."Self-serve HRIS workflowsRippling / Gusto / Deel + Slack approval bot
"Comp benchmarking is a guessing game."Comp data + leveling automationPave / Figures + LLM offer-letter generator
"Compliance is always a fire drill."Continuous compliance automationDrata / Vanta / Secureframe + evidence agent
"Legal needs to approve everything."Self-serve legal playbook + AI redlineIronclad / Spellbook + LLM playbook-routing agent
"NDAs take a week."NDA auto-execution with playbookIronclad / LinkSquares + LLM clause-check
"Vendor security reviews kill deals."Auto-completed security questionnaires from policy KBLoopio / Responsive + RAG over policies
"Marketing requests pile up in JIRA."Creative ops intake + auto-routingAsana / Wrike + LLM brief-completeness agent
"Brief never matches what's delivered."Structured brief intake with required-field validationForm + LLM brief-quality scorer + approval workflow
"We can't attribute pipeline to campaigns."Multi-touch attribution + LLM commentaryHubSpot / Bizible / Dreamdata + revenue-impact agent
"Content calendar is in three places."Single content ops hubNotion/Airtable + Asana + LLM repurpose agent
"SEO content takes forever to brief."AI brief generator from SERP + competitor dataAhrefs/SEMrush API + LLM brief-writer
"Newsletter is always late."Auto-curated newsletter from RSS + KB + LLM editorBeehiiv/Substack + LLM curation agent
"PagerDuty is keeping me awake."Incident triage + auto-runbook executionPagerDuty + Rootly / FireHydrant + LLM runbook agent
"Code review takes days."AI PR reviewer + auto-summarizationGitHub + Greptile / CodeRabbit / Graphite + LLM reviewer
"Onboarding a new engineer takes a month."AI codebase tour + setup automationCody / Cursor / internal RAG agent over repos
"IT tickets for password resets are 40% of volume."Self-service IT botMoveworks / Atomicwork / custom Slack agent
"Shadow IT is everywhere."SaaS discovery + spend managementZylo / Torii / Vendr + LLM categorization
"Procurement is a forest."Intake-to-pay workflow with policy routingZip / Tropic + LLM contract-summary agent
"Vendor renewals just happen."Contract repository + renewal-alert agentVendr / Tropic + LLM negotiation-prep agent
"We don't know what we're paying for."SaaS spend visibility + usage-based rightsizingZylo / Torii + usage-pull agent
"I'm the bottleneck for everything."EA agent + delegation trackingLindy / Cora / custom agent for inbox, scheduling, follow-up
"Time tracking is hated."Passive time-capture from calendar + toolsReclaim / Memtime + LLM categorization
"My calendar is on fire."AI scheduler with preference learningReclaim / Motion / Clockwise
"I have no visibility into what my team is actually doing."Cross-tool activity dashboard + LLM weekly synthLinear + GitHub + Slack + LLM exec digest
"Customers ask for the same thing in DMs."DM intake → ticket + auto-FAQSlack/Telegram listener + LLM categorizer + Notion FAQ writer
"Insurance claims sit for weeks." (insurance ops)"Claims triage + document extractionHyperscience / Sensible + claims-LLM agent + Guidewire write
"Field techs forget to update job status." (field services)"Mobile-first job update + voice-note → CRMServiceTitan / Jobber + voice-transcription LLM
"Patient intake forms are a disaster." (healthcare ops)"Pre-visit intake + EHR sync (HIPAA-scoped)Phreesia / Tebra + form-LLM with PHI controls
"Construction RFIs take forever to answer." (construction)"RFI auto-draft from drawings + specProcore + RAG over project docs + LLM draft agent
"Estimates eat our weekends." (trades / contracting)"AI estimator from photos + scope notesBuildxact / custom + vision-LLM takeoff agent
"Restaurant scheduling is chaos." (hospitality)"AI shift scheduler with demand forecast7shifts / Homebase + forecasting agent
"Legal intake is a Google form graveyard." (law firms)"Client-intake agent + matter creationClio + LLM intake bot + conflict-check agent
§12

Discovery question bank

decay: 12mo

The intake script the audit agent runs. Conditional branches by function. Wrap in <sysprompt> in deployment.

System prompt block — intake agent

You are the AutomationAudit intake agent. For each function the prospect operates in, run the matching block below. Always quantify with numbers. Always end with the red-flag check (§16) before recommending automation.

Sales

Openers (ask all)

  1. Walk me through what happens between a lead landing and a closed deal.
  2. What's a typical week for your reps from Monday to Friday?
  3. If we deleted one task from every rep's day tomorrow, what should it be — and why?
  4. Where do deals usually slip or die?
  5. What's the single number — pipeline, win rate, cycle time, ramp — your VP gets graded on this quarter?

Conditional branches

  • If they mention manual lead routingWho does the routing? Rules today? How long does a lead sit before first touch?
  • If they mention CRM hygieneWhich CRM? Last audit when? Which fields stalest? Who owns?
  • If they mention slow follow-upCurrent speed-to-lead? Target? What blocks reps from being faster?
  • If they mention forecasting painHow does CRO arrive at the number — spreadsheet, Clari, gut?
  • If they mention notes don't make CRMCall recording in use? Which tool? % of calls actually logged?
  • If they mention proposal/quote delaysWalk through a quote from request to signed. Longest pause?
  • If they mention outbound at volumeHow personalized is each touch — templates, sequences, manual?

Quantify (force numbers)

  • Leads/week? Avg response time vs goal?
  • Calls per rep per day? How many logged?
  • Time on CRM update per rep per day × reps × workdays.
  • Win rate? Cycle time? ACV? Stage-to-stage slip rate?

Tool stack probes

  • CRM — Salesforce, HubSpot, Pipedrive, Close, Attio, custom?
  • Engagement — Outreach, Salesloft, Apollo, Clay?
  • Call intel — Gong, Chorus, Fathom, Granola?
  • Enrichment — Clearbit, Apollo, ZoomInfo, Cognism?
  • CPQ — Salesforce CPQ, DealHub, PandaDoc, Google Docs?

Red flags (stop or scope down)

  • Who owns CRM data quality? — if 'no one,' flag.
  • Documented sales process? — if no, scope to one stage only.
  • How many CRMs? — if &gt;1, data consolidation pre-req.
  • Anyone hired/fired around this in 6 mo? — political risk.

Customer Success

Openers (ask all)

  1. Walk me through a customer's first 90 days post-sale.
  2. How do you decide which customer to call today?
  3. What does a QBR look like end-to-end?
  4. How do you find out a customer is unhappy — before they tell you?
  5. What's your NRR target and the gap today?

Conditional branches

  • If they mention onboarding painWritten onboarding plan? Who tracks? Longest phase?
  • If they mention renewals sneaking upHow far in advance does renewal motion start? Who triggers?
  • If they mention health is gut feelWhat signals — usage, support volume, exec engagement? Where do they live?
  • If they mention QBR overheadHow long does QBR deck take? Who builds? % of QBRs that actually happen?
  • If they mention churn surprisesWalk last 3 churns. Were there leading indicators in hindsight?

Quantify (force numbers)

  • Accounts per CSM?
  • QBR prep hrs/week × QBRs/qtr.
  • Renewal trigger to close time?
  • % accounts with updated success plan?

Tool stack probes

  • CS platform — Gainsight, Catalyst, Vitally, ChurnZero, Planhat?
  • Product analytics — Mixpanel, Amplitude, Heap, Pendo?
  • Where does account data live — CRM, CSP, sheet?

Red flags (stop or scope down)

  • Customer owner clear post-sale? — if not, flag.
  • Renewal forecast exists? — if no, scope renewal visibility first.
  • Product analytics wired up? — if no events, automation has nothing to read.

Customer Support

Openers (ask all)

  1. What's a typical day in your queue?
  2. What % of tickets are repeat questions?
  3. First-response and resolution time today?
  4. Where does your KB live and how stale is it?
  5. What gets escalated to eng/product, how often?

Conditional branches

  • If they mention repeat questionsTop 5? KB article for each? Current?
  • If they mention overnight pileup24/7 staffed? Overnight vs day volume?
  • If they mention macro decayWho owns macros? Update cadence?
  • If they mention agent KB-search overheadTime per ticket to find answer?
  • If they mention escalation overflow% tickets escalating? Tier-1 confidence threshold?

Quantify (force numbers)

  • Tickets/day, /week?
  • AHT vs target?
  • Deflection rate (self-serve)?
  • % of volume in top 10 categories?

Tool stack probes

  • Helpdesk — Zendesk, Intercom, Front, Help Scout, Freshdesk?
  • Existing AI — Fin, Zendesk AI, Forethought?
  • KB — Notion, Confluence, helpdesk-native, scattered?

Red flags (stop or scope down)

  • KB exists? — if no, build content first.
  • Public or behind login? — affects RAG strategy.
  • Regulated industry? — scope-down on what AI can answer.

Finance / Accounting / FP&A

Openers (ask all)

  1. Walk me through month-end close day 1 → final.
  2. Where do you lose most time — recon, AP, AR, reporting?
  3. Biggest manual lift between systems?
  4. What does your CFO ask for that always takes too long?
  5. Where do errors usually slip in?

Conditional branches

  • If they mention AR/collectionsOpen invoices? DSO? Who follows up, how?
  • If they mention AP/invoicesWhere do invoices arrive? Who codes? 3-way match by hand?
  • If they mention reconciliationWhich accounts? Auto vs manual? Exception handling?
  • If they mention close speedDays? Longest single task?
  • If they mention reportingRecurring report? Who builds, how long?
  • If they mention expense mgmtCard program? Manual reimbursements? Approval flow?

Quantify (force numbers)

  • DSO, DPO, close days, # recons/month
  • Hours/month on AR follow-ups
  • % of invoices auto-matched
  • # of JEs/close

Tool stack probes

  • ERP/GL — QuickBooks, NetSuite, Sage Intacct, Xero, Dynamics?
  • AP — Bill.com, Ramp Bill Pay, Tipalti, Stampli?
  • AR — Upflow, Chaser, Versapay?
  • Close — FloQast, Numeric, BlackLine?
  • FP&A — Mosaic, Pigment, Cube, Adaptive?

Red flags (stop or scope down)

  • Chart of accounts clean? — if messy, automation accelerates the mess.
  • Books current? — if behind, fix first.
  • Controller or just outsourced bookkeeper? — affects safe scope.
  • Audited or SOX? — if yes, advisory + HITL only in critical controls.

Operations / RevOps / BizOps

Openers (ask all)

  1. Most painful handoff in the business right now?
  2. Worst spreadsheet — the one holding everything together?
  3. Two systems that should be talking but aren't?
  4. What's a 'status update' look like today?
  5. Monday morning report you wish you didn't build?

Conditional branches

  • If they mention handoff failuresBetween which teams? What gets dropped? How caught today?
  • If they mention spreadsheet hellHow many people touch? How often breaks? Single source of truth elsewhere?
  • If they mention tool sprawlHow many SaaS tools? Which integrate? Where copy/paste?
  • If they mention manual reportingCadence? Audience? Decisions driven?
  • If they mention undocumented processRunbook anywhere? Or in heads?

Quantify (force numbers)

  • Hours/week on handoff tracking, report assembly
  • # tools in stack, integrations vs manual
  • Fields copied per transaction × transactions/day

Tool stack probes

  • iPaaS — Workato, Tray, Zapier, n8n, Make?
  • Work mgmt — Asana, Linear, Monday, ClickUp, Jira?
  • Data layer — warehouse? Reverse ETL? dbt?
  • Reporting — Looker, Mode, Hex, Sheets?

Red flags (stop or scope down)

  • Data warehouse exists? — if no & analytical need, scope it as pre-req.
  • Ops owner? — if none, change-mgmt risk.
  • Recent reorg in 90d? — scope down.

Marketing

Openers (ask all)

  1. How does a campaign go from idea to live?
  2. Where do creative requests get stuck?
  3. How do you attribute pipeline back to marketing?
  4. Recurring report you wish you didn't build?
  5. Content production pipeline?

Conditional branches

  • If they mention creative intake chaosWhere requests come — form, Slack, email? Triage owner?
  • If they mention attribution gapsFirst-touch, multi-touch, none? Data home? Who runs it?
  • If they mention content slownessTypical lifecycle — brief→draft→review→publish?
  • If they mention newsletter/lifecycleToday's automation? Open/click rates? Personalization?

Quantify (force numbers)

  • Requests/week, time-to-fulfill, time-to-publish
  • % of pipeline attributed
  • Content pieces/month, hours per piece

Tool stack probes

  • Marketing automation — HubSpot, Marketo, Customer.io, Iterable, Braze?
  • CMS — Webflow, WordPress, Sanity, Contentful?
  • Attribution — GA4, Bizible, Dreamdata, HockeyStack?

Red flags (stop or scope down)

  • Marketing & sales on same CRM? — if no, scope sync first.
  • Brand guidelines documented? — if no, AI content goes off-rails.

HR / People / Recruiting

Openers (ask all)

  1. Noisiest part of your hiring funnel?
  2. Most repetitive HR ticket?
  3. Onboarding a new hire from offer → day 30?
  4. How do performance reviews actually run?
  5. Payroll / comp / benefits admin time per cycle?

Conditional branches

  • If they mention recruiting black holeWhere do candidates wait longest — source-to-screen, screen-to-offer?
  • If they mention ghosting candidatesWho owns comms? SLA?
  • If they mention onboarding ad-hocChecklist exists? Owner? Tech vs people task split?
  • If they mention review overheadCadence, inputs, hours per report?
  • If they mention employee ticketsInbound channels, volume, top categories?

Quantify (force numbers)

  • Open reqs, time-to-fill, candidates per req
  • Onboarding completion rate, time-to-productive
  • HR ticket volume/month, top 5 categories

Tool stack probes

  • ATS — Greenhouse, Ashby, Lever, Workable?
  • HRIS — Rippling, Gusto, Deel, BambooHR, Workday, ADP?
  • Performance — Lattice, 15Five, Culture Amp, Leapsome?
  • Comp — Pave, Figures, Carta?

Red flags (stop or scope down)

  • Job ladder/leveling documented? — if no, AI-assisted reviews premature.
  • HR records consistent across systems? — if no, scope cleanup first.
  • Sensitive data flows (comp/PHI/PII)? — extra scoping.

Legal / Compliance / GRC

Openers (ask all)

  1. Where does legal slow business most today?
  2. % of contracts on playbook vs custom redline?
  3. How do you handle vendor security questionnaires?
  4. Compliance evidence collection at audit time?

Conditional branches

  • If they mention NDA/MSA delaysAvg turnaround? Stuck where? Self-serve possible?
  • If they mention questionnaire loadHow many/qtr? Who fills? Where answers live?
  • If they mention audit prepContinuous controls or fire drill? Evidence repo exists?
  • If they mention playbook coverageDocumented? Maintainer?

Quantify (force numbers)

  • Contracts/month, avg turnaround, % auto-executed
  • Questionnaires/qtr, hours each
  • Audit prep hrs/year

Tool stack probes

  • CLM — Ironclad, LinkSquares, Lexion, ContractWorks, DocuSign CLM?
  • Compliance — Drata, Vanta, Secureframe?
  • Questionnaires — Loopio, Responsive, Conga?

Red flags (stop or scope down)

  • GC in-house or fractional? — affects velocity.
  • Regulated industry? — tight scope; advisory mode.
  • Privilege concerns sending docs to LLM? — must scope private deploy.

Engineering / IT / DevOps

Openers (ask all)

  1. Most annoying IT/eng ticket that won't go away?
  2. Where does PR / code review get stuck?
  3. Time for a new engineer to ship their first PR?
  4. Incidents — runbook, tribal, both?
  5. IT request volume?

Conditional branches

  • If they mention L1 ticket volumeTop 5 categories, % of total, avg resolution time
  • If they mention PR slownessReview turnaround, tooling, bottleneck people
  • If they mention incident responseMTTR, runbook coverage, oncall rotation
  • If they mention onboarding lagSetup checklist? Internal docs? Codebase tour?
  • If they mention shadow ITSaaS visibility? Approvals?

Quantify (force numbers)

  • Tickets/week, % auto-resolved
  • PR cycle time, PRs/week
  • MTTR, incidents/month
  • Time-to-first-PR

Tool stack probes

  • Source — GitHub, GitLab, Bitbucket?
  • Observability — Datadog, NewRelic, Honeycomb, Grafana?
  • Incident — PagerDuty, Rootly, FireHydrant, Incident.io?
  • IT helpdesk — Jira SM, Atomicwork, Moveworks, Slack?
  • SaaS mgmt — Zylo, Torii, BetterCloud?

Red flags (stop or scope down)

  • Engineers willing to use AI agent? — culture check.
  • Sensitive code / regulated data in repos? — what models can touch.
  • Oncall burnout? — careful adding agent noise.

Exec / Owner-led SMB

Openers (ask all)

  1. Three things stealing the most hours from your week?
  2. What can only you do — and what shouldn't only you do?
  3. Walk me through a normal Monday.
  4. What do you wish you could see daily that you don't?
  5. Where do you trust your team, and where don't you?

Conditional branches

  • If they mention inbox overloadVolume? % needing personal response? EA?
  • If they mention calendar chaosWho manages? Tools? Recurring meetings to cut?
  • If they mention lack of visibilityWhat single metric would tell you Monday morning the business is fine?
  • If they mention approval bottleneckWhat approvals can be policy-based vs judgment?
  • If they mention customer DMsVolume, channels, triage possible?

Quantify (force numbers)

  • Hrs/week on email, calendar, approvals, customer DMs, status meetings
  • # recurring meetings, # approvals/week

Tool stack probes

  • Email — Gmail, Outlook, Superhuman, Shortwave?
  • Calendar — Reclaim, Motion, Clockwise, Cal.com?
  • Comms — Slack, Telegram, iMessage, WhatsApp, SMS?

Red flags (stop or scope down)

  • Anyone on the team side who could own the agent? — if no, system rots.
  • SOPs documented? — if no, scope SOPs first.
  • Prior automation attempt? — uncover failure pattern.
§13

Automation readiness scoring rubric

decay: 12mo

Score each candidate process 1–5 on seven axes. Composite threshold determines audit recommendation.

Scoring rubric — to embed in the audit agent

For each candidate process, output a row with these seven scores (1–5 each, 5 best) plus rationale. Composite = sum, max 35. Recommend automation if composite ≥ 22AND no axis <2. Otherwise, route to anti-list or fix-process-first.

AxisWhat it measuresSignals that score 5Signals that score 1
FrequencyHow often the process runs≥ 5×/day per user; multiple users≤ 1×/month, one user
Time costLoaded $ / month spent on this manually≥ $2,000/mo recoverable< $200/mo recoverable
AmbiguityHow well-specified the work isDeterministic rules cover > 90% of casesEach case is a judgment call
Stakes-of-errorCost of a wrong actionLow (internal report, easily reversed)High (legal, financial, customer-facing, irreversible)
Data availabilityRequired inputs already exist in systemsAll inputs in APIs / structuredInputs in heads or scattered docs
Tool-fitCurrent stack supports automation cleanlyNative APIs, mature integration platformWalled-garden tool, no API, no webhooks
Change-mgmt costEffort to get humans to adopt new flowSole owner who wants the helpMulti-team rollout, prior failed attempts, no champion

Composite interpretation:
≥ 28 — recommend full agent build, rank top of report.
22–27 — recommend with caveats; spec HITL gates.
16–21 — recommend deterministic workflow (Zapier/n8n) instead of agent.
11–15 — fix-process-first; surface SOP or data hygiene need.
≤ 10 — leave human; place on anti-list with reasoning.

§14

Decision tree — pick the right tool

decay: 12mo

Given a scored process, decide between: full agent, deterministic workflow, RPA, fix-process-first, or leave-human.

Decision tree — system prompt block

  1. Is composite ≤ 10 or any axis = 1? → leave-human. Recommend it goes on the anti-list with reasoning.
  2. Is data availability ≤ 2? → fix-process-first. Recommend a data-consolidation / SOP-extraction project as Phase 1. Don't sell the agent yet.
  3. Is ambiguity ≥ 4 AND change-mgmt ≥ 4? → deterministic workflow (Zapier / n8n / Make). Lower ceiling on failure modes; faster adoption.
  4. Is the system fully closed (no API/webhook) but UI is stable? → browser-automation agent (Stagehand + Browserbase) only when there's no other path. Tag HITL.
  5. Is the task primarily document-in / structured-out (PDF / invoice / contract clause)? → Claude API + Skills (or Gumloop for batch). Default to deterministic for high-volume.
  6. Is the task conversational and customer-facing (support, outbound)? → buy the deflection layer (Intercom Fin / Zendesk AI / Retell). Don't internal-build.
  7. Is the task multi-step, multi-tool, with state across steps? → LangGraph + Temporal (or Mistral Workflows-pattern). Anchor durability before any agent logic.
  8. Is the task long-horizon (> tens of minutes per run)? → re-scope to smaller checkpoints. Long-horizon at scale is not reliable in 2026.
  9. Default fallback if ≥ 22 composite, no blockers: full agent with HITL on send/spend/legal/customer-visible writes.
§15

Recommendation deliverable template

decay: 12mo

What an audit output looks like. Drop into the audit agent's report-generation prompt verbatim.

Recommendation template — emit one per top-5 process

# {{process_name}}
**Composite readiness:** {{composite}}/35 · **Confidence:** {{confidence}}%
**Owner / role:** {{owner_role}}
**Status:** {{status}}
**Reasoning:** {{one_sentence}}

## What you're doing today
{{manual_workflow_3_to_6_steps}}

## What you'll save
- Time: {{time_saved}} hrs/mo
- Money: ${{cost_saved}}/mo ({{cost_saved_assumption}})
- Annualized: ${{cost_saved_year}}

## How we'd build it
**Stack:** {{tools_needed}}
**Approach:** {{automation_approach}}
**Reliability today:** {{reliability_today}}/5 — {{reliability_rationale}}
**HITL:** {{hitl_mode}} — {{hitl_reason}}
**Estimated build:** {{build_hours}} hours · {{build_calendar_days}} calendar days

## Risks & failure modes
- {{primary_failure_mode}}
- {{mitigation}}

## What you'd own going forward
- {{maintenance_estimate}} per month · {{maintenance_owner}}

## Choose your path
- [ ] DIY — full build instructions inside
- [ ] We build for one shot: **${{one_shot_price}}**
- [ ] We build + maintain monthly: **${{monthly_price}}/mo** (recommended)

Report header block (above all top-5 items)

# Audit for {{company_name}}, {{role}}
**Total recoverable:** ${{total_savings}}/mo · {{total_hours}} hrs/mo back
**Top 5 automations · {{anti_count}} traps avoided · {{wedge_count}} wedge plays**
Generated {{date}} · Confidence-weighted · See methodology at /research
§16

Red flags during intake

decay: 12mo

Signals to scope-down or walk away. Walk-away signals override everything else.

Red flag detector — system prompt block

  1. "We don't have one source of truth for X." → scope-down to data consolidation Phase 1, or walk if they refuse.
  2. "The last team / agency got fired." → political risk. Get explicit on what failed, why, and who is the new sponsor. If you can't get a straight answer, walk.
  3. "We're in healthcare/store PHI / process payments / under SOX." → scope down. Advisory-only, or BAA-eligible infra (Azure OpenAI w/ BAA, AWS Bedrock, on-prem).
  4. "We need it deployed by Friday." → scope-down to one demo-able slice; set realistic 6–8 weeks; if they won't move, walk.
  5. "We tried [Zapier / a chatbot] and it failed." → spend 20 min on the post-mortem before any pitch.
  6. "Our team will adopt this — they're excited." → ask for proof (existing adoption rates, training plans, change champion). If absent, pilot with one team only.
  7. "Our process is too unique to automate." → probe for the one stable, repeated slice. Scope there or walk.
  8. "We don't want HITL — we want full automation." → educate on failure modes. If they refuse and stakes are high, walk.
  9. "Our IT team will help build it." → ask for named engineer with allocated hours. If vague, scope to end-to-end delivery.
  10. "Legal needs to approve everything before we move."→ map approval path, get named contact & turnaround estimate. Scope contract terms minimally.
  11. No clear process owner. → insist on a single named owner with decision authority before signing. If they can't name one, walk.
  12. No data inventory. → scope a paid 1–2 week discovery. Don't promise outcomes pre-mapping.
  13. No KPI for the process. → define one with them on the call. If they can't / won't commit, scope down or walk.
  14. Regulated industry + no compliance lead. → walk unless we have compliance expertise in-house.
  15. Multiple stakeholders openly disagree on the goal. → insist on a single decision-maker. Defer until they align internally.
  16. Recent reorg (< 90 days). → scope down to quick win that survives reorg risk, or wait.
  17. "We just need an MVP" with no defined outcome. → force a measurable target ("X goes from Y to Z in N weeks").
  18. "We're talking to 5 other agencies." → if they'll decide on lowest price, walk.
  19. "Process is in someone's head." → Phase 1 = SOP extraction. Charge for it. Or walk if they refuse.
  20. "We just got new leadership." → confirm new leader is sponsor. If not, wait.
  21. "We don't have budget approved." → real budget conversation before discovery. If they won't, walk.
  22. "Our customers are very different — every account is special." → scope to internal-ops automation, not customer-facing.
  23. "Can you guarantee X% accuracy?" → educate on evals, confidence thresholds, HITL. If they won't engage, walk.
  24. "Can it replace [person]?" → reframe to "augment + redirect." Be honest about what AI can/can't do for that role.
  25. "We want to own the IP / build in-house eventually." → price for transfer or be explicit on licensed vs transferred. Don't be surprised.
  26. Founder/CEO is on the call alone — no operator counterpart. → require an operator in discovery #2. If refused, scope tiny or walk.
  27. "We tried hiring and couldn't find anyone." → probe: is the role judgment-heavy (scope down) or repetitive (good fit)?
  28. "We don't have observability / logs / event tracking." → scope eval infra and instrumentation in Phase 1. Charge for it.
  29. Procurement-led conversation, no business owner present. → request a business stakeholder. If denied, walk.
  30. "We don't have time for discovery — just build it." → walk, or charge a non-refundable discovery fee that filters them out.
§17

Pricing & business reality

decay: 6mo

What AI automation consultancies actually charge in 2025-2026. What converts. Cycle realities.

To populate from synthesis

The pricing-and-business research returns from a dedicated agent (raw-business-pricing.md). Synthesized into this section: shop benchmarks, engagement structures, hourly/per-build rates, sales cycle reality, free-audit conversion benchmarks, and the competitive map. Until that lands, the published number to anchor is Zapier Experts hourly $75–$250 zapier.com/experts, Make.com partner agency engagements $1k–$15k/build, and productized AI ops subscription $99–$2,000/mo range across emerging shops on r/AI_Agents.

§18

What I can't verify yet · decay-watch

decay: 1mo

Open questions. Honest holes. What gets stale fastest.

Holes in this draft

  • Unit economics for the audit itself are modelled, not measured. Free-audit → paid conversion rate is a target, not a benchmark from a comparable shop. Need 100 audits to validate.
  • Voice agent cost in production assumes Retell at $0.05/min — confirm at actual call durations (we expect drift to ~$0.08–$0.10/min with multimodal LLM costs included).
  • Wedge-tier process count. 30+ is the target; the true number is whatever survives sales contact. Some never-automated entries will fail sales appeal; some will reveal they're rarely-automated by a niche vendor we missed.
  • Pricing for "Pro Audit" ($499) is untested. The $99/mo subscription will validate first; the deeper paid audit comes second.
  • Customer profile. "25–150 person services businesses" is the bet; validation requires 20+ paid customers before we know if it holds vs e-comm or solo SaaS.

Decay watch — re-run these first

  • 1 month:capability matrix & vendor map (§6, §7). New SOTA every quarter; reliability scores drift.
  • 3 months: failure-mode list (§9). Prompt-injection surface expands; mitigations mature.
  • 6 months: process database (§3), wedge tier (§4), anti-list (§5), pricing-reality (§17). Wedges close when other vendors copy us.
  • 12 months: thesis (§1), audit-job spec (§2), role pains (§10), discovery bank (§12), scoring rubric (§13), decision tree (§14), template (§15), red flags (§16). Stable, but check after first 100 paid engagements.
  • Stable: operator ↔ vendor glossary (§11). Language drifts slowly; the underlying pains are decade-stable.

Source bibliography: /research/sources.md. Raw research files live in /research/raw-*.md for traceability.