Skip to content

Commit 5ae59f9

Browse files
committed
docs: add agentic architecture layer to WhatsApp bot research
Adds section V proposing four agentic layers on top of the deterministic v1: LLM reply parsing (Haiku 4.5 + structured outputs), conversational replies (Sonnet 5 + tool use), agentic admin drafting and chef ranking, and a scheduled ops loop that chases non-responders. Makes the LLM/code boundary explicit — regex and SQL keep the deterministic work, the model only handles judgment calls. Deterministic path stays the fast path so an API outage degrades to v1 behavior. Includes model choices with current pricing, per-layer cost estimates (under $1/month total), a sequenced build order, and human approval gates on every outbound-message path.
1 parent 7f5802c commit 5ae59f9

1 file changed

Lines changed: 258 additions & 0 deletions

File tree

‎docs/whatsapp-bot-research.md‎

Lines changed: 258 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -8,6 +8,9 @@
88
Research spike to pick a WhatsApp API provider for the catering-opportunity bot and
99
propose the infrastructure. No bot code in this PR.
1010

11+
Sections I–IV cover the provider decision and a deterministic v1. Section V proposes the
12+
agentic layer built on top of it.
13+
1114
---
1215

1316
## TL;DR
@@ -26,6 +29,9 @@ propose the infrastructure. No bot code in this PR.
2629
plus a `wa.me` deep link.
2730
4. **Pricing changes in two days.** Service messages and in-window utility templates
2831
become billable **October 1, 2026**. Budget accordingly.
32+
5. **An agentic layer is proposed in §V** — LLM reply parsing, conversational replies,
33+
admin drafting, and an autonomous ops loop — layered on top of the deterministic v1,
34+
not replacing it. Adds **under $1/month**. Ship v1 first.
2935

3036
---
3137

@@ -522,6 +528,258 @@ a way to skip it — only a way to not be _blocked_ by it.
522528

523529
---
524530

531+
## V. Agentic architecture
532+
533+
Sections I–IV describe a **deterministic** bot: regex for the ref code, keyword match
534+
for yes/no. That is the right v1 and it should ship first. This section describes the
535+
agentic layer to build on top of it, and — more importantly — **where the line between
536+
LLM and plain code sits**, because getting that wrong is how this becomes expensive and
537+
flaky.
538+
539+
### The governing rule
540+
541+
> Use the model for judgment. Use code for everything deterministic.
542+
543+
Concretely:
544+
545+
| Job | Owner | Why |
546+
| ------------------------------------------------- | ----------------------------------- | --------------------------------------------------------------------------- |
547+
| Parse `OB12` out of a message | **code** (regex) | Deterministic. An LLM here is slower, costlier, and can hallucinate a code. |
548+
| Look up chef by phone number | **code** (SQL) | It's a database join. |
549+
| Dedupe on `provider_message_id` | **code** (unique index) | Correctness primitive, not a judgment. |
550+
| Decide whether "cant do sat, sun works?" is a yes | **LLM** | Genuine ambiguity. Regex cannot do this. |
551+
| Draft broadcast copy for a gig | **LLM** | Writing. |
552+
| Rank which chefs to contact | **LLM over code-fetched data** | Judgment, but the _data_ comes from SQL. |
553+
| Decide to send a message | **code** (with human approval gate) | Irreversible side effect. Never let the model do this unsupervised. |
554+
555+
The deterministic path stays the **fast path**: if the regex finds `OB12` and the body
556+
matches `/^(yes|yeah|y|in|i'?m in)$/i`, we save it and never call the model. The LLM is
557+
the **fallback for the messy 30%**, not the default path. This keeps cost near zero and
558+
means an API outage degrades to v1 behavior rather than breaking the bot.
559+
560+
### Layer 1 — Smart reply parsing
561+
562+
Replaces `interest: 'unclear' → needs_review` with structured extraction.
563+
564+
Today's regex fails on nearly every real message: _"cant do sat but sunday works"_,
565+
_"how many ppl again?"_, _"im in if maya is doing it too"_, _"yes but only if it's not
566+
the taco thing again"_. All of those currently land in the triage queue, which means a
567+
human reads them anyway and the bot saved nobody any work.
568+
569+
Use **structured outputs** so the response is schema-valid rather than parsed prose:
570+
571+
```ts
572+
// actions/whatsapp/parse.ts
573+
import Anthropic from "@anthropic-ai/sdk";
574+
575+
const client = new Anthropic();
576+
577+
const REPLY_SCHEMA = {
578+
type: "object",
579+
additionalProperties: false,
580+
required: ["interest", "confidence"],
581+
properties: {
582+
interest: { enum: ["yes", "no", "conditional", "question", "unrelated"] },
583+
confidence: { enum: ["high", "low"] },
584+
availableDates: { type: "array", items: { type: "string" } },
585+
conditions: { type: "array", items: { type: "string" } },
586+
question: { type: "string" },
587+
refCode: { type: "string" },
588+
},
589+
} as const;
590+
591+
export async function parseReply(body: string, context: OpportunityContext) {
592+
const response = await client.messages.create({
593+
model: "claude-haiku-4-5",
594+
max_tokens: 1024,
595+
system: [
596+
{
597+
type: "text",
598+
text: PARSE_PROMPT,
599+
cache_control: { type: "ephemeral" },
600+
},
601+
],
602+
output_config: { format: { type: "json_schema", schema: REPLY_SCHEMA } },
603+
messages: [{ role: "user", content: formatContext(body, context) }],
604+
});
605+
return extractJson(response);
606+
}
607+
```
608+
609+
**Model choice: Haiku 4.5** ($1/$5 per MTok). This is classification with a fixed
610+
schema — the easiest possible LLM task. Opus would be waste. Cache the system prompt so
611+
the instruction block is ~90% cheaper on every call after the first.
612+
613+
**Cost:** a reply is roughly 400 input + 100 output tokens ≈ **$0.0009**. At 120 replies
614+
a month that's **≈ $0.11/month**. Against the WhatsApp bill of ~$0.41 this is noise.
615+
616+
**`confidence: "low"` still routes to `needs_review`.** The LLM shrinks the triage queue;
617+
it does not eliminate human review. A wrong "yes" books a chef for a gig they declined,
618+
so the bar for auto-accepting stays high.
619+
620+
### Layer 2 — Conversational replies
621+
622+
The chef asks _"how much does it pay?"_ and the bot answers from the opportunity record.
623+
624+
This needs **tool use** — the model reads data, it doesn't invent it:
625+
626+
```ts
627+
const tools = [
628+
{
629+
name: "get_opportunity",
630+
description: "Fetch full details for one catering opportunity by ref code.",
631+
input_schema: {/* … */},
632+
strict: true,
633+
},
634+
{
635+
name: "record_response",
636+
description:
637+
"Save this chef's interest. Call only when intent is unambiguous.",
638+
input_schema: {/* … */},
639+
strict: true,
640+
},
641+
];
642+
```
643+
644+
Use the SDK's **tool runner** (`client.beta.messages.toolRunner`) rather than
645+
hand-writing the `while (stop_reason === "tool_use")` loop — its per-turn hooks are
646+
where the approval gate goes.
647+
648+
Three hard constraints:
649+
650+
1. **Read tools are free to run; write tools need a gate.** `get_opportunity` executes
651+
immediately. `record_response` is fine to auto-run (it's reversible and reviewable).
652+
Anything that _sends a message_ or _confirms a booking_ goes through human approval.
653+
2. **The 24-hour window is a hard wall.** The bot can only free-form reply inside the
654+
window opened by the chef's message. Outside it, a template is required — so the bot
655+
physically cannot hold an open-ended conversation days later. Design around this;
656+
don't discover it in production.
657+
3. **Never let the model quote money or commit the org to anything.** Pay rate comes out
658+
of the `opportunities` row verbatim via tool call, never from model generation.
659+
660+
**Model:** Sonnet 5 ($2/$10). Needs more judgment than classification but isn't
661+
reasoning-heavy. Bump to Opus 5 only if Sonnet visibly mishandles real transcripts.
662+
663+
### Layer 3 — Agentic admin side
664+
665+
Admin types _"taco catering next fri, 80ppl, downtown oakland, $600"_ and gets a drafted
666+
opportunity, drafted broadcast copy, and a ranked chef list.
667+
668+
```
669+
admin free-text
670+
↓
671+
[LLM] extract structured opportunity fields ← structured outputs, Sonnet 5
672+
↓
673+
[code] INSERT opportunity, generate ref_code ← deterministic
674+
↓
675+
[code] SQL: chefs matching cuisine, not booked that date, past reliability
676+
↓
677+
[LLM] rank + explain: "Maya — 3 taco gigs, all on time"
678+
↓
679+
[LLM] draft broadcast copy for the template variables
680+
↓
681+
[HUMAN] admin reviews list + copy, edits, approves ← REQUIRED GATE
682+
↓
683+
[code] send
684+
```
685+
686+
The ranking is the one genuinely interesting use of the model: it weighs cuisine match,
687+
past reliability, how recently someone got offered work, and fairness of distribution —
688+
fuzzy tradeoffs that are miserable to encode as SQL `ORDER BY`. But the **candidate set
689+
is fetched by SQL**, not chosen by the model. The model ranks and explains; it never
690+
queries freely.
691+
692+
**Note this needs schema the current data model doesn't have** — `chefs.cuisines`,
693+
`chefs.reliability_score` (or derived from a `gigs` table), `opportunities.pay_rate`.
694+
Worth adding in the same migration if we're building toward this.
695+
696+
**The approval gate is non-negotiable.** An agent that auto-messages 30 chefs on a bad
697+
extraction is a real-world incident with real people, not a bug report.
698+
699+
### Layer 4 — Autonomous ops loop
700+
701+
A scheduled agent that chases silence, because the failure mode of this whole system
702+
isn't a crash — it's a gig quietly going unfilled because nobody replied and nobody
703+
noticed.
704+
705+
```
706+
daily cron
707+
↓
708+
[code] SQL: open opportunities, event_date within 7 days, response counts
709+
↓
710+
[LLM] for each at-risk gig, decide: nudge / escalate / widen / alert admin
711+
↓
712+
[HUMAN] admin approves the batch ← gate stays
713+
↓
714+
[code] send nudges via template
715+
```
716+
717+
Escalation ladder: silent chefs get one nudge; zero yeses 3 days out alerts the admin and
718+
suggests widening to a backup roster; 24h out with nothing gets flagged urgent.
719+
720+
**Two options for running this:**
721+
722+
- **Vercel cron** hitting a route handler. Fits the existing deployment, no new infra,
723+
and the whole loop is ~50 lines. **Recommended for v1.**
724+
- **Managed Agents** with a scheduled deployment — Anthropic runs the loop and hosts the
725+
sandbox, with session budgets as a hard dollar cap. Worth it only if the loop grows
726+
genuinely multi-step. Overkill for "query, decide, queue nudges."
727+
728+
**Hard rule: one nudge per chef per opportunity, enforced in code** with a unique
729+
constraint, not by prompting the model to remember. An agent that re-nudges on every
730+
cron tick because a retry lost state is how you get your number reported for spam —
731+
which, per section I, is an unrecoverable outcome.
732+
733+
### Cost summary
734+
735+
Assuming ~30 chefs, 4 opportunities/month, ~120 replies:
736+
737+
| Layer | Model | Est. monthly |
738+
| ------------------ | --------- | -------------- |
739+
| 1 — reply parsing | Haiku 4.5 | ~$0.11 |
740+
| 2 — conversation | Sonnet 5 | ~$0.40 |
741+
| 3 — admin drafting | Sonnet 5 | ~$0.15 |
742+
| 4 — ops loop | Sonnet 5 | ~$0.10 |
743+
| **LLM total** | | **< $1/month** |
744+
| WhatsApp (from §I) | | ~$0.41–1.61 |
745+
746+
Estimates, not quotes — output tokens vary. But the order of magnitude holds: **the
747+
agentic layer costs less than the messaging does**, and both are under $5/month. Cost is
748+
not a reason to avoid this. Complexity and failure modes are the real budget.
749+
750+
### Build order
751+
752+
Ship these in sequence, not at once:
753+
754+
1. **v1 deterministic** (§II–III) — regex + keywords. Proves webhook, matching, dedupe.
755+
2. **Layer 1** — swap the classifier. Smallest, highest-value change; shrinks triage.
756+
3. **Layer 3** — admin drafting. Pure upside, human-gated, no inbound risk.
757+
4. **Layer 2** — conversation. Most user-visible risk; needs real transcripts to tune.
758+
5. **Layer 4** — ops loop. Only once there's enough history to know what "at risk" means.
759+
760+
Each step is independently shippable and independently revertable. If layer 2 makes
761+
chefs feel like they're talking to a machine, roll it back without touching 1 or 3.
762+
763+
### What to verify before building this
764+
765+
- **Does Oakland Bloom want a bot that talks back?** A chef expecting a human and getting
766+
an LLM is a trust problem, not a feature. Ask before building layer 2.
767+
- **Disclosure.** Chefs should know they're messaging an automated system. Ethical
768+
baseline, and cheap to do in the template copy.
769+
- **Retention.** Chef messages go to Anthropic's API. Fine for this use case, but say so
770+
out loud to the org rather than assuming.
771+
- **Template approval friction.** Every nudge/escalation variant in layer 4 is a separate
772+
Meta-approved template with a ~24h review. Batch the submissions.
773+
774+
### Links
775+
776+
- [Claude API: tool use](https://docs.anthropic.com/en/docs/build-with-claude/tool-use) ·
777+
[structured outputs](https://docs.anthropic.com/en/docs/build-with-claude/structured-outputs) ·
778+
[prompt caching](https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching) ·
779+
[pricing](https://www.anthropic.com/pricing#api)
780+
781+
---
782+
525783
## Links
526784

527785
**Meta**

0 commit comments

Comments
 (0)