88Research spike to pick a WhatsApp API provider for the catering-opportunity bot and
99propose the infrastructure. No bot code in this PR.
1010
11+ Sections I–IV cover the provider decision and a deterministic v1. Section V proposes the
12+ agentic layer built on top of it.
13+
1114---
1215
1316## TL;DR
@@ -26,6 +29,9 @@ propose the infrastructure. No bot code in this PR.
2629 plus a ` wa.me ` deep link.
27304 . ** Pricing changes in two days.** Service messages and in-window utility templates
2831 become billable ** October 1, 2026** . Budget accordingly.
32+ 5 . ** An agentic layer is proposed in §V** — LLM reply parsing, conversational replies,
33+ admin drafting, and an autonomous ops loop — layered on top of the deterministic v1,
34+ not replacing it. Adds ** under $1/month** . Ship v1 first.
2935
3036---
3137
@@ -522,6 +528,258 @@ a way to skip it — only a way to not be _blocked_ by it.
522528
523529---
524530
531+ ## V. Agentic architecture
532+
533+ Sections I–IV describe a ** deterministic** bot: regex for the ref code, keyword match
534+ for yes/no. That is the right v1 and it should ship first. This section describes the
535+ agentic layer to build on top of it, and — more importantly — ** where the line between
536+ LLM and plain code sits** , because getting that wrong is how this becomes expensive and
537+ flaky.
538+
539+ ### The governing rule
540+
541+ > Use the model for judgment. Use code for everything deterministic.
542+
543+ Concretely:
544+
545+ | Job | Owner | Why |
546+ | ------------------------------------------------- | ----------------------------------- | --------------------------------------------------------------------------- |
547+ | Parse ` OB12 ` out of a message | ** code** (regex) | Deterministic. An LLM here is slower, costlier, and can hallucinate a code. |
548+ | Look up chef by phone number | ** code** (SQL) | It's a database join. |
549+ | Dedupe on ` provider_message_id ` | ** code** (unique index) | Correctness primitive, not a judgment. |
550+ | Decide whether "cant do sat, sun works?" is a yes | ** LLM** | Genuine ambiguity. Regex cannot do this. |
551+ | Draft broadcast copy for a gig | ** LLM** | Writing. |
552+ | Rank which chefs to contact | ** LLM over code-fetched data** | Judgment, but the _ data_ comes from SQL. |
553+ | Decide to send a message | ** code** (with human approval gate) | Irreversible side effect. Never let the model do this unsupervised. |
554+
555+ The deterministic path stays the ** fast path** : if the regex finds ` OB12 ` and the body
556+ matches ` /^(yes|yeah|y|in|i'?m in)$/i ` , we save it and never call the model. The LLM is
557+ the ** fallback for the messy 30%** , not the default path. This keeps cost near zero and
558+ means an API outage degrades to v1 behavior rather than breaking the bot.
559+
560+ ### Layer 1 — Smart reply parsing
561+
562+ Replaces ` interest: 'unclear' → needs_review ` with structured extraction.
563+
564+ Today's regex fails on nearly every real message: _ "cant do sat but sunday works"_ ,
565+ _ "how many ppl again?"_ , _ "im in if maya is doing it too"_ , _ "yes but only if it's not
566+ the taco thing again"_ . All of those currently land in the triage queue, which means a
567+ human reads them anyway and the bot saved nobody any work.
568+
569+ Use ** structured outputs** so the response is schema-valid rather than parsed prose:
570+
571+ ``` ts
572+ // actions/whatsapp/parse.ts
573+ import Anthropic from " @anthropic-ai/sdk" ;
574+
575+ const client = new Anthropic ();
576+
577+ const REPLY_SCHEMA = {
578+ type: " object" ,
579+ additionalProperties: false ,
580+ required: [" interest" , " confidence" ],
581+ properties: {
582+ interest: { enum: [" yes" , " no" , " conditional" , " question" , " unrelated" ] },
583+ confidence: { enum: [" high" , " low" ] },
584+ availableDates: { type: " array" , items: { type: " string" } },
585+ conditions: { type: " array" , items: { type: " string" } },
586+ question: { type: " string" },
587+ refCode: { type: " string" },
588+ },
589+ } as const ;
590+
591+ export async function parseReply(body : string , context : OpportunityContext ) {
592+ const response = await client .messages .create ({
593+ model: " claude-haiku-4-5" ,
594+ max_tokens: 1024 ,
595+ system: [
596+ {
597+ type: " text" ,
598+ text: PARSE_PROMPT ,
599+ cache_control: { type: " ephemeral" },
600+ },
601+ ],
602+ output_config: { format: { type: " json_schema" , schema: REPLY_SCHEMA } },
603+ messages: [{ role: " user" , content: formatContext (body , context ) }],
604+ });
605+ return extractJson (response );
606+ }
607+ ```
608+
609+ ** Model choice: Haiku 4.5** ($1/$5 per MTok). This is classification with a fixed
610+ schema — the easiest possible LLM task. Opus would be waste. Cache the system prompt so
611+ the instruction block is ~ 90% cheaper on every call after the first.
612+
613+ ** Cost:** a reply is roughly 400 input + 100 output tokens ≈ ** $0.0009** . At 120 replies
614+ a month that's ** ≈ $0.11/month** . Against the WhatsApp bill of ~ $0.41 this is noise.
615+
616+ ** ` confidence: "low" ` still routes to ` needs_review ` .** The LLM shrinks the triage queue;
617+ it does not eliminate human review. A wrong "yes" books a chef for a gig they declined,
618+ so the bar for auto-accepting stays high.
619+
620+ ### Layer 2 — Conversational replies
621+
622+ The chef asks _ "how much does it pay?"_ and the bot answers from the opportunity record.
623+
624+ This needs ** tool use** — the model reads data, it doesn't invent it:
625+
626+ ``` ts
627+ const tools = [
628+ {
629+ name: " get_opportunity" ,
630+ description: " Fetch full details for one catering opportunity by ref code." ,
631+ input_schema: {/* … */ },
632+ strict: true ,
633+ },
634+ {
635+ name: " record_response" ,
636+ description:
637+ " Save this chef's interest. Call only when intent is unambiguous." ,
638+ input_schema: {/* … */ },
639+ strict: true ,
640+ },
641+ ];
642+ ```
643+
644+ Use the SDK's ** tool runner** (` client.beta.messages.toolRunner ` ) rather than
645+ hand-writing the ` while (stop_reason === "tool_use") ` loop — its per-turn hooks are
646+ where the approval gate goes.
647+
648+ Three hard constraints:
649+
650+ 1 . ** Read tools are free to run; write tools need a gate.** ` get_opportunity ` executes
651+ immediately. ` record_response ` is fine to auto-run (it's reversible and reviewable).
652+ Anything that _ sends a message_ or _ confirms a booking_ goes through human approval.
653+ 2 . ** The 24-hour window is a hard wall.** The bot can only free-form reply inside the
654+ window opened by the chef's message. Outside it, a template is required — so the bot
655+ physically cannot hold an open-ended conversation days later. Design around this;
656+ don't discover it in production.
657+ 3 . ** Never let the model quote money or commit the org to anything.** Pay rate comes out
658+ of the ` opportunities ` row verbatim via tool call, never from model generation.
659+
660+ ** Model:** Sonnet 5 ($2/$10). Needs more judgment than classification but isn't
661+ reasoning-heavy. Bump to Opus 5 only if Sonnet visibly mishandles real transcripts.
662+
663+ ### Layer 3 — Agentic admin side
664+
665+ Admin types _ "taco catering next fri, 80ppl, downtown oakland, $600"_ and gets a drafted
666+ opportunity, drafted broadcast copy, and a ranked chef list.
667+
668+ ```
669+ admin free-text
670+ ↓
671+ [LLM] extract structured opportunity fields ← structured outputs, Sonnet 5
672+ ↓
673+ [code] INSERT opportunity, generate ref_code ← deterministic
674+ ↓
675+ [code] SQL: chefs matching cuisine, not booked that date, past reliability
676+ ↓
677+ [LLM] rank + explain: "Maya — 3 taco gigs, all on time"
678+ ↓
679+ [LLM] draft broadcast copy for the template variables
680+ ↓
681+ [HUMAN] admin reviews list + copy, edits, approves ← REQUIRED GATE
682+ ↓
683+ [code] send
684+ ```
685+
686+ The ranking is the one genuinely interesting use of the model: it weighs cuisine match,
687+ past reliability, how recently someone got offered work, and fairness of distribution —
688+ fuzzy tradeoffs that are miserable to encode as SQL ` ORDER BY ` . But the ** candidate set
689+ is fetched by SQL** , not chosen by the model. The model ranks and explains; it never
690+ queries freely.
691+
692+ ** Note this needs schema the current data model doesn't have** — ` chefs.cuisines ` ,
693+ ` chefs.reliability_score ` (or derived from a ` gigs ` table), ` opportunities.pay_rate ` .
694+ Worth adding in the same migration if we're building toward this.
695+
696+ ** The approval gate is non-negotiable.** An agent that auto-messages 30 chefs on a bad
697+ extraction is a real-world incident with real people, not a bug report.
698+
699+ ### Layer 4 — Autonomous ops loop
700+
701+ A scheduled agent that chases silence, because the failure mode of this whole system
702+ isn't a crash — it's a gig quietly going unfilled because nobody replied and nobody
703+ noticed.
704+
705+ ```
706+ daily cron
707+ ↓
708+ [code] SQL: open opportunities, event_date within 7 days, response counts
709+ ↓
710+ [LLM] for each at-risk gig, decide: nudge / escalate / widen / alert admin
711+ ↓
712+ [HUMAN] admin approves the batch ← gate stays
713+ ↓
714+ [code] send nudges via template
715+ ```
716+
717+ Escalation ladder: silent chefs get one nudge; zero yeses 3 days out alerts the admin and
718+ suggests widening to a backup roster; 24h out with nothing gets flagged urgent.
719+
720+ ** Two options for running this:**
721+
722+ - ** Vercel cron** hitting a route handler. Fits the existing deployment, no new infra,
723+ and the whole loop is ~ 50 lines. ** Recommended for v1.**
724+ - ** Managed Agents** with a scheduled deployment — Anthropic runs the loop and hosts the
725+ sandbox, with session budgets as a hard dollar cap. Worth it only if the loop grows
726+ genuinely multi-step. Overkill for "query, decide, queue nudges."
727+
728+ ** Hard rule: one nudge per chef per opportunity, enforced in code** with a unique
729+ constraint, not by prompting the model to remember. An agent that re-nudges on every
730+ cron tick because a retry lost state is how you get your number reported for spam —
731+ which, per section I, is an unrecoverable outcome.
732+
733+ ### Cost summary
734+
735+ Assuming ~ 30 chefs, 4 opportunities/month, ~ 120 replies:
736+
737+ | Layer | Model | Est. monthly |
738+ | ------------------ | --------- | -------------- |
739+ | 1 — reply parsing | Haiku 4.5 | ~ $0.11 |
740+ | 2 — conversation | Sonnet 5 | ~ $0.40 |
741+ | 3 — admin drafting | Sonnet 5 | ~ $0.15 |
742+ | 4 — ops loop | Sonnet 5 | ~ $0.10 |
743+ | ** LLM total** | | ** < $1/month** |
744+ | WhatsApp (from §I) | | ~ $0.41–1.61 |
745+
746+ Estimates, not quotes — output tokens vary. But the order of magnitude holds: ** the
747+ agentic layer costs less than the messaging does** , and both are under $5/month. Cost is
748+ not a reason to avoid this. Complexity and failure modes are the real budget.
749+
750+ ### Build order
751+
752+ Ship these in sequence, not at once:
753+
754+ 1 . ** v1 deterministic** (§II–III) — regex + keywords. Proves webhook, matching, dedupe.
755+ 2 . ** Layer 1** — swap the classifier. Smallest, highest-value change; shrinks triage.
756+ 3 . ** Layer 3** — admin drafting. Pure upside, human-gated, no inbound risk.
757+ 4 . ** Layer 2** — conversation. Most user-visible risk; needs real transcripts to tune.
758+ 5 . ** Layer 4** — ops loop. Only once there's enough history to know what "at risk" means.
759+
760+ Each step is independently shippable and independently revertable. If layer 2 makes
761+ chefs feel like they're talking to a machine, roll it back without touching 1 or 3.
762+
763+ ### What to verify before building this
764+
765+ - ** Does Oakland Bloom want a bot that talks back?** A chef expecting a human and getting
766+ an LLM is a trust problem, not a feature. Ask before building layer 2.
767+ - ** Disclosure.** Chefs should know they're messaging an automated system. Ethical
768+ baseline, and cheap to do in the template copy.
769+ - ** Retention.** Chef messages go to Anthropic's API. Fine for this use case, but say so
770+ out loud to the org rather than assuming.
771+ - ** Template approval friction.** Every nudge/escalation variant in layer 4 is a separate
772+ Meta-approved template with a ~ 24h review. Batch the submissions.
773+
774+ ### Links
775+
776+ - [ Claude API: tool use] ( https://docs.anthropic.com/en/docs/build-with-claude/tool-use ) ·
777+ [ structured outputs] ( https://docs.anthropic.com/en/docs/build-with-claude/structured-outputs ) ·
778+ [ prompt caching] ( https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching ) ·
779+ [ pricing] ( https://www.anthropic.com/pricing#api )
780+
781+ ---
782+
525783## Links
526784
527785** Meta**
0 commit comments