This project was built with extensive AI assistance, as encouraged by the project guidelines. AI tools were used across every phase: planning, code generation, deployment, debugging, and documentation. Below is an honest account of what each tool contributed — including what worked well and what required iteration.
Role: End-to-end project orchestrator — planning, code generation, deployment automation, debugging, and documentation.
What it did well:
- Project planning: Generated a comprehensive project plan with architecture decisions, folder structure, and implementation sequence before writing any code.
- Synthetic corpus generation: Created 10 coherent, cross-referenced policy documents for the fictional "Acme Corp" company. Documents are internally consistent — the PTO policy references the holiday schedule, the expense policy references the remote work equipment stipend, and policy IDs cross-reference correctly.
- Full-stack code generation: Produced all Python source code (Flask app, LangChain RAG pipeline, ChromaDB ingestion, guardrails, SSE streaming), frontend code (HTML/CSS/JS chat interface with real-time streaming and dynamic citation linking), test suites (31 tests), CI/CD workflows, and evaluation scripts.
- Deployment migration: Diagnosed Render OOM root cause (runtime re-ingestion), migrated to DigitalOcean App Platform using
doctlCLI, rewrote.do/app.yaml, and verified the build/runtime split. - Citation system architecture: Designed the "Dynamic Client-Side Source Linking" approach — emitting source metadata before the text stream so the frontend can badge/link citations regardless of LLM formatting variation.
- Documentation: Created and maintained
README.md,design-and-evaluation.md,ai-tooling.md,deployed.md, andAGENTS.md.
What required iteration:
- Initial test assertions needed minor adjustments (e.g., string length boundary in truncation test was off by ~11 characters).
- ChromaDB API compatibility required investigation —
list_collections()return type changed between v0.5.x and v0.6.x; fixed by removing.nameattribute access. - Citation robustness required multiple rounds: ASCII bracket prompt rules → full-width bracket fallback → frontend dynamic repair for any bracket style.
- Render → DigitalOcean migration required researching DO buildpack Python version support after
3.12.13failed to download.
Role: Used for complex architectural decisions and debugging sessions where deep reasoning was needed.
What worked well:
- Excellent at reasoning through multi-step problems (e.g., diagnosing why ChromaDB ingest was OOMing on Render vs. build-time).
- Strong at producing well-structured, nuanced technical explanations — useful for
design-and-evaluation.mdrationale sections. - Best model for "why is this failing" debugging when the root cause wasn't obvious.
What didn't work as well:
- Slower than Sonnet for straightforward code generation tasks where reasoning depth wasn't needed.
- Occasionally over-explained when a concise code edit was all that was required.
Role: Primary model for most coding sessions — fast, reliable, and well-calibrated for technical tasks.
What worked well:
- Fastest high-quality code generation of any model used; handled Flask SSE implementation, frontend JS streaming, and test refactors cleanly.
- Excellent at surgical edits — understood "change only this, preserve everything else" instructions reliably.
- Best model for documentation writing: README updates, AGENTS.md, transcript drafting.
- Strong at iterating on UI/UX changes (Quantic branding, skeleton loaders, full-width layout) with minimal back-and-forth.
What didn't work as well:
- On very long multi-file refactors, occasionally lost track of context from earlier in the session.
- Less strong than Opus on novel architectural decisions requiring first-principles reasoning.
Role: Used for the full UI/UX redesign, web-grounded research, documentation review, and cross-checking decisions against current best practices.
What worked well:
- UI/UX redesign: Drove the complete frontend overhaul — Quantic branding (maroon/gold color scheme), full-width responsive layout, skeleton loaders during retrieval, improved message bubbles, and the source citation badge styling. Produced clean, self-contained HTML/CSS/JS changes with minimal back-and-forth.
- Citation badge UI: Helped finalize the visual design of the dynamic source badges — styling, hover states, and placement within the streamed response.
- Excellent for web-grounded queries (e.g., confirming FastEmbed MTEB benchmark scores, DigitalOcean buildpack behavior, shields.io badge syntax).
- 1M token context window was useful for pasting large file collections and asking "is anything inconsistent?".
- Strong at identifying stale references across documentation (caught Groq/MiniLM/Render references still present in README and design doc).
What didn't work as well:
- Less precise on surgical code edits compared to Claude models outside of UI work.
- Occasionally verbose in answers where a single sentence would suffice.
Role: Real-time confirmation of technology decisions, API compatibility, and best practices.
Specific uses:
- Verified OpenRouter and Groq free tier availability and rate limits
- Found open-source HR policy templates (OpenGov Foundation, Center for Open Science) as structural references for the synthetic corpus
- Confirmed ChromaDB v0.6.x API changes and FastEmbed MTEB retrieval benchmark scores
- Researched DigitalOcean App Platform buildpack behavior and Python runtime version support
- Looked up shields.io badge syntax and GitHub topic tag best practices
| Component | AI Contribution | Human Review |
|---|---|---|
| Project plan & architecture | 95% AI-generated | Reviewed and approved |
| Policy corpus (11 docs) | 100% AI-generated | Reviewed for consistency |
| Python source code | 95% AI-generated | Tested and verified |
| Frontend (HTML/CSS/JS) | 100% AI-generated | Visual review |
| Tests (31 cases) | 95% AI-generated | Run and verified |
| CI/CD workflows | 90% AI-generated, 10% adapted from reference | Verified against reference |
| Evaluation questions | 100% AI-generated | Reviewed against corpus |
| Documentation | 95% AI-generated | Reviewed and edited |
| Deployment migration (DO) | 90% AI-generated | CLI commands reviewed before execution |
Where AI tools excelled:
- Scaffolding: Rapidly creating project structure, boilerplate, and configuration files
- Content generation: Creating realistic synthetic data (policy documents) with internal consistency
- Test writing: Generating comprehensive test suites that cover edge cases
- Debugging: Diagnosing non-obvious failures (OOM root cause, ChromaDB API version breaks, citation formatting edge cases)
- Documentation: Producing clear, well-structured documentation that tracks architectural changes over time
Where human oversight remained essential:
- Architecture decisions: AI proposes, human validates against requirements and constraints
- Quality verification: Running tests, checking for regressions, verifying deployment
- Security review: Ensuring no secrets are committed, API keys are properly managed
- Business logic validation: Confirming corpus content is internally consistent and covers evaluation domains
- Model selection: Choosing the right AI tool for each task type (reasoning vs. generation vs. research)