🇰🇷 한국어 문서는 README.ko-KR.md를 참고하세요.
The Personal Multi-LLM Review Automation Tool is a CLI-based auxiliary system designed to automate the cross-validation of responses from multiple Large Language Models (LLMs).
In the process of cross-reviewing answers from various LLMs, manually switching between platforms (the "tab-switching hell") causes severe context-switching and fatigue. This tool orchestrates multiple AI models to evaluate a single prompt, logging the answers, metadata, and locally-computed costs into structured formats (JSONL and SQLite).
This project starts as a personal research automation tool to assist with business and investment decisions. Inspired by the philosophy of multi-agent orchestration systems, it has been strictly optimized into a Minimum Viable Product (MVP) for personal use.
- AI Majority Vote ≠ Absolute Truth: The consensus of multiple AIs does not guarantee the truth. The final decision always belongs to the human user.
- Reducing Blind Spots: The true purpose of this tool is not to delegate finding the right answer to AI. Instead, it is designed to automatically reveal missing evidence, counterarguments, and risk factors—effectively minimizing the user's cognitive blind spots before making critical decisions.
- Language: Python 3.14 (100%)
- CLI Framework: Typer (For rapid, type-hinted CLI generation)
- Data Validation: Pydantic (Strict, provider-neutral schemas with
extra="forbid") - Providers: OpenAI ✅, Anthropic ✅, Google/Gemini ✅
- Database: SQLite3 (For local logging and analytical queries)
- Environment & Dependency Management: uv (locked via
uv.lock), withpython-dotenvfor secure API key handling
Dependencies and the Python toolchain are managed by uv. Install it once:
# Windows (PowerShell)
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"# macOS/Linux
curl -LsSf https://astral.sh/uv/install.sh | shThen, from the project root:
uv syncThat single command reads .python-version (3.14) and fetches that interpreter if it is missing, creates .venv/, and installs every dependency at the exact version pinned in uv.lock—there is no separate venv or pip install step. The test and lint tooling (pytest, ruff) lives in the dev dependency group and is installed by default; uv sync --no-dev gives a runtime-only environment.
The commands below use the uv run prefix, which runs them inside that environment without activating it. If you would rather activate the venv (.venv\Scripts\activate on Windows, source .venv/bin/activate elsewhere), drop the prefix.
There is no requirements.txt in the repo—pyproject.toml and uv.lock are the single source of truth. If you need one for a pip-only workflow, generate it from the lockfile:
uv export --no-dev --no-hashes --format requirements.txt -o requirements.txtThe tool reads three keys—OPENAI_API_KEY, ANTHROPIC_API_KEY, GEMINI_API_KEY. All of them are optional: a missing key simply disables that one provider instead of crashing the tool.
Anthropic also accepts ANTHROPIC_AUTH_TOKEN in place of ANTHROPIC_API_KEY—the SDK uses whichever is set, so either one on its own configures the provider.
There are two supported ways to supply them. OS environment variables are read first and always win; .env only fills in what they leave unset.
Recommended—OS environment variables. They live outside the project directory, so they cannot be committed, zipped, or shared along with the repo:
# Windows (PowerShell): setx writes to your user profile. Open a NEW shell afterwards—
# setx does not affect the session you run it in.
setx OPENAI_API_KEY "your-openai-api-key-here"# macOS/Linux: add to ~/.zshrc or ~/.bashrc
export OPENAI_API_KEY="your-openai-api-key-here"Fallback—.env file. Copy .env.example to .env in the project root and fill in the keys you have:
OPENAI_API_KEY=your-openai-api-key-here
ANTHROPIC_API_KEY=your-anthropic-api-key-here
GEMINI_API_KEY=your-gemini-api-key-here.env is gitignored, but it is still a plaintext file inside the repo—the usual leak is a folder that gets zipped or shared, not a commit. Neither method encrypts the key at rest; both are readable by any process running as you.
The .env file is loaded once at startup (resources/env.py, called from resources/__init__.py) before any provider client is constructed, so both methods behave identically for every command.
Validate your setup at any time with:
uv run python -m resources.cli check-envCost is computed locally from data files—no extra API calls. Rates live per provider in config/prices/prices_{openai,anthropic,gemini}.json as USD per 1M tokens. Each file also carries updated_at/source metadata, and dated model IDs can reuse a base entry via alias_of:
{
"provider": "OpenAI",
"currency": "USD",
"unit": "per_1m_tokens",
"updated_at": "2026-05-31",
"source": "https://developers.openai.com/api/docs/pricing",
"models": {
"gpt-4o-mini": {
"input": 0.15,
"cached_input": 0.075,
"output": 0.60
},
"gpt-4o-mini-2024-07-18": {
"alias_of": "gpt-4o-mini"
}
}
}The SQLite file _db/llm_responses.db is created/seeded once from _db/_create_table.sql (run it with the sqlite3 CLI or any client). The tool connects to an existing DB and only ALTERs in missing audit columns—it does not create the tables itself.
_db/_create_table.sql is the single source of truth for the schema (_db/llm_responses.db itself is git-ignored). If you ever delete the database and want to recreate it from scratch, re-seed it from that file—do not restore from an ad-hoc DB-client export, which can silently drift from the tracked schema:
# Recreate an empty, correctly-seeded database
rm _db/llm_responses.db # optional: remove the old file first
sqlite3 _db/llm_responses.db < _db/_create_table.sqlThe statements use CREATE TABLE/INDEX IF NOT EXISTS, so running the file against an existing DB is safe (it only fills in what is missing). In a GUI client (e.g. DB Browser for SQLite), paste the same file into the Execute SQL tab and run it.
Run from the project root as a module (not cd resources). ask and compare make real (paid) API calls and automatically persist the logs (JSONL) and metadata (SQLite).
Ask a single provider/model:
uv run python -m resources.cli ask "<system_prompt>" "<user_question>"
# defaults: --provider openai --model gpt-5.6-luna
uv run python -m resources.cli ask "<system_prompt>" "<user_question>" \
--provider anthropic --model claude-haiku-4-5Compare several models (the core "review" feature)—requires at least one --target/-t provider:model, and the same provider:model may not be repeated. Answers print in the order you listed the targets. Each call gets its own run_id, all tied together by one shared group_id:
uv run python -m resources.cli compare "<system_prompt>" "<user_question>" \
-t openai:gpt-5.6-terra -t anthropic:claude-haiku-4-5Other commands:
uv run python -m resources.cli check-env # validate .env keys (interactive)
uv run python -m resources.cli list-models # list models across configured providers
uv run python -m resources.cli history # show recent calls (newest first)
uv run python -m resources.cli history -n 20 # show the last 20 calls
uv run python -m resources.cli history --group <group_id> # show one comparison's callssystem_promptis an instruction that predetermines how the LLM should respond.user_questionis the message to the LLM (the more specific and clear, the better).
The code is a single layered architecture under resources/, used as a package with package-absolute imports (from resources.schemas import ...) and run from the project root:
cli.py # thin Typer layer: parse args → delegate → render
└─ services/service_ask.py # owns ids (run_id/response_id), builds logs, collects errors, archives
└─ providers/registry.py # name → ChatProvider instance
└─ providers/provider_*.py # per-provider API specifics (openai, anthropic, google)
└─ providers/runner.py # run_chat(): shared call pipeline, provider differences injected as callbacks
└─ count_cost.py / schemas.py / storage_json.py / storage_sqlite.pycli.py— thin Typer layer; parses arguments, delegates to the service layer, renders results.services/service_ask.py— orchestration: mintsrun_id/response_id, resolves providers, builds the audit log, collectscomparefailures as data, and persists.providers/registry.py— the only place that maps a provider name to a concreteChatProvider.providers/provider_*.py— per-provider API specifics (OpenAI, Anthropic, Google).providers/runner.py— the shared chat pipeline (preflight → paid call → parse → best-effort cost → assemble); each provider only supplies_call_apiand_parse_responsecallbacks.schemas.py— strict, provider-neutral Pydantic models (LLMRequest,LLMCallResult,LLMCallLog, etc.).storage_json.py&storage_sqlite.py— data persistence to JSONL and SQLite.count_cost.py— calculates token usage costs locally (inDecimal) without extra API calls.
Adding a provider = add a provider_*.py + one line in registry.PROVIDERS.
The pytest suite lives in tests/ and makes no paid calls—provider parsing is tested with fake SDK responses, and the service layer by injecting fakes into the registry. Run from the project root:
uv run pytestSince this is currently a personal MVP, direct pull requests to the core logic may be limited. However, contributions and forks are highly welcome in the following areas:
- Adding New Providers: Implementing a
LocalLLMProvider(e.g., Ollama), or adding other cloud providers. - Evaluator Prompts: Enhancing the prompts used to detect conflicts and highlight missing citations across model answers.
- Cost Analytics: Creating SQL views or Pandas scripts to analyze model cost-efficiency over time.
Feel free to fork the repository, experiment with local LLM integrations, and open an issue if you discover a robust prompting strategy!
Other Info: I live in S.Korea. I am not very good at English, so please understand that a translator was used to write this README.