A small, privacy-first coding-agent harness written in Rust for Ollama and any OpenAI-compatible endpoint. It's a single binary that runs a real agent loop — tools, plan mode, verification, workspace isolation — against a model endpoint you control. No cloud auth layer, no MCP marketplace, no telemetry, no personality layer. Just the loop, the tools, and your workspace.
I built it to learn Rust and to have a harness that only contains the pieces I actually use. It's small enough to read end-to-end, and it runs fine on a Raspberry Pi.
I wanted a coding agent I could actually understand and trust. The big harnesses are powerful, but they bring a lot I don't need — managed auth, plugin marketplaces, telemetry, personality layers. Raven keeps the parts that matter for real work and drops the rest.
What it's like to use:
- Local first — runs against Ollama on your machine by default; dial in OpenRouter when a task needs a bigger model. Only LLM requests leave your network.
- No telemetry — ever. No usage tracking, no phone-home, no cloud sync. All session state stays on disk, locally.
- Auditable — about 21K lines of Rust. You can read the whole harness. No proprietary plugins, no external service dependencies.
- Small footprint — a single binary that runs comfortably on a Raspberry Pi. No daemon, no background indexing.
- No personality layer — no "soul file". It keeps a plain
.raven/MEMORY.mdwith exactly what you tell it to remember. - Production-grade safety — workspace confinement (Landlock + seccomp on Linux), shell command filters, git-worktree isolation, and a verify-before-commit gate.
It's built for focused, supervised work — not an autonomous loop trying to complete entire projects unsupervised. I review everything it produces.
Linux / Raspberry Pi:
curl -fsSL https://raw.githubusercontent.com/raythurman2386/raven/master/install.sh | shWindows (PowerShell):
irm https://raw.githubusercontent.com/raythurman2386/raven/master/install.ps1 | iexThe script detects your platform, downloads the latest prebuilt binary from GitHub Releases, verifies the SHA-256 checksum, and installs to ~/.cargo/bin. Build from source with cargo build --release or cargo install --path ..
Requirements: Rust 1.88+ (MSRV) and a model endpoint — local Ollama, Ollama Cloud, OpenRouter, or any OpenAI-compatible /v1 API.
# Interactive TUI (default when no task is given).
# On first run Raven walks you through provider + model selection.
raven
# One-shot tasks (quotes required for multi-word tasks)
raven -p "Explain the structure of this repository"
raven -p "Add a README and .gitignore"
# Plan mode (default): propose → approve → execute
raven -p "Fix the failing tests"
# Agent mode: full tools immediately, no plan step
raven --mode agent -p "Refactor utils"
# Fully autonomous (no confirmations)
raven --yolo -p "Write unit tests for auth"
# Pick a provider / model for this session
raven --provider openrouter -m x-ai/grok-4.5 -p "Task"
# Continue a previous session
raven --resume # resume latest
raven --list-sessions # browse all sessionsRaven speaks Agent Client Protocol (ACP) over stdio (raven --acp), so it runs
as an external agent inside Zed. It advertises a model
session config option listing every configured provider's models, so you can
switch providers and models from the editor's model selector. See
docs/zed_connection.md for the complete setup.
| In scope | Intentionally out of scope |
|---|---|
Streaming agent loop (OpenAI-compatible /v1/chat/completions) |
MCP marketplace — Raven stays lean |
25 tools: list_dir, read_file, search_replace, write_file, grep, run_shell, search_code, todo_write, goal_set, delegate_task, think, memory_update, memory_search, git_status, git_diff, git_log, git_commit, apply_patch, run_tests, run_lint, ask_user, web_search, web_fetch, skill_search, skill_load |
Remote config sync |
| Small footprint — runs on a Raspberry Pi | Personality / "soul file" |
Document extraction (read_file → Markdown via anydoc: .docx, .pdf, .xlsx, …) |
Multi-model routing |
| Workspace sandbox (path confinement + dangerous-command filter) | GUI / web frontend |
| OS-level subprocess confinement (Landlock, seccomp, rlimits; Windows Job Object) | Container/VM isolation |
| Git worktree isolation (isolated branches per task) | Cloud sync of sessions |
| Structured plan mode (parse → approve → revise → execute) | Native IDE integration beyond ACP |
Skills (SKILL.md discovery + skill_search/skill_load) |
Plugins / marketplace auth |
Repo symbol map (<repo_map> for large workspaces) |
Managed workflow orchestration |
| Parallel tool execution within a single model turn | Telemetry / usage tracking |
| Context-window inference + automatic compaction | |
JSONL session persistence + --resume / --list-sessions |
|
Cross-session project memory (.raven/MEMORY.md) |
|
ACP v1 stdio (raven --acp) for editor attachment |
|
| ratatui TUI + headless CLI |
| Guide | Audience | Contents |
|---|---|---|
| docs/usage.md | Users | Day-to-day workflows, plan mode, parallel sub-agents |
| docs/configuration.md | Users | Config, env vars, providers, API keys, AGENTS.md |
| docs/example.config.toml | Users | Fully-commented reference config |
| docs/zed_connection.md | Users | Connect Raven to Zed via ACP |
| docs/troubleshooting.md | Users | Common failure modes (Windows, streams, sandbox, ACP) |
| docs/tools.md | Users + Contributors | Tool contracts, parameters, sandbox rules |
| docs/security.md | Security reviewers | Threat model, defense layers, platform caveats |
| docs/architecture.md | Contributors | Design, agent loop, compaction, sandbox |
| docs/contributing.md | Contributors | Build, style, how to add a tool or event |
| docs/testing.md | Contributors | Test structure, coverage, mutation testing |
See also the full docs index and CHANGELOG.md.
Raven keeps a plain, editable file at .raven/MEMORY.md. The first 25KB is injected into the system prompt on each run. The memory_update tool lets the agent persist conventions, decisions, and context across sessions — but it's just a Markdown file you can read and edit yourself. There's no hidden state, no "soul" or persona. It remembers exactly what you (or the agent) put in that file, and nothing else.
The agent tracks token usage with a built-in token estimator (no external vocab file needed). When the conversation approaches the context window limit, it automatically:
- Prunes old tool results — soft-trims tool outputs older than 3 turns (keeps head + tail with a truncation marker)
- Compacts the conversation — summarizes the middle of the conversation, preserving the system message, a short facts block (goal, open todos, key paths, last verification), and the last ~40% of the context budget for recent messages. The TUI shows a one-line "what was compacted" note.
Context window sizes are fetched from the model's actual metadata via Ollama's /api/show endpoint. This returns the real context_length from the model file (e.g. gemma4 → 128K, qwen3.5 → 256K, deepseek-v4-pro:cloud → 1M). If the API is unreachable (Ollama not running, model not found), a name-based heuristic is used as fallback:
glm:cloud,deepseek-v4:cloud(flash and pro) → 1Mqwen3.5→ 256Kgemma4,gemma3,qwen2.5,qwen3,llama3.1,llama3.2,deepseek,codestral,glm→ 128Kllama3,codellama,"32k"in name → 32Kmistral,"8k"in name → 8K- Unknown models → 32K (safe default)
Sessions are stored as JSONL under .raven/sessions/:
.raven/sessions/
2026-08-17T12-34-56-12345-0001/ # collision-proof ID (timestamp + PID + counter)
summary.json # metadata: id, model, timestamps, title
messages.jsonl # one ChatMessage per line (append-only)
debug-events.jsonl # local-only event log (model changes, saves, etc.)
last.patch # git diff snapshot (for audit/rollback)
Local-only guarantees:
- All writes are atomic (temp file + rename) for crash safety.
- Debug events (model changes, saves, etc.) are logged locally for reproducible debugging — never networked.
- Patch snapshots (
last.patch) are created after each session, recording the full git diff for audit or rollback decisions. - No telemetry, no remote reporting, no cloud sync.
Usage:
raven --resume # continue the most recent session
raven --resume <id> # continue a specific session by ID
raven --list-sessions # browse all sessions and their metadata
raven --export # write a local Markdown/JSON bundle of the latest session
raven --export <id> # export a specific session (see also TUI `/export`)Plan mode is the recommended workflow for important changes:
- Propose — agent creates a step-by-step plan
- Review — you read and approve (or revise)
- Execute — agent runs the plan with full tools
raven -p "Add type safety to this handler"
# (agent proposes a plan)
# ── Approve? [Y]es / [n]o / [r]evise ──
# y
# (agent executes)Skip the plan step for quick, exploratory tasks:
raven --mode agent -p "Refactor this function"/model— switch models or check the live model list for the active provider/provider— switch providers (ollama, openrouter, etc.) — slash-command autocomplete shows all available/clear— start a fresh turn (keeps session history)/retry— re-run the last user prompt after a failed turn/loop [N]— show or set the max iteration budget for new turns/steer <message>— redirect the running agent (restarts the turn with your direction appended)/cleanup <days> [--yes]— prune sessions older than N days (dry-run unless--yes; never deletes the current session)^C— stop the current task (session auto-saves)- Up/Down — recall previous prompts (keep pressing Up to walk back through history; Down returns toward the live input; typing resets). Home jumps to the top of the transcript, End returns to the live tail
- Mouse drag — select text in the transcript to copy it to your clipboard
- The footer below the input box shows context-sensitive keyhints (approve / answer / interrupt / idle)
Raven injects a repo symbol map for files >50KB. The map helps the agent navigate structure without reading entire files:
raven --context-window 131072 -p "Find all database queries and optimize them"If the agent seems stuck compacting, raise the threshold:
raven --compact-threshold 0.85 -p "Task"Raven keeps a plain .raven/MEMORY.md file across sessions. The agent can read and update it:
- Use
memory_searchto find past decisions - Use
memory_updateto record conventions or context
Example: "Remember we prefer async/await over promises in this codebase"
It's just a Markdown file — open it, edit it, or delete it. It only holds what you put there.
Create an AGENTS.md or CLAUDE.md file in your repo root:
# Coding Guidelines
- Always write tests for new features
- Use TypeScript, not JavaScript
- Follow the style guide in docs/STYLE.mdRaven auto-loads this and injects it into every session. You can also override per-session:
raven --rules "Use Python 3.11+; no type hints optional." -p "Task"Use --parallel to spawn multiple focused agents and gather results in parallel:
raven --parallel \
"Summarize the architecture" \
"List all TODOs" \
"Check for secrets in git history"The run_shell tool uses two complementary filters, neither of which is a security boundary:
-
Denylist — a regex that blocks obviously destructive patterns (recursive root deletes, fork bombs,
curl | sh, etc.). This is a best-effort guard, not a security boundary. A denylist is inherently incomplete — it can always be bypassed (e.g.rm -rf ~is not blocked even thoughrm -rf /is). -
Allowlist — a regex that matches known-safe development commands (
cargo,git,npm,ls,grep, etc.). Whenconfirm_shellis enabled (the default, non---yolopath), commands matching the allowlist run without a confirmation prompt. Anything outside the allowlist requires explicit user approval. Commands whose first token is allowlisted and contain no shell metacharacters run via direct exec (nosh -c), removing the shell-injection surface for the common case.
The --yolo flag disables confirmation entirely, but the denylist still applies as a last-resort filter. In addition to these filters, confined subprocesses run under OS-level sandboxing: Landlock (filesystem confinement) and seccomp (network-block) on Linux, plus resource limits (CPU / file size / fds) on Linux + macOS, and Job Object confinement on Windows. See docs/security.md for the full threat model.
cargo test # offline unit + integration tests
cargo test eval_suite # Layer A (fake model) eval harness
cargo clippy # zero warnings
cargo clippy -- -W clippy::pedantic # stricter lintingThe eval suite runs real agent tasks against a live model endpoint and grades the results. See evals/README.md for full details. I run this to decide how well new models could run in this harness for your average usage, not as a hard evaluation of strength.
cargo build --release
# Against Ollama Cloud (needs RAVEN_API_KEY)
python3 evals/run.py --model qwen3.8 --host https://api.ollama.ai/api/v1
# Against local Ollama
python3 evals/run.py --model qwen3.8:latest --host http://127.0.0.1:11434/v1
# Against OpenRouter (needs RAVEN_API_KEY)
python3 evals/run.py --model grok-4.5 --host https://openrouter.ai/api/v1
# View results
cat evals/out/<run-id>.mdTop-performing models (current):
- Ollama Cloud (daily-use recommended):
kimi-k3:cloud(latest, excellent),deepseek-v4-pro:cloud(high quality),deepseek-v4-flash:cloud(efficient),glm-5.2:cloud(long-horizon) - OpenRouter (frontier):
x-ai/grok-4.5(best reasoning, multimodal),x-ai/grok-4.6(frontier),Stealth/ox-alpha(pretty noice) - Local (when cloud unavailable):
qwen3.8:latest
See docs/testing.md for coverage and mutation testing details.
Troubleshooting: common failure modes (Windows .exe, stream errors,
sandbox denies, SearXNG fallback, ACP) are covered in
docs/troubleshooting.md.
No telemetry, ever. Nothing is collected, even anonymously. All agent state stays on your machine (except LLM requests to your chosen endpoint). The only network access is:
- LLM requests to your endpoint (Ollama local, OpenRouter cloud, etc.) — you control this.
- Optional
web_searchrequests to DuckDuckGo or a self-hosted SearXNG instance.
Everything else stays local. Sessions are stored as plain JSONL under .raven/sessions/ with local debug-event logs and git-diff snapshots for audit/rollback.
cargo test # offline unit + integration tests
cargo clippy # zero warningsThe live eval suite runs real agent tasks against a live model endpoint and grades results — see evals/README.md.
src/
main.rs # CLI, TUI, headless runner, session management
lib.rs # Library re-exports for benchmarks/integration tests
agent/ # Core agent loop (core, tools_exec, stream, parallel)
commands/ # Slash-command registry + parser (/retry, /loop, /steer, /cleanup, ...)
tools/ # Tool implementations (25 total) + sandbox/
tui/ # ratatui TUI (render, markdown, completion)
config/ # Layered config.toml loading, provider presets
context.rs # Context-window management and compaction
tokenizer.rs # Pure-Rust token counter (no vocab file)
session.rs # JSONL persistence, resume, list
plan.rs # Structured plan mode
memory.rs # Cross-session `.raven/MEMORY.md`
skills.rs # SKILL.md discovery
repomap/ # Repo symbol map for large codebases
web.rs # web_search / web_fetch
evals/ # Agent evaluation suite
docs/ # Documentation
MIT