- Overview
- Installation
- Configuration
- Dashboard Guide
- MCP Tools Reference
- REST API Reference
- Knowledge Base Management
- Session Search
- Troubleshooting
- FAQ
agent-knowledge is an MCP (Model Context Protocol) server that provides cross-session memory and recall for AI coding agents. It combines a git-synced knowledge base with session search, knowledge graphs, and hybrid semantic+TF-IDF search:
- Knowledge base -- Markdown files with YAML frontmatter stored in a git repository. Categories:
projects,people,decisions,workflows,notes. - Git sync -- automatic pull before reads, push after writes. Knowledge persists across machines.
- Hybrid search -- combines semantic vector similarity (via embeddings) with TF-IDF scoring for accurate retrieval.
- Session search -- search across past AI coding sessions from Claude Code, Cursor, and OpenCode.
- Scoped recall -- targeted search within domains:
errors,plans,configs,tools,files,decisions. - Knowledge graph -- typed edges between entries (11 relationship types including code structure) with directed BFS traversal.
- Confidence scoring -- entries gain maturity (
candidate>established>proven) based on access frequency. - Auto-linking -- new entries are automatically linked to similar existing entries via cosine similarity.
- Duplicate detection -- warns when writing entries that are similar to existing ones.
- Session distillation -- past sessions are auto-distilled into knowledge entries on server startup.
- Confidence tagging -- entries are marked
extracted(user-written) orinferred(auto-distilled). Inferred entries are down-weighted in search ranking so explicit user knowledge is preferred. - Knowledge analysis -- find most-connected concepts (god nodes), cross-category bridges, and isolated entries (gaps).
- Knowledge brief -- compact ~200 token summary of the knowledge base state for session-start orientation.
- Pre-extracted session insights -- distillation extracts git commits, error patterns, URLs, and package changes from session transcripts via regex (no LLM cost).
- Real-time dashboard -- web UI showing knowledge entries, sessions, and search results.
agent-knowledge has two entry points:
| Entry Point | File | Purpose |
|---|---|---|
| MCP stdio server | dist/index.js |
Communicates with the AI agent via JSON-RPC over stdin/stdout. Auto-starts the dashboard. |
| HTTP server | dist/server.js |
Standalone dashboard + REST API + WebSocket. |
Internally:
knowledge/ Store (Markdown CRUD), search (TF-IDF), git sync, distillation, graph, scoring, consolidation, reflection
sessions/ Parser (multi-format), indexer, search, scopes, summary, adapters (Claude Code, Cursor, OpenCode)
search/ TF-IDF engine, fuzzy matching, excerpt generation
embeddings/ Provider registry (Claude/Voyage, OpenAI, Gemini, local fallback)
vectorstore/ SQLite-backed vector storage with cosine similarity, document chunking
No framework dependencies. Pure Node.js + TypeScript.
- Node.js 20.11.0 or later
- npm (comes with Node.js)
- Git (for knowledge base sync)
npm install -g agent-knowledgegit clone https://github.com/keshrath/agent-knowledge.git
cd agent-knowledge
npm install
npm run buildnpx agent-knowledge| Variable | Default | Description |
|---|---|---|
AGENT_KNOWLEDGE_PORT |
3423 |
Dashboard HTTP/WebSocket port |
AGENT_KNOWLEDGE_DATA_DIR |
(platform config) | Override primary host data root (auto-detected) |
AGENT_KNOWLEDGE_GIT_URL |
(none) | Git remote URL for knowledge base sync |
AGENT_KNOWLEDGE_MEMORY_DIR |
~/agent-knowledge |
Local knowledge base directory |
AGENT_KNOWLEDGE_AUTO_DISTILL |
true |
Enable/disable session auto-distillation |
AGENT_KNOWLEDGE_EXTRA_SESSION_ROOTS |
(none) | Comma-separated additional session directories |
Add to ~/.claude.json:
{
"mcpServers": {
"agent-knowledge": {
"command": "npx",
"args": ["agent-knowledge"]
}
}
}The dashboard auto-starts at http://localhost:3423.
{
"permissions": {
"allow": ["mcp__agent-knowledge__*"]
}
}Configuration can also be set via the knowledge_admin tool with action: "config". Persisted config is stored at a tool-agnostic location and environment variables override persisted settings.
knowledge_admin with action "config", git_url "https://github.com/user/memory.git"
knowledge_admin with action "config", memory_dir "/custom/path/knowledge"
knowledge_admin with action "config", auto_distill false
agent-knowledge ships six lifecycle hook scripts that integrate with Claude Code's event system. Running node scripts/setup.js installs all six into ~/.claude/settings.json; on other hosts that don't support lifecycle hooks (Cursor, Windsurf), the MCP tools still work — you just lose the automatic context injection.
| Script | Event | Purpose |
|---|---|---|
session-start.js |
SessionStart | Dashboard URL + auto-loads token-budgeted wakeup payload into context |
session-start-ingest.mjs |
SessionStart | Detects project + reports knowledge-ingest cache drift via SHA256 diff |
first-prompt-inject.mjs |
UserPromptSubmit | Query-targeted knowledge hits injected on the session's first real prompt |
precompact-flush.mjs |
PreCompact | Rich session summary on disk + save-unsaved-context nudge into context |
precompact-distill.mjs |
PreCompact | Lightweight text snapshot of recent user prompts |
sessionend-distill.mjs |
SessionEnd | Final summary (turn counts, tool uses, first 20 prompts) |
Every hook fails open — if a script errors, it logs to stderr and the session continues. Each hook has env-var toggles (AGENT_KNOWLEDGE_AUTOWAKE, AGENT_KNOWLEDGE_FIRSTPROMPT_INJECT, AGENT_KNOWLEDGE_PRECOMPACT_NUDGE, etc.) — see docs/HOOKS.md for the full reference, test coverage, and budget/threshold knobs.
If your host has a per-session memory system (Claude Code writes ~/.claude/projects/*/memory/ files; Cursor has its own analogue), route durable facts to agent-knowledge instead:
- User preferences / feedback rules →
knowledge(action: write, category: "workflows", ..., evergreen: true) - User profile facts →
knowledge(action: write, category: "people", ...) - Project context →
knowledge(action: write, category: "projects", ...)
Host auto-memory is machine-local and invisible to other sessions and other machines. agent-knowledge is git-synced, searchable via knowledge_search, shows up in wakeup, and survives machine swaps. For Claude Code specifically, add a rule to your global ~/.claude/CLAUDE.md so Claude honors the redirect automatically:
## Persistent memory: always agent-knowledge, never auto-memory
Every durable fact goes to agent-knowledge via `knowledge(action: write)`. Never write
to `~/.claude/projects/*/memory/` — auto-memory is machine-local and invisible to other
sessions.The dashboard is available at http://localhost:3423 (or the port configured via AGENT_KNOWLEDGE_PORT).
Shows all knowledge base entries organized by category. Each entry card displays:
- Title extracted from the Markdown content or filename.
- Category badge (
projects,people,decisions,workflows,notes). - Tags from YAML frontmatter.
- Maturity level (
candidate,established,proven). - Access count -- how many times the entry has been read.
- Last accessed timestamp.
Click an entry to view its full Markdown content, related entries (from the knowledge graph), and score data.
The Knowledge tab header includes analysis buttons:
- Duplicates — scans entries for near-duplicates (TF-IDF similarity)
- Reflect — finds entries with no graph connections, generates a structured prompt for the agent
- God Nodes — opens a panel showing the most-connected entries (your core concepts)
- Bridges — shows entries that connect different categories with a
whyexplanation - Gaps — lists isolated entries (0-1 edges) sorted by maturity
- Brief — displays the knowledge base summary suitable for session-start orientation
Lists discovered sessions from all configured sources (Claude Code, Cursor, OpenCode). Each session shows:
- Session ID (UUID).
- Project name.
- Message count and duration.
- Last modified timestamp.
Click a session to view its messages.
The search bar performs hybrid search across both knowledge entries and sessions. Results are ranked by relevance combining TF-IDF and semantic similarity scores.
Displays vector store statistics: total entries, knowledge entries, session entries, database size, embedding provider, and dimensions.
Dark and light themes available. Preference saved in localStorage.
The dashboard connects via WebSocket. On connect, it receives the full state snapshot. A file watcher monitors the UI directory for hot-reload during development. The state snapshot is cached for 30 seconds to avoid expensive disk and database scans.
agent-knowledge exposes 6 MCP tools, each with multiple actions.
Knowledge base CRUD, git sync, and session-start hydration.
Actions: list, read, write, delete, sync, wakeup
Parameters:
| Name | Type | Required | Description |
|---|---|---|---|
action |
string | Yes | One of: list, read, write, delete, sync, wakeup |
category |
string | No | [list/wakeup] Filter by category; [write] Target directory. One of: projects, people, decisions, workflows, notes |
tag |
string | No | [list] Filter by tag |
path |
string | No | [read/delete] Relative path, e.g. projects/my-project.md |
filename |
string | No | [write] Filename with or without .md extension |
content |
string | No | [write] Full Markdown content (max 1MB) |
token_budget |
number | No | [wakeup] Max tokens for the rendered blob (chars/4 estimate, default 800, range 50–8000) |
sections |
string | No | [wakeup, v1.8.1] Comma-separated section list in priority order. See below for valid names. |
section_budgets |
object | No | [wakeup, v1.8.1] Per-section token budgets, e.g. {"identity": 200, "top_weighted": 400}. Unspecified sections share evenly. |
Action: list
Browse knowledge entries, optionally filtered by category or tag.
knowledge with action "list"
knowledge with action "list", category "projects"
knowledge with action "list", tag "architecture"
Action: read
Read an entry's content. Records an access for scoring. Returns content, score data, and related entries from the knowledge graph.
knowledge with action "read", path "projects/my-project.md"
Action: write
Create or update an entry. Auto-syncs to git (pull before write, push after). Auto-links to similar entries via vector search. Warns about potential duplicates.
knowledge with action "write", category "decisions", filename "use-jwt.md", content "---\ntags: [auth, security]\n---\n# Decision: Use JWT\n\nWe chose JWT over session cookies because..."
Response includes:
path: Where the file was written.git: Push result.autoLinked: Similar entries that were auto-linked (cosine > 0.7).similarEntries: Potential duplicates found.
Action: delete
Remove an entry and sync to git.
knowledge with action "delete", path "notes/old-note.md"
Action: sync
Manual git pull + push.
knowledge with action "sync"
Action: wakeup
Return a token-budgeted, section-priority context bundle. Call once at session start so the agent has its world loaded before issuing any real search.
Section-priority packing (v1.8.1): the bundle is assembled from an ordered list of sections. Each section gets a per-section token budget; the packer fills them top-to-bottom, redistributes any unused budget to later sections, and stops once the global token_budget is hit.
Default section order (drag-to-reorder in a later UI pass):
identity— L0 identity from~/agent-knowledge/identity.md(always first).active_tasks— placeholder pointing at agent-tasks (no in-repo task data today).recent_decisions— newest entries indecisions/, up to K.known_gotchas— entries taggedgotchaacross any category.last_session_summary— condensed meta of the most recent indexed session.top_weighted— the legacy L1 logic:recency × log(size+1).semantic_fallback— pure top-weighted catch-all if earlier sections under-filled the budget.
Empty sections emit a short placeholder rather than disappearing silently, so downstream UIs keep a stable shape.
knowledge with action "wakeup"
knowledge with action "wakeup", token_budget 1200
knowledge with action "wakeup", category "projects"
knowledge with action "wakeup", sections "identity,recent_decisions,top_weighted"
knowledge with action "wakeup", token_budget 2000, section_budgets { "identity": 200, "recent_decisions": 400, "top_weighted": 800 }
Backwards compat: calling with only token_budget (+ optional category) and no sections / section_budgets routes to the v1.8.0 identity + top_weighted behaviour — the rendered output stays shape-compatible for existing callers.
Response:
identity— text fromidentity.md, or a default placeholder.entries— array of{ path, title, weight, excerpt }from thetop_weightedsection (back-compat field).sections— array of{ name, content, budget, used, truncated, empty }, one per emitted section.rendered— the assembled Markdown blob ready to paste into a system prompt.token_estimate—chars / 4.truncated—trueif any section was cut to fit the budget.
Search across sessions and knowledge entries. Supports general search and scoped recall.
Response shape (v1.8): { mode, sessions, knowledge }.
mode: "general"— hybrid TF-IDF + semantic across both sources.mode: "scoped"— sessions-only filtered recall.knowledgeis always[]in this mode (scoped search is session-only by design — the response now says so instead of the v1.7 silent polymorphism).
Parameters:
| Name | Type | Required | Description |
|---|---|---|---|
query |
string | Yes | Search query -- supports keywords and phrases |
scope |
string | No | Search scope: errors, plans, configs, tools, files, decisions, all. When provided, performs scoped recall in sessions only. |
project |
string | No | Restrict to sessions from this project |
role |
string | No | Filter by message role: user, assistant, all (default: all, ignored when scope is set) |
max_results |
number | No | Maximum results (default: 20) |
ranked |
boolean | No | Use TF-IDF ranking (default: true, ignored when scope is set) |
semantic |
boolean | No | Blend semantic similarity with TF-IDF (default: true) |
category |
string | No | Category hint for the knowledge leg: projects, people, decisions, workflows, notes |
category_mode |
string | No | How category is applied: boost (v1.8 default, non-matches kept but matching entries get a 1.25× boost) or filter (legacy hard-filter) |
mmr |
boolean | No | Apply Maximal Marginal Relevance re-ranking to knowledge hits. Default false. Kills near-duplicate clusters in the top-K. |
mmr_lambda |
number | No | MMR tradeoff, 0–1 (default 0.7). 1.0 = pure relevance; 0.0 = pure diversity. |
explain |
boolean | No | When true, each knowledge hit carries a score_components breakdown: {bm25, decay, maturity, confidence, category_boost, mmr_penalty}. |
General search (no scope):
knowledge_search with query "authentication error handling"
knowledge_search with query "database migration", project "backend"
knowledge_search with query "why postgres", category "decisions", mmr true, explain true
Returns { mode: "general", sessions, knowledge } — both session matches and knowledge base matches.
Scoped recall (with scope):
knowledge_search with query "ECONNREFUSED", scope "errors"
knowledge_search with query "microservices", scope "plans"
knowledge_search with query "DATABASE_URL", scope "configs"
knowledge_search with query "comm_register", scope "tools"
knowledge_search with query "src/auth.ts", scope "files"
knowledge_search with query "JWT vs sessions", scope "decisions"
Scope descriptions:
| Scope | What it searches |
|---|---|
errors |
Stack traces, error messages, debugging sessions |
plans |
Architecture discussions, TODOs, planning messages |
configs |
Settings, environment variables, configuration |
tools |
MCP tool calls, tool results |
files |
File paths, code references |
decisions |
Trade-offs, choices, "chose X over Y" |
all |
No filter (same as general search within sessions) |
Session operations.
Parameters:
| Name | Type | Required | Description |
|---|---|---|---|
action |
string | Yes | One of: list, get, summary |
session_id |
string | Varies | Session UUID (required for get, summary) |
project |
string | No | Filter by project name (substring match) |
include_tools |
boolean | No | [get] Include tool_use and tool_result messages (default: false) |
tail |
number | No | [get] Only return the last N messages |
limit |
number | No | [list] Max sessions (default: 20, max: 500) |
offset |
number | No | [list] Skip first N sessions (default: 0) |
Action: list
Browse sessions, optionally filtered by project.
knowledge_session with action "list"
knowledge_session with action "list", project "backend", limit 50
Action: get
Retrieve full conversation messages for a session.
knowledge_session with action "get", session_id "abc-123-def"
knowledge_session with action "get", session_id "abc-123-def", include_tools true, tail 20
Action: summary
Get a quick overview of a session (topics, file paths, message counts).
knowledge_session with action "summary", session_id "abc-123-def"
Admin operations for configuration, index stats, embedding management, and the v1.8 scored promoter.
Parameters:
| Name | Type | Required | Description |
|---|---|---|---|
action |
string | Yes | One of: status, config, rebuild_embeddings, prune_orphans, vacuum, promote |
git_url |
string | No | [config] Git remote URL (empty string to remove) |
memory_dir |
string | No | [config] Local knowledge base directory (empty to reset) |
auto_distill |
boolean | No | [config] Enable/disable scheduled promotion (governs the promoter, not the legacy regex distiller) |
vacuum |
boolean | No | [prune_orphans] Run VACUUM after pruning (default true) |
force_vacuum |
boolean | No | [prune_orphans] Run VACUUM even when no orphans were pruned |
promote_mode |
string | No | [promote] explain (default — read-only, returns score breakdowns) or apply (write + git-commit) |
min_score |
number | No | [promote] Override minScore gate (default 0.5) |
min_recall_count |
number | No | [promote] Override minRecallCount gate (default 2) |
min_unique_queries |
number | No | [promote] Override minUniqueQueries gate (default 2) |
Action: status
View vector store statistics.
knowledge_admin with action "status"
Returns total entries, knowledge entries, session entries, unique sessions, database size, provider name, and dimensions.
Action: config
View or update configuration. Without update params, returns current config.
# View config
knowledge_admin with action "config"
# Update git URL
knowledge_admin with action "config", git_url "https://github.com/user/memory.git"
# Disable auto-promotion (the v1.8 scored promoter runs on backgroundIndex)
knowledge_admin with action "config", auto_distill false
Action: rebuild_embeddings
Re-embed all knowledge entries. Useful when switching embedding providers.
knowledge_admin with action "rebuild_embeddings"
Wipes existing vectors, re-creates the store with the current provider's dimensions, and re-embeds all entries. Returns processed/failed counts.
Action: prune_orphans
Delete vector embeddings for sessions no longer present on disk. Optionally follows with VACUUM to reclaim pages.
knowledge_admin with action "prune_orphans"
knowledge_admin with action "prune_orphans", force_vacuum true
Action: vacuum
Reclaim free SQLite pages in the vector store.
knowledge_admin with action "vacuum"
Action: promote (v1.8)
Run the scored + gated promoter. Every project-level candidate is scored on six weighted signals (relevance 0.30, frequency 0.24, queryDiversity 0.15, recency 0.15, consolidation 0.10, conceptualRichness 0.06) and tested against three independent gates (minScore, minRecallCount, minUniqueQueries — all must pass). Candidates that pass are promoted to projects/<id>.md; every run drops an audit trail in ~/agent-knowledge/.dreams/YYYY-MM-DD.md.
Grounded rehydration: if a candidate's source session file is no longer on disk, the promoter skips it. Prevents writing content the user has deleted.
# Score candidates + write the diary, DO NOT touch the KB
knowledge_admin with action "promote"
# Actually promote + git-commit
knowledge_admin with action "promote", promote_mode "apply"
# Tighten gates for a conservative run
knowledge_admin with action "promote", promote_mode "apply", min_score 0.7, min_recall_count 3
Returns { runStartedAt, runFinishedAt, mode, candidates[], promotedPaths[], skippedIds[], diaryPath, totals }.
Knowledge graph operations with temporal validity and code-structure support. Seven actions cover creation, removal, time-scoped invalidation, directed BFS traversal, and bulk ingestion for code graphs.
Parameters:
| Name | Type | Required | Description |
|---|---|---|---|
action |
string | Yes | One of: link, unlink, invalidate, list, traverse, bulk_link, unlink_by_origin |
source |
string | Varies | [link/unlink/invalidate] Source entry path |
target |
string | Varies | [link/unlink/invalidate] Target entry path |
entry |
string | No | [list] Filter by entry path; [traverse] BFS start node |
rel_type |
string | Varies | Relationship type (required for link, optional filter for unlink/invalidate/list/traverse) |
strength |
number | No | [link] Edge strength 0-1 (default: 0.5) |
depth |
number | No | [traverse] Max traversal depth in hops (default: 2) |
valid_from |
string | No | [link] ISO date the fact became true. Null/omitted = unbounded. |
valid_to |
string | No | [link/invalidate] ISO date the fact stopped being true. For invalidate, defaults to today. |
as_of |
string | No | [list/traverse] Only return edges valid at this ISO date. |
direction |
string | No | [traverse] outbound (source→target), inbound (target→source), or both (default — undirected). |
edges |
array | No | [bulk_link] Batch of {source, target, rel_type, strength?, origin?} edges. Used by knowledge-ingest for code graphs. |
origin |
string | No | [unlink_by_origin] Origin tag to delete by (e.g. tree-sitter to clear the code graph before re-ingest). Also accepted on link. |
Relationship types:
| Type | Description |
|---|---|
related_to |
General association |
supersedes |
Entry replaces another |
depends_on |
Entry depends on another |
contradicts |
Entries have conflicting information |
specializes |
Entry is a specific case of another |
part_of |
Entry is a component of another |
alternative_to |
Entry is an alternative to another |
builds_on |
Entry extends or builds on another |
calls |
Function/method X calls Y (code structure, v1.7+, code: node IDs) |
imports |
Module X imports Y (code structure) |
inherits |
Class X inherits from Y (code structure) |
Action: link
Create an edge between two entries.
knowledge_graph with action "link", source "projects/backend.md", target "decisions/use-jwt.md", rel_type "depends_on", strength 0.8
Action: unlink
Remove an edge.
knowledge_graph with action "unlink", source "projects/backend.md", target "decisions/use-jwt.md"
knowledge_graph with action "unlink", source "projects/backend.md", target "decisions/use-jwt.md", rel_type "depends_on"
Action: list
List edges, optionally filtered by entry or relationship type.
knowledge_graph with action "list"
knowledge_graph with action "list", entry "projects/backend.md"
knowledge_graph with action "list", rel_type "depends_on"
Action: invalidate
Set valid_to on one or more edges without deleting them. Preserves history for temporal queries.
knowledge_graph with action "invalidate", source "projects/backend.md", target "decisions/use-jwt.md"
knowledge_graph with action "invalidate", source "projects/backend.md", target "decisions/use-jwt.md", valid_to "2026-03-15"
Action: traverse
Directed BFS traversal from a starting entry. Returns a graph of connected entries.
knowledge_graph with action "traverse", entry "projects/backend.md"
knowledge_graph with action "traverse", entry "projects/backend.md", depth 3
knowledge_graph with action "traverse", entry "code:src/auth.ts::login", direction "inbound", rel_type "calls"
knowledge_graph with action "traverse", entry "projects/backend.md", as_of "2026-01-01"
Action: bulk_link
Batch-create edges in a single transaction. Used by knowledge-ingest for code graphs (hundreds to thousands of edges per ingest).
knowledge_graph with action "bulk_link", edges [
{ "source": "code:src/a.ts", "target": "code:src/b.ts", "rel_type": "imports", "origin": "tree-sitter" },
...
]
Action: unlink_by_origin
Delete every edge matching a specific origin tag. Used to clear stale code-graph edges before a re-ingest.
knowledge_graph with action "unlink_by_origin", origin "tree-sitter"
Analysis tools for knowledge base maintenance.
Parameters:
| Name | Type | Required | Description |
|---|---|---|---|
action |
string | Yes | One of: consolidate, reflect, god_nodes, bridges, gaps, brief, search_gaps |
category |
string | No | Scan only this category (omit for all) |
threshold |
number | No | [consolidate] Similarity threshold 0-1 (default: 0.5) |
max_entries |
number | No | [reflect/gaps] Max entries to include (default: 20 for reflect, 30 for gaps) |
top_n |
number | No | [god_nodes/bridges] Max entries to return (default: 10 for god_nodes, 5 bridges) |
since_days |
number | No | [search_gaps] Lookback window in days (default: 30) |
min_count |
number | No | [search_gaps] Minimum occurrence count per merged group (default: 1) |
group_similarity |
number | No | [search_gaps] Jaccard token similarity threshold for grouping (default: 0.7) |
Action: consolidate
Find near-duplicate entries using TF-IDF similarity.
knowledge_analyze with action "consolidate"
knowledge_analyze with action "consolidate", category "projects", threshold 0.7
Returns clusters of similar entries that may be candidates for merging.
Action: reflect
Find entries with no graph connections and generate a structured prompt for linking them.
knowledge_analyze with action "reflect"
knowledge_analyze with action "reflect", category "decisions", max_entries 10
Returns unconnected entries and suggested relationships to investigate.
Action: god_nodes
Returns the most-connected entries in the knowledge graph, ranked by degree centrality.
| Parameter | Type | Default | Description |
|---|---|---|---|
top_n |
number | 10 | Maximum entries to return |
Returns: array of { path, title, category, degree, confidence } sorted by degree descending. Entries with only auto-link edges and the auto-distilled tag are excluded.
Action: bridges
Returns entries that bridge different categories, ranked by betweenness centrality.
| Parameter | Type | Default | Description |
|---|---|---|---|
top_n |
number | 5 | Maximum bridges to return |
Returns: array of { path, title, betweenness, connects, why }. Only entries connecting at least 2 different categories are returned.
Action: gaps
Returns entries with 0-1 graph edges (weakly connected or isolated).
| Parameter | Type | Default | Description |
|---|---|---|---|
max_entries |
number | 30 | Maximum entries to return |
Returns: array of { path, title, category, degree, maturity, daysSinceAccess } sorted with proven entries first (most concerning gaps), then by degree ascending.
Action: brief
Returns a compact summary of the knowledge base state (cached 1 hour, invalidated on write/delete/link/unlink).
No parameters.
Returns: { total_entries, total_edges, core_concepts, active_projects, recent_decisions, stale_count, gap_count, generated_at, text }. The text field is a ~200 token plain-text summary suitable for injection into agent prompts.
Action: search_gaps
Returns zero-result knowledge_search queries grouped by similarity — the single strongest signal for "what knowledge entries should I write next?". Every knowledge_search call is logged to a small query_log SQLite table (query text scrubbed for secrets before insertion). Queries that returned zero results are merged into groups via token-set Jaccard similarity, so "gitlab token" and "gitlab credentials" collapse into one row.
| Parameter | Type | Default | Description |
|---|---|---|---|
since_days |
number | 30 | Lookback window in days |
min_count |
number | 1 | Minimum occurrence count per merged group (raise to filter noise) |
group_similarity |
number | 0.7 | Jaccard token similarity threshold for merging queries (0-1) |
Returns: array of { query, count, last_seen, similar_queries? } sorted by count descending, then by last_seen descending. query is the most-recent query text in the group; similar_queries lists the other variants collapsed into the same group (omitted when the group has only one member).
knowledge_analyze with action "search_gaps"
knowledge_analyze with action "search_gaps", since_days 7, min_count 3
The REST API is served by the dashboard. All responses include CORS headers. Rate limited: 100 requests/minute general, 20 requests/minute for search/analyze endpoints.
GET /health Status, version, uptime, knowledge entry count
GET /api/knowledge List entries (?category=&tag=)
GET /api/knowledge/:path Read entry content (enriched with score data)
GET /api/knowledge/:path/links Get graph edges for an entry
GET /api/knowledge/search Search entries (?q=&category=&max_results=)
GET /api/knowledge/consolidate Find duplicates (?category=&threshold=)
GET /api/knowledge/reflect Find unconnected entries (?category=&max_entries=)
GET /api/knowledge/god-nodes Most-connected entries (?top_n=)
GET /api/knowledge/bridges Cross-category connectors (?top_n=)
GET /api/knowledge/gaps Isolated entries (?max_entries=)
GET /api/knowledge/brief Cached knowledge base brief
Query params: top_n (number, default 10).
Returns: same as knowledge_analyze(action: "god_nodes").
Query params: top_n (number, default 5).
Returns: same as knowledge_analyze(action: "bridges").
Query params: max_entries (number, default 30).
Returns: same as knowledge_analyze(action: "gaps").
No query params. Returns the cached knowledge brief.
GET /api/sessions List sessions (?project=&limit=&offset=)
GET /api/sessions/:id Get session messages (?project=&include_tools=&tail=)
GET /api/sessions/:id/summary Get session summary (?project=)
GET /api/sessions/search Search sessions (?q=&role=&max_results=&ranked=&project=&semantic=)
GET /api/sessions/recall Scoped recall (?scope=&q=&max_results=&project=)
GET /api/index-status Vector store statistics
# Health check
curl http://localhost:3423/health
# List all knowledge entries
curl http://localhost:3423/api/knowledge
# Read a specific entry
curl http://localhost:3423/api/knowledge/projects/backend.md
# Search knowledge
curl "http://localhost:3423/api/knowledge/search?q=authentication&max_results=5"
# Search sessions
curl "http://localhost:3423/api/sessions/search?q=deploy+error&role=assistant&max_results=10"
# Scoped recall
curl "http://localhost:3423/api/sessions/recall?scope=errors&q=ECONNREFUSED"
# List sessions for a project
curl "http://localhost:3423/api/sessions?project=backend&limit=20"Knowledge entries are Markdown files with optional YAML frontmatter:
---
title: JWT Authentication
tags: [auth, security, jwt]
updated: 2026-04-18
confidence: extracted
evergreen: true
---
# JWT Authentication
We use JWT with refresh tokens for stateless authentication...Recognised frontmatter fields:
| Field | Type | Meaning |
|---|---|---|
title |
string | Display title; derived from filename when omitted. |
tags |
string[] | Inline array form, e.g. [auth, security]. Filterable via knowledge(action="list", tag=…). |
updated |
string (date) | Human-readable last-edited date. |
confidence |
string | extracted (user-written, 1.0× rank) or inferred (auto-distilled / auto-promoted, 0.85× rank). |
confidence_score |
number | Optional 0–1 float carrying a finer-grained model certainty for distilled content. |
evergreen |
boolean | v1.8: when true, the entry skips time-based decay in ranking AND the scored promoter appends to rather than overwrites the entry. Use for durable decisions / identity. |
Entries are organized into categories (directories):
| Category | Purpose |
|---|---|
projects |
Project context, architecture, team info |
people |
Team members, contacts, roles |
decisions |
Architecture decisions, trade-offs |
workflows |
Processes, CI/CD, deployment procedures |
notes |
General notes, observations |
The knowledge base is optionally backed by a git repository. When configured:
knowledge(action: "read")does agit pullbefore reading.knowledge(action: "write")does agit pullbefore writing, thengit pushafter.knowledge(action: "sync")does a manual pull + push.
Configure the git remote:
knowledge_admin with action "config", git_url "https://github.com/user/memory.git"
Entries have a confidence score that increases with access:
| Maturity | Description |
|---|---|
candidate |
New entry, not yet validated |
established |
Accessed multiple times, gaining trust |
proven |
Frequently accessed, high confidence |
Scores include a decay factor -- entries not accessed recently receive lower rankings in search results. The decay is based on time since last access.
When writing a new entry, agent-knowledge automatically:
- Generates embeddings for the content.
- Searches for similar existing entries via vector similarity.
- Creates
related_toedges for entries with cosine similarity > 0.7 (up to 3 links).
Auto-linked entries appear in the write response.
On write, TF-IDF similarity is checked against existing entries. If similar entries are found, a similarEntries warning is included in the response. This helps avoid knowledge fragmentation.
On server startup (when auto_distill is enabled), past sessions are automatically scanned and distilled into knowledge entries. The distillation process:
- Parses session files from all configured sources.
- Extracts key insights, decisions, and patterns.
- Scrubs secrets (API keys, tokens, passwords) from the content.
- Writes distilled entries to the
projects/category.
Sessions are auto-discovered from installed AI coding tools:
| Tool | Location |
|---|---|
| Claude Code | $AGENT_KNOWLEDGE_DATA_DIR/projects/ (JSONL files) |
| Cursor | ~/.cursor/projects/*/agent-transcripts/ (JSONL) |
| OpenCode | ~/.local/share/opencode/opencode.db (SQLite) |
Additional roots can be added via AGENT_KNOWLEDGE_EXTRA_SESSION_ROOTS (comma-separated paths).
General search finds matches across all sessions and knowledge entries:
knowledge_search with query "authentication"
Scoped recall targets specific domains within sessions:
knowledge_search with query "ECONNREFUSED", scope "errors"
Results are ranked using a hybrid approach:
- TF-IDF: Term frequency-inverse document frequency scoring.
- Semantic: Cosine similarity between query and document embeddings (when available).
- Combined: Weighted blend of both scores (configurable via
embeddingAlphain config).
Set semantic: false to use pure TF-IDF, or ranked: false for regex-based matching.
Semantic search requires an embedding provider. Supported providers (auto-detected):
| Provider | Environment Variable |
|---|---|
| Claude/Voyage | ANTHROPIC_API_KEY |
| OpenAI | OPENAI_API_KEY |
| Gemini | GOOGLE_API_KEY |
| Local | (fallback, no API key) |
If no provider is available, search falls back to pure TF-IDF.
Symptom: Port 3423 already in use.
Solutions:
- Set
AGENT_KNOWLEDGE_PORT=3424in the MCP config env. - Multiple MCP instances share the same database. Only one serves the dashboard.
Causes and solutions:
- No sessions found: Check that
AGENT_KNOWLEDGE_DATA_DIRpoints to the correct directory. Verify session files exist. - Embeddings not available: Semantic search requires an API key. Check the embedding provider configuration.
- Index not built: Background indexing runs 5 seconds after startup. Wait for it to complete. Check stderr for
[knowledge] Background indexmessages.
Symptom: Push or pull errors.
Solutions:
- Verify the git remote URL is correct and accessible.
- Ensure git credentials are configured (SSH keys or credential helper).
- Git operations have a 30-second timeout to prevent hangs.
- Run
knowledge(action: "sync")manually to diagnose.
Symptom: Large database file, slow searches, or embedding errors.
Solutions:
- The vector store uses SQLite and can grow large with many sessions. Check size via
knowledge_admin(action: "status"). - If the embedding provider changes, rebuild embeddings:
knowledge_admin(action: "rebuild_embeddings"). - If the database is corrupted, delete
knowledge-vectors.dband associated WAL files. Embeddings will be rebuilt on next startup.
Symptom: Errors about missing knowledge base directory.
Solutions:
- The default directory is
~/agent-knowledge/. Create it manually or let the git clone create it. - Override via
AGENT_KNOWLEDGE_MEMORY_DIRorknowledge_admin(action: "config", memory_dir: "/path").
Yes. agent-knowledge is a standard MCP server. It also auto-discovers sessions from Cursor and OpenCode installations, so you can search across all your AI coding sessions regardless of which tool created them.
In ~/agent-knowledge/ by default (configurable via AGENT_KNOWLEDGE_MEMORY_DIR). Entries are plain Markdown files with YAML frontmatter, organized in category subdirectories.
When a git URL is configured, the knowledge directory is a git repository. Reads trigger git pull, writes trigger git pull + git push. This keeps the knowledge base synchronized across machines.
Search falls back to pure TF-IDF (keyword-based). Semantic features (vector similarity, auto-linking) are disabled. The system still works well for exact and fuzzy keyword matching.
Implement the SessionAdapter interface in src/sessions/adapters/. The adapter registry auto-detects new adapters on startup. Each adapter defines how to find and parse session files for a specific tool.
Yes. Configure a git remote via knowledge_admin(action: "config", git_url: "..."). All writes are automatically pushed to the remote. On startup and before reads, the latest changes are pulled.
The vector store grows with the number of indexed entries (knowledge + sessions). Each entry generates multiple chunks with embedding vectors. The knowledge_admin(action: "status") tool shows the current database size and entry counts.
knowledge_searchis for finding information: it searches across both sessions and knowledge entries, ranking results by relevance.knowledge_sessionis for browsing sessions: listing them, reading their messages, or getting summaries.
The distillation process applies regex patterns to detect and remove common secret formats (API keys, tokens, passwords, connection strings) before writing knowledge entries. This prevents accidental persistence of sensitive data.
Yes. Set AGENT_KNOWLEDGE_AUTO_DISTILL=false as an environment variable, or use knowledge_admin(action: "config", auto_distill: false).