Skip to content

Latest commit

 

History

History
399 lines (339 loc) · 22.7 KB

File metadata and controls

399 lines (339 loc) · 22.7 KB

SocratiCode — code exploration reference

Detail doc for the ## Code Exploration Policy block in AGENTS.md. That block carries the rule; this file carries the table.

Generated by init-socraticode — do not hand-edit above the END marker at the foot of this file. A re-run replaces everything between the markers. Anything written after the END marker survives, so repo-specific exploration notes — this repo's measured graph yield, its real artifact list, why an .socraticodeignore entry is there — go under a ## Repo-specific notes heading down there, not in AGENTS.md.

When to use each tool

Goal Tool
Where is X defined / how does Y work / what files touch Z codebase_search
Exact string/regex match (errors, log lines, known symbols) grep / rg
Blast radius of changing/deleting a file or function codebase_impact
What does an entry point actually do? codebase_flow
Callers and callees of a function codebase_symbol
Every symbol declared in a file codebase_symbols
Imports/dependents of a file codebase_graph_query
Import cycles codebase_graph_circular
Documented contracts (logging, DPV map, Pub 28, USPS API specs, style, dependency policy), design history — the declared artifacts codebase_context / codebase_context_search
Current DB schema, allowed values, migrations codebase_search or the migrations (alembic/versions/)
Path-pattern walks ("all *.py under src/address_validator/routers/") the Explore subagent

Prefetch

The codebase_* MCP tools are deferred: their schemas are not in the session until a ToolSearch prefetch loads them, and calling one before that fails validation. The SessionStart hook prints the select: query each session. If it did not fire, run the hook by hand and use the line it prints:

bash .claude/hooks/socraticode-reminder.sh

The query is deliberately not copied here: the hook is a vendored symlink, so upstream can change which tools it selects, and a copy in this file goes stale silently — the hook's output cannot drift from itself.

Per-tool notes

  • codebase_search takes a natural-language query, not a regex. It ranks by embedding similarity, so an empty result means "nothing scored above the threshold", not "no such code" — retry with minScore: 0 before concluding absence. With includeLinked: true it also searches linked projects, but only those whose paths resolve: a missing path — no checkout, no stub — is dropped without a word. The daily health check names any that do not.
  • codebase_impact / codebase_graph_query read the AST dependency graph, which is built separately from the embeddings. If the graph is stale or low-yield they answer empty rather than erroring — see Graph health.
  • codebase_flow traces from an entry point; give it a real file path, not a symbol name.
  • codebase_context_search only sees files listed in .socraticodecontextartifacts.json — and of those, only the ones actually indexed. A path that does not resolve is skipped silently, so a missing answer is often a manifest problem. But the same silence has a second cause a correct manifest cannot rule out: the path resolved, the run completed, and the artifact still is not indexed. Ask codebase_context, which is the only per-artifact index status there is — codebase_status gives a count and never a name — then run codebase_update. The once-per-day health check reports this gap too, and names the artifact. A third diagnosis has no empty result to warn you at all: the artifact is indexed, the answer arrives, and it is stale — behind its source. An edited file, or a new file under a directory artifact like docs/plans/, leaves the count at N/N with the superseded chunks still embedded. Measured: three of one repo's fourteen artifacts were behind their sources at a moment this check reported 14/14. What that costs is usually a wait, not a wrong answer: codebase_context_search re-indexes changed artifacts before it searches, so the first search after an edit pays that re-embed inline and then answers from current chunks. Old chunks reach an answer only when that staleness check itself errors — it is logged and the search proceeds anyway. codebase_context does not re-index, which is why a listing can sit at N/N while artifacts are behind. It prints each artifact's index time beside its status; compare it against the source, and for a directory against its newest file, not the directory's own timestamp. The daily check does exactly that and names the stale artifacts. codebase_update repairs them out of band, so the next search is not the one that pays. And every artifact competes in one ranking: a large directory of dated prose outranks a small current file, and a plan answers with the value it was written against. Set artifactName to search one artifact.
  • codebase_update is the incremental catch-up, and the repair for a stale artifact. It re-indexes changed files and re-embeds only the artifacts whose content hash moved, synchronously — seconds, on a repo the watcher has been following.
  • codebase_context_index is not. It re-embeds every artifact unconditionally — no content-hash skip, no progress notifications — so on a large manifest against a shared CPU embedder it can outlast Claude Code's 1800 s tool idle timeout. Measured: one repo's ~1,500 chunks took 77 minutes and the session gave up at 30. That timeout is not evidence the index failed — the server runs on after the client aborts, so check the project's lastIndexedAt in the socraticode_metadata collection before re-running. Keep it for a first index, or a manifest whose artifacts all changed.
  • The file watcher is ephemeral. It lives only while an MCP server process is running. After a long gap, or after a reboot, re-run codebase_index rather than trusting the index to be current.

Graph health

codebase_graph_status reporting READY does not mean the graph resolved anything — READY is reachable with a handful of edges across hundreds of files. Check yield, not status:

node skills-vendor/gregoryfoster-skills/skills/init-socraticode/scripts/mcp-driver.mjs health-check \
  "$(dirname "$(git rev-parse --path-format=absolute --git-common-dir)")" \
  --probe src/address_validator/services/validation/pipeline.py

Name the checkout, not the cwd. SocratiCode indexes by absolute project path, so a literal . run from a git worktree asks about a project the server never saw and reports a healthy index as broken. The spelling above is the one socraticode-health.sh uses; --show-toplevel is the near miss, because in a worktree it yields the worktree. Current drivers resolve a relative argument this way themselves, so . also works — the explicit form is here because it works against an older vendored driver too, and because it says out loud which path is being measured.

verdict: "low" means dependency questions must go to grep, and the AGENTS.md block should be on its degraded variant.

unresolvedPct is a statistic beside the verdict, never evidence for it. The same check — and the daily socraticode-health.sh run — report the figure whenever it clears the threshold, on a healthy graph too, and word it from the verdict. Beside ok: graph unresolved N% (> 50%) — share of captured symbol edges (calls, imports, re-exports, type or value references) matching no project symbol; edges into builtins and external libraries count by construction, so it runs high on healthy code — verdict is ok, so this is a statistic, not a defect. Beside low or unknown, where the verdict already stands on the yield arithmetic and the server's advisory: graph unresolved N% (> 50%) — share of captured symbol edges (calls, imports, re-exports, type or value references) matching no project symbol — reported beside the verdict, not as evidence for it, since edges into builtins and external libraries count by construction. Either way it is filed as a note — it appears as note: graph unresolved N% … and does not set the exit code, so a repo whose only finding is this one stays silent through the daily hook. The denominator is the server's own: since v1.14.0 codebase_graph_status says the same thing and adds that the share "is not a resolver failure rate". A repo that leans on frameworks, the stdlib and SDKs runs high by construction, because those symbols are not in the repo — no re-index brings them in and none lowers the figure. Judge the graph on verdict and on edges/file, which is what the gate keys on. A high unresolvedPct beside verdict: "ok" is normal; the src-layout resolver defect it can be mistaken for (giancarloerra/SocratiCode#107) shows up instead as near-zero edges/file. Do not cite the figure as the cause of an under-reporting graph query; test the import graph instead (#308).

If you suspect the import graph, test the import graph. Take a file you know has first-party importers, run codebase_graph_query on it, and compare the result against an rg sweep over every spelling that import could be written as. If the two sets match, the import graph is exact, and whatever unresolvedPct counts, it is not your first-party imports. Prefer that differential to any figure written into this file, which is repo- and day-specific.

Since SocratiCode 1.13.0 the server states the yield itself, and that is the signal to read. When resolution collapses, codebase_graph_status prints an advisory beneath the edge count:

Import resolution: 35 of 2959 captured imports resolved to project files (1.2%)
  Most imports did not resolve, so codebase_graph_query, codebase_graph_stats
  and codebase_impact will under-report dependencies — an empty answer there
  means unresolved, not independent.

That ratio is resolved-over-captured, which is a better measure than the edges/file floor this skill computes locally: it does not move with repo size, and it does not read as broken on a repo that is merely orphan-heavy. It is also not unresolvedPct, which counts every captured symbol edge, external ones included (see above).

Believe a present advisory when Built by: is current. It reports what the builder that cut this graph resolved, so on a stale graph it judges an older resolver and a rebuild may clear it. Either way the Built by: line below is what tells you whether the reading — advisory or silence — is about the resolvers you are actually running.

Its silence, in particular, is only meaningful if the graph is new enough to produce it:

Built by: what a missing advisory means
v<current server> the server measured and found nothing wrong — trust it
v<older> — STALE the graph predates the running resolvers; rebuild before judging
unknown (persisted before…) same, from a graph cut before the stamp existed
(line absent) server older than 1.13.0 — or no built graph here at all; fall back to edges/file

The middle two are the trap, and it is not hypothetical: a graph sitting at 37 edges across 621 files looked like a resolver collapse for over a week and was merely stale — rebuilt on 1.13.1 the same repo yields 2156 edges across 627 files. Run codebase_graph_build before concluding anything from a graph whose builder is stale or unstamped. health-check reports this as its own defect, and reports which measure ruled (source: "server" or "local") in its JSON.

Read row 1 as the server did not call it stale, not as it is current. The annotation is the server's to volunteer, and every way of not volunteering it used to land here: CannObserv/cannabis.observer-wordpress#803 spent three rounds concluding a PSR-4 composer.json declaration "would not help", from a graph cut by v1.10.0 — PSR-4 resolution having shipped in v1.11.0, so the graph predated the feature under discussion. READY throughout, and nothing said so. Since #297 health-check keeps the running server's serverInfo.version from the MCP handshake and makes the comparison itself, so row 1 now means both parties checked. Reading the stamp by hand, you do not have that second opinion: compare it against the server you are running before you trust it.

The server that matters is the one answering your queries. A rebuild runs through the session's server — under Claude Code, the plugin's — which need not be the one health-check launched to measure. Where the plugin's definition fixes a version, the check judges the graph against that one, and a graph matching it while trailing the check's own server is a note: rebuilding would re-stamp the same version, so the fix is to update the plugin, restart Claude Code so its MCP server reloads, then rebuild. Where the definition floats (socraticode@latest) the session's version cannot be read, and the finding says so; if a rebuild leaves the stamp unchanged, restart and rebuild (#305). The JSON records every version compared: graph.builderCheck holds the builder, checkServer, sessionServer and which of the two ruled; server is the check's own launch and sessionServer says how the session's version was known.

Stale is not the same as unmeasured, and the two answer different questions. A graph a release or two behind still carries the import counts its builder recorded, so the running server reads them and its advisory — or its silence — is a real ruling about resolution; health-check keeps that ruling (source: "server") and reports the staleness beside it. Only a graph cut before 1.13.0, which recorded no counts at all, leaves the server with nothing to measure and sends the verdict back to the edges/file fallback (source: "local"). Do not read "this graph is old" as "this graph is broken": on an orphan-heavy repo the local floor reads LOW on a graph that is perfectly fine, which is the reading that writes variant B.

Rebuild the graph after a SocratiCode upgrade. A stored graph reports READY forever, whatever cut it, and every resolver fix shipped since is absent from it. That rule does not need to live in your AGENTS.md — the once-per-day hook reports it, names both versions and names codebase_graph_build.

Index scope

Two stores, two controls. The repo-root .socraticodeignore (gitignore syntax, layered on the built-in defaults and .gitignore) governs the code index and the graph. The context store is governed by the manifest, .socraticodecontextartifacts.json. A path excluded from one stays searchable in the other: leaving docs/plans/ out of the code index does not take it out of codebase_context_search.

A directory artifact (socraticode 1.13+) honours the built-in defaults, the .gitignore files inside it, nested ones included, and a .socraticodeignore placed at the top of the artifact directory. The repo-root .socraticodeignore does not reach it. The defaults (build, dist, vendor, coverage, *.lock, __pycache__…) drop those names inside an artifact silently, so check each artifact's subtree for any you meant to keep.

To… Change
Trim what code search and the graph see the repo-root .socraticodeignore
Trim a directory artifact a .gitignore anywhere inside that artifact's directory, or a .socraticodeignore at its top
Drop an artifact its entry in the manifest

Editing .socraticodeignore affects subsequent scans only — re-index to apply it. Vendored trees dominate the index if left in, and vendored prose outranks first-party code in codebase_search results.

Repo-specific notes

Everything below the END marker survives an init-socraticode re-run. Measured figures carry the date they were taken; re-measure rather than trusting them.

Schemas: no schema artifact, so never context search

No schema, migration or source file is a context artifact here, so the generated schema row above has no artifactName half: schema, status vocabularies and migrations go to codebase_search or alembic/versions/. An unscoped codebase_context_search answers them from design-plans, and wrongly: "model_training_candidates status values allowed" returned 5 of 5 hits from plans asserting a stored 'assigned' status, which migration 014 dropped and SENSITIVE-AREAS.md records as a read-time rollup. The same question to codebase_search puts migration 014 in its top three (re-checked 2026-09-23). This is the measurement upstream's schema row cites (gregoryfoster/skills#315).

Context search is the right call for documented contracts — logging, DPV map, Pub 28, style, dependency policy — where a curated artifact exists and wins cleanly, and it is the only path to docs/plans rationale. Treat every design-plans hit as a dated snapshot.

The three collections, and what the generated Index scope leaves out

Measured 2026-09-22 against the live Qdrant (localhost:16333), SocratiCode 1.14.0. The generated section's "two stores" are three collections:

Store Collection Built by Governed by
Code index codebase_<id> codebase_index / codebase_update / watcher .socraticodeignore
Context store context_<id> codebase_context_index (auto on first search) .socraticodecontextartifacts.json
Graph <id>_symgraph_* codebase_graph_build (auto after a full index) .socraticodeignore

The generated section covers directory artifacts. It says nothing about single-file artifacts: those are read verbatim and skip every ignore file. Upstream calls that deliberate: "a declared path is an explicit instruction". Verified by calling readArtifactContent directly. A fixture whose root .socraticodeignore held docs/plans/, with a directory artifact at ./docs/plans, returned both files and exclusions.ignored = 0.

docs/plans/ and docs/research/ are excluded from the code index

Since GH #218, .socraticodeignore drops both — the directories AGENTS.md calls "dated snapshots, never current guidance". Both stay in the context store as the design-plans (directory) and address-validation-research (single-file) artifacts, so nothing became unsearchable.

Keep both lines on a re-run. Phase 4 asks about dated prose only where it is a directory artifact, and recommends against excluding a directory that is not one. docs/research/ is neither case: its one indexable file is a single-file artifact (the ISO 19160 PDF beside it was never indexed), so the exclusion loses nothing. Check git diff .socraticodeignore is empty afterwards.

Before the exclusion (2026-09-21, 1.14.0) plans were 981 of 2,677 code chunks (36.6%), 2.3x all of src/; after it the code index fell to 1,672. The context store stays about 88% design-plans: #218 closed on option E — keep the artifact and route schema questions away from it (above, and the first bullet of AGENTS.md's Code Exploration Notes). That routing is what makes E hold. Do not drop it from AGENTS.md without reopening #218.

Measured graph yield (2026-09-23, SocratiCode v1.14.0)

verdict: "ok" — 377 edges across 243 files, 1.55 per file, graph built by 1.14.0 (current), with unresolvedPct 72.2% (1,779 symbols, 8,425 call edges). That percentage counts edges into the stdlib and frameworks by construction; it is not a statement about imports. The import-graph differential above was run 2026-08-22 and again 2026-09-09 with the same outcome: codebase_graph_query on src/address_validator/services/validation/pipeline.py returned exactly two importers — src/address_validator/routers/v2/validate.py and tests/unit/validation/test_pipeline.py — matching an rg sweep over every spelling of that import. The import graph is exact here: an empty codebase_graph_query / codebase_impact answer means no importers.

The daily health hook is silent when clean

On a clean day the hook prints nothing. It prints only when the driver exits non-zero, i.e. on a defect or a crashed check. One note is written to .git/socraticode-health.log every run and never reaches the session: graph unresolved 72.2% (> 50%) — the statistic above, beside verdict: ok.

Both launches are pinned, to one version — the driver by GH #214 (~/.socraticode/pin), the session by GH #223 (SOCRATICODE_SPEC from VS Code's claudeCode.environmentVariables, declared in .claude/settings.json); HOST-MEMORY.md. So the pin-drift note (pinned at <version>; the plugin's 'socraticode@latest' resolves to …) no longer appears: that check measures only a floating session. That silence is not evidence. The hook reads SOCRATICODE_SPEC from its own environment, never from the session's launch, and stays silent over a session running @latest (gregoryfoster/skills#332); only the process table shows the launch. Unset the variable and the note returns — a patch gap as a note, a minor or major one as a defect naming the re-pin command. Re-pin both together, as a decision, not on a schedule; preflight.sh --check warns when the declared values disagree.

Earlier revisions of this file expected both lines in the session every day; since they became notes, the session sees neither. Output that says FAILED TO RUN or NOT measured means the check did not run. It is not an all-clear.

Context artifacts

11 declared, all documentation — AGENTS.md, the two vendored USPS OpenAPI specs, Pub 28, the provider / logging / style / dependency references, the research doc, docs/plans/, and this file. No source, schema or migration is registered. Validate with:

node skills-vendor/gregoryfoster-skills/skills/init-socraticode/scripts/mcp-driver.mjs \
  validate-manifest "$(dirname "$(git rev-parse --path-format=absolute --git-common-dir)")"

From a worktree, pass "$PWD" instead to validate the worktree's own manifest.

Index-scope exclusions

This repo excludes skills-vendor/, .claude/skills/ and .skills/, but deliberately not skills/ — first-party skills (brainstorming/, train-model/, writing-plans/) live there beside the vendor symlinks, and the indexer does not follow symlinks, so excluding skills-vendor/ already drops the vendored content.

Duplicate-config trap

If a session shows BOTH mcp__plugin_socraticode_socraticode__* and a standalone mcp__socraticode__*, remove the standalone — the plugin already provides the server: claude mcp remove socraticode.