-
Notifications
You must be signed in to change notification settings - Fork 3
Expand file tree
/
Copy path.deepreview
More file actions
462 lines (399 loc) · 20.5 KB
/
Copy path.deepreview
File metadata and controls
462 lines (399 loc) · 20.5 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
prompt_best_practices:
description: "Review prompt/instruction markdown files for Anthropic prompt engineering best practices."
match:
# Add any known prompt files to the include list
include:
- "**/CLAUDE.md"
- "**/AGENTS.md"
- ".claude/**/*.md"
- ".deepwork/review/*.md"
- ".deepwork/jobs/**/*.md"
- "plugins/**/skills/**/*.md"
- "learning_agents/skills/**/*.md" # learning_agents plugin skills are prompt-heavy
- "learning_agents/agents/**/*.md" # agent persona definitions
- "plugins/**/agents/**/*.md" # plugin agent definitions are prompt-heavy
- "platform/**/*.md"
- "src/deepwork/standard_jobs/**/*.md"
- "library/jobs/**/*.md" # library job step instructions are prompt-heavy files
- "library/jobs/*/job.yml" # inline step instructions in job.yml are prompt-heavy
- "src/deepwork/standard_jobs/*/job.yml" # standard job inline instructions
- "**/.deepreview" # .deepreview files contain inline agent instructions — prompt-heavy content
review:
strategy: individual
instructions:
file: .deepwork/review/prompt_best_practices.md
python_code_review:
description: "Review Python files against project conventions and best practices."
match:
include:
- "**/*.py"
exclude:
- "tests/fixtures/**"
review:
strategy: individual
instructions: |
Review this Python file against the project's coding conventions
documented in doc/code_review_standards.md and the ruff/mypy config
in pyproject.toml.
Check for:
- Adherence to project naming conventions (snake_case for functions/variables,
PascalCase for classes)
- Proper error handling following project patterns
- Import ordering (stdlib, third-party, local — enforced by ruff isort)
- Type hints on function signatures (enforced by mypy disallow_untyped_defs)
- Security concerns (injection, unsafe deserialization, path traversal)
- Performance issues (unnecessary computation, resource leaks)
Additionally, always check:
- **DRY violations**: Is there duplicated logic or repeated patterns that
should be extracted into a shared function, utility, or module? Look for
copy-pasted code with minor variations and similar code blocks differing
only in variable names.
- **Comment accuracy**: Are all comments, docstrings, and inline
documentation still accurate after the changes? Flag any comments that
describe behavior that no longer matches the code.
- **PR-local comments**: Flag any comments or docstrings that reference
specific source code line numbers (e.g., "covers lines 107-108",
"tests the branch at line 65"), PR numbers, or other transient context
that will not make sense to a future reader. Comments should describe
*what* and *why*, not *where in the current diff* something is. This
applies to both production code and test files.
Output Format:
- PASS: No issues found.
- FAIL: Issues found. List each with file, line, severity (high/medium/low),
and a concise description.
python_lint:
description: "Auto-fix lint and format issues, then report anything unresolvable."
match:
include:
- "**/*.py"
exclude:
- "tests/fixtures/**"
review:
strategy: matches_together
precomputed_info_for_reviewer_bash_command: make lint
instructions: |
This is an auto-fix rule — you SHOULD edit files to resolve issues.
The precomputed lint output above shows the results of running `make lint`
(ruff format, ruff check --fix, mypy) on the entire project.
If there are remaining errors:
1. Fix any issues you can resolve (e.g., adding type annotations,
renaming ambiguous variables).
2. Re-run `make lint` to confirm your fixes are clean.
Output Format:
- PASS: All checks pass (no remaining issues).
- FAIL: Unfixable lint errors, type errors, DRY violations, or stale
comments remain. List each with file, line, and details.
suggest_new_reviews:
description: "Analyze all changes and suggest new review rules that would catch issues going forward."
match:
include:
- "**/*"
exclude:
- ".github/**"
review:
strategy: matches_together
instructions:
file: .deepwork/review/suggest_new_reviews.md
requirements_traceability:
description: "Verify requirements traceability between specs, code, tests, review rules, and DeepSchemas."
match:
include:
- "**/*"
exclude:
- ".github/**"
review:
strategy: all_changed_files
agent:
claude: requirements-reviewer
precomputed_info_for_reviewer_bash_command: .deepwork/requirements_traceability_info.sh
instructions: |
Review the changed files for requirements traceability.
Requirements live in `doc/specs/` as `{PREFIX}-REQ-NNN-<topic>.md`
(prefixes: DW-REQ, JOBS-REQ, REVIEW-REQ, LA-REQ, PLUG-REQ), with
individually numbered items (e.g. JOBS-REQ-004.1). Each requirement
must be validated by automated tests, DeepSchemas, or `.deepreview` rules.
## Choosing the right validation mechanism
Pick the mechanism that matches the requirement type. The wrong choice
creates false confidence or wastes reviewer judgment.
**Anonymous DeepSchemas** (`.deepschema.<filename>.yml`): when the
requirement targets a specific file (structural or semantic). Use
`json_schema_path` / `verification_bash_command` for exact checks,
or judgment-based criteria for prose. Prefer DeepSchemas over tests
and `.deepreview` rules for single-file requirements.
**Automated tests** (`tests/`): for concrete, machine-verifiable facts
spanning multiple files (function return values, CLI output, cross-file
structure). Tests reference requirement IDs via docstrings/comments.
**`.deepreview` rules**: when evaluation requires judgment AND applies
broadly across many files of a type (coding standards, documentation
accuracy, prompt conventions). Rules and DeepSchemas reference
requirement IDs in `description`, `instructions`, or `requirements`.
## Anti-patterns to flag
- **Fragile keyword tests for judgment requirements**: e.g.
`"parallel" in content` for "MUST launch tasks in parallel" — use
a review rule instead. See doc/specs/validating_requirements_with_rules.md.
- **Review rules for machine-verifiable requirements**: e.g. asking a
reviewer to check for a specific flag — use a test instead.
## Review checklist
1. New/changed end-user functionality has a requirement in `doc/specs/`.
2. Every touched requirement has a test, DeepSchema, or `.deepreview`
rule. Verify the mechanism matches the requirement type.
3. Flag test modifications where the underlying requirement didn't change.
4. For rule-validated requirements, verify the rule references the
requirement ID and its scope covers the requirement's intent.
Produce a structured review with Coverage Gaps, Test Stability
Violations, and a Summary with PASS/FAIL verdicts.
test_file_quality:
description: "Verify test traceability comments are present, correctly formatted, and reference valid requirement IDs."
match:
include:
- "tests/**/*.py"
exclude:
- "tests/fixtures/**"
review:
strategy: individual
instructions: |
Review this test file for traceability comment quality.
Tests that validate a formal requirement MUST have a two-line
traceability comment immediately before the `def` line:
```
# THIS TEST VALIDATES A HARD REQUIREMENT (JOBS-REQ-004.5).
# YOU MUST NOT MODIFY THIS TEST UNLESS THE REQUIREMENT CHANGES
def test_something(self):
```
Check the following:
1. **Both lines present**: Every test with a `THIS TEST VALIDATES`
comment must also have the `YOU MUST NOT MODIFY` line. Flag any
test where only one of the two lines is present.
2. **Placement**: The two-line comment must appear immediately before
the `def` line (not inside the method body, not separated by
blank lines or other comments).
3. **Valid REQ ID**: The requirement ID in parentheses must follow the
pattern `{PREFIX}-REQ-NNN.M` where PREFIX is one of: DW-REQ,
JOBS-REQ, REVIEW-REQ, LA-REQ, PLUG-REQ.
4. **Consistency**: If a test references a REQ ID only in a docstring
but is missing the formal two-line comment block, flag it.
Do NOT flag tests that lack any requirement reference — only tests
that are utility/edge-case tests without a requirement mapping do not
need traceability comments.
Output Format:
- PASS: All traceability comments are correctly formatted.
- FAIL: Issues found. List each with the test function name, line
number, and a concise description of the issue.
update_documents_relating_to_src_deepwork:
description: "Ensure project documentation stays current when DeepWork source files, plugins, or platform content change."
match:
include:
- "src/deepwork/**"
- "plugins/**"
- "platform/**"
- "pyproject.toml"
- "README.md"
- "README_REVIEWS.md" # Documents the review system — must stay in sync with src/deepwork/review/
- "CLAUDE.md"
- "doc/architecture.md"
- "doc/mcp_interface.md"
- "doc/job_yml_guidance.md" # Describes job.yml runtime behavior — must stay in sync with parser/schema
- "doc/doc-specs.md"
- "doc/platforms/**"
- "src/deepwork/hooks/README.md"
exclude:
- "**/__pycache__/**"
review:
strategy: matches_together
instructions: |
When source files in src/deepwork/, plugins/, or platform/ change, check whether
the following documentation files need updating:
- README.md (install instructions, feature descriptions, project overview)
- README_REVIEWS.md (review system usage, .deepreview config format, strategies, examples)
- CLAUDE.md (project structure, tech stack, CLI commands, architecture)
- doc/architecture.md (detailed architecture, module descriptions, data flow)
- doc/mcp_interface.md (MCP tool parameters, return values, server config)
- doc/job_yml_guidance.md (job.yml field descriptions, runtime behavior, review cascade)
- doc/doc-specs.md (doc spec format, parser behavior, quality criteria)
- doc/platforms/claude/ (Claude-specific hooks, CLI config, learnings)
- doc/platforms/gemini/ (Gemini-specific hooks, CLI config, learnings)
- src/deepwork/hooks/README.md (hook system, event mapping, tool mapping)
Read each documentation file and compare its content against the changed source
files. Pay particular attention to:
- Directory tree listings (do they still match the actual filesystem?)
- CLI command descriptions (do flags, behavior descriptions still match?)
- MCP tool parameters and return values (do they match the implementation?)
- Module/class/function descriptions (do they match what the code actually does?)
- Dependency lists (does pyproject.toml still match what the doc says?)
- Install/usage instructions (do they still work?)
- Hook event/tool mappings (do tables match the implementation?)
Flag any sections that are now outdated or inaccurate due to the changes.
If the documentation file itself was changed, verify the updates are correct
and consistent with the source files.
additional_context:
unchanged_matching_files: true
update_learning_agents_architecture:
description: "Ensure learning agents documentation stays current when the learning_agents plugin changes."
match:
include:
- "learning_agents/**"
- "doc/learning_agents_architecture.md"
exclude:
- "learning_agents/**/transcripts/**"
review:
strategy: matches_together
instructions: |
When files in learning_agents/ change, check whether doc/learning_agents_architecture.md
needs updating. Read the architecture doc and compare against the changed source files.
Pay attention to:
- Plugin structure listings (do they match the actual directory layout?)
- Skill descriptions (do they match what skills actually do?)
- Learning cycle descriptions (does the flow still match the implementation?)
- File format descriptions (do they match actual file formats used?)
- Configuration examples (do they still work?)
Flag any sections that are now outdated or inaccurate due to the changes.
If the doc itself was changed, verify the updates are correct.
additional_context:
unchanged_matching_files: true
agent_tools_fully_qualified:
description: "Claude Code agent definition files must list every tool by fully-qualified name, never by wildcard."
match:
include:
- "**/agents/*.md"
review:
strategy: individual
instructions: |
This file defines a Claude Code subagent (the YAML frontmatter at the
top of the file configures the agent's name, description, model, and
tool grants). Check the `tools:` frontmatter field.
Rule: every tool entry MUST be a fully-qualified tool name. Wildcard
or glob patterns (any entry containing `*`, `?`, `[`, or ending in
`__*`) are non-conforming.
Rationale: wildcard patterns in subagent `tools:` frontmatter do not
reliably match deferred MCP tools at runtime. Observed failure mode:
a frontmatter entry like `mcp__deepwork-dev__*` is accepted at parse
time, but when the subagent tries to invoke
`mcp__deepwork-dev__mark_review_as_passed` the runtime responds with
`Error: No such tool available`. Enumerating each MCP tool by its
full name (e.g., `mcp__deepwork-dev__mark_review_as_passed`) avoids
this failure mode and also documents the exact tool surface the
agent relies on. This caught us in the e2e merge-queue run that
failed PR #390.
This rule does NOT apply to files that are not Claude Code agent
definitions — if the file has no YAML frontmatter with name/tools
fields (e.g., it is a context-injection document or a skill body
that happens to live under an `agents/` path), this review passes
vacuously.
Check for:
- Any entry in `tools:` containing `*`, `?`, or `[...]` glob syntax.
- Any entry ending in `__*` (the common MCP-wildcard form).
- Any quoted glob pattern such as `"mcp__<server>__*"`.
Output Format:
- PASS: `tools:` field is missing, empty, or every entry is a
fully-qualified name.
- FAIL: list each wildcard entry with its line number and the
fully-qualified tool names it should be replaced with (if
determinable from the agent's body; otherwise flag the need for
enumeration).
shell_code_review:
description: "Review shell scripts for correctness, safety, and project conventions."
match:
include:
- "**/*.sh"
review:
strategy: individual
instructions: |
Review this shell script against the project's shell scripting conventions.
Project shell conventions (observed from existing scripts):
- Shebang: `#!/usr/bin/env bash` (NOT `#!/bin/bash` — required for NixOS compatibility)
- Safety: use `set -e` or `set -euo pipefail`
- Variables: UPPERCASE_WITH_UNDERSCORES, always quoted ("$VAR")
- JSON parsing: use jq with fallback defaults (// empty, // "default")
- Header comment block: description, usage, input/output, exit codes
- Section markers: # ==== blocks for logical sections in longer scripts
- Exit codes: 0 for success, 1 for usage errors, 2 for blocking errors
Check for:
- Hardcoded shebang paths (e.g., `#!/bin/bash`) — must use `#!/usr/bin/env bash`
- Unquoted variables that could cause word splitting or globbing
- Missing error handling (commands that could fail silently)
- Unsafe use of eval, unvalidated user input, or command injection risks
- Portability issues (bashisms that break on other shells, if applicable)
- Proper use of jq for JSON parsing (not grep/sed on JSON)
Additionally, always check:
- **DRY violations**: Is there duplicated logic or repeated patterns across
sections that should be extracted into a function?
- **Comment accuracy**: Are all comments (especially header blocks and
section markers) still accurate after the changes? Flag any comments
that describe behavior that no longer matches the code.
Output Format:
- PASS: No issues found.
- FAIL: List each issue with file, line, severity (high/medium/low), and a concise description.
agents_md_claude_md_symlink:
description: "Ensure every AGENTS.md file has a sibling CLAUDE.md symlink pointing to it, because Claude Code reads CLAUDE.md but ignores AGENTS.md."
match:
include:
- "**/AGENTS.md"
review:
strategy: individual
instructions: |
This is an auto-fix rule — you SHOULD edit files to resolve issues.
Claude Code reads CLAUDE.md for project context but does not read AGENTS.md.
Every AGENTS.md file MUST have a sibling CLAUDE.md that is a symlink pointing
to AGENTS.md, so that Claude Code picks up the same content.
Check whether a CLAUDE.md symlink exists in the same directory as the matched
AGENTS.md file. If it is missing or is not a symlink to AGENTS.md, create it:
ln -sf AGENTS.md <dir>/CLAUDE.md
The symlink target MUST be the relative path "AGENTS.md" (not an absolute path),
so the link remains valid after the directory is moved or cloned.
Output Format:
- PASS: CLAUDE.md symlink exists and points to AGENTS.md.
- FAIL: Symlink missing or incorrect — describe the current state and the fix applied.
nix_claude_wrapper:
description: "Ensure flake.nix always wraps the claude command with the required plugin dirs."
match:
include:
- "flake.nix"
- ".envrc"
review:
strategy: matches_together
instructions: |
The nix dev shell must ensure that running `claude` locally automatically
loads the project's plugin directories via `--plugin-dir` flags. Verify:
1. **Wrapper exists**: flake.nix creates a wrapper (script or function)
that invokes the real `claude` binary with extra arguments.
2. **Required plugin dirs**: The wrapper MUST pass both of these
`--plugin-dir` flags:
- `--plugin-dir "$REPO_ROOT/plugins/claude"`
- `--plugin-dir "$REPO_ROOT/learning_agents"`
3. **PATH setup**: The wrapper must be discoverable — either via a
script placed on PATH (e.g. `.venv/bin/claude`) with `.envrc`
adding that directory to PATH, or via a shell function/alias.
4. **Real binary resolution**: The wrapper must resolve the real
`claude` binary correctly, avoiding infinite recursion (e.g. by
stripping the wrapper's directory from PATH before lookup).
Output Format:
- PASS: The claude wrapper is correctly configured with both plugin dirs.
- FAIL: Describe what is missing or broken.
mcp_interface_consumer_sync:
description: "When MCP tool interfaces change, verify that skills, hooks, and platform content still reference correct parameter names, types, and behavior."
match:
include:
- "src/deepwork/jobs/mcp/tools.py"
- "src/deepwork/jobs/mcp/schemas.py"
- "src/deepwork/jobs/mcp/server.py"
review:
strategy: matches_together
additional_context:
unchanged_matching_files: true
instructions: |
When the MCP tool interfaces (tools.py, schemas.py, server.py) change,
verify that downstream consumers still accurately describe the MCP
interface. Check each of these files for stale parameter names, removed
fields, renamed concepts, or outdated behavior descriptions:
- platform/skill-body.md
- plugins/claude/skills/deepwork/SKILL.md
- plugins/gemini/skills/deepwork/SKILL.md
- plugins/claude/hooks/post_compact.sh
- doc/mcp_interface.md
For each file, compare any MCP parameter names, tool names, response
fields, or workflow behavior descriptions against the actual schemas
and tool implementations. Flag anything that no longer matches.
Output Format:
- PASS: All consumer files are consistent with the current MCP interface.
- FAIL: List each inconsistency with the file, the stale content, and
what it should say based on the current implementation.