Anchors are verbatim transcript excerpts and the store is committed plaintext, so any secret that lands in a session must be redacted before the session is written. This doc enumerates exactly what the scrubber catches and what it doesn't — so you can make an informed call about whether you also want a human review step before pushing .lore/.
The scrubbing runs at ingest, before storage, before the LLM call, before anchors are written. It covers both message content and event meta — tool-call arguments (e.g. a token passed to a Bash command) live in meta, so meta is walked recursively and its string leaves are scrubbed too. See lore.scrub.scrub_text and lore.scrub.scrub_events for the implementation.
Each pattern is high-precision (anchored by a distinctive prefix or shape) to keep false positives low. The replacement is a labelled token so a reader can see what was redacted, not just that something was.
| # | Secret class | Recognizer | Replacement |
|---|---|---|---|
| 1 | RSA / EC / OpenSSH private-key blocks | -----BEGIN ... PRIVATE KEY----- … -----END ... PRIVATE KEY----- |
[REDACTED:private-key] |
| 2 | OpenAI / Anthropic / generic sk-* API keys |
sk-[A-Za-z0-9_-]{16,} |
[REDACTED:api-key] |
| 3 | AWS access-key IDs — long-term (AKIA) and STS temporary (ASIA) |
(?:AKIA|ASIA)[0-9A-Z]{16} |
[REDACTED:aws-key] |
| 4 | GitHub classic PAT | ghp_[A-Za-z0-9]{20,} |
[REDACTED:github-token] |
| 5 | GitHub fine-grained PAT | github_pat_[A-Za-z0-9_]{22,} |
[REDACTED:github-token] |
| 6 | Google API key | AIza[A-Za-z0-9_-]{35} |
[REDACTED:google-api-key] |
| 7 | Slack tokens (bot/user/app/refresh/config — any xox?-) |
xox[a-z]-[A-Za-z0-9-]{10,} |
[REDACTED:slack-token] |
| 8 | HuggingFace user-access tokens | hf_[A-Za-z0-9]{30,} |
[REDACTED:hf-token] |
| 9 | JWTs (3 base64url segments, starting with eyJ) |
eyJ[A-Za-z0-9_-]+\.[A-Za-z0-9_-]+\.[A-Za-z0-9_-]+ |
[REDACTED:jwt] |
| 10 | Connection-string passwords (postgres://, mongodb://, mysql://, redis://, amqp(s)://, mssql://) |
password group between user: and @host |
[REDACTED:uri-password] (scheme + user + host preserved) |
| 11 | Generic assignment shapes — password/passwd/secret/token/api_key/aws_secret_access_key/aws_session_token = value, where the value may be a quoted multi-word string or a bare token |
(?i)(?:password|passwd|secret|token|api[_-]?key|aws_secret_access_key|aws_session_token)\s*[:=]\s*(?:"…"|'…'|[^\s'"]{6,}) |
[REDACTED:secret] |
Patterns are tested individually in tests/unit/test_scrub.py (one test per class above, plus negative tests that ordinary prose, clean URLs, and clean tool-call meta survive untouched).
The scrubber is not a DLP system. It raises the floor on what reaches the model and the store; it is not a substitute for a security review when sensitive content is in play. Specifically:
- Custom or in-house token formats (e.g. an internal service that issues opaque tokens with no distinctive prefix). If your team has one, add a pattern to
src/lore/scrub.pyand a test intests/unit/test_scrub.py. - A raw AWS secret access key on its own — the 40-char secret is caught when it appears as
AWS_SECRET_ACCESS_KEY=…(pattern 11), but a bare 40-char base64-ish string with no surrounding keyword is indistinguishable from ordinary data and is left alone to avoid false positives. Same for bare high-entropy strings generally. - Free-form prose containing the secret value without one of the recognized markers (e.g. someone literally types
the password is hunter2— the generic assignment pattern catchespassword = …but not free prose). - Already-base64'd or otherwise encoded secrets the model received as opaque blobs.
- PII (emails, phone numbers, names) — out of scope for this module; if you need PII handling, layer a separate pass.
If you want a quick visual audit before sharing .lore/ with anyone, grep -RE 'REDACTED' .lore/ shows everything the scrubber caught (good sign that it is catching things), and git diff .lore/ shows exactly what changed since your last commit (manual review surface).
crewlore's default .gitignore excludes .lore/sessions/ — your captured transcripts never leave your machine. Only the compiled claims (.lore/claims/) and the rendered book (.lore/knowledge/) are committed by default, and those go through the scrubber first.
The recipe is small:
- Add a pattern + replacement to
_PATTERNSinsrc/lore/scrub.py. Put it before the generic-assignment pattern (which is last) so a more specific shape claims its match first. - Add a positive test (
test_redacts_<thing>) and verify a non-secret string isn't false-positive-redacted. - Update this doc (add a row to the table above).
If your pattern is rare / domain-specific (e.g. an internal-only token shape), keep it local; if it's a widespread public format that anyone running crewlore would benefit from, open a PR.
- Approve-before-push gate — a
lore reviewcommand that surfaces the diff in.lore/claims/since the last commit and asks the developer to acknowledge before staging. Tracked as the strong complement to automated scrubbing; not in v0.1 scope. - Pattern-coverage badge — a small CI job that fuzzes random strings against the pattern set and reports any new secret shapes that snuck through.
- Pluggable scrubbers — projects with their own DLP tooling could chain
crewlore's scrubber with theirs.
If you find a real-world secret shape that bypasses the current set, report it as a security issue rather than filing a public bug — give me a chance to ship the pattern before it's on Hacker News.