Presidio is a Python-based data protection and de-identification SDK with multiple components for detecting and anonymizing PII (Personally Identifiable Information) in text and images.
Data Privacy is Paramount - This is a PII detection and anonymization system used in sensitive contexts. Security and correctness are non-negotiable.
Key Principles:
- Accuracy First: False negatives (missed PII) and false positives (incorrect detections) both damage trust
- Security by Default: Never log PII values, use non-reversible anonymization, validate all inputs
- Cross-Component Awareness: Presidio is a multi-component system - changes ripple across boundaries
- Stateless Design: Presidio is designed for scalability - avoid adding unnecessary state
- Documentation Integrity: Code and docs must stay synchronized - outdated docs are dangerous
Use these guidelines when generating or writing code for Presidio.
Data Flow (Unidirectional):
Analyzer (detect PII) → Anonymizer (transform PII) → Output
↓
nlp_engine → recognizers → context
Design Patterns:
- Registry Pattern:
RecognizerRegistryfor dynamic recognizer management - Provider Pattern:
NlpEngineProvider,RecognizerRegistryProviderfor configuration
1. Choose the Right Base Class:
from presidio_analyzer import PatternRecognizer, LocalRecognizer, RemoteRecognizer
# For regex-based detection
class MyPatternRecognizer(PatternRecognizer):
pass
# For custom logic (NLP, ML)
class MyCustomRecognizer(LocalRecognizer):
def load(self): ...
def analyze(self, text, entities, nlp_artifacts): ...
# For calling remote services
class MyRemoteRecognizer(RemoteRecognizer):
def analyze(self, text, entities, nlp_artifacts): ...2. Predefined Recognizers Location Matters:
- Country-specific:
presidio-analyzer/presidio_analyzer/predefined_recognizers/country_specific/{country}/ - Generic patterns:
.../predefined_recognizers/generic/ - NLP/ML-based:
.../predefined_recognizers/nlp_engine_recognizers/or.../ner/ - Third-party:
.../predefined_recognizers/third_party/
Directory names are the full lowercase country name (south_africa, philippines,
canada), not the ISO country code. The only exceptions are the pre-existing us
and uk directories. Do not add new abbreviated directories.
Language codes are different: supported_language and the YAML supported_languages
key take ISO 639-1 language codes (ko for Korean), not country codes (kr). A
mismatch here produces a recognizer that never loads.
3. Pattern Design Best Practices:
# ❌ BAD: Too broad - matches month names as persons
pattern = r"\b[A-Z][a-z]+\b"
# ✅ GOOD: Specific pattern with context
PATTERNS = [
Pattern(
"SSN (medium)",
r"\b\d{3}-\d{2}-\d{4}\b",
0.3
)
]
CONTEXT = ["ssn", "social security", "tax id"]Score bands. The score must reflect how much the pattern alone narrows the space, independent of any threshold applied downstream.
| Score | Use when | Name the pattern |
|---|---|---|
| 0.05 - 0.1 | Bare digit or alphanumeric runs with no structure | "(very weak)" |
| 0.1 - 0.3 | Some structure: delimiters, a prefix, a length constraint | "(weak)" |
| 0.3 - 0.5 | Distinctive format, no validation | "(medium)" |
| 0.5+ | Distinctive format | "(strong)" |
Assigning 0.3 to a pattern that also matches covid19 and sha256 overstates it.
Compare against existing recognizers before choosing: UsPassportRecognizer uses
0.05 for nine bare digits, UsBankRecognizer uses a weak score for 8-17 digits.
Suppress with thresholds, not by requiring context. Use score_thresholds to
keep low-confidence matches out of default results. Do not design a recognizer that
cannot fire without context: presidio-structured counts matches per column and has
no surrounding context to work with.
Context words are matched as substrings. LemmaContextAwareEnhancer defaults to
matching_mode="substring", so short context words fire on unrelated tokens.
# ❌ BAD: "member" matches "remember", "auth" matches "author" and "OAuth",
# "claim" matches "disclaimer"
CONTEXT = ["member", "auth", "claim"]
# ✅ GOOD: long enough to be unambiguous
CONTEXT = ["member id", "subscriber", "prior authorization"]Context is prefix-only by default (context_prefix_count=5, context_suffix_count=0),
so a context word appearing after the match does not boost the score. Test both
placements.
4. Document Pattern Sources:
"""
Recognizes US Social Security Numbers.
Pattern based on SSA Publication No. 05-10633:
https://www.ssa.gov/history/ssn/geocard.html
Validation uses SSN format rules: AAA-GG-SSSS
- AAA: Area number (001-899, excluding 666)
- GG: Group number (01-99)
- SSSS: Serial number (0001-9999)
"""5. Required Configuration Updates:
# Update all of these:
# 1. presidio_analyzer/predefined_recognizers/__init__.py
from .country_specific.us.my_recognizer import MyRecognizer
__all__ = [..., "MyRecognizer"]
# 2. presidio_analyzer/predefined_recognizers/country_specific/us/__init__.py
from .my_recognizer import MyRecognizer
__all__ = [..., "MyRecognizer"]
# 3. presidio_analyzer/conf/default_recognizers.yaml
recognizers:
- name: MyRecognizer
supported_languages: ["en"]
type: predefined
enabled: false
country_code: us
# 4. docs/supported_entities.md (add row to appropriate table)Enabled by default or not. The question is false-positive surface, not which
country the entity belongs to. Default to enabled: false and justify anything else
in the PR description.
A recognizer may ship enabled only when all of these hold:
- The pattern is structurally distinctive (delimiters, fixed prefixes, or a checksum that is mandatory rather than regional)
- A coincidental match on ordinary text is implausible, not merely unlikely
- Failing validation lowers the score rather than leaving the base score intact
validate_result returns None on mismatch, the base score stands and every lookalike
token still surfaces at that score.
6. Test the Configuration Path, Not Just the Constructor:
Most predefined recognizers ship enabled: false, so the default test run never
constructs them from configuration. A recognizer that works when built in Python can
still be unreachable, or crash, when a user enables it in YAML. Every new recognizer
needs at least one test that goes through the registry.
def test_recognizer_loads_and_detects_when_enabled_in_yaml(tmp_path):
"""Detection must work through the path users actually configure."""
conf = tmp_path / "recognizers.yaml"
conf.write_text(
"""
supported_languages:
- en
recognizers:
- name: MyRecognizer
supported_languages:
- en
type: predefined
enabled: true
country_code: us
"""
)
registry = RecognizerRegistryProvider(
conf_file=conf
).create_recognizer_registry()
analyzer = AnalyzerEngine(registry=registry, nlp_engine=nlp_engine)
results = analyzer.analyze("Member ID ABC123456", language="en")
assert [result.entity_type for result in results] == ["MY_ENTITY"]This catches, at minimum:
- Constructor signatures incompatible with the keys the loader passes (
name,supported_entity,context) - Class name typos and missing
__init__.pyexports country_codemismatches between the class attribute and the YAML entry- Class-level defaults (thresholds, context) that configuration silently discards
- A recognizer whose declared languages are excluded by the top-level
supported_languagesfilter, which loads nothing and reports no error
supported_languages key in the config
acts as a global filter, and the shipped default is ["en"]. A recognizer supporting
only de will not load from that config, silently. State the required top-level
languages in the PR description and cover this in the test above.
Behavior must not depend on how the recognizer was constructed. Building it
directly, adding it with registry.add_recognizer(), and loading it from configuration
must all produce the same recognizer. Any defaulting or validation applied on one path
belongs on all of them.
7. Comprehensive Test Coverage:
@pytest.mark.parametrize("text, expected_len, expected_positions", [
# True positives - valid formats
("SSN: 456-78-9012", 1, ((5, 16),)),
("My SSN is 456-78-9012", 1, ((10, 21),)),
# True negatives - invalid formats
("SSN: 000-00-0000", 0, ()), # Invalid area
("SSN: 666-12-3456", 0, ()), # Excluded area
("SSN: 123-45-6789", 0, ()), # Well-known sample SSN, denylisted
# Boundary testing - embedded in text
("Contact: 456-78-9012 for info", 1, ((9, 20),)),
])
def test_ssn_detection(text, expected_len, expected_positions, recognizer):
results = recognizer.analyze(text, ["US_SSN"])
assert len(results) == expected_len
for result, (start, end) in zip(results, expected_positions):
assert result.start == start
assert result.end == end123-45-6789, 078-05-1120) are denylisted by UsSsnRecognizer, so a test using them
as true positives fails, and one using them as a false-positive case passes for the
wrong reason.
Assert exact scores, not ranges. A range assertion passes even when the logic that produces the score breaks entirely.
# ❌ BAD: still passes if checksum promotion stops working
assert 0.5 <= result.score <= 1.0
# ✅ GOOD: pins the behavior under test
assert result.score == pytest.approx(EntityRecognizer.MAX_SCORE)Include a lookalike negative. The false-positive surface is the thing worth testing, not the happy path. Add a case proving that a plausible non-PII token of the same shape is not flagged: a 17-character order ID for a VIN, a legal citation for a bank account number, a build tag for an alphanumeric member ID.
Exercise context enhancement. A recognizer that defines CONTEXT needs a test
showing the score changes between text with and without a context word. A suite that
never triggers the enhancer does not test the context words at all.
Presidio is a library. Changes to shared classes alter results for users who have written no new code. Before changing anything outside a new file, state in the PR description what existing behavior changes.
These count as behavior changes even without a signature change:
- Default values on shared base classes (
Noneto[]changes truthiness for every subclass) - Properties on abstract interfaces, since custom implementations inherit the new default and may break
- Scores, context lists, or patterns on existing recognizers
- Anything altering which entities are returned for text that previously worked
Surface new scoring inputs in explainability. Anything that changes how a score is
derived (context, negative context, thresholds) must be reflected in
AnalysisExplanation, or users cannot tell why a result scored the way it did.
Prefer warnings over exceptions when the caller cannot fix the condition. Raising on a configuration a user did not write, and cannot change, turns a degraded result into a hard failure. Where a lookup falls back to a default instead of failing, add a debug log so the fallback is discoverable.
Prefer a property on the base class over a maintained list of class names. Lists drift as recognizers are added, and users installing from PyPI cannot extend them.
1. Implement the Operator Interface:
from presidio_anonymizer.operators import Operator, OperatorType
class MyOperator(Operator):
"""Custom anonymization operator."""
def operate(self, text: str, params: dict = None) -> str:
"""Transform the detected PII."""
# Ensure non-reversible transformation
import uuid
return f"<{params.get('entity_type', 'REDACTED')}_{uuid.uuid4().hex[:8]}>"
def validate(self, params: dict = None) -> None:
"""Validate operator parameters before use."""
if params and 'entity_type' not in params:
raise ValueError("entity_type is required")
def operator_name(self) -> str:
return "my_operator"
def operator_type(self) -> OperatorType:
return OperatorType.Anonymize2. Security Checklist:
- ✅ Non-reversible: Cannot recover original PII from anonymized output
- ✅ Entropy: Uses random/unpredictable values (not deterministic hashing)
- ✅ No PII leakage: Doesn't preserve PII characteristics (length, format)
# ❌ BAD: Reversible via rainbow tables
def operate(self, text, params):
return hashlib.md5(text.encode()).hexdigest()
# ✅ GOOD: Non-reversible with entropy
def operate(self, text, params):
return f"<{params['entity_type']}_{uuid.uuid4().hex[:8]}>"3. Test Anonymization Quality:
def test_operator_is_non_reversible():
"""Verify same input produces different output."""
operator = MyOperator()
result1 = operator.operate("John Doe", {"entity_type": "PERSON"})
result2 = operator.operate("John Doe", {"entity_type": "PERSON"})
assert result1 != result2 # Different each time
def test_operator_preserves_structure():
"""Verify anonymized text maintains sentence structure."""
text = "Email: john@example.com, Phone: 555-1234"
# After anonymization
expected = "Email: <EMAIL_xxx>, Phone: <PHONE_yyy>"
# Structure preserved, PII replaced1. Maintain Backward Compatibility:
# ❌ BAD: Breaking change
def analyze(text: str, language: str, entities: List[str]):
...
# ✅ GOOD: Optional parameter with default
def analyze(
text: str,
language: str,
entities: Optional[List[str]] = None
) -> List[RecognizerResult]:
...2. Required Updates for API Changes:
# 1. Update OpenAPI schema
docs/api-docs/api-docs.yml
# 2. Add E2E tests
e2e-tests/tests/test_new_endpoint.py
# 3. Add usage example
docs/samples/python/new_feature_example.ipynb
When modifying shared interfaces:
- Identify all consumers:
# RecognizerResult is consumed by:
# - presidio-anonymizer (takes analyzer results)
# - presidio-cli (displays results)
# - presidio-structured (processes tabular data)
# - docs/samples/* (user examples)- Update all components in same changeset:
# If adding field to RecognizerResult:
# 1. presidio-analyzer: Add field and populate
# 2. presidio-anonymizer: Handle new field (or ignore safely)
# 3. presidio-cli: Display new field (optional)
# 4. Tests: Update expectations
# 5. Docs: Document new field
# 6. e2e-tests: Add integration test for new field- Respect component boundaries:
# ❌ BAD: Anonymizer importing analyzer internals
from presidio_analyzer.predefined_recognizers import UsSsnRecognizer
# ✅ GOOD: Use public interfaces only
from presidio_analyzer import RecognizerResult1. Cache Compiled Regexes:
from functools import lru_cache
@lru_cache(maxsize=128)
def _compile_pattern(pattern_str: str) -> re.Pattern:
return re.compile(pattern_str, re.IGNORECASE)2. Avoid Catastrophic Backtracking:
# ❌ BAD: O(2^n) on "aaaa...ab"
pattern = r"(a+)+"
# ✅ GOOD: Atomic grouping
pattern = r"(?>a+)"3. Batch NLP Processing:
# ❌ BAD: Process one at a time
for text in texts:
doc = nlp(text)
...
# ✅ GOOD: Use spaCy pipe for batching
for doc in nlp.pipe(texts, batch_size=50):
...Test Naming Convention:
# ✅ GOOD: Descriptive, intention-revealing
def test_when_valid_ssn_then_detect_with_correct_boundaries()
def test_when_invalid_checksum_then_no_match()
def test_when_context_missing_then_low_confidence()
# ❌ BAD: Non-descriptive
def test_ssn_1()
def test_case2()1. Required Documentation Updates:
When adding a feature, update ALL of:
✅ docs/supported_entities.md - For new entity types
✅ docs/api-docs/api-docs.yml - For API changes
✅ README.md - For major features
✅ Docstrings - All public classes/methods
✅ docs/samples/ - Usage examples for complex features
✅ Update docstrings based on the reST docstring format (:param:, :return:, :raises:, :example:)Do not update CHANGELOG.md in a PR. Before each version bump, changelog
entries for the current release are generated from merged PRs. Per-PR changelog
edits create unnecessary merge conflicts.
2. Pattern Source Documentation:
# In recognizer docstring or comments
"""
Pattern based on Royal Mail PAF specification:
https://www.royalmail.com/find-a-postcode
UK postcodes follow 6 formats:
- A9 9AA (e.g., M1 1AA)
- A99 9AA (e.g., M60 1NW)
- AA9 9AA (e.g., CR2 6XH)
- AA99 9AA (e.g., DN55 1PT)
- A9A 9AA (e.g., W1A 1HQ)
- AA9A 9AA (e.g., EC1A 1BB)
Plus special case: GIR 0AA
"""Use these guidelines when reviewing pull requests for Presidio.
- Only comment when you have HIGH CONFIDENCE (>80%) that an issue exists
- Be concise: one sentence per comment when possible
- Focus on actionable feedback, not observations
- Data privacy is paramount - this is a PII detection/anonymization system
- All modules in Presidio which process records are stateless - avoid suggesting stateful solutions
- Presidio is a multi-component system - consider cross-component impacts of changes
- Don't reinvent the wheel - check for existing patterns, functions and best practices in the codebase before suggesting new approaches
Focus on issues in this order of importance:
PII-Specific Risks:
- PII leakage in logs, error messages, or debug output - Never log detected PII values, only entity types and positions
- Regex injection vulnerabilities - User-provided patterns must be validated before compilation
- Inadequate anonymization - Reversible transformations, weak masking, deterministic fake data without proper context
- Side-channel leaks - Timing attacks revealing PII presence, cache-based information disclosure
General Security:
- Hardcoded secrets, API keys, credentials (especially for NLP model endpoints, cloud services)
- Command injection (especially in CLI component)
- Unsafe deserialization (pickle files, untrusted NLP models)
- Missing input validation on API endpoints (analyzer, anonymizer, image-redactor)
- Path traversal in file operations
- Insecure random number generation for fake data
PII Detection Accuracy:
- False positives - Overly broad regex patterns matching non-PII or other entity types with med/high confidence
- False negatives - Missing valid cases for a given entity type, especially edge cases or common variations
- Incorrect entity boundaries - Off-by-one errors in start/end positions causing malformed anonymization
- Confidence score miscalculation - Scores outside [0.0, 1.0], incorrect aggregation of multiple detection methods
General Logic:
- Race conditions in multi-threaded analysis
- Resource leaks (NLP models not released, file handles, network connections)
- Null/None handling in entity detection chains
- Incorrect error handling that silently fails to detect PII
- Adding state where unnecessary (Presidio is designed to be stateless for scalability)
PII Detection Specific:
- Inefficient regex patterns - Catastrophic backtracking (e.g.,
(a+)+bon "aaaa...a") - Redundant passes - Running the same logic multiple times on same text
- Unbounded batch processing - Loading entire datasets into memory
- Missing regex compilation caching - Recompiling patterns on every call
- Unnecessary model loads - Loading the same model multiple times instead of reusing instances
General Performance:
- O(n²) or worse algorithms when O(n) exists
- Blocking I/O on critical API paths
- Missing database indexes for entity result storage
- Inefficient image processing (loading full image when bounding box would suffice)
Respect the Natural Data Flow:
- Presidio follows a unidirectional flow: Analyzer → Anonymizer → Output
- Downstream components (CLI, structured, image-redactor) consume analyzer/anonymizer, never the reverse
- Breaking this flow creates circular dependencies and tight coupling
- Changes should propagate forward through the data pipeline, not backward
Module Reuse Guidelines:
- Reuse code by importing from shared modules, not by copying code across components
- Shared data models (RecognizerResult, OperatorConfig) should be treated as contracts - changes require coordinated updates across all consumers
- When adding functionality, check if it belongs in an existing shared module rather than duplicating logic
- If multiple components need the same feature, extract it to a common location rather than implementing it multiple times
- Backward compatibility is critical when modifying shared modules - ensure existing consumers continue to work without changes
Avoid Cross-Component Side Effects:
- Changes to internal implementation should not affect other components' behavior
- Modifying shared configuration files requires understanding impact on all components that consume them
- Registry and provider patterns exist to decouple components - bypassing them creates hidden dependencies
- Component boundaries must be respected: anonymizer should never import from analyzer internals, only public interfaces
- Providing a solution specific to one component in a shared module instead of providing a general solution that can be used by multiple components creates tight coupling and maintenance challenges
When Making Changes Across Components:
- Identify all components that consume the interface you're modifying
- Update dependent components in the same changeset to maintain system consistency
- Ensure configuration files, API schemas, and documentation stay synchronized
- Test the complete integration path, not just individual components in isolation in unit tests, integration tests, and the e2e test suite
- Communicate changes clearly in the PR description, especially if they affect multiple components or require coordinated
Presidio Patterns:
- Recognizer design violations - Not inheriting from
EntityRecognizer, missingload()oranalyze() - Operator design violations - Not implementing
OperatorTypeinterface correctly - Registry pattern misuse - Bypassing
RecognizerRegistry, hardcoding recognizer lists - Provider pattern violations - Not following
NlpEngineProviderorRecognizerRegistryProviderpatterns - Tight coupling - Recognizers depending on specific NLP engine implementation details
General Design:
- Circular dependencies between modules
- Missing abstraction for third-party service integrations
- Breaking existing public APIs without deprecation warnings
- Inconsistent error handling strategies (mixing exceptions and error codes)
Input Validation:
- Missing validation of user-provided entity types
- Accepting arbitrary regex patterns without safety checks
- No length limits on input text (DoS via memory exhaustion)
- Missing validation of parameters
- Unchecked file uploads
Output Validation:
- Confidence scores outside valid range
- Overlapping entity spans not handled correctly
- Missing entity type in anonymization results
Presidio-Specific Testing:
- Missing tests for new recognizers - Must include: true positives, true negatives, edge cases, false positive scenarios, entity within larger context
- No validation of entity boundaries - Tests only check entity type, not exact start/end positions
- Missing multilingual tests - Recognizers claiming multi-language support without language-specific tests
- Anonymization reversibility not tested - No verification that anonymized data can't be de-anonymized
- Missing E2E analyzer→anonymizer tests - Testing components in isolation without integration validation
- No configuration-path test for a new recognizer - Tests construct the recognizer in Python only. Recognizers shipping
enabled: falseare never exercised throughRecognizerRegistryProvider, which is how users enable them. Require one test that enables the recognizer in a YAML config and asserts detection - Construction paths that disagree - Behavior differs depending on whether a recognizer is built directly, added via
registry.add_recognizer(), or loaded from configuration. Flag defaulting or validation logic applied on one path but not the others - Score assertions using ranges -
assert 0.5 <= score <= 1.0passes even when checksum promotion or context enhancement breaks. Require exact assertions - No lookalike negative - Tests cover valid values and malformed values, but not a plausible non-PII token of the same shape, which is the actual false-positive surface
General Testing:
- Missing tests for critical business logic (PII detection, anonymization)
- Flaky tests due to non-deterministic NLP/ML models (use fixed random seeds)
- Tests that don't validate behavior (checking implementation details instead)
- Missing regex pattern edge cases (empty strings, special characters, unicode)
Code-Documentation Consistency:
- Code changes must be reflected in documentation - outdated docs are misleading and dangerous
- Implementation must not contradict existing documentation - if conflict exists, either update docs or reconsider implementation
- API documentation is auto-generated from docstrings - formatting errors break the build
Terminology:
- Use "threshold", not "cutoff", to match the codebase
- Use ISO 639-1 language codes in docs and configuration examples
Docstring Quality:
- All public classes, methods, and functions must have docstrings
- Docstrings must follow consistent format (Args, Returns, Raises, Examples)
- No formatting issues that break API doc generation (malformed RST/Markdown, incorrect indentation)
- Include type information in docstrings when not obvious from type hints
Documentation for New Features:
- New recognizers must document pattern sources - link to official standards, government specifications, or authoritative references
- Complex additions require usage examples in
docs/samples/- show common use cases, not just API reference - New entity types must be added to
docs/supported_entities.mdwith description and example - API changes require updates to
docs/api-docs/api-docs.yml(OpenAPI schema)
Pattern Recognizer Documentation:
- Explain the logic source: "Based on ISO standard X", "Follows format defined by Y government agency"
- Document regex pattern rationale - why specific character classes, lookaheads, or groups are needed
- Include references to validation algorithms (e.g., "Luhn checksum validation per ISO/IEC 7812")
- Note any limitations or known edge cases in the pattern
- Overly complex recognizer logic (>50 lines in
analyze()method, >3 nesting levels) - Misleading variable names (e.g.,
patternfor compiled regex, should becompiled_pattern) - Missing docstrings on public recognizer/operator classes
- Incomplete type hints on public APIs (especially
analyze(),anonymize()signatures)
DO NOT comment on these (handled by automated tools):
- ❌ Code formatting, line length, indentation (handled by
ruff format) - ❌ Import ordering (handled by
ruff check --select I) - ❌ Trailing commas, whitespace (handled by
ruff) - ❌ Type hint style preferences (
List[str]vslist[str]- both valid for Python 3.9-3.12 support)
DO NOT comment on style preferences that don't affect correctness:
- Personal preferences for syntax variations
- Subjective naming when current name is clear in PII context
- Minor refactoring suggestions that don't fix bugs or improve accuracy
- Unnecessary abstractions "for future flexibility" in recognizers
Be specific and actionable:
✅ GOOD: "🔴 CRITICAL: Line 45 logs detected PII value. Change logger.info(f'Found: {entity.text}')
to logger.info(f'Found entity type: {entity.entity_type}')"
❌ BAD: "Don't log PII"
Provide context:
✅ GOOD: "🟡 Important: This regex has catastrophic backtracking on input 'aaaaaa...b' (O(2^n) time).
Use atomic grouping: (?>a+)b or possessive quantifier a++b"
❌ BAD: "This regex is slow"
Differentiate severity:
- 🔴 CRITICAL - Security, data leakage, correctness bugs affecting PII detection accuracy
- 🟡 Important - Performance issues, cross-component breaks, missing tests for new recognizers
- 💡 Suggestion - Code quality improvements, better error messages, optimization opportunities
Acknowledge good practices:
- Well-tested recognizers with comprehensive edge case coverage
- Proper use of context validation (NLP + regex)
- Good error handling with informative messages
- Performance optimizations (regex caching, batch processing)
- Python -
requires-python = ">=3.10,<3.15". Code must run on every version in that range - uv - Dependency management and installation (not pip or Poetry). Each package commits a
uv.lock;poetry-coreis retained only as the build backend for now. - Ruff - Linting and formatting (replaces flake8, black, isort)
- spaCy - Default NLP engine (en_core_web_lg for production), although one can use other NLP engines via provider pattern
- Docker - Deployment via GitHub Container Registry (
ghcr.io/data-privacy-stack)
RecognizerResult- Shared analyzer output formatOperatorConfig- Anonymizer operator configurationconf/default_recognizers.yaml- System-wide recognizer registrydocs/supported_entities.md- Public entity type documentation- API schemas in
docs/api-docs/
# Setup (uv reads the committed uv.lock; --locked fails if it is stale)
cd presidio-analyzer # or presidio-anonymizer, presidio-cli, etc.
uv sync --locked --all-extras --group dev
uv run python -m spacy download en_core_web_lg # For analyzer/CLI only
# Run tests
uv run pytest -xvv # Stop on first failure with verbose output
uv run pytest tests/test_us_ssn_recognizer.py -k "test_valid" # Specific test
# Lint
uv run ruff check .
uv run ruff format .Dependency changes: whenever you edit a package's
pyproject.tomldependencies (add/remove/bump[project]deps, extras, or[dependency-groups]), you MUST regenerate and commit that package'suv.lockin the same change (cd <package> && uv lock). CI installs withuv sync --lockedand fails ifpyproject.tomlanduv.lockare out of sync, so an updatedpyproject.tomlwithout its matchinguv.lockwill break the build.
# Quick test with pre-built images
docker pull ghcr.io/data-privacy-stack/presidio-analyzer:latest
docker run -d -p 5002:3000 --name analyzer ghcr.io/data-privacy-stack/presidio-analyzer:latest
curl http://localhost:5002/health
# Full build from source (takes 15+ minutes)
docker compose up --build -ddocker-compose up -d # Start all services
cd e2e-tests
python -m venv venv && source venv/bin/activate
pip install -r requirements.txt
pytest -v # Run all E2E tests- Stale
uv.lock- Ifuv sync --lockedfails with "lockfile needs to be updated", runuv lockin that package and commit the result. - Missing spaCy models - Download en_core_web_lg before running tests
- AHDS test skips - Expected when AHDS_ENDPOINT not set
- Transformers test failures - Expected without HuggingFace access in restricted environments
- Recognizer enabled in YAML but never loads - Its declared languages are not in the top-level
supported_languages, which defaults to["en"]. The loader drops it silently TypeError: unexpected keyword argumenton registry construction - The recognizer's__init__does not accept a key the loader passes through from the YAML entry, most oftenname- Class defaults missing after loading from YAML - Check whether the loader assigns the value after construction rather than passing it to
__init__
- Logging PII values - Never log
entity.text, onlyentity.entity_type - Hardcoded language assumptions - Use
context.languageparameter - Missing None checks - NLP engines return None for empty/invalid text
- Unbounded regex backtracking - Test patterns with long strings
- Confidence score > 1.0 - Validate score normalization logic
See section 8 in Review Priorities above for comprehensive documentation guidelines.
When adding features, update:
- docs/supported_entities.md - For new entity types
- docs/api-docs/api-docs.yml - For API changes
- README.md - For major features
- Docstrings - All public classes and methods (ensure proper formatting for API doc generation)
- docs/samples/ - Add usage examples for complex new features
Do not update CHANGELOG.md in a PR. Its current-release entries are generated
from merged PRs before each version bump, avoiding conflicts between concurrent
contributions.
Consult these for detailed guidance:
- CONTRIBUTING.md - PR process, CLA, code of conduct
- docs/development.md - Build process, testing, CI/CD
- docs/analyzer/developing_recognizers.md - Recognizer best practices
- docs/analyzer/adding_recognizers.md - Step-by-step recognizer guide
- docs/anonymizer/adding_operators.md - Operator development guide
Summary for Code Review: Prioritize security (PII leakage), correctness (detection accuracy), and performance (regex efficiency). Ensure comprehensive testing for all recognizers. Let automated tools handle formatting. Focus on actionable, specific feedback with concrete fixes.