MCP tools are registered at startup after loading repository-root .env values:
- Always expose
process_binary_file - Expose
analyze_imageonly when vision configuration is available (VISION_API_KEYorOPENAI_API_KEY)
| Tool | Purpose |
|---|---|
process_binary_file |
Converts supported binary/document/archive files to structured output and saves to .local_read_mcp/. |
analyze_image |
Vision API analysis of images, result saved to .local_read_mcp/analysis/. Only registered when vision is enabled at startup. |
All output is written to .local_read_mcp/ in the current working directory. No files are written outside the working directory.
process_binary_file(file)
│
├─ format detection → BackendRegistry.select_best()
│ priority: VLM_HYBRID > SIMPLE
│
├─ SIMPLE (local converters, no model/API dependency)
│ └─ Built-in converters + MarkItDown fallback: PyMuPDF, mammoth, openpyxl, python-pptx, etc.
│
├─ VLM_HYBRID (requires MinerU + models, PDF only)
│ └─ MinerU hybrid-auto-engine
│ VLM layout → pipeline OCR/formula/table → middle_json
│ Engine: vLLM > LMDeploy > MLX-VLM > transformers (auto)
│
└─ chapter_split (internal, triggers for large PDFs)
├─ TocExtractor: PyMuPDF TOC → page label calibration → fixed chunk fallback
├─ ChunkPlanner: page ranges with overlap
├─ per-chunk backend processing → sliced PDF → intermediate.json
└─ merged output.md + structural_toc.json
Result → .local_read_mcp/<file>_<timestamp>/
├── intermediate.json (structured block representation)
├── output.md (markdown conversion)
├── index.json (section/table/figure index)
└── images/ (extracted images)
class BackendType(Enum):
AUTO = "auto"
SIMPLE = "simple"
VLM_HYBRID = "vlm-hybrid"- SIMPLE: Always available. Handles supported formats through built-in converters and the MarkItDown fallback.
- VLM_HYBRID: PDF only, requires MinerU + downloaded models. Calls
hybrid_analyze.doc_analyze()directly — nodo_parsecallback/tempdir pattern. - Selection:
VLM_HYBRID > SIMPLE(by available + format support).
Built into process_binary_file. Triggers when chapter_split != False and format is PDF.
Calibration of logical page numbers (TOC) to physical page indices:
page.get_label()— PDF /PageLabels structure- Heuristic text scan — match chapter titles in candidate pages
logical - 1— fallback
Chunk planning with configurable overlap and min_chunk_pages. Falls back to fixed-size chunks when no TOC or headings are detected.
MinerU is an external dependency (pip install local-read-mcp[mineru]). Models (~4.5GB total) are downloaded by MinerU's own tool, configured via mineru.json in the project root. The backend sets MINERU_TOOLS_CONFIG_JSON automatically at import time.
Calls MinerU APIs directly:
hybrid_analyze.doc_analyze()for VLM-HYBRID- Engine auto-selection is handled inside MinerU
- Output converted from MinerU's
middle_jsontoIntermediateJSON
| File | Purpose | Tracked |
|---|---|---|
.env |
Vision API key, base URL, model (used at startup for tool registration) | gitignored |
.env.example |
Template for .env | yes |
mineru.json |
MinerU model paths | gitignored |
mineru.json.template |
Template for mineru.json | yes |
The server loads .env from the project root, independent of the current working directory.
Existing process environment variables are not overwritten by .env values.
src/local_read_mcp/
├── server/
│ ├── app.py MCP tools registration entrypoint
│ │ process_binary_file always, analyze_image optional
│ ├── orchestrator.py Chunk planning, processing, merging
│ ├── utils.py Parameter compatibility helper
│ └── vision.py Vision API integration
├── backends/
│ ├── base.py BackendType enum, registry
│ ├── simple.py SimpleBackend (all formats)
│ └── mineru.py VlmHybridBackend (MinerU hybrid)
├── segmenter/
│ ├── toc_extractor.py TOC extraction + page calibration
│ └── chunk_planner.py Page range planning + overlap
├── converters/
│ ├── _compat.py Optional dependency imports (centralized)
│ ├── base.py DocumentConverterResult + constants
│ ├── utils.py apply_content_limit, html_to_markdown_result
│ ├── section_extractor.py extract_sections_from_markdown
│ ├── latex_fixer.py fix_latex_formulas
│ ├── pdf.py, docx.py, xlsx.py, pptx.py, html.py, ...
│ └── simple.py TextConverter, JsonConverter, YamlConverter, ...
├── output_manager.py .local_read_mcp/ directory management
├── markdown_converter.py Intermediate JSON → markdown
├── index_generator.py Section/table/figure index
└── intermediate_json.py Structured intermediate representation