English | 中文
Multi-phase PDF capture pipeline as an MCP server. Convert PDF documents into high-quality structured Markdown — with formula recognition, table extraction, and a built-in quality gate.
- Triple extraction engines
- pymupdf4llm (built-in) — zero setup, fast, always available
- marker (recommended) — highest quality for complex layouts
- MinerU (optional) — best for multi-column/InDesign PDFs, auto-managed in an isolated venv
- 14 MCP tools —
pdf_to_markdown,get_job_status,download_models,extract_tables,classify_document,pdf_info,setup_vlm,check_environment,install_engine - Async job mode — large PDFs convert in a background job (no MCP client timeouts); model downloads can be pre-fetched without time limits
- Optional VLM enhancement — plug in any vision-capable model (Qwen-VL, GLM-4V, MiniMax, Moonshot, OpenAI, local Ollama…) for better table/formula extraction. No extra dependencies needed — works out of the box with the base install.
- Quality gate — multi-dimensional QC (text completeness, heading structure, formula integrity, table coverage) plus content-aware audit rules that catch defects statistical checks miss (control chars, torn numeric columns, fused table headers, content loss)
- Progressive setup — works immediately with zero config; enhance with marker/VLM on demand
- Privacy-first — API keys are stored locally with
chmod 600and never echoed back in responses
# Base package (pymupdf engine + VLM support, ~80MB) — works immediately
pip install pdf-capture-mcp
# With marker engine (recommended for complex PDFs, includes PyTorch, ~2.5GB)
pip install "pdf-capture-mcp[marker]"
# Everything (marker + TATR table detection + DePlot charts, ~3GB)
pip install "pdf-capture-mcp[all]"For Qoder / Claude Desktop / Cursor, add to your mcp.json:
{
"mcpServers": {
"pdf-capture": {
"command": "uvx",
"args": ["pdf-capture-mcp"]
}
}
}With marker engine (recommended):
{
"mcpServers": {
"pdf-capture": {
"command": "uvx",
"args": ["pdf-capture-mcp[marker]"]
}
}
}Or if installed via pip:
{
"mcpServers": {
"pdf-capture": {
"command": "pdf-capture-mcp"
}
}
}Ask your AI agent things like:
"Convert ~/Downloads/paper.pdf to Markdown" "Extract all tables from this report" "Is this PDF a scanned document?"
The server works immediately with the built-in pymupdf engine. For higher quality on complex layouts, the agent can install marker on demand (or you can pre-install it).
| Tool | Description |
|---|---|
pdf_to_markdown |
Full pipeline: extract → clean → QC → repair → self-describing knowledge package (async for large PDFs) |
export_to_obsidian |
Copy a knowledge package into an Obsidian vault as a whole unit (idempotent) |
setup_embedding |
Configure an OpenAI-compatible embedding endpoint (OpenAI / MiniMax / BGE / Ollama) |
build_vector_index |
Index a knowledge package into embedded Qdrant (incremental, content-addressed) |
search_corpus |
Semantic search with metadata filters — this tool IS the RAG API |
batch_convert |
Directory-scale conversion (async job): dedup by doc_id, optional vault export + indexing |
get_job_status |
Poll background jobs (large conversions / model downloads) |
download_models |
Pre-download marker models (recommended on slow networks) |
extract_tables |
Table extraction (pdfplumber rules + optional TATR deep learning) |
classify_document |
Document type detection (academic paper, consulting report, …) |
pdf_info |
Fast metadata: page count, text layer, scanned detection |
setup_vlm |
Configure optional VLM enhancement (any vision-capable provider) |
check_environment |
Verify engines, dependencies, model cache, and network config |
install_engine |
Install marker/ml engines on behalf of the user |
Every conversion now produces a self-describing knowledge package — a folder any LLM agent can understand from one README read, and that drops into an Obsidian vault as a unit:
<out_dir>/<slug>/
├── <slug>.md main document — same name as folder ([[slug]] works),
│ YAML frontmatter included (title/doc_id/pages/qc_verdict)
├── README.md entry map: summary, file table, chunks schema
├── images/ extracted figures (relative refs — render everywhere)
├── tables/ p<page>_table_<n>.csv — independent extraction channel
└── data/
├── chunks.jsonl semantic chunks: heading_path, page, chunk_type,
│ content (hashed for chunk_id) + embed_text (context)
├── metadata.json doc-level metadata, manifest, content_hash
└── qc_report.json archived audit & repair record
Key properties (from two design-audit rounds):
- Content-addressed identity:
doc_id = sha256(pdf)[:16]; re-converting the same PDF overwrites its package (idempotent), regardless of filename. - Chunk ids are content-addressed too — editing one section re-embeds only that section when you rebuild a vector index later.
- Hierarchical chunking: heading-path metadata, tables/code as dedicated chunks (oversized tables split with the header repeated per part), figures chunked only when they carry a VLM description, page numbers anchored via a monotonic scan of the PDF text layer (scanned PDFs honestly report null).
- MD-110 cross-page table merge: adjacent same-column tables merge ONLY when PDF geometry proves a page break (bottom-edge + top-edge + no caption between) — same-looking but separate tables are never merged.
- Default output root:
$PDF_CAPTURE_OUTPUT_ROOTor~/Documents/pdf-capture. Passpackage=Falsefor the bare markdown layout of earlier versions.
Converting a big document (e.g. a 75-page paper) can take longer than most MCP
client timeouts. pdf_to_markdown handles this automatically:
mode="auto"(default): PDFs ≤ 15 pages return inline; larger PDFs start a background job and immediately return ajob_idwith an ETA — no client timeout.mode="async": always return ajob_id.mode="sync": always inline (previous behavior; may time out on large files).- Poll with
get_job_status(job_id)— it reports the current stage (classify → extracting → table_extraction → qc → done) and, when finished, themarkdown_pathplus a content preview. - The result is always written to
<out_dir>/extraction/full_text.md, so even if a client disconnects, nothing is lost. - Need a fast preview? Pass
page_range="0-9"to convert only the first pages.
The marker engine downloads ~2GB of models on first use — inside a 300s startup window. On slow networks this fails with an opaque timeout. Avoid it:
- Pre-download models first (no time limit, runs as a background job):
ask your agent to run
download_models, then pollget_job_status. - huggingface.co unreachable? Use a mirror — the Xet-incompatibility
workaround (
HF_HUB_DISABLE_XET=1) is applied automatically:export HF_ENDPOINT=https://hf-mirror.com - Using an HTTP proxy? Localhost must bypass it, or internal inference
health checks fail. The server enforces
NO_PROXY=localhost,127.0.0.1automatically at startup — but check your client config if you override env. Note: some setups excludehuggingface.cofrom the proxy viaNO_PROXY; remove that entry if you want HF downloads to go through the proxy. - All models cached? Go fully offline for reliable startups:
export HF_HUB_OFFLINE=1
check_environment reports per-model cache status (models_ready) and the
current network configuration, so your agent can diagnose this in one call.
Every pdf_to_markdown run finishes with a two-layer quality check whose
results are returned in qc_report (verdict, dimension scores, issues, fixes):
- Statistical gate — text completeness (chars/page), heading structure, formula integrity, table coverage. Catches gross failures.
- Content-aware audit — rules born from a real 75-page paper audit where every actual defect passed the statistical gate unnoticed:
| Rule | Detects | Severity | Auto-fix |
|---|---|---|---|
MD-101 |
Garbled chars (U+FFFD, private-use area) | critical | — |
MD-102 |
C0 control chars at in-cell word wraps (e.g. En\x02lightenment) |
critical | ✅ removed, words rejoined |
MD-103 |
Table with an all-empty header row (misread multi-column layout) | warn | — |
MD-104 |
Numeric column tearing — scientific notation split across cells (6 | 0 | × 10 | − 4, decimal point lost) |
critical | — |
MD-105 |
Table header fused with the first data row (header cells contain standalone numbers) | critical | — |
MD-106 |
Empty <span></span> placeholder cells |
info | ✅ removed |
MD-107 |
Flattened multi-row group headers (group labels glued onto repeated metric names in wide tables) | warn | — |
MD-108 |
Math-delimiter collision — escaped citation-link brackets ([\[…\]]() that MathJax/KaTeX misread as display-math openers, breaking formula rendering |
warn | ✅ de-escaped (link-scoped only) |
MD-109 |
Image-link integrity — references that do not resolve from the markdown's directory (e.g. bare filenames while files live under images/) |
warn | ✅ rewritten when the file exists under images/<name> |
MD-201 |
Body content loss — hyphenation-normalized token comparison against the PDF text layer (pymupdf, independent of the engine's layout analysis); figure text excluded | info/warn/critical by ratio | — |
MD-202 |
Figure-text omission — text embedded in vector figures never reaches the markdown body (expected; enable VLM enrichment to transcribe) | info | — |
Auto-fix policy: only deterministic, information-preserving fixes are
applied automatically (the sanitized markdown is written back to
full_text.md). Structural defects (MD-103/104/105) are located precisely
but never rewritten — automated guessing could corrupt values further.
Cross-channel repair (repair-or-report): with auto_repair=True
(default), structural defects get a repair attempt against the PDF text
layer (pymupdf word geometry — independent of the engine's layout analysis).
Each repair must pass a machine-checkable verification gate:
MD-104: recovered value minus decimal points must equal the joined fragments — digits are never altered, only the lost.restored.MD-105: the table is rebuilt from word geometry; the token multiset must be conserved — content is rearranged, never invented.MD-201: missing word runs are injected next to anchors present in the markdown; no token may exceed its PDF count (over-injection rolls back), and bulk deficits (>5%) are refused.
Gate passed → patched and re-audited (issue disappears). Gate failed → the
defect stays reported in qc_report.repairs with recovered candidates.
VLM arbitration (auto-activated): defects beyond geometric reach — merged
cells, multi-row group headers — escalate to the configured vision model
(setup_vlm). The broken table's source region is re-read from a hi-res
page render and replaced with an HTML <table> (Markdown tables cannot
express rowspan/colspan). NUMERIC GATE: the VLM may neither invent numbers
absent from the region's text layer nor drop numbers from the broken block.
Activation follows a tri-state model — configuring a VLM is itself the
opt-in: enable_table_enrich and enrich_figures default to 'auto',
which activates them whenever a VLM is configured and its stored policy
allows (setup_vlm policy='full' — the default — enables both; use
'tables_only' to keep figure descriptions off). Explicit 'on'/'off'
per call always overrides. Every response carries a features section
reporting exactly which capabilities ran and how to unlock the rest.
enrich_figures injects VLM descriptions under figure images so
figure-embedded content (MD-202) becomes retrievable by text-only RAG.
VLM calls consume API tokens.
Remaining escalation path:
- Cross-check the affected region with
extract_tables(pdfplumber — an independent extraction channel that bypasses layout analysis). - Re-run with VLM table enrichment (
setup_vlm, thenenable_table_enrich=True). - Any
criticalfinding escalates aPASSverdict toWARN, so agents know to inspectqc_report.audit_issuesbefore trusting the output.
VLM re-extracts complex tables and broken formulas from page images. Works with any provider whose model supports image input:
| Provider | Example model | API base |
|---|---|---|
| Alibaba Qwen | qwen-vl-max |
https://dashscope.aliyuncs.com/compatible-mode/v1 |
| Zhipu AI | glm-4v |
https://open.bigmodel.cn/api/paas/v4 |
| MiniMax | minimax-m3 |
https://api.minimaxi.com/v1 |
| Moonshot | moonshot-v1-vision |
https://api.moonshot.cn/v1 |
| OpenAI | gpt-4o |
https://api.openai.com/v1 |
| Ollama (local) | llama3.2-vision |
http://localhost:11434/v1 |
Set your key via environment variable (recommended — never typed into chat):
export PDF_CAPTURE_VLM_API_KEY=your_key_hereThen tell your agent: "Enable VLM with qwen-vl-max" — it validates vision capability before saving. Note: using VLM consumes your API tokens.
For the highest extraction quality on complex layouts (multi-column, InDesign PDFs):
pdf-capture-mcp setup-mineru # requires Python 3.11 on PATHModels (~2GB) auto-download from ModelScope on first extraction.
Image-only scans are detected automatically (is_scanned in pdf_info and
conversion responses): every page is force-OCRed, segment budgets triple,
and page windows that still fail OCR are reported in missing_segments
with explicit placeholders — loss is always visible, never silent.
Interrupted multi-hour jobs resume from finished segment checkpoints when
re-run with the same out_dir.
For very large scans (hundreds of pages), probe your machine's OCR throughput first, then submit the full document:
pdf_to_markdown(pdf_path="book.pdf", page_range="0-19") # ~20-page probe
Expect roughly 1.5 min/page on Apple Silicon for full-page OCR — a 900-page scan is an overnight job (the async job has no timeout and survives via checkpoints).
| Variable | Default | Description |
|---|---|---|
PDF_CAPTURE_ENGINE |
auto |
Default engine: marker / mineru / pymupdf / auto |
PDF_CAPTURE_VLM_API_KEY |
— | VLM API key (preferred over passing in chat) |
PDF_CAPTURE_CACHE_DIR |
~/.cache/pdf-capture-mcp |
Model & config cache (also stores job state) |
PDF_CAPTURE_MINERU_VENV |
<cache>/venv-mineru |
MinerU venv location |
PDF_CAPTURE_LOG_LEVEL |
INFO |
Logging level |
PDF_CAPTURE_SEGMENT_TIMEOUT_S |
1200 |
Per-segment extraction budget for oversized documents (scans get 3x automatically) |
MINERU_MODEL_SOURCE |
modelscope |
MinerU model source: modelscope / huggingface / local |
HF_ENDPOINT |
huggingface.co | HuggingFace mirror for model downloads (e.g. https://hf-mirror.com) |
HF_HUB_DISABLE_XET |
— | Set 1 when using a mirror (auto-set by download_models) |
HF_HUB_OFFLINE |
— | Set 1 after all models are cached for fully offline runs |
NO_PROXY |
— | Must include localhost,127.0.0.1 when a proxy is set (auto-enforced) |
git clone https://github.com/ChenHongYu2026/pdf-capture-mcp.git
cd pdf-capture-mcp
uv sync --extra dev
uv run pytest tests/ -v
uv run ruff check src/ tests/
uv run mypy src/pdf_capture_mcp/| Problem | Solution |
|---|---|
Operation not supported during install |
External/exFAT drive: export UV_LINK_MODE=copy then retry |
uvx: command not found |
Install uv: curl -LsSf https://astral.sh/uv/install.sh | sh |
| MCP server not appearing in tools | Restart your MCP client; check mcp.json syntax |
| MCP call times out on a large PDF | Expected with mode="sync" — use the default mode="auto" and poll get_job_status; the result is still written to out_dir |
First conversion fails with fast_layout/ocr_error server failed to become healthy |
Model download exceeded the 300s startup window — run download_models first |
| Downloads stall on huggingface.co | Set HF_ENDPOINT=https://hf-mirror.com (see Slow Networks section) |
| All health checks fail behind a proxy | Ensure NO_PROXY includes localhost,127.0.0.1 (auto-enforced at startup) |
| marker engine slow on first run | Downloads ~2GB models on first use; run download_models ahead of time |
| Python version too low | Requires 3.11+: uv python install 3.11 |
MIT — see LICENSE.
Third-party notices: MinerU (AGPL-3.0, invoked as a separate subprocess, not linked), marker (Apache-2.0), pdfplumber (MIT), Table Transformer (MIT), pymupdf4llm (Apache-2.0).
English | 中文
多阶段 PDF 捕获管线,以 MCP 服务器形式提供。 将 PDF 文档转换为高质量结构化 Markdown —— 支持公式识别、表格提取、版面清洁和内置质量门控。
- 三提取引擎
- pymupdf4llm(内置)—— 零配置、快速、始终可用
- marker(推荐)—— 复杂版面提取质量最高
- MinerU(可选)—— 多栏/InDesign 排版最佳,自动管理独立虚拟环境
- 14 个 MCP 工具 ——
pdf_to_markdown、get_job_status、download_models、extract_tables、classify_document、pdf_info、setup_vlm、check_environment、install_engine - 异步任务模式 —— 大型 PDF 在后台任务中转换(不再触发 MCP 客户端超时);模型可提前预下载,不受时间窗口限制
- 可选 VLM 增强 —— 接入任何具备视觉能力的模型(通义千问 Qwen-VL、智谱 GLM-4V、MiniMax、月之暗面 Moonshot、OpenAI、本地 Ollama 等),提升表格/公式提取质量。无需额外依赖,基础安装即可使用。
- 质量门控 —— 多维度 QC 评估(文本完整度、标题结构、公式完好率、表格覆盖率),另含内容感知审计规则,捕获统计指标无法发现的缺陷(控制字符、数值列撕裂、表头融合、内容丢失)
- 渐进式配置 —— 零配置即可工作;按需增强 marker/VLM
- 隐私优先 —— API Key 以
chmod 600权限本地存储,绝不在响应中回显
# 基础包(pymupdf 引擎 + VLM 支持,约 80MB)—— 安装即可用
pip install pdf-capture-mcp
# 含 marker 引擎(推荐复杂 PDF,包含 PyTorch,约 2.5GB)
pip install "pdf-capture-mcp[marker]"
# 完整安装(marker + TATR 表格检测 + DePlot 图表提取,约 3GB)
pip install "pdf-capture-mcp[all]"Qoder / Claude Desktop / Cursor 用户,在 mcp.json 中添加:
{
"mcpServers": {
"pdf-capture": {
"command": "uvx",
"args": ["pdf-capture-mcp"]
}
}
}含 marker 引擎(推荐):
{
"mcpServers": {
"pdf-capture": {
"command": "uvx",
"args": ["pdf-capture-mcp[marker]"]
}
}
}如果已通过 pip 安装:
{
"mcpServers": {
"pdf-capture": {
"command": "pdf-capture-mcp"
}
}
}直接对你的 AI 助手说:
“把 ~/Downloads/论文.pdf 转成 Markdown” “提取这份报告里的所有表格” “这个 PDF 是扫描件吗?”
服务器使用内置 pymupdf 引擎即可立即工作。对于复杂版面,助手可按需安装 marker 引擎。
| 工具 | 说明 |
|---|---|
pdf_to_markdown |
完整管线:提取 → 清洁 → QC → 修复 → 自描述知识包(大文件自动异步) |
export_to_obsidian |
将知识包作为整体拷入 Obsidian vault(幂等) |
setup_embedding |
配置 OpenAI 兼容嵌入端点(OpenAI / MiniMax / BGE / Ollama) |
build_vector_index |
将知识包索引进嵌入式 Qdrant(增量、内容寻址) |
search_corpus |
带元数据过滤的语义检索 —— 该工具即 RAG API |
batch_convert |
目录级批量转换(异步任务):doc_id 去重,可选入 vault + 建索引 |
get_job_status |
轮询后台任务(大文件转换 / 模型下载) |
download_models |
预下载 marker 模型(慢速网络强烈推荐) |
extract_tables |
表格提取(pdfplumber 规则 + 可选 TATR 深度学习) |
classify_document |
文档类型检测(学术论文、咨询报告等) |
pdf_info |
快速元数据:页数、文本层、扫描件检测 |
setup_vlm |
配置可选的 VLM 增强(支持任何具备视觉能力的供应商) |
check_environment |
校验引擎、依赖、模型缓存与网络配置 |
install_engine |
代用户安装 marker/ml 引擎 |
每次转换现在产出自描述知识包——任何 LLM Agent 读一遍 README 即可完全理解, 整个文件夹可直接拖入 Obsidian vault:
<out_dir>/<slug>/
├── <slug>.md 主文档,与目录同名([[slug]] 直达),出厂内置 frontmatter
├── README.md 入口地图:摘要、文件表、分块 schema
├── images/ 提取图片(相对引用)
├── tables/ p<页码>_table_<n>.csv —— 独立提取通道,可交叉验证
└── data/ chunks.jsonl + metadata.json + qc_report.json
关键设计(来自两轮设计审计):内容寻址身份(doc_id/chunk_id 均为内容哈希,
重转幂等覆盖、增量重嵌只触及变更块);层级分块(标题路径元数据、表格/代码
独立成块、大表分片重复表头、页码单调锚定);MD-110 跨页表合并(PDF 几何
三证门槛,同列数独立表绝不误合)。默认输出根目录 $PDF_CAPTURE_OUTPUT_ROOT
或 ~/Documents/pdf-capture;package=False 可回退旧布局。
转换大型文档(如 75 页论文)的耗时往往超过 MCP 客户端超时限制。
pdf_to_markdown 会自动处理:
mode="auto"(默认):≤ 15 页直接返回结果;更大的 PDF 自动转为后台任务,立即返回job_id和预估耗时 —— 不再触发客户端超时。mode="async":总是返回job_id。mode="sync":保持旧版同步行为(大文件可能超时)。- 用
get_job_status(job_id)轮询进度,可看到当前阶段(classify → extracting → table_extraction → qc → done);完成后返回markdown_path和内容预览。 - 结果始终写入
<out_dir>/extraction/full_text.md,即使客户端断开也不丢失。 - 需要快速预览?传
page_range="0-9"只转换前几页。
marker 引擎首次使用时会在 300 秒启动窗口内下载约 2GB 模型 —— 慢速网络下必然超时失败。规避方法:
- 先预下载模型(无时间限制,后台任务运行):让助手调用
download_models,再用get_job_status轮询。 - 连不上 huggingface.co? 使用镜像站(Xet 协议兼容问题会自动处理,即自动设置
HF_HUB_DISABLE_XET=1):export HF_ENDPOINT=https://hf-mirror.com - 使用 HTTP 代理? localhost 必须绕过代理,否则内部推理服务的健康检查会被代理劫持而失败。服务启动时会自动确保
NO_PROXY包含localhost,127.0.0.1。另注意:若你的环境把huggingface.co加入了NO_PROXY(即 HF 不走代理),想让 HF 下载走代理时需移除该条目。 - 模型全部缓存完成后,建议开启完全离线模式,启动更稳定:
export HF_HUB_OFFLINE=1
check_environment 会逐一报告模型缓存状态(models_ready)和当前网络配置,助手一次调用即可完成诊断。
每次 pdf_to_markdown 运行结束时都会执行双层质量检查,结果在 qc_report
中返回(结论、维度分数、问题清单、已修复项):
- 统计门控 —— 文本完整度(字符/页)、标题结构、公式完好率、表格覆盖率,捕获粗粒度失败。
- 内容感知审计 —— 源自一次真实的 75 页论文审计:当时所有实际缺陷都骗过了统计门控:
| 规则 | 检测内容 | 严重度 | 自动修复 |
|---|---|---|---|
MD-101 |
乱码字符(U+FFFD、私有区字符) | critical | — |
MD-102 |
单元格内换行处的 C0 控制字符(如 En\x02lightenment) |
critical | ✅ 移除并拼回断词 |
MD-103 |
全空表头行(多栏版式被误识为表格) | warn | — |
MD-104 |
数值列撕裂 —— 科学计数法被拆进多个单元格(6 | 0 | × 10 | − 4,小数点丢失) |
critical | — |
MD-105 |
表头与首行数据融合(表头单元格含独立数字) | critical | — |
MD-106 |
空 <span></span> 占位单元格 |
info | ✅ 移除 |
MD-107 |
多行分组表头被压平(宽表中分组名错接到重复的指标名上) | warn | — |
MD-108 |
数学定界符冲突 —— 引用锚点链接中的转义方括号([\[…\]]()被 MathJax/KaTeX 误读为块级数学开始符,导致公式渲染全面损坏 |
warn | ✅ 去转义(仅链接形态) |
MD-109 |
图片链接完整性 —— 引用路径相对 markdown 所在目录无法解析(如文件存在 images/ 子目录但引用是裸文件名) |
warn | ✅ 文件存在于 images/<name> 时自动改写 |
MD-201 |
正文内容丢失 —— 与 PDF 文本层(pymupdf,独立于引擎版面分析的通道)做去连字符归一化的 token 比对,图内文字已排除 | 按比例 info/warn/critical | — |
MD-202 |
图内文字省略 —— 矢量图内嵌入的文字不会进入正文(预期行为;需要时可开 VLM 增强转录) | info | — |
自动修复策略:仅自动应用确定性、信息无损的修复(修复后的 markdown 会回写
full_text.md)。
跨通道修复(repair-or-report):auto_repair=True(默认)时,结构性缺陷会
基于 PDF 文本层(pymupdf 词几何 —— 独立于引擎版面分析的通道)尝试修复,
每项修复必须通过机器可验证的门槛:
MD-104:恢复值去掉小数点后必须与碎片拼接完全相等 —— 数字永不改变,只找回丢失的.;MD-105:整表从词几何重建,token 多重集必须守恒 —— 只重排、绝不发明内容;MD-201:缺失词段注入到 markdown 中已存在的锚点旁;任何 token 不得超出 PDF 计数(超注入全量回滚),大体量缺失(>5%)拒绝自动注入。
门槛通过 → 打补丁并重新审计(问题从清单消失);门槛失败 → 缺陷保留在
qc_report.repairs 中并附恢复候选值。
VLM 仲裁(自动激活):几何修复触及不到的缺陷 —— 合并单元格、多行分组表头 ——
升级到已配置的视觉模型(setup_vlm):损坏表格的源区域以高清渲染重读,
替换为 HTML <table>(Markdown 表格无法表达 rowspan/colspan)。数值守恒门槛:
VLM 不得发明区域文本层没有的数字,也不得丢失损坏块里的数字。
激活遵循三态模型 —— 配置 VLM 本身就是授权:enable_table_enrich 与
enrich_figures 默认 'auto',VLM 已配置且其存储的 policy 允许时自动启用
(setup_vlm policy='full' 为默认,两者全开;'tables_only' 仅开表格修复)。
每次调用可用 'on'/'off' 显式覆盖。响应中的 features 段逐项报告哪些
能力实际运行、未运行的原因及解锁方式。enrich_figures 在图片下注入 VLM
描述,让图内信息(MD-202 缺口)可被纯文本 RAG 检索。VLM 调用消耗 API token。
后续升级路径:
- 用
extract_tables(pdfplumber —— 绕过版面分析的独立提取通道)交叉校验受影响区域; - 开启 VLM 表格增强重新转换(
setup_vlm+enable_table_enrich=True); - 任何
critical发现都会把PASS升级为WARN,Agent 应先检查qc_report.audit_issues再信任输出。
VLM 会从页面图像中重新提取复杂表格和损坏的公式。 支持任何模型具备图片输入能力的供应商:
| 供应商 | 示例模型 | API 地址 |
|---|---|---|
| 阿里通义千问 | qwen-vl-max |
https://dashscope.aliyuncs.com/compatible-mode/v1 |
| 智谱 AI | glm-4v |
https://open.bigmodel.cn/api/paas/v4 |
| MiniMax | minimax-m3 |
https://api.minimaxi.com/v1 |
| 月之暗面 | moonshot-v1-vision |
https://api.moonshot.cn/v1 |
| OpenAI | gpt-4o |
https://api.openai.com/v1 |
| Ollama(本地) | llama3.2-vision |
http://localhost:11434/v1 |
推荐通过环境变量设置 Key(避免在对话中输入):
export PDF_CAPTURE_VLM_API_KEY=你的密钥然后告诉助手:"启用 VLM,用 qwen-vl-max" —— 系统会先验证模型的视觉能力再保存配置。 注意:使用 VLM 功能会消耗你的 Token。
针对复杂版面(多栏、InDesign 排版 PDF)获得最高提取质量:
pdf-capture-mcp setup-mineru # 需要 PATH 中有 Python 3.11首次提取时会从 ModelScope 自动下载模型(约 2GB)。
纯图片扫描件会被自动识别(pdf_info 与转换响应中的 is_scanned):全页强制
OCR、段预算自动 3 倍;OCR 仍失败的页窗记入 missing_segments 并在正文插入
显式占位标记——内容丢失永远可见、绝不静默。多小时任务中断后,用同一
out_dir 重跑即可从已完成段的检查点续传。
超大扫描件(数百页)建议先实测本机 OCR 吞吐,再提交全量:
pdf_to_markdown(pdf_path="book.pdf", page_range="0-19") # 约 20 页预检
Apple Silicon 上全页 OCR 约 1.5 分钟/页——900 页扫描件是一个通宵级任务 (异步 job 无超时,检查点保证可中断恢复)。
| 变量 | 默认值 | 说明 |
|---|---|---|
PDF_CAPTURE_ENGINE |
auto |
默认引擎:marker / mineru / pymupdf / auto |
PDF_CAPTURE_VLM_API_KEY |
— | VLM API Key(推荐方式,避免对话中传递) |
PDF_CAPTURE_CACHE_DIR |
~/.cache/pdf-capture-mcp |
模型与配置缓存目录(同时存储任务状态) |
PDF_CAPTURE_MINERU_VENV |
<cache>/venv-mineru |
MinerU 虚拟环境位置 |
PDF_CAPTURE_LOG_LEVEL |
INFO |
日志级别 |
PDF_CAPTURE_SEGMENT_TIMEOUT_S |
1200 |
超大文档分段提取的单段预算(扫描件自动 3 倍) |
MINERU_MODEL_SOURCE |
modelscope |
MinerU 模型源:modelscope / huggingface / local |
HF_ENDPOINT |
huggingface.co | HuggingFace 镜像站(如 https://hf-mirror.com) |
HF_HUB_DISABLE_XET |
— | 使用镜像站时设为 1(download_models 会自动设置) |
HF_HUB_OFFLINE |
— | 模型全部缓存后设为 1,完全离线运行 |
NO_PROXY |
— | 设置代理时必须包含 localhost,127.0.0.1(启动时自动保障) |
git clone https://github.com/ChenHongYu2026/pdf-capture-mcp.git
cd pdf-capture-mcp
uv sync --extra dev
uv run pytest tests/ -v
uv run ruff check src/ tests/
uv run mypy src/pdf_capture_mcp/| 问题 | 解决方案 |
|---|---|
安装时报 Operation not supported |
外置/exFAT 磁盘:export UV_LINK_MODE=copy 后重试 |
uvx: command not found |
安装 uv:curl -LsSf https://astral.sh/uv/install.sh | sh |
| MCP 服务器未出现在工具列表 | 重启 MCP 客户端;检查 mcp.json 格式 |
| 大 PDF 转换时 MCP 调用超时 | mode="sync" 下属预期行为 —— 使用默认 mode="auto" 并轮询 get_job_status;结果仍会写入 out_dir |
首次转换报 fast_layout/ocr_error server failed to become healthy |
模型下载超出 300 秒启动窗口 —— 先运行 download_models |
| huggingface.co 下载卡住 | 设置 HF_ENDPOINT=https://hf-mirror.com(见“慢速网络”一节) |
| 挂代理后所有健康检查失败 | 确保 NO_PROXY 包含 localhost,127.0.0.1(启动时自动保障) |
| marker 引擎首次运行慢 | 首次使用需下载约 2GB 模型,建议提前运行 download_models |
| Python 版本过低 | 需要 3.11+:uv python install 3.11 |
MIT —— 详见 LICENSE。
第三方声明:MinerU(AGPL-3.0,通过独立子进程调用,未链接)、 marker(Apache-2.0)、pdfplumber(MIT)、Table Transformer(MIT)、pymupdf4llm(Apache-2.0)。