English | ็ฎไฝไธญๆ | Deutsch
Any format in, any language out โ layout, typesetting, and in-image text intact.
Translayer is an AI-native document localization middle layer. Existing translation tools treat document translation as a pure NLP problem, resulting in broken slide layouts, ignored text inside graphics, and formatting collapse. Translayer solves this by parsing documents into a unified intermediate representation (DocumentIR), enriching them with layout and OCR metadata, localizing them through human-in-the-loop and VLM-based workflows, and rendering them back losslessly โ including redrawing/editing in-image text.
- Local Tesseract OCR now upscales small screenshots, filters noise, merges wrapped lines, and erases text line by line.
- Dense text-only screenshots can be rebuilt as clean target-language text canvases that expand vertically when needed.
- Complete translations are preserved; layout fitting no longer creates incomplete words or silently drops meaning.
- PPTX text is fitted inside foreground images and visual panel boundaries.
- Existing bilingual target-language content is preserved while a matching, nearby source-language duplicate is removed.
- Provider URLs, keys, models, and engine selections persist in the current browser and can be cleared from the trilingual UI.
This section provides a machine-readable summary of the codebase structure, interfaces, and architecture to help LLMs and AI agents quickly index and modify this repository.
.
โโโ src/translayer/
โ โโโ api/ # Web server (FastAPI, Uvicorn, SQLite/JSON Job storage)
โ โโโ engines/ # Pluggable backends
โ โ โโโ image/ # Whole-image localization (Gemini, Cost Guard)
โ โ โโโ ocr/ # OCR Engines (Tesseract, Paddle OCR, Cloud Vision)
โ โ โโโ translation/ # Text translators (DeepL, Gemini, mock)
โ โโโ enrich/ # Enrichment pipeline (Tesseract screening, image text selection)
โ โโโ ir/ # Intermediate Representation definitions and JSON Schemas
โ โโโ localize/ # Quality Gates, explicit text mapping, and pipelines
โ โโโ parsers/ # Document parsing (PPTX, DOCX, HTML)
โ โโโ renderers/ # Lossless document generation and write-back
โ โโโ web/ # HTML5/JS Trilingual single-page App (default: English)
โโโ tests/ # Integration and unit tests
โโโ Dockerfile # Production environment with LibreOffice, Poppler, Tesseract
โโโ pyproject.toml # Project dependencies and entrypoints
- Primary Entrypoints:
translayer.cli:app(Typer CLI),translayer.api.app:app(FastAPI Server). - DocumentIR Specification: Standardized intermediate schema representing slide elements, layout constraints (EMUs), resources (fonts, images), translatable blocks, and OCR region mapping.
- Core Pipeline Contract:
[Input Doc] -> Parse() -> DocumentIR -> Enrich() -> Localize() -> Render() -> [Output Doc]
graph TD
A[input.pptx / docx / html] -->|Parse| B(DocumentIR)
B -->|Local Image Screening| C{Image Selector<br>Tesseract TSV Probe}
C -->|Decorative / Duplicate / Logo| D[Ignore Route]
C -->|Simple Labels| E[Inpaint + Text Box Route]
C -->|Text-Heavy Graphic| F[Whole-Image Gemini Route]
F -->|Pre-Gen OCR QA Gate| G{Has Text?}
G -->|No Text| H[Fallback / Local OCR]
G -->|Has Text| I[Translate OCR Regions first<br>Build Explicit Text Map]
I -->|Cost Guard Check| J{Within Budget?}
J -->|No| K[Manual Review Needed]
J -->|Yes| L[Gemini Interactions API<br>Native Image Edit]
L --> M[Post-Gen OCR QA Gate]
M -->|Source Leftovers / Missing Target| N[Fail Closed & Invalidate Cache<br>Return to Review/Retry]
M -->|Verified Pass| O[Merge Localized Image to IR]
B -->|Localize Text Blocks| P[Multi-Engine Translation<br>Context / Term / Constraints]
P --> Q(Localized DocumentIR)
O --> Q
E --> Q
Q -->|Lossless Write-back / Render| R[output.pptx / docx / html]
- Parse: Formats are converted into
DocumentIRby reading shapes, paragraph trees, SmartArt (ppt/diagrams/dataN.xml), text runs, and image assets. - Enrich: Enhances the document model with semantic roles, extracts terms, and runs Local Image Screening to filter decorative elements, icons, logos, and exact duplicates.
- Localize: Translates text blocks with context constraints, while routing qualified images to the specialized Gemini native image editing sub-pipeline.
- Render: Swaps in localized images and writes back translated texts using precise coordinate and font constraints, guaranteeing byte-for-byte preservation of unmodeled elements.
- ๐ Multilingual & Font-Aware OCR: Full support for English, Chinese, and German as source/target languages. Automatically provisions language-specific OCR engines (e.g. Tesseract langpacks) and selects proper rendering fonts.
- โก Human-in-the-Loop Image Screening: Uses local, zero-API Tesseract TSV probes (
--psm 11) to analyze images for text density, line count, and OCR confidence. Separates text-heavy graphics from simple labels, lets users manually select images for AI translation in the review UI, and filters out system icons. - ๐ก๏ธ Pre- & Post-Generation OCR Quality Gates:
- Pre-Gen: Runs text translation first to create an explicit
source -> targetmapping list. If OCR detects no text or mapping fails, the engine fails closed to prevent API waste. - Post-Gen: Re-OCRs the Gemini-generated image. If it detects un-erased source text (residual leftovers) or missing target translations, it invalidates the cache and returns the job to manual review.
- Pre-Gen: Runs text translation first to create an explicit
- ๐ฐ Cost Guard System: Configurable per-job cost limits, estimated provider costs, and a budget gate to prevent runaway API fees.
- ๐ฅ๏ธ Trilingual Web Interface: Single-page, modern layout (defaults to English, supports Chinese and German) featuring document upload, advanced settings, cost estimation, image review panels, live progress monitoring, and output download.
- ๐ Measurable Long-Running Jobs: Text progress reports slides and text blocks; approved image processing reports image counts and the current OCR, translation, redraw, generation, reuse, or validation stage.
- ๐ Reusable Provider Credentials: Users can configure any OpenAI-compatible translation endpoint, DeepL, and optional Gemini image editing directly in the web UI. Connection settings persist in browser-local storage, can be cleared from the UI, and are never returned by public job APIs.
- ๐ Moonshot Kimi Compatibility: Kimi K2.5/K2.6 requests use the model-compatible non-thinking payload automatically when the API base URL is a Moonshot endpoint.
- ๐ต Visible Image Cost Forecasts: The UI shows the configured per-image planning rate, projected paid calls, estimated total, and a hard approval budget before Gemini image editing starts.
- ๐ณ Production Docker Container: Fully-packaged image featuring LibreOffice, Poppler, Tesseract (with eng, chi_sim, deu packs), Python 3.11, and multilingual fonts.
Every document is represented by a standardized JSON object. Below is a simplified schema block for AI integration reference:
Windows users can download the Translayer-v0.2.3-windows-x64.zip asset from
the GitHub Release, unzip it, and double-click Translayer.exe. The executable
starts a local Translayer server and opens the browser UI automatically; it does
not compile the project or run tests on the user's machine.
The portable package includes Tesseract OCR, English/German/Simplified Chinese
OCR language data, and Poppler's pdftoppm.exe. For full PPTX slide previews,
install LibreOffice or make soffice.exe available on PATH.
Translayer requires a few system packages for document and image processing:
- LibreOffice: For document format conversions (PDF/PPTX).
- Poppler-utils: For
pdftoppmrasterization checks. - Tesseract OCR: For local screening and verification.
On macOS:
brew install libreoffice poppler tesseract tesseract-langOn Ubuntu/Debian:
sudo apt-get update && sudo apt-get install -y \
libreoffice \
poppler-utils \
tesseract-ocr \
tesseract-ocr-eng \
tesseract-ocr-chi-sim \
tesseract-ocr-deuWe recommend using uv for fast, robust virtual environment and dependency management:
# Create and activate environment
uv venv --python 3.11
source .venv/bin/activate
# Install package in editable development mode
uv pip install -e ".[dev]"The quickest way to run a production-ready server without installing system dependencies is via Docker:
# Build the production image
docker build -t translayer:latest .
# Run the trilingual API & Web Server
docker run -d \
-p 8000:8000 \
-e GEMINI_API_KEY="your_api_key_here" \
--name translayer-server \
translayer:latestTranslayer is built around a comprehensive CLI suite powered by typer.
Translate any .pptx, .docx, or .html document.
translayer translate input.pptx -o output.pptx \
--from en \
--to zh \
--engine openai \
--api-url http://llm.internal:8000/v1 \
--api-key optional-local-key \
--model my-local-model \
--ocr-engine tesseract \
--glossary my_terms.csvOptions:
--engine openai: Use any OpenAI-compatible Chat Completions API. Set--api-url,--model, and, when the server requires authentication,--api-key.--engine deepl --api-key ...: Use DeepL API Free or Pro; the endpoint is selected automatically from the key.--api-urlcan override the DeepL base URL.--no-images: Skip in-image text translation entirely.--no-image-screening: Force API translation for all images without screening (not recommended, increases cost).
Directly translate text inside a single raster image while preserving its visual design and background using Gemini's native image-editing engine, protected by dual OCR gates.
translayer translate-image input.png -o output_localized.png \
--from en \
--to de \
--allow-paid-api \
--max-cost-usd 0.10Inspect a document locally, evaluate images, estimate costs, and check against a budget without making any paid API calls:
translayer plan-images input.pptx -o cost_plan.json \
--from en \
--targets zh,de \
--budget-usd 1.50translayer serve --host 127.0.0.1 --port 8000The job form accepts credentials for OpenAI-compatible translation, DeepL, and optional Gemini image-text localization. Connection settings are remembered in browser-local storage and can be cleared from the UI. After submission, secrets remain in the in-memory job and are never returned by public job APIs. A Gemini API key is needed only when the approved image plan contains paid whole-image edits. The UI shows the configured per-image planning estimate and the projected total before approval; actual provider billing may differ from this safety estimate.
For Moonshot Kimi, select OpenAI-compatible API and use:
API base URL: https://api.moonshot.cn/v1
Model: kimi-k2.6
API key: your Moonshot API key
During text localization, the progress panel reports slides and text blocks.
After the image plan is approved, it reports completed images and the current
OCR, translation, inpainting, redrawing, generation, validation, or reuse stage.
The same data is available from GET /jobs/{job_id} under progress.text and
progress.images.
The Offline demo engine is only a workflow test: it prefixes source strings with a target-language marker and does not produce a semantic translation.
Translayer uses a strict, type-hinted registry model. You can extend any stage by registering custom adapters.
from translayer.plugins import registry
# Register a custom Translation Engine
@registry.register("translation", "my_custom_translator")
class CustomTranslationEngine:
name = "my_custom_translator"
def translate(self, texts: list[str], src: str, tgt: str, context: str | None = None) -> list[str]:
# Implement your custom LLM/translation logic here
return [f"[Custom Target] {t}" for t in texts]
# Register a custom Inpaint/Image Eraser
@registry.register("inpaint", "my_custom_inpaint")
class CustomInpaintEngine:
name = "my_custom_inpaint"
def inpaint(self, image_path: str, mask_path: str, out_path: str) -> None:
# Implement custom canvas replacement/erasing
passList all loaded and available plugins via the CLI:
translayer plugins- Parsing, initial image screening, and final rendering currently show named stages without item-level percentages.
- OCR accuracy varies with resolution, font style, contrast, and image complexity. Small or dense Chinese text can be misread or remain in a localized image, so image results should be reviewed before export.
- Local region redraw is best suited to a few clear text areas. Complex infographics generally need whole-image localization and still require human review.
- Format Interoperability:
- Lossless PPTX Parser/Renderer with full SmartArt namespace preservation.
- Modular HTML/DOCX Parsers.
- Support for native XLIFF imports and exports.
- Image Localization Engine:
- Zero-API screening & routing.
- Dual-Gate OCR Pre/Post QA validation.
- Gemini interactions API integration for design-aware in-painting.
- Fine-tuned local SDXL/Inpaint models to operate entirely offline.
- Human-in-the-Loop UX:
- Interactive web interface with visual side-by-side comparison.
- Inline manual override for translations and image selections.
Translayer is open-source software licensed under the Apache License 2.0.
{ "schema_version": "0.1.0", "meta": { "source_lang": "en", "target_lang": "zh", "doc_type": "pptx", "title": "Strategy Deck" }, "resources": { "images": [ { "id": "img-slide3-0", "media_type": "image/png", "data_ref": "assets/img-slide3-0.png", "width": 1280, "height": 720, "text_regions": [ { "id": "reg-0", "bbox": [100, 200, 300, 250], "source_text": "Executive Summary", "target_text": "ๆง่กๆ่ฆ", "translatable": true, "background_kind": "solid" } ], "localization_validation": { "status": "passed", // "passed", "failed", "not_run" "reason": null, // "residual_source_text", "missing_target_text", etc. "expected_texts": ["Executive Summary"], "expected_translations": ["ๆง่กๆ่ฆ"] } } ] }, "blocks": [ { "id": "s3-sh5-p1", "type": "shape_text", "source_text": "Core Metrics and Goals", "target_text": "ๆ ธๅฟๆๆ ไธ็ฎๆ ", "translatable": true, "constraints": { "max_chars": 50, "can_shrink_font": true, "min_font_size": 9 }, "source_ref": { "kind": "shape_text", "slide_index": 3, "shape_id": 5, "paragraph_index": 1 } } ] }