A synthetic document forge for training OCR and document-understanding models — Arabic first.
OCRSmith generates whole documents, not cropped text lines: multi-column pages with titles, tables, figures, forms and running headers, degraded to look like something a scanner or a phone actually produced — and it emits the ground truth for every objective those pages can supervise.
One rendered page gives you, from the same pass:
| Objective | What you get |
|---|---|
| Recognition | Line and word crops with logical-order text |
| Detection | Word, line and region boxes; polygons where the page is warped |
| Layout analysis | Typed regions (title, table, figure, key_value, …) in reading order |
| Document → markup | The page serialised back to Markdown and HTML |
| Table structure | Cell grid as HTML and OTSL |
Everything is reproducible from a seed, streamed rather than materialised, and validated before it reaches the dataset.
Most synthetic OCR data is a line of text on a noisy background. That teaches a model to read a crop; it does not teach it to read a page. And most generators built for Latin script get Arabic subtly wrong in ways that are invisible until training plateaus:
- Logical vs visual order. The label a model must predict is the logical string; the
pixels are in visual order. OCRSmith keeps them apart explicitly and never confuses
them — see
ocrsmith/text/shaping.py. - Missing glyphs. A font that cannot draw a character renders a blank or a tofu box
while the label still claims it. OCRSmith checks the font's
cmapbefore choosing it. - Boxes that do not follow the pixels. A rotation that moves ink but leaves the annotation behind produces a dataset that looks fine and trains a detector to be systematically wrong. Here every geometric degradation maps the annotation through the same transform.
- Silent truncation. A wrapper that drops the tail of a paragraph produces an image whose label claims text that was never drawn. OCRSmith's wrapping is lossless, and text that does not fit is reported, not discarded.
pip install -e ".[data,dev]"data adds pandas/pyarrow/datasets for tabular and Hugging Face corpora and for
Parquet output; the core install stays light.
Check the machine can produce correct Arabic:
ocrsmith doctorPillow built with Raqm delegates shaping to HarfBuzz; without it OCRSmith falls back to
arabic-reshaper + python-bidi. Both paths produce the same labels.
ocrsmith preview --count 3 --boxes --output outputs/previewocrsmith generate --num-samples 10000 --workers 8 --format webdataset -o data/trainocrsmith stats data/train --markdown data/train/DATASET_CARD.md
ocrsmith validate data/trainFrom Python:
from ocrsmith import load_config, run_generation
config = load_config("configs/darija_scan.yaml")
result = run_generation(config)
print(result.to_dict())Or drive the generator directly and keep the samples in memory:
from ocrsmith import SampleFactory, load_config
factory = SampleFactory(load_config())
for sample in factory.create(index=0):
sample.image.save(f"{sample.id}.png")
print(sample.page.to_markdown())
for word in sample.page.iter_words():
print(word.text, word.bbox.as_tuple())data/train/
├── images/ # one PNG/JPEG per page
├── annotations-00000.jsonl # one record per page
└── .shard-00000.done # completion marker, so a rerun resumes
Each record carries the full annotation tree plus both markup serialisations:
--format selects the writer: jsonl, parquet, webdataset, coco, paddleocr,
chat. See docs/formats.md.
config/ one validated GenerationConfig describing a whole corpus
↓
text/ script detection · normalisation policy · bidi + shaping · glyph coverage
core/documents/ content model → typography → flow layout → pages
core/rendering/ text drawn word by word, emitting word boxes in the same pass
core/degradations/ capture conditions; geometric ones move the annotation too
↓
domain/ Page → Region → Line → Word (+ Table), immutable and serialisable
↓
quality/ validate the page, then count what the corpus contains
datasets/ six writers behind one SampleSink protocol
evaluation/ CER · WER · NED · table similarity · detection P/R/F1
Longer version: docs/architecture.md.
A corpus is defined by its distributions, so every per-sample choice is a range or a weight map rather than a fixed value:
page:
papers: { a4: 4.0, a5: 1.0, letter: 1.0 }
dpi_range: [110, 200]
columns: { 1: 3.0, 2: 1.0 }
templates:
weights: { article: 3.0, report: 2.0, newspaper: 1.5, letter: 1.0, form: 1.0, invoice: 1.0 }
degradations:
presets: { clean: 1.0, scan: 4.0, photo: 3.0, fax: 0.5, archive: 1.0 }Override anything from the command line:
ocrsmith generate --set page.columns='{"2":1}' --set degradations.presets='{"photo":1}'Full reference: docs/configuration.md. Ready-made corpora: examples/configs.
Every sample's seed is derived, never drawn:
sample_seed(index) = f(config.seed, index)So sample 8 412 of a ten-million-page run can be regenerated on its own, months later, without replaying anything — which is what makes a synthetic dataset debuggable.
ocrsmith preview --set run.start_index=8412 --count 1 --boxes- Generation is a generator end to end; a ten-million-page run holds one page in memory.
- Work is sharded, and each worker writes its own shard rather than shipping images back through a pickle queue.
- A completed shard is marked done, so an interrupted job resumes instead of restarting.
- Font metrics are cached per
(font, size, shaper), which is most of the cost of laying out a page.
ocrsmith generate -n 1000000 --workers 32 --format webdataset -o /mnt/data/ocrocrsmith generate -n 2000 -o data/bench --set seed=777
# … run your model, write {sample_id: prediction} to predictions.json …
ocrsmith evaluate data/bench predictions.json --ignore-diacriticsfrom ocrsmith.evaluation import evaluate, load_references
report = evaluate(load_references("data/bench", target="markdown"), predictions)
print(report.to_markdown())
print(report.worst(10)) # where to start debuggingSee CONTRIBUTING.md. Tests are the specification: run pytest from a
fresh clone, no install required.
MIT — see LICENSE. Bundled fonts are distributed under their own licences (SIL OFL for the Noto, Amiri, Mada, Fustat, Kufam, Mirza and Vazirmatn families).
Built for the AtlasIA effort to bring Moroccan Darija into open models.
{ "id": "00000042_01", "image_path": "images/00000042_01.png", "text": "تقرير سنوي\nيحتوي هذا التقرير على جداول وأرقام…", "markdown": "# تقرير سنوي\n\nيحتوي هذا التقرير…", "html": "<h1>تقرير سنوي</h1>\n<p>يحتوي هذا التقرير…</p>", "page": { "width": 1240, "height": 1754, "direction": "rtl", "regions": [ { "type": "title", "bbox": [...], "reading_order": 0, "lines": [{ "text": "تقرير سنوي", "bbox": [...], "direction": "rtl", "words": [{ "text": "تقرير", "bbox": [...] }] }] }, { "type": "table", "bbox": [...], "table": { "rows": 4, "cols": 3, "cells": [...] } } ] }, "provenance": { "seed": 918273645, "template": "report", "font_path": ".../Amiri-Regular.ttf", "background": "paper", "degradations": [{ "name": "PerspectiveWarp", "magnitude": 0.03 }], "extra": { "index": 42, "page": 1, "preset": "photo", "dpi": 150, "columns": 2 } } }