Skip to content

Repository files navigation

OpenRead

OpenRead is an open-source, mobile-first reading layer for physical picture books. A caregiver takes one photo of a printed page, and OpenRead turns it into a layout-aware, structured read-aloud experience.

Live reader | Project site | Source

Try OpenRead

Open the live reader on a phone, allow camera access, and take a photo of one picture-book page. The public experience is designed around a phone camera and does not require an account.

The project site explains the initiative, its design principles, and the family reading problem it addresses.

Why OpenRead

Picture books are not ordinary OCR documents. Text may appear in speech bubbles, captions, curved regions, or several disconnected areas. An illustration may establish the speaker, complete a joke, or supply context that must be understood before the words make sense.

OCR can extract text. OpenRead compiles a visual page into a reading plan. Its Story Compiler uses Gemma 4 to reconcile layout, visible text, illustrations, and likely reading order before speech is generated.

The intermediate plan keeps recognized source text separate from optional illustration narration, keeps caregiver cues out of child-facing speech, and gives the frontend and voice engine a validated contract instead of an opaque paragraph.

OpenRead is built for a narrow, common family moment: a child is ready, a physical page is present, and voice, vision, language, or confidence gets in the way. It is web-first, caregiver-centered, self-hostable, and not tied to a proprietary book catalog, account, or reading toy.

Two Reading Modes

Read Page

Read Page is the default mode. The caregiver photographs one page and OpenRead:

  • reconstructs a natural reading order from the page layout;
  • preserves visible source text where possible;
  • adds brief illustration narration only when it contributes useful context;
  • returns ordered story beats, a child-facing script, TTS segments, and optional caregiver cues;
  • excludes caregiver cues from speech synthesis;
  • synthesizes the plan with Kokoro one segment at a time; and
  • begins playback when the first audio segment is ready, while later segments continue processing.

When synthesis finishes, a combined WAV remains available for replay until temporary media cleanup removes it.

Explore Word

Explore Word is a first-class camera mode for the moment when a child asks what one printed word says or means. A rectangular target selects the word. The browser maps that target to the source camera coordinates and uploads a tight crop with a modest amount of nearby context.

Gemma identifies the selected word and returns a structured explanation. OpenRead keeps the photographed display form separate from its pronunciation representation. The isolated word is spoken only when the response includes validated Kokoro/Misaki phonemes; uncertain or malformed pronunciation data fails visibly instead of being silently replaced with a guessed fallback.

The result includes a child-friendly meaning and example. For French words, the current contract provides:

  • the selected word pronounced with a French Kokoro voice;
  • a simple English meaning, spoken with an English voice;
  • a plain French equivalent, spoken with a French voice;
  • a French example sentence, spoken with a French voice;
  • an English translation displayed on screen but not spoken; and
  • optional grammatical gender or usage guidance when relevant.

Architecture

The Story Compiler is the central technical layer between visual understanding and speech:

Physical picture-book page
             |
             v
        Phone camera
             |
             v
     Image normalization
             |
             v
          Gemma 4
             |
             v
+-----------------------------+
| Structured Story Plan       |
|                             |
| ordered beats               |
| visible source text         |
| illustration context        |
| caregiver cues              |
| TTS segments                |
| spoken_script               |
+-----------------------------+
             |
             v
           Kokoro
             |
             v
 Incremental audio segments
             |
             v
         Child hears

The structured representation makes model output validatable and auditable. It supports diagnostics, separates source material from generated narration, prevents caregiver cues from entering speech, and gives Kokoro a deterministic sequence that can be synthesized and played incrementally.

Read Page defaults to gemma_vision, which sends the normalized page image directly to Gemma 4. An optional ocr_assisted mode runs PaddleOCR first and supplies its evidence to Gemma for ordering and correction; PaddleOCR remains an assisted and diagnostic path rather than the primary architecture.

The default inference provider is Cerebras, using gemma-4-31b. Google GenAI is also supported through configuration. Provider selection is environment-driven, not an in-app user setting.

The FastAPI backend exposes these primary endpoints:

Endpoint Purpose
POST /api/read/jobs Start a page-reading job from an image
GET /api/read/jobs/{request_id} Poll story, progress, and available audio segments
POST /api/word/jobs Start an Explore Word job from a target crop
GET /api/word/jobs/{request_id} Poll the word result and generated audio
POST /api/read Synchronous image or text-to-speech compatibility API
POST /api/ocr PaddleOCR diagnostic API
GET /media/audio/{request_id} Serve temporary generated WAV audio

Privacy and Trust

Uploaded page photos are not durably stored by the OpenRead application.

The current implementation has explicit storage boundaries:

Data Current behavior
Uploaded image Validated and normalized in memory, sent to the configured inference provider, then cleared from the application job
Provider processing The normalized page or target crop crosses the configured Cerebras or Google GenAI boundary; provider retention is governed by that provider's terms and deployment settings
Job state Read and word jobs are held in application memory and do not survive a backend restart
Generated audio WAV files and media metadata are temporary; the default media TTL is one hour, with periodic and disk-budget cleanup
Diagnostics Stored under backend/var/diagnostics/gemma/ for seven days by default; may include model output, validation errors, timings, provider/model metadata, request IP, and final structured results, but not the uploaded image
Accounts The current public reader has no user account or profile requirement

Diagnostics are enabled for successful and failed model calls by default so operators can reproduce malformed output, latency, pronunciation, and abuse issues without retaining page photographs. Retention settings are configurable for self-hosted deployments.

Quick Start

Requirements:

  • Docker with Docker Compose
  • a Cerebras API key for the default provider
git clone https://github.com/windrider2010/OpenRead.git
cd OpenRead
cp .env.example .env
# Edit .env and set CEREBRAS_API_KEY.
docker compose up --build

Open http://127.0.0.1:8001.

To use Google GenAI instead, configure GEMINI_API_KEY and set both STORY_COMPILER_PROVIDER and WORD_EXPLORER_PROVIDER to google_genai in .env.

For development without Docker, see backend/README.md. The usual entry points are:

uv sync --project backend
uv run --directory backend python -m uvicorn app.main:app --host 0.0.0.0 --port 8000 --reload

cd web
npm ci
npm run dev

Kokoro pronunciation preparation uses eSpeak NG and Misaki; install the native eSpeak NG runtime when running the backend outside the provided container.

Development and Testing

# Backend tests
uv run --directory backend python -m pytest

# Frontend tests
npm --prefix web test

# Production frontend build
npm --prefix web run build

The Story Compiler benchmark compares direct Gemma vision with the PaddleOCR-assisted path over fixture pages:

uv run --directory backend python scripts/benchmark_story_compiler.py

The Word Explorer benchmark compares configured Google Gemma models and crop behavior over the included real-photo fixtures:

uv run --directory backend python scripts/benchmark_word_explorer.py

Both benchmark scripts currently use Google GenAI and require GEMINI_API_KEY; their output is written under backend/var/diagnostics/ and is ignored by Git.

Repository Layout

backend/   FastAPI application, model services, tests, and benchmarks
web/       Vue/Vite mobile reader and public project site
deploy/    Docker, Nginx, and systemd deployment examples
docs/      Technical write-up and project documentation

Governance

OpenRead is an independent open-source public-interest AI project for early-literacy access, maintained by Hewei Li with infrastructure support from Sperion LLC.

Issues and pull requests are welcome through the GitHub repository.

Origins and Acknowledgements

OpenRead originated in work prepared for the Kaggle Gemma 4 Good Hackathon. The project has since developed as an independent public-interest open-source initiative. See the technical write-up for the original design rationale and implementation discussion.

OpenRead builds on major open-source projects and model ecosystems including:

OpenRead's Apache-2.0 license applies to this repository's original code and documentation. Models, packages, and hosted inference services remain subject to their own licenses and terms; consult the lockfiles and linked upstream projects for exact dependency versions and notices.

Citation

If OpenRead supports research, evaluation, teaching, or public-interest deployment, please cite:

Li, Hewei. OpenRead: A layout-aware picture-book read-aloud system using Gemma 4 and Kokoro TTS. 2026. https://github.com/windrider2010/OpenRead

@software{li2026openread,
  author = {Li, Hewei},
  title = {OpenRead: A Layout-Aware Picture-Book Read-Aloud System Using Gemma 4 and Kokoro TTS},
  year = {2026},
  url = {https://github.com/windrider2010/OpenRead}
}

License

OpenRead is licensed under the Apache License 2.0.

About

An open source book reader for kids, built for my daughter

Resources

Stars

27 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages