Skip to content

Latest commit

 

History

History
53 lines (42 loc) · 5.74 KB

File metadata and controls

53 lines (42 loc) · 5.74 KB

OpenRead Backend

FastAPI backend for the OpenRead picture-book read-aloud and Word Explorer app.

Core endpoints:

  • POST /api/ocr
  • POST /api/read
  • POST /api/read/jobs
  • GET /api/read/jobs/{request_id}
  • POST /api/word/jobs
  • GET /api/word/jobs/{request_id}
  • GET /media/audio/{request_id}
  • GET /media/audio/{request_id}/segments/{segment_index}
  • GET /healthz

Endpoint notes:

  • /api/ocr accepts an uploaded image plus optional lang_hint and returns recognized text, OCR blocks, confidence values, boxes, and detected script names.
  • /api/read accepts either uploaded image or form text, never both. Image reads run through the Gemma story compiler before Kokoro TTS; text reads go directly to Kokoro. It returns JSON metadata by default or streams WAV bytes with response_mode=stream.
  • /api/read/jobs is the mobile UI path. It accepts optional compiler_mode=gemma_vision|ocr_assisted and compiler_provider=google_genai|cerebras, returns 202 with a request_id, and is polled through /api/read/jobs/{request_id} until the job is completed or failed.
  • Read-job stages are queued, story_compile, ocr, tts, completed, and failed. During tts, progress is exposed as paragraphs_completed and paragraphs_total; audio_segments grows whenever a Kokoro story segment is ready, allowing playback before the final WAV is complete.
  • Read and Word Explorer job status responses include millisecond timings for image normalization, queue wait, Gemma processing, TTS, media storage, and end-to-end total time.
  • Completed image jobs include a story payload with ordered story beats, caregiver cues, compiler diagnostics, and the spoken_script used for TTS.
  • /api/word/jobs is the Word Explorer path. It accepts an uploaded image plus optional lang_hint, returns 202 with a request_id, and is polled through /api/word/jobs/{request_id} until the job is completed or failed. French results require a French pronunciation, a simple English meaning, a plain-French equivalent, a French example, and its English translation; gender and usage guidance are optional. Kokoro speaks the French word/equivalent/example with DEFAULT_FR_VOICE and the English meaning with DEFAULT_EN_VOICE.
  • The browser crops Word Explorer photos to the visible rectangular target with a small margin before upload. Word jobs preserve that patch and make one configured Gemma request to identify and explain the centered word. The exact selected_word remains display text; a separate required pronunciation supplies raw Kokoro/Misaki phonemes for the isolated word clip. Meaning and example segments continue through normal text G2P. Cerebras gemma-4-31b is the default; Google GenAI can be selected with WORD_EXPLORER_PROVIDER=google_genai. No physical pointer is required.
  • Word-job stages are queued, word_detect, tts, completed, and failed. Completed word jobs include a word payload with the selected word, speech kind, required pronunciation, explanation, optional example, diagnostics, and display spoken_script. Missing, uncertain, malformed, or unsupported phonemes fail the job; there is no pronunciation substitution.
  • Ready story segments are served from /media/audio/{request_id}/segments/{segment_index}. The completed 24 kHz WAV remains available from /media/audio/{request_id} until the media TTL expires or disk-budget cleanup removes it.
  • lang_hint=en selects PaddleOCR English. Other values, including bilingual and zh, select PaddleOCR Chinese for mixed Chinese/English target pages.
  • Raw Gemma text outputs and validation diagnostics are temporarily written to backend/var/diagnostics/gemma/{request_id}.json when image story compilation or Word Explorer runs. These diagnostics include per-attempt image encoding, generation, parsing, pipeline, and total timings plus client IP for abuse investigation, but do not include uploaded images.
  • Page story compilation uses STORY_COMPILER_MAX_OUTPUT_TOKENS and STORY_COMPILER_TEMPERATURE to bound Gemma generation and reduce malformed runaway JSON. Keep the output-token cap high enough for one page, but low enough to avoid multi-minute repeated-character responses.
  • compiler_provider=cerebras uses Cerebras Chat Completions with image input and strict JSON schema output. It is the default for page reads; compiler_provider=google_genai remains available as an API/env fallback.
  • PRELOAD_MODELS=1 enables startup preload; PRELOAD_TTS=1 warms Kokoro, while PRELOAD_OCR=0 keeps PaddleOCR lazy-loaded by default.

Diagnostics:

  • uv run --directory backend python scripts/run_fixture_pipeline.py
  • Runs tests/ocr_voice_test.png through the real OCR + Kokoro TTS stack.
  • Saves .txt, .wav, and timing .json outputs under backend/var/diagnostics/.

Story compiler benchmark:

  • uv run --directory backend python scripts/benchmark_story_compiler.py
  • Runs gemma_vision and ocr_assisted over fixture images.
  • Saves JSON reports under backend/var/diagnostics/openread/.

Word Explorer model benchmark:

  • uv run --directory backend python scripts/benchmark_word_explorer.py --model gemma-4-31b-it --model gemma-4-26b-a4b-it
  • Uses the five real camera fixtures under tests/fixtures/word_explorer/ by default and applies a configurable benchmark center crop; pass --fixture-dir or repeated --fixture arguments to override them, or --no-center-crop for a full-frame comparison.
  • Runs the same photos through each model with production image normalization.
  • Saves per-run JSON, aggregate JSON, and CSV timing reports under backend/var/diagnostics/word-explorer-benchmark/ by default.
  • Measures Gemma image encoding, generation, structured-output parsing, and total service latency. It deliberately excludes TTS so model comparisons are not distorted by the voice engine.