FastAPI backend for the OpenRead picture-book read-aloud and Word Explorer app.
Core endpoints:
POST /api/ocrPOST /api/readPOST /api/read/jobsGET /api/read/jobs/{request_id}POST /api/word/jobsGET /api/word/jobs/{request_id}GET /media/audio/{request_id}GET /media/audio/{request_id}/segments/{segment_index}GET /healthz
Endpoint notes:
/api/ocraccepts an uploadedimageplus optionallang_hintand returns recognized text, OCR blocks, confidence values, boxes, and detected script names./api/readaccepts either uploadedimageor formtext, never both. Image reads run through the Gemma story compiler before Kokoro TTS; text reads go directly to Kokoro. It returns JSON metadata by default or streams WAV bytes withresponse_mode=stream./api/read/jobsis the mobile UI path. It accepts optionalcompiler_mode=gemma_vision|ocr_assistedandcompiler_provider=google_genai|cerebras, returns202with arequest_id, and is polled through/api/read/jobs/{request_id}until the job iscompletedorfailed.- Read-job stages are
queued,story_compile,ocr,tts,completed, andfailed. Duringtts, progress is exposed asparagraphs_completedandparagraphs_total;audio_segmentsgrows whenever a Kokoro story segment is ready, allowing playback before the final WAV is complete. - Read and Word Explorer job status responses include millisecond
timingsfor image normalization, queue wait, Gemma processing, TTS, media storage, and end-to-end total time. - Completed image jobs include a
storypayload with ordered story beats, caregiver cues, compiler diagnostics, and thespoken_scriptused for TTS. /api/word/jobsis the Word Explorer path. It accepts an uploadedimageplus optionallang_hint, returns202with arequest_id, and is polled through/api/word/jobs/{request_id}until the job iscompletedorfailed. French results require a French pronunciation, a simple English meaning, a plain-French equivalent, a French example, and its English translation; gender and usage guidance are optional. Kokoro speaks the French word/equivalent/example withDEFAULT_FR_VOICEand the English meaning withDEFAULT_EN_VOICE.- The browser crops Word Explorer photos to the visible rectangular target with a small margin before upload. Word jobs preserve that patch and make one configured Gemma request to identify and explain the centered word. The exact
selected_wordremains display text; a separate required pronunciation supplies raw Kokoro/Misaki phonemes for the isolated word clip. Meaning and example segments continue through normal text G2P. Cerebrasgemma-4-31bis the default; Google GenAI can be selected withWORD_EXPLORER_PROVIDER=google_genai. No physical pointer is required. - Word-job stages are
queued,word_detect,tts,completed, andfailed. Completed word jobs include awordpayload with the selected word, speech kind, required pronunciation, explanation, optional example, diagnostics, and displayspoken_script. Missing, uncertain, malformed, or unsupported phonemes fail the job; there is no pronunciation substitution. - Ready story segments are served from
/media/audio/{request_id}/segments/{segment_index}. The completed 24 kHz WAV remains available from/media/audio/{request_id}until the media TTL expires or disk-budget cleanup removes it. lang_hint=enselects PaddleOCR English. Other values, includingbilingualandzh, select PaddleOCR Chinese for mixed Chinese/English target pages.- Raw Gemma text outputs and validation diagnostics are temporarily written to
backend/var/diagnostics/gemma/{request_id}.jsonwhen image story compilation or Word Explorer runs. These diagnostics include per-attempt image encoding, generation, parsing, pipeline, and total timings plus client IP for abuse investigation, but do not include uploaded images. - Page story compilation uses
STORY_COMPILER_MAX_OUTPUT_TOKENSandSTORY_COMPILER_TEMPERATUREto bound Gemma generation and reduce malformed runaway JSON. Keep the output-token cap high enough for one page, but low enough to avoid multi-minute repeated-character responses. compiler_provider=cerebrasuses Cerebras Chat Completions with image input and strict JSON schema output. It is the default for page reads;compiler_provider=google_genairemains available as an API/env fallback.PRELOAD_MODELS=1enables startup preload;PRELOAD_TTS=1warms Kokoro, whilePRELOAD_OCR=0keeps PaddleOCR lazy-loaded by default.
Diagnostics:
uv run --directory backend python scripts/run_fixture_pipeline.py- Runs
tests/ocr_voice_test.pngthrough the real OCR + Kokoro TTS stack. - Saves
.txt,.wav, and timing.jsonoutputs underbackend/var/diagnostics/.
Story compiler benchmark:
uv run --directory backend python scripts/benchmark_story_compiler.py- Runs
gemma_visionandocr_assistedover fixture images. - Saves JSON reports under
backend/var/diagnostics/openread/.
Word Explorer model benchmark:
uv run --directory backend python scripts/benchmark_word_explorer.py --model gemma-4-31b-it --model gemma-4-26b-a4b-it- Uses the five real camera fixtures under
tests/fixtures/word_explorer/by default and applies a configurable benchmark center crop; pass--fixture-diror repeated--fixturearguments to override them, or--no-center-cropfor a full-frame comparison. - Runs the same photos through each model with production image normalization.
- Saves per-run JSON, aggregate JSON, and CSV timing reports under
backend/var/diagnostics/word-explorer-benchmark/by default. - Measures Gemma image encoding, generation, structured-output parsing, and total service latency. It deliberately excludes TTS so model comparisons are not distorted by the voice engine.