Skip to content

Latest commit

 

History

History
1115 lines (799 loc) · 143 KB

File metadata and controls

1115 lines (799 loc) · 143 KB

Multilingual API benchmark — 2026-08-29

All-language beta readiness is not established. The current opt-in implementation includes 27 MMS recognition/alignment routes with guarded native language detection. The exact 50fd5ecf… Faroese source is integrated after the checks below, retaining the earlier Haitian and sixteen-language routes. Defaults and deployment are unchanged. Across different model versions and uneven sample sizes, 98/99 base languages have some reference evidence. This is not 98 quality passes. The frozen sixteen-language expansion has lower error in all 32 language/mode groups, but still has two automatic failures and known long-audio timing warnings. New Haitian, Latin and Sanskrit measurements expose additional quality/detection gaps. Turkmen now has a separate 20-recording operational comparison, so every base language has some operational or reference-based evidence across these mixed versions; Turkmen still has no trusted reference score. Regional-alias accent/dialect evidence remains open.

Fresh Cap holdout — completed API comparison

The frozen current 50fd5ecf… worker ran against 49 previously unused Cap recordings, 1.842 audio hours, across 47 known owner groups. Fifty were originally selected; one unavailable recording remains in coverage, without replacement. Known owner/source/audio exclusions were applied, but thirteen earliest legacy clips lack owner IDs, so complete owner independence is not established. The 49 cached Cap transcripts remain sealed and unused; they are not human ground truth. Both providers received the same pinned mono PCM16 WAVs with unrestricted automatic language detection, identical shared options and fixed alternating order. No candidate was selected or tuned on this holdout.

All 98 submissions reached terminal states. Sandchest completed 44/49 and returned five errors; AssemblyAI completed 47/49 with two errors. There are 43 mutually completed pairs. Of the four local-error/provider-completed pairs, three provider transcripts are empty; the fourth has nine words, leaving a potential missed-speech case unresolved. Conversely, one local result has one word where AssemblyAI reports no spoken audio. These differences are neither proof of silence nor evidence that either provider is correct.

Fresh upload-to-terminal latency, 43 mutually completed pairs Sandchest, local M4 Max AssemblyAI, remote service
p50 1.966s 7.700s
p95 10.471s 17.161s
p99 17.244s 19.996s

This includes upload, queue, inference, persistence and polling, with startup separate. Inference/build/model-hashing workloads did not compete with the timed requests; small metadata/review work continued. It is not a same-hardware, load, remote-database/billing or production-capacity comparison. All failures remain reported outside completion-only latency percentiles.

Returned base-language labels agree on 39/43 mutually completed pairs; labels are predictions, not independently checked language truth. The frozen case-preserving codepoint comparison scores 40 pairs at 10.27% pooled output disagreement; three longer pairs exceed its predeclared scoring-size limit and remain unscored. This is not human WER/CER. Word-boundary provider agreement is available on 36 pairs and is not acoustic timing accuracy. All 44 completed local responses have valid word structure; six completed provider responses have structural defects but remain in completion/text reporting. The local run used Parakeet on 30 recordings and Whisper on 14; none exercised an MMS recognition route. It therefore does not validate all 27 MMS routes or all supported languages in Cap.

The isolated audit reconciles 49 jobs, one attempt each, 44 exact-duration usage events and five unbilled errors. All 49 same-ID replays, conflicting-request checks, completed/error output resources and listing checks pass. Ephemeral keys are revoked and both owned process groups exited cleanly. The frozen reporter initially failed four error-shape checks because it wrongly required null duration for the established no-spoken-audio error. A separate read-back audit accepts zero only for that exact error, with text/words/model still null; existing source and two actual provider errors independently confirm the convention. Nine focused checks pass. Original failed report, every response and all failure counts remain unchanged; no request was repeated. The corrected runtime contract gate passes, not a recognition or beta-readiness gate.

A separate single-call diagnostic on the unchanged worker reproduces the potential missed-speech case's VAD rejection: peak speech probability 0.3304 across 569 checked frames, below the existing 0.5 threshold, with zero encoder/transcription calls. The waveform, settings and threshold are unchanged; captions remain sealed and no provider call is repeated. This establishes the rejection path, not whether the nine provider words are real speech. Runtime decoded sample count is not exposed; input WAV frames are kept separate. The helper records expected library paths from lsof -Fn, which alone cannot distinguish open descriptors from memory mappings. Private evidence: cap-vad-diagnostic1/execution; all owned processes exited and all 158 full pins stayed unchanged.

Private evidence: cap-unused-validation1 (local role holdout) and cap-holdout-api1/{plan-v2.json,harness-freeze.json,execution/report.json,readback-audit1}. The Python-version preflight failure started no service or provider job and is preserved separately. General beta remains unapproved: potential missed speech, language disagreements, independently verified quality/timing, complete feature parity and deployment/capacity proof remain open. No deployment or production data mutation occurred.

Turkmen — first operational API comparison, no reference accuracy

Twenty publisher recordings from twenty source-video groups now have 80 fresh API results, covering explicit tk and unrestricted automatic detection on the same frozen PCM. Source audio totals 179.168 seconds; each provider receives every recording in both modes. The publisher's caption provenance is unknown and captions remain unused. Language and speech truth are not independently established; different source videos do not establish different speakers or training independence.

Mode Sandchest completions AssemblyAI completions Paired local p50 / p95 Paired remote p50 / p95
Explicit Turkmen 20/20 20/20 1.235s / 3.596s 4.298s / 10.229s
Automatic detection 18/20 20/20 1.260s / 2.383s 4.162s / 14.121s

Latency includes upload, queue, inference, persistence and polling. Each row uses mutually completed pairs, so the automatic row has eighteen pairs; the two local failures remain in completion counts. This serial local M4 Max versus remote-service study has a fixed local-first ordering and small samples. It is not a same-hardware, concurrent-load or production-capacity result.

Both services return tk for all explicit requests, which does not prove recognition correctness. Neither returns tk automatically. Automatic returned labels agree on only 1/18 mutually completed pairs. Frozen NFC/casefold/whitespace token disagreement is 95.22% explicitly and 98.12% automatically, with no exact transcript matches. These are provider disagreements, not WER or evidence that either service is right; they must not be averaged with the Cap study's different codepoint metric. Turkmen recognition and automatic routing remain quality gaps.

All 38 completed local responses have valid word structure; eleven completed provider responses have structural defects and their text still remains in comparison. The isolated audit verifies forty single-attempt local jobs, thirty-eight exact-duration usage events, two unbilled automatic errors, all forty replay/conflict/resource checks, revoked keys and clean owned process exits. The frozen operational gate passes. Root independently rehashed 220 input/asset files and nineteen harness files, checked all eighty outcome/terminal-body identities, and reproduced the per-mode disagreement and latency summaries.

The unchanged local candidate is 50fd5ecf…: explicit requests use Whisper Turbo; automatic completions use Whisper on seventeen clips and Parakeet on one. AssemblyAI uses Universal-2 for all explicit requests and Universal-2/Universal-3.5 Pro for sixteen/four automatic requests. No Turkmen MMS route was used or added. Private evidence: turkmen-operational1, turkmen-api1/execution/report.json and turkmen-api1/root-readback1/receipt.json. This fills an operational coverage gap, not the trusted-reference, timestamp or beta-readiness gates.

Private Turkmen MMS pilot — operational pass, no quality promotion

A separate, pinned facebook/mms-1b-all Latin-script Turkmen adapter completes all twenty existing development recordings. All twenty raw word structures pass, but only 17/20 satisfy the unchanged all-word MMS timing gate; the other three retain raw recognition and diagnostic timing without being counted as accepted alignment. No gate is relaxed. Silence and tone return empty text, two invalid requests are rejected, and the final spoken replay exactly retains text, confidence and word times. All 25 protocol requests finish; no API/provider job is added. The binary and loaded native libraries are pinned, and the worker/inspector process groups exit cleanly.

Native request latency is 202ms p50 / 361ms p95 on these short clips, excluding model startup. This is not full API latency, a direct provider speed comparison or concurrent capacity. Peak native active allocation is approximately 5.21 GiB, not whole-process RSS.

The candidate emits Latin script while the saved explicit AssemblyAI responses are almost entirely Cyrillic; prior Sandchest responses are mixed. The frozen raw token disagreement is 100% against AssemblyAI and 98.54% against prior Sandchest. A separate script audit flags all twenty candidate/provider pairs and eight candidate/prior-worker pairs. Script differences explain why raw string disagreement alone cannot establish comparative recognition quality; they do not establish the actual spoken language or which output is correct. Trusted human Turkmen references and word boundaries are still absent. No serving route or automatic detector is promoted. Private evidence: turkmen-mms-development1/{runtime-run-v2-1,root-runtime-v2-1,root-result-diagnostic1}.

Faroese — native route and unused validation

The opt-in native worker now adds a pinned Faroese MMS adapter with acoustic word alignment. Automatic handoff requires positive speech, agreement from both native detectors at raw confidence ≥0.8, and compatibility with the caller's expected languages. Low-confidence/disallowed predictions retain the prior route. The implementation is the exact private 50fd5ecff1244e1c23e52f5f24cb03027ebc163593ea3f2b3154fa70b751efb1 binary's six-file change; it does not change deployment or defaults.

The 20-recording development set spans 13 conservative publisher-prefix groups and 2.541 minutes. The 20 previously unused validation recordings span 13 different groups and 2.138 minutes, from eight prospectively selected windows of the same pinned Ravnursson publisher test split. Candidate weights/code/policy were frozen before validation collection, and comparison gates before any validation inference. Source/audio/development-text identities and prior groups were excluded. These are normalized read-speech references, not independently verified people, training independence, Cap-domain speech or manual word boundaries.

Cohort / mode Sandchest WER AssemblyAI WER Sandchest CER AssemblyAI CER Correct auto labels local / provider
Development explicit 51.56% (99/192) 92.71% (178/192) 14.47% 46.04%
Development automatic 53.65% (103/192) 101.56% (195/192) 16.37% 44.35% 19/20 / 0/20
Unused validation explicit 37.22% (67/180) 100.00% (180/180) 12.02% 46.47%
Unused validation automatic 42.22% (76/180) 107.78% (194/180) 15.49% 54.89% 18/20 / 2/20

The historical local development baseline had 96.35% explicit / 100.52% automatic WER and one explicit failure. That baseline and the 40 development provider responses were reused unchanged, not submitted again. The candidate adds 40 actual development API jobs. Validation has 80 fresh paired API jobs: 40/40 local and 40/40 provider complete. Each local job ran once and persisted one unique, duration-matched usage event; all temporary keys were revoked and owned processes stopped. Local word structures are valid in all 80 development/validation jobs. Two development and one validation provider responses have structure defects; their actual text remains scored.

Both validation modes pass the predefined recording-level bootstrap comparison (5,000 resamples). Candidate-minus-provider WER intervals are −72.90 to −52.94 percentage points explicitly and −85.63 to −47.49 automatically. This is a comparative improvement on this small corpus, not an absolute accuracy or all-language beta pass: Faroese error remains high, two validation automatic labels are wrong, punctuation/formatting and human word-timing accuracy remain unverified. No validation result was used to tune the route.

Fresh validation full API latency Sandchest explicit AssemblyAI explicit Sandchest automatic AssemblyAI automatic
p50 0.526s 1.979s 0.648s 2.380s
p95 0.784s 2.555s 0.779s 3.755s
p99 0.784s 2.820s 0.779s 4.848s

These are serial local M4 Max versus remote-service measurements, including upload, queue, inference, alignment, isolated PGlite and polling. Startup, real remote-database/billing, concurrent load and deployment effects are excluded; other agents performed CPU-only preparation during parts of the work. They are not matched-hardware or production-capacity measurements. Explicit AssemblyAI used Universal-2; automatic responses retain their actual Universal-2/Universal-3.5 Pro labels.

The exact candidate preserves 1,206 retained native outputs, including 50 Cap recordings and earlier Haitian/broad-language routes. It passes 208 API compatibility cases plus 208 same-ID retries: 179 completions/unique usage events, 29 expected unbilled errors, one attempt each, and exact multichannel accounting. Faroese controls cover hints, threshold rejection, spelling, presentation flags, stereo and a shared exact-PCM crop fixture. Two real one-second timeout failures recover exactly. Eight derived long checks have no words inside or spanning the inserted pauses; they reuse existing audio and do not resolve the earlier sixteen-language timing warnings. Rust suites pass 122 base, 147 Whisper, 171 CTC and 232 MMS tests.

Private evidence: faroese-{native1,lid1,auto1,api1,options1,validation1,validation-api1} and durable1/faroese-source-checkpoint1. Source promotion verified 9,177 pinned evidence files, preserved prior source archives, and left HEAD/index, services, shared databases and environment unchanged. The reusable readiness inventory is documented in the evaluation workflow; it does not turn historical mixed-version coverage into current quality approval.

Nynorsk Small — promising private development result

The pinned native NB-Whisper Small candidate 3ee8b034… completed twenty short Nynorsk development recordings and four longer joins derived from those same recordings. It uses explicit nn, Small DTW timing and the unchanged request settings. Existing Sandchest and AssemblyAI outputs are reused; there are no new provider jobs or fresh validation recordings in this study.

Fixed cohort and metric NB-Whisper Small Earlier Sandchest Cached AssemblyAI
20 short clips: WER 13.53% (23/170) 36.47% (62/170) 37.65% (64/170)
20 short clips: CER 6.11% (46/753) 11.69% (88/753) 11.02% (83/753)
4 derived joins: WER 18.24% (31/170) 35.29% (60/170) 34.71% (59/170)
4 derived joins: CER 10.89% (82/753) 11.55% (87/753) 10.62% (80/753)

An independent regex/jiwer audit reproduces all 72 per-case scores, every reference denominator and both comparison cohorts after verifying the frozen inputs and outputs. The two cohorts reuse the same reference content and must not be added together as independent observations. Longer-clip CER does not beat AssemblyAI. All 24 readable native outputs remain in the accuracy calculation, including the final 33.38-second clip whose initial checker wrongly treated token-boundary whitespace as a text-integrity defect. The original failed report remains unchanged.

A separately frozen five-call replay corrects only the checker to match the public API's NFKC/whitespace rule, preserving punctuation, case, marks and numbers. The long transcript and every word timestamp/confidence exactly match the original baseline. Silence, tone, invalid-language rejection and exact spoken recovery all pass. Root independently verifies all 346 input pins, 73 output pins and clean native/inspector/runner exits. This closes those operational checks, not fresh quality, automatic Nynorsk, human timestamp accuracy or API-speed parity. No serving route is promoted. Evidence: nynorsk-nb-small1/{runtime-run-v4-1,runtime-score-v4.json,root-scoring-v4-1} and nynorsk-nb-small-controls1/{runtime-run-v1-1,root-runtime-v1-1}. Independent score receipt: nynorsk-nb-small1/independent-score-readback1/receipt.json (c4ba866a…).

Automatic Hindi — lower development error, language detection still weak

The unchanged private ad46fd67… candidate adds opt-in Qwen recognition only after Whisper actually detects Hindi and returns a valid result. The same 60 consumed Hindi development recordings were tested with routing off and on, with 30 pairs in each order. All 120 primary native requests complete and pass their structure checks. No AssemblyAI request was made in this comparison.

Same 60 development recordings Automatic routing off On
WER 42.02% (619/1473) 31.36% (462/1473)
CER 25.71% (1533/5962) 21.91% (1306/5962)
Native HTTP median 0.977s 1.952s

An independent regex/jiwer audit reproduces every score and text-token identity. A fixed 10,000-draw clip bootstrap gives on-minus-off WER interval −13.90 to −7.81 percentage points, and CER −5.46 to −2.40. These are reused development clips, not independent speaker or population evidence. Native HTTP timings include observer work, exclude startup/upload/queue/persistence, and cannot be compared with provider API latency. Automatic routing adds a recognition pass and is slower.

The label distribution is unchanged: 52 Hindi, six Urdu, one Gujarati and one Indonesian. The eight non-Hindi labels account for 220 of the candidate's 462 word errors; Qwen is never used for them. For the 52 Hindi-labelled recordings, it is accepted on 40; twelve retain the whole prior Whisper result. The current automatic selector therefore remains an accuracy gap. The separate full API comparison below measures this automatic mode without inferring it from explicit results.

The original run stopped during controls because its private-native 422 checker expected error instead of the actual detail envelope. Its failed report is preserved. A separately frozen 26-call companion now passes all thirteen controls with the feature on and off: language/fallback distinctions, strict confidence threshold, nonspeech, Faroese/Haitian regressions, exact 30-second eligibility, long-audio exclusion, explicit anchors and bare automatic options. All four owned child groups and the runner exit cleanly; root verifies all 997 input and 2,024 output pins. Controls and development accuracy are separate evidence, with no human word-time, source-promotion or beta approval. Evidence: durable1/hindi-qwen-auto-runtime1/{run-v4-1,independent-score-readback1} and durable1/hindi-qwen-auto-controls1/{run-1,root-runtime-1}.

Exact automatic-Hindi build — retained Cap regression passed

The same ad46fd67… build now passes 100/100 native requests on the original 50 retained Cap development recordings: both Hindi flags off, then both on. Every result matches its earlier b502e2df… semantic baseline, including complete text, words, word/token times, confidences, language metadata, model identity and alignment decisions. All 50 off/on pairs also match. Only declared elapsed-time and allocator observations are excluded from equality.

The cohort has 32 Parakeet-English and 18 Turbo outcomes, including five empty results; none previously selected Hindi, so no new Hindi route is expected here. These are retained model predictions, not independently checked language truth. This is exact-build regression evidence, not 50 new recordings, human accuracy, word-boundary accuracy or provider-speed parity. No provider, API or reference request is made. Root verifies 1,533 input and 2,309 output pins, actual exit0 and absence of both native groups, both inspectors and the runner. Evidence: durable1/hindi-auto-cap-regression1/{run-1,root-runtime-1}; runtime report 25b9b1c1…, root readback f3eeb2bf…. No serving defaults or deployment changed.

Exact automatic-Hindi build — advanced API regression passed

A separate 22-job API replay on unchanged ad46fd67… now passes all frozen checks: 20 completed jobs, two expected errors, exactly one attempt each and 20 unique usage events. Both Hindi flags are enabled. Twenty cases preserve their historical stable API fields; two automatic-Hindi cases match predeclared whole-Qwen outputs while preserving their own language metadata. These two are intended changes, not historical-baseline matches or new accuracy evidence.

All 88 upload/submit/replay/conflict POSTs and 111 resource/list GETs reconcile. The suite covers crop offsets, spelling/format flags, silence, malformed inputs, long-audio exclusion and stereo resources. Word conservation preserves every field, duplicate and channel-local order without imposing a false global stereo order. Root verifies 556 input and 2,024 output pins and actual clean exits of native, API, inspector and runner groups. Evidence: durable1/hindi-auto-advanced-api1/{execution,root-execution1}; report 0d5e8528…, root readback 277dcdcd…. This uses private PGlite with external billing disabled; no provider call, reference scoring, deployment or production reliability claim. Full AssemblyAI feature parity remains separate.

Hindi detector contrast screen — limited evidence, no routing change

A fixed follow-up runs both existing native detectors on 24 reused development recordings: the eight Hindi misses, four correctly labelled Hindi examples, and four each of Urdu, Gujarati and Indonesian. All 48 classifier requests finish, with complete 107/126-class distributions retained. Inputs, label mappings and individual 0.8 confidence thresholds are fixed before results; there is no ASR request, provider call, reference scoring or threshold adjustment.

Selected group Joint Hindi vote at ≥0.8 in both detectors
Previously missed Hindi 1/8
Previously correct Hindi controls 3/4
Urdu, Gujarati and Indonesian contrasts 0/12

Each detector individually casts a high-confidence Hindi vote for a different Urdu control. Requiring both prevents those two false votes in this small screen, but does not prove production precision or correct ASR output. Only one missed recording satisfies the joint rule; this is not eight recovered transcripts and no routing policy is promoted. Broader negative controls and an integrated recognition test are still required. Standalone detector timings do not measure serving/API performance. Root verifies all 302 input and 168 output pins and clean exits of both probes, both inspectors and the runner. Evidence: durable1/hindi-lid-screen1/runtime-preparation1/{run-1,root-run-1}; report b5571e01…. An independent arithmetic/raw-stdout audit reproduces all 48 requests and group votes after rechecking the 168 frozen outputs; receipt durable1/hindi-lid-screen1/independent-readback1/receipt.json (750b815e…).

Completed automatic Hindi API comparison

The unchanged ad46fd67… candidate now completes a separately frozen 80-job API comparison on the same 40 recordings previously consumed for explicit validation. This is not a fresh holdout. Both services receive identical WAV bytes and language_detection: true, without a language code, expected-language list or fallback hint. Twenty pairs run local first and twenty provider first. No model, threshold, option, sample or metric is changed after freezing.

All 40 recordings per provider Sandchest automatic Fresh AssemblyAI automatic
Completed 40/40 40/40
WER 21.76% (223/1025) 28.39% (291/1025)
CER 13.76% (575/4179) 17.11% (715/4179)
Hindi language labels 36/40 36/40
Full API p50 2.583s 4.639s
Full API p95 7.171s 8.386s
Full API p99 8.130s 12.482s

An independent tokenizer/edit-distance implementation reproduces all 80 raw and secondary scores and the fixed 10,000-resample intervals. Candidate-minus-provider WER is −6.63 percentage points, 95% interval −13.04 to −0.52. CER is −3.35 points, interval −9.79 to +2.18; the CER interval crosses zero. These are selected read-speech clip resamples, not independent-speaker or population guarantees. All completed text remains included, including structurally imperfect provider results. Sandchest's other labels are three Urdu and one Gujarati; AssemblyAI's are four Urdu. Full automatic language accuracy is not established.

All forty local output structures, same-ID replays, conflicting replays, 200 resource GETs and listing checks pass. Each job has one attempt and one unique usage event, totaling exactly 486,480 ms. Three AssemblyAI responses fail the frozen word-time structure checks: two extend past decoded audio duration and one has a zero-duration word. The combined runtime report remains failed for those three checks; no response is dropped or silently repaired. Root separately proves all 613 input and 4,071 output pins and clean native/API/inspector/runner exits.

Latency is the recorded serial local-versus-remote upload-to-terminal-body timer, including queue/persistence/polling and observer work, with startup separate. The terminal clock and upload wire span are independently checked; the original pre-upload start is not separately persisted, so that start depends on the frozen runner implementation. It is not a matched-hardware, load-capacity or production latency result. Human word-boundary accuracy, longer/conversational Hindi and all-language/API parity remain open. The candidate stays private and unpromoted. Evidence: hindi-auto-api-validation1/{execution,root-execution1,scoring,root-scoring1,independent-score-readback1}; independent receipt fc9cab98…, original runtime report 79a372f3….

Hindi unused validation — lower error and faster local API

The unchanged private 5cf89521… candidate completed a fresh 80-job paired API comparison on 40 previously unused Hindi recordings / 486.48 seconds. Selection was fixed before extracting audio/references and after pinning the candidate. Source IDs, source/canonical audio, PCM and reference hashes are disjoint from the recorded earlier experiments; all forty selected cases were available, with no replacements. This is FLEURS read speech, not fresh Cap conversational audio. Speaker identity and absence from model training remain unproved.

Same 40 published human-reference recordings Sandchest Fresh AssemblyAI
Completed requests 40/40 40/40
Raw-reference word error 13.56% (139/1025) 19.51% (200/1025)
Raw-reference character error 5.98% (250/4179) 7.87% (329/4179)
Secondary normalized-reference word error 13.17% (135/1025) 19.22% (197/1025)
Full API p50 1.74s 4.91s
Full API p95 3.88s 6.41s
Full API p99 8.04s 11.22s

Both services receive the identical canonical WAV and explicit-Hindi settings, with twenty local-first and twenty provider-first pairs. Sandchest uses accepted Qwen recognition on 35 clips and whole-Whisper fallback on five; all forty provider responses identify Universal-3.5 Pro. These are serial local-versus-remote upload/submit/poll measurements, not same-hardware performance or production capacity. The sample percentiles include request inference/alignment/fallback overhead; server/model startup is separate.

Every selected reference remains in the primary denominator, with unsuccessful outputs prospectively scored as empty hypotheses. Here all forty pairs complete, so failure-inclusive and completed-pair cohorts coincide. A fixed 10,000-draw paired clip bootstrap gives candidate-minus-provider intervals of −8.93 to −2.92 percentage points for WER and −3.35 to −0.51 for CER. These are descriptive clip resamples, not proof of independent speakers or general Hindi performance. An independent prior regex/jiwer implementation reproduces every raw/secondary score, all token hashes and both intervals after outputs are frozen. No model, threshold, sample or metric was tuned on these results.

All forty local jobs use one attempt and persist exactly one usage event for the exact duration; all replay/conflict/resource/list checks pass, keys are revoked and owned processes exit cleanly. Local word structures have no defects. One completed provider response contains a zero-duration word, so the original combined structural gate remains false; its text is retained in every accuracy score. This run does not exercise failed-job billing or establish human word-boundary accuracy.

This supports the private short, explicit-Hindi candidate; it is not all-language beta approval. Automatic Hindi routing, longer audio, conversational/domain generalization, human timing and full API feature parity remain separate work. The Qwen composition is currently verified on macOS Metal only. Evidence: hindi-unused-validation2/{collection-v3-1,root-collection-readback1,paired/{execution,root-output-readback1,scoring,independent-score-readback1}}. No serving default, shared service/database, deployment or production traffic changed.

Hindi integrated recognition and alignment — private development pilot

The private b502e2df… candidate combines native Qwen3-ASR 1.7B 8-bit recognition with the existing canonical Hindi MMS timing gate. It is opt-in, limited to explicitly requested Hindi and at most 30 seconds of selected audio. A recognized decoding or alignment-quality rejection returns the complete prior Whisper result, without mixing one model's text with another model's word times. Runtime, cancellation and deadline errors still propagate. All four decoding assets, including the EOS tokenizer configuration, are pinned; the running process was verified to map one expected MLX library.

The first real integrated native HTTP run completes 60/60 reused Hindi development recordings: 45 accepted Qwen results and 15 whole-Whisper fallbacks. Every result matches the expected accepted/fallback decision, text and complete word fields from the pinned development prototypes. No response has a structural defect. Twelve additional feature-off/on Faroese/Haitian, nonspeech, malformed-request and recovery controls also pass, for 72/72 checks. Both owned native process groups exited. The 40 previously consumed Hindi holdout recordings and new Cap holdout were not used.

Same 60 human-reference development recordings WER CER
Integrated private candidate, measured now 20.84% (307/1473) 10.23% (610/5962)
Historical Whisper baseline 32.52% (479/1473) 14.32% (854/5962)
Historical AssemblyAI results 19.42% (286/1473) 7.85% (468/5962)

On these earlier sixty development recordings, the candidate improves on the prior recognizer but remains numerically worse than the cached AssemblyAI results. The earlier 20.84% projection is now reproduced by actual integrated requests; it is still development evidence, not an untouched-validation or parity pass. These lexical scores do not verify human word boundaries or calibrated confidence. Earlier headline comparisons that discarded completed provider text because of word-metadata defects must not replace these corrected provider scores.

Measured local native HTTP pipeline latency is 1.077s p50 / 3.624s p95, including recognition, acoustic alignment and fallback overhead, excluding the transcript API's upload/queue/persistence. Accepted Qwen requests have 0.952s / 1.468s; fallbacks have 2.465s / 6.794s. Model initialization is excluded. Cached Whisper IPC and remote AssemblyAI full API durations have different scopes and are not a fresh speed comparison.

A separate retained-Cap run passes 100/100 native HTTP requests: all 50 existing development recordings match their frozen historical text and complete word fields with the feature disabled and enabled. This is a regression check, not 50 new recordings or independent Cap accuracy evidence. Both owned native process groups exited; no provider call was made.

A health-metadata-only successor, private 5cf89521…, lists the loaded Qwen model correctly without changing recognition or alignment. It passes 46 advanced native calls, covering the prior worker, disabled/enabled feature states, crop offsets, stereo, presentation options, automatic/long retention and invalid-input recovery. Its 22-job full API run has twenty completions with twenty exact-duration usage events and two expected unbilled errors; all jobs have one attempt, and replay/conflict/listing checks pass. The first reporter incorrectly required globally interleaved resource words for two stereo cases. An additive audit verifies exact text/time/confidence/channel/speaker multisets, duplicate counts, channel-local order, canonical utterance order and span bounds, correcting only those four resource-order flags. Sixteen focused checks pass; the original failed report remains unchanged. Actual AssemblyAI resource ordering was not newly measured.

The separate ten-call fault study returns two real one-second 504 errors with no partial output, preserves exact recovery after the Qwen timeout, and passes two full-body client disconnects with exact Hindi recovery. Immediate silence after the long Whisper timeout returns 503, so that original immediate-recovery gate remains failed. A separately frozen three-call diagnostic observes readiness states 503, 503, 200: ready at 213.58 ms after the 504 response body and exact silence recovery by 226.72 ms. This measures transient refusal while cancellation finishes; it does not rewrite the first failure or change serving code. A separate four-job queued-API study also passes: Qwen-eligible timeout → silence, then long-Turbo timeout → silence, with no native-readiness wait or probe between jobs. Recovery uploads begin 7.85/10.26 ms after the client observes each API error body. Both timeouts are unbilled, both silences exactly match the frozen API/native baseline and each has one one-second usage event; all four jobs use one attempt. All sixteen upload/submit/replay/conflict operations and twenty-one resource/list GETs reconcile, keys are revoked, and owned processes exit cleanly. These silence recoveries prove service/queue/accounting behavior in the measured serial cases, not spoken Whisper decoding after its long timeout or concurrent capacity. The original blocked advanced API-deadline phase was not modified or run; this is separately frozen evidence under hindi-api-timeout1/{execution,root-readback1}.

A further frozen three-call spoken-Whisper recovery study passes. The original 7.14-second automatic-language clip produces the same fourteen words before and after a forced one-second timeout on the original 42.84-second request. Text, word boundaries/confidences, language and model fields match exactly. Readiness returns after 111.94 ms and the spoken recovery completes by 719.97 ms after the 504 body. The owned native process exits cleanly and all pins remain unchanged. This closes spoken decoding recovery after readiness in this one serial case; it does not turn the original immediate 503 into a pass or establish concurrent capacity or human accuracy. Private evidence: hindi-whisper-timeout-spoken1/{execution,root-readback1}.

A further alternate MMS fallback was stopped before integration. Actual canonical alignment accepts only 3/15 cached fallback transcripts; ten fail boundary concentration and two fail supported-character checks. All fifteen calls finish cleanly with unchanged pins. Keeping only the three accepted replacements would project 298/1473 word errors (20.23%) and 583/5962 character errors (9.78%) on the reused development set, still worse than the provider's 286/468 errors. This is a cached-text projection, not new integrated recognition or an accuracy improvement in serving. No timing threshold was relaxed.

The Hindi candidate remains private. The later unused explicit-Hindi validation is reported above; automatic/long-audio behavior, human word timing, broader language behavior and complete API parity remain open. Evidence: durable1/hindi-qwen-integration2, hindi-qwen-integration3, hindi-qwen-runtime1/{hindi-dev1,cap-retained1}, hindi-qwen-advanced1/{runs,stereo-readback-audit1}, hindi-timeout-drain1/run-v2-1 and hindi-mms-fallback1/{run-v3-1,root-alignment-readback1}. These earlier Hindi development/regression/reliability checks added no provider call; the later fresh comparison above adds forty. No shared database/service or deployment changed.

Sanskrit specialized small model — rejected

A private native candidate tested the pinned Bidwill/whisper-small-sanskrit checkpoint (13d318013cbb5fa305ddd2e76f26196a3fd4cde4) on the same 20 Sanskrit development recordings. The raw tensors and canonical whisper.cpp conversion were verified, and an offline native build passed before the model ran. The frozen pilot made 49 real native requests, including 40 scoring requests, silence/tone, invalid inputs and recovery controls. No provider request or untouched validation example was added.

Fixed decoding condition Successful transcripts WER, all 20 references CER, all 20 references
Explicit Sanskrit 7/20 95.19% 74.25%
Native default acoustic language selection 10/20 93.27% 65.60%

Twenty-three scoring requests returned explicit unreliable-transcription errors, retained as deletions in the denominators. Default decoding identified Sanskrit correctly on 0/20. The four nonspeech controls stayed empty and both malformed inputs were rejected; the selected warmup/recovery recording continued to return a transcription error, so successful recovery was not established. All owned processes exited and input/model/source pins stayed unchanged.

The standalone candidate is rejected and remains private. Its completion rate and character error regress badly despite superficially similar aggregate word error to AssemblyAI. A retrospective whole-Turbo fallback calculation is only a development projection and was not implemented or advertised as measured performance. No threshold was relaxed to manufacture a pass. The exact checkpoint's training prefix remains unproved; the default diagnostic is not an assertion of the publisher's training procedure or permission to substitute Hindi. Native request latency excludes upload/queue/persistence and cannot be compared directly to the earlier provider API timings. Evidence: durable1/sanskrit-small-pilot1/run1/{runtime-receipt.json,report.json,root-decision.json}. Sanskrit remains a quality gap.

Sixteen-language expansion — completed explicit API comparison

A private full-worker candidate, b1cd442b613ccaab04d5f9d451fe54b0c65012c7b6a483ec0a927f17b34ce0dc, adds native MMS recognition and acoustic word alignment for sixteen languages. It is not yet promoted. Its automatic selectors remain unchanged; the follow-on automatic-routing candidate and completed API comparison are documented below.

The corpus contains 320 newly selected development recordings: 20 per language, 67.432 audio minutes, from pinned FLEURS published validation/dev data. Selection excluded the previous 64 examples using source identity, normalized development text, audio and canonical PCM. No protected validation/holdout reference was used. Speaker identities and model-training overlap are unknown; these short read-speech recordings do not establish Cap-domain or dialect-wide quality.

The first full API run completed 230/320 requests on the current edbef6a2… worker and 320/320 on AssemblyAI. All 90 local errors were explicit unreliable-speech errors, ran once, and were unbilled. Requests preferred Universal-3.5 Pro with Universal-2 fallback; AssemblyAI actually used Universal-2 for all 320 results. The replacement then completed 320/320 real API requests, each with one inference attempt and one unique exact-duration usage event. Isolated database audits reconcile all jobs; ephemeral keys were revoked and owned processes stopped.

Lower is better. Failed transcripts retain their reference deletions; completed provider text with encoding defects is scored rather than discarded. CER is primary for scripts without reliable whitespace word boundaries. Rates above 100% include insertions.

Language Metric Previous worker Candidate, full API AssemblyAI Universal-2
Bengali WER 89.70% (331/369) 30.35% (112/369) 105.42% (389/369)
Hausa WER 94.80% (419/442) 28.96% (128/442) 90.05% (398/442)
Georgian WER 104.18% (324/311) 29.90% (93/311) 110.93% (345/311)
Khmer CER 100.00% (2679/2679) 15.68% (420/2679) 98.92% (2650/2679)
Kannada WER 78.05% (224/287) 25.44% (73/287) 72.47% (208/287)
Lao CER 100.91% (2206/2186) 23.19% (507/2186) 100.64% (2200/2186)
Malayalam WER 100.00% (285/285) 36.49% (104/285) 105.61% (301/285)
Mongolian WER 97.62% (328/336) 38.39% (129/336) 106.25% (357/336)
Marathi WER 83.65% (307/367) 31.61% (116/367) 86.65% (318/367)
Burmese CER 100.03% (2881/2880) 14.79% (426/2880) 99.97% (2879/2880)
Nepali WER 85.58% (279/326) 31.90% (104/326) 93.25% (304/326)
Pashto WER 84.98% (430/506) 37.75% (191/506) 89.13% (451/506)
Somali WER 95.38% (392/411) 48.66% (200/411) 100.00% (411/411)
Tajik WER 104.24% (516/495) 16.16% (80/495) 92.53% (458/495)
Uzbek (Latin) WER 96.91% (314/324) 23.77% (77/324) 88.58% (287/324)
Yoruba WER 113.84% (551/484) 50.21% (243/484) 98.76% (478/484)

Every language improves against both comparators on this development set. Paired-recording bootstrap intervals (5,000 resamples, fixed seed, ratio of summed edits to reference units) are negative against both comparators for every group. These intervals describe the selected clips, not independent speakers, unseen data or population superiority. Absolute error remains substantial in several languages, especially Somali and Yoruba.

The full worker and API preserve all 320 raw-recognition text scores. All 320 native/API responses have identical text and full word metadata, with zero local word/encoding structure defects. Khmer, Lao and Burmese now group actual CTC token intervals into ICU words; indivisible acoustic tokens are not split using invented timestamps. Synthetic short-posterior and audio-window-boundary tests pass. These checks establish implementation consistency and valid timelines, not manually measured timestamp accuracy.

All 320 mutually completed API pairs Local candidate Remote AssemblyAI
p50 0.518s 3.189s
p95 0.775s 5.085s
p99 0.782s 5.759s

Latency includes upload, queue, inference, acoustic alignment, isolated PGlite persistence and polling. It excludes startup and real remote-database/billing roundtrips. This is serial local Apple M4 Max versus the remote service, not matched hardware, sustained load or a production-tail estimate. The local replay followed the provider baseline; it is not an interleaved capacity experiment.

The candidate also preserves 517 retained native cases exactly, including all 50 Cap recordings, and recovers to the same English output after the run. Its native suites pass 122 base,147 Whisper,171 CTC and223 MMS tests. No shared app source, environment, database or deployment changed.

Completed automatic API comparison

The follow-on private binary 997048bf8be116fa46855b9aba0088cb0f9cc061adb27581a58a4cde982653ae preserves that explicit recognizer and adds guarded automatic routing. It is not yet promoted. The policy enables fourteen additional primary-detector routes at confidence ≥0.8; Kannada/Uzbek require both detectors to agree above that threshold. A second opinion after a weak primary prediction must agree with it or match one of four measured pairs: Hindi→Nepali/Pashto, Persian→Tajik, or Thai→Lao. These new weak-primary handoffs also respect the caller's expected-language list. Earlier hint/fallback behavior is unchanged; this is not a general fallback-compatibility fix.

A broader low-confidence policy was rejected because it changed a retained Cap response through a wrong new route. The narrowed policy preserves all 518 retained native checks, including 50 Cap recordings and that case with unrestricted hints. It completes all 320 automatic native cases and recovers exactly. These are development/regression results, not guarantees of future detector precision.

The full run completes 960/960 API jobs: 320 local explicit,320 local automatic and320 fresh AssemblyAI automatic. The unchanged 320 explicit provider responses above are reused, yielding 1280 comparison rows on the same 320 unique development recordings. All 640 local responses match the verified native text, full word metadata and language/model metadata. Each local job has one attempt and one unique exact-duration usage event; keys are revoked and owned processes stopped.

All sixteen groups have lower primary text error in both modes. Automatic language labels are correct on 312/320 locally versus 200/320 remotely. Lao remains 13/20 and Uzbek 19/20 locally; every other local group is 20/20. The local automatic path uses MMS on 312 clips and retains Whisper on eight. AssemblyAI actually uses Universal-2 on 295 auto clips and Universal-3.5 Pro on 25; explicit responses all used Universal-2.

Language Metric Sandchest auto AssemblyAI auto Correct auto labels: local / provider
Bengali WER 30.35% (112/369) 105.42% (389/369) 20/20 · 20/20
Hausa WER 28.96% (128/442) 98.87% (437/442) 20/20 · 0/20
Georgian WER 29.90% (93/311) 113.18% (352/311) 20/20 · 16/20
Khmer CER 15.68% (420/2679) 99.18% (2657/2679) 20/20 · 20/20
Kannada WER 25.44% (73/287) 74.22% (213/287) 20/20 · 17/20
Lao CER 40.62% (888/2186) 101.33% (2215/2186) 13/20 · 0/20
Malayalam WER 36.49% (104/285) 105.61% (301/285) 20/20 · 20/20
Mongolian WER 38.39% (129/336) 103.87% (349/336) 20/20 · 18/20
Marathi WER 31.61% (116/367) 91.01% (334/367) 20/20 · 15/20
Burmese CER 14.79% (426/2880) 99.97% (2879/2880) 20/20 · 19/20
Nepali WER 31.90% (104/326) 92.33% (301/326) 20/20 · 19/20
Pashto WER 37.75% (191/506) 93.87% (475/506) 20/20 · 5/20
Somali WER 48.66% (200/411) 101.70% (418/411) 20/20 · 11/20
Tajik WER 16.16% (80/495) 111.31% (551/495) 20/20 · 0/20
Uzbek (Latin) WER 25.00% (81/324) 110.80% (359/324) 19/20 · 1/20
Yoruba WER 50.21% (243/484) 98.55% (477/484) 20/20 · 19/20

The fixed paired-clip bootstrap intervals are negative in all 32 language/mode groups. They describe these development recordings, not independent speakers or population superiority. No local word/encoding structure defects occur; provider defects occur in 55 explicit and 54 automatic completed responses, including replacement-character and word-bound checks. Their actual text remains scored; these are not reported as failed API jobs or independent acoustic-timing measurements.

Full API latency Sandchest explicit AssemblyAI explicit, cached Sandchest auto AssemblyAI auto, fresh
p50 0.529s 3.189s 0.526s 3.575s
p95 0.784s 5.085s 1.026s 6.115s
p99 1.032s 5.759s 1.032s 12.705s

These are serial local M4 Max versus remote API measurements, including queue/alignment/PGlite/polling and excluding startup/real remoteDB/billing. Lightweight corpus preparation and offline collector checks ran on the host during parts of the study; this study started no concurrent inference/build. Do not treat these as isolated hardware, cloud capacity or stable production-tail benchmarks.

The candidate passes 122 base, 147 Whisper, 171 CTC and 228 MMS Rust tests. The 188 API option cases and 188 same-ID retries also pass, including all 50 exact Cap responses. There are 161 completions with 161 unique usage events and 27 expected unbilled errors; channel-duration accounting is exact. New cases cover Bengali, Khmer, Burmese, Lao, Pashto, Kannada and Uzbek with hints, confidence rejection, crops, spelling, presentation flags and stereo. Both real one-second deadline failures return 504 without partial text, then recover exactly into automatic Pashto; raw response read-back passes. These are mechanical/reused-data checks, not additional quality samples. Longer-audio verification completes all 256 API jobs on 64 derived controls (35.54–98.96 seconds), with 128 local jobs in one attempt and 128 unique duration-matched usage events. Both modes have lower primary text error than AssemblyAI for all 16 languages. Local automatic labels are correct on 62/64; one Lao and one Marathi recording fall back to the wrong language. These controls reuse all 320 development recordings, five per control, and add no independent recordings. A coarse audit flags 12 local intervals spanning complete inserted 200ms pauses: six MMS intervals on three recordings (duplicated across explicit/auto) and six Whisper fallback intervals on two recordings. Zero local words lie entirely inside an inserted pause; interval bounds are valid. These remain open recognition/timing warnings, not manual timestamp measurements. Diagnosis completed before the unchanged candidate's opt-in source integration; the warnings do not become quality passes. Evidence is under mms-broad-long2, including the preserved failed pause audit. A separate 320-recording validation set from the published FLEURS test split is now collected after freezing the candidate, excluding 384 earlier development clips. The unchanged candidate has now completed this validation comparison: 638/640 local jobs and 640/640 provider jobs, with lower primary text error in all 32 language/mode comparisons; see the frozen-validation table below. Ten native diagnostic replays locate the six MMS pause crossings in the recognizer output before alignment, while the other six come from wrong-language Whisper fallbacks. A stricter VAD confirmation proposal was not adopted because it did not remove the problematic joins. No validation result was used to make that decision. Private evidence is under mms-broad-auto3; the rejected attempt remains under mms-broad-auto2.

MMS recognition output still lacks a verified multilingual punctuation/number-formatting restoration layer. Lexical WER/CER does not establish parity for punctuation, formatting, filler handling, code switching or manual word boundaries. Those remain acceptance gates alongside broader Cap validation, Turkmen reference data and the new Haitian/Latin/Sanskrit quality gaps below.

Sixteen-language frozen validation — 320 unseen recordings

The unchanged 997048bf… candidate was frozen before collecting 320 previously unused recordings, 20 per language, from the pinned FLEURS published test split. Local validation excludes 384 earlier development recordings by source ID, normalized development text, source audio and canonical PCM. Publisher speaker identities and model-training overlap are unknown; there are no manual word-time references. None of these validation results was used to tune the candidate.

The full paired matrix contains 1,280 actual API jobs. Sandchest completes 638/640, AssemblyAI 640/640. The two local failures are automatic Burmese and automatic Uzbek requests returning an explicit unreliable-transcription error, with no partial transcript or usage charge. They remain in text-error denominators as deletions; no retry replaces them. All 640 local jobs have one inference attempt, and the 638 completions have exactly 638 duration-matched usage events. Owned processes exited and temporary keys were revoked.

Primary error is lower than AssemblyAI for all 16 languages in both modes; paired-recording bootstrap intervals are below zero for all 32 comparisons. This describes the sampled read speech, not general language quality or independent speakers. Lower is better; rates over 100% include insertions. Khmer, Lao and Burmese use character error rate (CER) because whitespace WER is unsuitable; all other rows use WER.

Language / metric Explicit Sandchest / AssemblyAI Automatic Sandchest / AssemblyAI Correct auto labels Sandchest / AssemblyAI
Bengali / WER 23.10% / 107.38% 23.10% / 107.14% 20/20 / 20/20
Hausa / WER 24.46% / 87.98% 24.46% / 101.50% 20/20 / 0/20
Georgian / WER 27.61% / 111.55% 27.61% / 104.79% 20/20 / 17/20
Khmer / CER 13.51% / 104.70% 13.51% / 103.95% 20/20 / 19/20
Kannada / WER 37.82% / 75.63% 41.18% / 75.07% 19/20 / 20/20
Lao / CER 20.23% / 100.87% 39.01% / 101.49% 14/20 / 0/20
Malayalam / WER 38.46% / 110.49% 38.46% / 110.14% 20/20 / 18/20
Mongolian / WER 32.23% / 101.93% 32.23% / 101.93% 20/20 / 20/20
Marathi / WER 28.29% / 81.39% 28.54% / 81.39% 19/20 / 15/20
Burmese / CER 15.93% / 100.30% 21.55% / 100.30% 19/20 / 19/20
Nepali / WER 35.83% / 95.28% 35.83% / 94.44% 20/20 / 19/20
Pashto / WER 43.29% / 92.63% 44.23% / 95.09% 19/20 / 10/20
Somali / WER 41.30% / 100.97% 41.30% / 103.38% 20/20 / 15/20
Tajik / WER 13.40% / 75.93% 13.40% / 117.87% 20/20 / 0/20
Uzbek / WER 24.87% / 88.86% 26.42% / 109.59% 19/20 / 0/20
Yoruba / WER 51.95% / 95.24% 51.95% / 94.59% 20/20 / 19/20

Automatic language selection is correct on 309/320 versus 211/320. The local denominator includes nine wrong labels and two failed jobs. Lao remains 14/20, and absolute text error remains high for several languages, including Yoruba and Somali. Local completed results have no word/encoding structure violations; 47 explicit and 57 automatic provider responses have at least one such violation. A structurally valid response is not necessarily accurate.

Full API latency, completed jobs Sandchest explicit AssemblyAI explicit Sandchest automatic AssemblyAI automatic
p50 0.651s 3.264s 0.525s 3.665s
p95 0.784s 5.122s 1.027s 5.281s
p99 0.800s 8.515s 1.037s 6.839s

AssemblyAI returned Universal-2 for all 320 explicit responses and 300 Universal-2 / 20 Universal-3.5 Pro for automatic responses. Both providers received the same canonical PCM and equivalent mode options. Serial local M4 Max versus remote-service timings include upload, queue, inference, local persistence and polling, but exclude startup, real remote database/billing, concurrent load and deployment effects. Light corpus-collection and harness-check activity occurred on the local host; this is not an idle-host or matched-hardware benchmark. Failed local jobs took 4.845s and 3.831s and are excluded from completed-job latency, not from accuracy.

The earlier 12 inserted-pause word crossings remain open and are not reclassified by this text validation. Multilingual punctuation, number formatting, code switching and manually annotated word timing remain separate gaps. No deployment, default change or all-language beta approval follows from this comparison.

Private immutable evidence: mms-broad-fresh-validation1 and mms-broad-fresh-validation-api1/{runtime-plan.json,api-audit.json,report.json,comparison.json}. An initial scorer invocation used a relative directory and stopped before producing comparison results; its plan is preserved in relative-invocation1. The unchanged scorer then completed using the absolute directory. No API request or result was changed by that correction.

Haitian Creole, Latin and Sanskrit — first paired baselines

60 distinct development recordings now have 240 actual paired explicit/automatic API results on the same canonical PCM. This raises historical reference coverage to 98/99 base languages, not 98 quality passes. The unchanged 997048bf… candidate uses existing routes; no Haitian, Latin or Sanskrit recognizer was added or tuned for this run.

Language / mode Sandchest WER AssemblyAI WER Sandchest CER AssemblyAI CER Correct auto labels local / provider
Haitian Creole explicit 63.99% 65.03% 19.92% 18.33%
Haitian Creole auto 66.08% 90.91% 22.99% 58.37% 16/20 / 3/20
Latin explicit 69.23% 65.55% 16.83% 16.29%
Latin auto 70.57% 65.55% 17.90% 16.65% 11/20 / 14/20
Sanskrit explicit 99.52% 95.67% 38.07% 27.00%
Sanskrit auto 116.83% 88.94% 56.82% 24.51% 1/20 / 7/20

Sandchest completes 119/120 local jobs, AssemblyAI 120/120. One explicit Sanskrit request fails, remains in the error denominator and creates no usage charge; the 119 successful local jobs have exactly 119 duration-matched usage events. All 120 local jobs have one attempt. All owned processes exit and temporary keys are revoked. No local completed response has a word/encoding structure violation; 15 explicit and 11 automatic Sanskrit provider responses do. Structural validity does not establish recognition quality.

Haitian automatic WER improves against the provider on this sample, but both explicit results are poor and their paired bootstrap interval spans zero. Latin is worse numerically in both modes, with intervals spanning zero. Sanskrit automatic recognition and language detection are materially worse; its paired bootstrap interval favors AssemblyAI. These are gaps to address, not parity passes. High whitespace WER with lower CER can reflect orthography and word segmentation; the exact references and both metrics are retained rather than changing normalization to favor a model.

The source limitations differ: Haitian has 20 recordings but only one publisher speaker ID; Latin has 20 publisher speaker IDs and potentially automated macronization of reference text; Sanskrit uses 20 publisher scholar/recitation_corpus training rows with unknown speaker IDs, excluding synthetic, automatically graded and protected test rows. These are publisher reference transcripts, not independently audited Cap speech, verified people or manual word-time annotations. Model-training overlap is unknown.

Local full API medians are 0.779s/0.773s explicit/auto for Haitian, 0.772s/0.771s for Latin and 1.030s/1.027s for Sanskrit; remote medians are 2.076s/2.802s, 2.406s/3.122s and 3.038s/3.691s respectively. These are completed-job serial local-M4-versus-remote measurements, not matched hardware, load or production latency. Explicit AssemblyAI responses all use Universal-2; auto responses include Universal-3.5 Pro fallbacks and retain their actual model labels.

Private evidence: latin-development1, remaining-catalogs2 and remaining-languages-api1/{runtime-plan.json,api-audit.json,report.json,comparison.json}. The two failed initial collection attempts remain preserved with decoder diagnoses; they are not extra recordings. Turkmen remains unmeasured: the official Common Voice source needs authenticated download access not configured locally, while another public source lacks sufficiently clear reference provenance.

Haitian and Latin native recognition follow-up

A separate, unpromoted MMS prototype reuses exactly the 40 Haitian/Latin development recordings and their unchanged reference text. Four CPU-reference/native numerical pairs pass, with exact decoded text and frame labels; maximum logit difference is 0.000442. All 40 recognition calls complete, invalid adapter/nonfinite input is rejected, and recovery is exact. No new provider requests or independent recordings are added.

Language Raw MMS WER Current Sandchest explicit API WER Cached AssemblyAI explicit API WER
Haitian Creole 26.57% (76/286) 63.99% (183/286) 65.03% (186/286)
Latin 72.58% (217/299) 69.23% (207/299) 65.55% (196/299)

Haitian CER is 4.98%, versus 19.92% current and 18.33% provider. This is a promising recognizer candidate, but the corpus still has only one publisher speaker ID. Those were open gates at the raw-prototype checkpoint; the later serving/API/validation work is below. Manual word-time accuracy remains unverified. Latin CER also regresses (19.10% versus 16.83%/16.29%), so that candidate is rejected. No recognition route changes from this screen.

Raw recognition medians are 94ms Haitian and 128ms Latin; these exclude word alignment, formatting and the API, and must not be advertised as end-to-end latency. Private evidence: remaining-mms1/{plan.json,numerical-report.json,controls.json,report.json,decision.json} and pinned assets/probe build in durable1/mms-remaining1. Historical source hashes are verified against preserved snapshots rather than silently replacing old receipts with the current source.

Haitian serving candidate — two-source API comparison

A private Rust candidate (230df53c852ef7104dfdce78d5877a1c13fd108c29114ca44f54eb4b0011076a) adds the exact Haitian MMS adapter, acoustic word alignment, and an automatic handoff only when two audio language detectors agree at the unchanged 0.8 threshold. The new handoff also respects the caller's expected-language list. This exact candidate is now integrated after the validation and control checks below; no all-language beta is approved.

Development source / mode Candidate WER Prior B5 WER AssemblyAI WER Correct auto labels: candidate / B5 / provider
CMU, 14 conservative filename groups / explicit 13.97% (56/401) 48.88% (196/401) 53.87% (216/401)
CMU, 14 conservative filename groups / auto 13.97% (56/401) 48.88% (196/401) 76.31% (306/401) 20/20 / 20/20 / 3/20
Zilora, one publisher speaker ID / explicit 26.57% (76/286) 63.99% (183/286) 65.03% (186/286)
Zilora, one publisher speaker ID / auto 36.01% (103/286) 66.08% (189/286) 90.91% (260/286) 18/20 / 16/20 / 3/20

There are 40 distinct Haitian recordings, 20 reused from the original development cohort and 20 newly collected CMU clips totaling 3.50 minutes. The broader current-B5/provider baseline makes 80 actual API jobs; the candidate then makes 80 local API jobs and reuses the exact matched provider/B5 results. All 160 new jobs complete. The 120 local jobs have one attempt and exactly 120 measured-duration usage events; temporary keys are revoked and owned API processes stop. Every candidate completed response has valid word/encoding structure. Paired-recording WER intervals favor the candidate in all four source/mode groups against both baselines, but are development estimates, not independent-speaker population guarantees.

CMU candidate explicit/auto CER is 6.12%, versus B5 12.32% and provider 13.76%/31.84%. Zilora candidate CER is 4.98%/8.16%, versus B5 19.92%/22.99% and provider 18.33%/58.37%. All 33 automatic MMS handoffs have the Haitian label; the seven fallback responses include two remaining wrong labels. No references, normalization rules or thresholds were altered to obtain these gains.

Candidate full-API medians are 0.519s explicit / 0.774s auto on CMU, versus remote 2.376s / 2.878s; Zilora is 0.521s / 0.524s, versus cached remote 2.076s / 2.802s. CMU candidate auto p95 is 1.036s, slower than the B5 baseline's 0.788s even though its median is similar; the additional detector is not free. These serial local-M4 versus remote timings exclude cold startup and are not production/load/capacity proof.

The CMU source is pinned CMU Haitian Creole Speech, revision 4569351079203f9188abf62f8c7a2371930df35f. Its speaker_id equals an utterance identifier, so the collector uses conservative filename-prefix groups and never counts those IDs as verified people. Literal published text is preserved. Five successfully retrieved metadata pages yield eligible recordings; an unavailable later page and rows with unreviewed filename patterns remain documented exclusions. Training overlap, naturally occurring Cap-domain performance and manual word times are unknown.

The preceding explicit-only integration reproduces all 20 raw recognition texts and 1,158 retained native outputs, including Cap. Four derived long controls reuse those recordings with exact inserted pauses and show zero words inside or spanning those pauses; these are not new recordings or manual timing gold. The automatic candidate passes 122 base / 147 Whisper / 171 CTC / 230 MMS tests and release compilation. Its final automatic replay preserves all 1,158 retained outputs, all 20 explicit Haitian targets, and seven unchanged fallbacks; 13 new automatic handoffs preserve exact explicit text and word intervals. All eight explicit/auto derived-long responses have no inserted-pause crossings. The option/deadline and unused-validation gates are detailed below.

Private evidence: haitian-native2, haitian-lid1, cmu-haitian-development4, cmu-haitian-baseline1, haitian-api1; candidate build durable1/haitian-serving-build4. Interrupted metadata/compiler/comparator attempts are retained with their original failure receipts; none is reported as a successful test. Only the six verified native source files are integrated; shared services, Cap, environment/defaults and deployment remain unchanged.

Haitian frozen validation and final controls

The candidate and acceptance criteria were frozen before any validation inference. The bounded source scan exposed only five unused conservative filename groups, so the original 20-clip plan was retained as incomplete and a separate 10-clip, two-per-group validation was collected. Prior source IDs, normalized development references, audio and all 14 CMU development filename groups were excluded. The ten clips total 1.40 minutes. Publisher split remains train; these are unused local validation recordings, not proven independent model-training data or verified distinct people.

Validation mode Sandchest WER AssemblyAI WER Sandchest CER AssemblyAI CER Auto labels: local / provider
Explicit 17.03% (31/182) 56.04% (102/182) 6.92% (45/650) 15.23% (99/650)
Automatic 17.03% (31/182) 84.62% (154/182) 6.92% (45/650) 45.08% (293/650) 10/10 / 3/10

All 40 actual paired API jobs complete. Each of the 20 local jobs has one attempt and one exact-duration usage event; no local word/encoding structure defect occurs. Explicit and automatic paired-recording WER intervals favor the candidate (95% upper differences −31.36 and −51.31 percentage points). These are descriptive ten-recording intervals, not population guarantees. Automatic detection meets the predeclared minimum nine correct labels. No model, detector threshold, reference or normalization was changed after validation.

Local full-API medians are 0.524s explicit / 0.773s automatic, versus remote 2.103s / 2.783s. Local automatic p95/p99 is 0.921s/1.012s versus 4.042s/4.327s remotely. Small serial local-M4 versus remote measurements do not establish cold-start, concurrent or production-tail behavior.

Final controls pass 198 logical API cases and 198 same-ID replays, including 50 unchanged Cap responses, caller hints (including explicit exclusion of Haitian), confidence errors, exact 1001ms crops, custom spelling, flags and stereo. A first crop fixture rounded 4249.25ms to 4249ms and correctly removed four samples, changing two confidences while text and word offsets stayed exact. That original completed probe remains visible; a separate whole-millisecond fixture passes complete text/word/confidence equality. Actual accounting is 199 jobs, 171 completed usage events, 28 expected unbilled errors, all one attempt. Two real one-second native timeouts recover exactly into automatic Haitian. Keys are revoked and all owned processes are stopped.

The source cutover changes exactly six files, with all 103 native source hashes matching the tested 230df53c… binary. Prior source bytes are preserved for historical evidence verification. Shared services, Cap, environment, defaults, HEAD, index and deployment are unchanged. Private proof: haitian-auto1, haitian-options1 (including api-recovery), cmu-haitian-validation2, haitian-validation-api1, and durable1/haitian-source-checkpoint1/receipt.json.

Fresh isolated checks on the integrated source pass 194 application tests with seven conditional model skips, 2,378 assertions, typecheck, full lint, 56 collector tests and five detector-preparation tests. All 290 application/native snapshot files match current source; no shared services are used.

This closes a measured Haitian recognition/detection improvement, not global language parity. Multilingual punctuation/numbers, fillers, code switching, confidence calibration, most manual word-time gold, broader Cap validation and the other documented language failures remain open.

Full Whisper follow-up for Latin and Sanskrit — not promoted

An isolated copy of the current B5 worker changes only the pinned Whisper model to full large-v3, its architecture-specific word-timing preset/layer check, model identity and matching test assertions. It passes the 122 base / 147 Whisper / 171 CTC / 228 MMS test suites. The actual API study reuses 40 development recordings and runs 80 new local explicit/auto jobs, comparing the same PCM/references/options with the cached B5 and AssemblyAI responses. No new provider submission or independent recording is added.

Language / mode Full-v3 candidate WER Current Turbo-path WER Cached AssemblyAI WER Correct auto labels full / current / provider
Latin explicit 65.55% 69.23% 65.55%
Latin automatic 68.90% 70.57% 65.55% 14/20 / 11/20 / 14/20
Sanskrit explicit 92.31% 99.52% 95.67%
Sanskrit automatic 117.79% 116.83% 88.94% 4/20 / 1/20 / 7/20

The candidate completes 78/80, versus the baseline's cached 79/80 and provider's 80/80. Its two Sanskrit failures remain in the accuracy denominator and are unbilled. All 80 jobs have one attempt; the 78 successful jobs have exactly 78 duration-matched usage events. Temporary keys are revoked and owned processes are stopped. No completed local response has a word/encoding structure defect.

Explicit Sanskrit CER improves from 38.07% to 20.77%, versus provider 27.00%, but automatic CER worsens from 56.82% to 78.83%, versus provider 24.51%. Paired-recording WER intervals include zero for Latin and explicit Sanskrit; automatic Sanskrit remains worse than the provider. These small, limited-source development results do not justify a global model replacement.

Completed-job API medians are 1.280s/1.032s for Latin explicit/auto and 2.296s/1.030s for Sanskrit. Sanskrit explicit p95/p99 rises to 5.384s/18.384s, versus baseline 1.039s/1.040s and cached provider 3.554s/5.059s. These serial local-versus-cached-remote timings are not a fresh matched-load comparison. The global full-model replacement is rejected; serving remains the verified opt-in B5 worker.

A separate, explicitly post-hoc Latin spelling diagnostic removes only vowel macrons from both reference and output, leaving the original scores above untouched. The publisher references contain 185 macrons and may include automatic macronization. Under that secondary spelling rule, explicit WER is 31.10% full / 36.12% current / 36.79% provider; automatic WER is 38.13% / 40.13% / 35.79%. The rejected Latin MMS candidate remains worse at 43.81%. This explains some orthographic sensitivity but does not redefine primary recognition accuracy, prove phonetic correctness or close a beta gate.

Private receipts: full-remaining-api1/{runtime-plan.json,api-audit.json,comparison.json,comparison-recovery.json,decision.json} and latin-orthography-diagnostic1.json. A syntax error in the original private comparison script occurred after the API audit completed; its source/log remain preserved. A separate corrected scorer reproduced all primary scores from the unchanged responses without retrying any API request.

Twenty-language development recognition screen

The serving candidate remains edbef6a2…; none of the new MMS candidates in this section is promoted. The screen reuses 100 recordings: 60 Hindi development clips, four Tibetan clips and two each in 18 other languages. Protected Hindi holdout and all later validation references remain unopened. Current-worker explicit requests, raw native MMS requests and cached AssemblyAI responses use identical audio. The current worker reproduces all 100 earlier text/completion outcomes; this is not 100 new independent recordings.

All 21 numerical probes pass against the pinned CPU reference, including one per adapter and a longer Hindi probe. The maximum logit difference is 0.003302, every frame argmax agrees, and decoded text matches exactly. The native MMS prototype completes 100/100 recognition requests; the current worker completes 88/100. Its 12 errors are four Tibetan and two each Georgian, Khmer, Malayalam and Burmese. Cached AssemblyAI actually completed 100/100; one legacy Kannada encoding failure was rescored from its preserved completed response, reproducing the existing corrected harness result. Encoding/word defects remain visible. Invalid-adapter/nonfinite controls reject input and recover exactly.

MMS lowers the harness's primary error metric in 19/20 groups versus the current worker and 17/20 versus AssemblyAI. Hebrew is worse than the current worker; Hebrew, Welsh and Hindi remain behind AssemblyAI. Most groups have only two examples, so these comparisons select candidates for larger development studies, not language-level winners or parity passes. Failed transcripts remain in the denominator. CER is primary for Khmer, Lao and Burmese; both WER and CER remain in the private report. All scores use the same Unicode-preserving harness normalizer.

Language Clips Metric Current worker Raw native MMS candidate Cached AssemblyAI
Bengali (bn) 2 WER 80.49% (33/41) 24.39% (10/41) 102.44% (42/41)
Tibetan (bo) 4 WER 100.00% (148/148) 58.11% (86/148) 100.00% (148/148)
Welsh (cy) 2 WER 42.19% (27/64) 32.81% (21/64) 20.31% (13/64)
Hausa (ha) 2 WER 178.12% (57/32) 34.38% (11/32) 115.62% (37/32)
Hebrew (he) 2 WER 66.67% (20/30) 73.33% (22/30) 60.00% (18/30)
Hindi (hi) 60 WER 32.52% (479/1473) 22.00% (324/1473) 19.42% (286/1473)
Georgian (ka) 2 WER 100.00% (28/28) 14.29% (4/28) 107.14% (30/28)
Khmer (km) 2 CER 100.00% (289/289) 15.22% (44/289) 99.31% (287/289)
Kannada (kn) 2 WER 54.55% (12/22) 27.27% (6/22) 68.18% (15/22)
Lao (lo) 2 CER 100.00% (161/161) 18.01% (29/161) 98.76% (159/161)
Malayalam (ml) 2 WER 100.00% (26/26) 46.15% (12/26) 107.69% (28/26)
Mongolian (mn) 2 WER 103.12% (33/32) 25.00% (8/32) 103.12% (33/32)
Marathi (mr) 2 WER 81.82% (36/44) 50.00% (22/44) 90.91% (40/44)
Burmese (my) 2 CER 100.00% (284/284) 4.58% (13/284) 100.00% (284/284)
Nepali (ne) 2 WER 77.78% (42/54) 20.37% (11/54) 88.89% (48/54)
Pashto (ps) 2 WER 98.08% (51/52) 51.92% (27/52) 96.15% (50/52)
Somali (so) 2 WER 188.57% (66/35) 40.00% (14/35) 100.00% (35/35)
Tajik (tg) 2 WER 76.32% (29/38) 15.79% (6/38) 65.79% (25/38)
Uzbek (Latin) (uz) 2 WER 114.71% (39/34) 41.18% (14/34) 100.00% (34/34)
Yoruba (yo) 2 WER 106.90% (31/29) 48.28% (14/29) 96.55% (28/29)

The larger Hindi set improves from 479/1,473 errors (32.52%) to 324/1,473 (22.00%), but AssemblyAI remains better at 286/1,473 (19.42%). Hindi CER is 14.32% current, 8.40% MMS and 7.85% AssemblyAI. No numbers, spelling or reference text are supplied to recognition. Uzbek uses only the Latin-script adapter; Nepali uses the publisher's npi adapter. Neither mapping establishes every script or language variety.

Raw MMS timing excludes word alignment, formatting, upload, queue and persistence, so it is not comparable to full API latency. Native timings are retained only as diagnostics; read-only preparation/scoring overlapped parts of the overall study. These prototypes still need larger samples, automatic routing, segmentation and alignment checks, long audio, fault/recovery tests and frozen validation before integration. The generic alignment language preflight must also accept each exact route; adapter availability alone is insufficient, especially for script-specific adapter names and languages without spaces.

The first numerical attempt stopped before native recognition because its reference-tokenizer checker only allowed an unused separator for Tibetan. The corrected checker verifies the same out-of-head separator rule for Khmer, Lao and Burmese, without resizing or remapping trained vocabularies. A separate report audit corrected the historical Kannada completed-text classification; no audio, model, native output, denominator or provider request changed. Original attempts remain preserved.

Private receipts: broad-mms-development2/{plan.json,numerical/report.json,collection-receipt.json,audited-report.json,scoring-audit-plan.json}; the stopped numerical attempt remains in broad-mms-development1. Pinned adapters and native probe build are under whisper-large-20260827/durable1/mms-broad1. No serving source, default, environment or deployment changed.

Faroese — first reference and API measurements

Twenty previously unmeasured Faroese recordings now have a full explicit/automatic API comparison. The public Ravnursson corpus is pinned to 03665210706cf1f49b068b3c1188942484964b05; four bounded publisher train windows produce 20 unique PCM recordings across 13 conservative speaker-prefix groups, with at most two clips per group. Session IDs are not counted as separate people. References are publisher-normalized human read-speech text, with no manual word times or known model-training exclusion. This is development evidence, not unseen validation or representative Cap speech.

Mode Sandchest WER AssemblyAI WER Sandchest CER AssemblyAI CER
Explicit 96.35% (185/192) 92.71% (178/192) 46.25% (438/947) 46.04% (436/947)
Automatic 100.52% (193/192) 101.56% (195/192) 44.67% (423/947) 44.35% (420/947)

Neither provider identifies Faroese in any of the 20 auto requests. Sandchest completes 39/40 jobs, all in one attempt, with exactly 39 duration-matched usage events; the one failed explicit job is not billed. AssemblyAI completes 40/40. Local completed responses have zero word/encoding structure defects; two provider explicit responses have defects. These structural checks do not measure acoustic timestamp accuracy.

The subsequent dedicated Faroese MMS prototype completes all 20 raw recognition requests and improves explicit recognition to 51.56% WER (99/192) and 14.47% CER (137/947). Two CPU/native numerical probes pass, with maximum logit difference 0.000516 and identical decoded text; invalid inputs reject and recovery is exact. This is promising but still high absolute error. The prototype is not integrated, has no verified automatic route or API word timestamps, and must not be presented as the current serving result.

The paired API run made 40 new AssemblyAI submissions, totaling 5.0816 provider audio minutes. Local API p50 is 0.773s explicit and 0.647s auto, versus 2.092s and 3.073s remotely. These are local-versus-remote diagnostics, exclude startup, and overlapped some read-only scoring; they are not dedicated speed, load or production-tail acceptance results. The MMS recognition-only median is 0.128s and has a different scope.

The reusable Faroese collector checks source/revision/split/row URLs, frozen row windows, reference/audio duplicates, duration/transfer budgets and speaker-group limits. For validation/holdout collection it excludes supplied protected speaker groups without opening protected reference text. Seven focused tests pass. The whole isolated app checkpoint passes 189 tests, seven conditional skips, typecheck and lint, plus 36 collector tests and five detector-preparation tests. Native serving sources still match the prior verified checkpoint.

At that Faroese checkpoint, reference coverage became 95/99, not 95 quality passes. The later Haitian/Latin/Sanskrit measurements above raise coverage to 98/99; Turkmen, regional-alias evidence, manual timing and broader Cap validation remain open. No all-language beta approval follows from these experiments.

Private evidence: faroese-development1/{selection-plan.json,runtime-plan.json,report.json,api-audit.json}, faroese-mms1/{plan.json,numerical-report.json,controls.json,report.json}, and broad-mms-app-checks1. All are under the ignored benchmark root; audio, transcripts, signed URLs and credentials are not published here.

Bashkir and Tatar — current secondary detector

The verified Rust/MLX binary is edbef6a2b60bd5ad30ce8eb2c92869f349728868b4fc7181b5beddedbf14cd79. Its six native source changes are integrated, with serving defaults and deployment unchanged. The optional MMS-LID126 classifier runs only when the primary VoxLingua detector predicts Bashkir, Tatar or Kazakh. A secondary Tatar handoff requires both detectors to agree confidently; secondary Bashkir cannot override a confident primary Tatar prediction. Otherwise the prior routing behavior remains. Both confidence gates stay at 0.8, with the existing positive-speech check and maximum 30-second cropped prefix.

The classifier has its own pinned 3,864,495,808-byte checkpoint and shares the native encoder implementation with MMS ASR. It retains all 126 output classes, mapping the 99 API codes without renormalizing away the other classes. Explicit requests bypass language detection. Language hints, fallbacks, confidence thresholds, caller deadlines and acoustic alignment remain in the existing request path. This does not enable automatic handoffs for every model language.

Development selection

Six CPU-reference/native probes, including a 30-second input, agree on the top class. Maximum logit difference is 0.00003815 and maximum probability difference is 0.00000757. After factoring the shared encoder, all six probes reproduce the earlier native tensors and probabilities exactly. Numerical agreement validates implementation math, not language quality.

The detector screen uses 815 reused development/control inputs, including short and derived ML-SUPERB audio, public development, retained English/Spanish/FLEURS, three Cap provider-label controls, two nonspeech controls and 40 Malagasy/Albanian development inputs. These are not 815 independent, human language-annotated recordings. The selected primary-label gate invokes the second detector on 48 inputs; the conservative policy selects 32 correct Bashkir/Tatar handoffs and no incorrect new handoff in this screen. A more aggressive policy improved development text further but produced four wrong language labels, so it was not selected.

Original clips, 20 per language Previous auto WER Integrated auto WER Cached AssemblyAI auto WER Correct auto labels: local / provider
Bashkir 90.35% (103/114) 43.86% (50/114) 112.28% (128/114) 13/20 · 0/20
Tatar 106.72% (143/134) 55.22% (74/134) 105.22% (141/134) 11/20 · 0/20

Explicit recognition remains 18.42%/25.37% WER for Bashkir/Tatar. On four derived joins per language, auto WER is 21.05%/32.84%, with all eight labels correct; the joins reuse the same originals and are not independent long-form examples. The first integration completes 192/192 local API jobs across Bashkir, Breton, Nynorsk and Tatar, all with one attempt and one exact-duration usage event. All 144 explicit or Breton/Nynorsk automatic responses preserve the prior API response, excluding IDs, URLs and creation time; all 48 Bashkir/Tatar automatic responses match the verified native outputs. Word/encoding structure defects are zero.

Frozen validation on unused recordings

The final candidate and policy were frozen before collecting 40 unused ML-SUPERB recordings: 20 per language, with source, normalized development-text, audio and canonical-PCM exclusions. These are publisher dev data used as local validation, not independent model-training data. Speaker IDs are unavailable. Bashkir clips are 4.500–9.756s; Tatar clips are 4.104–8.748s. The short duration and read-speech domain limit generalization. There are no manual word-time references.

Language / mode Sandchest WER AssemblyAI WER Sandchest CER AssemblyAI CER Correct auto labels
Bashkir explicit 10.14% (14/138) 98.55% (136/138) 1.66% (12/724) 45.17% (327/724)
Bashkir auto 32.61% (45/138) 104.35% (144/138) 23.48% (170/724) 88.40% (640/724) 14/20 · 0/20
Tatar explicit 24.03% (31/129) 96.90% (125/129) 4.21% (31/737) 49.80% (367/737)
Tatar auto 46.51% (60/129) 106.98% (138/129) 25.64% (189/737) 97.29% (717/737) 9/20 · 0/20

All 160 paired API jobs complete, including 80 fresh AssemblyAI submissions. The 80 local jobs each have one attempt and one exact-duration usage event. No local word/encoding structure defects occur; one provider explicit Tatar response has a structural defect. All 23 local auto results using MMS have the correct source language; the other 17 retain Whisper fallback and have incorrect labels. This is not proof of perfect detector precision outside this small validation set. No model or policy was tuned using these validation results.

Full API p50 Local explicit Remote explicit Local auto Remote auto
Bashkir 0.530s 2.189s 0.773s 2.540s
Tatar 0.529s 2.127s 0.774s 2.473s

Local/remote auto p95 is 1.033/3.012s for Bashkir and 1.359/3.526s for Tatar. Timing includes upload, queue, recognition, alignment, persistence and polling; excludes model startup; and compares an Apple M4 Max with the remote service. It is not matched hardware, sustained load, cloud capacity or production-tail evidence. The final diagnostic fix reports the combined cost of both detectors and retains each individual timing.

Regression and reliability checks

  • Initial integrated build 1f6c9c94880cdf98fe999e7fbdbcc7eaf067cd0094c675bcbf29c896aad3efb5: 59 flag-disabled checks match the previous worker; 469 retained native checks, including 50 Cap recordings and 80 Malagasy/Albanian variants, remain exact; 48 target auto outputs match the selected policy. The 192-job development API replay above uses this build.
  • The final binary changes only timing diagnostics from that build; its recognition, mapping, routing and tensor code are unchanged. It passes the fresh validation above and 132 API compatibility cases with 132 same-ID replays, including 50 exact Cap responses and new detector hints, fallback, threshold, crop, spelling, formatting and stereo cases.
  • The logical compatibility cohort has 112 completions and 20 expected unbilled errors. An interrupted test also left one completed stereo probe, preserved separately. Both isolated databases reconcile 133 actual jobs, 113 completions/unique usage events and 20 unbilled errors, all in one attempt; no completion is hidden or charged twice within a job.
  • Two real one-second native deadlines each recover exactly into automatic Tatar. The corrected combined-detector timing is checked on a real response. This is not sustained-load or every-internal-stage cancellation proof.
  • Final Rust suites pass 122/147/171/215 tests across base/Whisper/CTC/MMS configurations. All 29 offline collector tests pass, including validation/holdout roles. Fresh isolated app checks after promotion pass 189 tests with 7 conditional skips, typecheck, lint, 29 collector tests and 5 detector-preparation tests; the source snapshot matches the promoted code.

Two test assertions were corrected without changing recognition: the detector diagnostics are top-level in the native response, and converting 4.068 seconds to milliseconds has binary floating-point roundoff. Original failed/partial attempts remain preserved; the recovery reports distinguish logical test cases from actual jobs. These were harness issues, not failed transcription or incorrect usage accounting.

Private evidence: mms-lid126-development1, mms-lid126-policy-native1, secondary-lid-route-api1, secondary-lid-options1, secondary-lid-fresh-validation1 under the benchmark root. The source promotion receipt is .sandchest/experiments/whisper-large-20260827/durable1/secondary-lid-source-checkpoint1/receipt.json. No shared database, Cap source/data, environment, deployment or production traffic was changed. That checkpoint had five unmeasured languages; Faroese reduced the gap to four, and the newer Haitian/Latin/Sanskrit measurements leave Turkmen as the single base-language reference gap. Weak fallback recognition, Hindi/Nynorsk quality, broader Cap validation and manual word timing still need work.

Malagasy — current native route

The retained Malagasy checkpoint used Rust/MLX binary d926ed31fd6fe9f568628cd67fb13f058c9ccd1a056b39ebe2150903e24e1a1e; the latest candidate preserves this route. Its six source changes are integrated. The pinned MMS mlg adapter is used for explicit Malagasy and guarded automatic handoffs at the unchanged 0.8 confidence threshold with positive speech detection. Serving defaults and deployment are unchanged. It uses the existing acoustic alignment and bounded segmentation, with no new numerical kernels or SDK changes.

A private recognition screen completed all 40 reused Malagasy/Albanian development recordings. Four CPU-reference/native probes agree on decoded text and all 396 frame decisions; maximum logit difference is 0.003872. Malagasy improves to 355/716 word errors. Albanian worsens to 753/1,074 versus the existing 708/1,074, so its route is deliberately unchanged. These native model timings exclude alignment/API work and are not service latency.

Reused development data

Plateau Malagasy, 20 clips Earlier Sandchest WER Current Sandchest WER Cached AssemblyAI WER Current local completion Correct auto labels: local / provider
Explicit 104.19% (746/716) 49.58% (355/716) 100.42% (719/716) 20/20
Automatic 97.21% (696/716) 49.58% (355/716) 100.70% (721/716) 20/20 20/20 · 0/20

All seven former local failures now complete. Across both languages, 80/80 fresh local API jobs complete in one attempt with 80 unique exact-duration usage events and no word/encoding structure defects. All 40 Albanian explicit/auto responses preserve their recognition, confidence and word metadata. No new provider submissions or independent examples were added by this replay. Current Malagasy CER is 12.43%; local explicit/auto medians are 776/774ms.

Frozen validation on unused recordings

The candidate and policy were frozen before collecting 20 unused Plateau Malagasy recordings, excluding prior source IDs, normalized development text, audio hashes and canonical PCM. There are four publisher speaker IDs, five clips each, and all four also occur in development. These are publisher train recordings used as local validation, not unseen model-training data or independent speakers. Physical duration is 6.984–26.288s; 16/20 are at least 15s, which does not guarantee that much speech. No manual word boundaries are supplied.

Mode Sandchest WER AssemblyAI WER Sandchest CER AssemblyAI CER Completed local / provider Correct automatic labels
Explicit 46.58% (279/599) 96.66% (579/599) 12.39% (390/3,148) 47.49% (1,495/3,148) 20/20 · 20/20
Automatic 46.58% (279/599) 105.18% (630/599) 12.39% (390/3,148) 72.87% (2,294/3,148) 20/20 · 20/20 20/20 · 0/20

All 80 paired requests complete before reference scoring, with identical audio/options and 40 fresh provider submissions. Local completions use MMS; provider explicit results use Universal-2, while auto returns Universal-2 or Universal-3.5 Pro. Local word/encoding structure defects are zero; provider defects occur in two explicit and four auto responses. Structural validity is not acoustic timestamp accuracy. The high absolute WER and narrow variety/speaker coverage remain limitations despite the measured improvement.

Full API latency Local explicit Remote explicit Local automatic Remote automatic
p50 0.778s 2.743s 0.775s 3.485s
p95 0.821s 4.935s 1.031s 5.099s

Timing includes upload, submit, durable queue, native recognition/alignment, persistence and polling locally; excludes model startup; and compares Apple M4 Max hardware with the remote service. This is not matched hardware, sustained load, cloud capacity or production-tail evidence. The separate 50-Cap local replay retains all responses, with 45 speech completions and p50/p95/p99 of 1.645/8.170/16.688s. No fresh remote Cap latency was measured.

Integration verification

  • 389 exact retained native comparisons, including 50 Cap recordings, 100 existing MMS cases and the previously added ba/br/tt routes. Explicit/auto Malagasy smoke requests pass before the bulk run.
  • 112 API controls and 112 idempotent replays: 94 completions, 18 expected unbilled errors, exact physical/channel usage and all 50 Cap responses unchanged. New Malagasy crop, spelling, flags, stereo and language-hint cases pass.
  • Two real one-second native deadlines recover exactly into Malagasy.
  • Rust suites pass 122/147/171/211 tests across base/Whisper/CTC/MMS; isolated app checks pass 189 tests with seven conditional skips, 28 collector tests, five preparation tests, typecheck and lint.
  • No wrong new Malagasy handoff on 815 development/control inputs. Those are reused/derived inputs, not 815 independent validation recordings. The new validation does not tune the policy.

Private receipts: variety-mms1, malagasy-route-api1/{comparison.json,api-audit.json,regression/report.json}, malagasy-route-options1/{api/report.json,deadline/report.json}, and malagasy-fresh-validation2/{selection-plan.json,report.json,api-audit.json,validation-readback.json} under the benchmark root. The source promotion receipt is .sandchest/experiments/whisper-large-20260827/durable1/malagasy-source-checkpoint1/receipt.json. Preparation failures that ran no models or collected no data remain preserved; their corrected attempts do not count as extra evidence.

Retained three-language integration

The earlier integrated Rust/MLX binary was 23c1d82e3c9cba1eff7619f74446f33515cfd249e81d5659995139208f91887f. Its six source changes are integrated; serving defaults and deployment are unchanged. Explicit Bashkir, Breton and Tatar now use MMS recognition with the existing acoustic word-alignment pipeline. Nynorsk retains Whisper because its recognition experiment regressed on longer controls.

192/192 fresh local API jobs complete in one attempt, with 192 exact-duration usage events and no word/encoding structure defects. These replay the same 80 ML-SUPERB development recordings and 16 derived controls used below. The 192 recent AssemblyAI results are cached, not resubmitted; both studies use identical audio and request options. This is development evidence, not newly unseen validation or manual word-time accuracy.

Language, 20 original clips each Current explicit WER Cached AssemblyAI explicit WER Current auto WER Cached AssemblyAI auto WER Correct auto labels: Sandchest / AssemblyAI
Bashkir 18.42% (21/114) 101.75% (116/114) 90.35% (103/114) 112.28% (128/114) 3/20 · 0/20
Breton 58.75% (94/160) 104.38% (167/160) 70.00% (112/160) 107.50% (172/160) 12/20 · 3/20
Nynorsk, unchanged 36.47% (62/170) 37.65% (64/170) 51.18% (87/170) 37.06% (63/170) 0/20 · 1/20
Tatar 25.37% (34/134) 96.27% (129/134) 106.72% (143/134) 105.22% (141/134) 0/20 · 0/20

Explicit WER improves from 104.39%, 90.63%, 97.01% to 18.42%, 58.75%, 25.37% for ba/br/tt. Automatic selection remains weak: 15/80 correct labels versus 4/80 provider labels. The raw 0.8 detector screen produced five confident Bashkir-to-Tatar mistakes, so only ba/br receive new detector handoffs. There are no wrong ba/br handoffs among 815 detector inputs: 96 new-language development inputs, 679 cached controls and 40 additional Albanian/Malagasy inputs, before the positive-speech gate. Those include derived and reused samples; they are not 815 independent quality examples or a production error-rate estimate.

On the 16 derived controls, explicit WER is 21.05%/62.50%/35.29%/32.84% for ba/br/nn/tt; all 32 local explicit/auto jobs complete. The formerly failed Bashkir composite now takes the MMS route successfully. This avoids that Whisper decoder failure for this route; it does not resolve the underlying SDK error generally. Derived controls reuse the 80 originals and remain separate from independent recording counts.

Current explicit local API medians are 514–522ms on the original clips, including upload, queue, alignment, persistence and polling. The prior remote medians are 1.89–2.03s. These are separate local/remote runs, exclude startup and are not matched hardware, load, capacity or production-tail measurements. No isolated latency or matched-load benchmark was performed.

Verification: 122/147/171/211 Rust tests across base/Whisper/CTC/MMS configurations; three real new-route smoke requests; 386 exact retained native comparisons including 50 Cap recordings; 104 API controls with 104 idempotent replays, 50 exact prior Cap responses, 86 completions and 18 expected unbilled failures; two real 1s deadlines with exact Tatar recovery. Current isolated app checks pass 189 tests with 7 conditional model skips, 28 collector tests, 5 preparation tests, typecheck and lint.

The first private integration failed because the aligner did not accept the new language IDs. It was not promoted. The fix validates recognition/normalization/alignment compatibility at startup and extends the regression tests. Its failed/partial batch remains preserved. A later post-run auditor also needed a read-back correction for the existing zero-duration no-spoken-language convention; no application behavior, responses or jobs were changed to satisfy it.

Evidence under the private benchmark root: mms-expanded-api2/{comparison.json,api-audit.json,regression/report.json}, mms-expanded-options1/{api/report.json,deadline/report.json}, mlsuperb-lid1, omnilingual-varieties1/lid-screen. Source promotion receipt: .sandchest/experiments/whisper-large-20260827/durable1/mms-expanded-source-checkpoint1/receipt.json.

Gheg Albanian and Plateau Malagasy — first paired screen

This is the preserved baseline before the Malagasy route above. Its failures are historical results, not the current candidate's results.

The same binary completed a fresh 160-job paired study on 40 new public recordings: 20 Gheg Albanian clips from two publisher speaker IDs and 20 Plateau Malagasy clips from five. These are publisher train recordings used for local development, with unknown model-training overlap and no manual word boundaries. Gheg does not establish Tosk/all Albanian dialects; Plateau Malagasy does not establish every Malagasy variety. Physical duration is 13.1–28.8 seconds; 38/40 clips are at least 15 seconds, which does not guarantee 15 seconds of speech.

Variety, 20 clips each Sandchest explicit WER AssemblyAI explicit WER Sandchest auto WER AssemblyAI auto WER Local completed explicit / auto Correct auto labels: Sandchest / AssemblyAI
Gheg Albanian (sq) 65.92% (708/1,074) 69.09% (742/1,074) 65.92% (708/1,074) 69.27% (744/1,074) 20/20 · 20/20 20/20 · 20/20
Plateau Malagasy (mg) 104.19% (746/716) 100.42% (719/716) 97.21% (696/716) 100.70% (721/716) 17/20 · 16/20 0/20 · 0/20

Both systems have high surface word error on these data. Albanian CER is 21.70% locally versus 24.58% explicit/24.78% auto remotely; this small, two-speaker sample does not establish general superiority. At that checkpoint, seven local Malagasy jobs returned an unreliable-transcription error, and neither provider identified any of its 20 auto examples correctly. Failures contribute deletions to WER; a slightly lower failure-inclusive auto WER is not evidence of useful recognition.

AssemblyAI completes 80/80 jobs; Sandchest completes 73/80. All 80 local jobs take one attempt, only 73 completions create usage events, and failed jobs remain unbilled. Completed word/encoding structure defects: 0 local versus 9 provider responses. The additional native detector check predicts both source labels on 40/40 inputs, with 38/40 above 0.8, and makes no new ba/br handoff. Thus the then-new ba/br routes did not explain these failures; the latest Malagasy study above addresses the identified recognition/routing gap.

For Albanian, local explicit/auto medians are 1.035/1.160s versus 4.077/4.368s remotely. Malagasy completion-only medians are 1.040/0.778s versus 3.526/3.969s, but the local completed subset is smaller; those are not matched-pair speed comparisons. No startup, cloud hardware, sustained-load or production-tail conclusion follows. Full outcomes and returned model identities are in omnilingual-varieties1/{runtime-plan.json,report.json,api-audit.json,collection-audit.json}.

Fresh five-language validation — retained model paths

The earlier frozen Rust/MLX candidate (4b86362161929e4f7071d6e24dee508cab2fb35dc7c2dc366250a329ca425b6b) completed 400 fresh paired API results on 100 previously unused recordings, 20 each in five weak languages. These validation results retain their original binary identity; the new implementation retains the five model paths but does not count a new validation rerun. Both providers completed all200jobs. All200local completions have exactly one attempt and one duration-matched usage event; ephemeral keys were revoked and owned processes stopped.

The policy was frozen before collection from the pinned FLEURS published test split, stored as local validation. Prior source IDs, normalized development text, WAV hashes and canonical PCM were excluded. These recordings were unseen to this study; model-training overlap and speaker independence remain unknown. These are read-speech references, not representative Cap recordings or manual word-time labels.

Each provider received identical audio and options in explicit and unrestricted-auto modes, including speech_models: ["universal-3-5-pro", "universal-2"]. Returned explicit AssemblyAI models were Universal-2; auto used its actual returned model. Lower surface-text WER is better; percentages above 100 include insertions.

Language Sandchest explicit WER AssemblyAI explicit WER Sandchest auto WER AssemblyAI auto WER Correct auto labels: Sandchest / AssemblyAI
Amharic 27.20% (96/353) 113.88% (402/353) 27.20% (96/353) 113.03% (399/353) 20/20 · 13/20
Assamese 33.72% (116/344) 123.55% (425/344) 33.72% (116/344) 115.70% (398/344) 20/20 · 1/20
Gujarati 29.70% (128/431) 107.42% (463/431) 35.50% (153/431) 107.42% (463/431) 18/20 · 11/20
Tamil 38.75% (143/369) 41.19% (152/369) 38.75% (143/369) 46.61% (172/369) 20/20 · 17/20
Telugu 31.59% (115/364) 102.47% (373/364) 33.79% (123/364) 102.75% (374/364) 19/20 · 18/20

Automatic language identification is 97/100 versus 60/100. The remaining local errors are two Gujarati clips labeled Hindi and one Telugu clip labeled Tamil, all under 10 seconds. The large WER gaps partly reflect wrong-script provider output: in explicit mode, all 20 provider responses in each of Amharic, Assamese, Gujarati and Telugu lacked a majority of letters from the reference script. Tamil has a much smaller gap. The private script diagnostic is descriptive, not a phonetic/semantic accuracy measure or proof of statistical superiority.

The recorded local API medians are 523–533ms explicit and 529–785ms auto, versus remote 3.10–4.22s explicit and 3.30–4.68s auto. These include upload/queue/persistence/polling, exclude model startup and use different local/remote hardware. They are not cloud capacity or production-tail estimates. Completed word/encoding defects: 0 local; 12 explicit and 10 auto provider responses. Structurally valid word metadata does not establish acoustic boundary accuracy.

AssemblyAI recommends at least 15 seconds of spoken audio for reliable language detection. A disclosed supplemental analysis separates physical audio length: below 15 seconds, auto labels are 72/75 versus 44/75; at least 15 seconds, 25/25 versus 16/25. Physical duration does not guarantee 15 seconds of speech. This addendum was made after collection and early response metadata inspection but before reference scoring; all 100 remain in the primary results.

Receipts: .sandchest/benchmarks/multilingual-20260828/vox-lid-validation1/{plan.json,api/report.json,scored-report.json,duration-strata.json,script-diagnostic.json}. No new model tuning, serving default, deployment or beta approval follows from this result.

Four-language development screen — previous candidate

These languages are now measured, but the quality gate is not passed. The reusable public harness ran 320 paired API jobs on 80 new ML-SUPERB recordings: 20 each in Bashkir, Breton, Nynorsk and Tatar. Each provider completed all 160 original-clip jobs; one explicit Bashkir provider transcript was empty. Public dev references become local development data, with unknown speakers and training overlap, no manual word times, and only 4–8 seconds of audio per original clip.

Language Sandchest explicit WER AssemblyAI explicit WER Sandchest auto WER AssemblyAI auto WER Correct auto labels: Sandchest / AssemblyAI
Bashkir 104.39% (119/114) 101.75% (116/114) 108.77% (124/114) 112.28% (128/114) 0/20 · 0/20
Breton 90.63% (145/160) 104.38% (167/160) 101.88% (163/160) 107.50% (172/160) 0/20 · 3/20
Nynorsk 36.47% (62/170) 37.65% (64/170) 51.18% (87/170) 37.06% (63/170) 0/20 · 1/20
Tatar 97.01% (130/134) 96.27% (129/134) 106.72% (143/134) 105.22% (141/134) 0/20 · 0/20

Both systems struggle on this set. Nynorsk explicit recognition is close; Sandchest auto recognition is worse. Bashkir, Breton and Tatar need recognition improvements, not merely accepted language codes. Correct automatic labels are 0/80 versus 4/80; Nynorsk and Norwegian/Bokmål remain separate codes in this assessment. Word/encoding defects are 0 local versus 3 provider original-clip responses. These small groups are not universal language rankings.

A second, separately reported cohort joins each fixed group of five source recordings with 200ms pauses, producing 16 derived controls of 24.4–34.5 seconds. They reuse the same 80 recordings and are not independent speakers, natural long form or proof of 15 seconds of speech. Of the 64 paired control requests, local completion was 31/32 versus 32/32. Auto labels remain 0/16 versus 2/16; more duration did not solve the routing problem. One explicit Bashkir request returned a native 503 after three configured queue attempts, taking 33.95 seconds; it had no transcript content or usage charge. A fresh native CLI replay also fails with SC_INFERENCE_FAILED (code 7). The deeper SDK cause remains unresolved.

Across both cohorts, the audit verifies 192 local jobs, 191 unique completed usage events, one failed/unbilled job, revoked keys and stopped owned processes. All 191 successful jobs took one attempt. The initial post-run auditor incorrectly required one attempt for the retryable failure; a separate read-back audit corrected that assumption without changing responses, database contents, the model or the original driver. No request was resubmitted.

Original-clip local API medians are 523–776ms explicit / 647–770ms auto, versus remote 1.89–2.03s / 2.29–2.47s. All are serial local-versus-remote diagnostics. Public metadata downloads occurred concurrently, including two oversized metadata responses; these timings are not isolated performance or production-tail evidence. Failed-call latency remains visible rather than included in completion-only percentiles.

Private evidence: .sandchest/benchmarks/multilingual-20260828/mlsuperb-development1/{short-report.json,derived-report.json,api-audit.json,failure-analysis.json}. This previous candidate matches the five-language validation binary. Its ba/br/tt results are superseded only by the measured integration above; the original failures and scores remain preserved.

Four-language MMS recognition experiment — historical prototype

A further private Rust/MLX experiment reuses the pinned MMS base model with four additional language adapters. All four CPU-reference comparisons match decoded text and every one of396 frame argmax decisions; the largest logit difference is 0.001072. Two prototype unit tests pass. This checks implementation math, not independent language accuracy.

The experiment then completes 96/96 inputs with nonempty text: the same 80 original recordings and16derived controls. There are no new AssemblyAI submissions or new independent examples. Invalid-adapter and nonfinite-input controls reject safely and the following valid output recovers exactly. No server, word-alignment path or default configuration was changed.

Language Current explicit API WER Private MMS WER on 20 short clips Cached AssemblyAI WER MMS WER on 4 derived controls
Bashkir 104.39% 18.42% (21/114) 101.75% 21.05% (24/114)
Breton 90.63% 58.75% (94/160) 104.38% 62.50% (100/160)
Nynorsk 36.47% 32.94% (56/170) 37.65% 37.06% (63/170)
Tatar 97.01% 25.37% (34/134) 96.27% 32.84% (44/134)

Bashkir and Tatar are strong next integration candidates on this development set. Breton improves materially but still has high word error. Nynorsk gains only six words on short clips and regresses on longer controls:63 errors versus 60 current and 59 provider; a blanket replacement is not justified. The formerly failed Bashkir composite completes in this recognizer, but that does not fix the current API's decoder failure.

Short-clip native model-pipeline medians are 90–106ms; derived-control medians 360–424ms. These exclude word alignment, formatting, upload, queue, persistence and polling, so they must not replace the full API timings above. Automatic language routing is untested for these new adapters. Recognition improvement alone does not establish timestamp, API or beta readiness.

Private evidence: mlsuperb-mms1/{plan.json,numerical-report.json,report.json,controls.json,collection-receipt.json} under the benchmark root; pinned assets/build under .sandchest/experiments/whisper-large-20260827/durable1/mms-mlsuperb1. All owned model processes exited. The subsequent ba/br/tt integration above closes its serving/alignment/regression step; Nynorsk remains unpromoted, and unseen new-route validation and manual word-time evidence remain open.

Historical broad baseline

The following 82-language screen predates MMS integration. Its failures and five-language scores are retained as the original baseline, not presented as latest-candidate results. The subsequent integration and validation sections supersede only their measured routes and samples.

What was actually tested

  • 904 real API tasks on 276 distinct public recordings. The original 80-language screen contributed 672 tasks; 100 fresh recordings in five weak languages added 200 explicit-mode tasks; Hawaiian and Eastern Yiddish added 32 explicit/auto tasks on eight recordings.
  • Both providers received the same audio and options for each pair. Sandchest used upload → durable queue → Rust/Metal → persisted result → poll on isolated local storage. AssemblyAI used its live API.
  • Baseline: F16 Whisper large-v3-turbo, Parakeet for detected English, retained English CTC/Spanish MMS alignment. Binary SHA256 6b6a355d3614c393004bbe53d3c1f3ca122fb478e3004f95a2a74dc0802cf12a. No serving default changed.
  • Data: pinned FLEURS, OpenSLR and Meta Omnilingual public recordings. Human publisher references, local development only; no manual word times or independent Cap ground truth. Model-training overlap is unknown for Whisper/AssemblyAI. Omnilingual train data is not independent quality evidence for the Omnilingual model.
  • The fresh five-language set excludes prior audio/source IDs and normalized reference text. Hawaiian/Yiddish have four speakers each; Eastern Yiddish does not prove every dialect. Selected public asset hashes are retained; complete archive/Parquet hashes are not claimed.

Completion and language detection

Measure Sandchest AssemblyAI
Completed API tasks 357/452 451/452
Failed API tasks 95/452 1/452
Correct auto language label 96/176 110/176
Completed responses with encoding/word-metadata defects 0 42

All 452 local jobs had one attempt. Exactly 357 completions produced 357 unique usage events; failed jobs were not billed. All owned API/native processes were stopped and checked absent.

Metadata validity is not acoustic timestamp accuracy. These corpora contain no manual word boundaries. An explicit response language can echo the request; text scores are needed to assess recognition.

Recognition by language

Lower error is better. WER/CER remain separate; failed transcripts contribute deletions. Completed transcripts retain their text score even when metadata is invalid. Rates above 100% include insertions. A failed/empty result is not useful merely because it has fewer errors than another bad result.

Sample sizes vary. Amharic, Assamese, Gujarati, Tamil and Telugu have 22 explicit recordings but only two auto recordings. Other rows have two or four examples. All remain small development studies, not established language-level rankings.

Code Metric Sandchest explicit error AssemblyAI explicit error Sandchest completed explicit/auto AssemblyAI completed explicit/auto Auto labels correct: Sandchest / AssemblyAI
af WER 21.62% (8/37) 21.62% (8/37) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
am WER 106.83% (344/322) 116.77% (376/322) 3/22 · 2/2 22/22 · 2/2 0/2 · 0/2
ar WER 32.26% (10/31) 35.48% (11/31) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
as WER 107.46% (461/429) 111.19% (477/429) 14/22 · 1/2 22/22 · 2/2 0/2 · 0/2
az WER 25.00% (8/32) 25.00% (8/32) 2/2 · 2/2 2/2 · 2/2 2/2 · 1/2
ba See current four-language screen above
be WER 46.15% (24/52) 51.92% (27/52) 2/2 · 2/2 2/2 · 2/2 1/2 · 0/2
bg WER 30.77% (12/39) 28.21% (11/39) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
bn WER 80.49% (33/41) 102.44% (42/41) 2/2 · 1/2 2/2 · 2/2 1/2 · 2/2
bo WER 100.00% (148/148) 100.00% (148/148) 0/4 · 1/4 4/4 · 4/4 0/4 · 3/4
br See current four-language screen above
bs WER 5.13% (2/39) 15.38% (6/39) 2/2 · 2/2 2/2 · 2/2 0/2 · 1/2
ca WER 0.00% (0/37) 0.00% (0/37) 2/2 · 2/2 2/2 · 2/2 2/2 · 1/2
cs WER 10.00% (4/40) 7.50% (3/40) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
cy WER 42.19% (27/64) 20.31% (13/64) 2/2 · 2/2 2/2 · 2/2 2/2 · 1/2
da WER 4.08% (2/49) 0.00% (0/49) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
de WER 2.86% (1/35) 0.00% (0/35) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
el WER 12.82% (5/39) 15.38% (6/39) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
en Earlier separate study; not pooled
es Earlier separate study; not pooled
et WER 9.38% (3/32) 21.88% (7/32) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
eu WER 24.14% (7/29) 55.17% (16/29) 4/4 · 4/4 4/4 · 4/4 4/4 · 2/4
fa WER 28.21% (11/39) 35.90% (14/39) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
fi Earlier separate study; not pooled
fo Not measured
fr Earlier separate study; not pooled
gl WER 0.00% (0/32) 9.38% (3/32) 2/2 · 2/2 2/2 · 2/2 2/2 · 0/2
gu WER 100.00% (452/452) 107.74% (487/452) 0/22 · 1/2 22/22 · 2/2 0/2 · 1/2
ha WER 178.12% (57/32) 115.62% (37/32) 2/2 · 1/2 2/2 · 2/2 0/2 · 0/2
haw WER 79.33% (165/208) 68.27% (142/208) 4/4 · 4/4 4/4 · 4/4 0/4 · 0/4
he WER 66.67% (20/30) 60.00% (18/30) 2/2 · 2/2 2/2 · 2/2 1/2 · 1/2
hi Earlier separate study; not pooled
hr WER 13.79% (4/29) 20.69% (6/29) 2/2 · 2/2 2/2 · 2/2 2/2 · 1/2
ht Not measured
hu WER 24.24% (8/33) 36.36% (12/33) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
hy WER 47.62% (20/42) 45.24% (19/42) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
id Earlier separate study; not pooled
is WER 27.27% (9/33) 42.42% (14/33) 2/2 · 2/2 2/2 · 2/2 2/2 · 0/2
it WER 0.00% (0/43) 0.00% (0/43) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
ja CER 0.00% (0/115) 1.74% (2/115) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
jw WER 39.47% (15/38) 73.68% (28/38) 2/2 · 2/2 2/2 · 2/2 0/2 · 1/2
ka WER 100.00% (28/28) 107.14% (30/28) 0/2 · 1/2 2/2 · 2/2 0/2 · 2/2
kk WER 14.29% (4/28) 50.00% (14/28) 2/2 · 2/2 2/2 · 2/2 2/2 · 1/2
km CER 100.00% (289/289) 99.31% (287/289) 0/2 · 2/2 2/2 · 2/2 0/2 · 2/2
kn WER 54.55% (12/22) 68.18% (15/22) 2/2 · 2/2 2/2 · 2/2 0/2 · 1/2
ko WER 2.22% (1/45) 4.44% (2/45) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
la Not measured
lb WER 83.78% (31/37) 81.08% (30/37) 2/2 · 2/2 2/2 · 2/2 0/2 · 0/2
ln WER 55.00% (22/40) 82.50% (33/40) 2/2 · 2/2 2/2 · 2/2 0/2 · 0/2
lo CER 100.00% (161/161) 98.76% (159/161) 2/2 · 2/2 2/2 · 2/2 0/2 · 0/2
lt WER 26.47% (9/34) 61.76% (21/34) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
lv WER 12.90% (4/31) 19.35% (6/31) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
mg Not measured
mi WER 13.51% (5/37) 18.92% (7/37) 2/2 · 2/2 2/2 · 2/2 0/2 · 2/2
mk WER 20.83% (10/48) 22.92% (11/48) 2/2 · 2/2 2/2 · 2/2 2/2 · 1/2
ml WER 100.00% (26/26) 107.69% (28/26) 0/2 · 2/2 2/2 · 2/2 0/2 · 1/2
mn WER 103.12% (33/32) 103.12% (33/32) 2/2 · 1/2 2/2 · 2/2 0/2 · 2/2
mr WER 81.82% (36/44) 90.91% (40/44) 2/2 · 1/2 2/2 · 2/2 0/2 · 1/2
ms WER 30.43% (14/46) 36.96% (17/46) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
mt WER 76.32% (29/38) 65.79% (25/38) 2/2 · 2/2 2/2 · 2/2 2/2 · 0/2
my CER 100.00% (284/284) 100.00% (284/284) 0/2 · 0/2 2/2 · 2/2 0/2 · 2/2
ne WER 77.78% (42/54) 88.89% (48/54) 2/2 · 1/2 2/2 · 2/2 0/2 · 2/2
nl WER 13.04% (3/23) 8.70% (2/23) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
nn See current four-language screen above
no WER 2.08% (1/48) 6.25% (3/48) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
oc WER 78.05% (32/41) 75.61% (31/41) 2/2 · 2/2 2/2 · 2/2 0/2 · 0/2
pa WER 84.75% (50/59) 101.69% (60/59) 2/2 · 2/2 2/2 · 2/2 0/2 · 2/2
pl WER 7.14% (2/28) 3.57% (1/28) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
ps WER 98.08% (51/52) 96.15% (50/52) 2/2 · 0/2 2/2 · 2/2 0/2 · 1/2
pt WER 4.35% (1/23) 4.35% (1/23) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
ro WER 7.84% (4/51) 7.84% (4/51) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
ru WER 4.00% (1/25) 4.00% (1/25) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
sa Not measured
sd WER 100.00% (51/51) 109.80% (56/51) 2/2 · 2/2 2/2 · 2/2 0/2 · 0/2
si WER 100.00% (12/12) 100.00% (12/12) 2/4 · 4/4 4/4 · 4/4 0/4 · 0/4
sk WER 14.29% (6/42) 16.67% (7/42) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
sl WER 8.57% (3/35) 17.14% (6/35) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
sn WER 123.08% (32/26) 165.38% (43/26) 2/2 · 2/2 2/2 · 2/2 0/2 · 0/2
so WER 188.57% (66/35) 100.00% (35/35) 2/2 · 2/2 2/2 · 2/2 0/2 · 1/2
sq Not measured
sr WER 14.29% (4/28) 14.29% (4/28) 2/2 · 2/2 2/2 · 2/2 1/2 · 0/2
su WER 21.21% (7/33) 66.67% (22/33) 4/4 · 4/4 4/4 · 4/4 0/4 · 0/4
sv WER 10.26% (4/39) 5.13% (2/39) 2/2 · 2/2 2/2 · 2/2 2/2 · 1/2
sw WER 35.90% (14/39) 56.41% (22/39) 2/2 · 2/2 2/2 · 1/2 2/2 · 1/2
ta WER 63.58% (213/335) 43.28% (145/335) 22/22 · 2/2 22/22 · 2/2 2/2 · 2/2
te WER 96.21% (305/317) 105.68% (335/317) 6/22 · 2/2 22/22 · 2/2 0/2 · 2/2
tg WER 76.32% (29/38) 65.79% (25/38) 2/2 · 1/2 2/2 · 2/2 0/2 · 0/2
th CER 6.47% (11/170) 8.82% (15/170) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
tk Not measured
tl WER 8.70% (4/46) 21.74% (10/46) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
tr WER 5.00% (1/20) 0.00% (0/20) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
tt See current four-language screen above
uk WER 6.67% (2/30) 13.33% (4/30) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
ur WER 24.14% (14/58) 25.86% (15/58) 2/2 · 2/2 2/2 · 2/2 2/2 · 1/2
uz WER 114.71% (39/34) 100.00% (34/34) 2/2 · 2/2 2/2 · 2/2 0/2 · 0/2
vi WER 15.00% (6/40) 10.00% (4/40) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2
yi WER 101.02% (199/197) 95.94% (189/197) 4/4 · 4/4 4/4 · 4/4 0/4 · 3/4
yo WER 106.90% (31/29) 96.55% (28/29) 2/2 · 2/2 2/2 · 2/2 0/2 · 2/2
zh CER 0.00% (0/67) 0.00% (0/67) 2/2 · 2/2 2/2 · 2/2 2/2 · 2/2

Original native candidate study

Twenty new recordings per language, 100 total (17.918 minutes). The Rust MMS prototype returned nonempty text on 100/100. The current API completed 40/100 (39 nonempty); AssemblyAI completed 100/100, with 19 encoding/word-metadata defects. All comparisons below use exactly the same references/audio.

Language Current Sandchest WER Rust MMS prototype WER AssemblyAI WER Full Whisper large-v3 WER
Amharic 102.97% (312/303) 33.00% (100/303) 117.16% (355/303)
Assamese 106.12% (416/392) 30.10% (118/392) 110.20% (432/392)
Gujarati 100.00% (420/420) 26.67% (112/420) 107.38% (451/420)
Tamil 64.21% (192/299) 40.47% (121/299) 45.15% (135/299) 44.48% (133/299)
Telugu 97.14% (272/280) 33.57% (94/280) 106.79% (299/280)

Some large provider errors involve the wrong writing system: the private script diagnostic retains per-language counts. These are surface-text WER/CER results, not a measure of semantic understanding or phonetic transcription accuracy. Twenty examples are still insufficient for broad language-quality guarantees.

The MMS language-adapter model was implemented as a private Rust/MLX prototype. Five numerical probes matched reference decoded text and every frame argmax. A prior seven-language, 28-clip screen also completed all recordings; Arabic regressed, so a universal model replacement is not justified.

MMS remains an explicit-language native prototype, not the serving model. The follow-up below adds acoustic word timing, bounded long-audio processing and fault tests. Formatting, automatic language routing, full API integration and long-audio quality remain unresolved. The earlier roughly0.1–0.2-second raw-model timings exclude those later stages and are not end-to-end speed parity.

The full large-v3 comparison uses the same 20 fresh Tamil recordings, but its latency includes native HTTP/model/word timing rather than the full API. It remains a private candidate.

Native candidate: word timing and longer recordings

The final timing study completed 116/116 clips: the same100 weak-language recordings plus 16 previously used Spanish development clips from four speakers. Recognition text stayed exact on all100 weak-language clips. The independent MMS_FA aligner selected1,321 of1,677 weak-language words; remaining words retained CTC times. This is valid metadata, not manually proved timing accuracy in those languages.

Spanish has175 manually marked words. The following comparison uses the same162 matched words across all variants; it does not hide unmatched words or count these reused clips as new data. AssemblyAI responses were cached; no new provider calls were made.

Spanish development timing Direct CTC MMS ASR + acoustic aligner AssemblyAI
Mean start error 69.14ms 21.70ms 54.82ms
Mean end error 45.28ms 23.98ms 51.25ms
Both endpoints within80ms 90/162 (55.56%) 147/162 (90.74%) 107/162 (66.05%)

The retained Spanish serving path has the same timings on these162 words and better recognition:6/175 word errors versus MMS16/175. This is not a reason to replace Spanish recognition. Full CTC boundary medians did not improve the direct timings in this manual subset. No new validation or holdout examples were used.

The next native candidate preserved all116 short transcripts and word arrays exactly. It also completed seven constructed recordings of2.4–4.2minutes, using bounded overlapping acoustic windows, plus six digital-silence and one pure-tone control. Eight invalid-request/cancellation/deadline tests recovered exactly, including cancellation during recognition and after one completed window. The final run made148 native calls, with no new AssemblyAI calls. An exact-silence guard fixed an Amharic false word without treating quiet nonzero audio as silence. These controls do not establish general noise rejection.

The initial fixed-window candidate regressed on long audio. Each five-language composite joins the same20 development clips with350ms gaps; it is not natural long-form or independent data. Compared with recognizing those clips separately:

Language Separate clips WER Constructed long recording WER
Amharic 33.00% (100/303) 42.90% (130/303)
Assamese 30.10% (118/392) 43.37% (170/392)
Gujarati 26.67% (112/420) 31.43% (132/420)
Tamil 40.47% (121/299) 39.13% (117/299)
Telugu 33.57% (94/280) 38.93% (109/280)

In that initial candidate, some decoded words span entire inserted pauses, and alignment coverage drops sharply; the Assamese composite selected no refined word boundaries. The speech-partitioning follow-up below addresses this observed failure. The two Spanish composites retained roughly22ms start /23ms end error on common manually marked words, but repeat the same16 clips. Native processing took about4.1–5.8seconds for these2.4–4.2minute composites, including alignment; there was no full API, remote provider or load comparison.

Private builds passed186 library tests and11 focused tests. Failed attempts and their receipts remain preserved. The final successful studies are mms-word-timing2 and mms-windowed3; mms-timing-checkpoint1.json records their checksums, controls and limitations. Serving source/defaults and the904-task API benchmark above remain unchanged.

Speech-aware native follow-up

A private native Silero VAD now chooses break points between speech regions. It does not discard audio: every sample belongs to one partition, with the existing bounded windowing retained for any long partition. Each partition is recognized and aligned separately. Model text or reference transcripts do not choose the boundaries.

A256ms minimum-pause policy recovered much of the initial long-recording regression. A single512ms variant then reduced extra cuts and was selected for the next integration experiment. It passed a full130-case replay: all116 short text/word arrays stayed exact, all seven constructed long recordings completed, and silence/tone controls remained empty. Eight rejection/deadline/cancellation controls passed; recovery includes full long transcriptions after interrupted VAD and segmented recognition. Repeated long outputs are exact. Builds pass186 library and15 focused tests.

Language Initial fixed-window WER 512ms speech partitions WER Separate short clips WER
Amharic 42.90% (130/303) 33.99% (103/303) 33.00% (100/303)
Assamese 43.37% (170/392) 30.87% (121/392) 30.10% (118/392)
Gujarati 31.43% (132/420) 27.38% (115/420) 26.67% (112/420)
Tamil 39.13% (117/299) 39.46% (118/299) 40.47% (121/299)
Telugu 38.93% (109/280) 36.07% (101/280) 33.57% (94/280)

No decoded word spans an entire inserted pause in any of the seven composites. Assamese refined alignment rises from0/346 to337/388 words. This fixes the observed boundary problem, but the residual errors, normalization fallbacks and absence of human timing labels in these five languages remain. Tamil has one more word error than the initial fixed-window result; there is no claim that every metric improved. The Spanish composites retain roughly22ms start/end error on common manual words, with the same small reused-source caveat.

Native times for the512ms composites ranged about3.7–6.6seconds across the pilot and full replay. The later full replay was slower despite exact text/words, so the apparent timing benefit needs controlled measurement. These are local diagnostic times, not production or AssemblyAI latency parity.

Natural Cap stress, not independent accuracy

Both native candidates also completed three previously used Cap development recordings: one57-second Spanish recording and two English recordings of8.4 and9.0minutes. Total source audio is18.31minutes. Word metadata, exact recovery and a repeated9-minute output pass. Cached AssemblyAI labels select explicit adapters; this does not test automatic detection.

Word-edit disagreement with cached AssemblyAI text changes from55.91% to45.16% on the Spanish recording,27.40% to28.11% on one English recording, and17.45% to16.65% on the other. These are not WER or proof of correctness; there is no human reference. Native candidate times were about1.4seconds for Spanish and13.0–13.4seconds for the long English recordings. English/Spanish MMS are used here only to stress the shared long-audio machinery; their retained recognition models are not being replaced.

This checkpoint made348 native requests and zero new provider submissions. The natural-audio scorer now computes only the planned word disagreement using the existing bounded word-distance function; the first attempt unnecessarily computed character distance and reached its safety limit. Neither normalization nor limits were weakened. The full final replay is mms-segmented-gap16-full, the natural stress is mms-cap-natural1, and mms-segmentation-checkpoint1.json records checksums and limitations. All owned processes are stopped. Full API integration, automatic routing, formatting and remaining language/quality coverage are still open; no serving source/default or production setting changed.

Opt-in native/API integration — 2026-08-28

The selected native MMS recognition/timing implementation is now in the worker source, enabled only with SANDCHEST_WHISPER_MMS_ASR=true alongside the whisper-mms build and ctc-mms alignment. The tested binary is 0a078ef78725e7b782ca9cb8748612beb2a1ff78420a622e631d6f1640e92fbb; source bytes match that build. Existing defaults, Cap, shared services and deployment settings are unchanged. The original candidate sections above describe earlier experiments; this checkpoint supersedes their integration status.

All100 explicit-language requests completed through real HTTP upload, the application API, durable queue, native worker, persistence and polling. Text and word timestamps exactly match the selected native candidate. These are the same100 public development recordings, not additional independent samples. Surface-text WER remains:

Language Clips Integrated explicit WER Cached AssemblyAI explicit WER Local API p50 / p95
Amharic 20 33.00% (100/303) 117.16% (355/303) 456 /582ms
Assamese 20 30.10% (118/392) 110.20% (432/392) 457 /674ms
Gujarati 20 26.67% (112/420) 107.38% (451/420) 503 /582ms
Tamil 20 40.47% (121/299) 45.15% (135/299) 455 /677ms
Telugu 20 33.57% (94/280) 106.79% (299/280) 509 /663ms

WER can exceed100% because insertions count. Wrong-script/transliterated provider output contributes to these large gaps; these numbers do not establish general semantic superiority. No new AssemblyAI calls were made. Latency is serial, warm, local end-to-end timing on an M4 Max, not a same-hardware capacity or production comparison.

Unrestricted automatic detection is substantially weaker:

Language Completed Correct language Failure-inclusive auto WER
Amharic 16/20 12/20 59.74% (181/303)
Assamese 19/20 0/20 102.55% (402/392)
Gujarati 20/20 4/20 98.33% (413/420)
Tamil 20/20 20/20 40.47% (121/299)
Telugu 20/20 0/20 103.21% (289/280)

Thus 95/100 completion is only36/100 correct language selection. These automatic results are not compared to the cached provider's explicit-mode numbers as if request settings matched. A hinted response label can differ from the acoustic language actually used; the integration preserves the existing tested policy rather than forcing ASR from that label. Dedicated language-identification work is required before broad beta traffic.

Verification:203 native tests pass with MMS enabled; base, Whisper and CTC configurations also pass. Ten retained English/Spanish/natural Cap requests match the baseline with the flag both disabled and enabled. All100 short MMS requests, five constructed long recordings, five silence cases and one tone case match the selected prototype over native HTTP. Five guided cases also match explicit inference in their actual acoustic language, while exposing incorrect language identification rather than counting hints as proof.

The200-job API audit has195 completed jobs and exactly195 unique usage events; all five failed jobs are unbilled, every job used one attempt, and all temporary keys were revoked. A separate28-case API suite checks five crops with exact single timestamp offsets, custom spelling, punctuation flags, silence, invalid ranges, invalid media and two stereo cases. It has22 completions/22 usage events and six unbilled errors. Billing retains physical duration for mono/crops and measured channel-duration for stereo, including silent channels. Two real 1s native HTTP timeouts on long audio are followed by exact short-request recovery. Owned processes are stopped; no shared database or port3000 was used.

The private harness retained failed assertions and corrected three test assumptions: hints need not force acoustic selection, punctuation=false should remove a final danda, and failed ranges/stereo metering have their existing API semantics. No service behavior or billing rule was weakened to make those checks pass. An earlier interrupted eight-job options run was separately audited:seven completions/seven usage events, one unbilled error. It is not counted as new quality data.

Gates at that historical checkpoint included independent quality validation, manual word-time references in the five languages, punctuation/formatting, automatic language ID,11 unmeasured base languages and regional accents, other API features and later deployment/capacity validation. Later sections supersede only their specifically measured gates. No general 20% beta approval. Private receipts: mms-api-integration1, mms-api-options2, and mms-api-options1/interrupted-run-audit.json under the multilingual benchmark directory.

Other candidate evidence

A private decoded-text repetition check recovered 12/26 failures on the original 304 FLEURS audio/mode tasks: 278→290 completions, no word-structure defects, all previous completions retained and 277 unchanged text/word outputs. Several recovered transcripts still have high error. This candidate remains unpromoted.

Guarded native language detection — 2026-08-29

An additional disabled-by-default detector is now in the worker source. The tested binary is 4b86362161929e4f7071d6e24dee508cab2fb35dc7c2dc366250a329ca425b6b. It uses the pinned SpeechBrain VoxLingua107 ECAPA model, ported to Rust/MLX. All450 native predictions match the reference implementation; the largest absolute log-probability difference is0.0000763. Its107 labels include all99 API base codes after the documented legacy-code mappings; that inventory is not a99-language quality result.

The detector can hand off only to the five measured MMS recognizers, with raw confidence at least0.8 and speech detected in the same first30seconds of the requested crop. Other cases retain Whisper's existing detection. Explicit requests bypass it. Extra model classes are not renormalized away. The threshold was selected on development evidence: an unrestricted switch would have sent four Hindi controls and one English control to Gujarati; the guard rejects all five. This is not an independently calibrated confidence guarantee.

The repeated100-clip unrestricted-auto API run now completes100/100 and selects the correct language98/100, versus95 completions/36 correct previously:

Language Previous correct New correct New failure-inclusive auto WER Local API p50 / p95
Amharic 12/20 20/20 33.00% (100/303) 451 /565ms
Assamese 0/20 20/20 30.10% (118/392) 560 /683ms
Gujarati 4/20 18/20 34.76% (146/420) 554 /808ms
Tamil 20/20 20/20 40.47% (121/299) 554 /678ms
Telugu 0/20 20/20 33.57% (94/280) 552 /681ms

All100 explicit requests also complete with exactly the prior MMS text and word times. The two remaining automatic mistakes are Gujarati. These are reused development clips, not new independent samples. The cached AssemblyAI scores in the earlier table use explicit language selection, so they are not a matched automatic-detection comparison. No new provider submissions occurred.

Regression and reliability checks:

  • All286 retained native requests have identical text, words, confidence, language, model and duration with the detector off/on, including 50Cap recordings. This is output preservation, not new human-reference accuracy.
  • The200-job API run has200 unique usage events, one attempt per job, valid word metadata and revoked temporary keys.
  • A further90 API cases cover40 mechanical controls plus the same50Cap recordings. All50 prior Cap responses match, including45 speech completions and five no-speech errors; all90 idempotent retries return the same jobs. There are72 completions/72 usage events and18 expected unbilled errors. Crops apply their1001ms offset exactly once, including both range endpoints. Stereo accounting includes silent channels; failed threshold requests do not bypass the gate through fallback.
  • That detector-only Cap API replay measured median/p95/p99 of1.667/8.373/18.324seconds across45 speech completions. Serial local timings exclude model startup and remote database/billing; the earlier run-to-run tail variation remains unresolved. No fresh AssemblyAI or matched-load latency claim is made.
  • Two1-second native HTTP deadlines recover to exact automatic transcripts. Builds pass210 MMS tests and all other configurations; the app passes189 tests with7 conditional skips, typecheck and lint. Fourteen collector tests and five model-preparation tests pass. Offline preparation reproduces the exact pinned85,221,060-byte artifact; serving uses no Python.

The first option preparation stopped on a fixture-name typo before any inference. A separate five-job attempt exposed a wrong test expectation about confidence thresholds; four completions and one unbilled error were independently audited. The final90 response checks all passed, but its original post-run auditor incorrectly expected null rather than the existing0 duration for no-speech failures. A read-only audit corrected that expectation and rechecked all responses, replays, jobs and usage without altering application behavior or original evidence.

Receipts: vox-lid-reference1, vox-lid-native1, vox-lid-retained1, vox-lid-api1, vox-lid-options3 and vox-lid-app-checks1. This supersedes the earlier36/100 automatic result only when the new opt-in detector is enabled. At that checkpoint, manual word-time accuracy, unseen validation, code switching,11 missing base-language reference sets, regional aliases and remaining API features were open. The fresh validation and later language studies above retain their separate provenance. No default, shared service, Cap data or production traffic changed.

Diagnostic baseline latency

Serial local-Mac versus remote-cloud measurements, including completed and failed tasks. Background collection/checks overlapped parts of the study. These are not production latency/capacity acceptance results.

Provider p50 p95 p99
Sandchest 0.78s 6.60s 11.69s
AssemblyAI 3.37s 5.58s 8.48s

Next gates

Integrate only justified language-specific candidates, then test automatic routing, manual word timing, long audio, code switching, faults and Cap-domain examples. Turkmen still needs suitable accessible reference data; regional aliases need accent/dialect coverage. The newly measured Haitian Creole, Latin and Sanskrit baselines need recognition/detection improvements and independent validation. Expand the promising MMS candidates beyond two-clip screens, prioritize current recognition failures and weak automatic detection, and retain the dedicated Faroese/Hindi quality gaps; keep the new-route validation and manual word-time gates separate from code inventory. API feature parity and deployment/capacity checks remain distinct gates. No production settings or traffic changed.

Harness: MULTILINGUAL_EVALUATION.md. Earlier Cap/timestamp results: BENCHMARK_STATUS.md. Private immutable plans, sources, responses and audits remain under .sandchest/benchmarks/multilingual-20260828.