Working objective clarified by the user on 2026-08-26: keep Rust as the serving runtime and match AssemblyAI's supported transcription language list, not Cap's language selector. Continue improving independently measured accuracy and end-to-end performance as separate goals while preserving the Cap API workflows. These are acceptance criteria, not claims that the current worker meets them.
The frozen current worker completed a 98-job paired API study on 49 new Cap recordings / 1.842 hours. One of fifty originally selected recordings was unavailable and remains in the denominator. Sandchest completed 44/49, versus 47/49 for AssemblyAI. On the 43 mutually completed pairs, local p50/p95/p99 is 1.966/10.471/17.244 seconds, versus 7.700/17.161/19.996 seconds remotely. These are local-versus-remote serial measurements, not production capacity.
Three of four provider-completed/local-error pairs have empty provider text; one has nine provider words, so potential missed speech remains unresolved. Returned language labels agree on 39/43 mutually completed pairs, without independent language truth. Cached Cap reference transcripts remain sealed. This is provider comparison, not human accuracy or word-timing proof.
All 49 replay/conflict/resource checks and exact isolated accounting pass: 49 single-attempt jobs,44 unique usage events,five unbilled errors,revoked keys and clean process-group exits. A narrow additive auditor correction handles the provider-observed zero-duration no-spoken-audio error without allowing partial text or changing any recognition result. The original failed report is retained. General beta is not approved. See the fresh holdout evidence for sampling, failure counts, scoring exclusions and limits.
Turkmen now has eighty fresh paired API results on twenty publisher recordings. Explicit requests complete 20/20 on both services; automatic requests complete 18/20 locally versus 20/20 remotely. Neither returns Turkmen automatically, and output disagreement is high. Caption provenance is unknown, so no human-reference score or Turkmen quality-parity claim is available. The operational/accounting gate passes; this is not all-language beta approval.
Fresh unused Hindi validation now favors the unchanged private candidate on forty published read-speech recordings: 13.56% WER / 5.98% CER, versus fresh AssemblyAI 19.51% / 7.87%. Both complete40/40; local full-API p50/p95 is1.74/3.88s versus remote4.91/6.41s. Twenty AB/twenty BA pairs use identical audio/options. Independent scoring reproduces all counts and the fixed clip-bootstrap intervals. All local single-attempt jobs, exact usage, replays and resources reconcile. One provider zero-duration word keeps the combined structural gate failed without excluding its completed text. This is short explicit Hindi on the verified Mac Metal candidate, not automatic/long/conversational Hindi, human timing or all-language beta approval.
Automatic Hindi now has separate development evidence: the private ad46fd67…
candidate reduces WER from42.02% to31.36% on sixty reused recordings, with all120
feature-off/on primary requests completed. Twenty-six companion boundary/recovery
checks pass. Eight recordings retain another language label and escape the Hindi
route; native HTTP median rises from0.98s to1.95s. That native study is separate
from the now-completed forty-case automatic API comparison. On those already
consumed explicit-validation recordings, Sandchest achieves21.76% WER versus28.39%
for fresh AssemblyAI; full-API p50/p95 is2.58/7.17s versus4.64/8.39s. Both complete
40/40 and return36Hindi labels. Independent scoring confirms the WER interval
favors local; the CER interval crosses zero. All local structure/accounting/
replay/resources pass. Three provider word-time defects retain the failed combined
structural report and remain included in text scoring. This is not a fresh
holdout, human word-time validation or all-language beta approval.
The exact automatic-Hindi ad46fd67… build also passes100/100 retained Cap native
requests: all50 recordings match their earlier semantic baselines with both Hindi
flags off and on. Complete word times, confidence and language metadata remain
unchanged; only elapsed-time/allocator observations are excluded. This is reused
Cap regression evidence, not fresh accuracy or provider-latency evidence. All four
child groups and the runner exit cleanly; no provider or reference call is made.
The same exact build now passes the22-case advanced local API replay:20 completed, two expected errors,20 unique usage events and one attempt per job. Twenty cases match historical fields exactly; two automatic-Hindi cases match the predeclared Qwen changes. All88 POSTs and111 resource/list GETs pass, including crop and stereo word conservation. Native/API/inspector/runner groups exit cleanly. This proves these retained contract cases, not full AssemblyAI feature parity or production capacity; external billing is disabled in the isolated test.
A private NB-Whisper Small candidate improves explicit Nynorsk WER on twenty short DEV clips to13.53%, versus37.65% for cached AssemblyAI. Four joins of the same audio also improve WER, but their CER is slightly worse than AssemblyAI; they are not four independent new recordings. A checker whitespace correction passes a separate five-call long/nonspeech/invalid/recovery replay with exactly unchanged baseline word times. Automatic Nynorsk, fresh validation, human timing and integration are still unverified. Neither candidate is promoted. See the latest multilingual comparisons for denominators and limits.
The private Hindi successor passes 46 advanced native calls and a 22-job API run with twenty exact usage events and two expected unbilled errors. An additive audit corrects only the reporter's invalid global-order assumption for stereo resources; every word field, duplicate count and channel-local sequence remains exact.
An immediate request after a long native timeout still receives 503. A separate unchanged-candidate test measures readiness recovery at 214 ms and exact next-output recovery by 227 ms, without relabeling the original failure. A separate four-job queued-API test passes both timeout → immediate-silence sequences without a native readiness wait: four single-attempt jobs, two unbilled errors, two exact outputs and two one-second usage events. This proves the measured service/accounting recovery, not spoken decoding after the long Whisper timeout or concurrent capacity. A separate three-call native check now verifies spoken Whisper recovery after readiness: the same fourteen words, timestamps and confidences are reproduced after the long timeout. Readiness returns at 112 ms and the spoken response completes by 720 ms after the 504 body. The original immediate 503 remains failed; this is one serial recovery case, not concurrency or human-quality evidence. The alternate Hindi MMS fallback remains unintegrated: only 3/15 transcripts pass the unchanged timing gate and its small projected text gain still trails the provider on the earlier development set. The unused short explicit-Hindi quality comparison above is now complete; broader quality and manual timing remain open.
The extra Cap diagnostic reproduces the unresolved nine-word discrepancy as a speech-detector rejection (peak 0.330 versus threshold 0.5), before transcription. This does not establish missed speech or provider hallucination; no threshold, source audio, caption reference or production setting changed. See the multilingual evidence for exact scopes and receipts.
The unchanged Rust route completed 50 new Cap recordings / two hours through the actual local API. On 48 mutually completed pairs, local p50/p95 was 1.56/8.05 seconds, versus 6.61/15.19 seconds for fresh remote AssemblyAI. Two provider errors were no-speech cases that Sandchest returned empty, so the completion difference is a compatibility caveat, not a recognition win. Cap's actual editor/caption helpers preserved all returned word metadata.
Three apparent language misses were traced to the restricted Cap hint list: unrestricted and explicit requests make both providers agree on Indonesian or Ukrainian. Text/timing quality and fallback/no-speech semantics still have gaps. General beta is not approved. The further fifty selected recordings were reserved at this earlier checkpoint;49 have now been tested in the holdout above. The final fresh-Cap section preserves this earlier sample and its follow-ups.
The user explicitly requires a drop-in, 1:1 AssemblyAI prerecorded transcription contract. Matching Cap's happy path or accepting language codes is insufficient. Verify endpoints, request defaults/nulls/options, complete response resources, queued/processing/completed/error transitions, error status/body semantics, word timestamps and the unchanged official SDK. Retain Sandchest's tenant and billing safeguards; do not weaken them to imitate a provider error code.
The current application is not yet 1:1. Word search, sentences, paragraphs,
SRT, VTT, listing and deletion now have authenticated, tenant-scoped implementations
and unchanged AssemblyAI SDK coverage. The redacted-audio endpoint,
incomplete reflected response fields/defaults, and unsupported features including
diarization, redaction, prompting and localization remain open. New jobs now use
the provider's false filler default with English cleanup; multilingual filled-pause
behavior remains a quality gap.
Mono worker results now also require the transcript text to match the ordered
word text after Unicode compatibility normalization and whitespace removal.
Punctuation, case, accents and numbers remain significant; the API does not rewrite
either output. Contradictory successful responses fail permanently before content
persistence or billing. New multilingual positive cases and four mismatch cases
pass through the actual queue/API test fixture. The current local app suite passes
228 tests, with seven conditional model skips, after the adaptive-selector
regressions, plus typecheck and scoped lint.
This is an output-integrity safeguard, not new recognition or timing evidence.
Private receipts: durable1/mono-text-integrity-checks1/receipt.json and
durable1/adaptive-harness-checks1/receipt.json.
Functional implementations are required; filling response keys with placeholders would not establish parity. The configurable 100 MiB upload limit also differs from AssemblyAI's published limit. Keep these gaps visible before replacement.
This integration closes two concrete boundary bugs: unknown or unsupported
request fields are rejected instead of silently discarded, and
language_confidence_threshold reaches the stored request/native worker. A
compatibility GET now returns current state immediately, leaving polling to the
client SDK. Previous latency results using a hidden long poll are historical and
must not be presented as measurements of this changed contract.
Live provider requests on public audio confirmed that omitted/null language
settings default to automatic detection; an explicit language plus detection is
rejected; disabling detection without a language is rejected; regional codes are
retained. These current observations take precedence over older SDK default
annotations. They are frozen privately under provider-contract1/2.
The opt-in Whisper backend remains under verification. Its regional aliases map
to base acoustic models; dialect accuracy and locale spelling are not proven.
language_detection_options.localization and code-switching controls remain
explicit gaps, not silently accepted features.
Prerecorded completion/error callbacks now have a durable delivery queue and
reflected webhook_url, webhook_status_code, webhook_auth and
webhook_auth_header_name response fields. The custom secret is never returned.
The implementation follows the documented POST payload, 10-second response
deadline, ten total attempts, ten-second retry delay and terminal 4xx behavior.
AssemblyAI webhook protocol,
request/response fields.
Creation persists the delivery row with the transcript/job transaction. Only a
committed terminal parent can be claimed. Delivery runs independently of inference,
billing and retention, and a callback retry never reruns or bills transcription.
Attempt/lease checks fence stale HTTP responses. Deletion cancels future attempts,
erases credentials and prevents a late response from restoring state; an HTTP
request already in flight cannot be recalled. Delivery is at least once: receivers
must deduplicate by transcript_id, especially after a crash following receipt.
Callback destinations are resolved and pinned for every attempt. Private/metadata addresses and mixed DNS answers are blocked; redirects are not followed. A separate non-production flag enables owned loopback test receivers. Routing/framing header overrides are rejected. Callback secrets are encrypted with a purpose-separated key derived from the API-key pepper, with tenant/transcript authenticated context; only a keyed digest enters request fingerprints. Keep that pepper stable while callbacks are pending. Callback fields never reach native inference options.
The additive 0005_transcript_webhooks migration is generated and verified only
in isolated PGlite, including preservation of completed jobs and pending billing
rows. It has not been applied to the shared external database. Apply migrations
before starting the updated queue worker there; no deployment is part of this
checkpoint. Existing migrations and all billing fields remain intact.
A real native hybrid8 run used ten reused Cap recordings across six detected
languages, non-speech, long audio and two expected recognition errors. All text and
word timestamps matched the preceding native checkpoint. Eleven authenticated
callbacks arrived: ten terminal notifications plus one forced 503 retry. All jobs
used one inference attempt; eight completed jobs retained exactly one measured
usage event each, and errors had no usage events. A separate actual SIGKILL/restart
of the API/queue process recovered an in-flight callback with a second delivery,
unchanged transcript, one inference attempt and one usage event. All ephemeral
keys were revoked and owned processes stopped. These runs made no provider calls
and are not new accuracy holdouts or paired speed benchmarks. Private receipts:
webhook-native1 and webhook-crash1.
Unit/integration coverage also exercises the unchanged AssemblyAI SDK, all null option combinations, secret binding/tampering, concurrent claimers, retry caps, expired/stale leases, deletion races and the actual ten-second timeout. The callback protocol is implemented; exact provider validation/error wording for every URL and header edge case is not established. Redacted-audio notifications remain part of the unimplemented redaction feature. General 1:1 API and beta approval remain open.
English filled-pause cleanup and the public filler default are now implemented.
New jobs persist disfluencies: false; explicit true and legacy queued jobs
preserve their previous behavior. The native worker must acknowledge cleanup
before a job can complete or be billed. Native suites pass 112/159/175 tests,
and the app passes 161 tests with seven conditional skips, typecheck and lint.
A completed 74-request native regression preserves every verbatim result on
30 reused Cap recordings and controls. Cleanup removes 30 recognized um/uh
tokens without changing any retained word's text, timestamps or confidence.
The subsequent 111-job real local API run has 108 completions, three instances of the known uncertainty error, and 108 exact-duration usage events. All 37 false/omitted pairs match; 111 retries return their original IDs and 111 changed-setting retries conflict without another job or charge. The cached AssemblyAI result for that one uncertain Cap recording is also empty (completed, zero words), so this status difference is not evidence of missing spoken words. It is not human proof of silence either. General beta approval remains open; English cleanup does not prove multilingual disfluency or recognition parity.
The latest timestamp candidate is not promoted. A private Rust CUPE acoustic port passes its fixed numerical checks against the pinned PyTorch reference (maximum logit difference 0.000074, zero changed top classes), plus 32 rejection/recovery pairs. On the existing 16 Spanish development clips, however, the complete experimental alignment has 32.98/86.74 ms start/end MAE, versus 22.19/25.06 ms for the retained MMS envelope. Three fixed comparisons using the candidate's published decoder/boundary components also fail. The 48-clip validation split and Hindi timing data were not used for this follow-up. Native acoustic equivalence does not establish word-timing quality. Serving is unchanged; licensing remains outside this workstream.
The preceding API change accepts five observed no-op null options for speaker, channel and vocabulary settings, normalizing them before persistence and idempotency checks. 159 app tests pass, with seven conditional model skips; typecheck and lint pass. A fresh local 74-job API replay covers paired omitted and null options on 30 reused Cap recordings and controls: 72 completions, two instances of the same known Cap error, 72 exact usage events, and 74 idempotent cross-form retries without another job or charge. All 37 pairs preserve words, timestamps, recognition confidence, language and status. The native model and serving release are unchanged. This closes a request-boundary gap, not the broader recognition, filler-default, API-feature or beta-readiness gaps.
The retained timestamp change adds a bounded, development-selected Spanish word envelope around accepted MMS acoustic cores. On 481 matched words in the existing 48-clip speaker-separated validation split, actual-hypothesis start/end MAE falls from 36.09/36.00 ms to 21.32/24.47 ms. Cached AssemblyAI is 56.84/57.03 ms on the same words. Recognition remains unchanged and still trails AssemblyAI on this Spanish set. The full local API replay has 100 completions, one known Cap error and 100 exact usage events, covering 64 reference clips, 30 reused Cap recordings and controls. This is a measured Spanish timestamp improvement, not general language/recognition parity or approval for a 20% Cap beta. The final envelope section records the complete evidence and limitations.
The subsequent native Cohere Spanish recognition pilot is not promoted. On the unchanged 16-clip development set it makes 7/175 word errors, versus 6/175 for current Sandchest and 5/175 for cached AssemblyAI. All four English regression outputs remain exact, but the raw candidate generates text on silence, tone and noise. Its 62 native requests and nine rejection/recovery checks pass; the accuracy gate fails, so the 48-clip Spanish validation split was not used for this candidate. Serving remains at the verified envelope checkpoint.
The combined-QKV MMS optimization is also not promoted. Across 97 aligned Spanish/Hindi development and regression cases, all word intervals and timing screen decisions stay exact, but the median per-case acoustic runtime improves by only 0.4%, short of the fixed 5% gate. Numerical drift also exceeds the preset tolerance. This experiment does not change serving or establish an API speedup. Independent Hindi word-timing validation remains open; licensing is not a gate in this workstream.
The preceding serving change reuses native PCM16 WAV/FLAC decoding for word alignment. All 158 before/after native pairs preserve transcript fields and alignment choices, including 30 reused Cap MP3s. The full local API replay has 31 successes and one previously known error, with 31 exact usage events. On 64 PCM16 WAV clips, native request p50 falls from 101.73 to 79.18 ms. This is a format-specific improvement; MP3 tail latency is not established as improved. The final decoder section records the unchanged-quality evidence and limitations.
The native Hindi candidate combines Qwen recognition with MMS word timing and bounded whole-recording composition in a private harness. The latest numeric pronunciation mapper raises short Hindi alignment coverage from 61/70 to 69/70 and complete timing acceptance from 50/70 to 52/70, without changing the recognizer or the 0.95 timing screen. Timing remains a blocker: none of the 12 windows of the difficult five-minute Cap recording passes completely.
The subsequent acoustic pronunciation experiment passed 95 focused tests and 228 native calls, preserving all 136 recognition results and the 78 timing cases outside its scope. Four of six eligible large numbers chose a different reading; one additional numeric word passed the conditional screen, but no additional complete clip passed. This strategy is not promoted. Higher acoustic path scores also made two numeric boundaries less concentrated, so a better path score alone is not a timestamp-quality gain.
The previous peak diagnostic showed that 65/77 uncertain Cap words remain too broad at any boundary center, so moving timestamps is not the main repair. No serving model, API route or beta traffic was changed. See the final acoustic pronunciation, numeric-target and timing/recording sections.
OmniASR CTC300M v2 now has an isolated native implementation, an exact tokenizer check, and a 136-clip acoustic/timing replay. It is not promoted. Its formatting-aware word endpoints pass the conditional timing screen on 55/70 short Hindi clips versus 52/70 for the canonical MMS baseline, with eight gains and five regressions. The difficult Cap recording still passes in 0/12 windows. More concentrated timing does not establish accuracy: the candidate's word ends are farther from cached AssemblyAI boundaries on the same matched words. One long input also remains outside the original numerical tolerance. The final Omni section records these failures, the precision diagnosis, and the measured scope.
The preceding independent aligner comparison also favors MMS: on 175 manually timed Spanish development words, start/end mean absolute error is 36.66/37.21 ms, versus 46.83/45.16 ms for Omni. A new native Whisper attention aligner was tested with Turbo and Large-v3, two tokenizers and two path constraints. None improves the baseline sufficiently for promotion. The corrected Turbo candidate has 59.63/49.04 ms start/end error on the same words. Its 336 native protocol calls and 46 failure/recovery checks establish scoped implementation behavior, not timing accuracy or beta readiness. See the final teacher-alignment section.
The latest native candidate work adds real token/word recognition scores and cooperative cancellation. Its private checks passed 270 score calls and 375 control/regression calls, preserving all 124 development transcripts and scores. This does not change the serving model or establish Hindi word-timestamp parity. The final native score/control section below records the exact scope.
The latest recognition-only experiment adds 100 previously unused Hindi references: guarded native Qwen has 15.50% WER, current Rust Turbo 32.52% and fresh AssemblyAI Pro 20.73%, using every completed transcript. All three completed 100/100. This is an experimental text-quality result, not a serving promotion: Qwen still lacks validated word timing and full API integration. See the final Hindi section for the 40-clip holdout, latency and timestamp diagnostics.
The latest local app checkpoint has 161 passing app tests, seven conditional
model skips; whole-app typecheck and lint pass.
The final native source has 112 passing Rust tests without Whisper, 159
with whisper-ctc and 175 with whisper-mms. The SDK bundle and declarations
also build from an isolated exact source snapshot.
Custom spelling now passes 76 live-provider fixtures, including exact text, word text and intervals on all 57 accepted cases, plus 12 exact sentence/paragraph resources. The subsequent 108-job actual local API replay produced 104 completions and four expected errors, with exactly 104 duration-based usage events and no failed-job charges. Both the unchanged AssemblyAI SDK and Sandchest SDK submitted and retrieved real transcripts. The 30 reused Cap recordings completed 29/30 in both baseline and spelling variants. Local baseline p50/p95 was 1.56/6.11 seconds, versus 1.31/6.19 seconds with spelling. This fixed-order, single replay does not establish a speedup. Recognition and original alignment are unchanged; custom spelling is display correction, not an accuracy gain. See the final custom-spelling section for scope and retained failures.
The preceding punctuation/export checkpoint passed 92 actual API jobs: 88 completions and four expected errors (the same known Cap recognition failure and an intentional language-threshold rejection, each with punctuation on/off). All completed results and sentence/paragraph exports retained the native words and acoustic metadata. Its 30 reused Cap development recordings are not a new holdout. It also passed 208 paired native calls and 12 standalone Parakeet calls. Engine and alignment defaults remain unchanged. The earlier 94-job automatic language and 263-job MMS checkpoints remain separately recorded below. The 25-case real API range check passed with 22 expected completions and three intentional unbillable range errors; four exact sample-cut comparisons and five reused Cap development regressions preserved the expected words and timestamps.
The latest broader fresh Cap comparison remains 55/59 Sandchest completions versus 57/59 AssemblyAI, with matched-completion local/remote medians 1.551/7.707 seconds. It predates the subsequent language/range boundary changes; it was not rerun or relabeled as new-model evidence. The new 40-clip public Indonesian/Finnish development comparison is independently referenced, but is not a Cap holdout. Details and limitations are in the dated sections below.
The general 20% Cap beta is not approved. Deployment/production checks were not run, as requested. Earlier build, lifecycle and benchmark checkpoints below are preserved as historical evidence rather than current universal guarantees. Resource tests cover the actual SDK over HTTP, missing/foreign IDs, nonterminal states, invalid queries, silence, CJK spacing, script marks and caption escaping. Derived resources preserve stored word boundaries and do not mutate transcripts.
A separate real-model run used the public API, durable local queue, native worker, isolated PGlite and unchanged AssemblyAI SDK. Nineteen jobs produced eighteen completions and one intentional language-threshold error. All ninety derived resource reads succeeded; eighteen unique usage events retained exact measured durations, and the failed job was not billed. This run used the rebuilt opt-in Rust Whisper backend with constrained DTW and punctuation separation. It does not establish deployed performance or change the default Parakeet worker. External billing was disabled and all owned processes were stopped. The latest run also listed all nineteen jobs across four pages and deleted one completed and one failed job through the unchanged SDK. Repeated DELETE/GET returned tombstones; all eighteen usage events retained their exact durations.
The expanded timing development set contains all 64 frozen AMI development clips, seven non-English natural recordings and JFK, plus silence/tone and known-offset controls. The baseline Whisper route failed six AMI clips because DTW assigned zero-length lexical intervals; direct native probes confirmed this cause. A constrained-DTW fix now included in the opt-in backend completed all 76 HTTP cases, with identical transcript text, token IDs and token bytes. Direct tests of the actual native function passed 50 matrix cases and 1,395 assertions against an exhaustive cost oracle, including cancellation and recovery; the original recurrence fails the regression test. Full transcription checks additionally cancel inside CPU graph, median filter and DTW processing, discard partial output, and recover to an identical fresh transcript. It requires each token row to consume an acoustic frame rather than padding durations after alignment. This is a reliability result, not proof that every selected frame is the correct word time.
The Rust punctuation mapper now separates displayed formatting punctuation from lexical timing evidence, while preserving timing for spoken symbols such as C++ and currency. On 961 matched words from successful baseline clips, offline remapping reduced word-end disagreement with cached AssemblyAI from 113 ms to 104 ms on average (p95: 333 ms to 268 ms). Disagreement with AMI's automatic alignment increased, so this is provider agreement, not a gold accuracy win. AMI's orthographic text is manually transcribed; its word boundaries are forced alignment, not human timing annotations. Failed jobs remain counted separately. The timestamp experiment uses unformatted native output and cached formatted provider output; it is not a new fully settings-matched recognition benchmark.
The earlier DTW-only route exposed a timing blocker: first words could start at zero through leading silence, and intervals could absorb pauses. The opt-in CTC integration below improves measured boundary agreement without fabricating times. Positive intervals and provider agreement still do not establish human-gold timing accuracy. The default route and Cap routing have not changed.
An additional comparison reused the existing native English CTC aligner on frozen Cohere hypotheses. On 1,005 words matched to both the AMI reference and AssemblyAI, start/end disagreement with AssemblyAI averaged 59/84 ms. Only 36 of 64 clips passed the existing full acoustic/timing acceptance screen. These are conditional agreement metrics, not human-gold accuracy or a production routing decision. Fallback, unsupported text, quiet/noisy input, long audio and cancellation still need a serving integration; this path does not supply the full language list.
The prior controlled 48-request ABBA comparison reduced the median native check for 60 seconds of silence from 117 ms to 86 ms; speech timings were essentially flat. This remains a local optimization measurement, not an AssemblyAI comparison.
Current resource compatibility exceptions are explicit: empty comma-separated word-search terms are rejected rather than reproducing the provider's invalid intervals; sentence/paragraph/caption grouping is local; captions keep actual stored word bounds instead of the provider's observed artificial caption shifts. Raw subtitle responses follow the documented text/plain contract. Remaining transcription options/defaults, redacted audio, gold timing, Linux/CUDA and deployed end-to-end performance still require implementation or validation.
Listing follows the observed newest-first cursor rules, preserves PostgreSQL microseconds when resolving cursors, and filters by tenant/status/UTC date. Page URLs use the configured public origin, never an untrusted request Host. Metadata is selected without loading full text/word arrays. Explicit null model selectors are accepted as omitted defaults.
Deletion keeps a nullable deleted_at tombstone and preserves the usage outbox,
measured duration and idempotency reservation. Worker claim/load/completion/failure
are fenced against deletion and stale attempts. Asset locking prevents a concurrent
CREATE from reusing media after deletion intent commits; uploads referenced by any
other live transcript are preserved. Physical storage work runs outside database
locks, with durable expiry intent for retries after failures or crashes. Regression
tests cover shared uploads, repeated deletion, private option/error scrubbing,
storage failures before/after physical removal, and late worker results. These
race tests use PGlite plus real HTTP with controlled inference fixtures; they do
not establish distributed Postgres/load-test coverage.
The additive nullable-column migration was separately tested in isolation, then applied to the shared audit Postgres under a coordinated window with one connection and short transaction-local lock/statement timeouts. Existing application-table contents, indexes and migrations 0000–0003 were unchanged. No existing transcript was deleted. Deploy this migration and asset-locking CREATE everywhere, then drain old creation requests before enabling DELETE during a rolling deployment.
Deletion exceptions are deliberate: private error details and arbitrary model-name strings are erased instead of retaining them on failed jobs; word search on a deleted transcript returns a stable 400 instead of the provider's observed 500. Deleted idempotency keys return 409 without recreating work or charging again. Public model identifiers, language and non-content request flags may remain as metadata, consistent with the provider's completed/error GET behavior.
The current requested milestone is a locally verified, limited Cap beta, potentially 20% of eligible requests, using broader read-only Cap production recordings. Do not deploy, probe production readiness, or change Cap routing during this phase. This milestone does not erase the full API/language parity gaps above. A percentage alone is not an eligibility rule: unsupported options/languages and unverified live modes must retain AssemblyAI until their separate checks pass.
A fresh run of the current local API completed all 38 existing Cap development clips for both native candidates (957.926 seconds of audio). Parakeet measured 530 ms p50 / 808 ms p95; Whisper measured 788 ms p50 / 1,562 ms p95. These include local upload, durable queue and persistence, with a 250 ms polling interval, but exclude model startup and deployment network/remote database costs. The paired AssemblyAI outputs were cached from the same audio bytes and settings; its old latencies are historical, not a current speed comparison. Lexical disagreement was 10.13% for Parakeet and 9.19% for Whisper, not ground-truth WER. This owner-heavy, mostly English sample has already been used for development and is not a holdout.
Cap's actual editor/live pure helpers accepted all 76 baseline API responses and 4,322 words, preserving times and metadata through editor serialization, caption generation, all-word offsets, repeated-chunk idempotency and two-chunk promotion. Those chunk placements are simulated; they do not replace real fragmented-audio inference testing.
Successful inference responses are now bounded before parsing and validated before persistence or usage: maximum three-hour duration, finite probabilities, positive ordered audio-bounded word intervals, bounded strings/word counts, valid UTF-8 and no PostgreSQL-incompatible NULs. Malformed or oversized successes fail once without usage. A real HTTP 200 response interrupted mid-body remains retryable; the tested valid recovery completes once and records exactly one measured-duration event.
An earlier English acoustic-alignment development experiment used cardinal normalization and per-word posterior screening increased usable timing coverage without changing recognition. A two-state selection path never increases adjacent overlap or changes word order/text/confidence. On the same 2,051 paired words, its start/end disagreement with cached AssemblyAI averaged 168/107 ms (baseline Whisper 262/113 ms), with start p95 reduced from 1,040 to 587 ms. Only 1,190 of 2,166 original words used CTC intervals; longer/unsupported/rejected words retained DTW. Conditional posterior concentration is not a calibrated correctness probability, and these are development/provider agreement metrics, not proof of better true timestamps. The later frozen candidate has now been integrated into the opt-in serving worker, as detailed below.
The frozen collection contains 59 usable Cap development recordings and 59 usable holdout recordings, with disjoint owners between the two splits. Two truncated recordings were excluded under the preset decode rule, without replacement. The holdout recordings were kept untouched until the frozen hybrid5 comparison below. They are now evaluated, and cannot serve as untouched final evidence for a candidate tuned after viewing these results.
A fresh, settings-matched, serial ABBA run sent the exact Cap options and audio bytes through both APIs. Sandchest completed 59/59; AssemblyAI completed 56/59 (three language errors). Local upload-to-result latency was 1.83 s p50 / 9.65 s p95, versus AssemblyAI 6.22 s / 15.58 s. Lexical disagreement on the 56 successful pairs was 9.67%, not WER. This was the earlier Whisper route without CTC.
The CTC integration subsequently completed all 59 cases through the real local API and durable queue, preserved all non-time output fields and non-English bypasses, and produced one exact-duration usage event per case. Selected intervals matched the frozen alignment prototype. It adds English alignment cost: local latency was 2.33 s p50 / 21.79 s p95; the high tail did not reproduce in a smaller native-only profile. That profile does not erase the API tail observation or prove its cause. AssemblyAI outputs were cached for this second run; do not compare cached provider latency with the new candidate as a fresh paired speed result.
On 64 public AMI development clips with manually transcribed text and identical provider options, current Parakeet v2 had 141/1,228 lexical word errors (11.48%), AssemblyAI 223/1,228 (18.16%), and Whisper 243/1,228 (19.79%). This is development WER, not held-out evidence. CTC applied to the exact Parakeet words preserved text/confidence and selected 1,152/1,229 word intervals; on the same 1,049 matched AssemblyAI words, mean start/end disagreement improved from 107/124 ms to 62/87 ms. Timeout and cancellation produced no partial candidate, and fresh retries recovered identical output. AssemblyAI boundaries are a comparison, not human timing gold.
An opt-in native composition now routes English to Parakeet v2 and retains Whisper
for other languages. Explicit English passes VAD and skips the Whisper encoder;
auto English uses the existing acoustic language policy before routing. Non-English
requests remain in the same Whisper full call and retain its first encoder result.
The English override is SANDCHEST_WHISPER_ENGLISH_ENGINE=parakeet with
SANDCHEST_INFERENCE_ENGINE=whisper; CTC remains a separate explicit opt-in. Invalid
configuration and model failures do not silently select another model. The composed
route requires GPU execution for both models; legacy Parakeet CPU mode is unchanged.
Health reports Whisper as the multilingual model and lists both actual models. The new
binary builds and passes native tests. The real composed regression preserved all
64 AMI outputs exactly and all 17 non-English Cap outputs, including identical
encoder counts in a separate default-Whisper pass. That first run exposed five
long recordings with backward times from the inherited chunk merger. The repair
preserves already ordered output, and otherwise chooses one aligned join between
unchanged model tokens. It does not sort words, clamp times or fabricate confidence.
The next complete native run returned valid bounds for all 59 Cap cases, preserved
all 64 AMI outputs and again matched all 17 non-English outputs/encoder counts.
On all 56 cached completed provider references, lexical disagreement was 9.49%
(not WER), and paired start/end disagreement averaged 115/103 ms (not timing gold).
One short recording remained empty in Parakeet. Its alternative Whisper words had
no lexical agreement with AssemblyAI and occurred before the PCM activity, so an
experimental Whisper fallback was rejected before use. The current candidate
instead returns an explicit, permanent uncertainty error for positive VAD followed
by an empty English decode. Its internal code is speech_recognition_uncertain;
the public transcript remains error, without words, measured usage or retries.
This does not assert that the recording contains no spoken audio. Tests separately
cover cancellation while reading that error and the same code on a retryable 503.
A beta client still needs a verified, explicit provider fallback for these results.
Original Cap segment testing collected nine completed recordings (453 original fragments); a tenth frozen selection failed the manifest-duration rule and was not replaced. Actual Cap planning produced 126 chunks plus nine complete assemblies. All 135 inputs decoded. A fresh API comparison of 27 first/middle/tail chunks plus nine full files completed 36/36 locally and 30/36 at AssemblyAI. All local word intervals were valid; there were eight empty results on each side, with no empty mismatch. AssemblyAI returned a no-spoken-audio error for six of those, whereas Sandchest completed them empty. Cap handles these differently for full recordings; that semantic difference is unresolved. Do not infer provider billing from status. The local database recorded 36 unique exact-duration usage events, including empty completions. Cap's actual pure consumer helpers then accepted all 36 pairs: 66 VTT conversions and 49 completed-chunk mappings/replays, with 741 mapped words and no clipping, dropped words, changed offsets or replay drift. All Sandchest raw words were valid; AssemblyAI had six raw boundary issues handled by Cap's normal mapping. The silent full-file status difference appeared once (the other five errors were chunks). This is helper execution on saved API output, not a production workflow. On 30 successful pairs, local latency was 0.79 s p50 / 6.51 s p95, versus AssemblyAI 3.19 s / 13.58 s. Lexical disagreement was 9.03% overall and 16.71% for chunks, not WER. This run does not prove sticky-language workflow behavior.
The rebuilt English composition also ran through the same 36 original-fragment API cases with fresh AssemblyAI requests. It completed 35/36, with one uncharged uncertainty error; AssemblyAI completed 30/36. Every native completion had valid word bounds and exactly one measured-duration usage event. The 29 mutually completed pairs measured 0.78 s p50 / 4.15 s p95 locally versus 3.22 s / 17.13 s for AssemblyAI. These are local transport timings, not deployed capacity. Using the same fresh provider references for both native candidates and counting the failed native case as an empty hypothesis, disagreement was 324/3,623 tokens (8.94%) for the composition versus 332/3,623 (9.16%) for cached Whisper+CTC. This is provider disagreement, not WER; the common-reference comparison does not reuse old latency as a fresh benchmark. Both owned processes were reaped, all private keys revoked, and all frozen source/input hashes remained unchanged.
Earlier beta1 process-crash tests killed and restarted the API while a job was running, then separately killed and restarted native inference. Both recovered identical results in two attempts and charged exactly once. API crash recovery waited for the normal 120 s lease (128.4 s measured). Eight concurrent clients completed with local p95 3.15 s through one inference lane; all ten jobs in the combined run had exactly one usage event. Retry spacing is now 10 s then 20 s so ordinary model startup does not exhaust attempts. These are isolated PGlite/local process tests, not distributed PostgreSQL or deployed capacity proof. At that checkpoint the separate queue entrypoint still needed graceful-shutdown and SQS-outage work. The user subsequently authorized the edits; the new local verification below executes that entrypoint directly.
The immutable hybrid5 binary is SHA-256
368a64c8c6e5feb765a73f0faf6168cf46e954bf3ab9811d2cc0494e7b16856b.
Models, application/native sources, scoring, inputs and request settings were frozen
before both comparisons. Both APIs received the same audio bytes and settings;
provider order alternated AB/BA, with 250 ms polling. These are real local TCP
upload, durable queue, inference, persistence and polling measurements on M4 Max,
compared with remote AssemblyAI. The local host uses the actual API handler and
queue processor with isolated PGlite and external billing disabled. These timings
exclude model cold start, remote PostgreSQL, SQS and live credit-gate roundtrips;
they do not establish cloud or deployed latency.
| Held-out set | Sandchest | AssemblyAI |
|---|---|---|
| Cap completions, 59 recordings from 56 owners | 57/59 | 55/59 |
| Cap paired p50 / p95, 53 mutually completed | 1.31 s / 8.04 s | 6.76 s / 18.35 s |
| English AMI completions, 128 clips from four groups | 126/128 | 128/128 |
| AMI lexical word errors / 2,406 human-reference words | 173 (7.19%) | 293 (12.18%) |
| AMI English-normalized WER, secondary metric | 6.22% | 8.83% |
| AMI paired p50 / p95, 126 mutually completed | 0.53 s / 0.55 s | 3.15 s / 4.49 s |
AMI WER includes every planned clip: both providers' failed jobs would be scored as empty hypotheses. The two native uncertainty errors therefore contribute nine missing reference words rather than being removed. The paired group-bootstrap 95% interval for the native-minus-provider WER difference is -6.90 to -3.03 percentage points. Four underlying groups limit generalization; this is evidence for this English meeting set, not all Cap recordings, accents or languages.
Cap has no independent human text reference. The 53 mutually completed cases had 11.67% provider lexical disagreement (1,729/14,812 tokens), not WER, after an offline Unicode scoring correction. The original frozen report remains intact: its tokenizer split Hindi vowel marks from their words and reported 13.60% (2,059/15,142). Preserving attached Unicode marks fixes the units; it is not a model-quality improvement. No English AMI reference or hypothesis tokenization changed, and its 173 versus 293 errors remain exactly unchanged. The audit retained and verified 519 source/input hashes. CJK needs separate character-based scoring; this correction is not a universal language-specific word segmentation algorithm. Word-boundary disagreement on 13,141 paired words averaged 106 ms at starts and 110 ms at ends; p95 was 377/305 ms. These are comparisons with AssemblyAI, not human-gold timing errors. The English provider-labeled subgroup had 9.62% lexical disagreement; subgroup labels come from AssemblyAI and are not independent correctness judgments.
The Cap language breakdown blocks general automatic-language traffic: three recordings identified by AssemblyAI as Spanish, Indonesian and Hindi were returned as English by Sandchest. The Indonesian and Hindi text disagreements were 69% and 88% with the corrected scorer (the legacy Hindi number was 93%). One Indonesian response had an English language probability below 0.001. The caller's expected-language list excluded Indonesian, so the native hard-filter behavior and the provider's observed response diverged. AssemblyAI's official language-detection documentation says detection is restricted to the supplied list; saved request receipts and the provider's echoed options confirm that this run nevertheless returned Indonesian. Do not silently reinterpret expected languages as soft hints or tune a confidence threshold on this holdout and still call it untouched.
Every native completion in both holdouts passed the response/timestamp checks. The Cap and AMI databases contained respectively 57 and 126 unique measured-duration usage events; failed jobs retained no completed content and incurred no charge. All 59 and 128 jobs were terminal at one attempt. Ephemeral credentials were revoked, all owned processes were independently confirmed absent, and all 214/480 frozen input/source files remained unchanged. Provider response-boundary warnings are retained separately rather than attributed to native failures.
Private evidence is retained under paired-api-holdout-cap1,
paired-api-holdout-ami1 and unicode-scoring-audit1 on the session RAM volume;
customer media and text must not be published.
The frozen hybrid5 also passed its own real process-recovery run on the fixed nine development recordings. Nine fresh baselines, two crash jobs and eight concurrent submissions produced 19 terminal jobs, 17 exact-duration usage events and two intentional, uncharged uncertainty errors. Both crash jobs recovered the same complete stable resource (apart from IDs, upload URL and creation time) in exactly two attempts. Killing the combined API/embedded-queue process took 123.85 s to recover, including the normal 120 s lease; native kill-to-result was 13.78 s, including model restart. The eight-client burst had completed-result p50/p95 of 2.08/2.84 s, with seven completions and the expected one uncertainty error. Failed idempotent replay did not recreate a job. All 470 inputs stayed unchanged and all four owned PIDs were independently confirmed absent. This still does not exercise SQS, SIGTERM, the standalone queue entrypoint or separate database/worker hosts.
A ten-call, native-only post-holdout diagnostic reproduced all three language mismatches exactly. Whisper alone with the same automatic-language options did not resolve them. Explicit provider-language selection improved the Spanish case, but Indonesian remained far from provider agreement, and explicit Hindi returned native output-validation code 8. No model or default route was changed; these selected cases are development diagnostics, not another holdout.
Three bounded native diagnostic calls isolated the Hindi error: two of 24 segments contained invalid UTF-8, and the complete concatenated text was also invalid. The final diagnostic passed all prior C++ guards before strict JSON serialization failed with error 316. Joining segments cannot repair this output safely. No characters, words or timestamps were guessed, dropped, clamped or replaced.
The subsequent hybrid6 build only changes error classification: native output
integrity failures now become the same permanent, uncharged uncertainty result as
an empty English decode. Compute/allocation failures keep their existing retry
behavior, and cancellation checks still take precedence. Its SHA-256 is
36d667302af96f40879959c03241c8f5d3b7c5b37b10f49af8d35905e33a6b09.
A real four-job API regression produced three completions and one one-attempt
Hindi uncertainty error. Error replay returned the same resource without another
job or usage event; English and Spanish controls retained their exact successful
outputs, and the same worker recovered without restart. All 222 frozen files and
both owned process exits were checked independently. This is failure-handling
verification, not repaired Hindi recognition. The quality, latency and crash
benchmark tables above remain explicitly tied to frozen hybrid5, not a fresh
hybrid6 benchmark.
The application still needs a verified Cap eligibility/fallback integration. The user authorized edits to the pre-existing queue files on 2026-08-27; their original contents were preserved before claiming and changing them. The standalone queue defects described earlier are addressed and separately tested below. The older local API harness does not execute that entrypoint, so its prior benchmark results must not be relabeled as standalone-worker verification. An explicit-English full-recording beta is a narrower candidate scope, not yet a readiness verdict. Automatic-language and live/sticky-language traffic must remain with AssemblyAI until their separate gates pass. A proposed 20% is a fraction of eligible requests, not permission to route arbitrary requests.
No limited-beta readiness verdict has been granted.
Database draining now starts independently of the SQS listener, polls at most one second after an empty scan, and retains wakeup hints received during an active batch. Failed receive or acknowledgement calls cannot gate durable jobs. Creation returns after the database transaction without awaiting SQS. Notifications are coalesced, and each operation kind retains its slot until the actual SDK promise settles, including when middleware ignores cancellation. Request deadlines, transport cancellation and disabled SDK retries bound the notification path.
Billing and retention each run non-overlapping sweeps outside the inference lane. SIGINT/SIGTERM reaches the core processor and stops new claims. A graceful interruption refunds its retry allowance under the existing live lease fences, without reusing an attempt token or charging incomplete work. The regression suite specifically cancels the last allowed attempt and then verifies one billed recovery, including when the API inline mode was enabled. The standalone process explicitly disables that mode, and interruption cannot schedule an uncancelled inline retry. Classification uses the actual first cancellation cause captured before cleanup, so an ordinary network error followed by shutdown retains its cap and backoff. Ordinary deadlines and exhausted crash attempts also retain their retry caps. Shutdown joins active work before closing the database; an unref'ed 25-second watchdog covers abort-ignoring handles.
The private queue-entrypoint3 verifier launches the actual workers/queue.ts
process against isolated PGlite, a local SQS HTTP fixture using the real AWS SDK,
and synthetic inference responses. It does not run a model or real AWS:
- Three API-created jobs and three idempotent replays produced exactly three jobs. Creation took 3.79–24.48 ms despite a stalled SQS send; the burst sent one hint.
- During an SQS receive failure, SIGTERM cancelled active inference and the worker exited successfully in 44.13 ms. No next job was claimed and no usage was written.
- Restart processed the remaining work despite failed acknowledgement and a stalled long poll: two completions, one expected uncertainty error, exactly two 1,000 ms usage events, and no charge for the cancelled attempt or failed job.
- Shutdown while the long poll was stalled completed in 20.69 ms. All five child processes were independently checked absent, keys revoked, and 25 frozen files unchanged. The timestamps are lifecycle diagnostics, not production latency.
A separate queue-watchdog2 process test deliberately held an SDK promise and
referenced timer after abort. The real worker drained in 805 ms and the watchdog
terminated the retained handle with exit status 1 after 25.06 seconds. This is an
expected forced shutdown, not a successful graceful exit. The process was reaped.
These tests close the local entrypoint gap; distributed PostgreSQL, AWS SQS and
cloud deployment validation remain unperformed as requested.
The serving candidate remains hybrid6. None of the following private experiments has been promoted, and the frozen hybrid5 holdouts remain unchanged:
- Language-window/model probe: 34 native calls on eight known Cap recordings compared Turbo with the pinned large-v3 Q5 artifact. On the long Hindi case, Turbo selected English in the opening window but Hindi in the middle and final windows. The larger model selected Hindi in all three windows. Controls retained their English, Arabic, Portuguese and Spanish detections. This supports testing broader acoustic sampling; it does not establish the correct language for every bilingual recording.
- Larger model rejected as a default upgrade: the Hindi transcript completed, but had 423 lexical differences against 493 AssemblyAI reference tokens and took 75.6 seconds explicitly or 128.8 seconds with automatic language selection. These are diagnostic native timings, not paired API latency. Spanish and Indonesian results were mixed. Provider differences are not human-reference WER.
- UTF-8 sampling prototype: its strict byte-state tests passed 1,177,605 sequences, including every Unicode scalar and every byte split. Ten native baseline/constraint calls still reproduced the long Hindi output error. A separate numeric guard probe confirmed malformed segment text. The filter also changed an Arabic control and token probabilities; it cannot be treated as an output-only repair or a quality-neutral change.
- Decoder capacity probe: using the unused context budget when prompt history is disabled increased the private budget from 220 to 441 generated tokens. Both baseline and constrained variants still failed the Hindi case, across ten calls. This did not justify changing the serving decoder limit.
The private reports are multilingual-model-dev1, unicode-constraint-dev2,
unicode-constraint-diagnostic1 and context-window-dev1 on the owned CTC RAM
volume. Their input hashes and stopped model processes were independently checked.
The failed first byte-test attempt is retained separately: its incremental Python
oracle deferred rejecting a surrogate prefix, and no model calls occurred in that
attempt. The corrected oracle derives prefixes from canonical Unicode encodings.
A completed published-reference Hindi development pilot used ten predetermined Google FLEURS test recordings and two concatenation variants through both real HTTP APIs. It supplements the private Cap comparisons; it is not new Cap holdout evidence. The ten unique clips total 130.5 seconds. Their 139.5-second concatenation adds nine one-second gaps and is reported separately, not counted as independent additional material.
| Hindi development case | Sandchest lexical WER | AssemblyAI lexical WER |
|---|---|---|
| Ten individual clips, explicit Hindi | 30.50% (86/282) | 21.99% (62/282) |
| Same clips concatenated, explicit Hindi | 41.13% (116/282) | 21.28% (60/282) |
| Same concatenation, automatic language | 105.32% (297/282) | 20.21% (57/282) |
These scores use published raw transcripts with the same Unicode-mark-preserving lexical normalization for both providers; formatting differences remain included. The automatic Sandchest response selected Urdu (confidence 0.7423) and Arabic script, while the reference and AssemblyAI use Hindi/Devanagari. That script mismatch dominates the very high automatic WER. Explicit Hindi still trails AssemblyAI, so better language detection alone does not close the quality gap.
All 12 local jobs and all 12 AssemblyAI jobs completed. The local audit found one exact-duration usage event per job, one attempt each, no invalid native completed word contracts, revoked test keys and stopped owned runtimes. A separate read-only audit checked all 24 request/response identities, actual model attribution, provider-specific timestamp tolerances, exact published audio/reference bytes, ten distinct audio/source IDs, and the concatenation against newly decoded PCM. Wrong-ID, wrong-model and zero-interval negative controls were rejected. Original reports remain unchanged; the independent audit verifies 253 frozen files.
Single-clip local API p50 was 1.051 seconds versus 3.980 seconds for AssemblyAI; this ten-clip development sample is not a production latency estimate. The explicit concatenation took 7.556 versus 17.498 seconds. Fast completion does not offset the measured recognition gap. The local runtime still uses isolated PGlite and disables external billing; it is not the standalone SQS queue path.
A subsequent 26-call acoustic-language probe detected Hindi on eight of ten
individual clips with Turbo and nine of ten with large-v3 Q5. Averaging Turbo's
three fixed windows selected Hindi for the concatenation; large-v3 selected Hindi
in all three windows. This gives a concrete next experiment for automatic
language routing, but no new detector or three-window policy is enabled in
serving. Silence handling, cancellation, overhead and unseen Cap regressions must
be checked before promotion. Private receipts are fleurs-hi-api1/report.json,
fleurs-hi-api1/independent-audit.json and fleurs-hi-lid1/report.json.
The guarded native SDK no longer selects decoder zero when all beams fail. The original 220-token decoder budget could leave every beam failed, yet emit an unfinished sequence; the diagnostic observed seven failed selected beams on the known long Hindi case, including two ending in partial UTF-8. The new guard resets selection for each temperature, requires bounded live candidates and finite scores, and validates the exact emitted text prefixes and segment boundaries. It never repairs bytes or manufactures tokens. Exhausted candidates clear all partial segments and return an explicit output-integrity error. Existing no-speech suppression stays successful and empty; cancellation takes priority. The SDK extension is versioned and the Rust bridge refuses a stale SDK.
The patch includes model-free tests for every non-NUL Unicode scalar, malformed encodings, split scalars, segment boundaries, NUL, prefix bounds and score selection. Twelve calls through the actual decoder and bridge additionally test all-failed beams, a surviving lower beam, failure after an earlier completed window, temperature retry, nonfinite scores, invalid bounds, cancellation, an empty temperature list, no-speech versus emitted invalid text, and recovery. All expected outcomes passed and partial results were cleared.
The resulting hybrid7 binary is
a52a7937d38c0dba8c73fd4831f945df2ff1d68900b44e1845f34e31ad26a51e.
A full local HTTP/API/queue/PGlite replay used the same 59 previously evaluated Cap
recordings, plus an explicit Hindi failure and a following recovery request.
56/59 original recordings completed; every surviving completion retained its
exact previous text and word timestamps. The two earlier failures remained,
and one previously completed Spanish case now fails the output guard. That case
needs successful recovery or fallback before beta; safer rejection does not
establish sufficient availability. The explicit Hindi error was terminal at one
attempt, its idempotent replay returned the same resource, and the following
request recovered to unchanged output without restarting the worker.
All 61 jobs were terminal at one attempt, with exactly 57 duration-matched usage
events, no failed-job charges, no persisted partial text/words on failures,
revoked isolated keys and stopped owned processes. The loaded SDK path was
checked on the real worker process; 325 frozen files were unchanged. Completed
local API p50 was 1.175 seconds, but this is a regression replay with cached
AssemblyAI outputs, not a new paired hosted-speed comparison. The older hybrid5
quality and latency tables remain tied to that older binary. Receipts are
whisper-beam-guard2, whisper-guard-runtime1, cap-guard-regression1, and
beta-app-suite14 in the private local evidence directories.
The current hybrid8 candidate retains the output guard and adds at most six sampling temperatures (0 through 1 in increments of 0.2), used only when every candidate fails. It does not lower the entropy threshold, accept invalid bytes, expand the token context, or resample a valid first-pass candidate for a low average log probability. The SDK extension is now version 2; its opt-in policy preserves valid output, while the disabled policy retains the SDK's usual low-log-probability fallback. Both policies still reject exhausted invalid output. Cancellation and the existing request deadline remain in force across attempts.
The Spanish regression came from the existing repetition heuristic: all five first-window beams had entropy 2.293, below the 2.4 threshold. A guarded retry recovered the recording. Its lexical disagreement with cached AssemblyAI fell from 52 to 47 edits over 265 provider words; this is not ground-truth WER.
The exact current binary is
19495a30d707d969b5697618d9675bd11e51ead017fe8c7ae1926a919830982d.
A fresh local HTTP/API/durable-processor/PGlite replay completed 57/59 original
Cap recordings. The Spanish case is the only changed successful transcript;
the other 56 successful cases retain exact text and word timestamps. The two
original failures remain, and explicit Hindi still returns a permanent uncertainty
error. Both successful and failed idempotent replays preserve their original
resource, and a following English request returns unchanged output without a
worker restart. All 61 jobs finish at one attempt, with 58 unique duration-matched
usage events, zero failed-job charges, no persisted partial text/words on errors,
revoked test credentials and stopped owned processes. All 325 frozen inputs
remain unchanged. This reused development sample is not a new quality holdout,
and external billing was disabled.
The current local API replay measures 1.169 seconds p50, 21.383 seconds p95 and 38.927 seconds maximum among completions. The slow tail was inside inference, including unchanged English recordings. A separate counterbalanced old/new/new/old comparison then ran those four long recordings twice per build, with one warm-up per process. All 20 jobs completed once, with exact usage and unchanged text/word output. The eight measured requests per build had API medians of 8.394 seconds for hybrid7 and 5.809 seconds for hybrid8; individual repeats varied substantially. This did not reproduce a consistent regression in hybrid8, but the shared host and small repeated sample do not justify a speedup claim or erase the slower full-run tail. Neither run made fresh AssemblyAI requests or measured deployed latency.
The twelve prior direct decoder cases pass with this recovery policy. Eight more real decoder/bridge calls verify that valid low-log-probability output retains all semantic fields, including token probabilities and timestamps; the ordinary SDK policy still retries it; later-window recovery preserves earlier complete segments; cancellation between attempts wins; no-speech windows are not retried; and a following request recovers unchanged. Elapsed runtime telemetry is excluded from semantic equality. A first harness attempt incorrectly included that timing field; its failed receipt and original helper are preserved separately.
Private receipts are whisper-recovery1, whisper-recovery-runtime1,
whisper-recovery-edges2, cap-recovery-regression1, cap-recovery-tail1, and
beta-app-suite15. Context-budget expansion remains an unpromoted experiment:
it helps selected public Hindi clips, but did not recover the difficult long
Cap Hindi case and needs separate prompt/history/DTW bounds validation. The
current candidate is still not approved for a general 20% Cap beta; language
quality, word-timing validation, exact API features and safe fallback remain open.
Native Rust Qwen3-ASR-1.7B, pinned 8-bit weights and a corrected frontend, is a promising recognition candidate. It is not the current serving worker and has no validated word timestamps or full API integration. Explicit Hindi, greedy sampling and a 1,024-token ceiling were fixed before testing; completed requests reached EOS. References were never supplied as prompts. Both providers received the same canonical 16 kHz PCM audio.
| Public Hindi sample | Native Qwen lexical WER | AssemblyAI lexical WER |
|---|---|---|
| Original ten development clips | 14.54% (41/282) | 21.99% (62/282) |
| 64 new clips, selected before outputs | 15.41% (241/1,564) | 19.76% (309/1,564) |
All 64 new requests completed on both systems. Selection was deterministic from 418 FLEURS test rows, with unique source IDs and lexical references, excluding all ten earlier references and IDs. Independent review reread the source parquet, verified every original audio/reference byte stream and provider identity, and recomputed every score. Qwen improved 39 clips, tied eight and worsened 17. A 100,000-resample paired clip bootstrap gives an AssemblyAI-minus-Qwen WER difference of +4.35 percentage points, with percentile 95% interval +1.86 to +6.95 points. This interval applies to these selected public clips; speakers and corpus conditions can still repeat. It is not a Cap-domain confidence interval.
The native median for the original ten clips was 912 ms, excluding startup, decoding, IPC and all API work. No native-only timing is presented as hosted API parity. The 64-clip receipt froze 141 files; independent review additionally verified extracted source audio and runtime libraries. Future receipts should include those latter files directly in the before/after freeze.
Long-audio testing used audio-only quiet cuts between 20 and 29 seconds, with complete sample coverage and no dropped chunks. The duplicate public concatenation scored 38/282 (13.48%) against its published references. However, the known five-minute Hindi Cap recording still had 323/493 (65.52%) lexical disagreement with cached AssemblyAI, including 147 insertions; that is not human-reference WER and is far from reassuring provider agreement. It took 24.8 seconds of native calls. The unguarded Qwen prototype also emitted one word on each synthetic silence and tone control. Therefore this candidate is not promoted: it needs speech gating, multilingual word alignment, language routing, Cap-domain validation, cancellation and serving-grade error handling. The bounded long-audio reader now has an actual byte limit/deadline; the older short-pilot line reader remains a documented private-harness limitation.
The larger Whisper large-v3 Q5 alternative did not close the short-Hindi gap:
87/282 errors (30.85%), versus Turbo's 86/282 (30.50%), with native median 2.498
seconds versus 676 ms. A larger token budget reduced concatenated Turbo errors
from 116 to 80, while blind 15-second cuts worsened recognition. Neither decoder
change is enabled in serving. Private receipts are fleurs-model-quality1,
fleurs-window-quality1, qwen-hindi1, qwen-hindi-expanded1, and
qwen-hindi-long1. Public-corpus gains must not be substituted for the remaining
Cap, word-timing and full-contract acceptance gates.
A bounded Rust/MLX implementation of the pinned Hindi Wav2Vec2 checkpoint
(theainerd/Wav2Vec2-large-xlsr-hindi, revision
062f7f566e2671336992b011dcb9387cd3cffe5e) now matches an independent CPU
Transformers reference. All 424 tensors load strictly. The 19 acoustic cases
include ten reused public Hindi clips, silence/tone/noise, known offsets,
repetition, and the 400-sample/35-second bounds. All frame argmax IDs match;
maximum absolute log-score difference is 0.000500 and maximum RMSE is 0.0000221.
The predeclared tolerances were unchanged. Deadline/fresh recovery passes.
The ten speech clips have a native acoustic-only median of 82.4 ms. This excludes ASR, word alignment, media decoding, startup and the API; it is not hosted latency. The private acoustic driver now publishes each case through an owned staging directory. Partial writes and cancellation leave no published result and permit a same-ID retry; existing completed outputs are preserved. This is a sequential experiment driver, not a distributed idempotency claim.
The word mapper preserves original surfaces, Unicode scalar offsets and Hindi combining marks. Unsupported lexical characters reject the entire candidate; there is no silent deletion, guessed timing or interpolation. Of ten public Qwen hypotheses, eight align; two contain Hindi characters absent from this checkpoint's vocabulary (U+091E and U+0949). All eight aligned cases give identical CTC paths and word intervals from native versus reference acoustic scores. Adding 800 ms of silence preserves boundaries exactly; the repeated-utterance probe differs by at most 20 ms. These shift tests are not human timing labels, and their complete conditional timing screen does not pass.
Only four of ten public clips pass the conditional timing screen. Worse, one of six deliberately wrong-text controls also passes it. Posterior mass conditional on a supplied transcript is therefore not an acoustic correctness gate; no Hindi mismatch threshold was fitted or promoted from these reused examples.
The known five-minute Cap Hindi recording is a decisive coverage failure: only 3/12 chunks align, covering 42/610 whitespace-delimited hypothesis words; zero chunks pass the complete timing screen. Latin code-switching and missing Devanagari vocabulary account for the rejected chunks. The concatenated public sample aligns only 2/6 chunks, also with no complete timing-screen passes. Partial chunk output is diagnostic and cannot become a successful transcript.
Review found and fixed an evaluation error: the original driver grouped all CTC
failures with expected rejection. The corrected driver treats only unsupported
characters, explicit acoustic rejection and NoPath as expected; cancellation,
deadline, invalid input, allocation and invariant failures invalidate evaluation.
Nine private Rust tests pass. A 56-call replay binds all actual fault arrays and
payloads, rejects three injected fatal faults, verifies a typed no-path rejection
and fresh recovery, and reproduces all 51 timing results. Historical comparison
is explicitly to the replay-start snapshot, not an earlier immutable inventory.
The two public baseline/oracle failures must be exactly the two known codepoints;
paired failures can no longer disappear from the denominator.
Receipts are multilingual-ctc2, hindi-ctc-parity3, hindi-ctc-alignment1,
hindi-ctc-cap1 and hindi-ctc-alignment-checked2. This prototype does not change
the serving backend, application tests, API behavior or beta verdict.
The long concatenated public sample exposed a separate defect: the tokenizer's byte decoder silently replaced invalid UTF-8 with U+FFFD while reporting success. An isolated native variant now validates the concatenated token bytes before returning text. It preserves multibyte characters split across tokens, rejects negative/unknown token IDs and missing tokenizer configuration, and checks every accepted result against the exact loaded tokenizer. It does not repair, drop or replace invalid bytes. Model math, weights, vocabulary and sampling are unchanged.
Two focused Rust tests pass. A real native replay covers all 74 prior public
clips plus 20 prior Cap/concatenation/silence/tone chunks and a fresh recovery call.
The 93 valid prior outputs and recovery retain exact text. The one previously
corrupted output now returns a typed invalid_utf8 error with no partial text.
All 95 calls finish and the owned process exits; source, runtime, model and input
hashes remain unchanged. Receipts are qwen-utf8-guard1 and qwen-utf8-replay1.
This closes silent output corruption in the private candidate, not its silence
hallucinations, word timing, language routing, batched decoding, cancellation or
full API integration gaps. There is no serving promotion or new quality holdout.
A second Rust/MLX alignment prototype uses the official TorchAudio MMS_FA
checkpoint. Its 423 F32 tensors load strictly into the pinned official model;
conversion to safetensors preserves every retained tensor exactly and removes
only the three auxiliary classes specified by TorchAudio. The optional wildcard
class is disabled. Native code uses the model's waveform layer normalization
with epsilon 1e-5, not the different Hindi feature-extractor normalization.
Official MMS_FA bundle.
All 28 acoustic comparisons against the independent official CPU graph pass unchanged tolerances: ten reused public Hindi clips, eight windows from seven Cap recordings in seven languages, and ten silence/noise/offset/length controls. All frame argmax IDs match. Maximum absolute log-score error is 0.003437 and maximum RMSE is 0.0001569. Native acoustic-only median across the 18 speech windows is 64.6 ms (range 44.5–256.2 ms); it excludes recognition, alignment, startup, decoding and the API and is not an AssemblyAI latency comparison. The deadline and exact fresh-recovery check also passes.
The first 49-call word-alignment evaluation produces intervals for all ten public Hindi hypotheses (284 whitespace-delimited words), all 12 chunks of the difficult Cap Hindi recording (610 words), and all six other Cap windows (137 words). The previous Hindi-specific model covered only 42/610 words in that recording. This is improved alignment coverage of supplied hypotheses, not improved WER or verified timestamp accuracy. All ten public cases give identical CTC paths and word intervals from native and official acoustic scores. An 800 ms prefix shifts boundaries exactly; repeated speech differs by at most 40 ms.
Seven public clips and two of the six other-language Cap windows pass the complete conditional timing screen. No complete Hindi Cap chunk passes: 77 of 610 word intervals have an uncertain start or end under this screen. All six wrong-text controls and the tone/noise controls fail the screen; digital silence is rejected. This small negative set does not calibrate a correctness gate. Conditioning alignment on a transcript can still force incorrect words into plausible intervals. No acceptance threshold or production fallback changed.
A pinned Rust uroman mapper preserves surface words and Unicode-scalar offsets through many-to-many romanization edges. It does not zip two independently split word lists or manufacture durations. Numbers without validated spoken forms, unknown symbols, private-use characters, unapproved fallback deletions, lost lexical words, and edges crossing word boundaries reject the complete candidate. Only an explicit set of common accent marks and the two joiners may disappear through fallback rules. Named upstream romanization rules remain a separate trust boundary. No CJK word segmentation or numeric verbalizer is claimed. The experimental mapper is not the complete MMS training text-normalizer.
Six alignment tests and five acoustic-driver tests pass. A stricter 66-call replay retains exact semantics on all 49 original cases, verifies 12 explicit normalization rejections and four fatal faults, and recovers exactly afterward. Cancellation/deadline, malformed arrays, allocation and invariant errors are not quality skips; failures contain no partial intervals. The historical baseline is explicitly the result snapshot frozen at replay start, with the original source versions preserved by hash.
The raw romanizer was separately compared with pinned original Python 1.3.1.1 and Perl 1.2.8 implementations on 94 cases, including three input guards. Of the 91 oracle cases, Python raises its existing incomplete-Chinese-percentage error on one; 89 of the remaining 90 strings and 87 edge paths match. Differences include documented Tibetan behavior and numeric edge segmentation. All 28 hypotheses used above match Python strings and edges exactly, but only 18 match the legacy Perl strings. Therefore this is not a claim of exact model-era normalization parity or measured support for AssemblyAI's full language list.
The current native acoustic binary is c1e91164…, with model SHA-256
8d3966ea…; the guarded alignment binary is 23415d9c…. Private receipts live
under the owned MMS experiment storage (parity1, alignment1,
alignment-checked3, romanization2) with exact sources and full hashes. No
serving code, model selection, public API, shared database or deployment changed.
The general 20% Cap beta remains unapproved.
The follow-up retained the full cached Cap cohort and added independent public word-timing references. It did not fetch new Cap recordings, change serving code, deploy, access the shared database or change port 3000.
Hindi recognition is separate from timing certainty. The 74 previously measured Hindi clips received 444 new alignment requests: reference, Qwen, AssemblyAI and three deliberately corrupted transcripts per clip. All 72 rejections were typed normalization/unsupported-character failures, with no fatal failures. In the 64-clip validation group, the timing-mass screen accepted 19 inserted-phrase cases. A cutoff derived only from development references rejected those insertions, but still accepted 27 deleted-word cases. Neither screen is a recognition correctness guarantee; neither was promoted. The existing failure-inclusive Qwen/AssemblyAI WER remains 241/1,564 (15.41%) versus 309/1,564 (19.76%). Accepted-subset WER is not a whole-corpus improvement. The audit separately counts hypothesis words, reference tokens, deletions and rejected-case errors; the word-uncertainty table is not error-detection recall.
Full cached Cap alignment still has gaps. All 59 records remain in the inventory: 52 nonempty completions, five empty completions and two existing errors. The 317 windows produce 257 candidates and 60 normalization rejections, covering 11,212/14,165 original words; 1,117 candidate words have an uncertain boundary. Only 27 nonempty records have complete, monotone candidates, and four also pass the conditional timing screen throughout. Cross-window overlaps remain explicit failures rather than clamped timestamps.
Of 13,145 lexical pairs eligible before MMS, only 10,397 have candidates; 2,748 are excluded for missing candidates. Both timing methods are compared on those same pairs. The earlier provider validator flags word bounds in 13 provider records, so the audit also separates the unflagged subset: 6,474 candidate pairs from 42 provider records. On that subset, mean start/end disagreement changes from 78/102 ms to 84/103 ms, while p95 changes from 196/230 ms to 134/202 ms. These are provider-agreement diagnostics, not human timing truth or a uniform improvement. All 59 provider identities were reconciled to original audio hashes, provider report rows and frozen request plans/settings.
New manual Spanish timing references. A selection frozen before model outcomes contains 64 DIMEx100 clips from 16 speakers: 16 development clips and 48 validation clips from 12 different speakers. The 669 lexical intervals were manually annotated; original labels/times are retained, with reversible decoding of the corpus's accent notation. Full original audio is used. The data is short studio Mexican Spanish, not Cap speech; model-pretraining overlap is unknown. DIMEx100 primary description.
With exact reference text, MMS aligns all 669 words. On the 494 validation words, mean start/end error is 36.13/35.93 ms and p95 is 78.04/95.13 ms. This measures alignment given correct text, not recognition accuracy.
The same 64 original recordings then went through the unchanged hybrid8 local upload → HTTP API → queue → Rust inference → persistence → GET flow and fresh AssemblyAI requests, with identical explicit-Spanish options and serial ABBA provider order. Neither recognizer received the reference text. All 128 requests completed. The 48 validation clips give:
| Validation metric | Sandchest hybrid8 | AssemblyAI |
|---|---|---|
| Lexical errors / 494 reference words | 14 (2.83%) | 9 (1.82%) |
| Mean start/end error, same 481 matched words | 48.43 / 47.55 ms | 56.84 / 57.03 ms |
| p95 start/end error, same matched words | 150.63 / 147.01 ms | 128.30 / 134.97 ms |
| Median full request time | 640.58 ms | 2,381.10 ms |
| p95 full request time | 721.36 ms | 4,251.91 ms |
This is local-client/local-API versus remote-provider request latency, not a same-hardware or deployed speed comparison. Across all 64 clips, lexical WER is 20/669 (2.99%) versus 14/669 (2.09%). A descriptive paired-speaker bootstrap does not establish quality parity; the validation difference is +1.01 percentage points, with a 95% percentile interval of 0.00 to +2.25 points.
MMS improves actual-hypothesis timing on this Spanish set. Another 128 native alignment calls use the exact word hypotheses returned by the two APIs. All candidates align, with no word changes or guessed numeric pronunciations. On Sandchest's same 481 matched validation words, MMS reduces mean start/end error to 36.09/36.00 ms and p95 to 77.94/95.53 ms. Both boundaries are within 80 ms for 422/481 words, versus 318/481 with current serving timestamps. This stage is experimental and is not included in the API latency numbers above. It does not reduce recognition WER or establish long-audio/multilingual safety.
The 64 local jobs each ran once and persisted exactly one usage event with the
native-measured duration. The first harness finalization failed on two 1 ms
half-up versus ties-to-even rounding expectations; an offline audit reconciled
all API, telemetry, persistence and usage durations without repeating requests.
The original failure and source are retained. AssemblyAI's integer-second
audio_duration also fails the generic 1 ms duration check on 63 responses;
all 64 provider word bounds/text pass, so those are not word-timing failures.
Ephemeral local API keys were revoked and both owned processes exited cleanly.
Immutable additive audits bind every result, including rejected responses and
all derived record files, and retain denominators and failed harness attempts.
Private evidence is owner-readable only under quality1, cap-expanded2,
manual-dimex1, manual-dimex-mms1, manual-dimex-api1 and
manual-dimex-hyp-mms1, with a durable private checkpoint under the session's
Git metadata. No production or general 20% Cap beta readiness is claimed.
A further bounded experiment retained the pinned 8-bit Qwen model and the existing strict UTF-8/EOS decoding guards. It added explicit language requests, lossless canonical WAV input validation and an integrated native CPU speech detector. Python only prepared/scored the experiment; recognition and the detector ran in one native Rust process with a narrow C++ detector bridge.
The model was tested on the same 64 Spanish reference clips and every eligible nonempty complete recording within 30 seconds from the frozen 59-recording Cap inventory: 16 recordings, with all 43 exclusions retained. Cached language hints were supplied explicitly. This does not test automatic language identification, long recordings, new Cap speakers, or the full language list.
| Spanish speaker-disjoint validation | Qwen | Current Sandchest | AssemblyAI |
|---|---|---|---|
| Reference words / clips | 494 / 48 | 494 / 48 | 494 / 48 |
| Word errors | 15 | 14 | 9 |
| WER | 3.04% | 2.83% | 1.82% |
All 64 Spanish clips yielded 23/669 Qwen word errors, versus 20/669 for the current native route and 14/669 for AssemblyAI. The paired speaker bootstrap is descriptive and too imprecise to establish parity. Qwen is not an accuracy improvement for this Spanish set and is not being promoted as the default multilingual model.
On the 16 short Cap recordings, disagreement with cached AssemblyAI text was 60/421 tokens (14.25%) for Qwen and 136/421 (32.30%) for the current native result. This is provider disagreement, not WER. Language strata matter: the one Arabic recording was worse with Qwen (8 disagreements versus 2, out of 10 reference tokens); the aggregate must not hide that regression.
The raw model also emitted a lexical token for each of three synthetic nonspeech controls. The final experimental gate now:
- rejects malformed/nonfinite WAV samples and Qwen FFT-overflow inputs without partial text, preserving valid PCM instead of silently substituting zero;
- accepts normal float PCM with headroom up to absolute amplitude 8.0, including the existing 2.0 case, and rejects larger values before the detector without clipping; the Qwen FFT bound alone does not protect the detector's float math;
- performs two one-frame detector computations at startup, including a NaN sentinel, so loading weights alone cannot advertise readiness;
- resets detector state per request, uses the existing 0.5 threshold and bounded 32-frame batches, and skips ASR only when the whole request has no detected speech; it never trims or rewrites speech audio.
The corrected native gate passed 110 calls: 93 unchanged speech results including recoveries, 14 expected errors without partial text, and three empty nonspeech results with ASR bypassed. Eleven native entry tests pass. A separate bridge fault test covers six unusable-graph outcomes, including an unwritten probability and failed computation, and verifies context cleanup. The earlier raw and first-gate results remain preserved, including their defects; they are not relabeled as safe.
A further 12 calls compared raw/gated Qwen on all four eligible short Cap recordings from the seven cached empty/error results, with speech warmup/recovery for each variant. Three longer recordings were explicitly excluded. The gate bypassed ASR on the three cases where current Sandchest was empty and cached AssemblyAI had an error; raw Qwen produced text on all three. The remaining case contained two AssemblyAI words and one Qwen word, and the gate retained speech. Cached empty/error status is not a human annotation of silence, and this test does not establish that remaining transcript's accuracy.
On the original 64 Spanish clips the final gate's recognition-only median was 157.95 ms and p95 was 218.61 ms; the detector itself had median 1.76 ms and p95 2.09 ms. These exclude API upload/queue/persistence, word alignment and cold start. They must not be compared directly with the earlier full API latencies.
The independent audits re-decoded 84 original Spanish/Cap media files to verify
exact PCM identity, checked all 94 standalone detector/input pairs, recomputed
224 lexical scores from their actual source responses, verified speaker splits
and every selected/excluded Cap record, and checked sample counts across replay
outputs. The corrected binary is
431957081064e88d0eb0ad47da42dcadf6b6d115ba5806d4fc816df8333493a1;
source, build commands, failures, results and audits are preserved by the private
qwen-multilingual-checkpoint.json receipt in the existing session directory.
This remains an experimental candidate, not the serving route. Its staging dylib
paths are not a deployable package; its guards are at the experimental entry point,
not automatically applied to every Qwen library caller. The current serving binary
remains 19495a30d707d969b5697618d9675bd11e51ead017fe8c7ae1926a919830982d.
No API, deployment, shared database, port 3000, or Cap source/data was changed by
this experiment. Full API/language behavior, safe routing/fallback, multilingual
quality and longer timestamp behavior remain open. A 20% beta is not approved.
AssemblyAI is the source of truth for language scope. Its current prerecorded
model combination, universal-3-5-pro plus universal-2, covers 99 languages.
Match the published language codes, regional aliases, explicit selection and
automatic detection. Do not add requirements merely because a language appears
in Cap, or stop at Cap's smaller selector. Track the actual code list rather than
relying only on a headline count. AssemblyAI supported languages.
The current published table contains 99 base-language codes plus en_au, en_uk,
en_us and de_ch. Preserve these aliases in the compatibility mapping; an alias
mapping alone is not proof of dialect accuracy. Recheck the list before release.
This scope concerns the prerecorded /v2/transcript API used by Cap, including
Cap's fragment submissions; it does not add a new WebSocket streaming product.
The default Parakeet route still supports only 25 European languages. The opt-in Whisper route maps all 99 base codes and four aliases; it remains under real speech and API verification. Neither this mapping nor the legacy route has passed the complete deployed replacement gate.
Accepting a code is not evidence of transcription quality. No language has yet passed the complete deployed replacement gate. The default Parakeet route rejects broader automatic-language hints; selecting the opt-in multilingual backend is required for those requests.
Whisper large-v3-turbo is now implemented as an opt-in native backend behind the Rust service. Compare its complete language inventory with AssemblyAI's list, including any API aliases; matching Cap's smaller language selector is insufficient. Whisper language inventory. The native whisper.cpp engine has Metal and CUDA implementations and a C API; its word timestamps are explicitly experimental and must pass the caption/editor checks before adoption. Native engine documentation.
Keep Parakeet for routes where it wins the measured accuracy/performance tradeoff. Do not assume a single candidate will win every language. Qwen3-ASR and Cohere Transcribe do not independently cover AssemblyAI's full list, so neither closes this coverage gap alone. Qwen3-ASR, Cohere Transcribe.
- Handle every language on AssemblyAI's published prerecorded list and compatible automatic-language requests, including accents, regional codes, silence, short speech, and switching languages within a recording. Unknown or uncertain language must not be forced into an English transcript or silently translated.
- Preserve upload, asynchronous creation and polling, idempotency, genuine punctuation and fillers, language metadata, actual model identity, and word timestamps in milliseconds. Validate grouping for scripts without spaces; never manufacture timestamps by splitting caption segments evenly.
- Verify complete files, fragmented live MP4, live word offsets, captions and editable transcripts through the actual Cap workflows. Pure formatter tests are useful but do not prove those workflows run successfully.
- An unverified language, request mode, timeout, or unavailable native model must retain an explicitly configured AssemblyAI route during rollout. Use separate credentials, bounded retries, stable job ownership, and exactly-once local usage recording; do not run both paid providers without a declared shadow test. This fallback is service continuity, not evidence of independent native parity.
Freeze model revisions, decoding settings, input bytes, scoring rules and routing before the final comparison. Keep development and held-out evaluation separate, including original recordings, speakers/sessions where available, and duplicate audio. Existing Cap captions and provider agreement are not human ground truth.
Use a complete language-code mapping test and real speech across the supported scripts to catch unsupported or broken paths. Expand independently referenced examples where needed, including mixed-language, noisy and long Cap recordings. Do not require 100 clips per language or an accuracy win in every language merely to establish language support. Public FLEURS can supply multilingual development examples, but cannot establish performance on Cap screen recordings alone. FLEURS dataset.
Use character error rate for Chinese/Japanese and word error rate with frozen language-appropriate normalization for other languages; also report character error rate, names/numbers, omissions, hallucinations, fillers, language detection, and timestamp coverage/error. Do not apply English-only number normalization to other languages. Bootstrap paired results by original recording or speaker/session where available. Report each language and domain separately, as well as a macro average; a large English sample must not conceal a regression elsewhere.
The accuracy objective remains an independently measured overall improvement over AssemblyAI. Report per-language regressions honestly and fix broken recognition, but do not confuse language coverage with a claim of equal or better WER in every language. The earlier proposed per-language 0.5-point non-inferiority gate is not a requirement for language parity under the user's clarification. Preserve old plans and results; do not relabel an unpassed accuracy test as a pass.
Measure identical complete API workloads against deployed Sandchest and AssemblyAI from the same client location. Include upload, queue delay, inference, database and storage roundtrips, result persistence, and polling. Report cold start/model loading separately, and keep warm-path timing honest. Local Metal versus remote AssemblyAI does not establish deployed parity.
For this goal, "comparable" means the upper paired 95% confidence bound for the Sandchest/AssemblyAI p50 and p95 completion-time ratios is at most 1.10 on the representative matched workload. Break results down by language and duration; this performance goal is separate from matching the supported language list. Target ratios below 1.0. Also report p99, effective throughput, failure/timeout rates, and cost per audio hour at equal offered load. A faster median does not compensate for a worse tail or dropped work. Do not infer cloud GPU performance from Apple Silicon.
Profile before optimizing: queue wakeups, HTTP polling, storage/database calls, audio decoding, encoder/decoder work, repeated model passes and device transfers. Evaluate warmed models, buffer reuse, safe batching and mixed precision only with accuracy, memory, cancellation and tail-latency checks. Batching must not add an unbounded wait to short recordings; quantization must pass each language's gate.
Run sustained concurrency and burst tests, plus worker death/restart, overload, expired leases, cancellation, invalid media, silence and long recordings. Verify bounded memory, recovery, no stranded jobs, tenant isolation, and one persisted usage event per completed job. Real billing and payment settlement remain a separate integration gate.
Deploy web, durable queue and native inference together; a reachable web API alone does not prove queued jobs are processed. Run the actual Cap workflows against the deployment, then shadow a bounded authorized sample. Enable a small canary only for routes that pass their gates, with an immediate rollback and an observed AssemblyAI fallback. Expand by verified language and request mode. A full switch requires every Cap route to pass; deployment alone is not approval to replace it.
Current evidence and private benchmark manifests must retain their sample counts, hardware, revisions, settings, failures and limitations. Never publish customer recordings, transcript text, identifiers, credentials or signed media URLs.
A host restart removed the prior RAM volumes. Completed private archives survived and were read back before recovery. The guarded pinned Whisper SDK, a baseline binary from the unchanged serving sources, an F16 large-v3 comparison binary, and the public evaluation inputs were rebuilt/restored on persistent disk. All 64 Spanish WAV/annotation pairs match the original hashes and selection; all ten Hindi WAV/reference pairs and three nonspeech controls also match. This restoration does not create a new holdout.
Full F16 large-v3 is not an improvement. A fixed 64-call trial used the actual Rust media decode, guarded recognition and native word conversion, with the correct large-v3 DTW heads and verified model artifact. Each model received the same 16 Spanish development clips, ten Hindi development clips, three nonspeech controls, an invalid-language request, and warmup/recovery requests. English routing and CTC alignment were disabled for this non-English model comparison.
| Development metric | Current Turbo | Full large-v3 F16 | Cached AssemblyAI |
|---|---|---|---|
| Spanish word errors / 175 words | 6 | 6 | 5 |
| Hindi word errors / 282 words, failures included | 107 | 106 | 62 |
| Hindi completed clips / 10 | 9 | 9 | 10 |
| Spanish native pipeline median | 370.5 ms | 734 ms | Not measured |
| Hindi native pipeline median, completions | 687 ms | 2,381 ms | Not measured |
| Spanish mean start/end error, same 169 gold-matched words | 47.43 / 48.85 ms | 116.18 / 115.55 ms | 54.43 / 51.86 ms |
| Spanish p95 start/end error, same matched words | 117.39 / 165.74 ms | 327.64 / 294.20 ms | 122.10 / 132.80 ms |
These are reused public development recordings, not new Cap recordings. Provider results are cached and identity/settings matched; no fresh provider latency was measured. Pipeline timing excludes API queue, persistence, network and cold model startup. The larger model is slower and has worse Spanish timestamps, so the serving model was not changed.
Bounded context recovery is promising but remains experimental. Both models failed the same dense Hindi recording. Earlier decoder diagnostics implicated the ordinary 220-token limit. A private SDK/worker candidate now tries at most one additional 441-token attempt only after all six ordinary temperature attempts fail and at least one decoder was explicitly failed by the SDK's token-limit/repetition branch. Prompted/history-conditioned requests are ineligible. A completed EOT candidate later rejected for invalid output does not independently qualify as a capacity failure.
The candidate preserves the existing UTF-8/output guards and no-speech policy, checks prompt/KV and separate DTW context bounds, and retains cancellation checks. It never resamples a successful ordinary result. Source review found that the first prototype also shortened some ordinary prompted budgets; V2 corrects that and preserves the exact original ordinary budget regardless of prompt length. Both attempts and their evidence remain retained.
V2's 32-call replay completes all 26 public speech clips. Hindi error count falls from 107 to 82/282 (37.94% to 29.08%), including failures in both denominators, versus AssemblyAI's cached 62/282 (21.99%). The recovered clip itself has 19/44 word errors and takes 3.94 seconds in the native pipeline. Recovery is not proof of a correct transcript or provider parity. All 30 previously successful responses, including controls and warmup/recovery, preserve their complete semantic output; elapsed telemetry is excluded from equality. Nonspeech remains empty, the invalid language remains an error, and subsequent requests recover without a process restart.
Two paired 16-call trials then cover six previously tested Cap recordings (8.06 minutes total), including the difficult long Hindi and recovered Spanish cases, plus warmup/recovery. Both versions preserve all five successful Cap transcripts and word timestamps exactly. The long Hindi case still fails. Explicit cached-provider language hints were supplied equally to both native variants; this does not test automatic language identification or compare matched provider request settings. No new Cap recording, provider request or API/job/billing flow is represented by these trials.
Each of the four binary builds passed 124 Rust tests. Context helper tests also passed AddressSanitizer/UndefinedBehaviorSanitizer checks. Those helper tests and source review do not replace actual decoder fault, prompted-call, boundary and cancellation tests; those remain necessary before serving integration. Across the five trials, 160 native calls were made on repeated inputs, not 160 unique recordings. An offline audit rereads responses, recomputes word distance, checks input/runtime fingerprints and verifies that owned model processes are absent.
The app regression checkpoint remains 142 passing tests, seven conditional
real-model skips and 1,171 assertions, with live service credentials disabled
and isolated PGlite/storage; typecheck and lint also pass. No serving source,
Cap data, shared database,
deployment, port 3000, index or HEAD was changed. Neither the F16 model nor
context recovery is promoted, and the general 20% beta remains unapproved.
Persistent receipts are under the private .sandchest/experiments/whisper-large-20260827/durable1
directory; a verified archive is stored under this session's Git metadata.
The Rust implementation is improved, but the general 20% Cap beta is still not approved. Licensing is not part of the active acceptance work. This checkpoint supersedes the earlier experimental-only status of bounded context recovery; the larger F16 model remains unpromoted.
The canonical pinned Whisper SDK patch now includes the bounded context retry tested above, a guard against oversized DTW input, and allocation cleanup fixes. Two injected allocation failures reproduced SIGSEGV/SIGBUS in disposable legacy processes. The fixed canonical code passed 31 actual decoder/bridge tests, including seven simulated allocation exceptions, cancellation before/during retry and alignment, exact token-budget boundaries, partial-result removal, and recovery without restarting the model. No actual memory exhaustion was induced. Context/output helper tests passed ASan/UBSan; the full Metal runtime was not sanitizer tested. The bridge requires private extension version 3, so an older SDK cannot silently bypass the required behavior.
Startup now owns partial state and backend handles until successful transfer, initializes batch pointers, and frees KV buffers idempotently. A failed KV reallocation no longer frees caller-owned state. The bridge always discards a failed state; arbitrary reuse of a failed public whisper.cpp state is outside this verified path and remains a caveat.
A full HTTP pilot exercised the unchanged application handler and durable queue with isolated PGlite/storage, actual Rust inference, persisted output, and usage events. The recovered Hindi clip completes through this path. Controls remain empty, the known long Hindi failure remains explicit, and a following request succeeds. Seven completions generated seven unique measured usage events; the failed job generated none. Transcript exports, sentences, paragraphs, search, list, invalid authentication, and idempotent replay were also checked. This is local application verification, not a shared Postgres, payment, deployment, or cloud-GPU test.
The pilot found invalid explicit/hinted/fallback language identifiers were
being queued. Three fresh public-audio AssemblyAI requests confirmed immediate
HTTP 400 responses. The application now validates these identifiers before
job creation. Source review caught and fixed an ALL-sentinel regression.
The exact expected_languages: ["all"] form and fallback_language: "auto"
remain accepted; mixed ["all", "en"], unsupported identifiers, and "all"
as an explicit/fallback code are rejected. A subsequent real HTTP test produced
identical words/text for explicit English and automatic ALL detection, rejected
seven malformed requests, and left exactly two completed jobs and two usage
events. The tests explicitly check that invalid inputs create no queue rows.
This does not establish all other request-option or error-message parity.
The fixed selection contained 60 recordings from 56 owners, disjoint from previously selected/evaluated video and owner hashes. One holdout source had truncated decoded audio and was excluded before inference. The evaluated set contains 59 recordings from 55 owners, 2.192 hours: 30 development and 29 holdout recordings. Exact input bytes, decoded PCM, duration, API settings, native models, libraries, and application source were frozen before inference. Neither production media nor Cap source/database rows were changed.
Both providers received identical Cap request settings and audio bytes through fresh upload/submit/poll requests, alternating serial provider order. Sandchest used the actual local TCP API, durable queue, native worker, persistence and result polling. Model startup is outside the timed boundary; ordinary first requests are included. This compares a local M4 Max to the remote AssemblyAI service, and does not establish deployed latency or same-hardware throughput.
| Full API result | Sandchest | AssemblyAI |
|---|---|---|
| Completed / 59 | 55 | 57 |
| Terminal errors | 4 | 2 |
| Median, same 53 completed pairs | 1.551 s | 7.707 s |
| p95, same completed pairs | 7.126 s | 17.800 s |
| p99, same completed pairs | 11.199 s | 43.279 s |
| Holdout completed / 29 | 26 | 28 |
| Holdout median, same 25 completed pairs | 1.816 s | 9.314 s |
| Holdout p95, same completed pairs | 7.698 s | 25.813 s |
Sandchest was faster on all 53 matched completions. Paired owner-bootstrap 95% intervals for the p50 and p95 latency ratios were respectively 0.152–0.253 and 0.123–0.678. These small-sample intervals apply only to this local-versus-remote run; failures are reported separately, not removed from the completion denominator.
All 55 native completions passed strict checks for text/word consistency, finite confidence, integer positive word intervals, monotonic bounds and measured duration. Fourteen AssemblyAI completions failed this same strict word-bound check. That is not a finding of inferior acoustic timing: these checks are stricter than merely returning integer timestamps. Neither structural validity nor agreement with provider timestamps measures absolute word-timestamp accuracy.
All four native errors were on recordings AssemblyAI labeled English. Three AssemblyAI results were empty; the fourth contained 13 normalized words. AssemblyAI's own two errors reported no speech. Preserve all these cases: an empty provider result is not independent confirmation of silence.
Normalized token disagreement against the provider was 3,263/17,643 (18.49%), including native failures where a provider transcript exists; the holdout rate was 21.59%. These are provider-disagreement rates, not WER. There are four language-code disagreements where both transcripts are nonempty: Indonesian/English, Finnish/English, Hindi/English and Turkish/Slovak. The first two provider languages were outside the supplied Cap hint list. Other language-code differences include empty transcripts and must not be counted as proven spoken-language errors. Independent references are still needed to determine which transcript/language decisions are correct.
The audit verified all 118 request identities, exact provider inputs/settings, and independently recomputed token edit distances. All 59 local jobs reached a terminal state in one attempt; 55 unique usage events match measured audio milliseconds exactly, and no failed job has completed content or usage. Ephemeral keys were revoked and both owned processes were independently absent.
The one later application change restores the ALL hint sentinel. It leaves the 25-code Cap request branch and all native model/decoding bytes unchanged. The exact application sources used for the 59-recording comparison are retained privately; the final sentinel behavior was separately verified over real HTTP. No quality or latency number was silently relabeled as a run of modified code.
A 14-call development-only diagnosis reproduced one native error with both automatic and explicit English. Whisper-only recognition completed that same clip with two words, while the cached provider result was empty. That result does not justify promoting a fallback or labeling the added words correct. Silence, tone, noise, and subsequent speech controls still pass. An initial diagnostic harness compared memory/time telemetry as if it were transcript content; the retained corrected run excludes only that telemetry and still requires exact text, word, confidence and alignment-result equality.
Current checks: 124 Rust tests; 143 app tests, seven conditional model skips, 1,406 assertions; whole-app typecheck and lint pass. The broad API is still not 1:1: punctuation/disfluency options, cropping, custom spelling, diarization, multichannel and other documented features remain gaps. Recognition continuity, automatic language selection, and independently referenced multilingual and timestamp quality remain acceptance work. No deployment or production checks were run, as requested.
Private durable receipts: integration1, integration-faults1,
api-comparison2, api-analysis2, language-all-fix1, and
fresh-dev-diagnosis2 under the persistent experiment directory.
Still no 20% beta sign-off. This follow-up fixes request semantics and preserves the existing recognition and alignment implementation. It does not establish multilingual quality parity, eliminate the earlier speech failures, or complete the remaining STT feature set.
Fifty-two bounded submissions used public English/Spanish audio only: 39 returned HTTP 200, ten returned 400, and three returned 500. HTTP 200 is submission acceptance, not necessarily successful transcription. Every intent, response, terminal result and failed attempt is retained privately; ambiguous POSTs were not retried.
The provider constrains automatic selection to the requested candidate list in these public probes. That agrees with the official language-detection documentation. We did not convert a restrictive request into a soft hint to hide the earlier out-of-list provider results on Cap audio.
Canonical application and Rust validation now agree on these observed cases:
- An empty candidate list with an explicit
"auto"fallback is accepted as unrestricted detection. This worked for both public English and Spanish. - An empty list with a fixed or omitted fallback is rejected before queueing.
- Within a supplied options object, omitted fallback behaves as English in the
live provider's validation. Thus
["es"]without a fallback is rejected, whereas explicit"auto"or"es"is accepted. This contradicts the current documentation's stated automatic default and is recorded as such. - English locale identifiers are folded to
enfor detection candidates and fallbacks; explicitlanguage_codeand stored request fingerprints are not rewritten. German and Swiss German remain distinct candidate identifiers. ["all"]accepts supported regional fallbacks, including Swiss German. A fixed fallback outside a restrictive candidate list is rejected.
Three nested-null probes returned provider internal-server errors. Sandchest preserves its existing null-as-omitted normalization rather than reproducing those server failures. This remains a documented behavior difference.
A frozen, development-only diagnostic made 22 native calls: seven reused Cap recordings with original versus unrestricted language candidates, public warmup/recovery, and silence/tone/noise under both settings.
For the two recordings whose provider language was outside Cap's 25-code list, unrestricted native detection selected Indonesian (confidence 0.99679) and Finnish (0.99823), matching the provider labels. The five other development controls retained exactly the same text, words, language, model and duration. All 22 calls completed and passed structural checks; no model or routing policy was promoted from this diagnostic.
Provider disagreement fell from 126/177 to 102/177 tokens on the Indonesian recording and from 395/406 to 88/406 on the Finnish recording when their languages were allowed. These are provider disagreement, not WER. The remaining disagreement, especially on Indonesian, still requires independently referenced accuracy work. These altered requests are not a like-for-like benchmark against the provider's original restrictive requests. No heldout recording was used for tuning or rerun in this follow-up.
The new canonical binary SHA-256 is
ed966f09b5a6606d28bae8869005979f840e09f2b3891cceaf77e5d34b93c19e.
Relative to the preceding canonical build, the only changed native source is
src/whisper/language.rs. SDK libraries, decoder recovery, model weights and
word alignment are unchanged.
Real local TCP upload → API → isolated durable queue → native inference → persisted results completed 20/20 requests. This covered English/Spanish empty-list and locale cases, seven previously used Cap development recordings, two additional unrestricted requests, and speech recovery after invalid input.
All 20 completions passed word/text/confidence/duration checks; 20 unique usage events matched measured milliseconds, with one job attempt each. Ten invalid requests created no extra jobs. Idempotent replay, authentication rejection, list, subtitles, sentences, paragraphs and word search also passed. English variants/recovery and Spanish variants returned identical content and timing. All seven original Cap development requests preserved exact text, words, language, model and duration from the earlier frozen full-API run.
125 Rust tests; 144 app tests, seven conditional model skips, 1,436 assertions; whole-app typecheck and lint pass. The initial validation attempt was too strict for empty-auto and English locales; public probes caught that before the final API run. Intermediate type/import-check failures are also retained. Temporary API keys were revoked and owned native/API processes stopped.
The earlier 59-recording performance numbers remain attached to their original binary and app snapshot; they are not relabeled as new measurements of this build. No deployment, shared database, port 3000 or Cap source/storage write was performed.
Public cropping probes established that word timestamps use original-file
positions while audio_duration reports the cropped interval. End positions
past EOF clamp, and starts past EOF or empty/inverted intervals become terminal
errors. Fractional positions are rejected at submission. Negative start input
was accepted by the provider, so a future implementation must explicitly handle
or document that behavior rather than assume all invalid-looking ranges get
HTTP 400. The cropping documentation
defines the millisecond input positions but does not resolve all these edges.
Cropping remains unsupported in Sandchest. It requires the same sample slice for recognition and floating-point alignment, correct original-file offsets, and integrity checks that distinguish selected duration from timestamp origin. No crop-related serving change was made in this checkpoint.
A read-only review identified an API/native mismatch for lists over 103 entries. The API now rejects those before persistence, matching the existing native resource bound; this is not claimed as an AssemblyAI limit. A separate full HTTP check accepted all 103 declared identifiers, rejected 104 entries without another job, completed one real transcription and persisted exactly one measured usage event. The native binary was unchanged.
The final app checkpoint remains 144 passing tests and seven conditional skips, now 1,440 assertions, with passing typecheck, lint and scoped checks. The 20-request receipt above retains the immediately preceding application snapshot; the capacity receipt freezes the final application source.
Six further public requests examined omitted, fixed-English and automatic fallback on Spanish and Hindi. Spanish selected English in all three cases; Hindi selected Urdu in all three. These do not distinguish the provider's underlying fallback-selection algorithm. The default-English validation is directly observed, while applying that default during native selection remains an inference and a parity risk to investigate. The original and unrestricted language policies still need broader independent accuracy validation.
The review also noted the provider's zero-probability Swiss-German error. We have not generalized that one observation into a new rejection rule across all native language probabilities; it remains a specific unverified boundary. No additional serving policy was changed to manufacture provider agreement.
Private receipts: language-hints-provider1/2, crop-provider1,
language-boundaries-provider3/4, language-dev-diagnosis1,
language-options-fix1, language-integration1, language-api1,
language-default-selection1/2, language-capacity-fix1,
language-capacity-http1, and app-checks8 in the persistent experiment
directory.
The serving model remains Turbo. A larger model is not automatically a better replacement, and these results do not establish full-language parity or beta readiness.
A new frozen public development set contains 40 FLEURS validation recordings, 20 Indonesian and 20 Finnish, totaling 8.311 minutes. Selection was deterministic, with unique source IDs within each language and unique decoded PCM hashes, before any model output was inspected. Published transcripts are the references. These are read-speech development examples, not Cap recordings or a heldout test; speaker independence is not established.
The same canonical audio bytes were submitted with each language specified and with unrestricted automatic detection to current native Turbo, full-size Whisper large-v3 F16, and AssemblyAI. All 240 development responses completed. Automatic detection selected the expected language on all 40 clips per backend. The explicit and automatic runs reuse the same 40 recordings, so they do not double the independent sample count. Native controls and recovery also passed.
| Language | Turbo lexical WER | Full-size F16 WER | AssemblyAI WER |
|---|---|---|---|
| Indonesian, 20 clips | 21/388 = 5.41% | 19/388 = 4.90% | 21/388 = 5.41% |
| Finnish, 20 clips | 29/323 = 8.98% | 29/323 = 8.98% | 39/323 = 12.07% |
Both request modes have the same aggregate WER shown above. Scores preserve Unicode letters, marks and numbers; punctuation is excluded. Scoring against the second published raw-transcript field gives the same aggregate rates. An independent dynamic-programming implementation recomputed all edit counts, and the audit verified every saved request's audio hash and options.
These are promising point estimates, not a demonstrated general superiority: the paired source-ID bootstrap 95 intervals for Turbo minus AssemblyAI WER cross zero (explicit Indonesian -1.57 to +1.98 percentage points; Finnish -7.61 to +0.96 points). There are no independent word-boundary annotations for this set, so valid timestamp structure is not a timestamp-accuracy result.
Native-only median recognition/alignment time for explicit requests was 421 ms versus 1,143 ms on Indonesian and 481 ms versus 2,322 ms on Finnish (Turbo versus full-size F16). These exclude upload, queue, persistence and startup; they are not hosted API latency comparisons. The four-language development experiment also retained Spanish WER at 6/175 but worsened its gold-matched boundary error with F16: start/end MAE 116/116 ms versus 47/49 ms. Hindi improved only from 82/282 to 78/282 errors, while native median latency rose from 707 ms to 2,527 ms. Therefore F16 was not promoted.
All models used the same current native guards and language parser. The F16
candidate changed only the pinned artifact, identity, correct 32-layer DTW
preset/guard and matching test identities in an isolated source copy.
Canonical serving weights, recognition policy and alignment remain unchanged.
Private receipts: large-current1, current-public-pilot1,
fleurs-id-fi-dev1, and id-fi-comparison1 under the durable experiment
directory. The six bounded crop edge probes are a separate API-semantics
experiment and were not included in these accuracy denominators.
Nonnegative integer audio_start_from / audio_end_at requests now work through
the native worker and API. Null means omission; end 0 follows the observed
provider behavior of using the rest of the file. End values beyond EOF clamp
to the measured audio. Empty/inverted ranges and starts beyond EOF become
terminal errors, with no transcript content or billable usage.
The request fields and units are documented in
AssemblyAI's range API.
Selection happens on decoded 16 kHz samples before language detection and recognition. English CTC alignment applies the identical sample range to its separately decoded float waveform and rejects a source-length mismatch. Internal words stay relative to the selection. The worker returns sample-count evidence; the queue validates that evidence against the requested range and measured duration before adding the original recording offset exactly once. Persisted words, sentences, paragraphs, word search, SRT and VTT therefore retain that absolute timeline. Usage remains the selected measured duration, not the last absolute word endpoint. Completion and unique usage still commit atomically.
An older worker that ignores range options cannot silently return a successful full-file transcript: missing or inconsistent selection evidence is rejected without content, billing or repeated inference. Decoder bounds, cancellation, request fingerprints, tenant isolation and stale/deleted-job fencing remain in place. The crop does not enable faster container seeking; media decode retains its existing bounded full-source validation, while recognition/alignment operate only on the selected samples.
The isolated real TCP verification passed 25 requests: 22 expected completions and three intentional range errors. Four English/Spanish range results match physically identical manual sample cuts exactly for text, words after removing the offset, language, model, confidence and selected duration. Five reused Cap development recordings retain the preceding text, word times and metadata. Silence/tone/noise crops remain empty; subsequent speech recovers unchanged. All 22 completions have one exact measured usage event, and the errors have none. Replay returns the same transcript; a changed range under the same key conflicts. Five malformed inputs create no extra jobs. Derived resources preserve offsets.
Current checks: 148 app tests, seven conditional model skips, 1,568 assertions; 128 Rust tests; whole-app typecheck, lint and scoped checks pass. A separate read-only code review found no blocking defect. The first app test run contained an incorrect empty-global-queue assumption, and the first HTTP harness expected an hours field in short VTT timestamps. Both failed receipts are retained; corrected tests ran against the implementation without changing those behaviors.
Known differences remain explicit: negative ranges return 400 rather than the provider's observed negative word timestamps; integers beyond JavaScript's safe range return 400 rather than its observed 500. A 1 ms provider selection returns a too-short error, but the exact minimum duration and equivalent native small-window behavior are not established. This is not a claim of every range edge case or complete STT API parity.
Private receipts are crop-fix1, crop-integration1, app-checks-crop1/2,
crop-api1/2, and crop-provider-edges1. The current native binary is
84c774230b71ba7535292e289b71bfdb92c21b3772a1f171a515aac1526b8fe5.
Owned API/native processes exited, ephemeral keys were revoked, and no shared
database, migration, deployment, port 3000 or Cap production data was changed.
This verification is a compatibility/regression check, not a new accuracy
holdout or a deployed performance benchmark. The general 20% beta is still not
approved; multilingual/Cap quality, timestamp evidence and remaining STT
features stay part of the active goal.
The native Qwen prototype was restored from its final guarded source checkpoint,
not the older unguarded source copy. The exact eight-bit model
(bf304b009cc7eca79283056f787b44c952d24ac22cec787b39732bba3c23c13c)
and reconstructed source were verified. The new binary
(40e3de4b8d7a5a375a861ee98e091f1bcb72003bc0359a6fe896cafa2d076f03)
adds Indonesian/Finnish request names only. Strict UTF-8, completion, finite-sample,
bounded-input and VAD safeguards remain in place. There is no serving promotion.
The fixed comparison made 144 native calls:72 per model, with40 public Indonesian/Finnish clips,10 previously used Hindi regressions,14 chunks covering two previously used Cap development recordings, six nonspeech controls, warmup and recovery. All calls completed; nonspeech returned empty text, recovery was stable, and all10 Hindi Qwen outputs reproduced the previous guarded prototype. The public ID/FI clips are from FLEURS's published validation partition, used as our internal development set. This is not a new Cap holdout.
| Independent reference set | Qwen word errors | Current Turbo word errors | Cached AssemblyAI word errors |
|---|---|---|---|
| Indonesian20 clips | 11/388 (2.84%) | 21/388 (5.41%) | 21/388 (5.41%) |
| Finnish20 clips | 93/323 (28.79%) | 29/323 (8.98%) | 39/323 (12.07%) |
| Hindi10 reused clips | 41/282 (14.54%) | 82/282 (29.08%) | Not rescored in this run |
The Indonesian result merits further evaluation. A paired source-clip bootstrap gives a Qwen-minus-AssemblyAI difference of-2.58 percentage points (95% interval-5.09 to-0.24); against Turbo the interval includes zero (-5.90 to+0.22). These intervals do not establish speaker-independent or Cap-wide performance. Finnish is clearly worse and rules out a global Qwen substitution.
Native-only median times were321/465ms for Indonesian,423/562ms for Finnish and 709/893ms for Hindi (Qwen/Turbo). Qwen here supplies no word alignment and has no public-API integration, whereas Turbo includes DTW. These are not comparable completed-pipeline latency measurements.
On the reused Indonesian Cap recording, Qwen had99 word edits against the cached AssemblyAI transcript, versus105 for identically chunked Turbo and102 for cached full-audio Turbo. On Finnish the counts were223,94 and88 respectively. These are provider disagreement, not WER. Neither the Indonesian disagreement nor the quality of its word times is resolved by this experiment.
An additive audit checked all144 raw returned languages and44 cached request
identities. The original driver omitted a returned-language assertion and did not
freeze the old ID/FI mapping plan; the audit verified every response and matched
that plan to the immutable pre-evaluation archive. The build plan also omitted
the VAD hash: the evaluation did freeze it, and a new explicitly frozen replay
passed12 entry tests, two strict-decoding tests and the bridge startup/fault tests.
The original provenance omissions and failed audit attempt are retained, not
rewritten. Private evidence: qwen-idfi-eval1, qwen-idfi-audit1/2,
qwen-recovery1/build1/2. The current serving binary is unchanged and general
beta approval remains open.
Spanish acoustic word timing is now an opt-in native Rust implementation in
the worker and normal startup script. Defaults are unchanged. The tested binary is
d33ac305fe15a1791970a1bc4244e12ee665f124ec74ff9aa77266827fc3f836;
the main worker source matches its verified source snapshot byte-for-byte.
The mode combines the existing English CTC route with MMS timing for Spanish. It preserves recognition text, word identities, confidence, language decisions, crop offsets and billed audio duration. Unsupported normalization, uncertain boundaries and conflicting window candidates retain the original native word times. Numbers are not guessed or silently dropped. Runtime/model failures and cancellation remain failures rather than successful fallback. Conditional timing probability is not treated as recognition correctness.
The pinned MMS artifact is
8d3966ea92a61832fe509eb0a0c3c783d070b71c7f0f25887bf37df9214b1c20
(1,261,919,144 bytes). Its original official source bytes and all 423 F32 tensors
were checked during recovery. The original converted artifact was recovered
exactly by restoring metadata ordering; tensor bytes were unchanged. The Rust
worker does not call Python. This mode requires an explicitly supplied, verified
MMS artifact; automatic distribution of that converted artifact remains open.
All 64 previously frozen public DIMEx clips were replayed through both real local APIs. These are short studio Spanish clips, not a fresh Cap holdout. The original split of 16 development and 48 validation clips from different speakers was retained, without new tuning. Every matched lexical word is scored, including unchanged fallback intervals; there is no posterior-selected subset replacing the primary denominator.
On the same 481 matching validation words out of 494 annotated words:
| Metric | Current native baseline | Integrated MMS | Cached AssemblyAI |
|---|---|---|---|
| Mean start error | 48.43 ms | 36.09 ms | 56.84 ms |
| Mean end error | 47.55 ms | 36.00 ms | 57.03 ms |
| p95 start error | 150.63 ms | 77.94 ms | 128.30 ms |
| p95 end error | 147.01 ms | 95.53 ms | 134.97 ms |
| Both boundaries within 80 ms | 318/481 | 422/481 | 300/481 |
Recognition is unchanged:14/494 word errors (2.83%) versus cached AssemblyAI's 9/494 (1.82%). The timing improvement does not establish recognition parity. AssemblyAI has 486 matching validation words overall; the table uses the 481 common pairs, while the private report retains both complete provider denominators.
The hardened candidate's validation API median/p95 was777.74/785.16 ms, versus 536.64/800.58 ms for the unchanged baseline. This is one local run per version with 250 ms polling, not deployed capacity or a fresh remote speed comparison. The added alignment increases median completion time here; the small p95 difference is not evidence of a general speed improvement.
The initial candidate comparison, hardened replay and longer-recording checks made 263 actual API test requests:260 completions and three repetitions of the same expected, unbillable recognition error. This is not 263 unique recordings. The data includes 64 public reference clips, 15 reused Cap development recordings, nonspeech controls, exact crop/manual-cut pairs and warmup/recovery. Five additional native diagnostic replays were outside the application queue and were not billed.
All expected outcomes passed. Every completed job had one measured-duration usage event; every job used one attempt; failed jobs had no usage event. Idempotent replay, derived resources, authentication rejection, sample-exact crop offsets and recovery passed. All ephemeral API keys were revoked and owned processes were checked absent.
The 12 newer Cap development regressions preserved all recognition and confidence;
non-Spanish word times stayed unchanged. The Indonesian/Finnish development cases
used their already diagnosed expected_languages: ["all"] setting. They do not
establish parity for Cap's original 25-language hint list.
Three additional, previously frozen Spanish Cap recordings of 68.267, 118.464 and 270.037 seconds retained all 525 words and their confidence values. MMS changed 456 word intervals across multiple windows; all final intervals remained valid. The candidate API took 2.842, 3.361 and 7.199 seconds respectively. These recordings lack independent human word boundaries, so structural validity and changed intervals are not claims of better Cap timing accuracy.
Use SANDCHEST_INFERENCE_ENGINE=whisper,
SANDCHEST_WHISPER_WORD_ALIGNMENT=ctc-mms and
SANDCHEST_MMS_MODEL=/absolute/path/to/the/pinned/model.safetensors.
The normal launcher selects the whisper-mms Cargo feature. Existing English
routing configuration is preserved. A wrong engine, missing model or bad model
identity fails startup; the option cannot silently run the legacy Parakeet engine.
The final source passes 137 Rust tests with whisper-mms and 129 with the existing
whisper-ctc feature. The app passes 148 tests, with seven conditional model skips
and 1,568 assertions; typecheck, full lint and scoped startup checks pass. The normal
startup wrapper was exercised with the matching prebuilt binary and pinned cached
models, including both failure cases. No server or deployment was started by that
startup check.
Review fixes include explicit engine validation, typed acoustic-rejection counters and the active alignment model/SHA/language in native diagnostics. All 83 hardened API outcomes reproduced the first candidate's exact recognition, timestamps and confidence. A failed intermediate build caused by a duplicate test module is retained; the corrected test extends the existing module.
Two public duration labels differed by 1 ms because their original metadata used a different half-millisecond rounding convention. The audit independently decoded all input samples before the hardened run and checked usage against exact sample durations. Original plans and raw responses were retained. No transcript, model output or measured usage was changed to satisfy the audit.
Private evidence: mms-serving1/2/3, mms-api1/2, mms-cap-long1,
mms-source-promotion1 and mms-app-checks1. Validation is macOS/Metal only.
Language-selection compatibility, broader recognition quality, remaining
AssemblyAI API features and general 20% beta approval remain open. No shared
database migration, Cap write, deployment or production verification was performed.
The tested native fix is integrated in src/whisper/mod.rs; the language-selection
test name is clarified in src/whisper/language.rs. The matching whisper-mms
binary SHA-256 is
8b929ffd55fc1420ebc539a8b41e3a7d57e97db2041cc2ace6346db647669724.
Engine, model, alignment and environment defaults are unchanged. This does not
complete AssemblyAI API parity or approve an unrestricted 20% beta.
Previously, automatic detection could identify non-English speech, then the request's restricted hint list could select English as a fallback. That fallback label also selected the English-only Parakeet recognizer. Simply bypassing Parakeet while still forcing Whisper to decode English failed the development comparison: both long Cap examples showed greater provider disagreement. That variant was not promoted.
The integrated fix keeps existing hint selection, confidence thresholds and
response-language metadata, but uses the acoustic detector's language for native
recognition, engine routing, formatting and alignment. Explicit language_code
requests keep their existing behavior. Genuine detected English still uses the
fast Parakeet route and English CTC; detected Spanish can still use MMS. The
acoustic detector is a model prediction, not ground truth.
The AssemblyAI language-detection documentation describes expected languages as a restriction. Eighteen fresh submissions of two fixed public clips, reusing verified uploads, showed more varied live behavior: correct spoken-language text can coexist with a fallback response-language label, and some responses return languages outside the list. Model-specific probes were diagnostic comparisons, not unchanged Cap requests. Their receipts, settings, audio hashes and submitted/final transcript IDs were checked. These observations do not establish AssemblyAI's internal implementation or a universal override rule.
The frozen comparison made 190 native calls: 95 per build. It used all 30 previously selected Cap development recordings, 40 public Indonesian/Finnish clips, 16 public Spanish clips, two English hint-conflict cases, nonspeech and warmup/recovery/invalid/threshold controls. No heldout recording was used for tuning, and these are not new independent Cap recordings.
The Cap requests retained their original 25-language hint list; it was not changed
to all. Both builds completed 29/30 Cap recordings, retaining the same known
unreliable-speech error. Against 29 completed cached AssemblyAI references, counting
a native failure as missing words, provider disagreement fell from 1127/7748
(14.55%) to 796/7748 (10.27%). This is provider disagreement, not WER. Only the
two previously diagnosed Cap recordings changed; the other 28 outcomes stayed
unchanged. The full 59-recording fresh comparison above was not rerun or relabeled.
Independent reference transcripts also tested the failure mode deliberately:
| Spoken language excluded by hints | Clips | Previous word errors | Fixed word errors |
|---|---|---|---|
| Indonesian, original Cap25 hints | 20 | 356/388 | 21/388 (5.41%) |
| Finnish, original Cap25 hints | 20 | 329/323 | 29/323 (8.98%) |
| Spanish, English-only hints | 16 | 91/175 | 6/175 (3.43%) |
These are hint-conflict stress tests, not average production accuracy estimates. Insertions can make word errors exceed the reference count. Two public clips that previously failed now complete; no successful native case became an error.
All 32 successful cases unaffected by fallback preserved their text, words,
confidence, model and timing exactly. Across 90 paired successes, response
language_code, language_confidence and the diagnostic fallback flag stayed
unchanged. Model/backend identity intentionally changed on 58 cases; it would be
incorrect to describe all metadata as unchanged. The original audit's broad field
name is corrected by an additive second audit, without rewriting frozen results.
Actual native routing was checked on 86 successful speech responses. Both English cases with non-English response labels retained the exact English words, confidence and CTC timing from the English control. All 16 Spanish cases retained the exact words, confidence and timestamps of their prior explicit-Spanish MMS results. Existing acoustic rejection and word-boundary safeguards remain active.
The hardened binary then processed 94 actual local API jobs, including the same 30 Cap recordings with unchanged requests. The invalid-language case was covered by an HTTP400 endpoint check rather than a queued job. There were 92 completions and two expected errors: the known Cap recognition failure and the requested language-confidence threshold rejection.
Every completed API text, word interval, confidence, response language and model matched its frozen native result. All 94 jobs used one attempt. Exactly 92 unique usage events recorded the measured durations; neither error was charged. Upload, authentication rejection, idempotent replay, polling, derived resources and warmup/recovery checks passed. The isolated API keys were revoked and both owned processes were independently checked absent. The first harness attempt stopped before starting a process or submitting a job because its exact private runtime path lease was missing; that preflight failure is retained separately.
For the 29 completed Cap jobs, local upload-to-result median/p95 was 1.563/6.100 seconds, with 250 ms polling. No fresh remote Cap latency comparison, deployed capacity test, shared database operation or production check was run.
The source passes 140 Rust tests with whisper-mms and 132 with whisper-ctc.
The app passes 148 tests, seven conditional model skips and 1,568 assertions;
typecheck, full lint and the normal startup wrapper pass. Missing-model and
wrong-engine startup failures still fail closed. Callback tests now separately
cover successful detection, cancellation, expired deadlines and confidence
rejection before recognition.
The response language for the two Cap hint-conflict examples still differs from AssemblyAI, even though recognition improved. Hint-driven disambiguation of ambiguous speech, broader accents and language coverage still require evaluation. This does not establish measured quality or accurate timestamps for all published languages. The remaining speech-to-text options and full API contract gaps listed above also remain open. The general 20% Cap beta is not approved; licensing is not being treated as a blocker or an active workstream.
The API and TypeScript SDK now accept punctuate: false. Both native paths apply
presentation after recognition/refinement and optional CTC/MMS alignment. Default
and explicit true transcription text/words remain unchanged. Lowercasing and
edge punctuation/symbol removal preserve internal contractions, hyphens, decimals,
times, email addresses and URLs. The observed provider behavior also removes
edge symbols such as those in C++, C#, negative numbers and percentages.
Text and words are transformed together. Surviving word timestamps, confidence, speaker and channel values are unchanged. Empty nonspeech stays empty; incoherent text/word mappings, oversized output or a nonempty transcript reduced entirely to punctuation fail instead of publishing partial or fabricated speech. Memory growth is checked before append, with bounded work and cooperative cancellation.
A punctuation-off success includes an explicit native acknowledgement. The queue worker rejects a missing/wrong acknowledgement, so older workers cannot silently ignore the option. Invalid presentation is a sanitized, terminal 422 failure, with no inference retry amplification or billable completion. Request flags are stored and echoed; changing the flag under the same idempotency key conflicts. Use the matching native worker when enabling the updated API.
There were 32 fresh provider submissions on fixed public/authored audio:
31 completed and one all-null boolean request returned the expected 400.
No private Cap recording was submitted in these option probes. Twelve same-model,
same-format punctuation on/off pairs match the renderer's text and word text
exactly. Formatting and filler behavior were measured separately; this does not
implement disfluencies: false or full inverse text normalization.
The main native comparison made 208 calls across 104 task pairs, including all 30 frozen Cap development recordings, public language references, controls and text-option fixtures. It completed 204 calls, retaining the two expected error cases in both flag states. Across 102 successful pairs, all 8,743 words retained exact acoustic metadata and transcript confidence. Default transcription text, words, model, language, duration and confidence matched the preceding checkpoint. Twelve additional calls verified standalone Parakeet v2 English/v3 Portuguese, both formatting settings and the new acknowledgement.
An initial private diagnostic check failed because serde JSON parsing moved some f64 values by one ULP. That failed attempt and the diagnosis are preserved. Only the diagnostic round-trip comparison permits that one-ULP tolerance; direct native on/off and real API/native comparisons remain exact. No model-confidence tolerance or accuracy threshold was relaxed.
Twenty read-only GETs inspected existing Pro/U2 provider resources. This exposed three export gaps, now fixed: lexical URL punctuation no longer splits a punctuation-off transcript; paragraphs contain at most five sentence groups while retaining pause/speaker boundaries; absent speaker/channel fields are omitted from resource words. Non-null metadata and the original transcript word array are preserved. The paragraph cap also agrees with the provider's published paragraph behavior.
All 16 complete saved Pro sentence/paragraph response fixtures now match exactly, including punctuation/formatting combinations, Portuguese and Hindi. The additional U2 checks confirm the same null-field omission. This is scoped response evidence, not proof of identical semantic paragraph boundaries for every recording or language.
The final native binary SHA-256 is e1b961cc08b707aaefecbb94d4c6c45b8835c325abfb6603f3c79b94dd8d1d86.
Two 92-job local API runs were retained; the second used the final export fixes.
Its 88 completions matched native text, words, model, language, confidence and
duration exactly. Every completed sentence/paragraph export preserved all words
and their metadata; punctuation-off speech had one group and paragraphs respected
the sentence cap. Exports, idempotent replay, changed-option conflicts and
authentication also passed. There were exactly 88 unique duration-based usage
events, one attempt per job, no failed-job usage and no duplicate charges.
Ephemeral API keys were revoked and all owned processes stopped.
| Final local API, 29 successful Cap recordings per flag | Median | p95 |
|---|---|---|
| Punctuation on | 1.554 s | 6.316 s |
| Punctuation off | 1.559 s | 6.343 s |
These are serial local M4 Max API measurements with 250 ms polling, not production capacity, a statistical speedup claim or a new remote AssemblyAI timing comparison. Some earlier native functional checks overlapped compilation; their timing is not used as performance evidence. The broader 59-recording comparison above is historical and was not rerun.
Evidence remains private under whisper-large-20260827/durable1: provider option
probes, punctuation-build1/2/3, punctuation-eval1/2, punctuation-api1/2,
punctuation-provider-resources1, punctuation-resource-fix1, punctuation-app1/2,
punctuation-legacy1, punctuation-ctc1 and punctuation-sdk1. Failed attempts and
exact prior source versions are retained; earlier plans are not rewritten after
intentional source changes.
The general 20% Cap beta remains unapproved. This closes concrete option/export gaps; it does not improve measured recognition accuracy or establish all-language word-timing parity. Multilingual quality, known recognition failures, filler removal/defaults, complete formatting, custom vocabulary, diarization, multichannel, redaction and other unsupported STT options remain work. No deployment, shared database migration, production checks or Cap source/data changes were performed. Licensing is not a blocking workstream.
The fixed guarded Qwen candidate is a stronger Hindi recognition candidate than
our current Turbo path on this new public reference set. It is not serving Cap
requests yet. No decoder, model weight, serving source or safety gate changed in
this experiment. The current native/API binary remains
e1b961cc08b707aaefecbb94d4c6c45b8835c325abfb6603f3c79b94dd8d1d86.
The Qwen candidate remains
40e3de4b8d7a5a375a861ee98e091f1bcb72003bc0359a6fe896cafa2d076f03,
using the existing 1.7B 8-bit artifact
bf304b009cc7eca79283056f787b44c952d24ac22cec787b39732bba3c23c13c.
The corpus has 100 unique Hindi FLEURS recordings, 19.039 minutes of audio, 60 published validation clips for development and 40 published test clips reserved for this fixed-candidate holdout. All 74 previously tested examples (10 pilot plus 64 expanded) were excluded by source ID, normalized references and decoded PCM. Selection used a fixed SHA256 rank and equal published gender counts within each split, before inference. Audio lengths were 3.66–26.28 seconds. There is no claim of speaker independence or exclusion from either model's training data.
The exact dataset revision is
70bb2e84b976b7e960aa89f1c648e09c59f894dd. Both complete source Parquet files were
verified against published size and SHA256. One bounded download did not finish;
its prefix and attempt receipt were retained, then an explicit range resume
completed the same immutable file. No replacement sample was selected. An
independent audit rechecked 100 audio/reference pairs, 74 prior exclusions, split
separation, all source IDs and PCM hashes, and private file permissions.
All three systems received identical canonical audio bytes and explicit Hindi.
The provider and current worker used Cap's punctuation/formatting/disfluency flags;
Qwen used its fixed temperature 0, 1,024-token decoder. There were 210 native calls
(100 references plus warmup, recovery and 3 nonspeech controls per backend), and
100 fresh AssemblyAI uploads/transcriptions. Every provider result used
universal-3-5-pro. Qwen is still a text-only prototype, so this does not claim
that it implements those presentation flags or the full STT response contract.
WER uses every actual completed transcript against the published raw reference, with Unicode combining marks retained. There is no number expansion, spelling repair or model-output-driven normalization. An independent dynamic-programming scorer agrees with all 300 recognition scores.
| System | Development 60 | Holdout 40 | All 100 / 2,426 reference words |
|---|---|---|---|
| Current Rust Turbo | 479/1,473 = 32.52% | 310/953 = 32.53% | 789/2,426 = 32.52% |
| Guarded native Qwen candidate | 235/1,473 = 15.95% | 141/953 = 14.80% | 376/2,426 = 15.50% |
| Fresh AssemblyAI Pro | 286/1,473 = 19.42% | 217/953 = 22.77% | 503/2,426 = 20.73% |
Qwen produced 25.25% fewer reference word errors than AssemblyAI across these 100 clips and 52.34% fewer than current Turbo. On the 40-clip holdout, the Qwen-minus-AAI WER difference is −7.97 percentage points, with a paired clip bootstrap 95% interval of [−11.05,−4.86]. The development interval [−6.89,+0.77] includes zero. These intervals resample clips within gender, not independently identified speakers; they do not establish all-language or Cap-domain superiority. The holdout was not used to choose a new decoder setting, model, threshold or normalization.
The separately published normalized transcript gives WER across all 100 clips of 14.96% Qwen, 32.36% Turbo and 20.40% AssemblyAI. It is secondary evidence, not a replacement for the raw-reference result after inspecting outputs.
All three systems actually completed 100/100 reference clips. The strict local
word validator flagged 10 zero-duration AssemblyAI words across 9 results. Those
are timestamp diagnostics, not provider transcription failures. The first frozen
report scored those responses as empty under its stricter usability rule,
producing a 27.70% gated score for AssemblyAI; that is not pure recognition WER
and must not be used as the headline accuracy comparison. The already-planned
text_only_score preserved all actual provider words. The additive
recognition-audit.json uses those full transcripts, yielding 20.73%, and retains
the original gated report unchanged. No clips, words or timings were fabricated
or repaired to improve a system's reported result.
Both native backends passed warmup/recovery and silence/tone/noise checks. Qwen reported no ASR invocation for nonspeech; Turbo diagnostics showed VAD rejection, zero encoder/full calls and no segments. All native results retained valid text, finite word metadata where supplied, exact duration and no embedded NUL or replacement scalar. Both owned native processes were independently confirmed absent and their runtime lease released.
| Measured boundary, 100 clips | Median | p95 |
|---|---|---|
| Current Rust recognition + word timing | 785.5 ms | 1,224.6 ms |
| Qwen native text only | 864.6 ms | 1,478.6 ms |
| Fresh AssemblyAI upload→completed result | 4,687.9 ms | 7,737.8 ms |
Qwen was slightly slower than the current worker even before adding alignment. The provider row includes network/upload/queue/polling and is not a like-for-like native speed comparison. This run did not exercise a Qwen HTTP/queue/persistence path, production concurrency, NVIDIA/CUDA, deployment or Cap's live rollout.
This larger disjoint result supports continuing a Hindi-specific Qwen integration, not replacing the multilingual serving model globally. The earlier ID/FI tests still showed a Finnish regression with Qwen. Next work is real native recognition confidence, reliable word alignment, bounded long-audio handling and explicit/auto language routing, followed by full API and Cap development/independent tests.
Existing Hindi alignment evidence is still inadequate: MMS technically aligned 610 supplied words across 12 windows of the difficult Cap recording, but 77 words had uncertain boundaries and 0/12 windows passed the existing conditional screen. That posterior is conditional on supplied text and is not human timing gold. The official FLEURS schema contains no word-boundary timestamps. A real Hindi timing reference is still needed; assigning plausible intervals is not sufficient verification.
Private receipts are under whisper-large-20260827/durable1/fleurs-hindi-fresh1,
hindi-fresh-comparison1 (unused initial plan), hindi-fresh-comparison2 and
hindi-fresh-doc1. The plan SHA256 pinned outside the experiment before calls is
e37fbc49f92062af287e9afaa7164a08e4e9f0f4a2656258d1093b89b4c35b15.
The existing checkpoint of 152 app tests and 148 Rust tests with MMS and 92-job API run above are
unchanged; they were not rerun or relabeled as Qwen serving verification.
The general 20% Cap beta remains unapproved. This is a measurable candidate quality gain, not a deployed implementation gain. No Cap source, media, database, shared service or production deployment changed. Licensing remains outside scope.
The Hindi candidate now has native greedy-token probabilities, exact UTF-8 byte mapping to its surface words, and cooperative request interruption. These are private, opt-in candidate interfaces, not a change to the current STT API or its serving model. Licensing remains outside scope.
The scorer checks finite raw logits, computes a stable F32 log-softmax for the chosen token, and retains strict token IDs and byte spans. It reconstructs split UTF-8 sequences before mapping words. Shared tokens are counted once in the transcript aggregate; special/whitespace-only tokens and EOS do not inflate word scores. EOS is recorded separately. Empty nonspeech has no fabricated confidence.
The value is the uncalibrated geometric mean of contributing raw token probabilities. It is not a calibrated probability that a word is correct, and not the aligner's conditional boundary confidence. Byte offsets are text coordinates, not millisecond timestamps. They are not exposed as API word times.
The score checkpoint passed 20 focused tests and 270 actual native calls. All 124 development pairs kept exact prior default transcripts and identical text with scores enabled. An independent tokenizer/byte/probability audit checked 11,439 token scores and 3,030 words. This uses 70 Hindi development references, 40 Indonesian/Finnish references, and 14 existing chunks from two Cap development recordings. No heldout clip or new provider call was used.
Exact EOS at the measured 144-token limit succeeded with identical scores. Limits of 143 and one token rejected without partial output. Invalid requests rejected; silence, tone and noise bypassed ASR; the next valid request recovered exactly. Synthetic tests cover split Unicode scalars and cross-word shared tokens.
| Language | Cases | Default native p50 | Scored native p50 | Paired overhead p50 |
|---|---|---|---|---|
| Hindi | 70 | 897.7 ms | 911.7 ms | 29.3 ms |
| Indonesian | 27 | 356.6 ms | 364.9 ms | 7.8 ms |
| Finnish | 27 | 522.4 ms | 538.9 ms | 16.0 ms |
These paired times cover native recognition/scoring only, without acoustic word alignment or the HTTP/upload/queue/persistence path. They are not a production capacity or AssemblyAI speed comparison.
Both controlled entrypoints enforce finite mono-16k audio of at most 30 seconds, amplitude at most eight, greedy decoding and 1–1,024 tokens. The caller supplies a small callback using its original deadline and cancellation state. Checkpoints follow mel computation, synchronous encoder/embedding evaluation, every token step, token/EOS scoring and final text assembly. Errors discard all partial text/scores and the request-local KV cache on the inference thread.
This is cooperative cancellation: it stops the next stage after the current GPU
operation finishes, not a running GPU kernel. The private harness's
cancel_after_ms schedules an external atomic signal and is a test probe, not a
hard timer guarantee. timeout_ms checks elapsed time from the original request,
including audio load/VAD. A final check after joining the timer covers response
completion races and early nonspeech returns.
The final controlled build passed 28 focused tests and 375 actual native calls on one model instance:
- 21 cancellations at fixed encoder, decoder, scoring and output checkpoints;
- 12 immediate or real 250/500 ms cancellation/deadline cases;
- 66 immediate clean recovery requests after those 33 interruptions;
- all 124 default/scored development pairs with exact cached text and scores;
- 12 nonspeech checks, eight invalid control requests, and long timer/deadline success checks proving timer threads are promptly joined.
Every interruption returned no partial transcript or scores. Each immediate recovery matched the prior clean result exactly. Nonzero timer/deadline cases interrupted actual token generation. The largest observed response overrun past the requested delay was 11.5 ms in this run; this is a local observation, not a latency SLA. The process exited normally and its PID was independently checked absent.
An initial unit run exposed unknown fields accepted by a serialized checkpoint variant; this was corrected and the failed attempt retained. Source review then found a response-completion timer race and missing bounds in the new controlled text method; both were fixed before the actual 375-call run and covered by tests. The legacy unscored method remains available with its prior behavior. The new controlled methods are the bounded integration surface.
Private sources, patches, binary hashes, plans and receipts are in
whisper-large-20260827/durable1/qwen-scored1, qwen-scored-eval1,
qwen-controlled3 and qwen-controlled-eval2. Superseded attempts remain recorded.
The controlled binary SHA256 is
ab97ebdf4f795aa2927bfee2d79cc7bd9e0c724ca78539bbf307e299b925fd38.
This preserves the previous Hindi recognition result; it is not new WER evidence. Reliable Hindi acoustic word timing, complete long-audio merging, explicit/auto language routing and the full API/Cap regression still remain. The existing chunk helper that skips failed chunks is not used as a serving implementation. The current Rust worker and its previous app/API checkpoint remain unchanged; those app checks were not rerun or reclassified as Qwen API verification.
The general 20% Cap beta is still not ready. No Cap source/data, shared database, shared service, deployment, default model or live traffic was changed.
Historical checkpoint: the compiler stall subsequently cleared. The completed recording verification in the following section supersedes this build status.
Licensing is outside scope. This checkpoint keeps completed measurements separate from the new code that could not yet be built and tested.
The private candidate combines the exact existing Qwen recognizer and MMS acoustic aligner in one process, against one actually mapped MLX library and one allocator policy. It preserves each display word and its original recognition confidence. MMS provides sample intervals directly; no placeholder times are assigned.
The unchanged screen requires at least 0.95 conditional boundary mass within 80 ms for every word. One uncertain word rejects the complete result; no partial word list or transcript is accepted. Private candidate boundaries are diagnostics, not successful API words. Conditional timing mass is not recognition confidence or a measured probability of human timing correctness.
The combined build passes 28 focused tests. 251 real native calls preserve all 136 recognition comparisons exactly, including every score field. Four cancellation/deadline cases discard content and recover exactly on the next request. Nonspeech bypasses recognition and acoustic alignment.
| Hindi group | Cases | Aligned | Accepted | Uncertain words |
|---|---|---|---|---|
| Earlier short development references | 10 | 10 | 7 | 4 |
| Fresh-corpus development references | 60 | 51 | 43 | 26 |
| Difficult reused Cap recording windows | 12 | 12 | 0 | 77 |
Nine short cases are rejected because decimal digits are not supported by the MMS target normalizer. Digits must not be dropped or display text rewritten to hide this gap. The difficult Cap recording has broader interior alignment uncertainty; source review did not find a CTC or sample-to-ms arithmetic defect. There is still no verified human-labelled Hindi word-timing result.
The accepted short clips contain 1154 words. Native recognition-plus-alignment p50/p95 is 1253.3/ 3227.7 ms; alignment alone has p50 135.0 ms. These successful-case observations are not full-API latency or a controlled throughput comparison.
For the 43 accepted clips with cached AssemblyAI results, 828 common words have start/end mean absolute differences of 66.25/96.43 ms; both boundaries are within 80 ms for 483 words. These are provider differences, not human timing errors. Rejected candidates are reported separately. No heldout clip or new provider/Cap request was used.
The new private recording module is designed to process finite mono-16k recordings up to one hour through bounded, contiguous windows. It retains all audio samples, offsets UTF-8 spans and token references, preserves word confidence, computes aggregate confidence from contributing tokens once each, and keeps EOS per chunk. It does not deduplicate repeated words or fabricate missing times. Any failed, uncertain or interrupted chunk rejects the recording instead of returning partial success. One original deadline covers the whole request.
Review caught a long-request trace-buffer limitation. Long diagnostic tracing is now explicitly rejected; deterministic chunk-boundary cancellation remains a private test option. The implementation also states that quiet cuts are not validated speech boundaries: selecting a low-energy block can split a spoken word. It must remain experimental until boundary quality is validated.
The prepared test corpus contains three complete reused Cap recordings, 633.824 seconds total, three constructed Hindi speech/silence recordings, 136 short recognition comparisons and 82 short timing comparisons. None of the new recording tests has run. Corpus preparation and source review are not runtime verification.
Cargo first waited on a shared cache used by other Cap builds. An isolated
dependency cache was prepared without changing those builds. Compilation then
stalled before four helper programs entered their code: samples show
_dyld_start, negligible memory/CPU and no child process. Code-signature validation
succeeds; an exact-byte copy preserving metadata also times out. Memory pressure
is not indicated (96% available, no swap). The previously verified combined
worker still starts normally. The precise loader cause is unresolved.
Only owned stalled processes were stopped. No shared service, foreign build or security setting was changed. An attempted direct metadata check against cached dependencies failed to resolve transitive crates; it is not a successful typecheck. All attempts and diagnostics are retained, and there is no successful recording build receipt.
Private completed evidence: whisper-large-20260827/durable1/qwen-mms7,
qwen-mms-eval1 and qwen-mms-audit1. Pending recording source and corpus:
qwen-long4 and qwen-long-eval2. The verified combined binary SHA256 is
aadf811a7f6c56b414a955f69f3942cec86d7ac8b297172c497c932e770e717d.
A general 20% Cap beta is still not ready. Numeric alignment, reliable Hindi word timing on difficult Cap audio, speech-safe long-recording boundaries, explicit/auto model routing and actual Qwen API integration remain open, alongside the wider AssemblyAI API/language gaps above. The earlier Hindi WER improvement is unchanged; these checks add no new human-reference accuracy claim.
Current serving source and prior app-check inputs still match their verified hashes. App/API checks were not rerun or counted as Qwen API tests. No Cap source/data, shared database/service, deployment or live traffic was changed.
Licensing is outside this workstream. This checkpoint is experimental and does not change the current Rust serving worker, HTTP API or default model.
The combined process loads the existing pinned Qwen and MMS models against one actual mapped MLX library and one allocator policy. Qwen owns exact display text and recognition confidence; MMS supplies acoustic intervals for those same words. No initial placeholder word intervals are needed. Every recognized word must survive target normalization and alignment, including its original confidence.
The existing 0.95 conditional boundary-mass screen within 80 ms is unchanged. If any word fails it, the request returns no accepted transcript or word list. Candidate boundaries remain private diagnostics. Conditional timing mass is not recognition confidence, and neither is human-validated timing accuracy.
The combined build passed 28 focused tests and 251 actual native calls. All 136 recognition comparisons preserve exact text and every score field. Four cancellation/deadline cases return no partial content and immediately recover to the exact clean result. Nonspeech bypasses both recognition and alignment.
| Hindi group | Cases | Aligned | Accepted completely | Uncertain words |
|---|---|---|---|---|
| Earlier short development references | 10 | 10 | 7 | 4 |
| Fresh-corpus development references | 60 | 51 | 43 | 26 |
| Reused difficult Cap recording windows | 12 | 12 | 0 | 77 |
The nine unaligned short cases contain decimal digits, unsupported by the current MMS target normalizer. Digit stripping or changing the recognized display text is not an acceptable repair. Source review found broader acoustic/target mismatch in the Cap recording, mainly interior words, not a sample-to-ms arithmetic defect. A numeric pronunciation path and a better validated alignment route remain work.
The 50 accepted short clips contain 1154 words. Their native recognition-plus-alignment p50 is 1253.3 ms and p95 3227.7 ms; alignment alone has p50 135.0 ms. These are local successful-case observations, not full-API latency or capacity measurements.
Only 43 accepted clips have cached AssemblyAI boundary comparisons: on 828 common words, start/end mean absolute differences are 66.25/96.43 ms, and both boundaries differ by at most 80 ms for 483 words. This is provider agreement, not human timing accuracy. Failed/uncertain candidates are tabulated separately in the private audit and do not inflate acceptance. Human-labelled Hindi word timing has not been verified.
The private long_audio mode requires recognition scores and accepts canonical,
finite mono-16k audio up to one hour. Native recognition still runs in at most
30-second windows on one serial model lane. A single original deadline and
cancellation signal cover loading, planning, all chunks, merging and final output.
Every audio sample belongs to exactly one contiguous window. Text is joined without trimming or word deduplication. UTF-8 byte ranges and token references are offset exactly once; confidence is preserved per word, and aggregate confidence uses contributing native tokens once each. EOS remains per chunk, never a made-up single whole-recording score. Accepted word times receive the original sample offset once; no interval is interpolated, stretched or clamped.
Any missing, failed, uncertain or interrupted chunk rejects the entire recording without content. The ledger cannot be reused after a failed merge. Silence adds no text or confidence. Token/text/word counts and memory-facing inputs are bounded. Long diagnostic tracing is rejected explicitly, avoiding the short trace buffer's global limit; this does not disable actual cancellation or deadlines.
Quiet cuts are not yet validated speech boundaries. The current policy selects the quietest 100 ms block between 20 and 29 seconds and can split a word. Diagnostics state this limitation explicitly. Exact audio coverage and exact chunk merging do not prove whole-recording recognition quality; this route remains experimental.
The final recording build passed 39 focused tests and 254 actual native calls:
- 136 short recognition and 82 short timing regressions preserve previous output;
- three reused complete Cap recordings, 633.824 seconds in total, preserve every standalone chunk's exact text, tokens and word confidence;
- the difficult Hindi recording still rejects when word timing is required;
- 2/3 constructed multi-window Hindi speech/silence cases return complete timed words, checked against fresh standalone calls on the exact same chunks;
- a 61-second silence recording produces empty words and no confidence;
- cancellation after one and two completed chunks, a real elapsed deadline and non-EOS termination discard all partial content, each followed by exact recovery;
- invalid long-mode/trace requests reject before inference.
Constructed recordings are test fixtures, not new natural Cap samples. These runs reuse development data; no heldout data, new provider call or Cap access was used. The local compiler-helper stall cleared without an OS or security-setting change; the exact same frozen source then built successfully. Other Cargo builds had finished before the recording run. These serial native observations still are not a full-API or paired-provider throughput benchmark. A private offline Cargo cache avoids shared cache locks. Superseded attempts remain in receipts.
| Reused Cap language | Audio duration | Native recognition and merge | Words |
|---|---|---|---|
| Indonesian | 156.544 s | 4.279 s | 177 |
| Finnish | 177.206 s | 11.083 s | 399 |
| Hindi | 300.075 s | 24.819 s | 610 |
These complete Cap runs have timing disabled and do not establish AssemblyAI latency parity. The Hindi text path takes about 24.8 seconds for five minutes of audio on this local machine before acoustic word timing.
Private sources, binaries, hashes, plans and receipts are under
whisper-large-20260827/durable1/qwen-mms7, qwen-mms-eval1,
qwen-mms-audit1, qwen-long4 and qwen-long-eval2.
The recording binary SHA256 is 03a46f16a47038bb7982198dce2c68a57f9e44c0ce377fb91a0d18a06f4b285a.
A general 20% Cap beta is still not ready. Hindi numeric targets, reliable timing on difficult real audio, speech-safe long-recording boundaries, explicit/ automatic model routing and actual Qwen API integration remain open. Wider AssemblyAI language/feature parity is not established by these three-language regressions. The earlier Hindi reference WER improvement still stands; this checkpoint adds no new human-reference accuracy claim.
The prior serving source and app verification inputs still match their recorded hashes. Those app/API checks were not rerun or counted as Qwen API tests. No Cap source/data, shared database/service, deployment or live traffic was changed.
This is a private diagnostic build, not a timing repair or serving promotion.
For each original word boundary, it measures the maximum existing conditional
posterior mass in any clipped ±4-frame window. The end center remains the last
occupied frame, end_frame - 1. Equal maxima prefer the center nearest the
original Viterbi boundary, then the lower frame. Actual clipped bounds are
included, so an edge window is not mistaken for a full nine-frame window.
The original Viterbi path, display words, recognition confidence, intervals, uncertainty and acceptance are unchanged. The ordinary alignment entry point keeps diagnostics off. Independent maxima are not a joint CTC path and are not human timing evidence; they are never substituted for output timestamps.
The build passed 75 focused tests (44 alignment, 31 entry/control/recording), including six new diagnostic tests, the existing exhaustive 105-lattice oracle, clipped edges, tie order, token-row isolation, invalid probabilities and mid-scan cancellation. 228 actual native calls then preserved all 136 recognition results and 82 Hindi timing results exactly, excluding elapsed-time fields. Four cancellation/deadline cases returned no partial content and recovered to the exact clean result on the same instance. The one mapped MLX library was verified and the owned process exited normally.
| Group | Uncertain words | Both independent peak windows can reach 0.95 | Still broad at best center | Additional whole cases theoretically possible |
|---|---|---|---|---|
| Earlier short Hindi development | 4 | 1 | 3 | 1 |
| Fresh-corpus Hindi development | 26 | 8 | 18 | 2 |
| Difficult Cap recording windows | 77 | 12 | 65 | 0 |
The theoretical case count is only an upper bound: independently chosen centers can violate the ordering of adjacent words. Across the three groups there are 11 such adjacent overlaps, despite no individual word having inverted centers. A constrained path would be necessary before considering any changed times. It cannot solve the 65 broad Cap words or the nine digit-normalization rejections.
Decision: do not spend the next iteration merely moving boundaries or relaxing the 0.95 screen. Focus on the acoustic alignment and pronunciation targets first. The 50/70 short-clip acceptance result and the rejection of all 12 difficult Cap windows remain unchanged. Human Hindi timing accuracy, automatic language routing, complete API parity and general beta readiness remain unproven.
Private reproducible artifacts are durable1/qwen-peak1 and
durable1/qwen-peak-eval1. The executable SHA-256 is
3f748cbe06fde723895023b9820b7fc07eb58e5580dcfdb73032ed942595fcec.
Only existing development audio was used: no new provider request, holdout use,
Cap access/write, shared service/database change, deployment or live traffic.
Licensing remains outside this workstream.
This private candidate fixes a target-coverage gap without rewriting recognized text. The original model words, confidence and Unicode code point offsets remain intact. Canonical numeric pronunciation components are aligned as one group and collapse into exactly one original display word. Inserted phonetic spaces do not become transcript words. No timestamp is interpolated or fabricated.
The lexical tables are generated from pinned Unicode CLDR revision
08512075888f5f7263004e3381be69406831c79d:
Hindi number rules
and Hindi percent wording.
They define canonical targets, not the actual pronunciation of a year,
identifier or code-switched number.
The bounded parser supports nonnegative integers through 999,999,999 using one
ASCII or Devanagari digit script, strict Indian or Western comma grouping,
explicit attached वाँ/वें/वी ordinal forms, and attached integer percentages.
It retains irregular 0–6 ordinal pronunciations. Unambiguous wrapping punctuation
is preserved. Leading zeros, mixed digit scripts, malformed grouping, signs,
decimals, fractions, currency, dates/ranges/times, unknown suffixes and trailing
periods/commas remain unsupported rather than being silently stripped.
If any numeric surface is unsupported, the entire request takes the original raw target path and retains its original rejection. Expansion bytes/scalars, target labels, source ownership and original word ordering are bounded and checked; cancellation propagates without partial output. Other languages and nonnumeric target mappings remain exact. The original 0.95 conditional boundary screen still rejects the whole transcript if any word is uncertain.
The native build passed 83 focused tests (52 alignment and 31 entry/control/ recording), including eight new numeric mapping tests. The actual replay made 228 calls, preserving all 136 recognition results and 74 unchanged Hindi timing results. Four cancellation/deadline failures returned no partial content and recovered on the same instance. The mapped MLX library, normal exit and process absence were verified.
| Hindi group | Cases | Aligned before → after | Accepted completely before → after |
|---|---|---|---|
| Short development clips | 70 | 61 → 69 | 50 → 52 |
| Newly supported numeric clips | 8 | 0 → 8 | 0 → 2 |
| Difficult Cap recording windows | 12 | 12 → 12 | 0 → 0 |
The eight newly aligned clips contain 217 recognized words, including 16 numeric words. Eleven numeric words pass the conditional timing screen; five remain uncertain. Five additional nonnumeric words in those clips are uncertain, so successful numeric normalization alone cannot make those recordings acceptable. One mixed date-range clip still has the original target rejection. The difficult Cap recording remains unchanged with 77 uncertain words.
For the two newly accepted clips, cached AssemblyAI comparison finds 39 common words, with start/end mean absolute differences of 100.51/145.13 ms. This is provider agreement, not human timing accuracy. The numeric subset is selected from the full-transcript alignment, preserving context when numbers repeat; it is not independently rematched against a shortened transcript.
Remaining work: larger numbers can be spoken using different constructions from canonical spellout. The remaining numeric failures need pronunciation and acoustic investigation; the current screen must not simply be loosened. Hindi human word timing, general acoustic coverage, automatic language routing, full API compatibility and beta readiness remain open. No default model or API change is justified by this experiment alone.
Private receipts: durable1/qwen-numeric1, durable1/qwen-numeric-eval1 and
durable1/qwen-numeric-audit1. Executable SHA-256:
85415118fd1252587ddb0a5f53a2085bc3230f589b598778329d229e30b144b1.
No new model weights, provider requests, Cap access/writes, heldout clips,
shared service/database changes, deployment or live traffic were involved.
Licensing remains outside this workstream.
This private experiment offers an equivalent hundreds reading for cardinal values 1000–9999, using only the already pinned CLDR small-number and hundred lexemes. The canonical reading remains available. A bounded CTC graph chooses all word pronunciations jointly by maximum complete acoustic path log probability, before any timestamp screening. Exact full-score ties retain the complete canonical target. No pronunciation is selected because it makes a timing check pass.
The graph keeps alternatives separate inside each word, requires blanks between repeated labels, and audits transitions, complete token occupancy, raw CTC collapse and acoustic score. The chosen target then passes through the original linear CTC aligner, whose score must equal the graph score exactly. Original transcript text, recognition confidence, numeric values and Unicode source spans remain unchanged. The 0.95 conditional boundary screen is unchanged; uncertain output still returns no partial transcript. The posterior is conditional on the selected pronunciation, not a calibrated distribution over all possible readings.
Allocation, input dimensions and cancellation are bounded. The graph permits at most four alternatives per word, 1024 labels on any complete path and 4096 labels across alternatives. At 2000 frames the predecessor table is bounded at about 32.8 MB; graph nodes, edges, score rows and other request allocations are additional. If the alternative-label budget is exceeded, the canonical target is retained with an explicit private diagnostic. Source and control failures are not suppressed.
The native build passed 95 focused tests: 64 alignment tests and 31 entry, control and recording tests. Twelve new tests cover exhaustive tiny raw-label paths, repeated labels across words, branch isolation, canonical ties, quoted numeric spans, malformed source metadata and cancellation. A separate read-only code review found no correctness blocker. These checks establish implementation behavior, not acoustic quality.
The full frozen replay made 228 actual native calls. All 136 recognition results remained exact. All 78 timing cases outside the alternative scope and their 77 available peak diagnostics remained exact. Four cancellation and deadline cases returned no partial content and recovered on the same instance; the mapped MLX library, normal process exit and absence were verified.
| Measure | Canonical numeric target | Acoustic pronunciation selection |
|---|---|---|
| Large numeric words passing conditional timing | 1/6 | 2/6 |
| All numeric words passing conditional timing | 11/16 | 12/16 |
| Eligible clips passing completely | 1/4 | 1/4 |
| Short Hindi clips passing completely | 52/70 | 52/70 |
| Difficult Cap recording windows passing completely | 0/12 | 0/12 |
Four large numeric words chose an alternative, with no complete-clip acceptance gain or regression. Two alternatives scored better acoustically but their minimum boundary masses fell from approximately 0.190 to 0.103 and 0.806 to 0.096. The only accepted eligible clip retained its canonical target exactly. Its 17 common AssemblyAI words are unchanged evidence, not a new accuracy gain; no numeric boundary match was available under full-transcript correspondence.
Decision: retain this as a reproducible, unpromoted experiment. Do not replace the canonical baseline or add timing-based pronunciation selection. The remaining work needs better acoustic alignment and human timing validation, alongside the open language, API and reliability gates. This result does not make the general 20% Cap beta ready.
Private receipts: durable1/qwen-pronunciation1,
durable1/qwen-pronunciation-eval1 and durable1/qwen-pronunciation-audit1.
Executable SHA-256:
3e12f6fea9711e1e59cecc3516c0a81b8c89008b66aa00ca9e2c13425059acca.
No new weights, provider submissions, Cap fetches or writes, holdout use, serving/API
changes, shared database or runtime changes, deployment or live traffic were
involved. Licensing remains outside this workstream.
The pronunciation experiment did not improve complete-recording acceptance, so the next investigation targets the acoustic model. This section is feasibility evidence, not a new recognition or timestamp result.
The NVIDIA Hindi Conformer CTC model's published archive configuration contains 128 tokens and no ordinary Latin-letter tokens. In the existing difficult Cap recording, 124 of 610 recognized words across five windows contain Latin letters. Those original words cannot be represented literally by that inventory. This is a concrete limitation for direct alignment of the unchanged mixed-script text; it is not a measurement of NVIDIA's acoustic accuracy.
Meta's official CTC architecture uses a Wav2Vec2 encoder and linear output projection. The 300M v2 configuration is close to the existing native MMS backbone: seven convolution layers, 24 Transformer layers, hidden dimension 1024 and 16 attention heads. The exact loader, 10288-class head, tokenizer, blank handling and waveform semantics still need independent native/reference verification.
A bounded parser inspected the two official tokenizer artifacts without executing remote code. The v2 inventory represents every character in all 136 existing development transcripts, including the 26 windows from three Cap recordings, without lowercasing, punctuation removal or unknown-token substitution. The v1 inventory lacks uppercase Latin characters. This is literal inventory coverage, not execution of the official tokenizer, declared-language verification or proof that the acoustic model can align those characters accurately.
The v2 checkpoint was downloaded from the URL in the pinned official asset card.
Its length is 1,304,065,508 bytes and its recorded full SHA-256 is
8ce340ada22435d189908a8af67e3bb04899b6f2f5329d036248b9c3e38d2b50.
The response ETag and length matched the preflight, but the card provides no
separately published v2 SHA-256; the v1 Hugging Face checksum is not v2 provenance.
Safe CPU-only deserialization found 423 F32 tensors / 325,982,896 elements.
It used the existing verified Torch 2.9.1 runtime, with weights_only=True and
memory mapping. No model forward or GPU call occurred.
The source reference is pinned to OmniASR commit
81f51e224ce9e74b02cc2a3eaf21b2d91d743455 and fairseq2 0.6.0 commit
6fa0aaf178db437bde0fae125b36105dc123119d. The next step is exact tokenizer and
F32 numerical equivalence on identical predecoded mono 16 kHz inputs, followed
by the existing difficult recordings and new independent quality evaluation.
Any official fairseq2 runtime must have matching native-extension/PyTorch versions.
Language coverage, human word timing, complete API parity and beta readiness
remain unproven.
Private artifacts: durable1/hindi-ctc-feasibility1,
durable1/omni-ctc-feasibility1, durable1/ctc-vocabulary-audit1,
durable1/omni-ctc-model1 and durable1/omni-reference-source1.
There were no new provider submissions, Cap fetches or writes, holdout use,
serving changes or deployment. Licensing remains outside this workstream.
This extends the preceding metadata investigation. All work remains in private experiment directories; the serving model and API implementation are unchanged. The 136 frozen development transcripts comprise 110 short public/development clips and 26 windows from three previously used Cap recordings. They are not 136 new Cap recordings. Recognition hypotheses were reused, not rerun or edited.
The actual SentencePiece core, pinned to commit
58f256cf6f01bb86e6fa634a5cc560de5bd1667d, matched the IDs and decoded text of
all 136 transcripts / 21,234 characters. Its 167 protocol calls also checked
whitespace, unknown characters, literal control-token spellings, decoding,
malformed requests, and recovery. This executed the tokenizer core, not the
fairseq2n extension. The native timing helper deliberately accepts only the
canonical single-ASCII-space subset verified here; unsupported inputs reject.
Conversion preserved all 423 F32 tensors with a strict bijective mapping and
exact value round-trip. The converted 1,303,985,280-byte artifact has SHA-256
904f96dbaab1eb01560f0a848ebae871e41947628bd5afab53ff3e39a8ba4bef.
No output rows were removed or quantized. The private acoustic executable is
omni-native1/omni-ctc-probe, SHA-256
d4cd363aeb8528cd8da32371f9ad2b0ffb12ce44668bb151599772f81b516a80.
The numerical reference loads the original checkpoint through an independent inverse mapping into the unchanged TorchAudio Wav2Vec2 implementation. It uses F32, CPU SDPA MATH, evaluation mode, one unpadded mono waveform, and explicit waveform normalization. Source review found the effective architecture equivalent for this scope. It is not an execution of official fairseq2/fairseq2n inference, nor proof of padded/batched behavior or equivalence to the official BF16 default.
Nineteen inputs include public and Cap Hindi/Indonesian/Finnish audio, silence, tone, seeded noise, short frame-count boundaries, and the 35-second input limit. Each runs through native raw-input and shared-normalized-input paths. Of the 38 comparisons, 36 pass the predeclared maximum-absolute 0.01 and RMSE 0.001 tolerances. All frame argmax IDs match. The same 18-second Hindi input fails both variants: maximum log-probability differences are 0.04233 and 0.03857, concentrated near two frames. Maximum RMSE is 0.000220. The original gate remains failed.
A trace-only build preserved that input's native scores exactly while exposing all 24 encoder layers and the effective positional weights. CPU F64 diagnosis shows final-score maximum differences of 0.02545 from native F32 and 0.01312 from CPU F32. Small initial differences grow through the encoder; no static shape, mapping, or positional-weight defect was found. This supports floating-point sensitivity, not a waived gate or a replacement correctness oracle. On this input's 33 words, native raw, native normalized, CPU F32, and CPU F64 scores rounded to the common F32 CTC input produced identical word intervals and conditional decisions. Boundary-mass differences were at most 4.06e-8. That result is limited to one clip and does not establish human timestamp accuracy.
The first native trial stopped after warmup because its library allowlist omitted MLX. The corrected immutable plan pins the previously verified MLX libraries; the model, cases, and tolerances are unchanged. The completed trial made 43 protocol calls, including deadline/nonfinite failures and exact same-instance recovery. Original failures and all successful receipts are retained.
The first literal-character timing pass aligned all 136 transcripts, but none passed completely. It included written punctuation as the first or last timed token of a surface word. Of 370 uncertain words, 274 ended in punctuation and 356 had uncertain end boundaries.
The follow-up retains the entire target sequence, including punctuation, spaces, and digits. It selects first/last lexical endpoint tokens inside each unchanged surface word, excluding only an explicit narrow list of edge formatting marks. It does not strip internal punctuation, combining marks, percent signs, currency, or arithmetic symbols. Formatting-only words reject. No transliteration, number guessing, interpolation, threshold change, or acoustic rerun is involved.
The new helper passes 14 focused Rust tests and 16 additional synthetic
score-matrix calls, including contractions, C++, C#, hyphenated words, URLs,
decimals, percentages, currency, combining marks, and formatting-only rejection.
The real replay makes 148 CTC requests for each endpoint variant. The final audit
verifies all target IDs, acoustic score hashes, greedy IDs, path scores, original
words, confidence, and source offsets unchanged. The 3,313 unaffected words have
exactly unchanged boundaries and masses; 327 words use formatting-aware endpoints.
The executable SHA-256 is
6f70bbbc9667c94332f88fb51b32799ad1c1cd3ebf471991455d2baa872c621f.
| Conditional timing measure | Canonical MMS baseline | Omni with lexical endpoints |
|---|---|---|
| Short Hindi clips passing completely | 52/70 | 55/70 |
| Difficult Cap Hindi windows passing completely | 0/12 | 0/12 |
| Uncertain words in those same 610 Cap Hindi words | 77 | 35 |
The short Hindi result contains eight gains and five regressions, not uniform improvement. Omni aligns all 70 short clips; the canonical baseline aligns 69. The newly representable clip must not be silently removed from the denominator. The 40 Indonesian/Finnish short clips pass in 30 cases, and their 14 Cap windows pass in one case. These are alignment diagnostics on frozen recognition, not recognition quality or full-language coverage measurements.
Six deliberately wrong transcripts all align but fail the conditional screen. This is a small negative-control result, not a calibrated false-positive rate. Three error/recovery pairs return no partial alignment content and recover exactly on the same process. A posterior conditioned on supplied text is not confidence that the supplied text is correct.
The comparison preserves full-transcript word correspondence, uses the same matched words for both aligners, and restores original sample offsets before matching the complete Cap recording. Public provider audio identities and cached provider digests are verified through the frozen original corpus plan. There are no new provider calls. These are timestamp disagreements with AssemblyAI, not errors against human timing truth:
| Same-word diagnostic | Canonical MMS | Omni |
|---|---|---|
| Public Hindi, 1,162 common words: mean start/end disagreement | 69.69 / 101.20 ms | 71.63 / 125.65 ms |
| Public Hindi: both boundaries within 80 ms of AssemblyAI | 652/1,162 | 527/1,162 |
| Cap Hindi, 302 common words: mean start/end disagreement | 156.29 / 191.75 ms | 155.83 / 214.99 ms |
| Cap Hindi: both boundaries within 80 ms of AssemblyAI | 156/302 | 141/302 |
The public comparison includes all common words for which both aligners produced diagnostics, including rejected clips; one unmapped MMS clip is excluded from that paired comparison. Restricting to clips where both pass the conditional screen does not reverse the end-boundary result: on 754 common words, mean end disagreement is 94.91 ms for MMS versus 119.82 ms for Omni. All Cap windows still fail, so their intervals are private diagnostics and must not be published as completed transcript words.
The 136 native acoustic calls plus warmup completed. On this M4 Max, acoustic compute p50/p95 is 122.62 / 192.27 ms; separate CTC compute p50/p95 is 29.83 / 106.19 ms. Audio duration p50/p95 is 12.66 / 27.45 seconds. These exclude recognition, upload, queues, persistence, file I/O, and deployment effects; they are not API latency or a speed comparison with AssemblyAI.
Decision: do not promote the Omni timing candidate or replace the canonical baseline. Better text coverage, fast acoustics, and fewer uncertain words have not demonstrated better timestamps. Remaining work includes resolving numerical acceptance, improving long-recording timing without these regressions, independent human timing evidence, language coverage, API parity, and reliability gates. The general 20% Cap beta remains unapproved. Licensing is outside this workstream.
Private receipts: durable1/omni-sentencepiece-eval1, omni-conversion1,
omni-native1, omni-acoustic-eval1, omni-acoustic-eval2,
omni-numerical-diagnosis1, omni-word-eval1, omni-lexical-build1,
omni-lexical-eval1, and omni-word-audit1 under the same durable directory.
Provider-harness schema/provenance corrections and the premature precision-probe
preflight are retained alongside the completed receipts. No serving/API source,
shared database, port 3000, Cap data, provider account, or deployment was changed.
Decision: retain the existing MMS baseline. Neither Omni nor the tested Whisper attention strategies establishes better word timing. The general 20% Cap beta remains unapproved. Licensing is outside the workstream.
The existing DIMEx development split contains 16 clips from four speakers and 175 manually annotated words. Audio, annotation and prior MMS-result hashes were verified; the same correct transcript was supplied to every aligner. This tests alignment given correct text, not recognition or complete API accuracy. The 48-clip Spanish validation split and 40-clip Hindi recognition holdout were not used by these experiments.
| Aligner, on the same 175 development words | Start MAE | End MAE | Both boundaries within 80 ms |
|---|---|---|---|
| Archived MMS baseline | 36.66 ms | 37.21 ms | 148/175 |
| Omni, unchanged lexical endpoint strategy | 46.83 ms | 45.16 ms | 143/175 |
| Turbo attention, canonical tokens and corrected prefix, strict path | 59.63 ms | 49.04 ms | 127/175 |
| Large-v3 attention, canonical tokens and corrected prefix, strict path | 119.01 ms | 118.02 ms | 34/175 |
Omni starts are late by 46.43 ms on average and ends early by 43.78 ms; MMS is late/early by 35.80/35.21 ms. This is interval compression, not a single shared clock offset. The earlier signed Hindi provider comparison also has late starts and early ends, but those provider boundaries are not human ground truth.
A private C++ extension to the pinned Whisper SDK and a serialized Rust driver run one audio encoder pass and one teacher-forced decoder pass. They keep the supplied transcript and confidence values; they do not run autoregressive ASR. Inputs are bounded to one 160 ms–30 second, mono 16 kHz window and the model's text context. Token IDs must be non-special, in the model vocabulary and decode to the exact supplied UTF-8 text. Word mapping retains original byte/character offsets, combining marks and numeric symbols; ambiguous cross-word tokens fail.
The experiment separates the legacy prefix from a prefix containing the transcription task token. Captured attention is already softmaxed. The corrected recipe renormalizes the retained audio crop rather than applying softmax again, then standardizes across tokens, median-filters seven frames and computes DTW. The strict path consumes a frame for each subtoken. A second path permits zero-duration subtokens, but rejects zero-duration complete words. Neither interpolates durations or pads invalid words into positive intervals.
This is derived from the official Whisper alignment method, but is not claimed to be numerically identical to the complete PyTorch model: normalization accumulates in f64 with the pinned SDK's variance epsilon of 1e-9, and attention comes from the native backend.
An explicit check against the official tokenizer found different token IDs for 22/28 development texts: all 16 Spanish, all three Indonesian and all three Finnish texts. The SDK uses greedy longest-piece matching, whereas the reference uses ranked byte-pair encoding and Unicode splitting. Reconstructing the same text was insufficient to prove tokenizer equivalence. Both SDK alignment-head presets match the official masks exactly.
A second native build accepts frozen reference token IDs and independently checks their model-vocabulary roundtrip. It replays the same data and fixes the tokenizer comparison, but does not improve measured timing. The reference tokenizer is a private Python oracle, not a new serving dependency or a completed native tokenizer implementation. All six Hindi texts already had identical token IDs; all 24 Hindi attention/cost/word/likelihood captures remained exactly unchanged across the two tokenization builds, models and prefixes.
The 28 unique development clips comprise the 16 manual Spanish clips, three public Hindi clips, four Indonesian/Finnish clips and five windows from three Cap recordings. Four configurations replay them: Turbo/Large-v3 with SDK/canonical token IDs. These are reused development recordings, not a new production sample or a 28-recording Cap holdout.
The provider comparison preserves full-recording word correspondence and original Cap window offsets. On 40 identical matched public-Hindi words, corrected Turbo start/end disagreement with cached AssemblyAI is 91.45/113.45 ms, versus 69.50/148.95 ms for MMS. On the 13 matched words in the three selected Cap Hindi windows, Turbo is worse: 1907.69/1346.15 ms versus 766.15/1218.46 ms. The small, difficult subset diagnoses a failure; it is not a population accuracy estimate. No human Hindi timing claim or fresh provider-speed comparison follows from it.
Both native builds pass the same ten focused Rust tests. The four runs completed 336 protocol calls, including 46 expected failure/recovery pairs. These cover expired/cancelled requests, invalid language, silence/nonfinite audio, transcript/word mismatch, size limits and duplicate result IDs. The canonical-ID build additionally rejects negative, special and wrong-text token IDs. Failed requests publish no partial content; subsequent requests reproduce the preceding valid features and intervals exactly. Existing result files survive duplicate-ID attempts. All four owned native processes exited and were checked absent.
An independent NumPy implementation exactly reproduces all 224 captured normalization matrices and 336 development DTW paths. This validates this implementation's declared arithmetic; it does not establish acoustic correctness. Four deliberately mismatched audio/transcript pairs, replayed in four configurations, still produce ordered, positive full-word intervals in strict mode. Monotonic timestamps alone are therefore not an acceptance screen. The vertical variant rejects some of these pairs and some genuine Cap words; it is not a calibrated replacement screen either.
The SDK's CPU and Metal cancellation callbacks are registered for request work. Mel extraction remains a bounded routine with checks before/after, not an interruptible inner computation. These local checks do not establish crash, concurrency or deployment reliability for an integrated service.
Canonical Turbo's local request computation before artifact publication has median/p95 295.37/361.88 ms, including state creation, mel, encoder, decoder and CPU alignment. Audio duration median/p95 is 5.43/25.23 seconds. These exclude recognition, upload, API queues, persistence and response publication. They are functional local measurements, not an AssemblyAI speed comparison. The later canonical Large-v3 run overlapped a CPU analysis process, so its timing is not used for a comparative performance claim.
A separately recorded development diagnostic averaged each MMS and canonical Turbo boundary with fixed equal weights, rounded down to an integer sample. This arithmetic combination lowers start MAE to 30.03 ms, but worsens end MAE to 40.78 ms, start/end p95 and both-boundaries-within-80-ms coverage (142/175 versus 148/175). It is also not promoted; no coefficients or thresholds were fitted and no validation data was opened to rescue it.
No serving model, API source, shared database/schema, port 3000, Cap data or deployment was changed. The existing full-API results above remain historical; these native diagnostic runs do not replace full API verification. Remaining gates include reliable timing on difficult recordings, broader language/quality coverage and the explicit speech-to-text API feature gaps at the top of this file.
Private evidence under whisper-large-20260827/durable1: omni-dimex-develop1,
omni-duration-diagnosis1, teacher-build1, teacher-eval1,
teacher-large-eval1, teacher-provider1, teacher-tokenizer1,
teacher-token-build1, teacher-token-eval1, teacher-token-large-eval1,
teacher-token-invariance1 and teacher-consensus-develop1. Original and effective
driver versions are retained; the first evaluation driver's oversized negative
request was corrected before execution to stay within the line protocol and
actually exercise recoverable request rejection.
Implemented and locally verified; general beta remains unapproved. Licensing is outside this workstream. This closes the custom-spelling feature gap, without claiming recognition, timestamp-quality, all-language or complete API parity.
Four public/authored audio inputs, 76 transcript submissions and 12 derived-resource GETs established the observed provider behavior. Nineteen submissions were rejected by validation; 57 completed. No Cap recording was sent by these new provider probes. The Rust parser/renderer and TypeScript contract match all 76 fixtures; accepted text, word surfaces and intervals match exactly. An independent word-level oracle also matches all 54 punctuation-enabled cases. Confidence merging is checked against the source words, not treated as bitwise evidence across independent ASR executions. The 12 saved sentence/paragraph resources match exactly.
The implementation handles case-insensitive aliases, phrases, longest-match and last-duplicate precedence, and a single non-cascading pass. Internal punctuation remains literal; the observed ASCII punctuation stripping and final-mark behavior are preserved, including append-before-trim for targets with outer whitespace. The API reflects normalized aliases without changing stored idempotency inputs. Spelling runs after acoustic alignment and punctuation presentation in both native backends. Phrase replacements retain the first start, final end, mean confidence, and compatible speaker/channel metadata; no synthetic timing is introduced.
Matching indexes possible phrase lengths with rolling hashes and verifies every candidate byte for byte. Input/output sizes, comparison work and cancellation are bounded. Expansion beyond the transcript limits fails without publishing partial text. A 100,000-word fixture with 180 distinct phrase lengths and a 37,641-byte option body completed unchanged in 572 ms on this Mac. This is postprocessing only, not model or API throughput. The 26 nonempty Cap development fixtures all matched their requested phrase; three empty results remained empty and the known uncertain recording remained an error.
The release build passes 105 default, 152 CTC and 160 MMS Rust tests. The exact app snapshot passes 157 tests with seven conditional skips, 1,724 assertions, typecheck, lint, and an isolated SDK bundle/declaration build. The default Parakeet engine also passed eight actual native calls: English/Portuguese baseline and spelling variants, punctuation on/off, invalid spelling before media access, and recovery.
The full local HTTP replay completed 108 jobs: 104 successes and four expected errors. It includes baseline/custom pairs for 30 reused Cap development recordings, public/authored option fixtures, nonspeech, the known recognition failure, an intentional language-threshold rejection, and recovery. Two jobs used the unchanged AssemblyAI 4.36.7 SDK and Sandchest SDK; both SDKs read identical stored responses. Nineteen malformed spelling requests returned the exact safe validation errors without creating jobs. Idempotent replay kept the same ID; changed rules returned 409. Sentence/paragraph exports preserved every resulting word and acoustic value.
All 108 jobs used one inference attempt. There are exactly 104 unique usage events, each with the measured duration, and none for failed jobs. Unit regressions also verify that an older worker lacking spelling acknowledgement cannot complete or bill the job, and deletion removes private spelling vocabulary while preserving billing history. Release the API and matching native worker together.
| Reused Cap recordings, local upload through persisted result | Completed | p50 | p95 |
|---|---|---|---|
| Baseline options | 29/30 | 1.559 s | 6.111 s |
| Same recordings with custom spelling | 29/30 | 1.309 s | 6.189 s |
Polling is 250 ms and baseline precedes custom for each recording. These numbers do not prove a speedup, sustained capacity, deployed performance, or a new paired latency win over AssemblyAI. Baseline text/timestamps match the prior native checkpoint; custom output matches the expected transformation. No new accuracy holdout was opened, and recognition/alignment defaults were not changed.
The first API-harness attempt stopped before audio inference or HTTP because its 100,000-word stress response exceeded the diagnostic reader's 1 MiB limit. A new version used a separately bounded stress process and completed the replay. A resource-fixture reader initially expected an envelope around already-decoded JSON; its corrected version passed. Both failed harnesses and the initial lint failure are retained; they were not overwritten or counted as successful runs.
Private evidence is under whisper-large-20260827/durable1:
custom-spelling-provider1/2/3, custom-spelling-differential1,
custom-spelling-contract2, custom-spelling-build1, custom-spelling-app2,
custom-spelling-e2e2, and custom-spelling-verification1.json. An independent
read-back verified source/model/SDK fingerprints, usage rows, revoked ephemeral
keys, absent owned processes, vacant private ports, and unchanged HEAD/index.
The runtime lease was released. No shared database, port 3000, Cap data, service
configuration, deployment or beta traffic was changed. Difficult-recording timing,
broader quality/language coverage and the remaining STT API features are still gates.
A small decoder optimization is implemented and locally verified. The new recognition-selection rule is not promoted; general beta remains unapproved.
A private Rust forward algorithm computes complete CTC sequence marginal scores using two rows of working memory. It passes 15 focused CPU tests, including an independent exhaustive-path oracle, agreement with the existing posterior calculation, repeated labels, empty targets, impossible paths, cancellation and input bounds. All 128 archived native alignments are reproduced, and their saved audio matches the development manifest sample for sample. Both hypotheses use the same acoustic matrix, tokenizer and blank. No length normalization or new threshold was fitted. Timing concentration remains conditional on the transcript, not a probability that the words are correct.
The fixed rule chooses the saved Cohere hypothesis only when its marginal exceeds the baseline by more than 1e-8 and the existing acoustic/timing gate accepts it. Unsupported targets retain the baseline. A fresh run of the current full native English route completed all 128 development clips, plus warmup and recovery. The 64 cached AssemblyAI jobs were rechecked against their audio and result hashes and rescored; there were no provider calls or newly opened holdouts.
| Human-reference development set | Current worker WER | Offline guarded selection WER | Cached AssemblyAI WER |
|---|---|---|---|
| AMI, 64 clips / 1,228 words | 11.48% (141 errors) | 10.75% (132 errors) | 18.16% (223 errors) |
| LibriSpeech, 32 clean + 32 other / 1,276 words | 1.72% (22 errors) | 1.80% (23 errors) | Not compared here |
The selection changes 19 clips: nine improve, three worsen and seven tie in word-error count. It does not clear the read-speech regression gate. These are development results from saved hypotheses, not a new combined serving model or end-to-end speed measurement. Neither the selector nor Cohere was enabled.
decode_f32_controlled now tries the existing bounded native decoder before
FFmpeg. Eligibility stays limited to mono 16 kHz PCM16 WAV/FLAC, whose sample
values convert exactly to F32. Float, 24-bit, resampled, multichannel and unusual
containers retain the previous floating-point FFmpeg path. Recognition is not
requantized, and malformed eligible media fails consistently with recognition.
The existing regular-file checks, cancellation, metadata/output bounds and
subprocess isolation remain in force.
Expanded existing Rust tests cover all 65,536 signed PCM16 values, random samples, exact F32 bit comparisons, wider precision/resampling fallback and half-millisecond duration edges. For native input, alignment now shares recognition's exact sample-count ties-to-even duration. This intentionally corrects four demonstrated one-millisecond rounding discrepancies; no samples are added or removed. Quiet float detail and nonfinite-sample rejection remain covered. The release passes 105 default, 152 CTC and 160 MMS Rust tests and scoped formatting checks.
The private decoder probe links the exact previous audio.rs for comparison.
All 158 public/Cap inputs have identical F32 sample hashes and durations before
and after. Eight additional format/duration fixtures pass. With FFmpeg absent
from the process PATH, eligible WAV and FLAC still decode, while float and 24-bit
fixtures correctly require fallback; cancellation and subsequent recovery pass.
| 64 PCM16 WAV development clips | Previous | New |
|---|---|---|
| Alignment decoder p50 | 21.63 ms | 0.084 ms |
| Complete native request p50 | 101.73 ms | 79.18 ms |
The median paired native saving is 22.72 ms. Across the mixed 128 public clips, native p50 is 119.74 versus 118.39 ms, because half are float WAVs and retain the fallback. All 158 native pairs, 320 calls including warmup/recovery, preserve all public transcript fields, word boundaries and alignment choices exactly. The 30 Cap MP3s retain 29 successes and the same known recognition error.
No MP3 speedup is claimed. One alternating native comparison measured Cap p50 1.178 versus 1.156 s and p95 8.254 versus 9.338 s, including the failed case. The unchanged MP3 fallback's timing varies, and this run does not establish whether the worse tail is noise. Repeated latency measurements remain necessary before making a performance or beta-readiness claim for Cap traffic.
The matching release binary processed 32 actual HTTP jobs: two public PCM16 WAVs and all 30 reused Cap MP3 recordings. Upload, authentication, durable queue, inference, PGlite persistence, polling and word exports were exercised. There were 31 completions and one expected recognition error, every job used one attempt, and exactly 31 unique usage events matched measured durations. Failed work was not charged. Idempotent replay returned the same job; invalid language and authentication requests were rejected. Sentences, paragraphs and word metadata match the verified native responses.
The 29 successful Cap jobs had local upload-to-persisted-result p50 1.309 s and p95 8.067 s, with 250 ms polling. This is a functional replay, not a paired AssemblyAI benchmark or a demonstrated improvement over earlier API runs. The app's 246 unchanged source/dependency files match its prior verified snapshot; only two documentation files and this separately tested native file differ. The app's prior test/build results are not counted as fresh executions here.
All owned processes stopped, ephemeral API keys were revoked and the private runtime lease was released. No shared database/schema, port 3000, Cap source/data, provider configuration, deployment or beta traffic was changed. Remaining gates include difficult-recording word timing, wider language/quality evidence, repeated Cap tail-latency checks and the STT feature gaps listed at the top of this file.
Private evidence under whisper-large-20260827/durable1: cohere-selection1/2,
cohere-current1, alignment-decode-build1, alignment-decode-eval1 and
alignment-decode-api1. The first CPU replay compilation failed on an ambiguous
Rust trait method and is retained. Its corrected version passes; the current-worker
driver's pre-execution diagnostic correction is also preserved with both plans.
The decoder comparison was repeated on the same 30 Cap MP3s with separate old/new native processes, one full warm pass and fixed ABBA/BAAB request order. All 184 native calls preserve the expected outputs, including the same known error. There are 60 warm and 120 measured calls, plus startup/recovery checks. The median per-recording candidate/baseline ratio is 0.996. Among completed measured requests, p50 is 1.900 versus 1.955 s and p95 is 8.078 versus 8.590 s. The completed p95 is about 6.3% higher; this does not establish MP3 tail improvement or nonregression. The worst earlier slow case did not invoke the changed alignment decoder at all, but that observation is not a complete explanation of run-to-run variation.
A separate unpromoted native experiment re-scores uncertain MMS words in fixed local audio context, without changing recognition. On 98 existing timing cases, it makes 113 protocol calls and reproduces the original timing diagnostics. The 16 manually timed Spanish development clips change no word bounds: 148/175 words still have both endpoints within 80 ms. Short Hindi conditional acceptance increases from 52/70 to 54/70, while the difficult Cap recording remains at 0/12 complete windows. There is no manual Hindi word-time reference, so the two gained clips are not proof of improved timestamp accuracy. Cancellation/recovery and four wrong-text controls pass without new complete false acceptance.
This experiment is not promoted. Its Hindi baseline check proves equal reported
diagnostics, not score-matrix identity. Review also found that its redundant
target encoding was not compared as a complete Target; no numeric-preservation
claim is strengthened by that prototype. The copied build source hashes are
independently checked at the final checkpoint, while the initial run verified
the binary without rechecking the entire source map. Private evidence:
cap-latency-repeat1, mms-local-build1, and mms-local-eval1.
The manual development references showed MMS cores tending to start late and end early. A single fixed rule was chosen before measuring its output: extend an already accepted word by at most one 20 ms model frame at each end. Neighboring words share only their original gap, at a millisecond-aligned midpoint; added outer digital silence is removed, and the original acoustic core never contracts. Words below the unchanged 0.95 conditional timing screen retain their original bounds. No coefficients or thresholds were adjusted on validation or Cap audio.
This is an empirical boundary envelope, not a new acoustic model, posterior or
calibrated confidence estimate. Recognition text, confidence, model weights,
CTC targets and scores are unchanged. The native diagnostic labels the policy
mms-core-plus-one-frame-envelope-v1; this diagnostic is not added to the public
AssemblyAI-compatible transcript resource.
On the 16 development clips/175 manually annotated words, the fixed diagnostic reduces start/end MAE from 36.66/37.21 to 22.19/25.06 ms and increases both-endpoints within 80 ms from 148 to 156 words. It was then applied once, unchanged, to the existing 48 validation clips from 12 different speakers: 494 words, start/end MAE 36.13/35.93 to 21.40/24.28 ms, and 435 to 466 words within 80 ms at both endpoints. These diagnostic runs supplied correct reference text. The validation split had been used in earlier work; it is not a newly acquired or unseen corpus.
The integrated Rust worker was then compared with the exact preceding decoder release, using actual recognition without supplying reference text. All 101 case pairs preserve text, confidence, language, model and expected status. All non-Spanish word times remain exact. The cases comprise 64 public Spanish clips, 30 reused Cap recordings, three nonspeech controls, two crop cases and separate warmup/recovery requests. There are 205 actual native calls, including two extra startup checks and one response whose initial harness assertion failed before it was saved; 204 response records are retained and 101 complete pairs are verified.
On the same 481 matching validation words, out of 494 manually annotated words, the results are:
| Metric | Previous MMS serving | New envelope | Cached AssemblyAI |
|---|---|---|---|
| Mean start error | 36.09 ms | 21.32 ms | 56.84 ms |
| Mean end error | 36.00 ms | 24.47 ms | 57.03 ms |
| p95 start error | 77.94 ms | 57.94 ms | 128.30 ms |
| p95 end error | 95.53 ms | 76.12 ms | 134.97 ms |
| Both endpoints within 80 ms | 422/481 | 453/481 | 300/481 |
Every matched lexical word is scored, including unchanged fallback intervals. The complete and common-provider denominators are retained separately; AssemblyAI has 486 matching validation words overall. Recognition WER is unchanged at 14/494 (2.83%) versus AssemblyAI's 9/494 (1.82%). Across all 64 clips, it remains 20/669 versus 14/669. Better timestamps do not close that recognition gap. Per-speaker statistics are diagnostic, not an extra fitted acceptance gate; the read-back found no speaker MAE, p95 or 80 ms coverage regression in this set.
The candidate then completed the real local upload → HTTP API → durable queue → native inference → PGlite persistence → polling flow for all 101 cases. 100 jobs completed and the same known Cap case failed, so Cap coverage remains 29/30. Every job used one attempt; the 100 completed jobs produced exactly 100 unique usage events with native-measured durations, and the failed job produced none. Public word/text/language/model fields and sentence/paragraph exports match native output, with the requested crop offset applied exactly once by the application. SRT/VTT, word search, listing, authentication rejection and idempotent replay pass. All ephemeral keys are revoked and both API-run processes exit normally and are independently checked absent. No shared database or external billing service is used.
The policy adds bounded CPU postprocessing and no extra model pass. A single paired native run measures Spanish p50/p95 at 491.70/566.99 ms before and 500.33/558.48 ms after. The 29 completed Cap inputs measure 1.282/9.273 s before and 1.250/9.472 s after. These mixed changes do not establish a speed improvement or tail nonregression. Candidate local API p50/p95 is 0.787/0.796 s for the 64 Spanish clips and 1.567/6.748 s for the 29 successful Cap clips, with 250 ms polling. There is no new remote-provider or deployed-capacity comparison.
The serving helper preserves the existing finite-input and 35-second window contract. If sample-valid cores cannot be separated on the public millisecond grid, it returns the original whole-window candidates for the existing selector. Cancellation, allocation and invariant failures remain errors, never partial success or quality fallback. Full target equality is checked before selection; word boundaries receive the window offset only once. The release passes 105 default, 152 CTC and 168 MMS Rust tests. The earlier standalone diagnostic passes seven tests and 84 protocol calls across development and validation.
The prior app snapshot has 244 unchanged source/dependency files after excluding two documentation files and three separately tested existing native files. The new helper is also covered by the native snapshot. Prior app suite results are not counted as freshly rerun tests. All 21 transitive harness helpers are frozen before the API phase; the initial native plan had omitted some helper hashes. The final audit explicitly reconciles the original request plan, every revision, audio/settings bindings, original plan hash, model/build identity and result hashes.
Harness failures are retained: a binary-stdin encoding error occurred before the first standalone request; the native crop check first used API offsets, then compared intentional source-range metadata with a physically cut clip. Neither changed the worker or timing rule. The four native processes from those stopped phases were intentionally killed and reaped; they are not reported as graceful exits. The corrected read-back verifies all retained pairs. An API setup attempt also stopped before startup because its new private output path lacked a lease; the owned path was then claimed and the complete API run passed.
This improvement is retained in the opt-in Spanish MMS implementation. General 20% Cap beta approval remains open: the known Cap recognition failure, wider language/quality evidence, difficult Hindi timing and the STT API feature gaps listed at the top still need work. Licensing review is outside this workstream, as requested. No production checks, deployment, Cap mutations, shared schema changes or control of port 3000 occurred.
Private evidence under whisper-large-20260827/durable1: mms-envelope-build1,
mms-envelope-develop1/2, mms-envelope-validation1, mms-envelope-serving1,
mms-envelope-serving-eval1, api-runtime/mms-envelope-api1, and
mms-envelope-verification1.json. The verified release SHA256 is
f83e7b4fc4fae4ce6b6d7af8a5b95fa4f323b4b1c1e3a820cd9ebce2e9750d5a.
The existing pinned 8-bit Cohere model was evaluated with explicit Spanish prompts in a private Rust entry. No serving source, model default, API, Cap data, provider configuration or deployed service changed. Licensing was out of scope. The candidate failed the frozen development accuracy gate; the 48-clip Spanish validation split was not evaluated or used to tune this candidate.
The model, encoder, decoder, frontend and tokenizer logic are unchanged from the
preserved native implementation. The private build only relocates identical
weight-header/test fixtures and adds bounded WAV input, exact file and PCM
hashes, cooperative cancellation and strict end-of-sequence enforcement. A
decoder that exhausts its token budget returns an error without partial text.
The release build passes two entry tests and all 23 library tests. The binary is
25339f195268d6a8ad999be0b9574b24efa8fc971cc31b02dcd7a2cbcae15a61;
the pinned model is
aadaf8d3388975853385400c9bd8dee92a71e12a83860f35bebaac345aa8af93.
Prompt verification uses the preserved reference processor's actual prompt method, isolated through its AST without loading a model. An independent parser reads all 16,384 tokens from the serialized SentencePiece model and checks them against the JSON vocabulary. All 28 native prompts match this oracle: 14 explicit languages, each with punctuation enabled and disabled. This proves prompt construction, not transcription quality in all 14 languages.
Sixteen existing Spanish development recordings from four speakers contain 175 manually referenced words. Each original recording is converted once with the serving decoder's FFmpeg settings to mono 16 kHz PCM16; both recognizers receive the same WAV. Fresh current-worker recognition matches the cached original-file text and word text for all 16 cases. The original WER normalizer and references are unchanged, including the spelling/acronym discrepancies seen in development.
| Development recognizer | Word errors / reference words | WER |
|---|---|---|
| Current Rust worker | 6 / 175 | 3.43% |
| Native Cohere candidate | 7 / 175 | 4.00% |
| Cached AssemblyAI Pro | 5 / 175 | 2.86% |
Four existing English development clips retain the exact cached Cohere text and token IDs. The raw model produces 13 lexical words on silence, one on a tone and one on noise; current Sandchest returns no words for all three. The candidate has no speech detector, word timestamps, recognition-confidence output, automatic language detection or full-recording composition in this entry. It cannot be substituted into serving based on its recognition-only runtime.
There are 62 actual native requests: 42 Cohere requests and 20 current-worker requests, plus the separate prompt-oracle invocation. Nine failure/recovery pairs cover wrong hashes, corrupt WAVs, unsupported/automatic language, zero deadline, immediate and encoder cancellation, token exhaustion and unknown fields. Every failure exposes no partial transcript; the next request on the same process returns the exact warm-up content. Both inference processes exit zero and are independently absent. Private runtime ownership is released.
Single-pass Cohere p50/p95 is 88.27/113.58 ms, but it measures raw recognition. Current-worker 378/411 ms includes speech detection and word alignment. These different stages are not a service-speed comparison, an API benchmark or evidence of parity with AssemblyAI latency. No new provider request was made.
The independent read-back binds every cached native result to the original audio hash and settings, checks the provider's identity/settings, Spanish language, punctuation, formatting, disfluencies and actual selected model, and recomputes all development scores. These checks close two metadata assertions omitted from the initial harness without changing its frozen plan or rerunning inference.
Failed attempts are preserved: build1 lacked a relocated weight-header include; build2 had a JSON macro syntax error; build3 passed the entry tests but lacked a relocated library-test fixture. Build4 passed. The first preparation stopped on the existing English float32 fixture format before model startup; preparation2 accepts both supported WAV encodings. None of these failed attempts contributed recognition results or were counted as successful runs.
Private evidence: durable1/cohere-spanish-build1 through build4,
cohere-spanish-develop1/2, and cohere-spanish-verification1.json under
whisper-large-20260827. The verified Spanish envelope serving checkpoint is
retained. This negative pilot does not close the remaining beta gates.
This private Rust experiment packs the existing F32 query, key and value weights into one evaluated matrix and bias, replacing three attention projections with one per layer. Q/K/V order, query scaling, model precision, waveform normalization, CTC targets, posterior computation and timing thresholds are unchanged. Original pinned tensors are validated before packing, and source arrays are released only after the packed weights have been materialized. No serving source changed.
The immutable release build passes two entry tests and 53 focused alignment tests; 118 unrelated library tests are filtered, not counted as newly tested. Two alternating paired repetitions cover 98 existing cases: 16 manually timed Spanish development clips, 60 Hindi development clips, 10 Hindi regression clips and 12 windows from the difficult Cap Hindi recording. One Hindi target remains unsupported; 97 cases align. All baseline score hashes and complete target reports match the preserved earlier F32 evaluation exactly. There are 214 actual protocol requests including warm-up, incorrect text and six failure/recovery pairs.
The gate was frozen before inference: maximum absolute log-probability drift at most 0.001, RMSE at most 0.0001, exact intervals and timing decisions, and at least 5% median per-case acoustic speed improvement without more than 2% p95 regression. The candidate fails both numerical and speed requirements:
| Metric | Observed | Frozen requirement |
|---|---|---|
| Maximum absolute log-probability change | 0.0020533 | <= 0.001 |
| Maximum RMSE across inputs | 0.0001444 | <= 0.0001 |
| Median per-case packed/baseline acoustic time | 0.9959 | <= 0.95 |
| Changed word intervals | 0 | 0 |
| Changed timing acceptance cases | 0 | 0 |
The Spanish subset is faster (median ratio 0.925), but the difficult Cap Hindi windows are slightly slower (1.003). Selecting the favorable subset after seeing results would not satisfy the overall gate. Pooled acoustic p50/p95 is 80.17/155.05 ms before and 78.12/153.23 ms after; these are model-stage timings, not end-to-end API latency. Stored active weight memory is effectively unchanged. The 175 Spanish manually timed development intervals remain exact. This private snapshot evaluates original MMS cores, not the newer serving boundary envelope; it is not evidence that serving timing has regressed.
Deadline, hash, language, nonfinite audio, silence and cancellation after the first model all return errors without partial output; each following request recovers exactly. Both repeated orders preserve output stability. The failed first build stopped at compilation because the diagnostic modules were private; its logs are retained. The second build passes, its inference process exits zero, and private runtime ownership is released. The independent verification binds the tested executable to Cargo's artifact record and rereads the frozen inputs, results, numerical changes and timing decisions. The candidate is rejected and the verified Spanish envelope serving release is retained.
A source audit of the released Multilingual Word Aligner (MWA) does not establish a usable
drop-in replacement. Its current sentence interface omits Hindi and Spanish, and
its CSV writer constructs each word's start from the previous word's end, rather
than producing independent starts and ends separated by pauses. Isolated NumPy
batching fixtures also reproduce ragged-batch failures near sequence boundaries
and an extra masked batch at exact sequence multiples. These are observations of
pinned source revision ed0ac8a5f0f8ba9370822737e8c4412cc47a1b5d, not a model
accuracy comparison. No MWA weights were downloaded or run, and these source
limitations do not prove its acoustic model is inaccurate.
Read-only reference research identified manually segmented Hindi data described in the Hindi Au-ToBI paper, the TIFR Hindi database paper and CMU's Hindi synthesis corpus description. No currently working public download of those human word/phone annotations was confirmed. This is an unresolved data-access gap, not a claim that no suitable corpus exists. Automatically forced-aligned datasets and AssemblyAI agreement must not be called human word-timing truth. No new provider calls or Cap data mutations occurred. Licensing remains out of scope, as requested.
Private evidence under whisper-large-20260827/durable1: mms-qkv-build1/2,
mms-qkv-eval1, and mms-qkv-verification1. The QKV result and source audit do not
close recognition, Hindi timing, full API parity or the general 20% beta gates.
Seventeen bounded live AssemblyAI submissions on one existing public audio
fixture tested defaults, individual null fields and combinations. Eleven jobs
completed and six requests were rejected. The provider accepts null for
speaker_labels, speakers_expected, multichannel, word_boost and
boost_param, with the same inactive settings and recognized text as the explicit
baseline. Provider word probabilities/times varied across submissions; this was
not an acoustic equality or accuracy experiment.
Sandchest previously rejected those five nulls with HTTP 400. The request schema now removes them as omitted values before storing options or computing an idempotency fingerprint. Explicit booleans, counts, vocabulary and boost settings are not discarded: unsupported feature requests still fail before creating work. Other types and unknown fields remain invalid. No handler, billing, database, queue, SDK, recognition or native timing implementation changed.
Null handling is deliberately limited to observed successful cases. Null punctuation and formatting values produced upstream HTTP 500, and null disfluencies produced HTTP 400; Sandchest keeps rejecting those invalid inputs. This does not claim exact provider error-status parity for those cases. Omitted filler handling is also still a documented compatibility gap: AssemblyAI defaults to removing fillers, while current Sandchest preserves them. Accepting or rejecting nulls does not implement genuine filler removal.
Two new regression tests cover individual/combined nulls, strict types and explicit
false values, both /v2/transcript and /api/v1/transcripts, exact stored options,
cross-form idempotency, mismatched-request conflicts and rejection of real
unsupported feature requests. Both tests first failed on the unchanged schema.
The initial green attempts exposed two mistakes in the new test itself: stored
options also retain audio_url, and the native endpoint is /api/v1/transcripts.
Those assertions were corrected without changing the schema fix or earlier
receipts. The final isolated 250-file app snapshot passes 159 tests, seven
conditional skips and 1,780 assertions, plus typecheck, lint and changed-source
checks. The model-dependent skips are not counted as native verification.
The exact retained Spanish-envelope release then served 74 new local jobs through real TCP upload, the unchanged API handler, durable queue, native inference, PGlite persistence and polling. The input set contains 30 reused Cap recordings, three nonspeech controls, two crop cases, warm-up and recovery, each submitted with omitted options and with all five nulls. Pair order alternates. It is not new quality data or a latency benchmark; no reference text is supplied to recognition.
All 37 pairs preserve recognized text, word intervals/confidence, language, model, status and error content. They also match the previously bound native results, including crop offsets applied exactly once. Sentence and paragraph exports preserve the same words; subtitle, listing and authentication checks also pass. There are 72 completions and two instances of the same known Cap error, so coverage remains 29/30 unique Cap recordings. All jobs use one inference attempt.
Each job is retried once under its existing idempotency key using the opposite null/omitted form. All 74 retries return the original ID. Read-back confirms only 74 stored jobs and exactly 72 unique measured-duration usage events, with none for the two errors. Five actual unsupported feature requests remain rejected and create no jobs. This uses isolated PGlite and no live billing service; it does not reverify shared Postgres or Autumn. Both owned runtime processes exit zero, all ephemeral keys are revoked, and private ports/ownership are released.
The independent audit rereads the frozen source, provider requests/results, all 74 local results and database evidence. It checks all 37 pairs, cached-native bindings, usage counts and fifteen app/native child processes absent. No Cap mutation, shared database migration, port 3000 control, deployment or production capacity check occurred. Licensing remains outside this workstream.
Private evidence under whisper-large-20260827/durable1: api-nulls-provider1,
api-nulls-fix1, api-nulls-app-red1, api-nulls-app-green1/2/3,
api-nulls-e2e1 and api-runtime/api-nulls1. The new native/API read-back is
api-nulls-e2e1/independent-readback.json. General 20% Cap beta approval remains
open for the quality, language/timing, filler-default and STT feature gaps above.
This experiment tests a different acoustic alignment family after the Hindi MMS timing coverage stopped improving. It does not change the native serving model, API, database, billing or queue. The candidate has declared Hindi and Spanish presets, but declaration is not measured quality or AssemblyAI language parity. Licensing was not evaluated and was not a gate.
The reference is pinned to Bournemouth Forced Aligner
and CUPE-2i model revision.
Both checkpoint downloads match the revision's published LFS SHA-256 and length.
The reference loads them with weights_only=True and strict state matching;
the inference tensors are exported unchanged as F32. Only checked integer
training counters are excluded. No unsafe pickle fallback is used.
The private Rust port implements the convolution, normalization, channel attention, temporal/spectral branches, transformer and phoneme/group heads. Its loader checks the complete tensor header, shapes, finite F32 values and file hash. An early attempt correctly failed startup because the existing Parakeet weight loader converts tensors to BF16; the final candidate uses a separate F32 loader. That failed run is preserved and is not counted as a successful inference test.
Two model presets each use eight frozen window fixtures: silence, impulses, tones, seeded noise at two batch sizes, and existing Spanish/Hindi/Cap audio windows. The 16 model/fixture combinations run twice. Across all 32 comparisons, maximum absolute logit difference is 0.0000743866, maximum RMSE is 0.00000929757, and there are zero argmax differences. This passes the predefined 0.001 absolute/0.0001 RMSE/zero-argmax gate. Repeated outputs are exact.
There are 98 actual native protocol requests, including 32 failure/recovery
pairs. Tests cover expired deadlines, cancellation at encoder/attention/last-layer
stages, invalid stages, hashes, language, batch sizes, shape, nonfinite and excessive
amplitude samples, unknown fields and malformed/oversized framing. Rejections
contain no partial scores, and the next valid request recovers exactly. Four
focused Rust tests and the release build pass. The verified acoustic binary is
63e16a60665207a4eedfcd348eac6bc8ecb1038f1fdce8ede554d41a802ed418.
These are acoustic and runtime checks, not proof of accurate words or timestamps.
An audit of the exact upstream window helper found two arithmetic problems in eight synthetic CPU fixtures: it hardcodes a half-window frame step even when the configured overlap differs, and cosine-edge normalization does not preserve a constant prediction. A single window also returns five stitched frames while a separate helper reports ten. These findings do not establish a particular timing drift in the full published pipeline, which also rescales by the clip length.
The offline feasibility harness therefore defines its own explicit clock: 1,920-sample windows, 1,152-sample steps, ten 192-sample output bins per window, positive centered-sine weights, and a clipped final bin. This 12 ms bin rule is an experimental interpretation, not verified provenance of the model's training labels. Thirteen boundary lengths verify coverage and constant/linear-signal composition. A strict CTC reference matches exhaustive path enumeration in 24 small cases, including repeated phones and impossible paths. Invalid waveform inputs are rejected.
eSpeak NG 1.52 was built in private storage, with no system installation or audio playback. Phonemization preserves each supplied word; unknown pronunciations fail instead of silently dropping phones. The acoustic work remains native Rust. The complete clock/CTC prototype and the published decoder-component comparisons run offline in Python; neither is a Python serving fallback or a completed native word aligner.
All 16 existing Spanish development clips, four speakers and 175 manually timed words are included with correct supplied text. A further 37 native requests produce 926 windows and repeat the first batch exactly. The same frozen scores are then evaluated with three predefined component variants, without another model run or a parameter sweep:
| Alignment on the same supplied words | Start MAE | End MAE | Both boundaries within 80 ms |
|---|---|---|---|
| Retained MMS envelope | 22.19 ms | 25.06 ms | 156/175 |
| CUPE, strict reference CTC | 32.98 ms | 86.74 ms | 99/175 |
| Reference CTC plus published soft boundaries | 35.77 ms | 76.83 ms | 111/175 |
| Published decoder cores | 46.17 ms | 73.13 ms | 110/175 |
| Published decoder plus soft boundaries | 49.82 ms | 68.77 ms | 114/175 |
Every variant fails the fixed development gate. Published decoder components use their default boosting/probability floor and boundary softness; they still use the experimental clock and word pronunciations above. This is not a reproduction of the full original BFA pipeline, nor a claim that the model cannot work elsewhere. No variant is promoted. The 48-clip Spanish validation set, Hindi alignment set and remaining Hindi recognition holdout are not opened for this comparison. There is no new recognition WER, provider comparison, API speed or capacity claim.
The native processes exit zero, and private runtime ownership is released. Independent read-back checks the frozen source/weights, numeric and recovery results, complete-clip score composition and timing summaries. No new Cap fetch, provider transcription, shared service, migration, deployment or production check is performed. The existing app/serving checkpoint remains unchanged.
Private evidence under whisper-large-20260827/durable1: bfa-source1,
bfa-reference1, bfa-build1/2/3/4, bfa-acoustic-eval1/2,
bfa-espeak1/2/3, bfa-full1, bfa-decoder1 and bfa-verification1.
Earlier unsuccessful build/download preparation attempts remain preserved.
General 20% Cap beta approval is still open for the previously documented
recognition, Hindi timing, filler-default and full STT API gaps.
The public API now accepts disfluencies: false. New jobs store that effective
default explicitly; true preserves verbatim model output as before. This follows
the documented provider default, verified
with 24 new AssemblyAI submissions on eight existing Cap development recordings:
four English clips and Spanish, Italian, Portuguese and Polish controls. All 24
complete. False and omitted have the same normalized words in all eight cases.
The English pairs remove 34 filler tokens without another normalized lexical
change; the four non-English control clips have no lexical change.
Provider behavior is not perfect ground truth. Its false results retain three
forms treated as fillers by the defined English cleanup, and word confidence and
some boundaries vary between submissions. The first presentation harness assumed
exact provider word equality and stopped on one retained um. That failed run is
preserved. The completed comparison requires identical nonfiller sequences and
reports the remaining filler/punctuation differences; it does not call these
fixtures human annotations or claim exact provider timing/confidence equality.
The Rust cleanup runs after recognition and acoustic alignment, before punctuation removal or custom spelling. It removes defined English filled-pause forms without changing any retained word's text, timestamp, confidence, speaker or channel. Quoted literals, uppercase letter/acronym forms and numeric metre units are protected, including a unit separated from its number by a pause. It validates the text-to-word mapping, bounds allocations and checks cancellation before publishing. A transcript consisting solely of removed fillers can intentionally become empty, with null aggregate confidence; the raw recognition uncertainty check still runs before this transformation. Aggregate confidence is the mean of the retained words, not a new acoustic estimate.
Explicit true and the internal legacy omitted option return the original
allocations. Other languages retain their model output: English um must not
erase a Portuguese number or German preposition. This is an English lexical
cleanup rule, not a multilingual semantic disfluency classifier. It does not
remove discourse words, repetitions or false starts. The four non-English
controls do not establish all-language filler behavior, and that quality gap
remains open. There is no additional encoder/decoder/alignment model pass.
Seven focused Rust tests cover these behaviors. The default, CTC and MMS suites
pass 112, 159 and 175 tests respectively, and the isolated release builds.
The completed native comparison makes 74 actual ASR requests on 37 existing
inputs: 30 Cap recordings, three nonspeech controls, two crops and warm-up/recovery.
Every explicit-true result matches the preceding serving checkpoint. False
removes 30 recognized tokens (15 um, 15 uh), with an exact ordered subsequence
of the original word objects. Language, model, duration and all retained acoustic
metadata remain unchanged. Another 28 native presentation requests cover the
eight provider fixtures and two rejection/recovery pairs.
An earlier partial native run stopped after 16 ASR requests because the harness
used Python's compensated sum for aggregate confidence while Rust uses a
sequential F64 sum. The difference was about 2e-16. The reference was corrected
to match the native arithmetic; no numeric tolerance, word metadata or model
output was loosened. Both failed attempts and their process exits are preserved.
This is a compatibility improvement, not new recognition WER or a speedup claim.
New omitted and explicit-false requests share an idempotent interpretation before the existing credit gate. Previously accepted jobs with an omitted option retain their old preserving behavior; retries do not rewrite or rerun them. A changed true/false request under an existing key conflicts. No migration is needed. The worker must return an explicit cleanup acknowledgement for false jobs. An older worker that ignores the option produces a permanent, unbilled failure. Upgrade native inference before enabling these application defaults in any future rollout; deployment is not part of this checkpoint.
Two new API/worker regression tests first fail on the previous implementation,
then pass with the new behavior. They cover both API route families, persisted
defaults, legacy replay, changed-setting conflicts, invalid acknowledgements,
retained words and deliberate empty completion with exact duration accounting.
The Sandchest SDK now types disfluencies as a boolean, and its existing upload
cache-expiry test verifies an explicit-false request replayed as omitted after a
new upload. This SDK source change alters only a TypeScript type, not emitted
runtime code. The existing compile-time assertion that fillers could not be
disabled was updated to assert the new supported capability. The final app
snapshot passes 161 tests and seven conditional
skips, whole-app typecheck/lint, SDK bundle/declaration builds and scoped checks.
The actual HTTP upload → queue → Rust → PGlite → polling run creates 111 jobs, with true/false/omitted variants for all 37 inputs. 108 complete; three are the same previously known uncertain-input error. All 37 false/omitted pairs match text, words, confidence, language, model and status; every result matches its bound native output, including crop offsets. Sentence/paragraph words, subtitles, listing and authentication checks pass. All 111 idempotent retries return the original ID, and all 111 changed-setting retries return 409 without another job. Read-back finds exactly 108 unique measured-duration usage events, one inference attempt per job, and no usage for errors. This is isolated PGlite, not another live Postgres/Autumn verification.
The cached AssemblyAI result for the one uncertain Cap recording completes with zero words on the same audio hash. That is an unresolved status difference, not demonstrated missing speech or human proof of silence. The conservative native uncertainty safeguard is unchanged. All owned native/API processes exit zero in the completed run, ephemeral keys are revoked and private runtime ownership is released. No Cap mutation, shared database, port 3000 control, deployment or production capacity check occurs.
Private evidence under whisper-large-20260827/durable1:
disfluencies-provider1, disfluencies-build1/2, disfluencies-native1/2/3,
disfluencies-app-red1, disfluencies-app-green1/2/3, disfluencies-e2e1,
api-runtime/disfluencies1 and disfluencies-verification1. General 20% beta
approval remains open for recognition, Hindi word timing, multilingual filler
behavior and the remaining STT API features. Licensing remains out of scope.
The new sample passes local API/persistence and Cap-consumer checks. It does not approve a general beta or establish all-language recognition/timestamp parity. No serving source, model, API implementation, Cap configuration or deployment changed in this expansion. Licensing remains out of scope.
A frozen selection contains 100 recordings from 96 owners, excluding the prior 242 prior video identifiers and 168 owners. It balances web/desktop MP4 sources and four duration ranges from eight seconds to fifteen minutes. These are public, completed recordings in Cap's standard storage; this is not a random sample of all traffic or a live-chunk test. Read-only metadata and media reads collected all 50 development recordings from 48 owners, totaling two hours. Their bytes and decoded PCM hashes are verified and unique. The other 50 recordings remain uncollected, with owners separate from development and all earlier samples.
The unchanged release d2fa00f8a6dbce9f297178fe9128a0c8de0af603f731f3eae6e35e2ac476270a
ran through real local TCP upload, API, durable queue processing, Rust inference,
PGlite persistence and result polling. Each recording was also sent to AssemblyAI
with identical MP3 bytes and Cap's current options, including its 25-language
hint list and disfluencies: true. Provider order follows a fixed serial ABBA
schedule; both use 250 ms polling. There are no ambiguous POST retries.
| Warm, single-pass development measurement | Local Sandchest | Remote AssemblyAI |
|---|---|---|
| Completed recordings | 50 / 50 | 48 / 50 |
| Paired end-to-end p50, 48 mutually completed | 1.564 s | 6.614 s |
| Paired end-to-end p95 | 8.048 s | 15.192 s |
| Paired end-to-end p99 | 19.550 s | 22.391 s |
AssemblyAI's two errors explicitly report no spoken audio; Sandchest returns empty completed transcripts for both. That is a status-semantics difference, not evidence that Sandchest recovered more speech or is more reliable. It remains part of the exact-compatibility gap. These timings compare a local Mac API with a remote cloud service, including their different network paths. They are not matched-hardware, deployed-load or sustained-tail measurements.
Including warmup, recovery and three nonspeech controls, all 55 local jobs have one attempt and exactly one duration-based usage event. All 55 idempotent replays return the original ID; changed filler settings conflict without another job or charge. Cap's unchanged pure edit-transcript and caption helpers consume all 15,570 returned words, preserve text, timestamps and confidence, round-trip the stored representation, and produce 3,428 valid caption cues. This does not execute Cap's production workflow or establish live-chunk correctness.
The provider-labelled English group has 908 lexical differences against 11,225 provider reference tokens (8.09%). This is provider disagreement, not human WER. Its 10,266 matched words have mean start/end disagreement of 59/91 ms and p95 of 167/234 ms. Other language groups contain larger disagreements; neither provider's output establishes which transcription or boundary is correct.
The one Chinese clip has 22 character differences against 32 normalized reference characters. Its whitespace-based lexical score is retained but is not a suitable Chinese quality metric. The tiny sample, orthographic differences, differing word segmentation and lack of human references prevent a general Chinese conclusion. Several non-English boundary comparisons have long tails. The earlier Spanish manual timing improvement cannot be generalized to these Cap recordings.
Three Cap-hinted requests return English response labels with very low confidence from Sandchest, versus Hindi/Hindi/Russian from AssemblyAI. The actual native recognition languages are Indonesian/Indonesian/Ukrainian, which are outside Cap's 25 hints. A fixed 17-call native diagnostic reproduces the API outputs exactly and tests opening/middle/ending 30-second windows: all nine windows retain those native language selections. Forcing the original provider labels does not resolve the large text disagreement.
Six additional complete API pairs reuse these three files with unrestricted automatic detection or an explicit Indonesian/Ukrainian code. Both services now agree on the language in all six pairs. Within each provider, the automatic and explicit requests have identical text and word objects. Sandchest's text, words and confidence are also identical to its original Cap-hinted responses; only its reported language metadata changes. The three unrestricted lexical disagreements are 11/130 (8.46%), 51/502 (10.16%) and 81/231 (35.06%). These are still provider comparisons, not human-reference accuracy wins.
This supports separating a restrictive caller hint list from recognition quality.
AssemblyAI documents that expected-language lists restrict detection, while an
omitted list permits all supported languages.
Automatic language detection,
supported languages.
Neither Cap's hints nor Sandchest's fallback implementation was changed here.
Six further provider-only controls keep Cap's hints and explicitly set fallback
auto or en; both still select the original Hindi/Russian labels. They do not
distinguish the provider's default fallback algorithm. The existing inferred
English fallback remains an explicit compatibility risk, not a resolved behavior.
This expansion contains 50 new recordings, not 62: subsequent requests reuse three of them. Totals are 66 actual local API jobs/66 unique usage events, 62 provider requests, and 17 standalone native diagnostic calls. Isolated runtime processes exit zero; ephemeral API keys are revoked; private ports and the runtime lease are released. Source/input/result and persistence read-backs are independent of the request driver. HEAD and the shared index remain unchanged. No shared Postgres, port 3000, Cap write, queue deployment or production check was performed.
Private evidence under whisper-large-20260827/durable1: cap-expansion1,
cap-expansion-control1, cap-expansion-api1, cap-expansion-consumer1,
cap-expansion-verification1, cap-expansion-language1, cap-expansion-scope1,
cap-expansion-fallback1, and cap-expansion-scope-verification1. One preparation
check initially compared two floating-point duration expressions; all 50 PCM
hashes were exact. The original failed driver and the exact-arithmetic correction
are preserved. No audio, model output or acceptance tolerance was changed.
Keep the untouched 50-recording validation set closed while addressing remaining quality, difficult-word timing and STT contract gaps. A language list and successful requests alone do not make a 20% beta safe. No candidate or routing policy was promoted by this evaluation.
The serving implementation is unchanged. Neither multilingual candidate is promoted, and this work does not approve a general beta. Licensing stays out of scope. These are further timing diagnostics on seven recordings already in the fresh Cap development sample, not seven additional recordings. The remaining 50-recording validation selection is still uncollected.
The first private native candidate applies the existing MMS weights, posterior screen, boundary envelope and interval selector to three French, two German and one Portuguese recording. The Spanish recording and warmup check the existing path. It receives frozen recognition words; it does not rerun recognition or change text, confidence, speaker or channel metadata. Native probes return only new bounds, preserving the original serialized word metadata outside the probe.
It passes 175 library tests, three probe tests, deterministic repeats, ten rejection/recovery pairs and a silence control. However, it fails the preset requirement of at least 20% lower mean start and end disagreement for both French and the combined new-language group, without worse p95. AssemblyAI agreement is a development diagnostic, not human timing truth.
A CPU diagnostic identifies 40 unique unsupported tokens in the two longer French recordings: 21 numeric/unsupported-character cases and 19 tokens with no lexical target. Those tokens cause 19 whole windows, covering 836 core words, to skip acoustic alignment. The second private candidate splits only rejected windows into contiguous supported sections. Unsupported words and two neighbors on each side keep baseline bounds; supported context is never clipped through an unsupported word's baseline interval. Sections require at least three candidate words and are capped at eight per original window. Recognized words are never deleted or replaced with guessed number pronunciations.
This candidate passes 182 library tests and three probe tests. The paired run contains 48 native calls, with deterministic repeats, seven rejection/recovery pairs and silence recovery. All 40 unsupported words and all 161 words in their surrounding guard regions retain baseline bounds in the recorded cases. The two long French recordings select 428 and 771 words for acoustic timing, compared with 242 and 508 without recovery. Increased coverage is not itself improved accuracy.
| French provider comparison, same 1,707 matched words | Mean start / end difference | p95 start / end difference |
|---|---|---|
| Current serving word bounds | 383 / 199 ms | 1,848 / 904 ms |
| Private MMS without recovery | 314 / 187 ms | 1,671 / 831 ms |
| Private MMS with section recovery | 282 / 190 ms | 1,541 / 804 ms |
The recovery candidate still fails the fixed improvement screen. End timing is slightly worse than the first candidate, despite improving against the current baseline. Combined median alignment-stage times across the seven cases total 10.05 seconds without recovery and 16.54 seconds with recovery; the fixed cost screen passes. These are alignment-only measurements. An unchanged German path also shows substantial timing variation, so this run cannot isolate a causal latency cost or establish a full API performance improvement.
One public 10.085-second MARC-Fr recording has explicitly named ManualAlign and
ManualTokensAlign tiers. Its 53 token intervals all end and start on supplied
manual phone boundaries. The corpus description specifically documents manual
phone alignment; ordinary TokensAlign files elsewhere in the corpus are not
silently treated as manual word references.
Official MARC-Fr description,
revision-one manual annotation.
The diagnostic excludes one non-speech event and joins only adjacent literal apostrophe clitics using their existing outer endpoints, leaving 45 reference tokens. No pronunciation guessing, G2P-generated boundaries or interpolation is used. Current native ASR and one fresh AssemblyAI request receive the same WAV and explicit French options. The timestamp candidates receive only the native ASR words, never the reference text or reference timestamps. Native recognition, repeated alignment and warmup recovery are deterministic in the recorded calls.
| Same 26 common manually referenced words | Mean start / end error | p95 start / end error |
|---|---|---|
| Current native word bounds | 84 / 51 ms | 430 / 132 ms |
| AssemblyAI | 75 / 69 ms | 152 / 248 ms |
| Private MMS | 68 / 36 ms | 430 / 104 ms |
The private MMS candidate improves average errors, but its start-time tail is still worse than AssemblyAI. This sample has no unsupported target windows, so it does not validate section recovery. Its size, matched-word coverage and annotation/tokenization differences prevent any general French quality claim. No thresholds were tuned on this sample. The full match counts and both common and provider-specific matched-word metrics remain in the private report.
The scoped review found an evidence wording ambiguity: the frozen recovery plan's "same posteriors" means the same calculation and selection policy, not identical posterior values. Different section audio/text produces new acoustic scores and new envelopes. An additional independent output comparison finds two changed words outside the original normalization-rejected cores in one French case: the global interval selector can reconsider earlier ordinary-window choices. The unchanged unsupported/guard-word checks do not cover that behavior. This needs a constrained selection change and regression coverage before promotion; the private candidate is retained as a failed experiment, not silently fixed. The mid-alignment cancellation checks also use the supported warm path; interruption inside recovery normalization remains untested.
An independent readback recomputes comparison metrics and verifies frozen inputs, source snapshots, binaries, libraries, word counts, interval order, overlap limits, guard regions, provider receipt identity and output hashes. It covers 87 Cap/control acoustic-alignment calls, four public French ASR calls, four public French alignment calls, ten CPU normalization requests and one new public provider request. All 16 owned build/worker processes are independently absent. The private runtime lease and ports are released. HEAD and the shared index remain unchanged; 889 session-owned fingerprints match.
Evidence under whisper-large-20260827/durable1: mms-languages-build1,
mms-languages-eval1, mms-normalization-build1, mms-normalization-eval1,
mms-recovery-build1, mms-recovery-eval1, marc-fr-manual2,
french-manual-timing1, mms-french-verification1, and mms-french-review1. The initial public
collector's Python CA-bundle failure is preserved in marc-fr-manual1; retry
used system TLS verification without disabling certificate checks.
Only this evidence document and private experiment files changed. The serving
binary remains d2fa00f8a6dbce9f297178fe9128a0c8de0af603f731f3eae6e35e2ac476270a.
No new application tests are claimed: the preceding app/native checkpoint remains
separate from the private-candidate tests above. No Cap writes, shared Postgres
changes, port-3000 control, deployment or production checks were performed.
Further work must resolve difficult-word timing and the remaining STT API
contract gaps; neither extra alignment coverage nor local speed closes them.
Multichannel transcription is implemented and locally verified. This does not approve a general 20% beta or establish complete AssemblyAI feature parity. Licensing remains outside this work's scope.
The Rust worker decodes each physical channel separately, uses the existing recognizer and acoustic timing for each channel, and merges overlapping words without downmixing. One-based channel/speaker labels and utterance ordering follow observed provider responses. Source text spacing is preserved, including unspaced scripts. All channels share a cancellation/deadline; a failed channel produces no partial transcript. The current resource guard is 32 channels, not a verified AssemblyAI maximum. Utterance segmentation uses our word-gap policy and need not match the provider's segment boundaries.
The API, response schema and SDK now support multichannel. The queue validates
channel acknowledgements, per-channel timelines, utterance word coverage and crop
receipts before completion. Crop offsets reach both top-level and utterance words
exactly once. A separate transcript_channels table stores output independently
of idempotency-bound request options. The generated migration is additive and was
applied only to isolated test databases. Existing mono reads do not query this table.
Completion persists the result, channel metadata and one unique usage event in the same fenced transaction. Public duration stays physical duration; billable duration is physical duration times decoded channel count. Failed/partial results remain unbilled. Deletion erases channel content while retaining usage/outbox history. No pricing, external billing service or shared database was changed.
- Ten real provider submissions on public/authored channel fixtures, plus 22 GETs, established mono/stereo/swapped/quiet/mixed-language/cropped behavior. These are contract probes, not Cap recognition ground truth. The provider accepts combined multichannel/diarization; our diarization option remains explicitly unsupported.
- Rust tests: 122 base, 169 CTC, 185 MMS, all passing. The first decoder attempt and an invalid space-delimited test assumption were retained and corrected.
- Ninety actual native calls: ten channel variants repeated, nine independent mono controls, all 50 existing Cap development recordings, and rejection/recovery checks. All 50 Cap outputs retain exact text, words, timing, confidence and language/model metadata versus the prior serving receipts. This is regression proof, not new WER.
- Sixty complete HTTP upload/queue/native/persistence jobs passed: ten channel variants and the same 50 Cap recordings. All completed once, with 60 exact idempotent replays, 60 changed-channel conflicts and 60 unique usage events.
- An adversarial overlap fixture exposed a reversed paragraph interval. The fix uses enclosing word bounds for multichannel paragraphs; word timestamps stay unchanged. Final code passed another 300 HTTP resource checks over a private copy of those 60 stored results, with zero additional inference jobs or mutations of the original test database. This group-envelope rule can differ from provider resource boundaries, which are not always enclosing on overlapping channels.
- Final application checkpoint: 165 tests passing, seven conditional model skips, 1,967 assertions; typecheck, lint and SDK bundle/type builds pass. Conditional skips are not represented as native passes; native/API proof is listed separately.
On these 50 reused Cap recordings, current local API p50/p95/max is 1.541 / 14.266 / 34.578 seconds. Historical local p50/p95 on the same 50 inputs was 1.551 / 7.586 seconds. The median is similar, but the current tail is worse; these were not contemporaneous controlled runs. Standalone native latency also varied substantially from HTTP latency. A controlled old/new serving comparison is required before claiming a performance improvement or non-regression.
The wider recognition/timestamp-quality gaps, diarization and remaining request features/limits documented above are still open. No deployment, shared database migration, normal app/worker startup, Cap source change or beta traffic switch was performed. All owned test processes exited and the private runtime lease was released.
Private evidence: multichannel-native-build2, multichannel-native-eval2,
multichannel-app-green3, multichannel-api1, multichannel-resource-replay1, and
multichannel-verification1 under the existing ignored durable1 experiment root.
The no-spoken-language API case is corrected and verified. The latency screen still fails, and general beta approval remains open.
Seven new provider transcripts on authored digital silence, plus one prequeue
rejection, confirmed the distinction between automatic and explicit language.
Automatic/default detection, fallback-language hints, multichannel detection and
Universal-2-only detection all returned error, zero duration, null content/model/
language metadata, and language_detection cannot be performed on files with no spoken audio. Explicit English returned a completed empty transcript with zero
confidence. Disabling detection without specifying a language was rejected before
queueing. No additional Cap recording was uploaded to the provider for this probe.
After structural validation, the queue worker now treats an empty result without a detected language as a permanent failure when automatic detection was requested. It persists zero duration and no content or usage, without retry amplification. Explicit-language silence remains valid. Successful empty API responses now report zero confidence; the native missing estimate and failed/deleted metadata are not changed. The native model and word-timing code did not change in this correction.
Two new focused tests failed before the fix. Final application checks pass: 167 tests, seven conditional skips, 2,106 assertions, plus typecheck, lint and SDK bundle/type builds. The genuine native/API replay covered 70 jobs: all 50 previously used Cap recordings, ten channel controls, seven authored silence configurations, two silent crops and automatic detection with speech beside a quiet channel. It produced 58 completions and 12 expected no-spoken-language failures, with exactly 58 usage events, zero failed-job usage, 70 exact idempotent replays and one attempt per job. All 45 nonempty Cap transcripts retained their exact words, boundaries, confidence and model/language metadata.
The five empty Cap recordings remain included. Two now match the provider's no-speech error. For three others, AssemblyAI returned empty successes with weak English language-confidence values (0.292–0.360); Sandchest did not detect a language and now returns an unbilled error. Both providers produced zero words on these three recordings. That status difference remains a compatibility gap, not an accuracy win or proof of no speech. No language label or confidence was invented to hide it.
A contemporaneous ABBA diagnostic ran 24 actual API jobs: four reused long Cap recordings and two public warm-ups in each of four isolated worker lifecycles. API code, model/library bytes, Metal selection, request options, polling cadence and inference concurrency were held fixed. All responses and usage were correct. The current/previous geometric-mean latency ratios were 1.252, 1.346, 1.185 and 1.213; the median case ratio was 1.232, triggering the predefined 1.10 screen. These selected recordings do not estimate a full-traffic p95.
Median API time beyond measured inference was 325 ms; inference accounted for 97.8% of measured latency. This points further profiling at inference rather than HTTP/database roundtrips. Substantial within-binary timing variation and changing background host activity were recorded. They do not establish the cause or excuse the failed screen. The slowdown remains unresolved.
A private Q8_0 artifact was generated from the already-pinned Turbo F16 model with the pinned upstream quantizer. It is 874,188,075 bytes versus 1,624,555,275 input bytes. Independent parsing checked all 587 tensor layouts, 233 quantized tensors, unchanged vocabulary/filter bytes, unchanged tensor payloads and exact EOF. A receipt-only wrapper exception was reconciled without rerunning conversion, and its missing import was corrected with the original invocation source preserved.
No Q8 transcription, accuracy, timestamp or latency test has run yet. The serving hash guard/configuration remains unchanged. Smaller files are not evidence of better inference performance. The next experiment must use an isolated worker and fixed development references before any promotion decision.
Private evidence: silence-contract-provider2, silence-contract-app-red1,
silence-contract-app-green1, silence-contract-api1, multichannel-latency1,
turbo-q8-model1, and silence-contract-verification1 under the ignored experiment
root. No held-out corpus, shared database, Cap source, normal runtime or deployment
was changed; all owned processes were reaped and the private runtime lease released.
Q8 has now been tested and is rejected for promotion. The current F16 serving model remains unchanged. General beta approval is still open. This supersedes the untested-artifact status in the preceding checkpoint.
The local-only Q8 worker differs from the current native worker by exactly two artifact size/hash constants. The default downloader metadata and production configuration were not changed. Its base, CTC and MMS Rust suites passed 122/169/185 tests, followed by a frozen development screen of 206 actual native HTTP requests. These produced 198 successful responses and eight deliberate 400/422 rejections, each followed by exact recovery. Both binaries rejected the other model's file at startup. All owned processes were reaped.
The screen used 81 evaluation recordings: 16 Spanish and 60 Hindi development clips with reference transcripts, one small French manual-timing diagnostic, and four previously used long Cap development recordings. Additional requests covered public warm-ups, silence/tone/noise, repeated long recordings and invalid inputs. No new provider requests or held-out audio/reference reads were made. These are native HTTP measurements, not new full public-API/queue/persistence comparisons.
| Reference set | Current F16 errors | Q8 errors | Cached AssemblyAI errors |
|---|---|---|---|
| Spanish: 16 clips, 175 words | 6 (3.43%) | 6 (3.43%) | 5 (2.86%) |
| Hindi: 60 clips, 1,473 words | 479 (32.52%) | 466 (31.64%) | 286 (19.42%) |
| French diagnostic: 1 clip, 45 words | 16 (35.56%) | 16 (35.56%) | 13 (28.89%) |
An independent dynamic-programming edit-distance audit reproduced every aggregate error count. Failures would remain in the denominator as empty hypotheses; none occurred on these reference clips. These development results do not establish all-language quality or statistical equivalence. The Hindi gap remains substantial; the single French clip cannot estimate general French accuracy.
On the same 169 manually annotated Spanish words matched by all three outputs, current F16 and Q8 both had 22.83 ms mean absolute boundary error and 80.24 ms p95, versus 53.14 ms and 130.07 ms for cached AssemblyAI. Both boundaries were within 80 ms on 152/169 words for Sandchest and 112/169 for AssemblyAI. On the 26 common French diagnostic words, F16/Q8/AssemblyAI mean errors were 67.58/67.20/71.97 ms; this small diagnostic does not close the French timing gate. Coverage is reported separately so omitted or mismatched words cannot improve the timing score.
The four long Cap cases ran in F16/Q8/Q8/F16 process order, with opposing case order and identical model settings, libraries, device and warm-ups. Q8/F16 geometric-mean native HTTP latency ratios were 1.355 (French Whisper), 1.284 (English Parakeet), 1.157 (German Whisper), and 1.076 (English Parakeet). The median ratio for the two Whisper cases was 1.256, failing the predefined 0.90 improvement target and per-case regression gates. Smaller weights did not improve performance here. Host activity was observed, not controlled; this targeted local screen is not a production p95 or a resolved explanation of the earlier old/new serving slowdown.
Repeated outputs were stable within each variant. Q8 changed words/confidence on the French and German Cap recordings; these remain provider-agreement diagnostics, not reference-transcript WER. English Parakeet words and timestamps stayed exact, but Q8 changed the acoustic language confidence, so the strictly frozen metadata identity gate also failed. No thresholds were relaxed to accept Q8.
The initial screen checker mistakenly looked for the stdio error field in native
HTTP errors. The pinned HTTP server emits detail and an optional code instead.
An independent audit verified all eight exact status/detail/code responses and
recovery results without changing or repeating any request. The original driver,
plan, results and initial report remain preserved alongside the corrected audit.
A separate private build added only a read-only getter for existing Whisper state counters and diagnostic JSON. It retained the original compiler/link options, identical GGML library bytes, F16 model, decoding and word-timing algorithms. Its 185 MMS-feature Rust tests passed. Nine actual native HTTP requests retained exact text, words, timestamps, confidence, language and model versus the frozen F16 responses, including all four long Cap cases and Spanish/Hindi/French controls.
Across the two long Whisper recordings, the weighted native-time breakdown was 46.38% encoding, 35.96% decoding, 11.08% sampling and 1.16% mel preparation. The remaining 5.42% is unassigned; it is not claimed to be alignment time. This makes encoder/decoder execution the next measured performance target. No optimisation or general throughput gain is claimed from adding counters.
The first private probe build failed on a copied include path containing an absent intermediate directory; a separate corrected build normalized that path. Its run then reached and saved all nine results before an overly strict probe assertion required a single-token decoder call. The last French sample validly used only batched/prompt decoding. Independent read-back verified all nine exact outputs, actual decoder activity, timing totals and process exits. Failed attempts and their receipts remain intact; no inference was rerun to hide the assertion failure.
Private evidence: turbo-q8-build1, turbo-q8-eval1/audited-report.json,
turbo-q8-eval1/supplement.json, and
native-stage-profile2/eval/audited-report.json under the ignored experiment root.
The application source still matches the preceding verified 70-job API checkpoint;
this experiment did not rerun or replace that public-API evidence. No Cap source,
shared database, normal runtime, deployment, model default, Git index or HEAD was
changed. Licensing is outside this task's acceptance gates.