You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
From the experimentation I've done with Whisper v3 large, accuracy is not that great when transcribing short segments even with highly articulate, high quality audio with no noise and native English speakers. It can be as low as 50-60% accurate with a latency of 1-3 secs. The latency is fine but quality is not.
Obviously accuracy improves significantly when transcribing larger segments such as a monologue or entire meeting. It's still not close to the high 90s %. With the latter it's probably in the 80s depending on the speaker and how articulate they are.
Has anyone gotten better results and if so what approach did you take? Prompts with vocab help a bit but it's very cumbersome and brittle. Thanks.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
-
From the experimentation I've done with Whisper v3 large, accuracy is not that great when transcribing short segments even with highly articulate, high quality audio with no noise and native English speakers. It can be as low as 50-60% accurate with a latency of 1-3 secs. The latency is fine but quality is not.
Obviously accuracy improves significantly when transcribing larger segments such as a monologue or entire meeting. It's still not close to the high 90s %. With the latter it's probably in the 80s depending on the speaker and how articulate they are.
Has anyone gotten better results and if so what approach did you take? Prompts with vocab help a bit but it's very cumbersome and brittle. Thanks.
Beta Was this translation helpful? Give feedback.
All reactions