

That depends much more on the model size than on the year. In the browser we’re limited to small models — tiny (English only), base and small, 40 to 250 MB — and on a German TV show those are exactly where Whisper gets shaky.
The desktop app and the CLI can run large-v3-turbo (1.6 GB, 99 languages), which is a different experience on non-English speech. If your last try was a small model, that’s the part worth changing before judging it.


Fair question about the model. The reason it’s Whisper is the runtime rather than the benchmark: the web version and the extension run on transformers.js, on WebGPU where it exists and WebAssembly where it doesn’t, and Whisper is what runs there today. The desktop and CLI builds use whisper.cpp with large-v3-turbo.
Granite Speech and Cohere Transcribe aren’t something I can run in a browser tab right now, so I can’t claim to have compared them on equal footing. If you’ve got benchmark numbers for them on long-form, multi-language audio, I’d read them.