- Whistle is a model that converts speech into text, and according to the developer Cactus, a single file is 16.9MB. The source says it runs on a CPU without separate dependencies.
- It supports 7 languages, including English, German, French, Spanish, Italian, Dutch, and Polish, and can process up to 30 seconds at a time. If no language is specified, it detects the language automatically.
- It provides start and end times and a probability for each word. Passing expressions such as names through keywords is said to raise the chance that those expressions are selected during search.
- The performance figures are results measured by the developer. On an Apple M4 Pro CPU, with 10 seconds of audio and a beam size of 5, the first token time was 11.1ms, and under the same conditions Whisper base was compared at 73.2ms.
- Accuracy varies by benchmark. The source states that Whistle leads on the averages for LibriSpeech, SPGISpeech, Earnings-22, and FLEURS, while Whisper base leads on the averages for TED-LIUM, AMI, and MLS.
The claim that Whisper base came out ahead on the TED-LIUM, AMI, and MLS averages stands out first. The speed gap measured on the same 10-second audio is also a figure the developer published, so how it actually feels on real devices needs to be checked separately. I'm curious whether the picture holds up if I save one meeting recording as a 16kHz WAV and run it through the pip-installed version, even when Korean is mixed in.