Whistle: Speech to Text in 16.9 MB
- skolos - 8114 sekunder sedanInteresting that this is here. I used whistle (and bunch of other things) to take ownership of my echo show. It now doesn't dial to Amazon at all - it does all processing locally with its own CPU and connects to my homeassistant for home automation. My initial setup involved qwen asr (1.7b model) running on rtx 5080. Compared to that, whistle was really bad (out of 170 messages, qwen recognized correctly 168, whistle - 70), but I adjusted whistle to work like jev - instead of free form transcription it recognizes only select set of templates (I trained tiny network with 10,000 generated utterances to translate whistle final state to probabilities within templates). The precision went up to 164/170 - almost matching qwen. By the way - I'm speaking with heavy accent.
- INTPenis - 18502 sekunder sedanI don't think the challenge with speech to text was size of the binary. In my experience the challenge is understanding my 84 year old Croatian father with a sagging mouth after a stroke, when he's trying to write his autobiography.
I just setup Windows speech to text for him last week and it's great to see how he can write an entire page in 10 minutes, it would take him days using the keyboard.
But every single sound he makes with his mouth ends up on the page too.
- albert_e - 15600 sekunder sedanWhat the demo does not do is show streaming output of transcribed text as we are speaking and recording (before we hit stop). That is an essential feature IMO for most general purpose live STT apps.
- flowerlad - 1059 sekunder sedanApple needs to incorporate this into iOS ASAP. This works much better than the speech recognition in iOS when you use technical terms. One of the most frustrating parts of iOS is speech-to-text in iMessage. For me no feature is more important in a phone.
Try this example: My website uses ASP.NET technology and I am using .NET 10.0. Works perfectly in Whistle, but not in iOS.
- zimpenfish - 13530 sekunder sedanTried it on a random TV episode and it seems to get stuck sometimes where it just outputs "Thank you." as a default - at one point emitting that for 60s of dialogue (and no, the episode does not have someone repeating "Thank you." for 60s.) Happens several times during the transcription.
- wkcheng - 15563 sekunder sedanHow does this compare with Parakeet? I've been using that locally in my projects on an M-series macbook and it's been working great. It's fast and accurate enough for my use cases (meeting transcription, audio transcription for demo videos, etc.)
This definitely seems lighter and faster. How does accuracy compare?
- andy_ppp - 18818 sekunder sedanWow certainly in English this is incredibly accurate I tried to break it and it understood me perfectly!
I know it's slightly off topic but surely it must be easy by now to train a spell checker that doesn't annoy the crap out of everyone using it (looking at you here Apple)!
- amelius - 5956 sekunder sedanLet's say I want to build a hardware product now, voice-controlled, with voice feedback, so STT, LLM, and TTS. All local. What are the best libraries to do this now, say with 8GB of GPU memory available?
- joewhale - 18315 sekunder sedanI initially read this as whistle to text, which would be way cooler.
- properbrew - 10070 sekunder sedanMight look into embedding Whistle into Whistle if it can make it 30x smaller (https://play.google.com/store/apps/details?id=com.blazingban...) - It's a shame there isn't as many languages supported though, I'm surprised at the amount of non-english downloads (I really shouldn't be, of course non-english speakers want dictation) of the app there is.
- TomGarden - 8181 sekunder sedanThe Qwen STT model that's leading the open weight leaderboard right now is excellent. Parakeet v2 is blazing fast and accurate enough on English. I feel like STT gains from this point on will be marginal, especially given that you can do a quick LLM pass afterwards with a small model
- sfpk - 12671 sekunder sedanIf you need speech2text try this, https://ccoreilly.github.io/vosk-browser/
- jakobov - 8188 sekunder sedanZWhispr seems to be the best for those who care about accuracy as it uses three SOTA models.
- MisterMunchkin - 13087 sekunder sedanFor something you could plausibly ship inside a webapp, it’s very good. I can definitely think of some cool uses for this.
- e12e - 15431 sekunder sedanHm. I saw language=detect and tried some Japanese - which (given the actual list of supported languages) unsurprisingly turned into some mangled Spanish.
Since it doesn't support Norwegian - I tried English - and it mis-transcribed "cleaning" for "training" - probably a failure due to context/training (Hello everyone, today we are going to do some cleaning).
So, reasonable, but limited?
- armcat - 17858 sekunder sedanThose are insane benchmarks at this size. Well done!
- kamranjon - 17350 sekunder sedanSooo I haven't really been super impressed with the needle models before, but this is very impressive. It transcribed multiple sentences I gave it with complex timing and words and in such a small footprint, I'm super impressed. Excited to see what types of things can be built with something like this, the performance seems very good.
- rafaelm - 14744 sekunder sedanHuh, this was really confusing. I already had an STT app called Whistle on my Android phone.
- rshemet - 11019 sekunder sedanhey, Roman here from Cactus, thank you for the feature!
opening this thread for questions/feedback if you have any
- jayshah5696 - 15795 sekunder sedanThis is actually a really great release. Congratulations team. I just tried few words. My Indian accent also was able to pick up.I'm gonna run it on my Linux Box.
- Centigonal - 14186 sekunder sedanThis is quite good, especially given the size and the fact that it runs on the CPU.
- mo2art - 16817 sekunder sedanRuntimeError: audio limit is 30 s
- mrkn1 - 17754 sekunder sedanlove seeing more sub-20MB, CPU-first models. if anyone wants a CLI built on the same ethos (no GPU, no cloud), been using yapsnap streaming Zipformer ASR, plus diarization and timestamps all on CPU! It supports 10 languages. Unlimited transcription for free.
- pzo - 11890 sekunder sedantested in polish and unless you speak very loud, clear and slow is not that good, parakeet definitely better.
- lab14 - 5778 sekunder sedanTried it a few times with English, Spanish and French and the quality/accuracy is pretty "meh". If the model doesn't really work, it doesn't matter if it fits in 1MB.
- saturn8601 - 18528 sekunder sedanInitial tests make this feel just like iPhone's terrible text to speech. It is the one thing I utterly hate about iPhone. Ive tried apps that try to embed themselves into the iPhone keyboard and they always don't work out well. Hopefully this gets better and we can somehow get it into the iPhone more seamlessly.
- - 10104 sekunder sedan
- aidotguru - 17314 sekunder sedaneager to see if working in android phones
- tecleandor - 18803 sekunder sedanSpanish is not good (seems to write non existing words and/or with terrible typos...) but English seem to work good even with my (Spanish) accent...
- paaloeye - 9759 sekunder sedanRIP Wispr Flow
- contingencies - 9410 sekunder sedanFor speech to text UX I currently use https://handy.computer/ as it's cross platform and open source. With that I am currently using Parakeet Unified EN 0.6B and finding it excellent. Often I use it to talk to AIs without giving them audio, which works very well. Honestly, I would never go back to typing now. Promised since ~Y2K, the tech is finally here. You really notice it when you wake up at 2AM and don't want to wake people ... it can get really annoying reverting to key-tapping. My long-gnawing fear of losing my hands to RSI is no longer a thing, and I can focus on losing them to another hobby: like sailing or machining! Just bought a band saw...
- try-working - 11232 sekunder sedanI built an STT plugin for DeepSeek Harness that uses this Whistle model as well as a larger one from Desert Ant Labs: https://github.com/try-works/dsh-stt
- - 9634 sekunder sedan
- - 9560 sekunder sedan
- thomkaar - 13069 sekunder sedanit can't even tell when i say "my booty thick"
- - 9588 sekunder sedan
- thomkaar - 13081 sekunder sedani can't even tell when i say my booty thick
- agilek - 16488 sekunder sedanCan we have more languages?
- isaac_laughs - 6784 sekunder sedan[flagged]
- - 17926 sekunder sedan
- promptspheree - 13477 sekunder sedan[dead]
- bassie123 - 9884 sekunder sedan[dead]
- theaxeonthanksg - 7645 sekunder sedan[dead]
- grezql - 16796 sekunder sedan[dead]
- eethrowaway - 16697 sekunder sedan[dead]
Nördnytt! 🤓