Congrats on the launch! I'm glad more progress is being made in this area.
Because of the hallucinations inherent in transformer models, I went looking for a transducer-based model with a low WER. I have been super pleased with parakeet-unified-en-0.6b. It's WER isn't as low as Canto, but it's about as low as you can get (~5-6.5%) with a non-transformer-based model as far as I'm aware.
I've been very pleased with its output.
I wasn't looking for this, but it's also lightweight enough to run on my little potato PC, which has an i5 8th gen processor, and still transcribe 9-10x faster than realtime.
I vibe-coded a little wrapper for it and use it on folders of audio or podcast rss feeds or even YT playlists and channels and it's been one of my new favorite tools.
Because of the hallucinations inherent in transformer models, I went looking for a transducer-based model with a low WER. I have been super pleased with parakeet-unified-en-0.6b. It's WER isn't as low as Canto, but it's about as low as you can get (~5-6.5%) with a non-transformer-based model as far as I'm aware.
I've been very pleased with its output.
I wasn't looking for this, but it's also lightweight enough to run on my little potato PC, which has an i5 8th gen processor, and still transcribe 9-10x faster than realtime.
I vibe-coded a little wrapper for it and use it on folders of audio or podcast rss feeds or even YT playlists and channels and it's been one of my new favorite tools.