Add support for Phonon-2 speech recognition #2174
spencerchubb
started this conversation in
Ideas
Replies: 2 comments
|
Yes, would love to use Phonon-2 |
0 replies
|
Please file an issue on transcribe.cpp |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Would there be interest in adding Fermion Research's Phonon-2 as an optional speech recognition model in Handy?
For English dictation, a smaller download with accuracy close to a larger model could make local transcription more accessible. According to the model card, Phonon-2 is a quantized version of NVIDIA's Parakeet TDT 0.6B v3 with:
These are the publisher's reported results; I haven't independently benchmarked it in Handy. The model is English-only, so it would be an additional option alongside the existing models.
The proposed feature is to make Phonon-2 available in Handy's model selector, with local download and inference. This seems consistent with Handy's emphasis on offline use, privacy, and accessibility. Integration feasibility, runtime dependencies, and performance in Handy would still need evaluation. The model weights are CC-BY-4.0, and the project's code is Apache-2.0.
Related alternatives include the existing Parakeet models and the Parakeet Redux / Ultra proposal (#2133). Phonon-2 seems worth considering as another compact English model.
I understand Handy is currently in a feature freeze, so I'm opening this discussion to gather community interest before any implementation work. Would others find this useful, and would support need to start in transcribe.cpp or another inference backend?
References:
All reactions