Auto-segment long recordings for transcription on low-memory or CPU-only machines #2176
DolphinDream
started this conversation in
Ideas
Replies: 1 comment
|
It's coming |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
The problem
I'm running Handy on a 2018 MacBook Pro (Intel, 16 GB RAM, no usable GPU), so transcription runs on the CPU. Recordings of around 13 minutes push memory use close to the full 16 GB. The machine stalls and I have to force quit the app, and the recording never gets transcribed.
What I did as a workaround
I had Claude split the saved audio file into shorter segments, update Handy's history database, and show the pieces in the app as separate entries. I then transcribed each segment individually, and that worked fine. Memory use seems to scale with recording length, so splitting a 13-minute file in half keeps usage well within limits. The splitting used ffmpeg to find pauses in the speech, so it didn't cut mid-word or mid-sentence. It works, but it's a manual, clunky process, and I'd rather Handy handle it.
Feature request: automatic segmentation behind the scenes
Related question: live overlay text vs. the final transcription
I use the live overlay that shows text as I dictate (using Nemotron streaming 3.5 model). When I finish, Handy seems to run a second transcription pass before pasting into the input field, and on long recordings that pass is what fails or stalls. A few questions and ideas:
Why it matters
Long dictation and brainstorming sessions are where Handy is most valuable, and they're exactly where older hardware breaks down. Automatic chunking and a fallback to the live transcript would make long recordings reliable for anyone without a powerful machine.
Happy to share more details about my setup or help test. Thanks for the great app!
All reactions