Replies: 38 comments 36 replies
|
I've realised we actually almost have the functionality for this already. I would suggest extending the custom words list function to include support for direct replacement (exact match only, not using Levenshtein distance) and regex replace. I think this also needs a UI for editing the custom words (replace) list. At the moment it's not very user friendly - all custom words are stored in settings_store.json, with no UI based way of seeing what you have added or removing words. This would meet all the requirements mentioned in this discussion, and is a natural improvement to the already existing custom words function. It would also work for #162 |
|
Just thought of one more stupid use case – replace with dynamic content like |
|
Another use case is the option for sending final text to LLMs for post processing (so refining after post processing I guess?) I know sending data to APIs is a bit out of scope for this project but I regularly use local server for the final cleanup with this command: |
|
I'm also strongly in favor of having support for basic post-processing via a configuration file, for example. Currently, this limitation significantly restricts handy's usefulness for my use case, as I absolutely need to go through a post-processing step. The basics I need are:
This seems basic, probably not too hard to implement, and would be very flexible while waiting for more advanced solutions. Perhaps this could be implemented as an advanced option to avoid overwhelming regular users? What do you think? |
|
It would be great to have the option to run any code over the text before it is inserted. Maybe more of an option for devellopers, but that'd allow for any sort of post processing, including sending that to an LLM for cleanup / formating |
|
I would love a feature to be able to prompt freely my post-processing. Use cases for me are:
EDIT: says here its merged but i don' have the feature in my macos client |
|
Hi, I made this: #455 Handy.Replacements.mp4Edit: This is a first version of English punctuation rules : handy-replacements-english punctuation-v1.0.json Edit2: This is a first version of French punctuation rules : handy-replacements-french-punctuation-v1.0.json You'll just have to click on "import" button in the replacements tab and select the file to start to use it. |
|
Guys, I am not sure if it should be here or not, but how are you handling removing filler words? like (uh uh etc) |
|
It seems difficult for me to understand. |
|
Would be great to have an option to insert space after transcription. So if I am transcribing two consecutive sentences, there's a space after the period. Without this option, we get results like:
|
|
Can we have more variables in prompt when post-processing:
|
|
Hey! I built a lightweight local server for exactly this use case: It provides an OpenAI-compatible Rule types:
Features:
Fair warning: I put this together pretty quickly with Claude Code, so it's definitely a beta. But it's been working well for my own use case (German punctuation rules). PRs welcome! |
|
I am trying to use this with Amazon Bedrock. However, both OpenAI models are thinking models insist on returning the "thinking back". IE: It would be helpful to have a custom script we can add to remove this, like some simple regex. |
|
Has anyone had an issue where, when using post-processing, it spits out text that is completely different to what you've said? I mean, a completely different topic. There are no words that are even the same. This happened twice in a row today. Then I did local processing only, and it was correct. Not sure whether to mark it as a bug at this stage or not. I did have a look through the existing bugs and I couldn't see one that was the same. |
|
This is indeed not the case, because I just ran a blank transcription (by mistake) for about one second and got the following after about 1-2 seconds. I'm using OpenRouter using Gemini 2.5 Flash Lite. I am also just running a slightly modified version of the original prompt, which are changes to do with formatting. " I hope this email finds you well. I'm writing to follow up on our discussion about the upcoming marketing campaign. I've reviewed the initial proposal, and I have a few suggestions before we proceed. Firstly, I think we should allocate a larger portion of the budget to digital advertising. Our target audience spends a significant amount of time online, and this channel has proven to be highly effective in previous campaigns. Secondly, I’d like to propose a collaboration with Influencer X. They have a strong following within our demographic, and their endorsement could significantly boost our reach. Finally, we need to finalise the key performance indicators for the campaign. It’s crucial that we have clear, measurable objectives to track our progress and ensure we're on the right track. Please let me know your thoughts on these points. I'm available to discuss this further at your convenience. Best regards, This has nothing to do with my work, and I have nothing to do with Tom or Sarah. |
|
Post-processing model returns modified text as expected - but Handy inserts the original transcribed text anyway. |
|
Any plans on adding this feature soon? |
|
Following up from the HN thread (https://news.ycombinator.com/item?id=48963879): I had the "reuse the loaded model for post-processing" idea spiked out — when the loaded model is an instruction-following LLM (Voxtral today), the post-processing pass runs on the weights already in memory. No endpoint, no API key, no second model. Working branches, validated end-to-end on Voxtral-Mini-3B-2507 Q4 (Metal):
Not trying to preempt the header design @cjpais said he's thinking about — happy to rework the branch to whatever surface he lands on, or drop it if he'd rather build it himself. Posting here per the feature-freeze guidance to gauge interest. (Disclosure: the implementation was AI-assisted; I direct and review the work.) |
|
This is good example of how post processing can help with prompting. Most of the time, when we use STT, we are dumping our raw thoughts. And if we have something like this, then we can craft very, very high quality prompts. Raw Thoughts as STT -> Refine (post processing) via AI again -> Paste |
|
I’d like to suggest exposing the reasoning effort / thinking level for the Custom post-processing provider. Handy already has |
|
Superwhisper just released S1-mini, a 0.6B local model made specifically for cleaning up ASR transcripts, with GGUF builds available. Could it be relevant for Handy’s local post-processing plans? |
|
I'm usingA pass processing combination key I'm using the provider Apple Intelligence. I'm on a MacBook M1 with iOS 27 beta 6, I'm using the Improve transcription prompt below: ${output}The above is a transcript generated by a speech-to-text model. Clean it by:
Preserve exact meaning and word order. Do not paraphrase or reorder content. If the transcript is empty, output nothing (a single space at most). Do not output messages like "The transcript is empty". Return only the cleaned text. But I'm having a hard time understanding how to actually use it, Because if I use a keyboard combination: what I get is the prompt above pasted on the input field. And I thought that what would happen is that Apple Intelligence would actually make better the transcript, can anyone tell me what I am doing wrong? |
|
+1 for explicit word replacements/correct spellings. Handy already lets users add custom words. It would be useful if adding one could optionally define a commonly transcribed misspelling and its intended spelling, similar to Wispr Flow’s “Correct a misspelling” option. |
|
I just started using handy, and found using it for sentence fragment edits was getting annoying, so this is the prompt I use after linking to a bionic LLM with post processing: === Maybe this will help others, so i figured I'd post it here, but there's probably a leaner way to do the same thing. (edits, tweaks) |
|
Post Processing Idea: Try this
|
|
Is there already a way to use Codex CLI for post-processing on Windows with an existing ChatGPT subscription, without separate API billing? If not, I'd love to see this supported—similar to WhisperBar. |
|
Hi, everyone, I know that new features are not accepted, so I would open a PR. |
|
I'd like to see window-title-aware post processing. I'd like to have a different system prompt for when I write a WhatsApp message than when I write an email. Additional idea: Give Handy "eyes". I could make screenshot and then OCR or VLLM that screenshot and give that as additional relevant context. |




Uh oh!
There was an error while loading. Please reload this page.
As raised within #157 it would be nice to have an option to post-process the transcript before it is inserted.
Use cases outlined so far:
"or new line with\n@jamaggsI initially envisioned this as a simple regex search-replace, but for other use cases it would not be that simple to configure. One proposed solution is to add an advanced option to specify CLI command for post-processing. CLI command should receive a transcript & it's context via
stdinand respond with a processed transcript usingstdout.For example, post-processing CLI script input could be:
{ "transcript": "{transcript}", "language": "en", "foregroundApplication": ".../chrome.exe" }All reactions