Translate documents using a local Ollama model. This tool extracts text from supported files, sends it to a local LLM for translation, and writes the translated content back to a new file. This tool currently uses Gemma3 for which we observed the best results (speed/accuracy) with low VRAM requirements (if locally, +3GB).
- Local translation with Ollama
- Supports multiple file formats:
.odt,.ods,.odp,.docx,.xlsx,.pptx,.txt,.csv,.xml,.html,.htm,.md,.tex - Preserves structure
- Verbose mode to show original and translated text
- Language selection for source and target, or agnostic source language
- Python 3.10+
- Ollama
If locally:
- +3GB VRAM
pip3 install -r requirements.txt./ollama-translate.py INPUT_LANG OUTPUT_LANG { INPUT_FILE | -t TEXT }Examples:
./ot.py en es document.docx # Translate file from English to Spanish
./ot.py - es document.docx # Translate file from any language to Spanish
./ot.py -r en es dir/ # Translate files recursively from English to Spanish
./ot.py en es -t "Hello world !" # Translate text from English to Spanish- Note: When translating from a specific language, any other language will be kept as-is.
List available languages:
./ot.py -l # Short
./ot.py -ll # FullShow original and translated texts ;
./ot.py en es document.docx -vTo translate your document faster and more accurately ; you might exclude words (insensitive case strings) that you know cannot/shouldn't be translated (e.g, Programming Language, Company Name, Conference Name) ;
./ot.py en es document.docx -e "Turing, Einstein, NoSQL, USENIX"If a string appear to not contain any relevant word, it will be kept as is without being passed to LLM.
- Note: Applicable to
-t/--text - Note 2: These exclusions are applied before any other default exclusion, be careful not to overwrite/overlap default exclusions, e.g,
Albert Einsteinwould overlapeinstein.com/ - Note 3: The word is expected to match word boundaries (
\b), it can't terminate in the middle of a word (e.g,Turinwon't matchTuring)
Default excluded expressions (regex) (cf. conf.py) :
- Emails
- URLs/domains
- Phone numbers
- Digits and capitals being more than 3 characters
- Numbers
If a string appear to contain relevant words even after preprocessing filter, prompt exclusion can be used to prevent expressions from being translated by the LLM. This would typically include expressions that have possible translation e.g, a conference name ;
./ot.py - es document.docx -E "Internet of Things Journal"
# Possible to use with `-e`
./ot.py en es document.docx -e "Turing" -E "Internet of Things Journal"- Note: List passed to
-Eis also used as a preprocessing filter (-e) - Note 2: Prefer
-ewhen possible