Skip to content

Repository files navigation

Ollama Translate

Translate documents using a local Ollama model. This tool extracts text from supported files, sends it to a local LLM for translation, and writes the translated content back to a new file. This tool currently uses Gemma3 for which we observed the best results (speed/accuracy) with low VRAM requirements (if locally, +3GB).

Features

  • Local translation with Ollama
  • Supports multiple file formats: .odt, .ods, .odp, .docx, .xlsx, .pptx, .txt, .csv, .xml, .html, .htm, .md, .tex
  • Preserves structure
  • Verbose mode to show original and translated text
  • Language selection for source and target, or agnostic source language

Requirements

Softwares

Hardware

If locally:

  • +3GB VRAM

Python dependencies

pip3 install -r requirements.txt

Usage

./ollama-translate.py INPUT_LANG OUTPUT_LANG { INPUT_FILE | -t TEXT }

Examples:

./ot.py en es document.docx	# Translate file from English to Spanish
./ot.py - es document.docx	# Translate file from any language to Spanish
./ot.py -r en es dir/		# Translate files recursively from English to Spanish

./ot.py en es -t "Hello world !" # Translate text from English to Spanish
  • Note: When translating from a specific language, any other language will be kept as-is.

List available languages:

./ot.py -l	# Short
./ot.py -ll	# Full

Advanced usage

Verbose

Show original and translated texts ;

./ot.py en es document.docx -v

Optimisation

Exclusions

Preprocessing filter

To translate your document faster and more accurately ; you might exclude words (insensitive case strings) that you know cannot/shouldn't be translated (e.g, Programming Language, Company Name, Conference Name) ;

./ot.py en es document.docx -e "Turing, Einstein, NoSQL, USENIX"

If a string appear to not contain any relevant word, it will be kept as is without being passed to LLM.

  • Note: Applicable to -t/--text
  • Note 2: These exclusions are applied before any other default exclusion, be careful not to overwrite/overlap default exclusions, e.g, Albert Einstein would overlap einstein.com/
  • Note 3: The word is expected to match word boundaries (\b), it can't terminate in the middle of a word (e.g, Turin won't match Turing)

Default excluded expressions (regex) (cf. conf.py) :

  • Emails
  • URLs/domains
  • Phone numbers
  • Digits and capitals being more than 3 characters
  • Numbers
Prompt exclusion ⚠️

If a string appear to contain relevant words even after preprocessing filter, prompt exclusion can be used to prevent expressions from being translated by the LLM. This would typically include expressions that have possible translation e.g, a conference name ;

./ot.py - es document.docx -E "Internet of Things Journal"
# Possible to use with `-e`
./ot.py en es document.docx -e "Turing" -E "Internet of Things Journal"
  • Note: List passed to -E is also used as a preprocessing filter (-e)
  • Note 2: Prefer -e when possible

About

Translate documents using a lightweight local Ollama model (+3GB VRAM): odt, docx, md, tex, etc.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages