Text Processing

Libraries for parsing and manipulating plain texts.

Mojibake like "café" is ftfy's to fix, and charset-normalizer detects encodings. For fuzzy matching, add RapidFuzz to your Python text processing libraries.

How to choose:

  • Text decoded with the wrong encoding: ftfy
  • Bytes in an unknown encoding: charset-normalizer
  • Fuzzy matching against a long list of strings: RapidFuzz
  • Diffs, or close matches with nothing to install: difflib
  • ASCII-art banners: pyfiglet
  • Translations, plus dates and numbers in each user's locale: Babel
  • Syntax highlighting: Pygments
  • A grammar written in Python code: pyparsing, or parsy for small languages
  • Splitting and formatting SQL: sqlparse
  • Phone numbers: phonenumbers
  • URL slugs: python-slugify, or Unidecode for plain ASCII
  • Short IDs that users will read: shortuuid

Listed in editorial order, grouped by use case. Click a column to re-sort the whole list.

Press / to search. Tap a tag to filter. Click any row for details.

Search and filter

Results

Row number Tags

Encoding and Unicode 3 projects

Universal character encoding detector, and a dependency of requests.
jawah/github.com/jawah/charset_normalizer / /1,194,739,615 downloads/month
Python character encoding detector.
chardet/github.com/chardet/chardet / /164,865,407 downloads/month
Makes Unicode text less broken and more consistent automagically.
rspeer/github.com/rspeer/python-ftfy / /14,252,070 downloads/month

Fuzzy Matching 1 project

Rapid fuzzy string matching using various string metrics, with a C++ core.
rapidfuzz/github.com/rapidfuzz/RapidFuzz / /119,847,030 downloads/month

General 2 projects

(Python standard library) Helpers for computing deltas.
An implementation of figlet written in Python.
pwaller/github.com/pwaller/pyfiglet / /5,965,288 downloads/month

Internationalization 1 project

An internationalization library for Python.
python-babel/github.com/python-babel/babel / /104,141,572 downloads/month

Parser 5 projects

A generic syntax highlighter.
pygments/github.com/pygments/pygments / /933,250,675 downloads/month
A Python library for creating PEG parsers.
pyparsing/github.com/pyparsing/pyparsing / /318,913,439 downloads/month
A non-validating SQL parser.
andialbrecht/github.com/andialbrecht/sqlparse / /109,968,141 downloads/month
Parsing, formatting, storing and validating international phone numbers.
daviddrysdale/github.com/daviddrysdale/python-phonenumbers / /35,038,572 downloads/month
Easy, generic parser combinator library for creating parsers.
python-parsy/github.com/python-parsy/parsy / /4,310,015 downloads/month

Transliteration and Slugs 2 projects

A Python slugify library that translates unicode to ASCII.
un33k/github.com/un33k/python-slugify / /65,171,523 downloads/month
ASCII transliterations of Unicode text.
avian2/github.com/avian2/unidecode / /26,219,209 downloads/month

Unique identifiers 1 project

A generator library for concise, unambiguous and URL-safe UUIDs.
skorokithakis/github.com/skorokithakis/shortuuid / /16,925,188 downloads/month

Text Processing guide

ftfy fixes Unicode that's broken in various ways: bad Unicode goes in, good Unicode comes out. It's a mojibake detector and fixer, not an encoding detector, so give it text you've already tried to decode, never bytes. The function you'll call most is ftfy.fix_text(), which applies every fix it can. ftfy aims never to change text that was decoded correctly.

charset-normalizer reads text from an unknown charset encoding. Call from_bytes(data).best(), or from_path(path).best() for a file, and str() of the result is your decoded text. chardet does the job differently: charset-normalizer decodes candidate encodings and scores the text, while chardet scores raw bytes against per-language models.

RapidFuzz does fuzzy string matching with various string metrics, mostly in C++, with a pure Python fallback for every algorithm. To compare one string against a list, use the process module, like process.extractOne().

difflib comes with Python and compares sequences, mostly lines of text, producing unified, context, and HTML diffs like the ones diff and git diff print. For a "did you mean" suggestion, get_close_matches() returns the best "good enough" matches for a word.

pyfiglet is a full port of FIGlet into pure Python: it renders text in ASCII-art fonts, as in pyfiglet.figlet_format("text", font="slant").

Babel does two jobs: it builds gettext message catalogs without the GNU gettext tools for common tasks, and it formats dates, numbers, and locale names from CLDR data. Run catalogs through the pybabel command: extract messages, init a catalog per language, update it, and compile it.

Pygments is a generic syntax highlighter for code hosting, forums, wikis, and other apps that show source code. If you highlight code your users submit, its docs recommend you stop the Pygments process after a short timeout.

pyparsing builds a grammar directly in Python code from a library of classes, instead of lex/yacc or regular expressions. It handles whitespace, quoted strings, and embedded comments for you. Its best practices start with writing a BNF before any code. Use results names to get tokens by field name rather than by position. parsy combines small parsers into larger ones and gives you building blocks only. Its docs say it excels at easy-to-read parsers for relatively small languages.

sqlparse is a non-validating SQL parser: it splits scripts into statements, formats them, and walks their token tree, whatever the dialect. Its three module-level functions, split(), format(), and parse(), cover most needs.

phonenumbers is a Python port of Google's libphonenumber. Parse a number together with the country it's dialed from, unless it's in E.164 format. Then check it with is_possible_number() or is_valid_number(), and print it with format_number().

python-slugify makes Unicode-aware slugs with your choice of transliteration backend: by default, Unidecode if it's installed, otherwise text-unidecode. Its own code is MIT, while Unidecode is GPL, so check that license before you install Unidecode next to it.

Unidecode makes lossy ASCII transliterations of Unicode text, close to what someone on a US keyboard would type. Use it for ASCII machine identifiers. Its README says it's best kept off strings your users see.

shortuuid generates concise, unambiguous, URL-safe UUIDs for IDs users will see: it takes UUIDs from Python's uuid module, writes them in base57, and drops look-alike characters such as l, 1, I, O, and 0. Call shortuuid.uuid() for a new ID.

Decode with the encoding you were told whenever there is one. charset-normalizer's FAQ calls detection a last resort, and ftfy's docs say to assume UTF-8 unless you have a specific reason to believe otherwise.

Know a project that belongs here?

Tell us what it does and why it stands out.