Using the parser

Requires Python 3.11+. pip install nameparser

Parse a name

>>> from nameparser import parse
>>> name = parse("Dr. Juan Q. Xavier de la Vega III")
>>> name.given, name.family
('Juan', 'de la Vega')
>>> name.title, name.middle, name.suffix
('Dr.', 'Q. Xavier', 'III')

A parsed name has seven fields: title, given, middle, family, suffix, nickname, and maiden. Parsing never raises; unparseable input yields a ParsedName with empty fields plus any ambiguities the parser noticed along the way (see When the parser had to guess below, and How the parser works for why they exist). A name with no fields set is falsy, which is how you tell “nothing parsed” from “parsed to something”:

>>> bool(parse("")), bool(parse("   ")), bool(parse("John"))
(False, False, True)

Input shapes

Three arrangements are understood by default, two more when you declare a family-first order, and two more by how they are written.

Understood by default

Every piece of each is optional:

  1. Title Given "Nickname" Middle Middle Family Suffix

  2. Family [Suffix], Title Given (Nickname) Middle Middle[,] Suffix [, Suffix or Title]

  3. Title Given Middle Family [Suffix], Suffix or Title [, Suffix or Title]

A title written with a period (Prof.) is also read at the end of a name in forms 1 and 3 and at the end of the given name in form 2; written bare there (Prof), it is a name word, because many title words are also surnames.

The last two differ in what the comma is doing. In form 2 it separates the family name from the rest, so the family name comes first; in form 3 it only sets off suffixes and titles, and the name before it is still given-then-family:

>>> parse("de la Vega, Juan Q. Xavier III").family   # form 2
'de la Vega'
>>> parse("Doe Jr., John").suffix                    # form 2, suffix before the comma
'Jr.'
>>> parse("John Doe, Jr.").family                    # form 3
'Doe'

A part after the comma in form 3, or after a further comma in either form, holds suffixes and titles, a word there that is a title rather than a suffix being read as the title (after a further comma, since 2.4):

>>> parse("Eric H. Holder, Attorney General").title
'Attorney General'
>>> parse("Eric H. Holder, Jr., Attorney General").title
'Attorney General'

Family-first forms you declare

Two more arrangements apply only when name_order declares family-first input — common outside Europe; see Family-first name order:

  1. Title Family Given Middle Middle [Particle] [, Suffix] (FAMILY_FIRST)

  2. Title Family Middle Middle Given [, Suffix] (FAMILY_FIRST_GIVEN_LAST)

A trailing particle earns a slot in form 4 alone because it is displaced from the family name it belongs to; form 5’s trailing word is the given name by the caller’s declaration, so there is nothing there to reinterpret.

Forms the script carries

Two more arrangements are native East Asian forms. The first needs no name_order at all — the script itself carries the reading — and the second is read the way the default order reads it, its own written order:

  1. Family Given [Honorific]

  2. Given[·Given]·Family / katakana transcription (source order)

Form 6 is the native family-first arrangement written in Han or Hangul — spaced or unspaced, with the honorific spaced or glued and landing in suffix. It has no title slot and no comma, because native CJK writing has neither convention. Form 7 is a transcription listing — Han divided by the 间隔号, or katakana joined by the nakaguro — never segmented, and kept in the order it was written as long as name_order is the default: the script settles no order for a transcription, so a declared FAMILY_FIRST reverses it as it reverses a Latin name.

A comma or a Latin wrapper around a CJK name — a listing comma, a Latin honorific or credential set beside it — is tolerated input: parsed best-effort, its handling changeable without notice.

Forms 6 and 7, and how a Latin wrapper around either is handled, are covered in full under East Asian names below.

Words that attach to their neighbors

Input shapes tell you where the fields sit. This tells you which words merge into one field instead of standing alone — between them, that is most of what decides a parse.

Most words stand alone. A few pull in a neighbor, and each kind pulls into one particular field — so this table is also a map of which fields get built by attachment rather than by position. Titles and suffixes attach only to adjacent words of their own kind (a run of titles chains into one, but a title never swallows a plain name); the rest reach forward to pull in the next word, whatever it is. Follow a row’s name to its full vocabulary set:

Words

Attach to

Field

Example

Titles

adjacent titles

title

Asst. Vice Chancellor John → Asst. Vice Chancellor

Suffixes

adjacent suffixes

suffix

John Smith PhD MD → PhD, MD

Bound given names

the following word

given

abdul salam ahmed → abdul salam

Particles

the following surname

family

Juan de la Vega → de la Vega

Maiden markers

the following name

maiden

Jane Smith née Jones → Jones

Conjunctions

the words on both sides

any

John and Jane Smith → John and Jane

Conjunctions are the exception the last row names: they take the field of whatever sits on both sides, so the same and pulls two given names together as easily as two surnames:

>>> parse("John and Jane Smith").given
'John and Jane'
>>> parse("Juan de la Vega y Rodriguez").family
'de la Vega y Rodriguez'

One shipped vocabulary works the other way round and so is not in the table above: surnames splits a word instead of merging two, and is covered under East Asian names. Customizing the parser covers how to change which words are in each of these sets.

A particle at the start of a name

Position matters in exactly one place: a particle standing on its own at the start of a name. It has no surname to attach to yet, so what decides the reading is whether it is one that can double as a given name: the particle either becomes the given name or turns the whole name into a surname. Only the first of those is name_order’s question — see Where the vocabulary answers first, and read the given name below as the default given-first order’s — since a particle that can never be a given name is the surname whatever order you declare:

>>> parse("van Gogh").given          # 'van' can be a given name
'van'
>>> parse("de Mesnil").given         # 'de' cannot
''
>>> parse("de Mesnil").family
'de Mesnil'

A comma gets there first. It names the surname outright, so a particle opening that surname has nothing left to decide and the part after the comma is the given name:

>>> parse("de Mesnil, Juan").given
'Juan'

Words that are also ordinary names covers changing which particles may double as given names.

East Asian names

A Chinese, Japanese, or Korean name written in its own script puts the family name first: 毛泽东 is MAO Zedong, 山田太郎 is YAMADA Taro, 김민준 is KIM Minjun. The family name is short — one Han character or one hangul syllable, occasionally two — and comes from a small closed set, while given names are open-ended. And in native writing the parts are usually not separated at all: the whole name is one unbroken run of characters. A parser therefore has two distinct jobs here: assign family and given to the right fields, and, when the name arrives as a single token, find the boundary inside it.

Which name comes first

Field assignment is automatic. A name written wholly in Han characters or hangul is assigned family-first, because every language written in those scripts orders names that way — Chinese and Japanese share little else, but they agree on this — so the assignment requires no knowledge of which language the name is in:

>>> parse("毛 泽东").family
'毛'
>>> parse("山田 太郎").family
'山田'

Splitting an unspaced name

Splitting an unspaced name is also automatic, but only for Korean. Hangul is written by exactly one language, and Korean family names are limited to a small closed set: the census surname list is part of the default vocabulary, and the longest listed surname at the start of an unspaced hangul token becomes the family name:

>>> minjun = parse("김민준")
>>> minjun.family, minjun.given
('김', '민준')

The same split is not automatic for Han text, because there the script does not identify the language: 高橋一郎 is a Japanese name whose family name is 高橋, but 高 alone is a common Chinese surname, so a Chinese surname list would split it in the wrong place. Declaring the language is up to you. When you know the data is Chinese, apply the zh locale pack, which carries the surname list the split needs:

>>> from nameparser import locales, parser_for
>>> parser_for(locales.ZH).parse("毛泽东").family
'毛'

Transcribed foreign names

Chinese also has a transcription convention, and it is written in the punctuation: a foreign name transcribed into Han characters keeps its source order and divides its parts with the 间隔号, the interpunct · (U+00B7) — 威廉·莎士比亚 is William Shakespeare, given name first. The dot itself is the marker, so nameparser reads U+00B7 as a token separator when it sits between classified-script characters (each side judged on its own — a hangul character beside a Han one qualifies), and a dot anywhere in the name reads the whole name as a transcription listing: it keeps the declared order — the order it was written in, under the default — and is never segmented — the role pure katakana plays for Japanese transcriptions, played here by the divider instead of the script:

>>> shakespeare = parse("威廉·莎士比亚")
>>> shakespeare.given, shakespeare.family
('威廉', '莎士比亚')

Because only classified characters on both sides make it a divider, the same codepoint interior to a Latin-script name — the Catalan punt volat in Gal·la — is untouched, and a dot with a classified character on just one side (王·Smith) stays part of the word undivided:

>>> parse("Gal·la Puig").given
'Gal·la'
>>> parse("王·Smith").given
'王·Smith'

The Japanese middle dot ・ (covered next) is a different mark carrying a different convention, and keeps its own reading.

Japanese

Japanese writes a name in three scripts at once. Family names and most given names are kanji — the same characters Chinese uses — but a given name is often written in one of the two kana syllabaries instead: hiragana (高橋みなみ) or katakana (山田エミ). The two syllabaries carry different information about whose name it is. Hiragana never transcribes a foreign name, so a name that holds hiragana, alone or beside kanji or katakana, is a Japanese person’s name, written in Japanese order. Katakana is ambiguous: native given names use it, but katakana is also how Japanese text writes a foreign name — マイケル・ジャクソン is Michael Jackson — and a transcription keeps the source language’s order, given name first, its parts divided by the middle dot ・ (the nakaguro, U+30FB) rather than by a space. Katakana also has a halfwidth form (タロウ for タロウ), which older systems that could not store kanji used for every name, Japanese or foreign; bank, payroll and CSV exports still carry it. nameparser reads halfwidth katakana exactly as it reads the full-width form.

What happens automatically

Two behaviors follow from that without any configuration. A name whose characters stay within kanji and kana, carry at least one kana, and are not katakana alone is assigned family-first, like any other native-script East Asian name: it cannot be Chinese, and it is not a transcription, because a transcription is written in katakana alone.

>>> minami = parse("高橋 みなみ")
>>> minami.family, minami.given
('高橋', 'みなみ')
>>> parse("山田 エミ").family
'山田'
>>> parse("山田 タロウ").family
'山田'

And the middle dot separates tokens the way a space does, so a transcribed foreign name divides into its parts — which, being wholly katakana, keep the order they were written in:

>>> michael = parse("マイケル・ジャクソン")
>>> michael.given, michael.family
('マイケル', 'ジャクソン')

Dividing an unspaced name: the segmenter

Dividing an unspaced Japanese name is a separate matter, and one no surname list can settle: family and given names draw on the same kanji, both sides run one to four characters, and the reading rather than the spelling decides most divisions. nameparser therefore takes a segmenter — a callable from token text to a division — and ships a factory wrapping namedivider-python, an optional dependency installed with the ja extra:

$ pip install "nameparser[ja]"
from nameparser import locales, parser_for

parser = parser_for(locales.JA, segmenter=locales.ja_segmenter())
parser.parse("山田太郎").family        # '山田'

Both halves are required, and they do different jobs: the ja pack activates division for Japanese text, and the segmenter performs it. ja_segmenter() wraps namedivider’s BasicNameDivider, which reads data bundled in the installed package; ja_segmenter(gbdt=True) selects its gradient-boosted divider instead, which is more accurate and downloads its model and surname files from the network on first use, worth knowing before deploying it somewhere sandboxed or air-gapped.

Forgetting the segmenter is loud: building the parser emits a UserWarning naming the scripts that could never divide and the call to pass, because the misconfigured parser would otherwise behave exactly like a working one minus the feature. The same check guards any configuration whose activated scripts nothing can serve. A from-scratch lexicon with no hangul surnames warns under the default policy, and Policy(segment_scripts=frozenset()) is the deactivation the message offers.

Decomposed text

Korean and Japanese text is sometimes stored decomposed, with a syllable held as its separate jamo rather than as one codepoint. macOS filenames are the common source. Everything above works on decomposed input: script classification normalizes to NFC before deciding, so a decomposed name gets the same order rule as its composed twin. Vocabulary lookup does the same before matching a word against titles, honorifics and the rest, so a decomposed Señor or née — macOS-origin data again — is recognized as readily as its composed spelling. Case repair reads a decomposed accent as part of its letter, so capitalized() gives a decomposed josé garcía the same José García as the composed spelling, still decomposed.

Splitting is the exception. An unspaced decomposed hangul name is ordered correctly but not split, because surname matching runs against the text as given and a decomposed name matches no entry in the census list. That is the deliberate choice: being unsplit is recoverable, whereas splitting in the wrong place is not.

One consequence is worth stating outright, because it looks like a bug. Parse output preserves the encoding it was given, so a field from a decomposed name is decomposed too, and comparing it against a composed literal fails even where the parse was correct:

decomposed = unicodedata.normalize("NFD", "김 민준")
parse(decomposed).family == "김"                          # False
unicodedata.normalize("NFC", parse(decomposed).family)    # '김'

Normalize both sides before comparing across encodings.

Boundaries

Several boundaries apply to all of the above.

When the script rules don’t apply

Romanized names (“Kim Min-jun”, “Yamada Taro”) are Latin script and follow the ordinary positional rules. Order genuinely varies in romanized data, so nothing script-based applies. A name written wholly in katakana stays positional for the reason given above, pack or no pack: it may be a transcription, already in the order it should be read in, or a Japanese name’s reading written family-first — a furigana field, or legacy halfwidth data — and the script cannot tell the two apart. If you know your katakana names are Japanese, map katakana to family-first yourself:

>>> from nameparser import (DEFAULT_SCRIPT_ORDERS, FAMILY_FIRST, Parser,
...                         Policy, Script)
>>> parse("ヤマダ タロウ").family
'タロウ'
>>> kana_family_first = Parser(policy=Policy(script_orders=(
...     *DEFAULT_SCRIPT_ORDERS, (Script.KATAKANA, FAMILY_FIRST))))
>>> kana_family_first.parse("ヤマダ タロウ").family
'ヤマダ'

This reaches full-width katakana too, transcriptions included (マイケル・ジャクソン would read family マイケル), which is why it is yours to choose rather than a default. name_order=FAMILY_FIRST would also work, but it reverses every name, Latin ones included.

A Han transcription written with a space instead of the 间隔号 (威廉 莎士比亚) carries nothing to distinguish it from a native two-token name, and keeps the family-first reading. The dot is the marker, and without it there is no signal. The same holds for a transcription typed with the Japanese middle dot (威廉・莎士比亚): each dot carries its own script’s convention, the nakaguro’s is Japanese roster formatting rather than transcription, so only the Chinese dot stands the family-first reading down.

Honorifics come off first

Honorifics and degrees follow a CJK name, and both the spaced and the glued forms are recognized as suffixes:

>>> parse("王小明 先生").suffix
'先生'
>>> glued = parse("김민준씨")
>>> glued.family, glued.given, glued.suffix
('김', '민준', '씨')

The honorific is split off the end of the last token of the name before the name is split or ordered, so those rules see the name without it. That is why 김민준씨 still divides into family 김 and given 민준, and why a configured Japanese segmenter is handed 山田太郎 rather than 山田太郎様.

The stop can be any width: the fullwidth . a Japanese or Chinese input method produces by default, the ideographic 。 and the halfwidth 。 all reach the honorific vocabulary as an ASCII period does, so 김민준 씨. gives suffix 씨., family 김, given 민준. A period glued to an ordinary name word, not an honorific, is likewise left where it was written rather than breaking the segmentation that follows it:

>>> parse("양. 지훈").family, parse("양. 지훈").given
('양.', '지훈')

Commas and Latin wrappers around a CJK name

A comma or a Latin credential set wrapped around a CJK name is tolerated input rather than contract: no native CJK writing uses either convention, so nameparser reads it best-effort and the handling can change without notice. Today:

  • A comma names the family and stops the split. 남궁민수, 지훈 reads family 남궁민수 whole, where the bare 남궁민수 alone would split into family 남궁 and given 민수.

  • A glued honorific peels off before the comma when what follows it is nothing but suffix words, as in 田中さん, PhD (suffix さん, PhD).

  • It stays glued when the comma is followed by a title or another name word, as in 田中さん, Dr. (family 田中さん).

  • Credentials after the comma land in given or suffix by spelling, and 田中さん, V. and 田中さん, Ph. D. give up さん exactly as 田中さん, PhD does. Policy(lenient_comma_suffixes=False) reads V. as name text instead, and さん then stays glued (family 田中さん, given V.).

Spacing, and where the name divides

Where a segmenter divides the name, the two spellings part company. A spaced honorific is a token boundary the writer typed, and the segmenter is asked only where an undivided name divides, so anything standing beside the name calls it off, honorific or not. Under the Japanese pack 佐藤 氏 keeps family 佐藤, where bare 佐藤 would have been divided 佐 + 藤.

That is a conservative reading rather than a principled one. A spaced honorific cannot be told apart from a spaced given name by position, and treating it as one keeps four real surnames whole (佐藤 氏, 田中 様, 鈴木 先生, 中村 教授) at the price of one division it declines to make (山田太郎 様). A glued honorific carries no boundary at all, its writer having drawn none anywhere, so 田中さん divides the way bare 田中 does, into family 田 and given 中, with さん in suffix. Writing the honorific spaced is therefore also how you ask for a family name to be kept whole on this path, short of declining the pack.

Where the VOCABULARY divides the name the two spellings never part company, because the peel hands the same remainder to the same surname match either way:

>>> parse("김민준 씨").family, parse("김민준씨").family
('김', '김')

Under the Chinese pack 王小明 先生 and 王小明先生 both give family 王, given 小明 and suffix 先生 for the same reason. Spacing the honorific is no lever there, and for Korean data there is no pack to decline either, hangul segmentation being on by default.

Which honorifics peel when glued

The glued reading is deliberately narrower than the spaced one. A spaced honorific sits behind a boundary its writer drew; a glued one has only itself to go on, so a word peels off the end of a name only if it could never BE the end of a name.

씨, 님, 박사, 박사님, 선생님, 교수님, さん, さま, くん, ちゃん, 様, 先生, 教授, 女士 and 小姐 qualify. 양, 군, 氏, 博士 and 殿 do not, because 김지양 is a given name, 田中博士 is Tanaka Hiroshi as readily as Doctor Tanaka, and some ninety Japanese surnames end in 殿 (鵜殿, 真殿). Those stay recognized in their spaced form, where position settles what the glued form leaves ambiguous. 君 is recognized in neither form, since 王君 is a complete Chinese name, though its kana spelling くん peels.

Exactly one honorific peels off a token, and the entries are whole honorifics rather than parts: 김민준박사님 gives up 박사님 entire, not 님 with 박사 left behind.

When a division was a judgment call

A division the parser had to choose is reported rather than hidden. When an unspaced name has more than one vocabulary-supported split, the longest surname wins and the parse records the decision as an AmbiguityKind.SEGMENTATION, described under When the parser had to guess. 남궁민수 is 남궁 + 민수 by the two-syllable surname but 남 + 궁민수 by the single-syllable one:

>>> parse("남궁민수").ambiguities
(Ambiguity('segmentation': '남궁'/'민수'),)

A name with only one possible split reports nothing.

A segmenter’s answer is reported on the same kind whenever its confidence falls below the stage’s floor, naming the division and the score in the report’s detail:

"'山田太郎' splits as '山田' + '太郎' on a segmenter answer scoring 0.44, under the 0.9 confidence floor"

With namedivider that line separates its two kinds of answer. A division it states as a rule, such as the kanji-to-kana boundary in 高橋みなみ, scores 1.0 and reports nothing; a division read off kanji statistics scores far below the floor and always reports. Read the report as a statement about the kind of answer, not as a measure of how likely this particular one is to be wrong.

A lone two-character kanji name divides one character to each side, on namedivider’s rule for that length. A name that short carries no evidence of where its own boundary falls, and one character each way is the presumption Japanese practice makes. That is an accepted presumption rather than a measurement, so a two-character token is the shape to check first if a division looks wrong.

When a segmenter misbehaves

A segmenter that answers outside the token it was given, with a cut at or past the end of the text, has violated the protocol, and the parse says so rather than hiding it. It raises ValueError naming the offending offset and the token’s length, the way an answer of the wrong type raises TypeError. Declining silently would leave an off-by-one segmenter invisible: every answer it gave would vanish and the name would merely look undivided.

The shipped ja_segmenter() does decline, returning None and leaving the token whole, for the cases that are not protocol violations: text outside the Japanese repertoire, text too short to divide, an answer that fails to reconstruct its input, and a score outside [0, 1].

Exceptions are the one thing that does not stay inside the parse: a segmenter is your code, so its errors propagate rather than being absorbed as content errors.

The command line takes the pack but not the segmenter: python -m nameparser --locale ja has no way to attach one, so it activates nothing by itself. Customizing the parser covers turning any of these behaviors off.

Aggregate views

>>> name.given_names          # given + middle
'Juan Q. Xavier'
>>> name.family_base, name.family_particles   # family, split apart
('Vega', 'de la')

surnames (middle + family) is the mirror-image aggregate. The plural is the tell: given and family are single fields, while given_names and surnames roll several fields together — the same sense in which a passport form asks for your “given names” as one blank that can hold more than one word.

family_base is the one to sort on. Dutch and Belgian directories file a name under the base surname and ignore the tussenvoegsel — the van, de or van der in front of it — and the same convention applies wherever particles are common:

>>> names = [parse(s) for s in
...          ["Vincent van Gogh", "Juan de la Vega", "John Smith"]]
>>> [n.family_base for n in sorted(names, key=lambda n: n.family_base.lower())]
['Gogh', 'Smith', 'Vega']

Sorting on family instead files those under “d”, “S” and “v”, which is the problem the split exists to solve:

>>> [n.family for n in sorted(names, key=lambda n: n.family.lower())]
['de la Vega', 'Smith', 'van Gogh']

Dicts and strings

>>> name.as_dict(include_empty=False)
{'title': 'Dr.', 'given': 'Juan', 'middle': 'Q. Xavier', 'family': 'de la Vega', 'suffix': 'III'}
>>> str(name)
'Dr. Juan Q. Xavier de la Vega III'
>>> name.render("{family}, {given}")
'de la Vega, Juan'
>>> name.initials()
'J. Q. X. V.'

Fixing case

>>> str(parse("juan de la vega").capitalized())
'Juan de la Vega'

The no-argument form above uses the DEFAULT lexicon; for a name parsed with a custom Parser, call Parser.capitalized() so the parser’s own vocabulary decides the exceptions.

Nicknames and maiden names

>>> parse("Jonathan 'Jack' Kennedy").nickname
'Jack'
>>> parse("Jane Smith née Jones").maiden
'Jones'

Both fields appear in the default str() rendering — the nickname quoted after the given name, the maiden name parenthesized after the family name:

>>> str(parse("Jane (Janie) Smith née Jones"))
'Jane "Janie" Smith (Jones)'

Surviving a reparse

The default str() rendering is built for display, and the two fields differ in what survives it. The quoted nickname reparses as a nickname; the parenthesized maiden name reparses as a nickname too, so a parse-render-reparse round trip silently loses it:

>>> parse('Jane "Janie" Smith (Jones)').maiden
''
>>> parse('Jane "Janie" Smith (Jones)').nickname
'Janie Jones'

If the rendered string has to survive a reparse — storing names as text and reading them back, for instance — render the marker explicitly instead of relying on the default:

>>> name = parse("Jane Smith née Jones")
>>> text = name.render("{given} {family} née {maiden}")
>>> text
'Jane Smith née Jones'
>>> parse(text).maiden
'Jones'

What bracketed content reads as

Delimited content is not always a nickname. If it opens with a marker word and has a word after it, the clause is a maiden name inside any configured delimiter pair, with nothing else configured — the clause has said which convention it means, so you do not have to:

>>> parse("Jane Smith (née Jones)").maiden
'Jones'
>>> parse('Jane Smith "née Jones"').maiden
'Jones'
>>> parse("Jane (née Jones) Smith").family
'Smith'

A marker with no name after it is just a word in brackets, and a clause with no marker at all stays a nickname — the parenthesized birth surname is a real convention, but nothing in the clause says so, and only you can declare that with maiden_delimiters (see Nicknames, maiden names, and brackets):

>>> parse("Jane Smith (née)").nickname
'née'
>>> parse("Cherice J. (Johnson) Williams").nickname
'Johnson'

If what’s inside is a known suffix, or simply ends in a period, it is read as a suffix instead — that reading is taken before the maiden one, and parenthesized credentials and retired ranks are far more common than parenthesized nicknames that happen to be credentials:

>>> parse("Andrew Perkins (MBA)").suffix
'MBA'
>>> parse("Andrew Perkins (Ret.)").suffix
'Ret.'

The exception to that exception is an ambiguous acronym — one that is also a plausible name or set of initials. Standing alone inside delimiters, it keeps the nickname reading:

>>> parse("JEFFREY (JD) BRICKEN").nickname
'JD'

JD is in suffix_acronyms_ambiguous; see Words that are also ordinary names for what that field marks and how to add to it. Nicknames, maiden names, and brackets covers the full order these readings are tried in, and how to configure the delimiter pairs.

Titles you didn’t configure

A trailing period marks an abbreviation, and the parts of a name that get abbreviated are the ones standing outside it — titles and post-nominals. So a word ending in a period, in no vocabulary list at all, is read as a title when it stands where a title stands: at the front of the part that carries the given name. That is what lets unfamiliar ranks, honorifics and abbreviations work without configuring anything:

>>> parse("Insp. Jane Morse").title
'Insp.'
>>> parse("Det. Insp. Jane Morse").title     # chains
'Det. Insp.'

Neither det nor insp is in the shipped vocabulary; the periods are doing all the work.

The part that carries the given name is not always the front of the string. After a family comma it is the part after the comma, and the rule applies there in exactly the same way:

>>> parse("Morse, Det. Insp. Jane").title
'Det. Insp.'

Where the inference stops

The rule is bounded in four ways, so it doesn’t swallow ordinary names:

  • Single initials are left alone (J.).

  • Abbreviations with interior periods are left alone (E.T.).

  • Only the leading run is read this way: the same word after the given name is a middle name.

  • A word carrying a Han, kana or hangul character is a name word even with a period, since those scripts write no abbreviation with one ("田中. 太郎" has family 田中.).

>>> parse("J. Smith").given
'J.'
>>> parse("E.T. Jones").given
'E.T.'
>>> parse("Jane Insp. Morse").middle
'Insp.'

Because this is structural rather than vocabulary-driven, emptying titles does not switch it off; see Turning title detection off.

At the end of a name

Only the leading slot INFERS. At the back of a name a period-marked word is read by vocabulary alone: a word the parser already knows as a post-nominal or as a title reads as one, and an unfamiliar abbreviation stays a name word.

>>> parse("John Smith Prof.").title
'Prof.'
>>> parse("John Smith Xyz.").family
'Xyz.'

That asymmetry is deliberate. The title vocabulary holds hundreds of words that are ordinary surnames — king, judge, bishop — so inferring a trailing title from the shape alone would cost real names their family field. The period is what separates the two: a trailing title word written without one is a name word, so parse("Mary Jane King").family is 'King'.

Comparing names

== is strict value equality — two ParsedName instances are equal only if every field matches exactly. For “is this the same name, allowing for order and case?” use matches() or comparison_key() instead. matches() parses a str argument with the DEFAULT parser; to compare against a string using a custom Parser’s vocabulary, call Parser.matches().

>>> parse("de la Vega, Juan").matches("Juan de la Vega")
True
>>> parse("JUAN DE LA VEGA").comparison_key() == parse("Juan de la Vega").comparison_key()
True

When the parser had to guess

Some names have no single correct reading. A leading Van could be a given name — it really is for the actor Van Johnson — or the start of a family name, as it is for President Van Buren. Both are the same shape, so no rule can tell them apart. 2.0 takes the more likely reading and records the choice on ambiguities rather than deciding silently:

>>> name = parse("Van Buren")
>>> name.given, name.family
('Van', 'Buren')
>>> for a in name.ambiguities:
...     print(a.kind.value, "-", a.detail)
particle-or-given - leading 'Van' may be a family-name particle; read as a given name

AmbiguityKind members are their own string values, so branching on a kind needs no import:

>>> name.ambiguities[0].kind == "particle-or-given"
True
>>> [t.text for t in name.ambiguities[0].tokens]
['Van']

Credentials that are also surnames

The post-nominals that double as ordinary surnames report an ambiguity the same way. Which reading a bare one gets depends on what the writing says:

  • Capitals in a name written in more than one case lean to the credential, even with no words to spare (Jack MA).

  • Any other cased spelling that is not wholly lowercase leans to the surname, even with words to spare (Jack Ma, John Smith Ma).

  • No signal — all lowercase, or a name written wholly in one case — leaves it to the count: the credential with two or more words before it, not counting a title or a nickname, the surname otherwise.

Either way the choice is recorded:

>>> parse("John Smith MA").suffix
'MA'
>>> [a.kind.value for a in parse("John Smith MA").ambiguities]
['suffix-or-name']
>>> parse("Jack MA").suffix
'MA'
>>> parse("Jack Ma").family
'Ma'
>>> parse("John Smith Ma").family
'Ma'

Words that are also ordinary names has the full rules, including how a comma changes them.

Two Policy switches extend the same class to words the vocabulary does not hold. unlisted_dotted_suffixes (on by default) reads a token of two or more period-separated chunks the same way — parse("John Smith X.Y.Z.").suffix is 'X.Y.Z.' — and unlisted_caps_suffixes does the same for an unlisted all-caps word, by default only after a comma behind a full name, where an all-caps surname is never written — parse("John Smith, XYZ").suffix is 'XYZ' — and at the end of a name only on request, since an all-caps surname is written there. See Credentials the vocabulary doesn’t list for both.

A reading the vocabulary settles on its own is not a guess and reports nothing — periods make M.A. unambiguously a credential:

>>> parse("John Smith M.A.").ambiguities
()

Most names report none. A non-empty ambiguities is a useful signal for routing a record to human review instead of trusting it silently.

Tokens and spans

The seven fields are the convenient view. Underneath, each one is backed by tokens carrying their exact offsets into the original string, so you can always get back to the text a field came from:

>>> name = parse("Juan de la Vega")
>>> for tok in name.tokens:
...     print(f"{tok.text!r:8} {tok.role.value:8} {tuple(tok.span)}")
'Juan'   given    (0, 4)
'de'     family   (5, 7)
'la'     family   (8, 10)
'Vega'   family   (11, 15)
>>> tok = name.tokens[1]
>>> name.original[tok.span.start:tok.span.end]
'de'

tokens_for() narrows that to a single role:

>>> from nameparser import Role
>>> [t.text for t in name.tokens_for(Role.FAMILY)]
['de', 'la', 'Vega']

A role’s string name works too: name.tokens_for("family").

Correcting a parse

ParsedName is immutable, so a correction is a new value. There are two ways to make one: replace() splices the new text in as written, and Parser.revise() reads it the way a parse would.

Splicing a value in with replace()

replace() returns a copy with the given fields changed. Untouched fields keep their tokens (and original is preserved), with one deliberate exception: an ambiguity that pointed into a replaced field is dropped — correcting the field that was flagged clears the flag, while correcting an unrelated field keeps it.

>>> name = parse("Juan de la Vega")
>>> corrected = name.replace(title="Dr.")
>>> str(corrected)
'Dr. Juan de la Vega'
>>> name.title
''
>>> flagged = parse("Van Buren")
>>> flagged.replace(given="Martin").ambiguities
()
>>> [a.kind.value for a in flagged.replace(family="Harrison").ambiguities]
['particle-or-given']

What a spliced value loses

replace() splits values on whitespace into plain, untagged tokens — the vocabulary knowledge a parse would have about the new text is not there. A token the parse never saw carries no decision to honor, so each view falls back on what it can answer without one:

  • capitalized() is handed a vocabulary, so it asks that whether a word is a conjunction or an initial, which a word answers on its own. Its other special cases never needed a reading: an exceptions-map mask applies wherever its word stands, and a listed credential acronym or a roman numeral is written in capitals wherever the field is the suffix, so replace(suffix="mba") repairs to MBA and suffix="vi" to VI. A credential recognised only by its dotted shape is the exception, since the shape is something the parse records. Particles are a second exception: whether a particle is acting as one is a fact about the whole part, and no word of a spliced field carries a reading to derive it from, so a family set to de la stays lowercase where the same words parsed are repaired to De La.

  • initials() takes no vocabulary at all, so every word of a spliced field contributes an initial.

  • family_particles and family_base are properties on the parsed name, which holds no vocabulary of its own either: family_particles empties and family_base takes the whole field.

Of the parsed name’s own views, capitalized() is the only one handed a vocabulary. (The v1 HumanName facade’s initials() is handed one too, and asks it whether a spliced word is a conjunction or a particle, so a spliced de la gives no initials there; it is not a method of the parsed name.)

>>> name.family_particles
'de la'
>>> replaced = name.replace(family="de la Vega Smith")
>>> replaced.family_particles
''
>>> replaced.family_base
'de la Vega Smith'
>>> replaced.initials()
'J. d. l. V. S.'
>>> name.replace(family="de la").capitalized(force=True).family
'de la'

Reading a value with Parser.revise()

Parser.revise() is the same operation with each value classified by the parser’s vocabulary, so the correction behaves like a fresh parse of the corrected name (which also means delimiters and marker words in the value are consumed as they would be in a parse):

>>> from nameparser import Parser
>>> parser = Parser()
>>> revised = parser.revise(name, family="de la Vega Smith")
>>> revised.family_particles
'de la'
>>> revised.initials()
'J. V. S.'

revise() has two siblings on Parser: Parser.matches() and Parser.capitalized(). Those two matter when you have built a custom parser — the ParsedName methods of the same names fall back to the default configuration for str or omitted arguments.

Command line

python -m nameparser parses one name and prints what it found. It is the quickest way to see how a particular string comes out — worth reaching for when a name parsed unexpectedly and you want to check a variation of it, or when trying a locale pack before wiring one into code:

$ python -m nameparser "dr. juan de la vega iii"
Parsed:
<ParsedName: [
    title: 'dr.'
    given: 'juan'
    family: 'de la vega'
    suffix: 'iii'
]>
Capitalized:
<ParsedName: [
    title: 'Dr.'
    given: 'Juan'
    family: 'de la Vega'
    suffix: 'III'
]>
Initials: j. v.

--json prints the fields as a single line instead, which is the form to pipe somewhere:

$ python -m nameparser --json "Doe, John"
{"title": "", "given": "John", "middle": "", "family": "Doe", "suffix": "", "nickname": "", "maiden": ""}

Add --locale to parse with a locale pack (for example --locale ru); see Locale packs.

Where next