Skip to content
Merged
Show file tree
Hide file tree
Changes from 1 commit
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Prev Previous commit
Next Next commit
fix(S2): #564 -- the contrast reads a word's last LETTER, not character
The previous commit tested the last character, so a name typed with
decomposed accents ending in an accented letter lost the contrast
('José André, XYZ', 'Lê Thị Hà, XYZ', 'René Noé, XYZ' read given where
their composed spellings read suffix), as did "Jones'", 'Smith2' and a
glued 'Smith)'. written_as_a_name now composes the word (NFC), passes
over trailing non-letters, and passes over a lowercase letter with no
single capital form (ß, ĸ), which generalizes the ß rule. One call, no
generator.

Measured against the previous commit: the four decomposed-accent names
and 'JOHN Smith)' regain the contrast and the ĸ record keeps its given
name; nothing else moves, and OFF equals master on 4292 inputs.

Also from the review: rules.md#C1 states the last-letter reading and
pins 'MÜLLER WEIß, HANS'; decisions.md#S2 records the true blast radius
of the previous commit (seven prefixes plus ß and Džokić, and
'JOHN O'NEILL's' / 'JEAN-pierre DUPONT' / 'MARY-kate OLSEN', now
accepted costs) with a recompute recipe; the cost prose says what it
is, about four frames per word before the comma and constant in a
word's letters; the fix(#564) ledger comments drop the wording
decisions.md calls wrong.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
  • Loading branch information
derek73 and claude committed Oct 2, 2026
commit 07c5662727954cf4ae5db1d9e0e1987e2d841df0
2 changes: 1 addition & 1 deletion docs/customize.rst
Original file line number Diff line number Diff line change
Expand Up @@ -532,7 +532,7 @@ listed below.
- ``CapsSuffixes``
- Where an unlisted all-caps word of two or more letters, with no
period in it, reads as a credential. The name must contrast it
with a word holding a capital and ending in a lowercase letter
with a word holding a capital whose last letter is lowercase
(``Smith``, ``DiCaprio``) that the vocabulary
does not claim as a title, particle or credential: a record
written wholly in capitals, or wholly in lowercase, keeps every
Expand Down
4 changes: 2 additions & 2 deletions docs/design/decisions.md

Large diffs are not rendered by default.

5 changes: 4 additions & 1 deletion docs/design/rules.md
Original file line number Diff line number Diff line change
Expand Up @@ -1726,7 +1726,9 @@ C1. Rationale: a credential run after the comma means the name is in
the class in such a part only where the name carries the
contrast: one of its own words before the comma written the way a
name is written in mixed case, holding a capital and ending in a
lowercase letter, with no period and not claimed by the vocabulary
lowercase letter (its last letter, past any trailing mark, and
setting aside a lowercase letter with no capital of its own, as ß
has none), with no period and not claimed by the vocabulary
as a title, particle, connective, credential or generation. A
surname written in capitals ends in a capital whatever is glued in
front of it, and a word written wholly in lowercase holds none, so
Expand Down Expand Up @@ -1874,6 +1876,7 @@ C1. Rationale: a credential run after the comma means the name is in
"John Smith, LEED AP" → suffix="LEED AP"
"Smith, XYZ" → given="XYZ" · boundary
"García Márquez, MJ" → given="MJ" · boundary
"MÜLLER WEIß, HANS" → given="HANS" · boundary
"García Márquez, MJ PhD" → given="MJ" · boundary
"García Márquez, MJ JK" → suffix="MJ JK"
"John Smith, PhD XYZ" → suffix="PhD XYZ"
Expand Down
2 changes: 1 addition & 1 deletion docs/release_log.rst
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,7 @@ Release Log

- **New Policy field unlisted_dotted_suffixes, on by default: a dotted acronym nobody has listed is read by position.** ``HumanName("John Smith X.Y.Z.")`` gives suffix ``X.Y.Z.`` where every release gave last ``X.Y.Z.``, while ``Jack X.Y.Z.`` keeps its surname, the same words-to-spare rule a listed acronym takes -- and both readings are reported. After a comma the count is of the words before it, and two dotted single letters are the exception: they are how a person's initials are written, and two words before a comma may be one surname, so ``García Márquez, G.J.`` keeps first ``G.J.`` and last ``García Márquez`` and reports the fork, unless an unambiguous post-nominal in front of the initials that is not also a title, or another unlisted dotted word beside them, says otherwise (``John Smith, PhD X.Y.`` gives suffix ``PhD X.Y.``, while ``García Márquez, Ms G.J.`` keeps title ``Ms``, first ``G.J.``). Three letters or more read by the count, so ``John Smith, X.Y.Z.`` gives suffix ``X.Y.Z.`` -- and so does ``García Márquez, G.J.R.``, the accepted cost of the line, since initials are conventionally written apart (``García Márquez, G. J. R.``), as separate words this rule does not read (#563). Case is irrelevant here: the periods are the signal, so ``john smith x.y.z.`` reads the same way. Words the vocabulary does know are untouched (``M.A.``, ``Ph.D.``, ``A.B.C.``), a single trailing period is still not this shape (``John Smith Xyz.`` keeps last ``Xyz.``), and a dotted run at the FRONT of a name is untouched (``J.R.R. Tolkien``). One accident retires with it: a dotted word whose only vocabulary matches were SINGLE ASCII CHARACTERS -- the roman numerals the suffix list holds, and the lone digit ``2`` -- was reading as a generational suffix, so ``Jack X.Y.I.`` gives last ``X.Y.I.`` again, as 1.4.0 read it, while ``Msc.Ed.``, ``JD.CPA`` and ``Lt.Gov.`` are unchanged. The digit is why a dotted VERSION STRING moves with them and moves SILENTLY: ``John Smith 1.4.2`` gives last ``1.4.2`` where 2.3 gave suffix ``1.4.2``, and ``John Smith, 1.4.2`` gives first ``1.4.2``, last ``John Smith``. Such a token reports nothing at any policy -- it is no acronym either, the shape reading wanting every chunk alphabetic -- and a version string read as a credential was the same accident this retirement removes. That retirement is NOT behind this switch and stands either way -- setting it to ``False`` reads an unlisted dotted word as name material by position instead (``John Smith X.Y.Z.`` keeps last ``X.Y.Z.``), the pre-2.4 reading for THAT half alone. See the ``S2`` and ``suffix-acronym-collisions`` entries of ``docs/design/decisions.md`` (closes #516)

- **New Policy field unlisted_caps_suffixes: an unlisted all-caps word reads as a credential after a comma by default, and elsewhere on request.** Its value is a ``CapsSuffixes``. The default, ``CapsSuffixes.AFTER_COMMA``, reads such a word in the part right after a comma behind two or more name words, alone or in a run with other credentials: ``HumanName("John Smith, XYZ")`` gives first ``John``, last ``Smith``, suffix ``XYZ``, where 1.4.0 through 2.3.0 gave first ``XYZ``, last ``John Smith``; ``John Smith, LEED AP`` and ``John Smith, PhD XYZ`` give suffix ``LEED AP`` and ``PhD XYZ`` the same way, and ``John Smith, RAI`` gives suffix ``RAI`` again, as it did before 2.3. The all-caps surname convention writes the capitals at the end of a name or before a comma (``Jean DUPONT``, ``DUPONT, Jean``) and never there. A word after a one-word surname stays the given name (``Smith, XYZ``), a two-letter word reads exactly as dotted initials do (``García Márquez, MJ`` and ``García Márquez, MJ PhD`` keep first ``MJ``), and the name has to contrast the capitals with a word of its own holding a capital and ending in a lowercase letter (``Smith``, ``DiCaprio``). A surname written in capitals ends in a capital whatever is glued in front of it, so such a record keeps its given name beside its lowercase particles, titles and maiden clauses and beside a mixed-case credential (``GISCARD d'ESTAING, VALÉRY``, ``LLOYD FitzGERALD, RONALD``, ``LLOYD WEBBER, ANDREW PhD``), as does a name written wholly in lowercase. ``CapsSuffixes.EVERYWHERE`` also reads the end of a name, the given part's last word after a family comma and the word ending a maiden marker's clause: ``.parse("John Smith XYZ")`` gives suffix ``XYZ``, and ``Jean Pierre DUPONT`` gives last ``Pierre``, suffix ``DUPONT`` -- why it is not the default. ``CapsSuffixes.OFF`` reads none of them and reports nothing; it is the way to keep a given name written in capitals after a two-word surname, which the default reads as a credential (``García Márquez, GABRIEL`` gives suffix ``GABRIEL``). The field reaches the core parser only, through ``Parser(policy=Policy(unlisted_caps_suffixes=...))``; a ``HumanName`` tracks the parser's defaults, so the comma reading reaches it and the other two settings cannot be chosen from there. Neither this field nor ``unlisted_dotted_suffixes`` has a v1 ``Constants`` manager. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` (closes #516, closes #564)
- **New Policy field unlisted_caps_suffixes: an unlisted all-caps word reads as a credential after a comma by default, and elsewhere on request.** Its value is a ``CapsSuffixes``. The default, ``CapsSuffixes.AFTER_COMMA``, reads such a word in the part right after a comma behind two or more name words, alone or in a run with other credentials: ``HumanName("John Smith, XYZ")`` gives first ``John``, last ``Smith``, suffix ``XYZ``, where 1.4.0 through 2.3.0 gave first ``XYZ``, last ``John Smith``; ``John Smith, LEED AP`` and ``John Smith, PhD XYZ`` give suffix ``LEED AP`` and ``PhD XYZ`` the same way, and ``John Smith, RAI`` gives suffix ``RAI`` again, as it did before 2.3. The all-caps surname convention writes the capitals at the end of a name or before a comma (``Jean DUPONT``, ``DUPONT, Jean``) and never there. A word after a one-word surname stays the given name (``Smith, XYZ``), a two-letter word reads exactly as dotted initials do (``García Márquez, MJ`` and ``García Márquez, MJ PhD`` keep first ``MJ``), and the name has to contrast the capitals with a word of its own holding a capital whose last letter is lowercase (``Smith``, ``DiCaprio``); a name typed with decomposed accents reads as its composed spelling. A surname written in capitals ends in a capital whatever is glued in front of it, so such a record keeps its given name beside its lowercase particles, titles and maiden clauses and beside a mixed-case credential (``GISCARD d'ESTAING, VALÉRY``, ``LLOYD FitzGERALD, RONALD``, ``LLOYD WEBBER, ANDREW PhD``), as does a name written wholly in lowercase. ``CapsSuffixes.EVERYWHERE`` also reads the end of a name, the given part's last word after a family comma and the word ending a maiden marker's clause: ``.parse("John Smith XYZ")`` gives suffix ``XYZ``, and ``Jean Pierre DUPONT`` gives last ``Pierre``, suffix ``DUPONT`` -- why it is not the default. ``CapsSuffixes.OFF`` reads none of them and reports nothing; it is the way to keep a given name written in capitals after a two-word surname, which the default reads as a credential (``García Márquez, GABRIEL`` gives suffix ``GABRIEL``). The field reaches the core parser only, through ``Parser(policy=Policy(unlisted_caps_suffixes=...))``; a ``HumanName`` tracks the parser's defaults, so the comma reading reaches it and the other two settings cannot be chosen from there. Neither this field nor ``unlisted_dotted_suffixes`` has a v1 ``Constants`` manager. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` (closes #516, closes #564)

- **The comma's own decision about an ambiguous credential is now reported.** ``parse("Smith, MA").ambiguities`` names ``suffix-or-name``, and so does every other decision at the ambiguous credential class -- before or after a comma, in either direction, with no new ``AmbiguityKind`` (the family-comma attachment fork already reported this way, e.g. ``parse("Berg, Jan vd")``). A flip of the comma in which no listed ambiguous acronym takes part is the exception and is made in silence: ``John Smith, X.Y.Z.`` and ``John Smith, PhD X.Y.`` report nothing, the only such word a reader takes for a name being a pair of initials, which the comma reads as the given name unless something beside it has already said otherwise. Two pairs speaking only for each other still make the credential run, and that flip reports: ``John Smith, X.Y. P.Q.`` gives last ``Smith``, suffix ``X.Y. P.Q.`` (#563). One report per decision: ``Smith, Ma`` reports that the word was kept as the given name just as ``Smith, MA`` reports that it was taken as a credential. The reading a SURNAME PARTICLE swallows is reported too, which no release before this one did: ``John van der Berg Ma`` gives last ``van der Berg Ma`` and names ``suffix-or-name``, where the chain took a word the credential reading had considered. ONE report goes away, because a comma segment the parser reads as a credential run is no longer called unrecognized: ``Steven Hardman, MD, DO, DDS`` no longer reports ``comma-structure``, on its written case. That is the whole of the losses over the differential corpora -- ``John Smith, MD, R.A.I.`` is quieted on its shape by the same change, but it never reported at 2.3.0 either, having only carried the flag inside this release's own development. The other movement an upgrader sees is a SWAP rather than a loss: ``Jack X.Y.I.`` reported ``given-or-family`` at 2.3.0 and reports ``suffix-or-name`` here, the dotted retirement above having handed it to the ambiguous class. Everything else at this class is a GAIN, which is what the rest of this bullet describes. Two slots this bullet left silent no longer are, and the two bullets below close them: a credential trailing the GIVEN part of a family-comma listing now reads as a credential and reports either way, and so does one ending a maiden marker's clause. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md``

Expand Down
9 changes: 5 additions & 4 deletions nameparser/_pipeline/_segment.py
Original file line number Diff line number Diff line change
Expand Up @@ -140,10 +140,11 @@ def texts(seg: tuple[int, ...]) -> list[str]:
# al-ASSAD', 'LLOYD WEBBER née Smith') and beside a mixed-case
# credential ('LLOYD WEBBER, ANDREW PhD').
# Asked only once a caps word is in hand: a part with no lowercase
# at all is settled in one C-level comparison; past that the walk
# costs a frame a word before the comma (the generator), plus a
# call, a fold and a wordlist test only where the C-level case
# test passes.
# at all is settled in one C-level comparison; past that it is
# linear in the words before the comma, about four frames a word
# (the generator, the case test's call, and the own-words walk's
# fold and marker test), plus a fold and a wordlist test where the
# case test passes -- and constant in a word's letters.
def name_contrast() -> bool:
before = "".join([state.tokens[i].text for i in groups[0]])
if before == before.upper():
Expand Down
30 changes: 25 additions & 5 deletions nameparser/_pipeline/_vocab.py
Original file line number Diff line number Diff line change
Expand Up @@ -722,16 +722,36 @@ def in_any_wordlist(n: str, lexicon: Lexicon) -> bool:

def written_as_a_name(text: str) -> bool:
"""Whether TEXT is written the way a name is written in mixed case:
it holds a capital (or titlecase letter) and its last letter is
it holds a capital (or titlecase letter) and its last LETTER is
lowercase -- #564's test for the name's case contrast (Derek).
'Smith', 'DiCaprio', 'IJzerman', 'al-Rashid', "d'Estaing",
'McDonald', 'Džokić' pass. A surname written in capitals fails
whatever is glued in front of it ("d'ESTAING", 'al-ASSAD',
'McDONALD', 'FitzGERALD', 'DeVITO', 'St-PIERRE'), and so do a lone
capital and a lowercase-only word. A trailing 'ß' is set aside,
having no single capital form: 'WEIß' is written in capitals and
'Weiß' is not. Two C-level checks, so a word costs no frame."""
return text != text.lower() and text.rstrip("ß")[-1:].islower()
capital and a lowercase-only word.

The last letter is found, not the last character: the word is
composed first (NFC), so a name typed with decomposed accents reads
as its composed spelling ('André' ends in 'é', not in the combining
accent), and trailing non-letters are passed over ("Jones'",
'Smith2', 'Smith)'). A lowercase letter with no single capital form
is passed over too, being no evidence of case: 'WEIß' and 'KAĸ' are
written in capitals, 'Weiß' is not. One call and no generator, so a
word costs one frame."""
if text == text.lower():
return False
word = unicodedata.normalize("NFC", text)
i = len(word)
while i:
i -= 1
ch = word[i]
if not ch.isalpha():
continue
upper = ch.upper()
if ch.islower() and (upper == ch or len(upper) != 1):
continue
return ch.islower()
return False


def claimed_as_non_name(n: str, lexicon: Lexicon) -> bool:
Expand Down
4 changes: 2 additions & 2 deletions nameparser/_policy.py
Original file line number Diff line number Diff line change
Expand Up @@ -709,8 +709,8 @@ class Policy:
unlisted_dotted_suffixes: bool = True
#: Where an UNLISTED all-caps word of two or more letters, with no
#: period in it, reads as a credential (:class:`CapsSuffixes`). The
#: name must contrast it: a word of the name holding a capital and
#: ending in a lowercase letter ("Smith", "DiCaprio") that the
#: name must contrast it: a word of the name holding a capital whose
#: last letter is lowercase ("Smith", "DiCaprio") that the
#: vocabulary does not claim as a title, particle
#: or credential -- a record written wholly in capitals or wholly in
#: lowercase keeps every word a name word. A listed member keeps its
Expand Down
10 changes: 10 additions & 0 deletions tests/v2/cases.py
Original file line number Diff line number Diff line change
Expand Up @@ -3022,6 +3022,16 @@ def _check_cjk_shape_purity(self) -> None:
"lowercase is the contrast, so an interior capital counts "
"('DiCaprio', 'IJzerman', 'al-Rashid'). "
"The `istitle()` draft read given 'XYZ' here"),
Case("a_name_typed_with_decomposed_accents_carries_the_contrast",
"Jose\u0301 Andre\u0301, XYZ",
{"given": "Jose\u0301", "family": "Andre\u0301", "suffix": "XYZ"},
classification="fix(#564)",
ambiguities=("suffix-or-name",),
notes="#564: the contrast test reads a word's last LETTER after "
"composing it, so 'André' typed with a combining accent "
"ends in 'é' and reads as its composed spelling does. A "
"draft reading the last character took the accent and "
"kept given 'XYZ'"),
Case("a_capitalized_given_name_behind_a_two_word_surname_is_the_accepted_cost",
"García Márquez, JUAN Jr.",
{"given": "García", "family": "Márquez", "suffix": "JUAN Jr."},
Expand Down
11 changes: 8 additions & 3 deletions tests/v2/pipeline/test_vocab.py
Original file line number Diff line number Diff line change
Expand Up @@ -896,7 +896,10 @@ def test_is_single_letter_numeral() -> None:
("Smith", True), ("DiCaprio", True), ("IJzerman", True),
("al-Rashid", True), ("d'Estaing", True), ("McDonald", True),
("MacLeod", True), ("Mack", True), ("O'Neil", True),
("Džokić", True), ("E\u0301lodie", True), ("Weiß", True),
("Džokić", True), ("E\u0301lodie", True), ("Andre\u0301", True),
("Ha\u0300", True), ("Jones'", True), ("Smith2", True),
("Smith)", True), ("Weiß", True), ("KAĸ", False),
("ANDRE\u0301", False),
("SMITH", False), ("smith", False), ("C", False), ("ap", False),
("d'ESTAING", False), ("al-ASSAD", False), ("McDONALD", False),
("MacDONALD", False), ("FitzGERALD", False), ("DeVITO", False),
Expand All @@ -906,6 +909,8 @@ def test_is_single_letter_numeral() -> None:
def test_written_as_a_name(text: str, expected: bool) -> None:
# #564 (Derek): the name's case contrast is a word holding a
# capital and ending in a lowercase letter -- a surname written in
# capitals ends in one whatever is glued in front of it, and a
# trailing ß, which has no single capital, is set aside.
# capitals ends in one whatever is glued in front of it. The last
# LETTER: composed first, so a decomposed accent at the end counts
# as its letter; trailing non-letters and a lowercase letter with
# no single capital (ß, ĸ) are passed over.
assert written_as_a_name(text) is expected
18 changes: 10 additions & 8 deletions tests/v2/test_ledger_guards.py
Original file line number Diff line number Diff line change
Expand Up @@ -3811,11 +3811,12 @@ def _claim(rule: dict) -> _Claim:
# case-row names, every one a comma name. Reach, verified
# name by name.
"fix(comma-family) lone post-comma piece routes to suffix/title, not first":
# 2026-10-01, #564: 423 -> 429, 'John Smith, XYZ', 'Smith,
# 2026-10-01, #564: 423 -> 430, 'John Smith, XYZ', 'Smith,
# XYZ', 'García Márquez, MJ', 'García Márquez, MJ PhD',
# 'García Márquez, MJ JK' and 'John Smith, PhD XYZ', #564's
# rules.md#C1 examples. Reach, verified name by name.
_Claim(429, ('given', 'suffix', 'title'), "2c336f3d3ae0", None),
# 'García Márquez, MJ JK', 'John Smith, PhD XYZ' and 'MÜLLER
# WEIß, HANS', #564's rules.md#C1 examples. Reach, verified
# name by name.
_Claim(430, ('given', 'suffix', 'title'), "dcfb3a9638a9", None),
"fix(comma-family) a comma followed only by titles keeps the given/family split":
_Claim(2, ('family', 'given'), "5bd9c6d96c38", None),
"fix(comma-family) a comma followed only by titles keeps the given/family split, the C1 example":
Expand Down Expand Up @@ -3919,11 +3920,12 @@ def _claim(rule: dict) -> _Claim:
# case-row names, every one a comma name. Reach, verified
# name by name.
"fix(comma-precomma-family) pre-comma run reads as family, not given":
# 2026-10-01, #564: 423 -> 429, 'John Smith, XYZ', 'Smith,
# 2026-10-01, #564: 423 -> 430, 'John Smith, XYZ', 'Smith,
# XYZ', 'García Márquez, MJ', 'García Márquez, MJ PhD',
# 'García Márquez, MJ JK' and 'John Smith, PhD XYZ', #564's
# rules.md#C1 examples. Reach, verified name by name.
_Claim(429, ('family', 'given'), "2c336f3d3ae0", None),
# 'García Márquez, MJ JK', 'John Smith, PhD XYZ' and 'MÜLLER
# WEIß, HANS', #564's rules.md#C1 examples. Reach, verified
# name by name.
_Claim(430, ('family', 'given'), "dcfb3a9638a9", None),
# 2026-10-01, #575: new, 4; 'De La Cruz, Ed', 'Freiherr von
# Berg, Ed', 'Van Buren, Ed', 'de la Cruz, Ma'.
"fix(#575) a particle surname before a comma is one name word":
Expand Down
Loading
Loading