Skip to content
Merged
Show file tree
Hide file tree
Changes from 1 commit
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Prev Previous commit
Next Next commit
fix(S2): #564 -- the name's contrast: a capital and a lowercase last …
…letter (Derek)

The capital-then-lowercase test leaked a whole class: any capitalized
prefix glued to a surname written in capitals supplies the pair
('LLOYD FitzGERALD, RONALD', 'PAOLO DeVITO, MARCO', 'JEAN LaFLEUR,
PIERRE', 'DICK VanDYKE, JOHN', 'PAUL DuBOIS, JEAN', 'JEAN St-PIERRE,
MARC', 'LLOYD SMITH-McDONALD, RONALD'), as does ß ('MÜLLER WEIß, HANS');
each lost its given name, the Mc/Mac skip being one member of the class.
It also missed titlecase digraphs ('Džokić Ljubić, XYZ') and cost a frame
per letter.

Derek's criterion: a word carries the contrast if it holds a capital
and its last letter is lowercase, a trailing ß aside. A surname written
in capitals ends in one whatever is glued in front, so no prefix list
is needed; 'McDonald', 'Džokić' and decomposed accents pass. Two C-level
checks, so the per-letter cost is gone (a record with lowercase
particles is a constant +22 over master at any length). Measured against
the previous commit: exactly those eight prefixes and Džokić moved, no
corpus or case-table name, OFF still equal to master.

rules.md#C1, the Policy docstring, customize.rst and the release log
drop the false universal; decisions.md#S2 records the chosen option's
own leaks and the accepted hyphenated case; case rows and a broader
parametrized test pin it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
  • Loading branch information
derek73 and claude committed Oct 2, 2026
commit bbe6b0f62e1507f680e92bcc54955f1c03a36e3d
2 changes: 1 addition & 1 deletion docs/customize.rst
Original file line number Diff line number Diff line change
Expand Up @@ -532,7 +532,7 @@ listed below.
- ``CapsSuffixes``
- Where an unlisted all-caps word of two or more letters, with no
period in it, reads as a credential. The name must contrast it
with a word holding a capital followed by a lowercase letter
with a word holding a capital and ending in a lowercase letter
(``Smith``, ``DiCaprio``) that the vocabulary
does not claim as a title, particle or credential: a record
written wholly in capitals, or wholly in lowercase, keeps every
Expand Down
4 changes: 2 additions & 2 deletions docs/design/decisions.md

Large diffs are not rendered by default.

18 changes: 9 additions & 9 deletions docs/design/rules.md
Original file line number Diff line number Diff line change
Expand Up @@ -1725,15 +1725,15 @@ C1. Rationale: a credential run after the comma means the name is in
initials as two dotted groups. An unlisted all-caps word joins
the class in such a part only where the name carries the
contrast: one of its own words before the comma written the way a
name is written in mixed case, a capital directly followed by a
lowercase letter (a leading Mc or Mac before a capital aside),
with no period and not
claimed by the vocabulary as a title, particle, connective,
credential or generation. A word written wholly in lowercase never
carries it, so a record that writes its surname in capitals keeps
its given name whatever else it writes in lowercase or beside it
("GISCARD d'ESTAING, VALÉRY"), and a name written wholly in
lowercase reads the listing form. Where the name carries it, an
name is written in mixed case, holding a capital and ending in a
lowercase letter, with no period and not claimed by the vocabulary
as a title, particle, connective, credential or generation. A
surname written in capitals ends in a capital whatever is glued in
front of it, and a word written wholly in lowercase holds none, so
such a record keeps its given name beside its lowercase particles,
titles and clauses ("GISCARD d'ESTAING, VALÉRY", 'LLOYD
FitzGERALD, RONALD'), and a name written wholly in lowercase reads
the listing form. Where the name carries it, an
unlisted word of two capitals after the comma is the paired
initials' shape undotted and is decided at this comma exactly as
they are, by the sentences above ('García Márquez, MJ' and 'García
Expand Down
2 changes: 1 addition & 1 deletion docs/release_log.rst
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,7 @@ Release Log

- **New Policy field unlisted_dotted_suffixes, on by default: a dotted acronym nobody has listed is read by position.** ``HumanName("John Smith X.Y.Z.")`` gives suffix ``X.Y.Z.`` where every release gave last ``X.Y.Z.``, while ``Jack X.Y.Z.`` keeps its surname, the same words-to-spare rule a listed acronym takes -- and both readings are reported. After a comma the count is of the words before it, and two dotted single letters are the exception: they are how a person's initials are written, and two words before a comma may be one surname, so ``García Márquez, G.J.`` keeps first ``G.J.`` and last ``García Márquez`` and reports the fork, unless an unambiguous post-nominal in front of the initials that is not also a title, or another unlisted dotted word beside them, says otherwise (``John Smith, PhD X.Y.`` gives suffix ``PhD X.Y.``, while ``García Márquez, Ms G.J.`` keeps title ``Ms``, first ``G.J.``). Three letters or more read by the count, so ``John Smith, X.Y.Z.`` gives suffix ``X.Y.Z.`` -- and so does ``García Márquez, G.J.R.``, the accepted cost of the line, since initials are conventionally written apart (``García Márquez, G. J. R.``), as separate words this rule does not read (#563). Case is irrelevant here: the periods are the signal, so ``john smith x.y.z.`` reads the same way. Words the vocabulary does know are untouched (``M.A.``, ``Ph.D.``, ``A.B.C.``), a single trailing period is still not this shape (``John Smith Xyz.`` keeps last ``Xyz.``), and a dotted run at the FRONT of a name is untouched (``J.R.R. Tolkien``). One accident retires with it: a dotted word whose only vocabulary matches were SINGLE ASCII CHARACTERS -- the roman numerals the suffix list holds, and the lone digit ``2`` -- was reading as a generational suffix, so ``Jack X.Y.I.`` gives last ``X.Y.I.`` again, as 1.4.0 read it, while ``Msc.Ed.``, ``JD.CPA`` and ``Lt.Gov.`` are unchanged. The digit is why a dotted VERSION STRING moves with them and moves SILENTLY: ``John Smith 1.4.2`` gives last ``1.4.2`` where 2.3 gave suffix ``1.4.2``, and ``John Smith, 1.4.2`` gives first ``1.4.2``, last ``John Smith``. Such a token reports nothing at any policy -- it is no acronym either, the shape reading wanting every chunk alphabetic -- and a version string read as a credential was the same accident this retirement removes. That retirement is NOT behind this switch and stands either way -- setting it to ``False`` reads an unlisted dotted word as name material by position instead (``John Smith X.Y.Z.`` keeps last ``X.Y.Z.``), the pre-2.4 reading for THAT half alone. See the ``S2`` and ``suffix-acronym-collisions`` entries of ``docs/design/decisions.md`` (closes #516)

- **New Policy field unlisted_caps_suffixes: an unlisted all-caps word reads as a credential after a comma by default, and elsewhere on request.** Its value is a ``CapsSuffixes``. The default, ``CapsSuffixes.AFTER_COMMA``, reads such a word in the part right after a comma behind two or more name words, alone or in a run with other credentials: ``HumanName("John Smith, XYZ")`` gives first ``John``, last ``Smith``, suffix ``XYZ``, where 1.4.0 through 2.3.0 gave first ``XYZ``, last ``John Smith``; ``John Smith, LEED AP`` and ``John Smith, PhD XYZ`` give suffix ``LEED AP`` and ``PhD XYZ`` the same way, and ``John Smith, RAI`` gives suffix ``RAI`` again, as it did before 2.3. The all-caps surname convention writes the capitals at the end of a name or before a comma (``Jean DUPONT``, ``DUPONT, Jean``) and never there. A word after a one-word surname stays the given name (``Smith, XYZ``), a two-letter word reads exactly as dotted initials do (``García Márquez, MJ`` and ``García Márquez, MJ PhD`` keep first ``MJ``), and the name has to contrast the capitals with a word of its own holding a capital followed by a lowercase letter (``Smith``, ``DiCaprio``), so a record that writes its surname in capitals keeps its given name whatever it writes beside it (``GISCARD d'ESTAING, VALÉRY``, ``LLOYD McDONALD, RONALD``, ``LLOYD WEBBER, ANDREW PhD``), as does a name written wholly in lowercase. ``CapsSuffixes.EVERYWHERE`` also reads the end of a name, the given part's last word after a family comma and the word ending a maiden marker's clause: ``.parse("John Smith XYZ")`` gives suffix ``XYZ``, and ``Jean Pierre DUPONT`` gives last ``Pierre``, suffix ``DUPONT`` -- why it is not the default. ``CapsSuffixes.OFF`` reads none of them and reports nothing; it is the way to keep a given name written in capitals after a two-word surname, which the default reads as a credential (``García Márquez, GABRIEL`` gives suffix ``GABRIEL``). The field reaches the core parser only, through ``Parser(policy=Policy(unlisted_caps_suffixes=...))``; a ``HumanName`` tracks the parser's defaults, so the comma reading reaches it and the other two settings cannot be chosen from there. Neither this field nor ``unlisted_dotted_suffixes`` has a v1 ``Constants`` manager. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` (closes #516, closes #564)
- **New Policy field unlisted_caps_suffixes: an unlisted all-caps word reads as a credential after a comma by default, and elsewhere on request.** Its value is a ``CapsSuffixes``. The default, ``CapsSuffixes.AFTER_COMMA``, reads such a word in the part right after a comma behind two or more name words, alone or in a run with other credentials: ``HumanName("John Smith, XYZ")`` gives first ``John``, last ``Smith``, suffix ``XYZ``, where 1.4.0 through 2.3.0 gave first ``XYZ``, last ``John Smith``; ``John Smith, LEED AP`` and ``John Smith, PhD XYZ`` give suffix ``LEED AP`` and ``PhD XYZ`` the same way, and ``John Smith, RAI`` gives suffix ``RAI`` again, as it did before 2.3. The all-caps surname convention writes the capitals at the end of a name or before a comma (``Jean DUPONT``, ``DUPONT, Jean``) and never there. A word after a one-word surname stays the given name (``Smith, XYZ``), a two-letter word reads exactly as dotted initials do (``García Márquez, MJ`` and ``García Márquez, MJ PhD`` keep first ``MJ``), and the name has to contrast the capitals with a word of its own holding a capital and ending in a lowercase letter (``Smith``, ``DiCaprio``). A surname written in capitals ends in a capital whatever is glued in front of it, so such a record keeps its given name beside its lowercase particles, titles and maiden clauses and beside a mixed-case credential (``GISCARD d'ESTAING, VALÉRY``, ``LLOYD FitzGERALD, RONALD``, ``LLOYD WEBBER, ANDREW PhD``), as does a name written wholly in lowercase. ``CapsSuffixes.EVERYWHERE`` also reads the end of a name, the given part's last word after a family comma and the word ending a maiden marker's clause: ``.parse("John Smith XYZ")`` gives suffix ``XYZ``, and ``Jean Pierre DUPONT`` gives last ``Pierre``, suffix ``DUPONT`` -- why it is not the default. ``CapsSuffixes.OFF`` reads none of them and reports nothing; it is the way to keep a given name written in capitals after a two-word surname, which the default reads as a credential (``García Márquez, GABRIEL`` gives suffix ``GABRIEL``). The field reaches the core parser only, through ``Parser(policy=Policy(unlisted_caps_suffixes=...))``; a ``HumanName`` tracks the parser's defaults, so the comma reading reaches it and the other two settings cannot be chosen from there. Neither this field nor ``unlisted_dotted_suffixes`` has a v1 ``Constants`` manager. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` (closes #516, closes #564)

- **The comma's own decision about an ambiguous credential is now reported.** ``parse("Smith, MA").ambiguities`` names ``suffix-or-name``, and so does every other decision at the ambiguous credential class -- before or after a comma, in either direction, with no new ``AmbiguityKind`` (the family-comma attachment fork already reported this way, e.g. ``parse("Berg, Jan vd")``). A flip of the comma in which no listed ambiguous acronym takes part is the exception and is made in silence: ``John Smith, X.Y.Z.`` and ``John Smith, PhD X.Y.`` report nothing, the only such word a reader takes for a name being a pair of initials, which the comma reads as the given name unless something beside it has already said otherwise. Two pairs speaking only for each other still make the credential run, and that flip reports: ``John Smith, X.Y. P.Q.`` gives last ``Smith``, suffix ``X.Y. P.Q.`` (#563). One report per decision: ``Smith, Ma`` reports that the word was kept as the given name just as ``Smith, MA`` reports that it was taken as a credential. The reading a SURNAME PARTICLE swallows is reported too, which no release before this one did: ``John van der Berg Ma`` gives last ``van der Berg Ma`` and names ``suffix-or-name``, where the chain took a word the credential reading had considered. ONE report goes away, because a comma segment the parser reads as a credential run is no longer called unrecognized: ``Steven Hardman, MD, DO, DDS`` no longer reports ``comma-structure``, on its written case. That is the whole of the losses over the differential corpora -- ``John Smith, MD, R.A.I.`` is quieted on its shape by the same change, but it never reported at 2.3.0 either, having only carried the flag inside this release's own development. The other movement an upgrader sees is a SWAP rather than a loss: ``Jack X.Y.I.`` reported ``given-or-family`` at 2.3.0 and reports ``suffix-or-name`` here, the dotted retirement above having handed it to the ambiguous class. Everything else at this class is a GAIN, which is what the rest of this bullet describes. Two slots this bullet left silent no longer are, and the two bullets below close them: a credential trailing the GIVEN part of a family-comma listing now reads as a credential and reports either way, and so does one ending a maiden marker's clause. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md``

Expand Down
23 changes: 12 additions & 11 deletions nameparser/_pipeline/_segment.py
Original file line number Diff line number Diff line change
Expand Up @@ -128,21 +128,22 @@ def texts(seg: tuple[int, ...]) -> list[str]:
# and the name carries it only if one of its own words before the
# comma (no maiden clause, no delimited content, as `own_words`
# defines them for `case_class` above) is written the way a name is
# written in mixed case -- a capital directly followed by a
# lowercase letter (`written_as_a_name`) -- and is not claimed as a
# written in mixed case -- a capital in it and its last letter
# lowercase (`written_as_a_name`) -- and is not claimed as a
# title, particle, connective, credential or generation
# (`claimed_as_non_name`); a word with a period is an abbreviation
# or an initial, not one. A lone capital ('de GAULLE C') or a
# lowercase-only word has no such pair, and nor has a lowercase
# prefix glued to a capitalized surname, so a record that writes
# its surname in capitals keeps its given name whatever else it
# writes ('LLOYD ap RHYS', 'HAFEZ al-ASSAD', "GISCARD d'ESTAING",
# 'LLOYD McDONALD', 'LLOYD WEBBER née Smith'), and so does a
# mixed-case credential ('LLOYD WEBBER, ANDREW PhD').
# or an initial, not one. A lone capital ('de GAULLE C'), a
# lowercase-only word and a surname written in capitals with
# anything glued in front ("d'ESTAING", 'FitzGERALD', 'McDONALD')
# all fail it, so such a record keeps its given name beside its
# lowercase particles, clauses and titles ('LLOYD ap RHYS', 'HAFEZ
# al-ASSAD', 'LLOYD WEBBER née Smith') and beside a mixed-case
# credential ('LLOYD WEBBER, ANDREW PhD').
# Asked only once a caps word is in hand: a part with no lowercase
# at all is settled in one C-level comparison; past that the walk
# costs a few frames a word before the comma, a fold and a
# wordlist test only where the case test passes.
# costs a frame a word before the comma (the generator), plus a
# call, a fold and a wordlist test only where the C-level case
# test passes.
def name_contrast() -> bool:
before = "".join([state.tokens[i].text for i in groups[0]])
if before == before.upper():
Expand Down
22 changes: 10 additions & 12 deletions nameparser/_pipeline/_vocab.py
Original file line number Diff line number Diff line change
Expand Up @@ -722,18 +722,16 @@ def in_any_wordlist(n: str, lexicon: Lexicon) -> bool:

def written_as_a_name(text: str) -> bool:
"""Whether TEXT is written the way a name is written in mixed case:
a capital directly followed by a lowercase letter ('Smith',
'DiCaprio', 'IJzerman', 'al-Rashid', "d'Estaing") -- #564's test
for the name's case contrast (Derek). A word written wholly in one
case has no such pair, and nor does a lowercase prefix glued to a
surname written in capitals ("d'ESTAING", 'al-ASSAD'). A leading
'Mc'/'Mac' is skipped where a capital follows it, so 'McDONALD'
has no pair while 'McDonald' keeps 'Do' and 'Mack' its own 'Ma'."""
if text.startswith("Mc") and text[2:3].isupper():
text = text[2:]
elif text.startswith("Mac") and text[3:4].isupper():
text = text[3:]
return any(a.isupper() and b.islower() for a, b in zip(text, text[1:]))
it holds a capital (or titlecase letter) and its last letter is
lowercase -- #564's test for the name's case contrast (Derek).
'Smith', 'DiCaprio', 'IJzerman', 'al-Rashid', "d'Estaing",
'McDonald', 'Džokić' pass. A surname written in capitals fails
whatever is glued in front of it ("d'ESTAING", 'al-ASSAD',
'McDONALD', 'FitzGERALD', 'DeVITO', 'St-PIERRE'), and so do a lone
capital and a lowercase-only word. A trailing 'ß' is set aside,
having no single capital form: 'WEIß' is written in capitals and
'Weiß' is not. Two C-level checks, so a word costs no frame."""
return text != text.lower() and text.rstrip("ß")[-1:].islower()


def claimed_as_non_name(n: str, lexicon: Lexicon) -> bool:
Expand Down
6 changes: 3 additions & 3 deletions nameparser/_policy.py
Original file line number Diff line number Diff line change
Expand Up @@ -709,9 +709,9 @@ class Policy:
unlisted_dotted_suffixes: bool = True
#: Where an UNLISTED all-caps word of two or more letters, with no
#: period in it, reads as a credential (:class:`CapsSuffixes`). The
#: name must contrast it: a word of the name with a capital followed
#: by a lowercase letter ("Smith", "DiCaprio") that the vocabulary
#: does not claim as a title, particle
#: name must contrast it: a word of the name holding a capital and
#: ending in a lowercase letter ("Smith", "DiCaprio") that the
#: vocabulary does not claim as a title, particle
#: or credential -- a record written wholly in capitals or wholly in
#: lowercase keeps every word a name word. A listed member keeps its
#: own case lean in every setting ("Jack MA" gives suffix ``MA``),
Expand Down
Loading
Loading