Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 9 additions & 1 deletion docs/design/decisions.md

Large diffs are not rendered by default.

47 changes: 42 additions & 5 deletions docs/design/rules.md
Original file line number Diff line number Diff line change
Expand Up @@ -455,6 +455,15 @@ P2. Rationale: a particle is written as part of the surname it
Where P1's fold has claimed the opening, the fold decides the
family instead — and may take only PART of the final group,
since it counts name words and the group is one part.
In the part before a family comma, which the comma named the
family, a particle that is also suffix vocabulary heads the name
word behind it, whatever case the name is written in, and is not
read as a suffix standing between two family words; with a suffix
word or nothing behind it, it reads as the suffix it is there. A
lone letter behind it reads as it does in the same part without
the particle: a capital is a name word ('SMITH VD V, JOHN' as
'SMITH V, JOHN'), a lowercase letter the numeral ('smith vd v,
john' as 'smith v, john').
"John van der Berg" → family="van der Berg"
"John van der Berg Smith" → family="van der Berg Smith"
"Vincent van Gogh van Beethoven" → middle="van Gogh"
Expand All @@ -471,6 +480,8 @@ P2. Rationale: a particle is written as part of the surname it
"Freiherr von Berg MA" → family="von Berg"
"Freiherr von Berg MA" → suffix="MA"
"Freiherr von Richthofen V" → suffix="V" · boundary
"SMITH VD MA, JOHN" → family="SMITH VD MA"
"SMITH VD, JOHN" → suffix="VD" · boundary
"John van der Berg née Jones" → family="van der Berg"
Accepted: a particle of the unambiguous suffix vocabulary too
(vd, mc) is a suffix piece to the peel, so where it opens the
Expand All @@ -496,7 +507,7 @@ P2. Rationale: a particle is written as part of the surname it
(#132's ask) has it as the surnames view rather than the
family field.
"Vincent van Gogh van Beethoven" → surnames="van Gogh van Beethoven"
history: decisions.md#P2 · interacts: P1, P4, H5, M2, S2, C2 · implemented: nameparser/_pipeline/_group.py, nameparser/_pipeline/_post_rules.py
history: decisions.md#P2 · interacts: P1, P4, H5, M2, S2, C2 · implemented: nameparser/_pipeline/_assign.py, nameparser/_pipeline/_group.py, nameparser/_pipeline/_post_rules.py

P3. Rationale: connective words ("y", "of the") bind name words into
one name part; but a single letter in a short name is more
Expand Down Expand Up @@ -786,7 +797,7 @@ P6. Rationale: a particle ending the name has nothing to link
the word is BOTH a particle and suffix vocabulary, this
attachment outranks the suffix reading (S2): a trailing
abbreviation after a family comma is the tussenvoegsel far more
often than the decoration it collides with. Two exceptions. The
often than the decoration it collides with. Three exceptions. The
first is where the capitals speak: a word of the AMBIGUOUS
credential class, written in capitals in a name written in more
than one case, reads as the credential and this attachment stands
Expand All @@ -795,7 +806,18 @@ P6. Rationale: a particle ending the name has nothing to link
The second is S2's company: such a word standing behind an
unambiguous credential in one run of suffix words reads as the
credential in any spelling and any case, reported as S2's
credential fork. Every other spelling of such a word attaches as
credential fork. The third is a lone word of both the particle and
the UNAMBIGUOUS suffix vocabulary standing INSIDE a run read as
post-nominals — one in front of it and another behind, a
credential or a generation, whatever made each one, and a title
between it and the one in front transparent (H5): the
tussenvoegsel stands right behind the given name, so a
post-nominal between them says the word is not one. It does not
end the name, and keeps the
post-nominal reading in any case, reported as S2's credential
fork. With nothing behind it, it ends the name and attaches, and
two such words side by side are one particle run, which this rule
takes whole. Every other spelling of such a word attaches as
it did before, and the kind rule below gives it this rule's
particle fork rather than S2's credential one. In a name written
wholly in one case, with nothing in front of the word to speak
Expand All @@ -814,6 +836,10 @@ P6. Rationale: a particle ending the name has nothing to link
"Doe, John Do" → family="Do Doe"
"SMITH, JOHN DO" → family="DO SMITH"
"Doe, John van DO" → family="van DO Doe"
"DOE, JANE PHD VD MA" → suffix="PHD VD MA"
"doe, jane phd vd ma" → suffix="phd vd ma"
"DOE, JANE PHD VD" → family="VD DOE" · boundary
"DOE, JANE MA VD PHD" → suffix="MA VD PHD"
Without a comma, a declared family-first order has named the
family in the same way and the attachment fires there too — but
only where the run ENDS the name and stands in a MIDDLE — the one
Expand Down Expand Up @@ -984,7 +1010,15 @@ S2. Rationale: generational suffixes and credentials are recognized
and the count is not what decides it there. The comma has already
named the family and the first name word after it is the given
name, so the words to spare are there by construction and the
count says nothing: a word that ENDS that part reads as the
count says nothing — unless that first word is a never-given
particle, which is no given name and needs the next name word to
attach to (P1): that word is the surname the particle heads, not
a word to spare, so it reads as a name unless its capitals in a
mixed-case name say credential, as they do with no words to spare
anywhere: `SMITH, VD MA` reads as `Smith, vd Ma` does, `Smith, de
MA` keeps suffix `MA`, and `SMITH, VAN MA`, whose particle can be
a given name, keeps it and reads suffix `MA`. Otherwise a word
that ENDS that part reads as the
credential unless its WRITING says otherwise, and a word that
does not end it is never asked. "Ending the given part" reaches
past the credentials behind it and past a trailing title, which
Expand Down Expand Up @@ -1124,6 +1158,9 @@ S2. Rationale: generational suffixes and credentials are recognized
"Doe, John MA Smith" → middle="MA Smith" · boundary
"Doe, John DO" → suffix="DO"
"SMITH, JOHN DO" → family="DO SMITH" · boundary
"SMITH, VD MA" → family="SMITH VD MA"
"smith, de ma" → family="smith de ma"
"SMITH, VD DE LA MA" → family="SMITH VD DE LA MA"
"Doe, John MA JD" → ambiguities=("suffix-or-name", "suffix-or-name")
"Doe, John MA Ma" → middle="MA Ma" · boundary
"Doe, John MA Ma" → ambiguities=("suffix-or-name",) · boundary
Expand Down Expand Up @@ -1193,7 +1230,7 @@ S2. Rationale: generational suffixes and credentials are recognized
and unchanged (decisions.md#v1-xfail-triage: `king` stays a
title, for the addressing forms).
"Dr Jr" → suffix="Jr"
history: decisions.md#S2 · interacts: H1, H2, H3, H5, C1, C2, S3, P2, P3, P5, P6, M2 · implemented: nameparser/_pipeline/_classify.py, nameparser/_pipeline/_group.py, nameparser/_pipeline/_pieces.py, nameparser/_pipeline/_vocab.py
history: decisions.md#S2 · interacts: H1, H2, H3, H5, C1, C2, S3, P1, P2, P3, P5, P6, M2 · implemented: nameparser/_pipeline/_classify.py, nameparser/_pipeline/_group.py, nameparser/_pipeline/_pieces.py, nameparser/_pipeline/_vocab.py

S3. Rationale: credentials are often written run together with
periods; the chunks between the periods are what carry the
Expand Down
2 changes: 2 additions & 0 deletions docs/release_log.rst
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,8 @@ Release Log

- **Fix a particle surname before a comma being split when a credential follows it.** ``HumanName("van der Berg, PhD")`` gives last ``van der Berg``, suffix ``PhD``, where 1.4.0 through 2.3.0 gave first ``van``, last ``der Berg``. Where the parser counts the words before a comma, a particle and the word it attaches to now count as one, so a particle surname reads as ``Berg, PhD`` does, and ``Abu Bakar, PhD`` gives last ``Abu Bakar`` where 1.4.0 through 2.3.0 gave first ``Abu``. The same count keeps a particle surname whole in front of a credential that is also a name: ``De La Cruz, Ed`` gives first ``Ed``, last ``De La Cruz``, as 2.3.0 read it, and ``van der Berg, MA`` gives last ``van der Berg``, suffix ``MA``, the reading ``Smith, MA`` gets below. Where the part is the surname alone, the comma also decides that a leading particle is not a first name: ``Van Johnson, Dr.`` gives last ``Van Johnson``, where 2.2 and 2.3 gave first ``Van``. A given name in front still makes two words, so ``John van Buren, Ed`` reads suffix ``Ed`` as ``John Smith, Ed`` does. A connective surname is not counted as one: ``Ortega y Gasset, PhD`` still gives first ``Ortega``, middle ``y``, as ``Ortega y Gasset`` reads on its own. See the ``C1`` entry of ``docs/design/decisions.md`` (closes #575)

- **Fix a surname particle that is also a credential (vd, mc) being split away from the surname or the credentials around it.** ``HumanName("SMITH VD MA, JOHN")`` gives last ``SMITH VD MA``, where 2.2 and 2.3 gave last ``SMITH MA``, suffix ``VD``, and ``HumanName("Doe, Jane PhD vd MA")`` gives suffix ``PhD vd MA``, where 2.3 gave middle ``MA``, suffix ``PhD vd``. A ``vd`` with nothing behind it still joins the last name: ``HumanName("Doe, Jane PhD vd")`` gives last ``vd Doe``. (closes #573)

- **Fix a one-letter connective joining a name that gives no sign it is a connective.** ``HumanName("jose e maria santos")`` gives first ``jose``, middle ``e maria``, last ``santos``, where 1.4.0 through 2.3.0 gave first ``jose e maria``; and ``JUAN GARCIA Y LOPEZ`` gives last ``GARCIA Y LOPEZ``, where every release since 1.4.0 read the bare capital as an initial and gave middle ``GARCIA Y``. A single letter is an initial where the writing says so -- a bare Latin capital in a name that is not written wholly in one case -- and a name written wholly in one case says nothing either way, so the reading comes from the vocabulary there: ``e`` reads as an initial and ``y`` joins. Mixed-case input is untouched in both directions: ``Jose e Maria Santos`` still gives first ``Jose e Maria`` and ``Jose E Maria Santos`` still gives middle ``E Maria``. Short names move in the derived views rather than the fields, P3's three-word carve-out being unchanged: ``parse("john e smith").initials()`` is ``j. e. s.`` where 2.3.0 gave ``j. s.``, and ``HumanName("john e smith").capitalize()`` gives ``John E Smith`` where 2.3.0 gave ``John e Smith``; ``JUAN Y GARCIA`` moves its ``capitalize()`` the same way in reverse, giving ``Juan y Garcia``, while its initials do not move at all: ``parse(...).initials()`` is ``J. Y. G.``, what 2.3.0 gave and what 1.4.0's own view gave, the ``Y`` holding its part alone and contributing an initial again under the connective-initials fix further down this list (closes #461). ``HumanName.initials()`` agrees with the core on both -- see the #528 bullet below, which closed a split this change opened and the same release closes. Seventeen names in the differential corpora are written in one case and carry a cased single-letter connective, and ten of them move something against 2.3.0. The Cyrillic reading is unchanged (``Хосе И Мария Сантос`` still gives first ``Хосе И Мария``), and Arabic ``و`` never enters the rule, having no case to be written against. A ``Lexicon`` knob decides which letters are marked, so the reading is configurable rather than fixed. See the ``P3`` entry of ``docs/design/decisions.md`` (closes #383, closes #479)

- **Fix HumanName.initials() reading a one-letter connective by vocabulary and written shape instead of by the parse.** ``HumanName("john e smith").initials()`` gives ``j. e. s.``, where every release from 1.4.0 through 2.3.0 gave ``j. s.``; ``JUAN GARCIA Y LOPEZ`` gives ``J. G. L.`` where 2.3.0 gave ``J. G. Y. L.``; ``JUAN Y GARCIA`` does not move at all this cycle, giving ``J. Y. G.`` on both surfaces as 2.3.0 and 1.4.0 did, since the connective-initials fix further down this list (closes #461) gives its ``Y`` an initial again. That first name read 1.4.0's way at 2.3.0 and only there: 2.0.0 through 2.2.0 already gave today's answer, by the unrelated bug the 2.3.0 note below records as fixed (the facade dropping a bare capital that is also a one-letter conjunction, #462), so against those three releases it does not move at all. The v1 facade decided whether a word was the connective by looking the word up and checking its shape, while ``parse(...).initials()`` read the tag the parse recorded -- so the change above, which reads a single letter in a one-case name from the vocabulary rather than from its case, moved one view and not the other. Both views of a parse now give the same answer. Mixed-case names are untouched on both, the writing having decided the letter: ``John E Smith`` is still ``J. E. S.`` and ``Scott E. Werner`` still ``S. E. W.``. So is a one-case name whose letter is outside the marked set -- ``maria y lopez`` is still ``m. l.``, ``y`` having joined before this release and after it. Two costs, and both match what ``capitalize()`` has always done: editing ``C.conjunctions`` after a name is parsed no longer changes its initials until ``full_name`` is assigned again, and a name restored from a pickle, copied with ``copy.copy``/``copy.deepcopy`` (the same state hooks), or built from keyword fields (``HumanName(first=..., middle=..., last=...)``) carries no tags, so its initials come from the vocabulary and can differ from a fresh parse of the same string. One private break, stated because a v1 subclass can hit it: an override of ``_process_initial`` written to v1's ``(name_part, firstname=False)`` signature now raises ``TypeError`` the first time ``initials()`` runs, since ``initials()`` passes the part's tokens. Such an override has to accept a ``tokens`` keyword *and pass it on* -- ``return super()._process_initial(name_part, firstname, tokens=tokens)`` -- to receive this fix. Widening the signature without forwarding still works, but on the pre-#528 STRING path: the token call hands the override the group's own text as ``name_part`` rather than an empty placeholder, so ``john e smith`` initials ``j. s.`` under such an override, not the ``j. e. s.`` above. A subclass overriding one of the public ``first_list``, ``middle_list`` or ``last_list`` properties keeps working too: that member takes the pre-2.4 vocabulary reading instead of the change above, while an un-overridden member still moves. See the ``R3`` entry of ``docs/design/decisions.md`` (closes #528)
Expand Down
66 changes: 63 additions & 3 deletions nameparser/_pipeline/_assign.py
Original file line number Diff line number Diff line change
Expand Up @@ -76,7 +76,8 @@
)
from nameparser._pipeline._pieces import (
anchor_in_reach, credential_at_the_given_slot, given_slot_anchors,
is_suffix_piece, leading_titles, peel_walk,
is_lone_never_given_particle, is_suffix_piece, leading_titles,
listed_lean, peel_walk,
segment_suffix_reading, tail_reading, trailing_titles,
)
from nameparser._pipeline._state import (
Expand Down Expand Up @@ -775,6 +776,49 @@ def trailing_floor(m: int, titled: frozenset[int]) -> int:
floors[titled] = (low, final)
return low

#: where the particle run opening the given part ends,
#: per `titled` value -- `attaches_the_lead`'s one walk
lead_run_ends: dict[frozenset[int], int] = {}

def attaches_the_lead(m: int, titled: frozenset[int]) -> bool:
"""Is piece `m` the name word a never-given particle
opening the given part attaches to? post_rules' P1
fold takes such a particle forward into the family and
needs a name word for it to attach to, so the given
name the #531 slot below counts on is not there, and
the first word past the particle run is the surname it
heads rather than a word to spare ('SMITH, VD MA'
reads as 'Smith, vd Ma' does, #573). Only where the
member's writing says nothing: capitals in a mixed-case
name still make it the credential ('Smith, de MA'
keeps suffix 'MA'), as they do with no words to spare
anywhere (rules.md#S2). Asked only of a class member
the slot would otherwise take."""
if not is_lone_never_given_particle(pieces[n], tokens):
return False
if listed_lean(tokens[pieces[m][0]],
state.one_case) == "credential":
return False
# The member is the particle run's word when it is the
# first kept piece past that run. Where the run ends is
# one answer per `titled` value, walked once and kept,
# so a run of members costs one walk rather than one
# each (the second review measured the per-member scan
# quadratic on 'SMITH, VD ' + 'DE '*k + 'MA '*k). Every
# token, not the piece's tags: a chain of particles
# ('DE LA') is one piece whose tags say nothing of it
# ('SMITH, VD DE LA MA').
run_end = lead_run_ends.get(titled)
if run_end is None:
run_end = n + 1
while run_end < len(pieces) and (
run_end in titled
or all("particle" in tokens[i].tags
for i in pieces[run_end])):
run_end += 1
lead_run_ends[titled] = run_end
return m <= run_end

def reads_as_a_suffix(m: int, titled: frozenset[int]) -> bool:
"""Does this segment's walk read piece `m` as a suffix?

Expand Down Expand Up @@ -843,7 +887,8 @@ def reads_as_a_suffix(m: int, titled: frozenset[int]) -> bool:
# A name word behind the member ends the
# reach, and the member is an ordinary middle
# name read in silence.
if m >= trailing_floor(m, titled):
if (m >= trailing_floor(m, titled)
and not attaches_the_lead(m, titled)):
# A member that is ALSO particle
# vocabulary reads as the credential only
# on a POSITIVE credential lean: P6's
Expand Down Expand Up @@ -1109,8 +1154,23 @@ def reads_as_a_suffix(m: int, titled: frozenset[int]) -> bool:
if positional:
order = _assign_main(0, state, tokens, ambiguities)
else:
# rules.md#P2: a particle "joins the words after it into one
# name part" -- so a particle that is suffix vocabulary too
# (vd, mc) with a name piece behind it in this part heads
# that name, and is not peeled from between two family words
# (#573: 'SMITH VD MA, JOHN', where the uniform case left the
# chain stopped before VD; 'SMITH VD JR, JOHN', a suffix word
# behind it, keeps suffix 'VD JR'). Group cannot make this
# call, not knowing which read this segment gets: a comma
# followed by no name word reads it positionally, trailing
# run and all ('Berg de MA, Prof.').
for k, piece in enumerate(fam_pieces):
if k > 0 and is_suffix_piece(piece, fam_tags[k], tokens):
if (k > 0 and is_suffix_piece(piece, fam_tags[k], tokens)
and not (k + 1 < len(fam_pieces)
and "particle" in tokens[piece[0]].tags
and not is_suffix_piece(
fam_pieces[k + 1], fam_tags[k + 1],
tokens))):
_set_roles(tokens, piece, Role.SUFFIX)
else:
_set_roles(tokens, piece, Role.FAMILY)
Expand Down
Loading
Loading