Skip to content
Merged
Show file tree
Hide file tree
Changes from 1 commit
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Prev Previous commit
docs(C1): Freiherr von Berg as the surname is the right reading (#575)
Since 1919 a former German noble title is part of the legal surname,
written between the given name and the particle, so 'Freiherr von
Berg, Ed' reading family 'Freiherr von Berg' is correct, not a cost --
the cost is the 'Prof. Cruz, Ed' half of the same listing-form rule.
rules.md#C1's Accepted block, decisions.md#C1 and the ledger comments
say so, and particles.py records why 'freiherr' is a particle at all
(added in 3e14ea2 without a stated reason).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
  • Loading branch information
derek73 and claude committed Oct 2, 2026
commit 733cc17e7e468f9f651e9d90d1985e23379ead7f
2 changes: 1 addition & 1 deletion docs/design/decisions.md
Original file line number Diff line number Diff line change
Expand Up @@ -884,7 +884,7 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py):
THE COMMA SETTLES P1's FORK, and that is the decision rather than a side effect. Standing alone, `Van Buren` is given 'Van', family 'Buren' with a `particle-or-given` report, so counting it as one surname makes the comma's reading disagree with the standalone one. Derek's call, after weighing the narrow alternative (count only a part P1 already reads as all surname, i.e. one led by a never-given particle like `de`): the listing form puts a surname before the comma, so the comma is evidence the standalone parse does not have, and the readings it gives are the ones a person takes — `Van Buren, Ed` given 'Ed', family 'Van Buren'; `van der Berg, MA` family 'van der Berg', suffix 'MA' (one name word, so C1 reads the case, and capitals in a mixed-case name make the credential). `Van Johnson, Dr.` moves with it, family 'Van Johnson', title 'Dr.', where 2.2.0 and 2.3.0 read given 'Van' and `tests/v2/pipeline/test_assign.py` had pinned the positional fork; that test now uses `Van Johnson Smith, Dr.`, which still has two name words.
THE PARTICLE REACHES ONE WORD, as P1's fold does, not to the end of the part as P2's chain does: `de Mesnil Jean, Dr.` under a family-first order is family 'de Mesnil', given 'Jean', two name words. The cost, and it is real for an ambiguous particle: a `van`-led part with a word after the surname counts two and keeps the positional read with P1's fork, so `van Buren John, Ed` reads given 'van', middle 'Buren', family 'John' and `van der Berg Smith, PhD` given 'van', family 'der Berg Smith', both unchanged from master. Under the default order a never-given particle loses nothing (P1's fold makes `de Mesnil Jean` all surname anyway). `_vocab.unit_ends` carries both reaches behind a `chain` flag, post_rules' fold reading the full chain as before.
ONLY PARTICLE CHAINS, not mechanisms.md#UNIT-PARTITION's full set — stated there as the Contract's one exception, and in P3's statement, whose "ONE name word wherever another rule counts them" now names C1's count as the exception. A bound given-name pair builds a GIVEN name, and P5 gives up a family word where the name has no other (`abdul Salam` alone is given 'abdul', family 'Salam'), so `abdul Salam, Ed` keeps the credential reading. Whether P3 joins a single-letter connective depends on the words of the whole name — `Carod i Rovira, Josep` joins as four words where `Ortega y Gasset` alone, three, does not — and the count is part of deciding what the whole name is, so segment could reach P3's answer only by copying P3; `Ortega y Gasset, Ed` keeps the credential reading. Approved as "particle chains and connective joins", narrowed to particles when P3's exception surfaced in implementation (the first draft joined `John e Smith, III` into one family name, caught by `tests/test_conjunctions.py`).
A WORD THAT IS ALSO A TITLE IS NO PARTICLE WHERE IT OPENS THE PART. `TITLES ∩ particles` is `{freiherr, st}` and `bound_given_names ∩ particles` is `{abu, أبو, ابو}` (recompute: `L = Parser().lexicon; L.particles & L.titles`, `L.particles & L.bound_given_names`). The first draft counted every particle, and the docs review found it reading `Freiherr von Berg, PhD` as family 'Freiherr von Berg' and `St John, PhD` as family 'St John', where every release reads the title. The first fix then excluded both overlaps in EVERY position, and the review of that fix found two errors in it. Position: inside a surname a title-particle chains as P1 and P2 chain it, and the blanket exclusion brought #575's own defect back for `de St Pierre, Ed`, `De St. Croix, Ed` and `de Abu Bakar, Ed` (family only, no given name, where 2.3.0 read given 'Ed'). And direction: 2.0 through 2.3 read `Abu Bakar, Ed` as given 'Ed', family 'Abu Bakar', so excluding the bound-given particle gave released behavior up rather than keeping it — and Abu Bakar is a common Malay surname. DECIDED: a title-particle is no particle only where it LEADS the part, and a bound given-name particle is a particle. So `Freiherr von Berg, PhD` and `St John, PhD` keep the title, `de St Pierre, Ed` reads given 'Ed', and `Abu Bakar, Ed` reads given 'Ed', family 'Abu Bakar' while `Abu Bakar, PhD` reads family 'Abu Bakar', suffix 'PhD' where 1.4.0 through 2.3.0 read given 'Abu'. The agreement test compares each token in both positions; it cannot see which position a caller passes, so case rows pin each count's flag: `de St Pierre, Ed` segment's, and `St John, Dr.` and `von St Johann, PhD` assign's, one per direction (the review of this round found assign's flag fixed to either constant passing every test, and both mutants now fail). ACCEPTED: `Freiherr von Berg, Ed` reads family 'Freiherr von Berg', given 'Ed' — 2.0.0's through 2.3.0's reading too, so no change against a release; only this cycle's count had read title 'Freiherr', family 'von Berg', suffix 'Ed' — the surname counts once and the title is no name word, so the count reads the listing form, which keeps a leading title in the family as it already did for `Prof. Cruz, Ed` and `Freiherr Berg, Ed` (rules.md#C1's Accepted line). A title in front is otherwise outside the fix: v1's count for an unambiguous credential counts the title as a word, so `Dr. van der Berg, PhD` reads as `Dr. van der Berg` does alone, given 'van'.
A WORD THAT IS ALSO A TITLE IS NO PARTICLE WHERE IT OPENS THE PART. `TITLES ∩ particles` is `{freiherr, st}` and `bound_given_names ∩ particles` is `{abu, أبو, ابو}` (recompute: `L = Parser().lexicon; L.particles & L.titles`, `L.particles & L.bound_given_names`). The first draft counted every particle, and the docs review found it reading `Freiherr von Berg, PhD` as family 'Freiherr von Berg' and `St John, PhD` as family 'St John', where every release reads the title. The first fix then excluded both overlaps in EVERY position, and the review of that fix found two errors in it. Position: inside a surname a title-particle chains as P1 and P2 chain it, and the blanket exclusion brought #575's own defect back for `de St Pierre, Ed`, `De St. Croix, Ed` and `de Abu Bakar, Ed` (family only, no given name, where 2.3.0 read given 'Ed'). And direction: 2.0 through 2.3 read `Abu Bakar, Ed` as given 'Ed', family 'Abu Bakar', so excluding the bound-given particle gave released behavior up rather than keeping it — and Abu Bakar is a common Malay surname. DECIDED: a title-particle is no particle only where it LEADS the part, and a bound given-name particle is a particle. So `Freiherr von Berg, PhD` and `St John, PhD` keep the title, `de St Pierre, Ed` reads given 'Ed', and `Abu Bakar, Ed` reads given 'Ed', family 'Abu Bakar' while `Abu Bakar, PhD` reads family 'Abu Bakar', suffix 'PhD' where 1.4.0 through 2.3.0 read given 'Abu'. The agreement test compares each token in both positions; it cannot see which position a caller passes, so case rows pin each count's flag: `de St Pierre, Ed` segment's, and `St John, Dr.` and `von St Johann, PhD` assign's, one per direction (the review of this round found assign's flag fixed to either constant passing every test, and both mutants now fail). `Freiherr von Berg, Ed` reads family 'Freiherr von Berg', given 'Ed' — 2.0.0's through 2.3.0's reading too, so no change against a release; only this cycle's count had read title 'Freiherr', family 'von Berg', suffix 'Ed'. The surname counts once and the title is no name word, so the count reads the listing form, which keeps a leading title in the family as it already did for `Prof. Cruz, Ed`. For `Prof.` that is a cost (rules.md#C1's Accepted line); for a German rank it is the right reading (Derek, 2026-10-01): since 1919 a former noble title is part of the legal surname, written between the given name and the particle (Karl-Theodor Freiherr von und zu Guttenberg), so `Freiherr von Berg` is the family name. That is also why `freiherr` is in PARTICLES at all — added with the German prefixes in 3e14ea20 (#18) without a stated reason, and the reason is this one. `Freiherr von Berg, PhD` keeps title 'Freiherr', the unambiguous credential's count taking the title as a word; both are readings of one name, and neither is changed here. A title in front is otherwise outside the fix: v1's count for an unambiguous credential counts the title as a word, so `Dr. van der Berg, PhD` reads as `Dr. van der Berg` does alone, given 'van'.
`De La Cruz, M.J. K.L.` reads given 'M.J.', suffix 'K.L.', reporting twice, as `Cruz, M.J. K.L.` does (Derek, 2026-10-01): with one name word before the comma the paired-initials count no longer reaches it, and C1's two-dotted-groups sentence now says "behind two or more name words". #563 had read it as a flipped credential run with no given name.
SEGMENT RUNS BEFORE CLASSIFY, so its count builds the two facts it reads (particle, suffix) from the vocabulary (`_vocab.surname_unit_tags`), while assign derives them from classify's tags (`_vocab.surname_unit_facts`) over every token of the part, so a suffix word stops a particle in both: the first draft's assign count dropped suffix pieces before walking and read `van Jr. Berg, Mr.` as family 'van Berg' (the code review). `test_classify.test_surname_unit_tags_agree_with_classify` sweeps every single-word vocabulary entry in three casings; its first run caught `JD.CPA` and `Msc.Ed.`, which classify tags as suffixes through the period-joined derivation, now mirrored, and that pair is the recorded negative control (`_SURNAME_UNIT_CONTROL`, asserted with the mirror patched out).
BLAST RADIUS, measured 2026-10-01: the gate exits 0 at all five baselines, and of the 1459 names in master's corpora exactly one moves, `De La Cruz, M.J. K.L.` (above, decided). Comparator: `parse(n).as_dict()`, the ambiguity kinds and `initials()` for every name in every `tools/differential/corpus*.jsonl` at `git archive origin/master`, under all three name orders, on master's tree against this one. The population that could move is small and the corpus is evidence about itself, not the rule: two corpus names have a part before the comma that is one particle surname of several words, with the part after the comma holding a credential — that one and `De La Cruz, M.J. PhD`, which keeps given 'M.J.'. Recompute: corpus names whose part before the first comma has more than one token and `_vocab.surname_unit_count` 1 (24 at master, nearly all `de la Vega, Juan`-type listings), then keep those whose second segment holds a word `_vocab.ambiguous_class_candidate` admits (a listed or dotted member of the credential class) — 2 at master. A filter on any suffix word keeps 17, the `de la Vega, Juan … III` listings among them. Against 2.3.0, `fix(#575)` classifies one name, `van der Berg, PhD`; the other particle-surname example names diff there only by this cycle's #289 count and report. Outside the corpora the move against 2.3.0 is a class, not a list: a leading ambiguous particle and one word, before a comma followed by an unambiguous credential or a title alone, now reads as one surname (`Abu Bakar, PhD`, `bin Laden, PhD`, `Mac Donald, PhD`, `van Gogh, Jr.`, `Van Johnson, Dr.`), where 2.3.0 read the particle as the given name; `Freiherr von Berg, Ed` and `Abu Bakar, Ed` move only against this cycle's master. The boundary case rows added in review carry no shape tag on purpose: their diffs against the older baselines come from earlier changes (#296's positional read, #289's count), so admitting them to the contract corpus would have stretched unrelated ledger rules over them; the case table asserts them either way.
Expand Down
9 changes: 7 additions & 2 deletions docs/design/rules.md
Original file line number Diff line number Diff line change
Expand Up @@ -1877,8 +1877,13 @@ C1. Rationale: a credential run after the comma means the name is in
counts once and the title is no name word, so the count reads the
listing form, and the listing form keeps a title before the comma
in the family, as it already did for any one-word surname ('Prof.
Cruz, Ed', 'Freiherr Berg, Ed'). The count had read 'von' as a
second name word and kept the title apart.
Cruz, Ed'). For a German rank that reading is the right one:
since 1919 a former noble title is part of the legal surname,
written between the given name and the particle, so 'Freiherr von
Berg' is the family name in 'Freiherr von Berg, Ed', as 2.0
through 2.3 read it. Before an unambiguous credential the same
words keep the title apart ('Freiherr von Berg, PhD'), that count
taking the title as a word; both are readings of one name.
"Freiherr von Berg, Ed" → family="Freiherr von Berg"
Accepted: three or more initials run together with periods behind
a surname of two words read as a credential. Initials are
Expand Down
5 changes: 4 additions & 1 deletion nameparser/config/particles.py
Original file line number Diff line number Diff line change
Expand Up @@ -214,7 +214,10 @@
# {do, freiherr, st} as load-bearing for the emitter
'du', # Du is a Chinese surname (Du Fu), leading under a
# family-first reading
'freiherr', # German noble title; also in TITLES, load-bearing
'freiherr', # German rank; since 1919 a former noble title is part
# of the legal surname, written before the particle
# ("Karl-Theodor Freiherr von und zu Guttenberg"), so
# it chains like one. Also in TITLES, load-bearing
'freiherrin', # as above
'heer',
'la', # La Shawn, La Toya: a given-name element. The Romance
Expand Down
10 changes: 5 additions & 5 deletions tools/differential/expected_since_1.4.0.toml
Original file line number Diff line number Diff line change
Expand Up @@ -783,11 +783,11 @@ issue = "fix(#575) a particle surname before a comma is one name word"
#
# Literal; the probe 'John van Buren, Ed' (two name words: the
# credential reading stays) is _MUST_NOT_MATCH.
# 'Freiherr von Berg, Ed' joins as rules.md#C1's accepted
# consequence: the surname counts once and 'Freiherr' is no name word,
# so the listing form reads it and keeps the title in the family, as
# it does for 'Prof. Cruz, Ed'. 1.4.0 read title 'Freiherr', first
# 'von Berg', suffix 'Ed'.
# 'Freiherr von Berg, Ed' joins, rules.md#C1's Accepted example: the
# surname counts once and 'Freiherr' is no name word, so the listing
# form reads it and keeps the rank in the family -- since 1919 part of
# the legal German surname. 1.4.0 read title 'Freiherr', first 'von
# Berg', suffix 'Ed'.
name_regex = "^(?:De La Cruz, Ed|Freiherr von Berg, Ed|Van Buren, Ed|de la Cruz, Ma)$"
fields = ["title", "given", "family", "suffix"]

Expand Down
2 changes: 1 addition & 1 deletion tools/differential/expected_since_2.0.0.toml
Original file line number Diff line number Diff line change
Expand Up @@ -2533,7 +2533,7 @@ issue = "fix(#289) a written case contrast decides a bare ambiguous acronym"
# is one, and its capitals make 'MA' the credential, as 'Smith, MA'.
# These baselines already read the particle surnames as one name;
# what #575 fixed is this cycle's count, which had not.
# 'Freiherr von Berg, Ed', #575's accepted consequence in rules.md#C1,
# 'Freiherr von Berg, Ed', rules.md#C1's Accepted example for #575,
# reads as these baselines read it and gains only the comma's report.
name_regex = "^(?:Davis Royce, Ed|De La Cruz, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|Freiherr von Berg, Ed|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van Buren, Ed|John van der Berg Ma|Ortega y Gasset, Ed|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|Van Buren, Ed|abdul Salam, Ed|abdul Smith Berg Ma|abdul Smith Ma|de la Cruz, Ma|john smith, ma|van der Berg, MA)$"
fields = ["family", "given", "middle", "suffix", "_ambiguities"]
Expand Down
2 changes: 1 addition & 1 deletion tools/differential/expected_since_2.1.0.toml
Original file line number Diff line number Diff line change
Expand Up @@ -2420,7 +2420,7 @@ issue = "fix(#289) a written case contrast decides a bare ambiguous acronym"
# is one, and its capitals make 'MA' the credential, as 'Smith, MA'.
# These baselines already read the particle surnames as one name;
# what #575 fixed is this cycle's count, which had not.
# 'Freiherr von Berg, Ed', #575's accepted consequence in rules.md#C1,
# 'Freiherr von Berg, Ed', rules.md#C1's Accepted example for #575,
# reads as these baselines read it and gains only the comma's report.
name_regex = "^(?:Davis Royce, Ed|De La Cruz, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|Freiherr von Berg, Ed|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van Buren, Ed|John van der Berg Ma|Ortega y Gasset, Ed|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|Van Buren, Ed|abdul Salam, Ed|abdul Smith Berg Ma|abdul Smith Ma|de la Cruz, Ma|john smith, ma|van der Berg, MA)$"
fields = ["family", "given", "middle", "suffix", "_ambiguities"]
Expand Down
2 changes: 1 addition & 1 deletion tools/differential/expected_since_2.2.0.toml
Original file line number Diff line number Diff line change
Expand Up @@ -1009,7 +1009,7 @@ issue = "fix(#289) a written case contrast decides a bare ambiguous acronym"
# is one, and its capitals make 'MA' the credential, as 'Smith, MA'.
# These baselines already read the particle surnames as one name;
# what #575 fixed is this cycle's count, which had not.
# 'Freiherr von Berg, Ed', #575's accepted consequence in rules.md#C1,
# 'Freiherr von Berg, Ed', rules.md#C1's Accepted example for #575,
# reads as these baselines read it and gains only the comma's report.
name_regex = "^(?:Davis Royce, Ed|De La Cruz, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|Freiherr von Berg, Ed|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van Buren, Ed|John van der Berg Ma|Ortega y Gasset, Ed|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|Van Buren, Ed|abdul Salam, Ed|abdul Smith Berg Ma|abdul Smith Ma|de la Cruz, Ma|john smith, ma|van der Berg, MA)$"
fields = ["family", "given", "middle", "suffix", "_ambiguities"]
Expand Down
2 changes: 1 addition & 1 deletion tools/differential/expected_since_2.3.0.toml
Original file line number Diff line number Diff line change
Expand Up @@ -297,7 +297,7 @@ issue = "fix(#289) a written case contrast decides a bare ambiguous acronym"
# is one, and its capitals make 'MA' the credential, as 'Smith, MA'.
# These baselines already read the particle surnames as one name;
# what #575 fixed is this cycle's count, which had not.
# 'Freiherr von Berg, Ed', #575's accepted consequence in rules.md#C1,
# 'Freiherr von Berg, Ed', rules.md#C1's Accepted example for #575,
# reads as these baselines read it and gains only the comma's report.
name_regex = "^(?:Davis Royce, Ed|De La Cruz, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|Freiherr von Berg, Ed|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van Buren, Ed|John van der Berg Ma|Ortega y Gasset, Ed|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|Van Buren, Ed|abdul Salam, Ed|abdul Smith Berg Ma|abdul Smith Ma|de la Cruz, Ma|john smith, ma|van der Berg, MA)$"
fields = ["family", "given", "middle", "suffix", "_ambiguities"]
Expand Down
Loading