Skip to content
Merged
Show file tree
Hide file tree
Changes from 1 commit
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Prev Previous commit
Next Next commit
test(C1): pin assign's leading flag for #575; fix three doc claims
The third review found assign's title-particle position flag unguarded:
fixed to False or to True, every test passed. 'St John, Dr.' and 'von
St Johann, PhD' now pin it one per direction, and both mutants fail.

decisions.md#C1: the blast-radius recipe's second step now reproduces
its "two" (ambiguous_class_candidate in the second segment; a filter on
any suffix word keeps 17); the out-of-corpus moves are stated as the
class they are, with examples measured against 2.3.0; the guard claim
names which rows pin which count. rules.md#C1 interacts gains H1.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
  • Loading branch information
derek73 and claude committed Oct 2, 2026
commit b4199f7de00450269d405ed41d6667400065f805
4 changes: 2 additions & 2 deletions docs/design/decisions.md
Original file line number Diff line number Diff line change
Expand Up @@ -884,10 +884,10 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py):
THE COMMA SETTLES P1's FORK, and that is the decision rather than a side effect. Standing alone, `Van Buren` is given 'Van', family 'Buren' with a `particle-or-given` report, so counting it as one surname makes the comma's reading disagree with the standalone one. Derek's call, after weighing the narrow alternative (count only a part P1 already reads as all surname, i.e. one led by a never-given particle like `de`): the listing form puts a surname before the comma, so the comma is evidence the standalone parse does not have, and the readings it gives are the ones a person takes — `Van Buren, Ed` given 'Ed', family 'Van Buren'; `van der Berg, MA` family 'van der Berg', suffix 'MA' (one name word, so C1 reads the case, and capitals in a mixed-case name make the credential). `Van Johnson, Dr.` moves with it, family 'Van Johnson', title 'Dr.', where 2.2.0 and 2.3.0 read given 'Van' and `tests/v2/pipeline/test_assign.py` had pinned the positional fork; that test now uses `Van Johnson Smith, Dr.`, which still has two name words.
THE PARTICLE REACHES ONE WORD, as P1's fold does, not to the end of the part as P2's chain does: `de Mesnil Jean, Dr.` under a family-first order is family 'de Mesnil', given 'Jean', two name words. The cost, and it is real for an ambiguous particle: a `van`-led part with a word after the surname counts two and keeps the positional read with P1's fork, so `van Buren John, Ed` reads given 'van', middle 'Buren', family 'John' and `van der Berg Smith, PhD` given 'van', family 'der Berg Smith', both unchanged from master. Under the default order a never-given particle loses nothing (P1's fold makes `de Mesnil Jean` all surname anyway). `_vocab.unit_ends` carries both reaches behind a `chain` flag, post_rules' fold reading the full chain as before.
ONLY PARTICLE CHAINS, not mechanisms.md#UNIT-PARTITION's full set — stated there as the Contract's one exception, and in P3's statement, whose "ONE name word wherever another rule counts them" now names C1's count as the exception. A bound given-name pair builds a GIVEN name, and P5 gives up a family word where the name has no other (`abdul Salam` alone is given 'abdul', family 'Salam'), so `abdul Salam, Ed` keeps the credential reading. Whether P3 joins a single-letter connective depends on the words of the whole name — `Carod i Rovira, Josep` joins as four words where `Ortega y Gasset` alone, three, does not — and the count is part of deciding what the whole name is, so segment could reach P3's answer only by copying P3; `Ortega y Gasset, Ed` keeps the credential reading. Approved as "particle chains and connective joins", narrowed to particles when P3's exception surfaced in implementation (the first draft joined `John e Smith, III` into one family name, caught by `tests/test_conjunctions.py`).
A WORD THAT IS ALSO A TITLE IS NO PARTICLE WHERE IT OPENS THE PART. `TITLES ∩ particles` is `{freiherr, st}` and `bound_given_names ∩ particles` is `{abu, أبو, ابو}` (recompute: `L = Parser().lexicon; L.particles & L.titles`, `L.particles & L.bound_given_names`). The first draft counted every particle, and the docs review found it reading `Freiherr von Berg, PhD` as family 'Freiherr von Berg' and `St John, PhD` as family 'St John', where every release reads the title. The first fix then excluded both overlaps in EVERY position, and the review of that fix found two errors in it. Position: inside a surname a title-particle chains as P1 and P2 chain it, and the blanket exclusion brought #575's own defect back for `de St Pierre, Ed`, `De St. Croix, Ed` and `de Abu Bakar, Ed` (family only, no given name, where 2.3.0 read given 'Ed'). And direction: 2.0 through 2.3 read `Abu Bakar, Ed` as given 'Ed', family 'Abu Bakar', so excluding the bound-given particle gave released behavior up rather than keeping it — and Abu Bakar is a common Malay surname. DECIDED: a title-particle is no particle only where it LEADS the part, and a bound given-name particle is a particle. So `Freiherr von Berg, PhD` and `St John, PhD` keep the title, `de St Pierre, Ed` reads given 'Ed', and `Abu Bakar, Ed` reads given 'Ed', family 'Abu Bakar' while `Abu Bakar, PhD` reads family 'Abu Bakar', suffix 'PhD' where 1.4.0 through 2.3.0 read given 'Abu'. The agreement test compares each token in both positions; it cannot see position itself, which the `de St Pierre, Ed` case row pins. ACCEPTED: `Freiherr von Berg, Ed` reads family 'Freiherr von Berg', given 'Ed' — 2.0.0's through 2.3.0's reading too, so no change against a release; only this cycle's count had read title 'Freiherr', family 'von Berg', suffix 'Ed' — the surname counts once and the title is no name word, so the count reads the listing form, which keeps a leading title in the family as it already did for `Prof. Cruz, Ed` and `Freiherr Berg, Ed` (rules.md#C1's Accepted line). A title in front is otherwise outside the fix: v1's count for an unambiguous credential counts the title as a word, so `Dr. van der Berg, PhD` reads as `Dr. van der Berg` does alone, given 'van'.
A WORD THAT IS ALSO A TITLE IS NO PARTICLE WHERE IT OPENS THE PART. `TITLES ∩ particles` is `{freiherr, st}` and `bound_given_names ∩ particles` is `{abu, أبو, ابو}` (recompute: `L = Parser().lexicon; L.particles & L.titles`, `L.particles & L.bound_given_names`). The first draft counted every particle, and the docs review found it reading `Freiherr von Berg, PhD` as family 'Freiherr von Berg' and `St John, PhD` as family 'St John', where every release reads the title. The first fix then excluded both overlaps in EVERY position, and the review of that fix found two errors in it. Position: inside a surname a title-particle chains as P1 and P2 chain it, and the blanket exclusion brought #575's own defect back for `de St Pierre, Ed`, `De St. Croix, Ed` and `de Abu Bakar, Ed` (family only, no given name, where 2.3.0 read given 'Ed'). And direction: 2.0 through 2.3 read `Abu Bakar, Ed` as given 'Ed', family 'Abu Bakar', so excluding the bound-given particle gave released behavior up rather than keeping it — and Abu Bakar is a common Malay surname. DECIDED: a title-particle is no particle only where it LEADS the part, and a bound given-name particle is a particle. So `Freiherr von Berg, PhD` and `St John, PhD` keep the title, `de St Pierre, Ed` reads given 'Ed', and `Abu Bakar, Ed` reads given 'Ed', family 'Abu Bakar' while `Abu Bakar, PhD` reads family 'Abu Bakar', suffix 'PhD' where 1.4.0 through 2.3.0 read given 'Abu'. The agreement test compares each token in both positions; it cannot see which position a caller passes, so case rows pin each count's flag: `de St Pierre, Ed` segment's, and `St John, Dr.` and `von St Johann, PhD` assign's, one per direction (the review of this round found assign's flag fixed to either constant passing every test, and both mutants now fail). ACCEPTED: `Freiherr von Berg, Ed` reads family 'Freiherr von Berg', given 'Ed' — 2.0.0's through 2.3.0's reading too, so no change against a release; only this cycle's count had read title 'Freiherr', family 'von Berg', suffix 'Ed' — the surname counts once and the title is no name word, so the count reads the listing form, which keeps a leading title in the family as it already did for `Prof. Cruz, Ed` and `Freiherr Berg, Ed` (rules.md#C1's Accepted line). A title in front is otherwise outside the fix: v1's count for an unambiguous credential counts the title as a word, so `Dr. van der Berg, PhD` reads as `Dr. van der Berg` does alone, given 'van'.
`De La Cruz, M.J. K.L.` reads given 'M.J.', suffix 'K.L.', reporting twice, as `Cruz, M.J. K.L.` does (Derek, 2026-10-01): with one name word before the comma the paired-initials count no longer reaches it, and C1's two-dotted-groups sentence now says "behind two or more name words". #563 had read it as a flipped credential run with no given name.
SEGMENT RUNS BEFORE CLASSIFY, so its count builds the two facts it reads (particle, suffix) from the vocabulary (`_vocab.surname_unit_tags`), while assign derives them from classify's tags (`_vocab.surname_unit_facts`) over every token of the part, so a suffix word stops a particle in both: the first draft's assign count dropped suffix pieces before walking and read `van Jr. Berg, Mr.` as family 'van Berg' (the code review). `test_classify.test_surname_unit_tags_agree_with_classify` sweeps every single-word vocabulary entry in three casings; its first run caught `JD.CPA` and `Msc.Ed.`, which classify tags as suffixes through the period-joined derivation, now mirrored, and that pair is the recorded negative control (`_SURNAME_UNIT_CONTROL`, asserted with the mirror patched out).
BLAST RADIUS, measured 2026-10-01: the gate exits 0 at all five baselines, and of the 1459 names in master's corpora exactly one moves, `De La Cruz, M.J. K.L.` (above, decided). Comparator: `parse(n).as_dict()`, the ambiguity kinds and `initials()` for every name in every `tools/differential/corpus*.jsonl` at `git archive origin/master`, under all three name orders, on master's tree against this one. The population that could move is small and the corpus is evidence about itself, not the rule: two corpus names have a part before the comma that is one particle surname of several words, with the part after the comma holding a credential — that one and `De La Cruz, M.J. PhD`, which keeps given 'M.J.'. Recompute: corpus names whose part before the first comma has more than one token and `_vocab.surname_unit_count` 1 (24 at master, nearly all `de la Vega, Juan`-type listings), then keep those whose part after it holds a suffix-vocabulary or dotted-credential word. Against 2.3.0, `fix(#575)` classifies one name, `van der Berg, PhD`; the other particle-surname example names diff there only by this cycle's #289 count and report. Moves outside the corpora against 2.3.0 are `Van Johnson, Dr.` and `Abu Bakar, PhD`, above; `Freiherr von Berg, Ed` and `Abu Bakar, Ed` move only against this cycle's master. The boundary case rows added in review carry no shape tag on purpose: their diffs against the older baselines come from earlier changes (#296's positional read, #289's count), so admitting them to the contract corpus would have stretched unrelated ledger rules over them; the case table asserts them either way.
BLAST RADIUS, measured 2026-10-01: the gate exits 0 at all five baselines, and of the 1459 names in master's corpora exactly one moves, `De La Cruz, M.J. K.L.` (above, decided). Comparator: `parse(n).as_dict()`, the ambiguity kinds and `initials()` for every name in every `tools/differential/corpus*.jsonl` at `git archive origin/master`, under all three name orders, on master's tree against this one. The population that could move is small and the corpus is evidence about itself, not the rule: two corpus names have a part before the comma that is one particle surname of several words, with the part after the comma holding a credential — that one and `De La Cruz, M.J. PhD`, which keeps given 'M.J.'. Recompute: corpus names whose part before the first comma has more than one token and `_vocab.surname_unit_count` 1 (24 at master, nearly all `de la Vega, Juan`-type listings), then keep those whose second segment holds a word `_vocab.ambiguous_class_candidate` admits (a listed or dotted member of the credential class) — 2 at master. A filter on any suffix word keeps 17, the `de la Vega, Juan … III` listings among them. Against 2.3.0, `fix(#575)` classifies one name, `van der Berg, PhD`; the other particle-surname example names diff there only by this cycle's #289 count and report. Outside the corpora the move against 2.3.0 is a class, not a list: a leading ambiguous particle and one word, before a comma followed by an unambiguous credential or a title alone, now reads as one surname (`Abu Bakar, PhD`, `bin Laden, PhD`, `Mac Donald, PhD`, `van Gogh, Jr.`, `Van Johnson, Dr.`), where 2.3.0 read the particle as the given name; `Freiherr von Berg, Ed` and `Abu Bakar, Ed` move only against this cycle's master. The boundary case rows added in review carry no shape tag on purpose: their diffs against the older baselines come from earlier changes (#296's positional read, #289's count), so admitting them to the contract corpus would have stretched unrelated ledger rules over them; the case table asserts them either way.
COST, measured 2026-10-01 by `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `John Smith, MA` 252 → 252, `John Smith, PhD` 209 → 217, `Dr. Juan de la Vega III` (the benchmark reference) 366 → 366. The +8 is the unambiguous credential path building its token list and asking each token whether it is a particle; a part with no particle never builds the units, and a frame-free approximation of `_normalize` was declined as a second spelling of it.

### T1 — separators, not joiners
Expand Down
2 changes: 1 addition & 1 deletion docs/design/rules.md
Original file line number Diff line number Diff line change
Expand Up @@ -1920,7 +1920,7 @@ C1. Rationale: a credential run after the comma means the name is in
V` reads the suffix and `Smith, John PhD I.` continues the run,
while adding a suffix comma after either turns that same letter
into the middle initial.
history: decisions.md#C1 · interacts: H2, P1, P2, P3, P5, P6, W3, S2, S3 · implemented: nameparser/_pipeline/_segment.py, nameparser/_pipeline/_assign.py, nameparser/_pipeline/_group.py
history: decisions.md#C1 · interacts: H1, H2, P1, P2, P3, P5, P6, W3, S2, S3 · implemented: nameparser/_pipeline/_segment.py, nameparser/_pipeline/_assign.py, nameparser/_pipeline/_group.py

C2. Rationale: text beyond the recognized comma parts should be
taken in without silent guessing.
Expand Down
18 changes: 18 additions & 0 deletions tests/v2/cases.py
Original file line number Diff line number Diff line change
Expand Up @@ -1118,6 +1118,24 @@ def _check_cjk_shape_purity(self) -> None:
"counts the title as a word, so the comma reads as a "
"credential comma and the part reads as it does alone, "
"P1's fork and all. Unchanged"),
# #575: assign's own count (the positional read after a comma
# followed by no name word) applies the title-particle exclusion
# by position too. These two pin that flag both ways -- a review
# mutant fixing it to False or True passed every other test.
Case("a_leading_title_particle_stays_a_title_before_a_title_only_comma",
"St John, Dr.",
{"title": "St Dr.", "family": "John"},
notes="#575 boundary: 'St' opens the part, so it is the title "
"it also is and 'John' alone is the name: two words do "
"not stand before the comma, the positional read holds. "
"2.3.0's reading"),
Case("an_inner_title_particle_chains_before_a_credential_only_comma",
"von St Johann, PhD",
{"family": "von St Johann", "suffix": "PhD"},
classification="fix(#575)",
notes="#575: 'St' inside the surname is a particle, so 'von St "
"Johann' is one name word and the family comma keeps it "
"whole. 2.3.0 read given 'von'"),
Case("the_anchor_does_not_reach_across_a_maiden_clause",
"Jane Doe Jr. nee Smith Ma",
{"given": "Jane", "family": "Doe", "suffix": "Jr.",
Expand Down
Loading