Skip to content

fix(C1): a particle surname before a comma is one name word (#575) - #576

Merged
derek73 merged 5 commits into
masterfrom
fix/issue-575-particle-count
Oct 2, 2026
Merged

derek73 merged 5 commits into
masterfrom
fix/issue-575-particle-count

Conversation

@derek73

@derek73 derek73 commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

Closes #575.

C1 decides whether a comma introduces the listing form (Family, Given) or trailing credentials (Name, PhD). Part of that decision is counting the words before the comma: two or more name words mean the name is already complete. All three places that count, the ones listed under Changes, counted tokens or pieces. So a particle surname counted as several words:

text 2.3.0 master this PR
De La Cruz, Ed given Ed, family De La Cruz family De La Cruz, suffix Ed, no given given Ed, family De La Cruz
van der Berg, MA given MA, family van der Berg given van, family der Berg, suffix MA family van der Berg, suffix MA
van der Berg, PhD given van, family der Berg same family van der Berg, suffix PhD
Van Buren, Ed given Ed, family Van Buren given Van, family Buren, suffix Ed given Ed, family Van Buren

The ambiguous-credential rows (Ed, MA) broke earlier in this 2.4 cycle and that breakage never shipped. The PhD row (an unambiguous credential) was wrong in 1.4.0 through 2.3.0.

The rule

A particle run and the one name word it attaches to count as one word. That's the reach P1's fold uses, not P2's whole chain, so de Mesnil Jean, Dr. under FAMILY_FIRST is still two words: family de Mesnil, given Jean.

  • The comma settles P1's fork for a leading particle that could be a first name (Derek's call). Standing alone, Van Buren reads given Van, but the listing form puts a surname before the comma. Van Johnson, Dr. moves the same way.
  • Particle chains only. Connective joins aren't counted as one word, because whether P3 joins depends on the whole name, which the count is helping to decide. Bound given-name pairs aren't either, because they build a given name. P3's statement and mechanisms.md#UNIT-PARTITION's Contract now name this count as their exception.
  • A title-particle (Freiherr, St) is no particle only where it opens the part. Freiherr von Berg, PhD keeps its title; de St Pierre, Ed stays one surname. Abu counts as a particle, so Abu Bakar, Ed keeps 2.0–2.3's reading.
  • De La Cruz, M.J. K.L. now reads given M.J., suffix K.L. (Derek's call), the same as Cruz, M.J. K.L..
  • Freiherr von Berg, Ed reads family Freiherr von Berg, which is the right reading: since 1919 a former German noble title is part of the legal surname. That's also 2.0–2.3's reading. The same listing-form rule is a cost for Prof. Cruz, Ed, recorded as accepted.
  • Out of scope: a title in front. Dr. van der Berg, PhD reads as its part does on its own.

Changes

  • _vocab.unit_ends: the unit walk moved here from post_rules, with a chain flag. Post_rules' fold uses the full chain; the C1 counts use the fold's one-word reach.
  • The C1 counts:
    • segment's legacy token count (surname_unit_count)
    • the ambiguous-class count (name_word_count, which has a no-particle fast path)
    • assign's positional-read test, which now walks every token so a suffix word still stops a particle
  • Facts: segment builds them from the vocabulary (surname_unit_tags) and assign derives them from classify's tags (surname_unit_facts). An agreement test sweeps every single-word vocabulary entry in three casings and both positions, with a stored negative control. Case rows pin each count's position flag; mutating either flag to a constant now fails.
  • Docs and ledgers: rules.md#C1 (plus P3's exception), decisions.md#C1, mechanisms.md#UNIT-PARTITION, the release log (including a stale Dotted initials after a two-word surname read as a credential: García Márquez, G.J. has no given name #563 example), and ledger entries at all five baselines.

Verification

  • Tests: full suite 10732 passed; mypy and ruff clean.
  • Gate: exits 0 at 1.4.0, 2.0.0, 2.1.0, 2.2.0 and 2.3.0. At 2.3.0, fix(#575) classifies one name.
  • Blast radius: of the 1459 names in master's corpora, under all three name orders, exactly one moves: De La Cruz, M.J. K.L., the decided move. The corpus holds only two names of this shape. Outside the corpora, the class that moves against 2.3.0 is a leading ambiguous particle plus one word before an unambiguous credential or a title alone (Abu Bakar, PhD, bin Laden, PhD, Van Johnson, Dr.).
  • Cost: John Smith, PhD goes from 209 to 217 frames per parse. Smith, John, John Smith, MA and the benchmark reference name are unchanged.
  • Review: four rounds — design docs, code, then two reviews of fix commits. Each fix round found defects in the one before: title-particles folded into the family, a suffix not stopping a particle, the title-particle exclusion applied regardless of position, Abu described backwards, an unguarded flag, false blast-radius claims. All are fixed and pinned. Details are in decisions.md#C1.

Unblocks #564.

🤖 Generated with Claude Code

derek73 and others added 4 commits October 1, 2026 20:15
rules.md#C1 counts the words before a comma three times -- v1's "more
than one word" for an unambiguous credential, the ambiguous class's
name-word count (#289, #544), and assign's two-name-word test for the
positional read -- and all three counted tokens or pieces. So
'De La Cruz, Ed' lost its given name and 'van der Berg, MA' split the
chain (an unreleased regression against 2.3.0), and every release split
'van der Berg, PhD'.

A particle and the name word it attaches to now count as one word, the
reach P1's fold uses, so 'de Mesnil Jean' stays two words under a
family-first order. The comma settles P1's fork for a leading 'van':
the listing form puts a surname before it. Bound given-name pairs and
connective joins are not counted as one (P5 builds a given name; P3
leaves 'Ortega y Gasset' three words), stated at _vocab.SURNAME_UNIT_TAGS.

The unit walk moves from post_rules into _vocab.unit_ends, shared by
post_rules' fold (full chain) and the C1 counts (fold reach). Segment
runs before classify, so it builds its facts from the vocabulary
(surname_unit_tags), held to classify's tags by a sweep test with a
recorded negative control. 'De La Cruz, M.J. K.L.' now reads given
'M.J.' (Derek's call); 'Van Johnson, Dr.' reads family 'Van Johnson'.

No pre-existing corpus name moves at any baseline; the gate exits 0 at
all five. Cost: 'John Smith, PhD' 209 -> 217 frames, other measured
names unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Code (both reviews):
- A word that is also title vocabulary or a bound given-name head
  ('Freiherr', 'St', 'Abu') is no particle to the count before a comma:
  the first draft folded 'Freiherr von Berg, PhD' into the family and
  moved 'St John, PhD' and 'Abu Bakar, Ed' against every release.
- Assign's count walks every token, so a suffix word still stops a
  particle: 'van Jr. Berg, Mr.' had read family 'van Berg'. Segment and
  assign now derive their facts through one function each from one
  set (surname_unit_tags / surname_unit_facts).
- The agreement test's negative control is stored and asserted with the
  period-joined mirror patched out; stale comments corrected.

Docs (design review):
- rules.md#C1: the reach is P1's fold, not P2's chain; title and bound
  particles excluded; the P3 reason restated (P3's join depends on the
  whole name); a title in front is out of scope; 'Freiherr von Berg, Ed'
  recorded as an accepted consequence (2.3.0's reading too). P3's
  statement and mechanisms.md#UNIT-PARTITION's Contract name C1's count
  as their exception.
- decisions.md#C1: corrected blast radius (one corpus name moves, of a
  population of two), the real cost of the one-word reach, and the
  version range.
- release_log: precise version range, the Van Johnson move, and the
  #563 bullet's stale 'De La Cruz, M.J. K.L.' example replaced; the
  #563 ledger comments updated to match.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ding

The first review round excluded title-particles and bound given-name
particles from the count in every position. The review of that fix
found it brought #575's own defect back inside a surname ('de St
Pierre, Ed' and 'De St. Croix, Ed' lost the given name again), and that
excluding 'Abu' gave up 2.0-2.3's reading of 'Abu Bakar, Ed' (given
'Ed', family 'Abu Bakar') rather than keeping it.

Now a title-particle is no particle only where it opens the part, and a
bound given-name particle is a particle. 'Freiherr von Berg, PhD' and
'St John, PhD' keep the title; 'de St Pierre, Ed' and 'Abu Bakar, Ed'
read given 'Ed'. The agreement test compares both positions; a case row
pins the position itself.

Also: the five ledger comments quoted a C1 sentence the first round
rewrote; the blast-radius recipe gains its credential filter and
comparator; the decisions entry says why the review's boundary rows
carry no shape tag, and lists Abu Bakar, PhD among the moves.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The third review found assign's title-particle position flag unguarded:
fixed to False or to True, every test passed. 'St John, Dr.' and 'von
St Johann, PhD' now pin it one per direction, and both mutants fail.

decisions.md#C1: the blast-radius recipe's second step now reproduces
its "two" (ambiguous_class_candidate in the second segment; a filter on
any suffix word keeps 17); the out-of-corpus moves are stated as the
class they are, with examples measured against 2.3.0; the guard claim
names which rows pin which count. rules.md#C1 interacts gains H1.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@derek73 derek73 added this to the 2.4 milestone Oct 2, 2026
@derek73 derek73 added the bug label Oct 2, 2026
@derek73 derek73 self-assigned this Oct 2, 2026
@codecov

codecov Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 98.89%. Comparing base (65afcdb) to head (733cc17).

Additional details and impacted files
@@            Coverage Diff             @@
##           master     #576      +/-   ##
==========================================
+ Coverage   98.87%   98.89%   +0.01%     
==========================================
  Files          45       45              
  Lines        4017     4070      +53     
==========================================
+ Hits         3972     4025      +53     
  Misses         45       45              

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Since 1919 a former German noble title is part of the legal surname,
written between the given name and the particle, so 'Freiherr von
Berg, Ed' reading family 'Freiherr von Berg' is correct, not a cost --
the cost is the 'Prof. Cruz, Ed' half of the same listing-form rule.
rules.md#C1's Accepted block, decisions.md#C1 and the ledger comments
say so, and particles.py records why 'freiherr' is a particle at all
(added in 3e14ea2 without a stated reason).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@derek73
derek73 merged commit 7394078 into master Oct 2, 2026
11 checks passed
@derek73
derek73 deleted the fix/issue-575-particle-count branch October 2, 2026 04:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Why does De La Cruz, Ed lose its given name?

1 participant