Skip to content
Merged
Show file tree
Hide file tree
Changes from 1 commit
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Prev Previous commit
Next Next commit
review round: PR #520 findings
P5's bound-given reserve counted a trailing period-marked title word
as a name word to spare, so `Prof. abdul rahman Prof.` joined the
bound pair and read family 'abdul rahman' where `Prof. abdul rahman`
reads given 'abdul', family 'rahman'. The reserve now subtracts the
H5 run from both views -- assign's second peel runs over the pieces
that walk LEFT, so the reserve has to read the same list -- and the
same-suffix comparison is read over the shortened lists too, which is
what stops `Sir abdul Prof.` joining a title word into the given name.
No corpus name moves: the 21-mover list against a0b93f0 is unchanged.
Frames are unchanged at 416 parse / 453 facade on py3.11; the
reference name never enters the reserve branch, having no bound given
word, so its two new calls cost nothing.

rules.md#P5 names the H5 walk in the sentence the code cites; H5 gains
the P2 and M2 boundaries (`John van der Berg Prof.` keeps family 'van
der Berg Prof.', `Mary Smith née Jones Prof.` maiden 'Jones Prof.')
and the P5 clause; H5/P2/M2/P5 gain each other's `interacts:`.

Tests: the script-order half of "set before the positional read"
(`毛 泽东 Dr.`), the conjunction-merged unit that actually pins the
one-word gate in `trailing_titles`, `Dr. Do Jr.`'s own test kept out
of the no-op-chain parametrization's silence claim, and case rows for
the two bound-given readings, `Smith, E.S.Q.`, the two maiden
orderings, `Smith, John, Prof.` and `John Smith Prof. and Dr.`.

Stale counts and false comments, all re-measured: 12/nineteen ->
14/twenty-one for the title-run bundle in the four ledgers and the
guard, the twelve-member alternation is fourteen, the release log's
bullet names the accepted `Mary Jane King.` cost, and the ledgers'
role paragraph says which baseline has three two-role movers and
which has two. `John Prof. MA` does report an ambiguity, so the
"none of the fourteen" sentence is corrected to the one that matters
(`title-or-name`, still none). The first suffix peel is provisional
only where the walk takes something; `trailing_titles` is entered by
every parse with a name word to place and not by all 1289 corpus
parses (52 return first); its `rest` is the caller's name pieces, and
group is now a third caller, which mechanisms.md's census records.
Two defensive branches are marked measured-inert over 191,146
generated inputs -- the folded-vs-raw last word and the titled-piece
skip -- and the third the review named, `walkable` starting at `n`,
is NOT inert: dropping it moves 24 of those inputs, so it is marked
load-bearing with the shape that moves.

compare.py's RECOMPUTE paragraph: the strict row counts are 38/34/33/8
over 52 names, 50 of them in corpus_issues.jsonl. The every-file pair
recorded on 2026-09-05 is retracted rather than bumped -- it exceeded
the strict pair, which is impossible -- and replaced with a derivation
a reader can run without a baseline wheel.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
  • Loading branch information
derek73 and claude committed Sep 9, 2026
commit e489dc162589d4069e54609370fd83caf7e5573a
2 changes: 1 addition & 1 deletion docs/design/mechanisms.md
Original file line number Diff line number Diff line change
Expand Up @@ -55,7 +55,7 @@ Problem shape. "Which stage does X?" — asked before attributing behavior in pr

## ONE-PREDICATE-PER-QUESTION — one predicate answers it, and every other site calls that

Problem shape. Two stages need the same answer about the same input, and the one that does not own the decision is about to test for it. Contract statement. Where two sites ask the same question, exactly one predicate answers it and every other site calls that one — never a condition written to match it. The predicate belongs to the QUESTION, not to whichever stage decides: it may sit in a leaf both stages import, and for the leading-title test it must, since the deciding stage is assign and group cannot import assign. How it works. A hand-written mirror agrees with its original only until one of them moves, and the drift is invisible in both directions: each site keeps passing its own tests while they disagree about an input neither covers. Five instances, every one found as a defect before it was found as a pattern — #319 lifted the wholly-suffix predicate into the vocabulary layer "so the comma decision and the honorific peel's segment test cannot drift apart"; #401/#421 lifted the trailing-numeral fork out of assign so the bound-given reserve stopped carrying a copy, its hand-written mirror having been falsified in review more than once — the lesson recorded there being that what must be mirrored is assign's WALK, not merely its condition; #425 replaced that reserve's hand re-derivation of the trailing peel with one function over the view the join would leave; #424 moved assign's leading-title test down because group's own `title()` does not see H2's unlisted abbreviations, so `Xyz. van Johnson` chained where `Dr. van Johnson` did not; #429 moved the no-name-segment test down because group asked by segment INDEX where assign asks by CONTENT. The destination follows the LAYER, not the topic: a predicate over token text goes to `_vocab`, one over pieces and tags to `_pieces`. Both are leaves the stages sit on. The piece layer got its own module only in #439 — until then those predicates collected in `_group`, not because grouping owned them but because `_assign` imports `_group` and cannot be imported back, so group was the one place both stages could reach; five had accumulated across four PRs before the module existed. Stage order is this mechanism's limit, and it forecloses the alternative: where the reader comes AFTER the decider, record the answer on the state instead — `ParseState.order` is that shape, "Recorded rather than recomputed downstream, because the two can differ" — which is unavailable whenever the EARLIER stage is the one asking. (The concrete assign→group import that forced the `_group` collection is gone since #439; what remains is the ordering it was a symptom of, and tests/v2/test_layering.py is where the leaf's contract is now written down.) The cost is a second evaluation of the same predicate, measured for #429 at 1.2–2.2% of a family-comma parse and 0% of every other; recording that number was the right answer there over plumbing a state field the two sites would not otherwise share. Lives in. nameparser/_pipeline/_vocab.py over text (is_wholly_suffix; is_trailing_numeral_suffix — the #401/#421 instance, whose only caller since #439 is the shared peel rather than a stage; and maiden_marker_run, the #434 instance and the clearest two-stage case, called by classify over token texts and by extract over a clause's whitespace words, with group reading the tags classify recorded because it runs later; and delimiter_cores, the #436/#437 instance, read by group where a tail segment DROPS a configured delimiter core and by post_rules where the suffix view's entry boundary asks whether a dropped token was one, with a third reader inside this same module, is_wholly_suffix, where a configured core counts as suffix-shaped) and nameparser/_pipeline/_pieces.py over pieces: is_suffix_piece, leading_titles, peel_walk and peel_trailing are called by both stages, while is_leading_title, is_title_piece and trailing_start are called by group alone (measured 2026-09-06 by call site: `is_leading_title` has no caller in `_assign.py`, which reads `leading_titles` instead — a first draft of this clause listed it among the shared ones) — `trailing_start` being the one to know, since it answers where the trailing run begins and is what P2's chain and M2's walk stop at — and segment_suffix_reading by assign alone since #436/#437, that last one being #430's instance, where THREE readers shared one answer until the render join, group's third, was replaced by a rule over the commas the writer typed (decisions.md#C1, 2026-09-06); it stays where it is, one call site being no reason to move a predicate that two sites will contest again. `trailing_titles` joins that last shape (2026-09-08, the #316/#489 bundle, rules.md#H5): assign alone calls it, at TWO sites — the main walk and the family-comma segment-1 walk — and it is in the leaf rather than inline because each site had been given a cheap frame-free gate written to match the walk's own first condition, which is a second implementation of the question and was removed in review; what the leaf costs is one frame per entry point, measured, and the walk's own first test is a compiled regex rather than a call, so an ordinary name pays a match and stops. Re-measured 2026-09-08 by call site over `_pipeline/*.py`, the whole census above holds unchanged: is_suffix_piece, leading_titles, peel_walk and peel_trailing shared, is_leading_title, is_title_piece and trailing_start group-only — assign still reads `leading_titles` and never `is_leading_title`, which is what keeps H2's shape inference out of the trailing slot. And nameparser/_pipeline/_post_rules.py over a state: suffix_entries, the #511 instance, the R1 entry pass as a function, the one instance living in a stage rather than in a leaf — it is a pass over a whole ParseState and no leaf takes one, and AGENTS.md names it as the exception — run by post_rules last in the stage (through its in-place worker) and by Parser.revise over a sub-parse whose roles it has forced, so a suffix value handed to revise() derives its entries by the rule a whole name uses rather than by a second reading of the value's commas (decisions.md#C1, 2026-09-06 #511). tests/v2/test_layering.py holds each module's contract, and a piece predicate growing a dependency on a STAGE shows up there as a widened entry. Reach for it when. You are about to write a condition that mirrors, matches or "does what X does" — or you find a comment saying one does. Grep for the other site's predicate and call it instead.
Problem shape. Two stages need the same answer about the same input, and the one that does not own the decision is about to test for it. Contract statement. Where two sites ask the same question, exactly one predicate answers it and every other site calls that one — never a condition written to match it. The predicate belongs to the QUESTION, not to whichever stage decides: it may sit in a leaf both stages import, and for the leading-title test it must, since the deciding stage is assign and group cannot import assign. How it works. A hand-written mirror agrees with its original only until one of them moves, and the drift is invisible in both directions: each site keeps passing its own tests while they disagree about an input neither covers. Five instances, every one found as a defect before it was found as a pattern — #319 lifted the wholly-suffix predicate into the vocabulary layer "so the comma decision and the honorific peel's segment test cannot drift apart"; #401/#421 lifted the trailing-numeral fork out of assign so the bound-given reserve stopped carrying a copy, its hand-written mirror having been falsified in review more than once — the lesson recorded there being that what must be mirrored is assign's WALK, not merely its condition; #425 replaced that reserve's hand re-derivation of the trailing peel with one function over the view the join would leave; #424 moved assign's leading-title test down because group's own `title()` does not see H2's unlisted abbreviations, so `Xyz. van Johnson` chained where `Dr. van Johnson` did not; #429 moved the no-name-segment test down because group asked by segment INDEX where assign asks by CONTENT. The destination follows the LAYER, not the topic: a predicate over token text goes to `_vocab`, one over pieces and tags to `_pieces`. Both are leaves the stages sit on. The piece layer got its own module only in #439 — until then those predicates collected in `_group`, not because grouping owned them but because `_assign` imports `_group` and cannot be imported back, so group was the one place both stages could reach; five had accumulated across four PRs before the module existed. Stage order is this mechanism's limit, and it forecloses the alternative: where the reader comes AFTER the decider, record the answer on the state instead — `ParseState.order` is that shape, "Recorded rather than recomputed downstream, because the two can differ" — which is unavailable whenever the EARLIER stage is the one asking. (The concrete assign→group import that forced the `_group` collection is gone since #439; what remains is the ordering it was a symptom of, and tests/v2/test_layering.py is where the leaf's contract is now written down.) The cost is a second evaluation of the same predicate, measured for #429 at 1.2–2.2% of a family-comma parse and 0% of every other; recording that number was the right answer there over plumbing a state field the two sites would not otherwise share. Lives in. nameparser/_pipeline/_vocab.py over text (is_wholly_suffix; is_trailing_numeral_suffix — the #401/#421 instance, whose only caller since #439 is the shared peel rather than a stage; and maiden_marker_run, the #434 instance and the clearest two-stage case, called by classify over token texts and by extract over a clause's whitespace words, with group reading the tags classify recorded because it runs later; and delimiter_cores, the #436/#437 instance, read by group where a tail segment DROPS a configured delimiter core and by post_rules where the suffix view's entry boundary asks whether a dropped token was one, with a third reader inside this same module, is_wholly_suffix, where a configured core counts as suffix-shaped) and nameparser/_pipeline/_pieces.py over pieces: is_suffix_piece, leading_titles, peel_walk and peel_trailing are called by both stages, while is_leading_title, is_title_piece and trailing_start are called by group alone (measured 2026-09-06 by call site: `is_leading_title` has no caller in `_assign.py`, which reads `leading_titles` instead — a first draft of this clause listed it among the shared ones) — `trailing_start` being the one to know, since it answers where the trailing run begins and is what P2's chain and M2's walk stop at — and segment_suffix_reading by assign alone since #436/#437, that last one being #430's instance, where THREE readers shared one answer until the render join, group's third, was replaced by a rule over the commas the writer typed (decisions.md#C1, 2026-09-06); it stays where it is, one call site being no reason to move a predicate that two sites will contest again. `trailing_titles` was that last shape for one day (2026-09-08, the #316/#489 bundle, rules.md#H5) and is a shared one since 2026-09-09: assign calls it at two sites — the main walk and the family-comma segment-1 walk — and group's bound-given reserve at a third, because that reserve counts the name words assign will leave and this walk is half of what leaves them (rules.md#P5; counting a trailing title word among them joined 'Prof. abdul rahman Prof.' where 'Prof. abdul rahman' does not). It is in the leaf rather than inline because each assign site had been given a cheap frame-free gate written to match the walk's own first condition, which is a second implementation of the question and was removed in review; what the leaf costs is one frame per entry point, measured, and the walk's own first test is a compiled regex rather than a call, so an ordinary name pays a match and stops. The reserve's two calls cost the reference name nothing — it never enters that branch, having no bound given word — and the parse and facade frame counts did not move (measured 2026-09-09). Re-measured 2026-09-09 by call site over `_pipeline/*.py`, the rest of the census above holds unchanged: is_suffix_piece, leading_titles, peel_walk, peel_trailing and now trailing_titles shared, is_leading_title, is_title_piece and trailing_start group-only — assign still reads `leading_titles` and never `is_leading_title`, which is what keeps H2's shape inference out of the trailing slot. And nameparser/_pipeline/_post_rules.py over a state: suffix_entries, the #511 instance, the R1 entry pass as a function, the one instance living in a stage rather than in a leaf — it is a pass over a whole ParseState and no leaf takes one, and AGENTS.md names it as the exception — run by post_rules last in the stage (through its in-place worker) and by Parser.revise over a sub-parse whose roles it has forced, so a suffix value handed to revise() derives its entries by the rule a whole name uses rather than by a second reading of the value's commas (decisions.md#C1, 2026-09-06 #511). tests/v2/test_layering.py holds each module's contract, and a piece predicate growing a dependency on a STAGE shows up there as a widened entry. Reach for it when. You are about to write a condition that mirrors, matches or "does what X does" — or you find a comment saying one does. Grep for the other site's predicate and call it instead.

## RENDER-HONORS-THE-PARSE — the parse decides it, the views honor it

Expand Down
27 changes: 20 additions & 7 deletions docs/design/rules.md
Original file line number Diff line number Diff line change
Expand Up @@ -288,7 +288,18 @@ H5. Rationale: a word abbreviated with a period at the END of a name
not a reading a reader would hesitate over — where the doubt is
real it is the word left STANDING that carries it, which is H4's
report and not this rule's.
history: decisions.md#H5 · interacts: H1, H2, H3, H4, S2, C1 · implemented: nameparser/_pipeline/_assign.py, nameparser/_pipeline/_pieces.py
Accepted: the chain reads PIECES, so a join that ran earlier
puts the word out of reach. A particle chain (P2) has already
taken the trailing word into the family name, and a maiden
marker (M2) has already taken it into the maiden name; in
neither is a title word standing in the trailing slot at all.
"John van der Berg Prof." → family="van der Berg Prof."
"Mary Smith née Jones Prof." → maiden="Jones Prof."
Accepted: what the chain leaves is also what counts as a name
word to spare (P5). A trailing title word is not one, so a bound
given-name word behind one joins exactly as it joins with the
title absent.
history: decisions.md#H5 · interacts: H1, H2, H3, H4, M2, P2, P5, S2, C1 · implemented: nameparser/_pipeline/_assign.py, nameparser/_pipeline/_pieces.py

## Particles & surname prefixes (P)

Expand Down Expand Up @@ -421,7 +432,7 @@ P2. Rationale: a particle is written as part of the surname it
(#132's ask) has it as the surnames view rather than the
family field.
"Vincent van Gogh van Beethoven" → surnames="van Gogh van Beethoven"
history: decisions.md#P2 · interacts: P1, P4, M2, S2 · implemented: nameparser/_pipeline/_group.py, nameparser/_pipeline/_post_rules.py
history: decisions.md#P2 · interacts: P1, P4, H5, M2, S2 · implemented: nameparser/_pipeline/_group.py, nameparser/_pipeline/_post_rules.py

P3. Rationale: connective words ("y", "of the") bind name words into
one name part; but a single letter in a short name is more
Expand Down Expand Up @@ -516,9 +527,11 @@ P5. Rationale: some given-name words are incomplete alone — "abdul"
particle's attachment (P6) sees the name. What
there is to spare is what
assign will leave: the join is tried on the pieces as it would
leave them, assign's trailing peel (S2) is read over that, and
the name words it leaves are the words to spare — a trailing
roman numeral, or a bare acronym the peel takes, is no
leave them, assign's trailing peel (S2) is read over that and its
trailing title run (H5) over what that peel leaves, and the name
words the two of them leave are the words to spare — a trailing
roman numeral, or a bare acronym the peel takes, or a trailing
title word the run takes, is no
word to spare. The join joins two name words into one and
changes no suffix reading: a word the peel reads as a suffix
unjoined must read so joined, or the join declines. After a
Expand Down Expand Up @@ -572,7 +585,7 @@ P5. Rationale: some given-name words are incomplete alone — "abdul"
"Sheik abdul salam" family-first → family="abdul salam"
"Sheik abdul salam" family-first → given=""
"Sheik abdul salam" family-first-given-last → family="abdul salam"
history: decisions.md#P5 · interacts: S2, M2, H1, P2, P4, P6 · implemented: nameparser/_pipeline/_group.py, nameparser/_pipeline/_post_rules.py
history: decisions.md#P5 · interacts: S2, M2, H1, H5, P2, P4, P6 · implemented: nameparser/_pipeline/_group.py, nameparser/_pipeline/_post_rules.py

P6. Rationale: a particle ending the name has nothing to link
forward to, so it is not doing a particle's work there. What it
Expand Down Expand Up @@ -968,7 +981,7 @@ M2. Rationale: a maiden marker announces that what follows it is the
is maiden text all the same — the count it needs includes the
very words the marker removes, so the reading is left to assign.
"John née Jones Smith Ma" → maiden="Jones Smith Ma"
history: decisions.md#M2 · interacts: P2, P3, P5, R2, M1, S2, H1 · implemented: nameparser/_pipeline/_group.py
history: decisions.md#M2 · interacts: P2, P3, P5, R2, M1, S2, H1, H5 · implemented: nameparser/_pipeline/_group.py

M3. Rationale: an enclosure says nothing about whether it means
maiden, but a recognized marker word inside it does — the clause
Expand Down
Loading