Skip to content
Merged
Show file tree
Hide file tree
Changes from 1 commit
Commits
Show all changes
36 commits
Select commit Hold shift + click to select a range
4f5b956
Add the glued-honorific tail vocabulary (#308)
derek73 Jul 31, 2026
2b0f007
Vet the glued tail set against real surname data (#308)
derek73 Jul 31, 2026
b9c58a2
Add Lexicon.honorific_tails, the glued-peel vocabulary (#308)
derek73 Jul 31, 2026
b72150a
Enforce the honorific-tail subset invariant (#308)
derek73 Jul 31, 2026
9075fde
Peel a glued CJK honorific off the last name token (#308)
derek73 Jul 31, 2026
17f9c39
Never treat a post-nominal as a name site (#308)
derek73 Jul 31, 2026
be97aec
Decline rather than scan past a post-nominal surname site (#308)
derek73 Jul 31, 2026
90be6ba
Pin the suffix-comma peel and name the maiden reach (#308)
derek73 Jul 31, 2026
59aac1f
Let a recognized honorific reach the segmenter, not block it (#308)
derek73 Jul 31, 2026
f71a2e4
Pin that an honorific does not excuse a real boundary (#308)
derek73 Jul 31, 2026
d56fff2
Classify the glued-honorific diffs in the differential (#308)
derek73 Jul 31, 2026
aba6c03
Document the glued honorific peel (#308)
derek73 Jul 31, 2026
1d46fc6
Say what the 君 row actually pins (#308)
derek73 Jul 31, 2026
73fb92f
Finish the #308 staleness sweep (#308)
derek73 Jul 31, 2026
fdd3527
Scope the doctrine paragraph's count to what it counts (#308)
derek73 Jul 31, 2026
227426d
Note the lone-honorific fix and the peel's off-switch (#308)
derek73 Jul 31, 2026
063fd1a
Exempt only the tail the peel manufactured (#308)
derek73 Jul 31, 2026
931ce0e
Ship 박사님, the missing third -님 honorific (#308)
derek73 Jul 31, 2026
cd626cb
Correct six claims the code makes about itself (#308)
derek73 Jul 31, 2026
a527bd6
Pin the v1 shim's honorific_tails translation (#308)
derek73 Jul 31, 2026
288fef2
Say what the peel actually does, in prose (#308)
derek73 Jul 31, 2026
acbfefc
Pin the peel's segment choice and its two boundaries (#308)
derek73 Jul 31, 2026
3e5dad6
Reach the peel from a drawn lexicon (#308)
derek73 Jul 31, 2026
b31f759
Replace the peel docstring's self-refuting no-peel-twice example (#308)
derek73 Jul 31, 2026
dde0c60
Reclassify the FAMILY-comma honorific row to comma-family (#308)
derek73 Jul 31, 2026
d6cd806
Justify the segmenter's spaced-honorific rule by its cost (#308)
derek73 Jul 31, 2026
7819a3d
Scope the spaced-honorific advice to the segmenter path (#308)
derek73 Jul 31, 2026
2189401
Pin _is_post_nominal's strict suffix test with the input it names (#308)
derek73 Jul 31, 2026
b8a9252
Give the spaced-honorific test its missing control (#308)
derek73 Jul 31, 2026
e6a6b69
Correct five accuracy nits around the peel (#308)
derek73 Jul 31, 2026
41588b4
Let _split tag the tail it just built (#308)
derek73 Jul 31, 2026
2c5edcc
Derive the peel test lexicon from the stage's own (#308)
derek73 Jul 31, 2026
6ae3351
Drop a test whose assertion cannot fail (#308)
derek73 Jul 31, 2026
3171267
Pin the 殿 exclusion, the one nothing held (#308)
derek73 Jul 31, 2026
46fdc3e
Put a non-ASCII shape in the scaling benchmark (#308)
derek73 Jul 31, 2026
53611b1
Name the escape hatch that actually works (#308)
derek73 Jul 31, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Prev Previous commit
Next Next commit
Pin the v1 shim's honorific_tails translation (#308)
honorific_tails appeared nowhere in tests/v2/test_config_shim.py, and
the snapshot-equality test passes VACUOUSLY here -- with the default
config the GLUED_HONORIFICS ∩ suffix_words intersection is a no-op,
which is exactly the failure mode AGENTS.md warns about. Pin the case
the translation decides: deleting a CJK honorific from
Constants.suffix_not_acronyms, the only v1 knob that reaches the
field, turns the peel off rather than raising the subset error or
leaving a tail that peels into a suffix field no longer recognizing
it.

AGENTS.md's roster of shim transformations said four and named four;
this branch added the fifth. Same-commit rule applied one commit late.
  • Loading branch information
derek73 committed Jul 31, 2026
commit a527bd6424a9602134a308b8ae18cdcf66c4fae5
2 changes: 1 addition & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -131,7 +131,7 @@ The 2.0 rewrite lands as underscore-private modules alongside the v1 code. These
- **Parser owns config-dependent conveniences**: `Parser.matches`/`Parser.capitalized`/`Parser.revise` exist because the `ParsedName` equivalents fall back to DEFAULT config for str/omitted arguments (documented loudly in both docstrings). `revise` harvests tokens from a full sub-parse of each replacement value (tags kept minus `FOLDED_TAG`, roles forced, ambiguities discarded); the merge tail is shared with `replace()` via `ParsedName._with_field_tokens`. `Parser.capitalized` delegates through `name.capitalized(self.lexicon)` specifically so `_parser` never imports `_render` — keep it that way.
- **Per-word vocabulary fields warn on multi-word entries** (`_normset`/`_normpairs` via `_warn_dead_entry`, UserWarning, never a raise — see the given_name_titles Gotcha for why raising is wrong). `given_name_titles` is the one multi-word-matched field and is exempt; `_edit` passes `warn=False` (add() warns once via the new instance's `__post_init__`; remove() stores nothing). The default vocabulary and every locale pack must stay warning-free (`test_default_lexicon_builds_warning_free`, `test_pack_vocabulary_entries_are_single_words`).
- **Invariants guard harm, not no-ops**: add a constructor check when violating it produces a *wrong parse*, not when it produces *nothing*. A false positive costs a working configuration; a true positive on an inert condition costs the user nothing, so that trade is never worth taking. `suffix_acronyms_ambiguous ∩ suffix_words` is guarded because the overlap loses a family name; `given_name_titles` is not, because an unreachable entry is simply never consulted (see Gotchas). Before adding one, construct the config it forbids and check what actually breaks.
- **The shim TRANSLATES; it never raises on a config v1 accepted, and never silently changes the parse**: `Constants._snapshot()` is a translation boundary between v1's model and v2's invariants, and every transformation there carries its v1-reachability argument in a comment. Four exist today — `first_name_titles` re-folded per word (v1 joins-then-`lc`, v2 normalizes-then-joins), `suffix_acronyms_ambiguous ∩ acronyms` (a provable no-op), `suffix_words − ambiguous` (v1 already accepts the word via the acronym branch, so the addition is inert there), and `particles_ambiguous ∪ (bound ∩ particles)` (a pinned deviation, `test_bound_never_given_prefix_deviates_on_two_pieces`). When a v1 config cannot satisfy a v2 invariant, work out what v1 actually *does* with it — usually nothing — and reproduce that; weakening the invariant or letting the raise through are both wrong. **Test the case the translation decides**, not one where both branches agree: a test using an input v1 parses identically with and without the config pins nothing.
- **The shim TRANSLATES; it never raises on a config v1 accepted, and never silently changes the parse**: `Constants._snapshot()` is a translation boundary between v1's model and v2's invariants, and every transformation there carries its v1-reachability argument in a comment. Five exist today — `first_name_titles` re-folded per word (v1 joins-then-`lc`, v2 normalizes-then-joins), `suffix_acronyms_ambiguous ∩ acronyms` (a provable no-op), `suffix_words − ambiguous` (v1 already accepts the word via the acronym branch, so the addition is inert there), `particles_ambiguous ∪ (bound ∩ particles)` (a pinned deviation, `test_bound_never_given_prefix_deviates_on_two_pieces`), and `honorific_tails = GLUED_HONORIFICS ∩ suffix_words` (#308 behavior with no v1 manager of its own, so the one v1 knob that reaches it is deleting the suffix word — which turns the peel off, `test_snapshot_removing_a_honorific_word_turns_the_peel_off`). When a v1 config cannot satisfy a v2 invariant, work out what v1 actually *does* with it — usually nothing — and reproduce that; weakening the invariant or letting the raise through are both wrong. **Test the case the translation decides**, not one where both branches agree: a test using an input v1 parses identically with and without the config pins nothing.
- **Reprs are bounded**: render which fields deviate from a named baseline and by how much, never contents (`Lexicon(default + titles: +2)`). `PolicyPatch`'s repr shows only set (non-UNSET) fields; `_order_repr` must never raise even on an unvalidated patch's garbage `name_order` (PolicyPatch defers validation to apply time); the sweep test in `tests/v2/test_reprs.py` pins that no config repr leaks the UNSET sentinel.
- **Typing/docs**: `from __future__ import annotations`; `frozen=True, slots=True` on every public dataclass; strict-profile mypy flags via per-module overrides in pyproject (`strict = true` itself is not valid per-module). Docstrings state contracts in prose with **no doctest blocks** — `--doctest-modules` makes every example a test; behavior examples go to unit tests per the lean-docs rule. **Document the positive direction of a partial property**: "a non-empty `ambiguities` is a signal to act on" is checkable, while "an empty one means no fork occurred" is a universal negative needing exhaustive verification -- that claim was written twice and falsified twice, at sites the author had not audited.
- **The segmenter contract**: the optional `Parser(segmenter=...)` hook is parse-totality's ONE exception (locales spec section 4). Everything inside that exception is a bug in USER CODE, never a fact about the name, so it is surfaced rather than absorbed: the segmenter's own exceptions propagate, and the two protocol violations the stage can detect for itself — an answer of the wrong type, and one cutting at or past the end of the token it was handed — raise `TypeError`/`ValueError` from `_script_segment` for the same reason. The line to hold when adding a check there: a protocol violation by the segmenter's AUTHOR raises, while an adapter's defense against its own third-party library (`locales/ja.py`'s repertoire, length, reconstruction and score guards) declines with `None`, because what those catch is a fact about the content.
Expand Down
20 changes: 20 additions & 0 deletions tests/v2/test_config_shim.py
Original file line number Diff line number Diff line change
Expand Up @@ -476,6 +476,26 @@ def test_snapshot_keeps_a_bound_never_given_prefix_parseable() -> None:
assert (name.first, name.last) == ("dos Santos", "Silva")


def test_snapshot_removing_a_honorific_word_turns_the_peel_off() -> None:
# The deciding case for honorific_tails, which has no v1 manager of
# its own: the snapshot intersects GLUED_HONORIFICS with the WORD
# set, so deleting 씨 from suffix_not_acronyms -- the only v1 knob
# that reaches it -- must make the tail stop mattering rather than
# raise the subset error or leave 씨 peeling into a suffix field
# that no longer recognizes it. With the default config the
# intersection is a no-op, so the equality test above pins nothing
# here.
c = Constants()
assert HumanName("김민준씨", constants=c).suffix == "씨" # baseline
c.suffix_not_acronyms.remove("씨")
lexicon, _, _ = c._snapshot() # must not raise
assert "씨" not in lexicon.honorific_tails
# the peel is off: the glued honorific goes back into the name,
# which is 2.0's answer for this input
name = HumanName("김민준씨", constants=c)
assert (name.first, name.last, name.suffix) == ("민준씨", "김", "")


def test_snapshot_field_translation() -> None:
c = Constants()
lexicon, policy, defaults = c._snapshot()
Expand Down