fix(lang): count letters no script range claims, instead of routing them to English - #169
Conversation
…hem to English `detect_script` counts an alphabetic character only when one of `_SCRIPT_RANGES` claims it. The loop had no `else`, so a character no range covers was counted nowhere: text written only in an unlisted script produced a total of 0, was reported as "unknown", and `analyse` treats "unknown" as English. It therefore reached the English checkpoint -- which has no tokens for it -- with the reason "no letters detected in state", for text that plainly has letters. 92,529 of Unicode's 136,104 alphabetic codepoints (68%) match no range, because the table names only the scripts the checkpoints were measured on. That includes the CJK ideograph extensions (60,331), the kana supplements, bopomofo, halfwidth katakana, hangul jamo extended-A/B, and dozens of smaller scripts. Confirmed on 573e5b6, all reporting unknown / is_english=True / model "english": halfwidth katakana, bopomofo, kana supplement, CJK Ext-B, hangul jamo ext-A, Cherokee, Mongolian, Syriac, Thaana, Tifinagh and Yi. Fullwidth Latin was also read as an unlisted script rather than `latin`. Both `_SCRIPT_RANGES` loops now count an unclaimed character under "other", which is not "latin" and so routes to the multilingual checkpoint -- the safe direction, matching the existing rule that an undecided Latin-script language is not English. Fullwidth Latin (U+FF21-FF3A, U+FF41-FF5A) counts as `latin`. `else` rather than a longer table because the defect is structural: no finite list fixes it, and any list that tries is wrong again at the next Unicode release. Letterless states keep the `unknown` path, so `route` still reports "no letters detected" for empty, digits-only and emoji-only input. tests/test_router.py: 23 new assertions. On 573e5b6 they give 166 passed, 23 failed; on this branch 189 passed, 0 failed. test_criteria.py 34/0, test_email.py 11/0, test_shortlist.py 81/0 and test_decision_model.py unchanged. ruff and compileall clean.
2325939 to
45e8c64
Compare
|
Thanks — agreed on both counts, and there is something you should have before you reconcile with #41: it contains the same core fix as this one, opened a day earlier (2026-09-20), and I did not spot it when I opened this. #41's else:
# Never treat an alphabetic script omitted from the table as an absence of language.
counts["other"] = counts.get("other", 0) + 1It also widens The difference is the case you asked aboutMeasured on
#41 widens the table but leaves the Latin test itself alone ( Proposed reconciliationIf #41 is the base, the fullwidth part is one condition in two places and I am happy to push it onto that branch myself, or to close this one in favour of it — your call. What I would avoid is #41 merging without it, since it is the case you flagged and it would need a second pass. Two practical notes while you decide:
I have rebased this branch onto |
|
Thank you for #168 and for this fix. Counting unclaimed letters under I checked it on real data together with the other routing changes: no English text changed route across 20,000 states. Merging. #41 took the enumeration route to the same goal, and I'll thank them there. |
Widen the Latin range in the shared script ranges to U+02B0 so the schwa 'ə' (U+0259, IPA Extensions) counts as a Latin letter. Uppercase 'Ə' (U+018F) already fell below the cutoff, so the same letter counted in one case and not the other. Since NandhaKishorM#169 the unclaimed 'ə' lands in "other", which reads Azerbaijani as partly non-Latin: "Müştəridən iki dəfə pul alınıb" reports non_latin_fraction 0.1111 and a lone 'ə' reports script "other" at 1.0, for text that is Latin script throughout. Add an Azerbaijani function-word list and the schwa to the non-English diacritic set. The diacritic signal does not reach Azerbaijani typed without diacritics, which is common in support text, so that text still went to the English checkpoint. All 41 words are unique to the list, so it names Azerbaijani only on evidence no other list claims, per the rule added in NandhaKishorM#178. Fold the dotted capital 'İ' to 'i' before lowercasing, after the NandhaKishorM#179 identifier mask. Python lowercases it to 'i' plus a combining dot above, which is not a word character, so 'İLƏ' split into 'i' and 'lə' and matched no word list; uppercase Azerbaijani now scores the same as lowercase. The remaining Azerbaijani letters were already handled correctly in both cases. Add regression cases for the lone schwa, uppercase input with the dotted I, script profiles, text typed without diacritics, multilingual routing and an explicit English override. The routing suite reports 7 failures before the fix and 338 passes after it. A sweep over 23 states changes exactly one route, the de-diacriticised Azerbaijani case; English with links, French, Spanish, Brazilian Portuguese, Italian, German, Romanian, Polish, Turkish, Hindi, Japanese, Armenian, Cherokee, fullwidth Latin and digit-only input are unchanged. Tests ran with a namespace loader and a stub torch that bypass the ML-dependent package initializer: test_router 338, test_email 46 and test_packaging 15 pass. test_criteria, test_decision_model, test_download and test_shortlist need transformers, safetensors or a real package import and were not run, and model inference was not exercised. ruff was not available locally. Byte compilation and git diff --check pass.
Widen the Latin range in the shared script ranges to U+02B0 so the schwa 'ə' (U+0259, IPA Extensions) counts as a Latin letter. Uppercase 'Ə' (U+018F) already fell below the cutoff, so the same letter counted in one case and not the other. Since #169 the unclaimed 'ə' lands in "other", which reads Azerbaijani as partly non-Latin: "Müştəridən iki dəfə pul alınıb" reports non_latin_fraction 0.1111 and a lone 'ə' reports script "other" at 1.0, for text that is Latin script throughout. Add an Azerbaijani function-word list and the schwa to the non-English diacritic set. The diacritic signal does not reach Azerbaijani typed without diacritics, which is common in support text, so that text still went to the English checkpoint. All 41 words are unique to the list, so it names Azerbaijani only on evidence no other list claims, per the rule added in #178. Fold the dotted capital 'İ' to 'i' before lowercasing, after the #179 identifier mask. Python lowercases it to 'i' plus a combining dot above, which is not a word character, so 'İLƏ' split into 'i' and 'lə' and matched no word list; uppercase Azerbaijani now scores the same as lowercase. The remaining Azerbaijani letters were already handled correctly in both cases. Add regression cases for the lone schwa, uppercase input with the dotted I, script profiles, text typed without diacritics, multilingual routing and an explicit English override. The routing suite reports 7 failures before the fix and 338 passes after it. A sweep over 23 states changes exactly one route, the de-diacriticised Azerbaijani case; English with links, French, Spanish, Brazilian Portuguese, Italian, German, Romanian, Polish, Turkish, Hindi, Japanese, Armenian, Cherokee, fullwidth Latin and digit-only input are unchanged. Tests ran with a namespace loader and a stub torch that bypass the ML-dependent package initializer: test_router 338, test_email 46 and test_packaging 15 pass. test_criteria, test_decision_model, test_download and test_shortlist need transformers, safetensors or a real package import and were not run, and model inference was not exercised. ruff was not available locally. Byte compilation and git diff --check pass. Co-authored-by: Nandakishor <48623612+NandhaKishorM@users.noreply.github.com>
…hem to English (NandhaKishorM#169) `detect_script` counts an alphabetic character only when one of `_SCRIPT_RANGES` claims it. The loop had no `else`, so a character no range covers was counted nowhere: text written only in an unlisted script produced a total of 0, was reported as "unknown", and `analyse` treats "unknown" as English. It therefore reached the English checkpoint -- which has no tokens for it -- with the reason "no letters detected in state", for text that plainly has letters. 92,529 of Unicode's 136,104 alphabetic codepoints (68%) match no range, because the table names only the scripts the checkpoints were measured on. That includes the CJK ideograph extensions (60,331), the kana supplements, bopomofo, halfwidth katakana, hangul jamo extended-A/B, and dozens of smaller scripts. Confirmed on 573e5b6, all reporting unknown / is_english=True / model "english": halfwidth katakana, bopomofo, kana supplement, CJK Ext-B, hangul jamo ext-A, Cherokee, Mongolian, Syriac, Thaana, Tifinagh and Yi. Fullwidth Latin was also read as an unlisted script rather than `latin`. Both `_SCRIPT_RANGES` loops now count an unclaimed character under "other", which is not "latin" and so routes to the multilingual checkpoint -- the safe direction, matching the existing rule that an undecided Latin-script language is not English. Fullwidth Latin (U+FF21-FF3A, U+FF41-FF5A) counts as `latin`. `else` rather than a longer table because the defect is structural: no finite list fixes it, and any list that tries is wrong again at the next Unicode release. Letterless states keep the `unknown` path, so `route` still reports "no letters detected" for empty, digits-only and emoji-only input. tests/test_router.py: 23 new assertions. On 573e5b6 they give 166 passed, 23 failed; on this branch 189 passed, 0 failed. test_criteria.py 34/0, test_email.py 11/0, test_shortlist.py 81/0 and test_decision_model.py unchanged. ruff and compileall clean. Co-authored-by: PerryLink <255665900+PerryLink@users.noreply.github.com>
…shorM#36) Widen the Latin range in the shared script ranges to U+02B0 so the schwa 'ə' (U+0259, IPA Extensions) counts as a Latin letter. Uppercase 'Ə' (U+018F) already fell below the cutoff, so the same letter counted in one case and not the other. Since NandhaKishorM#169 the unclaimed 'ə' lands in "other", which reads Azerbaijani as partly non-Latin: "Müştəridən iki dəfə pul alınıb" reports non_latin_fraction 0.1111 and a lone 'ə' reports script "other" at 1.0, for text that is Latin script throughout. Add an Azerbaijani function-word list and the schwa to the non-English diacritic set. The diacritic signal does not reach Azerbaijani typed without diacritics, which is common in support text, so that text still went to the English checkpoint. All 41 words are unique to the list, so it names Azerbaijani only on evidence no other list claims, per the rule added in NandhaKishorM#178. Fold the dotted capital 'İ' to 'i' before lowercasing, after the NandhaKishorM#179 identifier mask. Python lowercases it to 'i' plus a combining dot above, which is not a word character, so 'İLƏ' split into 'i' and 'lə' and matched no word list; uppercase Azerbaijani now scores the same as lowercase. The remaining Azerbaijani letters were already handled correctly in both cases. Add regression cases for the lone schwa, uppercase input with the dotted I, script profiles, text typed without diacritics, multilingual routing and an explicit English override. The routing suite reports 7 failures before the fix and 338 passes after it. A sweep over 23 states changes exactly one route, the de-diacriticised Azerbaijani case; English with links, French, Spanish, Brazilian Portuguese, Italian, German, Romanian, Polish, Turkish, Hindi, Japanese, Armenian, Cherokee, fullwidth Latin and digit-only input are unchanged. Tests ran with a namespace loader and a stub torch that bypass the ML-dependent package initializer: test_router 338, test_email 46 and test_packaging 15 pass. test_criteria, test_decision_model, test_download and test_shortlist need transformers, safetensors or a real package import and were not run, and model inference was not exercised. ruff was not available locally. Byte compilation and git diff --check pass. Co-authored-by: Nandakishor <48623612+NandhaKishorM@users.noreply.github.com>
Fixes #168
Text written in a script that
_SCRIPT_RANGESdoes not list is routed to the English checkpoint.detect_scriptcounts an alphabetic character only when a range claims it, and the loop has noelse, so an unclaimed character is counted nowhere:analysemapsscript == "unknown"tois_english: True(lang.py:218-221), androutesends that to the default checkpoint (router.py:292-294). The English checkpoint has no tokens for these scripts.This is not a rare corner. 92,529 of Unicode's 136,104 alphabetic codepoints (68%) match no range, because the list names only the scripts the checkpoints were measured on. Confirmed on
main, all reportingunknown/is_english=True/english:アリガトウhalfwidth katakanaㄆㄇㄈㄉbopomofoHELLOfullwidth Latinlatindetect_script's own docstring says it returns'unknown'"if there are no letters"."アリガトウ"has five.Change
laya/lang.py, two loops, 16 added / 3 removed lines:_SCRIPT_RANGESloops gain anelse, so an alphabetic character no range claims is counted under"other"instead of vanishing."other"is not"latin", so such text is non-Latin and routes to the multilingual checkpoint — the safe direction, consistent withanalyse's existing rule that an undecided Latin-script language is not English.U+FF21-FF3A,U+FF41-FF5A) is counted aslatinin both loops, soHELLOis Latin rather than unlisted.Two decisions worth naming
Why
elserather than extending the table. I counted what extending it would take: the unclassified codepoints sit in dozens of Unicode blocks, and the CJK ideograph extensions alone are 60,331 of them (Ext-B 42,720; Ext-C 4,154; Ext-D 222; Ext-E 5,762; Ext-F 7,473). The defect is structural — a character that no range claims must still be counted — so no finite extension fixes it, and any list that tries will be wrong again at the next Unicode release.Why
"other"rather than a named script. It is deliberately coarse: it answers "not Latin, not one of the scripts we name" and nothing more. Naming the script would need the full block table this change exists to avoid. Callers who need the name for a specific script should extend_SCRIPT_RANGESfor that script, which this change does not prevent.What does not move
latin→ englishunknown→ default (english)unknown→ englishlatin→ englishLetterless states keep the
unknownpath, soroutestill reports "no letters detected" for them.Verification
tests/test_router.py, with the new cases inserted before the report block:The 23 failures are the 11 unlisted scripts × 2 assertions, plus the fullwidth-Latin case. Full suite on this branch (torch 2.14.0+cpu, transformers 5.17.0, Python 3.12):
ruff check laya/andpython -m compileall laya/are clean.Diff:
laya/lang.py+16/-3,tests/test_router.py+47 (new cases only).Known limitation
I could not measure routing accuracy for these scripts end to end: that needs the checkpoints, which I did not download. What is verified is the routing decision itself —
is_englishandRouter().route(...)["model"]— not the multilingual checkpoint's accuracy on Cherokee or Yi. The claim here is only that the English checkpoint cannot read them and the multilingual one is the checkpoint selected for unreadable scripts, which is the same rulerouter.pyalready applies to unidentified Latin text."other"also meansscript_profilenow reports a key that is not a script name, soanalyse(...)["script"]can be"other". I do not think that is worth a second field, but it is a visible API change and worth flagging rather than burying.