Skip to content

fix(lang): count letters no script range claims, instead of routing them to English - #169

Merged
NandhaKishorM merged 1 commit into
NandhaKishorM:mainfrom
PerryLink:fix/unlisted-script-routes-to-english
Sep 23, 2026
Merged

NandhaKishorM merged 1 commit into
NandhaKishorM:mainfrom
PerryLink:fix/unlisted-script-routes-to-english

Conversation

@PerryLink

Copy link
Copy Markdown
Contributor

Fixes #168

Text written in a script that _SCRIPT_RANGES does not list is routed to the English checkpoint. detect_script counts an alphabetic character only when a range claims it, and the loop has no else, so an unclaimed character is counted nowhere:

from laya import Router, detect_script

text = "アリガトウ"                       # halfwidth katakana: five letters
detect_script(text)                    # 'unknown'
Router().route(text)["model"]          # 'english'
Router().route(text)["reason"]
# 'no letters detected in state; using default (english)'

analyse maps script == "unknown" to is_english: True (lang.py:218-221), and route sends that to the default checkpoint (router.py:292-294). The English checkpoint has no tokens for these scripts.

This is not a rare corner. 92,529 of Unicode's 136,104 alphabetic codepoints (68%) match no range, because the list names only the scripts the checkpoints were measured on. Confirmed on main, all reporting unknown / is_english=True / english:

input script
アリガトウ halfwidth katakana Japanese, extremely common in mobile text wrong
ㄆㄇㄈㄉ bopomofo wrong
kana supplement, small kana extension wrong
CJK Ext-B / C / D / E / F (60,331 codepoints) wrong
hangul jamo extended-A / B wrong
Cherokee, Mongolian, Syriac, Thaana, N'Ko, Tifinagh, Yi, Vai, Javanese, Balinese, Sundanese, Batak, Lepcha, Glagolitic, Coptic, Runic, Ogham, Deseret, Osage, Adlam, Canadian syllabics wrong
HELLO fullwidth Latin reads as an unlisted script instead of latin wrong

detect_script's own docstring says it returns 'unknown' "if there are no letters". "アリガトウ" has five.

Change

laya/lang.py, two loops, 16 added / 3 removed lines:

  • Both _SCRIPT_RANGES loops gain an else, so an alphabetic character no range claims is counted under "other" instead of vanishing. "other" is not "latin", so such text is non-Latin and routes to the multilingual checkpoint — the safe direction, consistent with analyse's existing rule that an undecided Latin-script language is not English.
  • Fullwidth Latin (U+FF21-FF3A, U+FF41-FF5A) is counted as latin in both loops, so HELLO is Latin rather than unlisted.

Two decisions worth naming

Why else rather than extending the table. I counted what extending it would take: the unclassified codepoints sit in dozens of Unicode blocks, and the CJK ideograph extensions alone are 60,331 of them (Ext-B 42,720; Ext-C 4,154; Ext-D 222; Ext-E 5,762; Ext-F 7,473). The defect is structural — a character that no range claims must still be counted — so no finite extension fixes it, and any list that tries will be wrong again at the next Unicode release.

Why "other" rather than a named script. It is deliberately coarse: it answers "not Latin, not one of the scripts we name" and nothing more. Naming the script would need the full block table this change exists to avoid. Callers who need the name for a specific script should extend _SCRIPT_RANGES for that script, which this change does not prevent.

What does not move

before after
English / German / French text latin → english unchanged
Hindi / Bengali / Tamil (devanagari, bengali, tamil) named → multilingual unchanged
Chinese (han), Japanese (kana), Korean (hangul) named → multilingual unchanged
Russian (cyrillic), Arabic, Hebrew, Greek, Armenian, Thai, Khmer, Georgian, Ethiopic, Sinhala, Myanmar named → multilingual unchanged
empty / digits-only / emoji-only states unknown → default (english) unchanged
fullwidth Latin unknown → english latin → english

Letterless states keep the unknown path, so route still reports "no letters detected" for them.

Verification

tests/test_router.py, with the new cases inserted before the report block:

main  (573e5b6)     166 passed, 23 failed
this branch         189 passed, 0 failed

The 23 failures are the 11 unlisted scripts × 2 assertions, plus the fullwidth-Latin case. Full suite on this branch (torch 2.14.0+cpu, transformers 5.17.0, Python 3.12):

python tests/test_router.py           ->  189 passed, 0 failed
python tests/test_criteria.py         ->   34 passed, 0 failed
python tests/test_email.py            ->   11 passed, 0 failed
python tests/test_shortlist.py        ->   81 passed, 0 failed
python tests/test_decision_model.py   ->  all decision model tests passed

ruff check laya/ and python -m compileall laya/ are clean.

Diff: laya/lang.py +16/-3, tests/test_router.py +47 (new cases only).

Known limitation

I could not measure routing accuracy for these scripts end to end: that needs the checkpoints, which I did not download. What is verified is the routing decision itself — is_english and Router().route(...)["model"] — not the multilingual checkpoint's accuracy on Cherokee or Yi. The claim here is only that the English checkpoint cannot read them and the multilingual one is the checkpoint selected for unreadable scripts, which is the same rule router.py already applies to unidentified Latin text.

"other" also means script_profile now reports a key that is not a script name, so analyse(...)["script"] can be "other". I do not think that is worth a second field, but it is a visible API change and worth flagging rather than burying.

…hem to English

`detect_script` counts an alphabetic character only when one of `_SCRIPT_RANGES`
claims it. The loop had no `else`, so a character no range covers was counted
nowhere: text written only in an unlisted script produced a total of 0, was
reported as "unknown", and `analyse` treats "unknown" as English. It therefore
reached the English checkpoint -- which has no tokens for it -- with the reason
"no letters detected in state", for text that plainly has letters.

92,529 of Unicode's 136,104 alphabetic codepoints (68%) match no range, because
the table names only the scripts the checkpoints were measured on. That includes
the CJK ideograph extensions (60,331), the kana supplements, bopomofo, halfwidth
katakana, hangul jamo extended-A/B, and dozens of smaller scripts.

Confirmed on 573e5b6, all reporting unknown / is_english=True / model "english":
halfwidth katakana, bopomofo, kana supplement, CJK Ext-B, hangul jamo ext-A,
Cherokee, Mongolian, Syriac, Thaana, Tifinagh and Yi. Fullwidth Latin was also
read as an unlisted script rather than `latin`.

Both `_SCRIPT_RANGES` loops now count an unclaimed character under "other",
which is not "latin" and so routes to the multilingual checkpoint -- the safe
direction, matching the existing rule that an undecided Latin-script language is
not English. Fullwidth Latin (U+FF21-FF3A, U+FF41-FF5A) counts as `latin`.

`else` rather than a longer table because the defect is structural: no finite
list fixes it, and any list that tries is wrong again at the next Unicode release.

Letterless states keep the `unknown` path, so `route` still reports "no letters
detected" for empty, digits-only and emoji-only input.

tests/test_router.py: 23 new assertions. On 573e5b6 they give 166 passed,
23 failed; on this branch 189 passed, 0 failed. test_criteria.py 34/0,
test_email.py 11/0, test_shortlist.py 81/0 and test_decision_model.py unchanged.
ruff and compileall clean.
@PerryLink
PerryLink force-pushed the fix/unlisted-script-routes-to-english branch from 2325939 to 45e8c64 Compare September 23, 2026 04:11
@PerryLink

Copy link
Copy Markdown
Contributor Author

Thanks — agreed on both counts, and there is something you should have before you reconcile with #41: it contains the same core fix as this one, opened a day earlier (2026-09-20), and I did not spot it when I opened this.

#41's laya/lang.py hunk is the same else clause, with the same intent in the comment:

        else:
            # Never treat an alphabetic script omitted from the table as an absence of language.
            counts["other"] = counts.get("other", 0) + 1

It also widens _SCRIPT_RANGES with named families (syriac, thaana, nko, cherokee, canadian_aboriginal, mongolian, tifinagh, bopomofo, yi, javanese) and adds halfwidth katakana to kana — and its tests are better than mine, asserting the script name per family where mine only assert other. The else clause is rajin-khan's, not mine, and I would rather say so than have it found later.

The difference is the case you asked about

Measured on 573e5b6, each variant in its own process so no import cache can leak between them:

input main #41 this branch
HELLO fullwidth Latin unknown → english other → multilingual latin → english
コンニチハ halfwidth katakana unknown → english kana → multilingual other → multilingual
ᎣᏏᎳᎩ Cherokee unknown → english cherokee → multilingual other → multilingual
𠀀𠀁 CJK Ext-B unknown → english other → multilingual other → multilingual

#41 widens the table but leaves the Latin test itself alone (cp < 0x0250 or 0x1E00 <= cp <= 0x1EFF), so fullwidth Latin still lands in other. This branch widens the Latin test, which is the normalisation you asked for. Neither is a superset: #41 names more scripts, this one gets fullwidth right.

Proposed reconciliation

If #41 is the base, the fullwidth part is one condition in two places and I am happy to push it onto that branch myself, or to close this one in favour of it — your call. What I would avoid is #41 merging without it, since it is the case you flagged and it would need a second pass.

Two practical notes while you decide:

I have rebased this branch onto c752770, so it now sits on top of #178. test_router.py is 240 passed, 0 failed; test_criteria.py 34, test_email.py 11, test_shortlist.py 81; ruff and compileall clean.

@NandhaKishorM

Copy link
Copy Markdown
Owner

Thank you for #168 and for this fix. Counting unclaimed letters under other is the maintainable answer: it covers the whole 68% of Unicode you measured instead of enumerating scripts one by one, and it makes detect_script match its own docstring. Fullwidth Latin now reads as Latin too.

I checked it on real data together with the other routing changes: no English text changed route across 20,000 states. Merging. #41 took the enumeration route to the same goal, and I'll thank them there.

@NandhaKishorM
NandhaKishorM merged commit 129007f into NandhaKishorM:main Sep 23, 2026
10 checks passed
ch4ki added a commit to ch4ki/laya that referenced this pull request Sep 23, 2026
Widen the Latin range in the shared script ranges to U+02B0 so the schwa 'ə' (U+0259, IPA Extensions) counts as a Latin letter. Uppercase 'Ə' (U+018F) already fell below the cutoff, so the same letter counted in one case and not the other. Since NandhaKishorM#169 the unclaimed 'ə' lands in "other", which reads Azerbaijani as partly non-Latin: "Müştəridən iki dəfə pul alınıb" reports non_latin_fraction 0.1111 and a lone 'ə' reports script "other" at 1.0, for text that is Latin script throughout.

Add an Azerbaijani function-word list and the schwa to the non-English diacritic set. The diacritic signal does not reach Azerbaijani typed without diacritics, which is common in support text, so that text still went to the English checkpoint. All 41 words are unique to the list, so it names Azerbaijani only on evidence no other list claims, per the rule added in NandhaKishorM#178.

Fold the dotted capital 'İ' to 'i' before lowercasing, after the NandhaKishorM#179 identifier mask. Python lowercases it to 'i' plus a combining dot above, which is not a word character, so 'İLƏ' split into 'i' and 'lə' and matched no word list; uppercase Azerbaijani now scores the same as lowercase. The remaining Azerbaijani letters were already handled correctly in both cases.

Add regression cases for the lone schwa, uppercase input with the dotted I, script profiles, text typed without diacritics, multilingual routing and an explicit English override. The routing suite reports 7 failures before the fix and 338 passes after it. A sweep over 23 states changes exactly one route, the de-diacriticised Azerbaijani case; English with links, French, Spanish, Brazilian Portuguese, Italian, German, Romanian, Polish, Turkish, Hindi, Japanese, Armenian, Cherokee, fullwidth Latin and digit-only input are unchanged.

Tests ran with a namespace loader and a stub torch that bypass the ML-dependent package initializer: test_router 338, test_email 46 and test_packaging 15 pass. test_criteria, test_decision_model, test_download and test_shortlist need transformers, safetensors or a real package import and were not run, and model inference was not exercised. ruff was not available locally. Byte compilation and git diff --check pass.
NandhaKishorM added a commit that referenced this pull request Sep 23, 2026
Widen the Latin range in the shared script ranges to U+02B0 so the schwa 'ə' (U+0259, IPA Extensions) counts as a Latin letter. Uppercase 'Ə' (U+018F) already fell below the cutoff, so the same letter counted in one case and not the other. Since #169 the unclaimed 'ə' lands in "other", which reads Azerbaijani as partly non-Latin: "Müştəridən iki dəfə pul alınıb" reports non_latin_fraction 0.1111 and a lone 'ə' reports script "other" at 1.0, for text that is Latin script throughout.

Add an Azerbaijani function-word list and the schwa to the non-English diacritic set. The diacritic signal does not reach Azerbaijani typed without diacritics, which is common in support text, so that text still went to the English checkpoint. All 41 words are unique to the list, so it names Azerbaijani only on evidence no other list claims, per the rule added in #178.

Fold the dotted capital 'İ' to 'i' before lowercasing, after the #179 identifier mask. Python lowercases it to 'i' plus a combining dot above, which is not a word character, so 'İLƏ' split into 'i' and 'lə' and matched no word list; uppercase Azerbaijani now scores the same as lowercase. The remaining Azerbaijani letters were already handled correctly in both cases.

Add regression cases for the lone schwa, uppercase input with the dotted I, script profiles, text typed without diacritics, multilingual routing and an explicit English override. The routing suite reports 7 failures before the fix and 338 passes after it. A sweep over 23 states changes exactly one route, the de-diacriticised Azerbaijani case; English with links, French, Spanish, Brazilian Portuguese, Italian, German, Romanian, Polish, Turkish, Hindi, Japanese, Armenian, Cherokee, fullwidth Latin and digit-only input are unchanged.

Tests ran with a namespace loader and a stub torch that bypass the ML-dependent package initializer: test_router 338, test_email 46 and test_packaging 15 pass. test_criteria, test_decision_model, test_download and test_shortlist need transformers, safetensors or a real package import and were not run, and model inference was not exercised. ruff was not available locally. Byte compilation and git diff --check pass.

Co-authored-by: Nandakishor <48623612+NandhaKishorM@users.noreply.github.com>
baninaveen pushed a commit to baninaveen/laya that referenced this pull request Sep 24, 2026
…hem to English (NandhaKishorM#169)

`detect_script` counts an alphabetic character only when one of `_SCRIPT_RANGES`
claims it. The loop had no `else`, so a character no range covers was counted
nowhere: text written only in an unlisted script produced a total of 0, was
reported as "unknown", and `analyse` treats "unknown" as English. It therefore
reached the English checkpoint -- which has no tokens for it -- with the reason
"no letters detected in state", for text that plainly has letters.

92,529 of Unicode's 136,104 alphabetic codepoints (68%) match no range, because
the table names only the scripts the checkpoints were measured on. That includes
the CJK ideograph extensions (60,331), the kana supplements, bopomofo, halfwidth
katakana, hangul jamo extended-A/B, and dozens of smaller scripts.

Confirmed on 573e5b6, all reporting unknown / is_english=True / model "english":
halfwidth katakana, bopomofo, kana supplement, CJK Ext-B, hangul jamo ext-A,
Cherokee, Mongolian, Syriac, Thaana, Tifinagh and Yi. Fullwidth Latin was also
read as an unlisted script rather than `latin`.

Both `_SCRIPT_RANGES` loops now count an unclaimed character under "other",
which is not "latin" and so routes to the multilingual checkpoint -- the safe
direction, matching the existing rule that an undecided Latin-script language is
not English. Fullwidth Latin (U+FF21-FF3A, U+FF41-FF5A) counts as `latin`.

`else` rather than a longer table because the defect is structural: no finite
list fixes it, and any list that tries is wrong again at the next Unicode release.

Letterless states keep the `unknown` path, so `route` still reports "no letters
detected" for empty, digits-only and emoji-only input.

tests/test_router.py: 23 new assertions. On 573e5b6 they give 166 passed,
23 failed; on this branch 189 passed, 0 failed. test_criteria.py 34/0,
test_email.py 11/0, test_shortlist.py 81/0 and test_decision_model.py unchanged.
ruff and compileall clean.

Co-authored-by: PerryLink <255665900+PerryLink@users.noreply.github.com>
baninaveen pushed a commit to baninaveen/laya that referenced this pull request Sep 24, 2026
…shorM#36)

Widen the Latin range in the shared script ranges to U+02B0 so the schwa 'ə' (U+0259, IPA Extensions) counts as a Latin letter. Uppercase 'Ə' (U+018F) already fell below the cutoff, so the same letter counted in one case and not the other. Since NandhaKishorM#169 the unclaimed 'ə' lands in "other", which reads Azerbaijani as partly non-Latin: "Müştəridən iki dəfə pul alınıb" reports non_latin_fraction 0.1111 and a lone 'ə' reports script "other" at 1.0, for text that is Latin script throughout.

Add an Azerbaijani function-word list and the schwa to the non-English diacritic set. The diacritic signal does not reach Azerbaijani typed without diacritics, which is common in support text, so that text still went to the English checkpoint. All 41 words are unique to the list, so it names Azerbaijani only on evidence no other list claims, per the rule added in NandhaKishorM#178.

Fold the dotted capital 'İ' to 'i' before lowercasing, after the NandhaKishorM#179 identifier mask. Python lowercases it to 'i' plus a combining dot above, which is not a word character, so 'İLƏ' split into 'i' and 'lə' and matched no word list; uppercase Azerbaijani now scores the same as lowercase. The remaining Azerbaijani letters were already handled correctly in both cases.

Add regression cases for the lone schwa, uppercase input with the dotted I, script profiles, text typed without diacritics, multilingual routing and an explicit English override. The routing suite reports 7 failures before the fix and 338 passes after it. A sweep over 23 states changes exactly one route, the de-diacriticised Azerbaijani case; English with links, French, Spanish, Brazilian Portuguese, Italian, German, Romanian, Polish, Turkish, Hindi, Japanese, Armenian, Cherokee, fullwidth Latin and digit-only input are unchanged.

Tests ran with a namespace loader and a stub torch that bypass the ML-dependent package initializer: test_router 338, test_email 46 and test_packaging 15 pass. test_criteria, test_decision_model, test_download and test_shortlist need transformers, safetensors or a real package import and were not run, and model inference was not exercised. ruff was not available locally. Byte compilation and git diff --check pass.

Co-authored-by: Nandakishor <48623612+NandhaKishorM@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Unlisted scripts are routed to the English checkpoint: 68% of Unicode's letters match no _SCRIPT_RANGES entry

2 participants