Skip to content

Latest commit

 

History

12 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

banglachhanda

A rule-based scanner for Bangla verse. It takes a line, works out its syllables and their structure, and derives matra counts and foot divisions under each of the three meters — অক্ষরবৃত্ত, মাত্রাবৃত্ত, স্বরবৃত্ত.

This is the baseline for the BanglaChhanda project, not the product. Its job is to be the honest floor that a learned scansion model has to beat, and to pre-fill annotation records so annotators correct a draft instead of starting from nothing.

>>> from chhanda import scan
>>> print(scan("শব্দ").summary())
শব্দ
  শব্[C] দ[O]
  অক্ষরবৃত্ত=2  মাত্রাবৃত্ত=3  স্বরবৃত্ত=2
  best: no pattern fits
  flags: final_schwa_kept_after_conjunct, conjunct_split

One written word, one structural analysis, three different matra readings — all produced by rule. That is the whole design in one example.

The design rule

Structure is annotated; matra is computed. A closed syllable is worth 2 matra in matrabritta, 1 in svarabritta, and 1 or 2 in akkharbritta depending on whether it ends a word. If annotators wrote matra counts by hand, the corpus would encode the textbook rules and could never be used to test them. So the annotation records syllabification, open/closed and foot boundaries; this library derives the rest.

That is why to_annotation() emits no matra field. It is not an omission.

Uncertainty is flagged, never resolved quietly

Schwa deletion and conjunct splitting are not fully regular. Where a rule could plausibly go either way, the scanner sets a flag and carries on rather than guessing:

Flag Meaning
schwa_medial_uncertain inherent vowel before another inherent vowel — variable, partly dialectal
final_schwa_kept_after_conjunct শব্দ stays shob-do rather than becoming shobd
final_schwa_kept_monosyllable nothing precedes it to attach a stranded consonant to
conjunct_split a written cluster divided across a syllable boundary
phala_kept_in_onset r-phala held in the onset: আক্রমণ is a-kro-mon
y_phala_gemination_check y-phala may geminate; guideline hard case 3 says a human decides
stranded_onset consonants with no syllable to attach to — usually malformed input

scan(line).needs_review is true when anything is flagged, no foot pattern fits, or the best fit is weakly aligned. Route those lines to a human first; they are where the rules are earning or failing.

Meter proposals rank, they do not classify

Under svarabritta every syllable is 1 matra, so any line whose syllable count divides by four fits exactly on the arithmetic alone — a deliberately unmetrical line drew exact fits under two meters. Each fit therefore carries alignment: the share of its foot boundaries landing at word boundaries, which is where Bangla feet tend to break. Candidates rank on it, and a fit below 0.5 is marked weak.

Mid-word boundaries stay legal (guideline hard case 9); they just score lower. This reranks rather than rejects — the invented line still fits, at 0.67. Rejecting non-metrical verse properly needs poem-level consistency, which is open work.

Rhyme

chhanda.rhyme works over the same syllable layer. Bangla rhyme is phonological, so শ/ষ/স fold to one sibilant, ণ/ন to one nasal, and vowel length is not contrastive — matching written characters would both miss real rhymes and invent false ones.

>>> from chhanda.rhyme import rhymes, rhyme_scheme
>>> rhymes("করে", "পড়ে")          # different onsets, rhyme on -e
True
>>> rhyme_scheme(["বুকে", "ভুলি", "মুখে", "তুলি"]).pattern
'abab'

An unrhymed line is labelled -, not a fresh letter: a rhyme class of one would make free verse look like an elaborate scheme.

Known limitations

These are real and should be measured by the pilot rather than patched by guesswork.

  • Final-schwa exceptions are not modelled. The rule deletes a word-final inherent vowel unless it follows a conjunct. That is right most of the time and wrong for a class of common words — ছোট is said cho-to, but the scanner gives one closed syllable. Fixing this properly needs an exception lexicon, and that lexicon should come out of annotation, not out of intuition. How often this rule is wrong is a headline number for the resource paper.
  • The foot patterns are a starting set, not a complete inventory of Bangla line forms. PATTERNS in meters.py holds payar, four matrabritta feet and one svarabritta foot. Tripadi and the longer forms are absent.
  • No authority is pinned yet. The matra table follows the common textbook description; §2 and §5 of the annotation guideline list what a prosody specialist must confirm before the numbers are trusted.
  • Orthography only. No pronunciation lexicon, no dialect handling, no performance timing.

A note on that first limitation, since it will come up: when the scanner disagrees with a reader about a well-known line, the scanner is the more likely one to be wrong. It fits আমাদের ছোট নদী চলে বাঁকে বাঁকে as payar, exactly; whether that is the accepted scansion of the line is precisely the sort of question the specialist review exists to settle.

Layout

chhanda/script.py     text -> orthographic units (aksharas); mechanical, no judgement
chhanda/syllable.py   units -> spoken syllables; schwa deletion and conjunct splitting
chhanda/meters.py     matra values, foot fitting, word-boundary alignment
chhanda/rhyme.py      phonological rhyme keys and scheme labelling
chhanda/scan.py       scan() and to_annotation()
corpus/               Wikisource fetcher, pilot builder, Hub sync
tests/                49 tests

No test asserts the "correct" meter of a canonical poem. Until the guideline is signed off, doing so would be inventing ground truth.

Running it

python3 -m venv .venv && .venv/bin/pip install pytest
.venv/bin/python -m pytest -q

Requires Python 3.10+ (uses X | None annotations). No runtime dependencies.

Corpus

corpus/fetch_wikisource.py pulls public-domain Tagore from Bengali Wikisource with per-poem provenance (page URL, revision id, retrieval date). corpus/build_pilot.py samples it poem-wise, stratified by closed-syllable density, and writes one pre-filled annotation record per line.

.venv/bin/python corpus/fetch_wikisource.py --collection "কণিকা (রবীন্দ্রনাথ ঠাকুর)" --limit 40
.venv/bin/python corpus/build_pilot.py --target 300

The corpus is 153 poems / 7,009 lines from two poets, and the current pilot is 316 lines from 19 poems — 191 Jibanananda, 125 Tagore. The scanner flags 80% of them for review, which is the honest state of the rules rather than a defect to tune away. Review rate barely differs between the two poets (76% vs 74% before weak fits were counted), so the rules are not simply falling over on looser verse.

A poet is added only once the copyright term has expired. attribution() resolves author and rights per collection from the POETS table and raises for an unlisted poet or one still in copyright — a wrong rights line would propagate into every record and then get redistributed. Kazi Nazrul Islam (d. 1976) is out until 2037 under life + 60; Wikisource carries his pre-1931 books as US public domain, which is not enough here. A test says so, so adding him means deleting it.

Citing this work

This dataset and code are released under CC BY-SA 4.0, which requires attribution. If you use the scanner, the guideline or any part of the corpus in research or in a derived resource, cite it:

@misc{morol2026banglachhanda,
  author    = {Morol, Md Kishor},
  orcid     = {0000-0002-4468-8260},
  title     = {{BanglaChhanda}: a rule-based scanner and pilot corpus for
               {Bangla} verse},
  year      = {2026},
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.22798431},
  url       = {https://doi.org/10.5281/zenodo.22798431}
}

That is the concept DOI: it always resolves to the newest version, which is what you want when citing the work in general. To pin the exact release you used, cite the version DOI instead — v0.1.0 is 10.5281/zenodo.22798432.

Author ORCID: 0000-0002-4468-8260.

Please also credit Bengali Wikisource for the transcriptions, as the CC BY-SA terms of the source text require. The poems themselves are public domain.

If you correct or extend the annotations, CC BY-SA also requires you to release the result under the same licence.

Related

The annotation guideline this implements is in docs/; the flag names above correspond to its numbered hard cases.

The corpus and pilot are published as a dataset: huggingface.co/datasets/kishormorol/banglachhanda. The dataset card leads with the warning that none of its labels are human-checked.

The Hub copy is a separate store, so it drifts the moment the corpus is rebuilt here. hf/README.md is the card under version control, and:

.venv/bin/python corpus/sync_hf.py --check   # compare; no token needed
.venv/bin/python corpus/sync_hf.py           # upload whatever differs

CI runs the --check half on every push that touches the corpus, the card or the guideline, and on every release. It holds no credential and cannot fix drift — it fails and tells a human to run the sync.

About

Rule-based Bangla scansion (chhanda) — annotation guideline, scanner baseline, and pilot corpus

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages