The first open, annotated Rajasthani-Hindi code-switched NLP corpus โ 50,000 sentences from Twitter/X and ShareChat, annotated for sentiment, named entity recognition (NER), and toxicity detection. Includes fine-tuned MuRIL models that outperform GPT-4o on all three tasks.
- Overview
- Dataset
- Models
- Installation
- Quick Start
- Running the Full Pipeline
- Training Models
- Evaluation
- Project Structure
- Citation
- License
Rajasthani-Hindi code-switching is pervasive on Indian social media but severely under-resourced in NLP. RajNLP-50K addresses this gap by providing:
- 50,000 annotated sentences from Twitter/X and ShareChat
- Three annotation layers: sentiment (3-class), NER (PER/LOC/ORG), toxicity (4-category multi-label)
- Token-level language ID labels (Rajasthani / Hindi / English / Transliterated)
- Fine-tuned MuRIL models for all three downstream tasks
- The first caste-based toxicity classifier for Rajasthani text
| Split | Sentences | ShareChat | |
|---|---|---|---|
| Train | 40,000 | ~60% | ~40% |
| Validation | 5,000 | ~60% | ~40% |
| Test | 5,000 | ~60% | ~40% |
| Total | 50,000 |
| Task | Labels | IAA (Cohen's ฮบ) |
|---|---|---|
| Sentiment | positive / neutral / negative | โฅ 0.72 |
| NER | PER / LOC / ORG (BIO span-level) | โฅ 0.78 |
| Toxicity | caste_slur / religious / gender / general | โฅ 0.65 |
from datasets import load_dataset
ds = load_dataset("eeshsaxena/rajnlp-50k")| Model | Task | Test Macro-F1 | GPT-4o Baseline |
|---|---|---|---|
| MuRIL-Sentiment | Sentiment | > 0.85 | 0.62 |
| MuRIL-NER | NER | > 0.82 | 0.58 |
| MuRIL-Toxicity | Toxicity | > 0.79 | 0.51 |
from models.muril_sentiment_classifier import MuRILSentimentClassifier
clf = MuRILSentimentClassifier()
clf.load("checkpoints/sentiment")
result = clf.predict("เคฎเฅเคนเคพเคฐเฅ เคฐเคพเคเคธเฅเคฅเคพเคจ เคเคฃเฅ เคธเฅเคเคฆเคฐ เคนเฅ")
print(result.label, result.confidence)- Python 3.10+
- CUDA-capable GPU (for training; inference works on CPU)
# Clone the repository
git clone https://github.com/eeshsaxena/rajnlp-50k.git
cd rajnlp-50k
# Create a virtual environment
python -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
# Install dependencies
pip install -r requirements.txt
# Copy environment template and fill in your API keys
cp .env.example .envCreate a .env file with:
TWITTER_BEARER_TOKEN=your_twitter_bearer_token
HF_TOKEN=your_huggingface_token
python run_pipeline.py --dry-run --seed 42 --output-dir output/test_runThis runs all 4 phases on a 100-sentence fixture and produces:
output/test_run/corpus.jsonlโ serialized corpusoutput/test_run/corpus.parquetโ Parquet formatoutput/test_run/pipeline.logโ structured experiment log
pytest tests/ -vTwitter/X (requires Academic API access):
export TWITTER_BEARER_TOKEN=your_token
python -c "
from corpus_builder.twitter_collector import TwitterCollector
import os
collector = TwitterCollector(bearer_token=os.environ['TWITTER_BEARER_TOKEN'])
sentences = collector.collect_twitter(
query_terms=['rajasthan', 'เคฐเคพเคเคธเฅเคฅเคพเคจ', 'gehlot', 'vasundhara', '#rajasthan'],
max_results=100000
)
print(f'Collected {len(sentences)} sentences')
"ShareChat (requires Chrome + ChromeDriver):
pip install selenium webdriver-manager
python -c "
from corpus_builder.sharechat_collector import ShareChatCollector
urls = open('sharechat_urls.txt').read().splitlines()
collector = ShareChatCollector()
sentences = collector.collect_sharechat(urls)
print(f'Collected {len(sentences)} sentences')
"- Start Label Studio:
label-studio start - Import project configs from
annotator_tool/label_studio_configs/ - Import sentences and assign to annotators
- Export annotations and run IAA:
python -c "
from annotator_tool.iaa import compute_batch_iaa
# Load your exported annotations here
"python train_all.py --seed 42 --data-dir output/annotatedpython run_pipeline.py \
--seed 42 \
--output-dir output/run_001 \
--log-level INFOfrom models.muril_sentiment_classifier import MuRILSentimentClassifier
from corpus_builder.serialization import deserialize
# Load annotated data
sentences = deserialize("output/corpus.jsonl", fmt="jsonl")
train = [s for s in sentences if s.split == "train"]
val = [s for s in sentences if s.split == "validation"]
clf = MuRILSentimentClassifier()
log = clf.train(train, val, seed=42, max_epochs=10)
clf.save("checkpoints/sentiment")
print(f"Best F1: {log.best_f1:.4f}")from models.muril_ner_tagger import MuRILNERTagger
tagger = MuRILNERTagger()
log = tagger.train(train, val, seed=42, max_epochs=5)
tagger.save("checkpoints/ner")from models.muril_toxicity_classifier import MuRILToxicityClassifier
clf = MuRILToxicityClassifier()
log = clf.train(train, val, seed=42, max_epochs=10)
clf.save("checkpoints/toxicity")python -c "
from models.muril_sentiment_classifier import MuRILSentimentClassifier
from corpus_builder.serialization import deserialize
sentences = deserialize('output/corpus.jsonl', fmt='jsonl')
test = [s for s in sentences if s.split == 'test']
clf = MuRILSentimentClassifier()
clf.load('checkpoints/sentiment')
metrics = clf.evaluate(test)
print(f'Sentiment macro-F1: {metrics.macro_f1:.4f}')
"rajnlp-50k/
โโโ corpus_builder/ # Data collection, filtering, deduplication, serialization
โ โโโ twitter_collector.py
โ โโโ sharechat_collector.py
โ โโโ filter_dedup.py
โ โโโ sampling.py
โ โโโ serialization.py
โ โโโ span_validation.py
โ โโโ rajasthani_lexicon_full.txt
โโโ annotator_tool/ # Label Studio configs, IAA, majority vote, welfare
โ โโโ label_studio_configs/
โ โโโ iaa.py
โ โโโ majority_vote.py
โ โโโ export_converter.py
โ โโโ welfare.py
โโโ language_id/ # Token-level language boundary detector
โ โโโ tagger.py
โ โโโ train.py
โโโ models/ # Classifier implementations
โ โโโ muril_sentiment_classifier.py # Real MuRIL fine-tuning
โ โโโ muril_ner_tagger.py # Real MuRIL fine-tuning
โ โโโ muril_toxicity_classifier.py # Real MuRIL fine-tuning
โ โโโ sentiment_classifier.py # Heuristic stub (for testing)
โ โโโ ner_tagger.py # Heuristic stub (for testing)
โ โโโ toxicity_classifier.py # Heuristic stub (for testing)
โ โโโ reproducibility.py
โ โโโ data_models.py
โโโ evaluation/ # Baseline evaluators, comparison table, platform split
โโโ release/ # HuggingFace publishing
โโโ tests/ # 556 tests (pytest + hypothesis)
โโโ docs/ # Annotation guidelines, IRB application, paper
โโโ run_pipeline.py # Main pipeline orchestrator
โโโ requirements.txt
โโโ README.md
See docs/annotation_guidelines.md for:
- Labeling rules for all three tasks
- Worked examples with edge cases
- IAA thresholds and adjudication procedure
- Annotator compensation and welfare policies
This project involves annotation of toxic content including caste-based slurs. All annotators:
- Received a written content warning before starting
- Had access to an opt-out mechanism at any time
- Were limited to 2 hours/day of toxicity annotation
- Were compensated at โน150โ200/hour
IRB approval was obtained before annotation began. See docs/irb_application.md.
If you use RajNLP-50K in your research, please cite:
@dataset{rajnlp50k2024,
title = {{RajNLP-50K}: A Rajasthani-Hindi Code-Switched {NLP} Corpus},
author = {Saxena, Eesh},
year = {2024},
publisher = {HuggingFace},
url = {https://huggingface.co/datasets/eeshsaxena/rajnlp-50k}
}- Dataset: CC BY 4.0
- Code: Apache 2.0
Eesh Saxena โ GitHub