Successfully developed a fine-tuned BERT transformer model which can accurately classify symptoms to their corresponding diseases upto an accuracy of 89%.
-
Updated
May 6, 2024 - Jupyter Notebook
Successfully developed a fine-tuned BERT transformer model which can accurately classify symptoms to their corresponding diseases upto an accuracy of 89%.
Python module designed to replace common internet slang and abbreviations with their full forms, enhancing the readability of informal text. It efficiently cleans text data from chats, social media, and online communication. The module also supports tokenization and integrates seamlessly with pandas for batch processing of text in DataFrames.
Natural language processing examples and automations
Tugas akhir (final year project). This source code was used in the bachelor thesis “Public Sentiment Analysis of Mental Disorder based on Twitter Texts using Support Vector Machine”.
News, full-text, and article metadata extraction in Python 3. Advanced docs:
Successfully developed a fine-tuned RoBERTa transformer model which can almost perfectly classify whether any given SMS is spam or not.
Unsupervised Machine Learning project for Netflix Movies and TV Shows Clustering. The main goal of this project is to create a content-based recommender system that recommends top 10 shows to users based on their viewing history.
Jupyter notebooks on Natural Language Processing.
A Scrapy package based web scraper for collecting Kurdish text data from websites. The tool recursively crawls specified domains, extracts article content using Trafilatura, and filters results by language using Facebook's FastText language identification model.
FastAPI service for Persian FAQ normalization, deduplication, chunking and retrieval evaluation. Synthetic examples, Docker and tests.
Text Preprocessing in Python
Twitter Sentiment Analysis of NBA Players
Implementation of text preprocessing impact analysis on named entity recognition (NER) based on conditional random field (CRF) in Indonesian text.
Explore the vast field of Natural Language Processing (NLP) with our comprehensive toolkit. From text preprocessing to advanced sentiment analysis and language modeling, this repository provides a range of tools and algorithms to empower your NLP projects. Dive into state-of-the-art techniques and resources curated to enhance your understanding.
This project extracts and analyzes textual data from given URLs using BeautifulSoup and NLTK. It performs sentiment analysis, word complexity assessment, and calculates average word length, saving results in text and CSV formats.
Leksara is a Python toolkit for cleaning and preprocessing Indonesian text data, focused on the E-commerce domain. It automates text cleaning tasks like punctuation removal, stopword filtering, and slang normalization, helping Data Scientists and ML Engineers save time and streamline workflows.
Train a custom Word2Vec model on Armenian news articles using the ilur-news-corpus. Includes preprocessing, Armenian-specific normalization, rare word filtering, training with Gensim, and saving for reuse
This repository contains code for a text classification project using Twitter and news datasets, where several classification models were evaluated and compared based on their performance metrics.
Documents classification using KNN Algorithm a graph based approach along with scrapped data
A deep learning model I built to detect toxic comments and promote safer online interactions.
To associate your repository with the text-preprocessing topic, visit your repo's landing page and select "manage topics."