Skip to content

Latest commit

 

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

LoRA Fine-Tuning of RoBERTa for IMDB Sentiment Analysis

GitLab pipeline status License: MIT

A project demonstrating efficient fine-tuning of a RoBERTa transformer model using Low-Rank Adaptation (LoRA) for sentiment classification. This approach achieves high performance on the IMDB movie review dataset while training only a fraction of the model's parameters.

Model Training Performance

Table of Contents

  1. Project Overview
  2. Key Features
  3. Methodology
  4. Performance
  5. How to Reproduce
  6. Exploratory Notebook

Project Overview

This project provides a complete pipeline for fine-tuning a roberta-base model for binary text classification. It showcases modern NLP techniques, including the use of the Hugging Face ecosystem and Parameter-Efficient Fine-Tuning (PEFT) with LoRA, making it possible to train large models on consumer-grade hardware.

Key Features

  • Efficient Fine-Tuning: Uses LoRA to drastically reduce the number of trainable parameters, leading to faster training and lower memory usage.
  • Reproducible Pipeline: The project is structured with scripts for model creation and training, along with a dependency list, ensuring the results can be easily replicated.
  • Clear & Structured Code: Logic is separated into a model definition (src/model.py) and a training script (src/train.py), following software engineering best practices.

Methodology

  1. Data Processing: The IMDB 50k movie review dataset is loaded, cleaned, and split into training (64%), validation (16%), and test (20%) sets.
  2. Model Architecture: A pre-trained roberta-base model is loaded and adapted for sequence classification. LoRA is applied to the query and value matrices of the attention layers.
  3. Training: The model is fine-tuned using the Hugging Face Trainer API, which handles the training loop, evaluation, and logging. Mixed-precision training (fp16) is used for further efficiency.
  4. Evaluation: Performance is measured on the held-out test set using accuracy as the primary metric.

Performance

The fine-tuned model achieves excellent performance on the test set. The final metrics after training are:

  • Test Accuracy: 93%

How to Reproduce

  1. Clone the repository:

    git clone https://gitlab.com/deep-learning-lc0/nlp/imdb-movie-classifier.git
    cd imdb-movie-classifier
  2. Set up the environment:

    python -m venv venv
    source venv/bin/activate
    pip install -r requirements.txt
  3. Download the data: get the data from HERE and place it in the data folder

  4. Run the training pipeline:

    python src/train.py

    The trained model and results will be saved in the roberta_lora_results/ directory, and predictions will be in roberta_lora_preds.csv.

Exploratory Notebook

For a detailed, step-by-step walkthrough of the data exploration and initial model building process, please see the Jupyter Notebook:

About

Sentiment analysis project to correctly classify movie reviews, you can explore the project on gitlab too: https://gitlab.com/deep-learning-lc0/nlp/imdb-movie-classifier

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages