This repository contains the data science pipeline and analysis for the Open University Learning Analytics Dataset (OULAD). It is structured to support a reproducible, bilingual (R and Python) workflow, drawing inspiration from established project structures like Cookiecutter Data Science.
Follow these steps to get the repo up and running:
# 1. Clone the repository
git clone https://github.com/Hariswara/TPSM-POS-LEARN.git
cd oulad-analysis
# 2. Install Git LFS (required once per machine)
git lfs install
# 3. Extract the raw dataset from the zip
unzip data/raw/OULAD.zip -d data/raw/
# 4. Install Python dependencies
pip install -r requirements.txt
# 5. (Optional) Restore R environment
# Open oulad-analysis.Rproj in RStudio, then run:
# renv::restore()Why unzip? To save bandwidth and storage, only the compressed dataset (
OULAD.zip, ~45 MB) is committed to Git via LFS. The extracted CSVs (~500 MB) are git-ignored and live only on your local machine.
The project is organized into the following directories to separate concerns, ensure reproducibility, and keep the workflow organized for collaborators.
oulad-analysis/
│
├── data/
│ ├── raw/ ← Original, immutable OULAD CSVs. Never modify these.
│ ├── processed/ ← Cleaned, merged, and transformed data files (generated).
│ └── features/ ← Engineered feature tables ready for modeling (generated).
│
├── notebooks/ ← Jupyter or R Markdown notebooks, prefixed by execution order.
│ ├── 00_data_pipeline/ ← Data loading, cleaning, and merging scripts (R or Python).
│ ├── 01_descriptive/ ← Exploratory Data Analysis (EDA), distributions, summaries (R).
│ ├── 02_inferential/ ← Hypothesis tests, ANOVA, chi-square, etc. (R).
│ ├── 03_predictive/ ← Machine learning models, regression, time series (Python).
│ └── 04_scratch/ ← Personal sandboxes for ad-hoc exploration (not reviewed).
│
├── src/ ← Reusable source code and helper modules.
│ ├── r/ ← Shared R helper functions used across notebooks/scripts.
│ └── python/ ← Shared Python modules/classes.
│
├── outputs/ ← Final generated artifacts.
│ ├── figures/ ← Saved plots and visualizations (generated).
│ ├── tables/ ← Exported result tables (generated).
│ └── models/ ← Serialized, saved model objects (generated).
│
├── report/ ← Source files for final written reports or presentations.
├── docs/ ← Team documentation, design decisions, and meeting notes.
│
├── .gitattributes ← Git LFS tracking rules (auto-generated, do not edit manually).
├── .gitignore ← Specifies intentionally untracked files to ignore.
├── oulad-analysis.Rproj ← RStudio Project file to set working directory natively.
├── README.md ← The top-level README for developers/collaborators.
├── requirements.txt ← Python dependencies file.
└── renv.lock ← R dependencies lockfile (renv).
raw/: ContainsOULAD.zip(tracked via Git LFS) and the extracted CSVs (git-ignored). After cloning, rununzip data/raw/OULAD.zip -d data/raw/to extract. Never modify these source files.processed/: Intermediate datasets that have been cleaned and prepared by the pipeline scripts.features/: The final formulated tables that are fed directly into statistical analysis and machine learning models.
Notebooks are numbered sequentially (00_ to 03_) to clearly indicate the order of operations. Anyone picking up the project should run them in this order.
04_scratch/: Use this for messy, experimental work. Notebooks here are not expected to run cleanly or be reviewed.
Any code that is used in multiple notebooks or scripts should be extracted into functions/classes and stored here based on the language. This keeps notebooks clean and analytical.
This folder is for generated assets. Everything in here should be deterministically reproducible from the data/ and the code in notebooks/ or src/. Large files in this folder are tracked via Git LFS (see below).
This project uses two separate package management systems due to the bilingual nature of the analysis:
- Python: Use
pip install -r requirements.txtto install the required Python packages. Maintain this file when adding new Python libraries. - R: Uses
renvfor reproducible environments. Therenv.locktracks the exact package versions used.
This repository is fully compatible with RStudio as an "R Project".
- Working Directory & Paths: Double-click
oulad-analysis.Rprojto open the project in RStudio. This automatically sets R's working directory to the project root, meaning you can load data securely using relative paths (e.g.,read.csv("data/raw/data.csv")). - Environment Management: Opening the
.Rprojfile will triggerrenvto restore the precise package versions specified inrenv.lock. - Seamless Development: You can create and knit
.Rmdfiles directly innotebooks/and seamlesslysource("src/r/...")helper functions.
This project uses Git LFS to handle large files. Instead of storing full data files directly in Git history, LFS replaces them with lightweight pointer files while the actual content is stored on the LFS server.
First-time setup (required once per machine):
git lfs installTracked file types:
| Category | Extensions |
|---|---|
| Data | .csv, .xlsx, .xls, .parquet, .feather |
| Models | .pkl, .rds, .h5, .hdf5 |
| Images | .png, .jpg, .jpeg, .svg |
| Documents | .pdf |
All files inside data/ and outputs/ are also tracked by LFS regardless of extension.
Note: After cloning, run
git lfs pullif large files appear as pointer files instead of actual content.