This project presents an Object-Oriented Programming (OOP) analysis of the Scikit-learn library. The main purpose of this project is to understand how Scikit-learn uses important OOP concepts such as inheritance, polymorphism, abstraction, and composition in its internal architecture. This repository also includes a custom machine learning extension built using standard Scikit-learn interfaces.
- Muhammad Bilal — F25BDATS1M02098
- Maryam Tahir — F25BDATS1M02051
- Jaweria Zafar — F25BDATS1M02090
Submitted to: Dr. Akmal Shahbaz
Department: Department of Data Science, The Islamia University of Bahawalpur
Subject: Object-Oriented Programming (BS Data Science — 2nd Semester — M2 Section)
Session: 2026
This project bridges the gap between software engineering theory and applied machine learning pipeline creation. It comprises two main artifacts:
- The Formal Report: A granular, code-level analysis tracking line-by-line engineering choices within Scikit-learn's underlying core system.
- The Custom Extension: A working, production-grade Python script implementing a custom transformer, classifier, and pipeline ecosystem that inherits from and complies with standard Scikit-learn Mixins.
The custom framework developed within this project explicitly showcases the four pillars of Object-Oriented Programming:
Both SmartClassifier and SmartTransformer inherit concurrently from Scikit-Learn's structural layout.
SmartClassifierusesBaseEstimatorandClassifierMixinto automatically gain hyperparameter utilities and uniform validation properties.SmartTransformerleveragesTransformerMixinto acquire automaticfit_transform()behavior through boilerplate code reuse.
Standard behavioral workflows (fit, predict, transform, score) are overridden to execute custom math patterns while maintaining plug-and-play compatibility with native estimators.
Granular mathematical manipulations (such as data matrix centering, state constraints tracking via check_array, and multi-class bound logic parsing) are encapsulated away from the terminal client layer behind generic interface loops.
Rather than using heavy, deep structural inheritance trees, SmartPipeline implements a strict Composition Pattern. It acts as an independent execution orchestrator that holds and controls standalone component objects (SmartTransformer and SmartClassifier), respecting the software axiom: "Favor object composition over class inheritance".
Below is the accurate deployment structure of our project directories:
SCIKIT-LEARN-OOP-ANALYSIS/
│
├── analyzed_files/ # Contains raw/processed data files used during system validation
├── Code/ # Core Python engine containing the custom_extension.py script
├── diagrams/ # Visual UML, inheritance trees, and pipeline architecture diagrams
├── Report/ # Formal academic technical analysis report document
│
├── .gitignore # Python environment and metadata runtime filter file
└── README.md # Project Blueprint & System Deployment Manual (This File)