A hands-on course on modern LLM architectures, training, and inference, with step-by-step PyTorch implementations and runnable notebooks.
-
Updated
Oct 1, 2026 - Jupyter Notebook
A hands-on course on modern LLM architectures, training, and inference, with step-by-step PyTorch implementations and runnable notebooks.
Repository for the companion Colab notebook of the Domain-Specific Small Language Models book.
Implementation for the different ML tasks on Kaggle platform with GPUs.
A collection of hand on notebook for LLMs practitioner
Here are Jupyter Notebooks that demonstrate pruning and quantization of the Fashion-MNIST Classification model in Pytorch. Also, using Torchscript, one can easily deploy these models on a C++ platform.
This repository contains example Jupyter notebooks demonstrating how to use the quantized versions of the PLLuM-8x7B-chat model in GGUF format
Code snippets and notebooks used in PDS classes.
A Tutorial Notebook to Quantization in Machine Learning
This repository contains example notebooks and homeworks demonstrating various techniques in model optimization for Edge ML.
Implementations and Quantization Notebooks of models for Edge AI!
Backprop-free learning study: spiking (LIF) neurons + Forward-Forward + JEPA + int4 QAT, with a full ablation notebook.
Notebooks for compiling a GTSRB MobileNet classifier for the Hailo edge AI accelerator
[INDEX] Course on LLMs with roadmaps and Colab notebooks.
This project fine-tunes Google's Gemma 2B model for Python code generation using LoRA and 4-bit quantization, implemented in a Google Colab notebook for efficient training.
A hands-on PyTorch notebook portfolio where each notebook runs end to end and is CI-validated via NNx — spanning classical ML, GNNs, NLP, transformer LMs with BPE, DDPM diffusion, DPO, JEPA, MoE, PEFT, quantization, pruning, and knowledge distillation.
End-to-end LLM engineering notebook series covering SFT, PEFT, preference optimization, RL tuning, quantization, inference, and deployment workflows.
GPT-OSS-120B fine-tuning notebook and setup guide using LoRA, quantization and DeepSpeed on a two-GPU configuration.
A set of notebooks analyzing neural network quantization, comparing symmetric and asymmetric schemes, calibration methods, and PTQ vs. QAT
Google Colab notebooks for deploying, quantizing (4-bit BitsAndBytes), serving (Flask/ngrok & Ollama), and inferencing LLMs (DeepSeek-R1, Mistral-7B, Qwen2.5-Coder).
Course-grade implementation and curated material for Fundamentals of Deep Learning and TinyML (MME 26849). Includes hands-on notebooks, slides, and practical experiments covering classical models (SVMs, perceptrons), modern deep nets (CNNs), and efficiency techniques (pruning, quantization) with a focus on size/latency-aware workflows
To associate your repository with the quantization topic, visit your repo's landing page and select "manage topics."