LeetCode for PyTorch — 65 ML/AI interview problems from real interviews at Google, Meta, Anthropic. Jupyter notebooks, an auto-grader, and an MCP AI tutor.
-
Updated
Sep 15, 2026 - Jupyter Notebook
LeetCode for PyTorch — 65 ML/AI interview problems from real interviews at Google, Meta, Anthropic. Jupyter notebooks, an auto-grader, and an MCP AI tutor.
20 RL basics notebooks + 10 advanced projects with Streamlit apps covering RLHF, Offline RL, Multi-Agent RL, Safe RL, World Models, RLAIF, Hierarchical RL, Meta-RL, Drug Discovery RL, and Sim-to-Real transfer
The LLM FineTuning and Evaluation project 🚀 enhances FLAN-T5 models for tasks like summarizing Spanish news articles 🇪🇸📰. It features detailed notebooks 📚 on fine-tuning and evaluating models to optimize performance for specific applications. 🔍✨
🏟️ Modern RL algorithms from scratch — from Q-Learning to GRPO — with clean PyTorch code and interactive notebooks. Compare PPO vs DPO vs GRPO for LLM alignment.
Comprehensive university-level study guide for LLMs, transformers, RLHF, and generative AI. 335+ pages (actively expanding), 47 visualizations, 4 notebooks. Regular updates with enhanced sections, new implementations, and expanded coverage.
An opinionated, end‑to‑end tutorial project for learning Reinforcement Learning (RL) from first principles to deployment. No notebooks. Everything is an explicit, inspectable Python script you can diff, profile, containerize, and ship.
23 hands-on Colab notebooks for LLM fine-tuning (LoRA, QLoRA, PEFT, RLHF), RAG pipelines, knowledge graphs, 1-bit quantization & MLflow evaluation — Llama 2, Mistral 7B, Falcon, Gemma 2, Phi-1.5, GPT-3.5 & more.
This repository contains Jupyter Notebooks, scripts, and datasets used in our finetuning experiments. The project focuses on Direct Preference Optimization (DPO), a method that simplifies the traditional finetuning process by using the model itself as a feedback mechanism.
Five-day value-added course on Reinforcement Learning and Language Model Alignment at Vel Tech, with interactive lecture pages, browser labs, Colab notebooks and Word lecture notes. From bandits and Q-learning to PPO, RLHF, DPO and honest evaluation, all implemented in NumPy.
An end-to-end pipeline for adapting FLAN-T5 for dialogue summarization, exploring the full spectrum of modern LLM tuning. Implements and compares Full Fine-Tuning, PEFT (LoRA), and Reinforcement Learning (RLHF) for performance and alignment. Features a PPO-tuned model to reduce toxicity, in-depth analysis notebooks, and interactive Streamlit demo.
To associate your repository with the rlhf topic, visit your repo's landing page and select "manage topics."