You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
This repository contains the implementation of a wide variety of Reinforcement Learning Projects in different applications of Bandit Algorithms, MDPs, Distributed RL and Deep RL. These projects include university projects and projects implemented due to interest in Reinforcement Learning.
AI project combining Monte Carlo racetrack control with on/off-policy learning, weighted importance sampling, and Sinkhorn optimal-transport face morphing with Wasserstein interpolation.
Q-learning is an off-policy temporal-difference control algorithm. It learns the value of the optimal action, independent of the action actually taken by the agent.
Off-policy Monte Carlo control (weighted importance sampling) solving Sutton & Barto's Racetrack problem — with a training CLI, wandb logging, weight saving, and side-by-side agent replays.
This repository contains all of the Reinforcement Learning-related projects I've worked on. The projects are part of the graduate course at the University of Tehran.
🧗Comparative Reinforcement Learning analysis implementing SARSA (On-Policy) vs Q-Learning (Off-Policy) from scratch on OpenAI Gymnasium CliffWalking-v1 with live side-by-side animations and policy heatmaps.