Skip to content
#

rloo

Here are 7 public repositories matching this topic...

Language: All
Filter by language

This training offers an intensive exploration into the frontier of reinforcement learning techniques with large language models (LLMs). We will explore advanced topics such as Reinforcement Learning with Human Feedback (RLHF), Reinforcement Learning from AI Feedback (RLAIF), Reasoning LLMs, and demonstrate practical applications such as fine-tuning

  • Updated Mar 9, 2026
  • Jupyter Notebook

Readable 25.7M-parameter GLM-5.3-Flash-style Hybrid-MoE built from scratch in PyTorch. Features hybrid linear/sparse attention, executable-reward RLOO, and recursive self-improvement on verified programs.

  • Updated Oct 10, 2026
  • Python

Add this topic to your repo

To associate your repository with the rloo topic, visit your repo's landing page and select "manage topics."

Learn more