Skip to content
View shehio's full-sized avatar

Highlights

  • Pro

Block or report shehio

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
shehio/README.md

Shehab Yasser

I work on reinforcement learning and the infrastructure it runs on. Lately that has mostly meant teaching agents to play games: a DQN agent that learns Crash Bandicoot from raw pixels on an emulated PlayStation 1, and a chess network distilled from Stockfish.

Before that I did backend engineering on distributed systems and built evaluation harnesses for LLM agents. My degree is in financial mathematics.

What I'm Working On

  • RSI (recursive self-improvement) and AI safety
  • RL agents that learn gameplay from pixels
  • Training infrastructure

Selected Work

  • world-models (site): Stockfish distilled into AlphaZero's 20x256 ResNet on ~46M positions. It plays chess at 2,301 Elo (95% CI [2,190, 2,601]). Includes single-variable ablations and the negative results.
  • ps-env: a headless PlayStation 1 wrapped as an RL environment, with a Nature-DQN agent that learns Crash Bandicoot from raw pixels
  • cassandra-playground: Cassandra-style anti-entropy repair in Go, using a gossip protocol and Merkle trees
  • rl / tabular-rl: RL agents with nothing abstracted away, from tabular methods up through PPO
  • rl-playbook (rlplaybook.com): a visual timeline of deep RL papers, starting from DQN
  • Project-Nash: Nash equilibria, Lemke–Howson, minimax, and simplex

Publications

Background

Distributed systems and infrastructure: gossip protocols, Temporal.io workflows, Terraform on AWS. Quantitative finance: portfolio optimization, stochastic programming, derivative pricing. Python, Go, Rust, TypeScript, Java.

Pinned Loading

  1. Everything-Financial-Engineering Everything-Financial-Engineering Public

    Links for the most relevant topics

    33 2

  2. monte-carlo-tree-search monte-carlo-tree-search Public

    Monte Carlo Tree Search with a clean interface for perfect-information games

    Python 16 2

  3. Project-Nash Project-Nash Public

    A panoply of algorithms in game theory, econometrics, linear programming, and Montecarlo simulations.

    TypeScript 13

  4. tabular-rl tabular-rl Public

    Reinforcement Learning algorithms with nothing abstracted away

    Python 7 1

  5. Computational-Biology Computational-Biology Public

    Bioinformatics algorithms from UW CSEP 527 — HMMs, Viterbi, Gillespie, and Smith-Waterman

    Jupyter Notebook 6

  6. rl rl Public

    Implementing RL agents, one algorithm at a time

    Python 9 2