I work on reinforcement learning and the infrastructure it runs on. Lately that has mostly meant teaching agents to play games: a DQN agent that learns Crash Bandicoot from raw pixels on an emulated PlayStation 1, and a chess network distilled from Stockfish.
Before that I did backend engineering on distributed systems and built evaluation harnesses for LLM agents. My degree is in financial mathematics.
- RSI (recursive self-improvement) and AI safety
- RL agents that learn gameplay from pixels
- Training infrastructure
- world-models (site): Stockfish distilled into AlphaZero's 20x256 ResNet on ~46M positions. It plays chess at 2,301 Elo (95% CI [2,190, 2,601]). Includes single-variable ablations and the negative results.
- ps-env: a headless PlayStation 1 wrapped as an RL environment, with a Nature-DQN agent that learns Crash Bandicoot from raw pixels
- cassandra-playground: Cassandra-style anti-entropy repair in Go, using a gossip protocol and Merkle trees
- rl / tabular-rl: RL agents with nothing abstracted away, from tabular methods up through PPO
- rl-playbook (rlplaybook.com): a visual timeline of deep RL papers, starting from DQN
- Project-Nash: Nash equilibria, Lemke–Howson, minimax, and simplex
Distributed systems and infrastructure: gossip protocols, Temporal.io workflows, Terraform on AWS. Quantitative finance: portfolio optimization, stochastic programming, derivative pricing. Python, Go, Rust, TypeScript, Java.




