selected researchall 8 →
- 2026-09-27 — agents / formal verification Agents turned a one-kernel Lean proof into a checker that certified 25 PyTorch kernels One Lean theorem and a small checker certified 32-bit size arguments for 25 torch.compile kernels, including 10 of 24 held out, with each proof bound to the live compilation. Two of nine tuned kernels got faster.
- 2026-09-26 — agents / formal verification A Lean proof let agents cut a compiled PyTorch workload’s runtime by 26% Agents declared a torch.compile kernel’s size arguments 32-bit, proved in Lean when that is exact, and cut a changing-shape workload’s time by 26% on one L4. The opt-in patch is an open PyTorch PR.
- 2026-09-25 — agents / formal verification 25 bugs in free-threaded CPython, found by agents Agents confirmed 25 bugs in the no-GIL build, with no earlier report found for 18. TLAPS and Lean proofs of the locking model held; replays of real runs showed five places CPython leaves it.
- 2026-09-24 — agents / formal verification I used formal verification and agents to make an algorithm 105× faster and prove it behaves the same Claude Opus 5.5 rewrote four deliberately slow Dafny routines under a frozen spec it could not edit. Seven rewrites verified, and one ran 105× faster on a workload it never saw.
- 2026-05-26 — rl / sampling bias A 50-step RL update reduces categorical sampling bias Fifty RL steps on one random-integer task moved nine untrained pick-one tasks toward uniform on Qwen3-30B-A3B-Instruct.
- 2026-04-25 — interp / probes Hidden-state probes outperform self-reported confidence A linear probe on Llama 3.1 8B’s hidden states ranks claim correctness better than the model’s stated confidence.
selected softwareall 25 →
-
A Claude Code skill: describe a task, and Claude writes a multi-agent workflow script, runs it on Codex agents instead of Claude subagents, and shows the run as a live map. Agents edit files without asking unless the run is read-only.
-
Train a small chatbot from scratch on Apple Silicon, from tokenizer training to a chat interface.
Why I built it: I wanted to train a chatbot from scratch on my MacBook without touching PyTorch or a cloud GPU.
-
Find duplicate and stale Safari tabs, search open pages, and group tabs by domain.
Why I built it: I had 80+ Safari tabs open and no way to make sense of them.
brew install --cask scasella/tap/tabpilotsource -
Adjust Night Shift warmth directly from the menu bar.
Why I built it: Night Shift warmth is one of the few display settings I actually want to tweak throughout the day, but Apple hides it behind too many clicks.
brew install --cask scasella/tap/sunshiftsource
model adapters
-
Random-choice adapter
A Qwen adapter for experiments with fixed-list sampling behavior.
-
Panel-reasoning adapter
The Qwen adapter used in the accuracy and completion-length comparison.
contact
I’m interested in replicating LLM experiments and building useful open-source tools. Get in touch about research ideas, corrections, or collaboration.
For bugs and code, open an issue on the repo in question.