I build reinforcement learning systems — mostly offline RL, policy evaluation, and the training infrastructure that makes experiments reproducible instead of a folder of one-off scripts.
Currently a research engineer at a robotics lab, working on manipulation policies. Previously built RLHF training infra at scale.
About
I spend most of my time on the gap between "works in a notebook" and "works reliably at scale." A lot of RL research dies to bad logging and irreproducible seeds — most of my tooling exists to fix that.
Outside of work: bouldering, and a running list of unfinished side projects.
A result you can't reproduce isn't a result — it's a rumor.