I keep things here.
Some of them work. Some of them worked once. Some of them probably shouldn't have worked at all.
Mostly around machine learning systems, model behaviour, interpretability, agents, RL environments, robotics / VLA, reasoning systems, research infrastructure, and open-source rabbit holes.
A few things I’ve left lying around:
gs-dronegym
photorealistic drone simulation + trajectory tooling for VLA research
axon
real-time SAE feature visualisation inside language models
BrainPatch
activation-space interventions, runtime steering and behavioural experiments
LEMMA
neuro-symbolic mathematical reasoning with search and rules
I also spend a fair amount of time inside other people's repositories.
A few detours have ended up around Apache Airflow, Qiskit, Kornia, uv, OpenTelemetry, Pandera, Litestar, Toqito, TorchGeo and Bespoke Curator.
Usually fixing something small enough to understand, but annoying enough that I couldn't leave it alone.
I care about whether an experiment survives controls, whether a benchmark is actually measuring what it claims to measure, and whether the code still works when the environment gets weird.
Usually Python. Sometimes Rust, Zig, TypeScript, Swift, or whatever gets the job done.
And yes, there are a lot of forks here.
Fixed it. Enlightened the repo. Disappeared.
“Always say thanks to your GPT at the end of the chat. You never know when it turns Skynet.” — me, 2026





