Technical AI Safety
I come from a customer service background — I spent several years as a Technical Support Engineer, working across SaaS products. In early 2022 I became interested in AI and decided to pursue a master’s degree in the field — an MSc in Advanced Computer Science with AI at the University of Strathclyde. One month into the programme, ChatGPT was released, and the world I’d just entered changed overnight.
For the next couple of years I leaned toward building — prototyping apps with AI, learning to ship things, picking up the engineering side. But the research questions kept pulling me back. In early 2025 I found my way to AI safety: the problem of making sure these systems do what we actually want, especially as they grow more capable.
In 2026 I joined the BlueDot Impact Technical AI Safety Project Sprint, where I completed a research project on activation-space interventions for jailbreak defence. That work is below.
My current work is on the infrastructure the field uses to read itself. AI Safety Map is a live map of where AI safety research is heading, built to be sliced by the question being asked rather than fixed to a single view. Explore: An AI Safety Wiki takes on the other half of the problem: research that is prerequisite-heavy and thin on exposition, which makes the field expensive to enter. Both began as tools I wanted and could not find, and I have come to think they address one problem rather than two — which is what most of the writing below is about.
Research generation is automating faster than research evaluation, field-building keeps widening the funnel, and a handful of correlated buyers price the field’s directions. Observations on a squeeze the field’s epistemics is not set up to handle.
A tool for reading AI safety research — papers unbundled into cited concepts and relations, anchored by a living, human-curated canon — and the field-level orientation layer it leaves behind.
How you actually build the self-observing map — version control for a field, a check pointed at sources and never at the canon, and a toolmaker who keeps it forkable and steps out of the way. A companion to The Loop No One Chose.
Why a funder willing to back neglected work still needs the map — and without it reproduces the very loop they mean to escape, only running it in reverse. A companion to The Loop No One Chose.
Reflexive dynamics in decentralized nonprofit fields, and why AI safety needs a live, actionable map that reflects the field back to itself instead of telling it what to do.
A case for using different directions to detect and correct jailbreaks. Splitting detection from correction nearly doubles refusal rates on jailbreak prompts at no capability cost.