Technical AI Safety
My first job was as a game programmer. After working for over a year, I had an epiphany and realized that trading cognitive skills for money was not a proposition I enjoyed. My belief at the time was that cognitive abilities are personal property, only meant to work on things that one personally enjoys.
In search of a job that demanded minimal cognitive skills, I decided to work in customer service, and nearly spent a decade working there.
However, in late 2021 I realized people are talking about AI and its potential impact on my job as a customer service executive, and decided to look into it. I realized quickly there is a lot of activity, and even heard about GPT-3. I was very convinced that it was a PR stunt, but at the same time, at the back of my mind, if this is true, this is my dream of having commoditized intelligence, where human intelligence is only used for purposes they actually wanted. The trend was hard to miss. There was a lot of activity, most of it was in Kaggle and ML for data science. I finally decided to get to the bottom of this, and decided to do a masters in AI from the University of Strathclyde in late 2022. Two months into my masters, ChatGPT got released, and I knew this is not hype but it is real.
I got interested in technical AI safety in early 2025. I realized that there is an asymmetrical risk associated with AI, where the risk of absolute ruin grows exponentially with capability, and that means the future of humanity gets increasingly dependent on how technical AI safety achieves its goal. I believed I had a responsibility to contribute, and then I attended a BlueDot course and I completed a project on persona vectors.
Now, I think about this differently. AI safety research is going to get automated very fast. There is not a lot of value in working on and being an expert on a very specific research area, as AI will be able to do that research better than anyone else at some point in the future, probably very soon. Should that be the case, it is going to be many folds more valuable to be a researcher who can make sense of generated research and find directions that meaningfully reduce catastrophic risks.
Generation of new research is getting cheaper, and evaluation is expected to get cheaper in the future. However, understanding the implications of new research, and how that research augments the rest of the field in reducing existential risks, are two areas that are not getting the attention they deserve, even though these two functions increase in value exponentially as model capabilities increase and are not expected to get automated very soon. Furthermore, whether to fully automate them or not is a highly debated question. I am of the strong opinion that handing over the responsibility of making sure that frontier models are safe to frontier models themselves is a terrible choice.
However, there are other epistemic aspects, like discovery and distribution of new work, and how the attention dynamics can bias the research efforts, where certain research directions get more attention because they are already legible and there is a critical mass of papers and talent, and some research directions are starving because of a lack of existing research and talent, like there is no existing research because there is not enough talent pool, or talent is discouraged because there is no ecosystem to support that type of research.
In general, I believe the amount of research that AI is going to generate is staggering, and we should have the right epistemic infrastructure in general to handle such a volume of output. I have written an article about why this is important.
I am currently focused on making sure we have the right epistemic infrastructure to handle the AI generated research.
An example of one such tool. This is a tool that I personally use, and I have also shared it publicly. Skimmaxxer helps with reading research papers that require you to understand complex prerequisites.
In future I may come up with other tools to support the epistemic infrastructure, possibly one that helps researchers understand how a new result augments the rest of the field, what its implications are, and what future possibilities it enables.
Generating research is getting cheap; evaluating it and putting it in context is not. Why that asymmetry, in a correlated field with no real-world feedback, rhymes with 2008 — and who gains if the field can see itself clearly.
A case for using different directions to detect and correct jailbreaks. Splitting detection from correction nearly doubles refusal rates on jailbreak prompts at no capability cost.
Research generation is automating faster than research evaluation, field-building keeps widening the funnel, and a handful of correlated buyers price the field’s directions. Observations on a squeeze the field’s epistemics is not set up to handle.
A tool for reading AI safety research — papers unbundled into cited concepts and relations, anchored by a living, human-curated canon — and the field-level orientation layer it leaves behind.
How you actually build the self-observing map — version control for a field, a check pointed at sources and never at the canon, and a toolmaker who keeps it forkable and steps out of the way. A companion to The Loop No One Chose.
Why a funder willing to back neglected work still needs the map — and without it reproduces the very loop they mean to escape, only running it in reverse. A companion to The Loop No One Chose.
Reflexive dynamics in decentralized nonprofit fields, and why AI safety needs a live, actionable map that reflects the field back to itself instead of telling it what to do.