Essay
What is the next big crisis AI safety is going to face very soon?
Before answering our central question, let me take you back to 2008 and tell you what caused the financial meltdown and what is so interesting about it. If you are wondering why on earth we are talking about a completely unrelated event that happened roughly 20 years back, instead of working on pressing research questions that could prevent AI from taking over the world, I would totally understand it. I will try to answer that question slowly but thoroughly in the coming paragraphs. First, let me explain what was unsettling about the crisis and why it may or may not have relevance to AI safety.
For hundreds of years, banks have lent money that people never paid back. But how the same thing could grind the entire global financial system to a halt is still a mystery to a layman, or sometimes even to an economist. I will try to unravel that mystery a little bit here.
History rarely repeats, but it often rhymes
This catastrophic meltdown was partly enabled by highly complex derivatives, such as synthetic CDOs and over-the-counter credit default swaps. They were so complex, so much so that even seasoned experts often struggled to explain how these instruments worked or what role they played in the crisis. Nonetheless, I believe that if you prompt the frontier models enough, you can hope to get a decent understanding of them, or alternatively, you can choose to watch ‘The Big Short’ movie, depending on your free time, attention span and the status of your Netflix subscription. Back in early 2007, policymakers and the lenders were regularly seeing defaults of individual loans. They quickly understood the risk those defaults posed to their own organisation. However, they were completely oblivious to the extent of the risks it posed to the interconnected global financial networks. So they low-key missed the bigger picture, and they missed it with spectacular flamboyance, as you all know.
There were valid reasons for this structural blindness that persisted among the regulators, lenders and bankers. Firstly, it was really hard to access granular data, and secondly, it was painstaking to analyse it even when they could access it. So the system relied on aggregate data and proxies, like agency ratings, algorithms, and mathematical models, in order to understand market direction and the extent of the risks they were taking. Regardless, Michael Burry, the investor who predicted and profited from the financial crisis, could read the situation clearly only because he was willing to put himself through the pain of reading documents that the rest of the system deemed too obscure to read. Such as the bond prospectuses (to understand the nature of the bonds) and the loan tape (to understand the individual borrower behaviours like delinquency). The point here is that individually, institutions were able to competently assess the risks to their own organisation, yet collectively no one really could see a faithful picture of the systemic risks except for a handful of people. So the obscurity of the critical information or inability to recognise the criticality of the visible information contributed directly to the crisis, as much as or even more than the intentional malpractices.
Now, back to 2026: does AI safety suffer from the same epistemic issues the 2008 financial crisis had? Absolutely not. In 2008, the rating agencies were intentionally rating bad bonds as AAA, and then the banks intentionally sold those AAA-rated CDOs to pension funds even though they believed it was subprime. I don’t think AI safety has such malicious practices; everyone is by default assumed to be acting in good faith, and it’s not centralised, meaning there are no rating agencies. If there are no malicious actors or centralised control, can we say AI safety is immune to such epistemic risks? Can we, at least now, focus on more relevant issues like prevention of takeover by a superintelligence? I am going to answer those questions in a while. We are almost there. Please bear with me. You have top-notch patience, by the way. You could easily work in Comcast customer service if it were 2008.
If AI safety is not going in the right direction, would we know?
Or, in other words, what would be the best way to check AI safety is going in the right direction?
Let’s just analyse the inputs into the field and the outputs it produces and see what we can come up with.
The inputs are effort put in by the researchers and field builders, funds deployed by the grantmakers, proposals created by the field strategists, etc.
Outputs are research output produced by the researchers, talent attracted and placed by field builders, performance of the funded orgs and individuals, etc.
The problem here is that we don’t have access to live, granular data on both; the second problem is we don’t have the right tools to parse it. The third problem is that even after collecting it and parsing it, we don’t have the right tools to present it.
What we have is a manual, slow process of people manually collecting the data and parsing it and presenting it as LessWrong posts and annual safety reports. So the current epistemic infrastructure directly relies on reports which are centralised, or on informal discussion in forums like LessWrong, or even on private interactions between the top researchers.
Moreover, since most of the technical safety work is done within a few orgs that are working very closely, it is possible that the research directions might skew to a particular direction not by choice but because of the interconnected nature of active organisations, i.e. frontier labs, AISIs etc. Correlation itself is not a bad sign; that is how the field moves fast and coordinates effortlessly with minimal bureaucracy.
The current epistemic infrastructure of the field
Can we meaningfully estimate the direction the field is heading from such practices and analysis?
In principle, yes, as long as the data is live and the analysis is bias-free, comprehensive and easy to understand. Often, in the real world, by the time these posts or reports come out, they lag behind the data, even though there is enormous effort behind this. Finally, as these are centralised efforts, they are prone to unconscious biases.
Usually, these reports are presented in a generic format, which may not be optimal from a specific perspective, because a grantmaker, a researcher, or a field builder expects different information from a report. A researcher wants to know which open problems are underexplored; a field builder wants to know where talent is bottlenecked. A grantmaker may want to know how his deployed capital is augmenting the rest of the field’s efforts.
Here, valuable information is being obscured by presentation and distorted by stale data. Another issue here is that we are relying on the creators of the report for insights regarding the health of the field, or on someone who publishes a post in LessWrong or a similar forum after doing a lot of this work individually.
One of the structural problems with AI safety is that it’s not real-world feedback-based, like civil engineering or aerospace engineering. In those fields, if the underlying science is wrong or the engineering is imprecise, we can observe that from the quality testing. We don’t have such avenues for AI research. Instead, what we have are proxies such as benchmarks and evaluations that the field believes accurately represent future real-world scenarios.
Another issue is that we might be collectively steering away from the safe path, even when we are able to verify individual papers through the peer review system. Because in peer reviews, they validate whether the claims are right or how the paper advances the research frontier, not whether the frontier research itself reduces the x-risk from a superintelligence. The latter part is done by the field as a whole through collective discourses. Current sensemaking efforts indirectly rely on the instincts of funders, the judgment of mentors, the thousand small decisions about what gets read, how it gets cited, and who gets hired, etc. All these are part of the epistemic infrastructure which decides the direction the field is going as a whole.
We might be collectively steering away from the safe path, even when we are able to verify individual papers through the peer review system.
There is no formal process for this. So how that consensus is reached through that collective discourse is of extreme importance if we want to make sure the field is actively contributing to reducing risks from AI capabilities. So what the rest of the AI safety field provides is direction, critique and prioritisation, and the main contribution of the AI safety field is of an epistemic nature rather than technical, directly through the discussions it facilitates and mentorship it provides and indirectly through the talent the labs or safety institutes absorb.
If the existing systems somehow work, why should we worry about this so-called epistemic infrastructure at all?
AI is increasingly automating generation of new research. The volume of research is going to increase exponentially, and we are already seeing signs of this in NeurIPS submissions and LessWrong posts[1].
Research has four parts: generation, verification (whether the new research is meaningful), exposition (how this contributes to the field) and distillation (communication of the new research to the rest of the field). Advances in AI capability have helped the generation disproportionately more than they have the other parts. So current epistemic infrastructure has to rely on scarce manual labour from even scarcer experts. Although it is functional as of now, it is fragile.
This asymmetry, where the generation of new research got really cheap, but the verification of that research did not, and the act of putting that research into the context of existing work also did not, is severely straining the current epistemic infrastructure, and it may very soon start affecting the quality of the research and, in turn, the future of humanity.
The closing of the loop, where a frontier model creates and evaluates its own research, has not started happening yet in AI safety as it has in capability research. It might happen after a while, even though it’s debatable whether automated AI safety is a reasonable choice or not. Regardless, the idea is to have safety research positioned ahead of capabilities research before models that are capable of recursive self-improvement arrive.
On top of this, the field is growing fast. A lot of field builders are working hard to attract talent[2]. The new talent is going to produce new automated research, and it is going to severely strain the already fragile system. The system was never designed to handle this level of volume.
Another issue with the research that is not contextualised to existing work is that the discovery of new relevant work becomes difficult, and very soon, attention would start becoming a cheap proxy for value, which becomes a loop where some genuine research directions starve and some directions get more attention than they deserve, just because they were visible or because it was able to generate a discussion around it[3]. Here, there are no malicious actors, no central authority to blame, no one to criticise, everyone is acting in good faith, yet the same information obscurity that threatened the global financial system quietly presents itself in a new form. The role of malicious actors is being played by the volume of LLM-generated research and the self-reinforcing attention loops[4].
The role of malicious actors is being played by the volume of LLM-generated research and the self-reinforcing attention loops.
What are the major epistemic risks
In scientific research, it is documented that proposals that confirm existing beliefs get approved easily[5], while what gets counted as a successful outcome once the experiment is done has to be something that challenges the existing belief[6]. This presents the researchers with a dilemma, where the proposal has to confirm existing beliefs and the actual experiment has to challenge the existing beliefs. This contradictory set of requirements restricts their research directions severely. Most of the time, instead of arguing with the system and defending their idea, the researchers end up self-censoring certain topics, which has either a chance to get rejected at the proposal stage for being too novel, or during the venue submission for ‘not’ being novel.
Here in AI safety, a response to a higher volume of applications or research output is almost always reactive rather than foresightful. Most obvious responses will be raising the bar to entry or resorting to easier evaluation techniques. For example, grantmakers might favour topics that are easier to evaluate or use proxies like personal references in order to aggressively filter the volume. If such patterns emerge, proxies like personal references work in isolation, but they will add to the correlation risk the field already has. It could cause the funded projects to look skewed in a particular direction. Similarly, if grantmakers favour a certain type of application or projects in order to combat the volume, even though it is a reasonable approach individually, it can easily become a Goodhart target, and researchers will self-censor the rest of the research directions and produce content that is more likely to get a grant. I am not saying that all of this is guaranteed to happen, but it is a possibility; its likelihood increases as the volume strains the current infrastructure.
We do have a niche category of field strategists who keep track of the field and make sure that everyone is working on the leveraged problems; they too are working with less-than-optimal visibility because of the same structural issues. Ironically, they lack the right epistemic tools needed to make sure their bets are made with all the available information. Right now, they talk to researchers or do their own parsing of public data, probably all of this manually. If two researchers have competing worldviews, they may have to rely on proxies like credentials, or put in a lot of manual effort. Field strategist work can be very difficult without a reference frame. Without a clear understanding of where the field stands now, it will be difficult to plot future trajectory. There should be a proper pipeline for extracting the tacit information from the researchers into a consolidated format in order for them to work consistently on identifying highly leveraged problems.
There are going to be competing worldviews among the researchers and funders, so a lack of transparency on how the consensus is made, whether there exists consensus at all, or how to double-check the reached consensus. All of this is going to be extremely important. This is a field that is going to be faced with multi-faceted risks such as skewed research directions by correlation, inaccurate proxies for future real-world scenarios, high velocity volume produced by automated research.
Who benefits
This is something I believe currently has the least attention, but it has the potential to be the highest leverage activity if we prioritise it. The consequences of the 2008 financial meltdown were actually exacerbated more by the fact that there was a lack of epistemic infrastructure than by malicious practices. If everyone had a clearer view and knew what was about to happen, it may not have devolved into a full-blown global meltdown. It was not a true black swan event; it could have been avoided. If we have a clear way to measure the progress of AI safety, it will actually help the capability research to accelerate rather than slow down, instead of having an endless debate on whether to pause AI or not, or we need more regulations or not, etc. It’s very easy for a highly correlated field which has no real-world feedback to drift off the safe path. We certainly don’t want this, since all agree that the stakes are too high.
If everyone had a clearer view and knew what was about to happen, it may not have devolved into a full-blown global meltdown.
Researchers — Easy to discover relevant new work, enabling fast collaborations or follow-up work. Similarly, relevant work does not have to wait in arXiv to gather enough citations to get discovered
Grantmakers — They get a reference frame to appraise the grant proposals they receive. The visibility of the rest of the field can help them accurately check if a proposal makes strategic sense
Field strategists — They will benefit from this enormously, as they get a clear view of the field, so their suggestions and plans are more grounded in reality, and it also helps them identify promising directions fast or understand why a certain direction failed to deliver. With the current infrastructure, it is extremely difficult to see why a certain direction failed to deliver or quickly spot a new promising research direction without putting in tremendous manual effort.
New talent — How to orient themselves to the field by understanding the state-of-the-art research. This can help them tremendously to get up to speed with existing work and start providing value immediately, instead of waiting on the sidelines or being a part of the slope.
Footnotes
- arXiv blog (Oct 2025). “Attention authors: updated practice for review articles and position papers in arXiv CS category.” https://blog.arxiv.org/2025/10/31/attention-authors-updated-practice-for-review-articles-and-position-papers-in-arxiv-cs-category/↩
- “AI safety talent needs in 2026: insights for field-building.” EA Forum. https://forum.effectivealtruism.org/posts/jwwrC4n9H53doRjRH/ai-safety-talent-needs-in-2026-insights-for-field-building↩
- Frank, R. H., & Cook, P. J. (1995). The Winner-Take-All Society. Free Press.↩
- Merton, R. K. (1968). “The Matthew Effect in Science.” Science, 159(3810), 56-63. https://www.science.org/doi/10.1126/science.159.3810.56↩
- Boudreau, K. J., Guinan, E. C., Lakhani, K. R., & Riedl, C. (2016). “Looking Across and Looking Beyond the Knowledge Frontier: Intellectual Distance, Novelty, and Resource Allocation in Science.” Management Science, 62(10), 2765-2783. https://pubsonline.informs.org/doi/10.1287/mnsc.2015.2285↩
- Gross, K., & Bergstrom, C. T. (2021). “Why ex post peer review encourages high-risk research while ex ante review discourages it.” PNAS, 118(51), e2111615118. https://www.pnas.org/doi/full/10.1073/pnas.2111615118↩
LLM policy
Unlike the other essays on this site, this one was written entirely by hand. No language model wrote or reworded any sentence of it.
The machine’s share of the page is everything around the prose: the HTML the essay is wrapped in, the diagrams, and a copyedit I asked for — spelling, punctuation and grammar only, with the sentences left as written. The diagram captions and the pull-quotes reuse the essay’s own sentences verbatim.
This note is the one exception: it was written by Claude.
Read next →
The Coming Epistemic Crisis AI Safety Is Not Ready For
The longer treatment of the same squeeze: research generation automating faster than evaluation, a widening talent funnel, and a handful of correlated buyers pricing the field’s directions.