Essay

The Coming Epistemic Crisis AI Safety Is Not Ready For

Observational· grounded in what the field is visibly doing nowSpeculative· based on untested yet promising ideas

Some observations on research automation and the epistemics of AI safety

Manu Xaviour Thaisseril Shaju · July 2026

In an ideal world, I do my research, get good results, contribute to the field’s efforts, and am content that I have done my part; that is roughly what an independent researcher who wants to understand how we make AI safe for humanity expects. Most of us proceed on that assumption, move on with our lives, and everything is fine. If you are willing to entertain a few questions about whether that worldview holds, you can see a different story unfolding. What follows is a set of observations. Individually, most of them are boring, like arXiv tightening submission rules, LLM use in peer review, and the competitiveness of the fellowships. Together, they point to a squeeze, which could cause epistemic misalignment. Concerningly, its severity depends almost entirely on collective epistemic infrastructure the field has not built yet. What is surprising is that every action described here is individually rational, and yet the aftereffects on the field’s strategic self-awareness converge on something close to singular.

cheap research the evaluators the buyers no external check
Research is getting cheap to generate faster than the field’s few evaluators can judge it, and a handful of buyers — grantmakers and labs — decide what it is worth. The dashed line is the missing floor: no external check, no bridge that stands or falls, to settle who is right.

1. Research generation is being automated faster than research evaluation

Language models are becoming genuinely useful across the research pipeline: literature review, experiment scaffolding, analysis, drafting. One underrated upside is that this lowers the cost of moving between fields. When the topic-specific grind gets cheap, a researcher’s ability transfers more easily; expertise becomes more topic-agnostic. For AI safety, which recruits most of its people from other disciplines, that is good news.

But the same tools lower the cost of producing plausible-looking research much faster than they lower the cost of judging it. Evaluation still runs on senior human attention: reading, checking, placing work in context, deciding what matters and what gets funded. That attention does not scale with model capability. This asymmetry, generation getting cheap while evaluation stays expensive, is the inflection point, and everything in this essay cascades from it.

Two ideas up front

  • The field’s collective epistemics is the machinery by which a field works out what is true and what is worth working on: peer review, the forums where work gets discussed, the instincts of funders, the judgment of mentors, and the many small decisions about what gets read, how it gets cited, and who gets hired.
  • Most fields can afford to run that machinery loosely, because reality eventually checks the answers. The bridge stands or it falls.
  • Its strategic self-awareness is what the same machinery produces when the field turns it on itself: what the field as a whole is betting on, what would have to be true for those bets to pay off, and what nobody is working on at all.
  • The reason to keep the two ideas separate is that a field’s strategic self-awareness depends on its collective epistemics.

Two ideas will carry a lot of weight in what follows, so it is worth being plain about them up front. The first is what I will call the field’s collective epistemics: the machinery by which a field works out what is true and what is worth working on. Some of it is formal, like peer review. Most of it is not: the forums where work gets discussed, the instincts of funders, the judgment of mentors, the thousand small decisions about what gets read, how it gets cited, and who gets hired. Every field has some version of this machinery, and most can afford to run it loosely, because the real world eventually checks the answers: the bridge stands or it falls, the app works or not, string theory is proven or not, you get the idea.

The second idea, strategic self-awareness, is what that machinery produces when the field turns it on itself. Not “is this paper good?” but “what are we, all together, betting on?” Which problems is the field pouring effort into, what would have to be true for those bets to pay off, and what is nobody working on at all?

The reason to keep the two ideas separate is that a field’s strategic self-awareness depends on its collective epistemics.

2. The volume shock is already visible, and institutions are responding by raising the bar

Detector-based analysis of ICLR 2026 found several hundred submissions written wholly by LLMs, and roughly one in ten with majority-LLM content. arXiv stopped accepting position papers and review articles in computer science without prior peer review, citing the volume. LessWrong holds LLM-generated content to a higher bar than human writing. Meanwhile, the same analysis found one in five reviews fully LLM-generated, closing a loop nobody designed, in which machine-written papers meet machine-written reviews and human judgment thins at both ends.

These responses are mostly reactionary: they gate work by raising the cost of entry. That collides with the upside from (1): cheap field-switching is worth little if nobody can afford to read what the switcher produces. Field-building runs into the same wall: it keeps widening the funnel on the assumption that newcomers will be integrated into the field, and expensive evaluation of their work is exactly what slows that integration down.

Nobody is making evaluation cheaper. Everyone is making entry more expensive.

3. Field-building trains more candidates than the field can absorb, and their output is straining evaluation capacity

Field-building programs are scaling admissions faster than the field scales absorption. Introductory programs plan to train 100,000 people by 2030; research fellowships like MATS accept applicants at rates in the low single digits, a hyper-selectivity that the field’s own needs assessment traces to scarce supervision capacity. The people who are trained but not placed do not vanish. Underemployed, motivated, and equipped with the same generation tools as everyone else, their output lands on the field’s open surfaces: grant applications, forum posts, conference submissions, cold emails to researchers.

The problem is not that weak work exists; weak work has always existed. The problem is that filtering consumes the scarcest resource the field has, the attention of the people qualified to filter: mentors, admissions committees, grant evaluators, forum moderators. Slop is better understood as the indiscriminate spending of evaluator attention than as pollution.

None of this requires the work to be bad. Suppose most of it is decent and some of it genuinely contributes: with no infrastructure to absorb the volume, the same systems get stressed, and a stressed filter does not select for the best work. It selects for the work that is cheapest to judge. Legible, agenda-adjacent work tends to clear the bar; work that would take an afternoon to understand often does not, however good it might be. That is not a neutral filter, and it is worth asking how it skews the field’s output. (4) takes up what follows.

4. Even though grantmakers have the most leverage, they ironically have the faintest signal

Consider what a funder can realistically consult: shallow reviews, forum and Twitter write-ups, yearly organizational reports, grant applications, and personal connections. Is that enough of a view of the field? Probably not, and a funder making the individually rational call on that signal is not necessarily making the call that serves the field. There is precedent for how this goes. Research on science funding describes a squeeze that operates from two directions at once. Ex ante review, which determines what gets funded, favors proposals that are not especially speculative and that are bold in the direction of what the field already believes, since those are the proposals a reviewer can assess quickly and score with confidence. Ex post evaluation, which determines what counts as a successful grant, favors novel results, and those generally require speculative, high-risk experiments that change prevailing beliefs rather than confirm them. Because the two criteria point in opposite directions, the space of work a researcher can realistically propose is squeezed from both sides, and the cheapest response, or even the sensible one, is to quietly self-censor whole topics rather than defend them. I am not claiming that exact dynamic will replicate here. But some version of research being squeezed toward a highly correlated direction seems likely, and without any view of the whole research landscape, there would be no way to notice it is happening.

That reflex is not confined to grant proposals. It would attach to any map of the field that became canonical, which is the failure Forethought files under epistemic misalignment: once a map is what funders and hiring committees consult, the rational move is to look legible on it, and what is easy to represent starts crowding out what is true. A small and correlated field makes that a stronger Goodhart target than a large one would, not a weaker one. It is the most serious objection to what (6) proposes, and the appendix answers it.

5. No external check, a handful of buyers, and no way to measure the correlation

AI alignment is not a field with external signals the way computer science or civil engineering are: most of its concerns are about systems that do not exist yet.

The field’s epistemic processes are not a support function. They are close to being the product.

Most direct safety work happens inside a handful of frontier labs and a few independent organizations. What the surrounding field mainly supplies is direction, critique, and prioritization. Its technical contributions are marginal. Its epistemic ones are not: directly through the discourse, indirectly through the talent the frontier labs absorb.

On the market side, there are a few frontier labs, a few independent safety orgs that supply the frontier labs, and a few governments. The nodes read the same forum, share advisors, and hire from one another; newer entrants such as national AI safety institutes staff themselves from the same pool. The talent pipeline compounds this: a small number of mentors’ taste is replicated into each cohort, a shared curriculum forms the canon, and a few venues host most of the deliberation.

Concentration is partly load-bearing: it is how a small field moves fast, maintains trust, and filters work. The convergence may even be correct; perhaps the dominant agendas deserve their current weight. In a normal field, false consensus eventually collides with reality; here it may not. That leaves provenance as the only remaining test: did several funders converge on the same bets independently, or because the bets share three advisors?

Nobody currently measures this. There is no concentration index for the field’s collective portfolio, no funder-overlap statistic, no way for any individual funder to know how correlated their bets are with everyone else’s. The pre-2008 financial system is the cautionary parallel: every institution looked diversified on its own books, and the correlation lived in a counterparty network no one had mapped. History never repeats, but it often rhymes.

6. Starting points and a general direction for what collective epistemic infrastructure could look like

The collective version is not mysterious, and it does not have to be a map. It starts with a data layer: continuously updated records of research output, funding flows, which organizations work on which agendas, and which venues host what. Not snapshots. The field changes faster than any annual review can track, which is part of why the annual reviews lost their charm (8).

On top of the data layer, we can build a tool layer that makes understanding and evaluating research cheaper for humans, without handing either over to the models. The asymmetry of (1) is that generation got cheap while evaluation stayed expensive; the fix is to multiply scarce evaluator attention, not to delegate the gate to the frontier models that produced the volume. (2) already shows how that loop closes.

The test for corpus-level analysis is speed. Today, anyone who wants to know where the field is heading either runs a three-month project with four FTEs or trusts someone who did. Seeing the field’s direction should be a query, not a research program. Funding data deserves the same standard: the concentration and overlap measures of (5) should be standing instruments anyone can consult, not bespoke studies commissioned after the worry has surfaced.

The labor for this already exists, and (3) described it: the residual cohort, trained and motivated and unplaced, ingesting new research, distilling it, situating it within the field. On first reading this is the worst answer on the table. The scarce thing is evaluator attention, and the proposal is to route the volume through people who have never been evaluated themselves, the same people whose output is straining the filters. Against a model that at least applies one standard consistently, it looks like a downgrade with a mission statement attached.

What that reading misses is why the cohort has no track record. It is not that they were assessed and found lacking. Almost nothing they produce is assessed at all publicly; they write in volume into a system with limited capacity to read them, and the missing record about their quality is a large part of why they stay unplaced, along with limited upstream opportunities. A structured channel changes that: a defined slice of the corpus to cover, a cadence, and experienced people reviewing what comes out. The reviewer’s time then buys three things where it used to buy one. Feedback the newcomer can act on. A public record of whose judgment holds up, which is what the field’s filters currently assess in private and only for the few who get far enough to be assessed at all. And the distillation itself.

That is the leverage, and it is why this is not the same as handing the reading to the models. Mentor attention spent on a submission today disappears into a yes or a no. Spent this way it compounds: it trains someone, it produces a record, and it scales across far more of the corpus than the mentor could have read directly. The work orients the cohort in return and pays down the research debt of (8).

Whether it converts is a question about the protocol rather than the premise. What fraction of the output gets reviewed, how a reviewer’s verdict is recorded and weighted, what happens when two reviewers disagree, how much of the corpus a distillation has to cover before anyone consults it: those choices decide whether mentor attention actually turns into collective epistemic quality, or only into a better-trained cohort. The second is worth having on its own. Only the first pays for the build.

And none of it works as a static artifact. It needs an active community and a channel: a standing place where new work gets absorbed as it arrives, and where the point is helping people find understanding and share their own.

There is modest precedent for funding work of this kind: worldview-investigation and forecasting projects were supported when the open question was AI timelines, those efforts were early, and recently, Forethought has written about strategic awareness and collective epistemics as a category. The combination of research automation and field-building residual cohorts makes the next year or two the period in which this stops being early.

That paper is also the nearest thing to a design document for a build like this, and it names the two ways such a build fails: work shaped to look good on the map, and one map becoming unquestionable. Both apply to what is described above, not only to the general program, and (5) makes both worse here rather than better. The appendix takes the objection at length, along with the partial answer available to it.

A field’s strategic self-awareness is a product of its collective epistemic infrastructure.

7. What happens if nothing is built

Sorted by confidence: the first items are visible trends continuing, the middle ones need assumptions, the last are long shots included for their size.

8. Is it worth it?

It is not going to cost much, in absolute terms or relative ones: a data pipeline, a few standing tools, a channel, and stipends for people the field has already paid to train. Field-building alone has absorbed roughly $320 million. The parts described here would cost a small fraction of that, and unlike a cohort or a conference, infrastructure keeps paying out for as long as the field uses it. And it goes straight at the problem in (4): funders have the most leverage and the worst view of the field, and this is the cheapest way to improve the view. Since, by (5), the field’s collective judgment is most of what it produces, better judgment raises the value of everything else. It may be the highest-leverage money the field can spend.

Chris Olah and Shan Carter named the underlying condition years ago: research debt. Distillation is undersupplied, so papers stay harder to read than they need to be, and the cost compounds with volume. In alignment, the symptom is concrete. The once-annual comprehensive review of the field’s literature was discontinued largely because no single person could read it all anymore, and its best-known successor is titled, candidly, a “shallow review.” The sense-making layer was already struggling with volume, even before the surge.

The $320 million figure comes from Elena Ericheva’s A genealogy of AI safety: how directions are born, and how they die, a solo reconstruction of two decades of the field’s directions and funding flows: careful, rigorous, and done by hand. That provenance makes the point by itself: this is what it currently takes to get a view of the field, and even then the result is a snapshot, not something live. Ericheva was not trying to build epistemic infrastructure. But the skill set and the process are exactly the ones the infrastructure needs, which means the people who can build the live version already exist.

Appendix: sources and notes

The essay compresses several empirical claims into single sentences. This appendix records what each source actually says, and what it adds beyond the sentence it supports.

The gates (2). The arXiv change (October 2025) requires that review articles and position papers in the CS category complete peer review at a journal or conference before they can be submitted; papers without documentation of that review are rejected. The stated reason is capacity: the category now receives hundreds of review articles a month, most of them, in the moderators’ description, closer to annotated bibliographies than to synthesis, and a volunteer moderation team cannot absorb the load. Note the mechanism. arXiv is not making evaluation cheaper; it is relocating the cost onto refereed venues. The gate works by making someone else’s gate a prerequisite.

The Pangram analysis (November 2025) covered roughly 19,000 ICLR 2026 submissions and 70,000 reviews. On the paper side: several hundred submissions fully LLM-generated, and about one in ten with majority-LLM content. On the review side: about one in five reviews fully LLM-generated, and more than half showing some level of AI involvement. Two second-order findings matter more than the headline. AI-heavy papers received lower scores; the flood is mostly not fooling anyone yet. But AI-written reviews assigned higher scores and ran longer, meaning part of the gate layer has quietly automated itself, and the automated part is systematically more permissive and less informative. That is the loop of (2) with numbers attached: cheap generation meeting cheap evaluation, and the cheap evaluation waving it through.

One caveat on provenance. Pangram sells AI detection, so the prevalence figures come from the party whose market grows if the prevalence is high; they are a vendor’s detector output rather than an audited count. Note also which way that cuts here: the numbers flatter this essay’s thesis, which is when a figure deserves more scrutiny rather than less. What the objection touches least is the pair of second-order findings, since those are comparisons made within one detector’s own output, where a consistent error rate largely cancels. The absolute prevalence is the part to hold loosely.

The LessWrong policy (March 2025, since revised) requires that LLM-assisted submissions add significant value beyond the model’s output and that a human vouch for every claim; first-time posters may not use AI-generated text at all; moderators reject LLM-looking content by default and report several such submissions arriving daily. Same immune response as arXiv, at forum scale: raise the price of entry, and spend scarce moderator attention patrolling the border.

The bycatch (3). The TARA retrospective (September 2025) is a useful specimen because it is explicitly a middle-of-funnel program: it exists to bridge the gap between awareness-level courses and selective fellowships. Its own framing cites BlueDot’s plan to train 100,000 people in AI safety fundamentals within about five years against MATS’s roughly 450 scholars across three and a half (the funnel’s two ends, more than two orders of magnitude apart), and TARA itself hopes to scale to over a thousand trained practitioners a year. To its credit, the report names its own weak point: insufficient support and connection after the program ends. Widening programs measure success at entry. Nobody owns the exit, and the exit is where bycatch is created.

The MATS talent-needs survey (March 2026; 23 interviews with hiring managers, research leads, and funders) is the field’s own diagnosis, and three things stand out when it is read against this essay. First, it concedes the evaluation bottleneck outright: organizations are constrained not by funding or by applicants but by senior researchers able to supervise, each of whom can mentor only a few juniors before quality degrades: the asymmetry of (1), restated as a hiring fact. Second, its guidance doubles down on the trust network. Because organizations cannot afford hiring mistakes, the dominant hiring signal becomes direct collaboration or a calibrated endorsement from an already-known researcher; credentials, interviews, and publications count for little, and fellowships function as extended vetting periods. Under scarce supervision this is individually rational. Collectively, it is the correlation machine of (5): a small set of mentors’ judgment becomes the field’s admission function, replicated into each cohort, a choice whose aggregate cost nobody can price, because the diagnostics that would price it do not exist. High-risk portfolio decisions are easiest to make, individually and rationally, precisely when no one can see the portfolio. Third, it recommends cultivating a broad strategic perspective (which threat models are plausible, which agendas address them, what assumptions each rests on) as an individual skill. That is a demand for precisely the infrastructure of (6), addressed to individuals, in a field with no commons to supply it. Each researcher is asked to privately reconstruct a strategic picture that no institution maintains.

The template (6). Forethought’s AI for AI for Epistemics is the closest existing thing to a design document for the commons version. Its argument: as R&D automates, strong epistemic tools become buildable by whoever spends the compute, and the binding problems shift from capability to trust, ground truth, and incumbency. Its risk register names the two ways a build like (6) fails: what it calls epistemic misalignment is the Goodhart worry, people optimizing to look legible on a canonical map, and what it calls trust lock-in is the worry that one such map becomes unquestionable. Its interventions (open and auditable infrastructure, development in incentive-compatible hands, anticipating future data needs) are the general form of what (6) states for this field in particular. What the template leaves open is the field-specific measurement layer: concentration indices, funder overlap, the provenance of convergence. The generic program says how to build a map people can trust. It does not say what this field’s map has to show.

The objection (6). Naming Forethought’s two failure modes and not applying them to this proposal would be a dodge, so: both apply to what (6) describes, not only to the general program. Epistemic misalignment is the Goodhart worry: once a map is what funders and hiring managers consult, work starts being shaped to read well on it. Trust lock-in is the worry that one map becomes the map, consulted rather than argued with. And (5) makes both worse here rather than better. A small, interconnected field is the easiest place for a single account of who is working on what to become load-bearing: few enough readers that they end up consulting the same one, correlated enough that they read it the same way. There are fewer maps, each carries more weight, and the same handful of buyers consult all of them. The scale that makes such a map cheap to build is the scale that makes it easy to make canonical.

The target, though, is mostly not the topic, and that half of the problem is one I have already written about at length. The Loop No One Chose is about this dynamic in a decentralized field, and the argument transfers to the map directly. The obvious way to play to a map is to pick whatever it marks as neglected, which is at least visible and roughly self-limiting: if enough people move, the gap closes and the map says so. The harder version runs on form. Whatever shape of work the measurement layer can see (an output with a venue, an agenda it belongs to, a unit a funder can price) is the shape work drifts toward, and work that fits no column does not appear as missing, because nothing is counting it. That is the failure (7) expects from building nothing, research bending toward whatever is cheapest to evaluate, arriving instead through the instrument meant to prevent it. It is reflexive in the way the rest of this essay is: the map reports where attention sits, work adjusts to be reportable, and the next reading is more confident for being backed by numbers. A map of the field’s attention ends up as one more thing competing for it.

I do not think this is disqualifying, but the answer is partial rather than clean, and anyone pitching this should be able to give it. The mitigations are the template’s own, made specific: keep the underlying records public and traceable, so anyone can re-slice them and get a different answer; publish what the measurement cannot see as carefully as what it can; keep a competing version cheap enough to build that the first one can be disagreed with rather than only trusted; and point the concentration measures of (5) at the map’s own readership. None of that buys immunity. A map the field consults is a map the field will optimize against, and the case for building one rests on that cost being smaller than the present arrangement, where the same bending happens with nothing in place to notice it.

The fork (6). Those mitigations are the defensive half. The structural answer is that the map’s authority should be impossible to hold in the first place, which is where the two companion essays go. The Loop No One Chose reaches the precaution from the reflexivity side: not a better mirror but more of them, several versions sharing a substrate so they can disagree about what it means, because versions that disagree cannot fail in unison and make a poorer target besides. How to Build a Mirror Without Holding It works out the mechanics. The load-bearing distinction there is between gating a merge and gating a fork. Gating merges is ordinary quality control: the checks keep the trunk faithful to its sources. Gating the right to fork hands the center a veto over dissent, and is the one move that would turn this design into the thing it was built to prevent. A fork inherits its parent’s checks and is free to rewrite them, relaxing one and adding another, and because forking is granular it can fork a single subtree rather than splitting the map. The diff between two check suites is then the most precise form a disagreement in this field can take: two versions of a discipline arguing over what counts as a valid claim. No map is the map. The field decides which one is worth using by using it, and forks when it does not agree, which is how open source has handled the same problem for thirty years.

The same essays answer the other objection to (6), that staffing this with the residual cohort of (3) would have the field mapped by its least experienced members. The design splits drafting from endorsement. Candidate entries are generated from the open record, arXiv, OpenReview, grant databases, and drafted by people who have the time to do it; the named researcher then confirms or corrects their own entry, and that confirmation, not the drafter’s judgment, is what the map rests on. The cohort supplies the labor the field will not otherwise pay for, the researcher supplies the authority, and the drafter accumulates something the field currently has no way to see: a public record of judgment, which is the quality (3) noted its filters can only assess in private.

The ledger (8). Elena Ericheva’s genealogy of AI safety (July 2026) reconstructs the field from 323 documented events, 129 actors, and 18 directions, 2005 through June 2026, each event tied to a primary source: arXiv preprints, lab reports, tax filings, regulatory pages, fund announcements, assembled in several passes over live sources rather than from retrospectives, part scripted and part backfilled by hand. Field-building is the largest itemized money flow in it, $326.1 million. Two of its findings bear on this essay directly, and both are about attention and money failing to track each other: interpretability’s scientific attention runs roughly 37 times ahead of what its approximately $1 million in grants would predict, while governance funding went from $0.4 million in 2018 to $18.4 million in 2023, a 46-fold jump in five years. The caveats are as informative as the figures. Recent years are marked as under-collected, arXiv counts are said to measure how much was recorded rather than how much happened, and the different kinds of money (grants, government budgets, venture equity, pledges) are kept in separate views and never summed. That care is the reason to trust the numbers and also the reason to notice what producing them cost. Nothing in it is wrong. The problem is only that a field this size should not have to wait on one person’s manual pass over primary sources to learn where its money and its attention went, and that by the time such a pass is finished it describes a field that has already moved.

The debt (8). Chris Olah and Shan Carter’s Research Debt (Distill, March 2017) defines the condition as “the accumulation of missing interpretive labor” and separates it into four kinds: poor exposition, undigested ideas, bad abstractions and notation, and the noise of more papers than anyone can read. The mechanism is a tradeoff the essay states plainly, between the energy put into explaining an idea and the energy needed to understand it: one explainer’s fixed cost against every future reader’s, which is why distillation pays off at scale and why almost nobody supplies it. Distillers, it notes, lack “a career path, places to learn, examples and role models,” because the work is not seen as a real research contribution. Two things follow here. The first is the date. 2017 is well before the volume shock, so cheap generation did not create research debt; it compounded a deficit the field already carried. The second is that (6) asks for funding for precisely the role Distill described as structurally unrewarded, and (3) names the first group with both the skills and the reason to take it.

LLM policy

The appendix was written entirely by a language model based on my ideas, but I have verified it carefully: every claim in it checked against the source it cites. The essay itself was LLM-assisted but edited word by word by me, so if you do not find it tasteful, you can confidently blame my literary faculties and sensibilities.

About the author

Manu Xaviour Thaisseril Shaju

I came to AI safety from customer service: several years as a technical support engineer, then a master’s in AI in 2022. My first piece of research was Cross-Axis Capping, an experiment on what happens when jailbreak detection and correction stop sharing a single direction.

Since then I have mostly been building the infrastructure I wished existed. I wanted a map of where the field is actually heading, and a parsable one: something I could slice by whatever I was trying to find out, so that a different question returns a different view of the same field. That is AI Safety Map, which came out of The Loop No One Chose and the two pieces that followed it, How to Build a Mirror Without Holding It on how such a map gets built and Open in Principle, Blind in Practice on why a funder needs one. I also wanted help with the papers themselves, the prerequisite-heavy ones written with almost no exposition, which is Explore, the research explorer that came out of Paying Down Research Debt.

At some point these stopped looking like two separate annoyances. They are symptoms of the same systemic problem, so I did my own digging and decided to write about it. I have not heard anyone discussing it in these terms, and I think it is worth discussing.

Two caveats I would rather state than hide. This essay may itself be part of the slop it warns about; if it is, it is at least a specimen that makes the point. And I do not know whether grantmakers hold a picture of the field that I cannot see. I think that is unlikely, since I have not found the discourse that would suggest it, but I may be wrong.

Read next →

The Loop No One Chose

Reflexive dynamics in decentralized nonprofit fields, and why AI safety needs a live, actionable map that reflects the field back to itself instead of telling it what to do.

Related →

Paying Down Research Debt

A tool for reading AI safety research, and the living canon it leaves behind: the undersupplied distillation layer this essay leans on in (8).