Essay
The Loop No One Chose
Reflexive dynamics in decentralized nonprofit fields, and why AI safety needs a live, actionable map
Decentralization is supposed to be a safeguard. Spread decisions across many independent actors, strip out the profit motive, and you are meant to be protected from the pathologies of centralized power: the captured agenda, the controlled narrative, the single point of failure. A field like technical AI safety (dozens of small nonprofits, scattered independent researchers, a handful of mission-driven funders) looks like exactly that safeguard in action. No one is in charge, so no one can steer it wrong.
This is true with one caveat, and the caveat is the danger. Decentralized, good-faith efforts are not immune to the dynamics we associate with central control. They are merely immune to the visible versions. The narrative still gets steered, resources still concentrate, the field still bends away from its most important work, but no one decides any of it, which is precisely what makes it so hard to see and so hard to fix.
The mechanism: attention as a proxy for importance
The dynamic has gone by many names in different fields: the Matthew effect (Robert K. Merton, in the sociology of science), reflexivity (George Soros, in financial markets), winner-take-all (Robert H. Frank and Philip J. Cook, in economics), the same terrain Nassim Taleb later mapped as “Extremistan.” Soros’s version is the sharpest about the mechanism. He used the word reflexivity for feedback loops in which the act of observing a system changes the system, which then changes what observers see. Markets are his classic example: a belief that an asset will rise can drive the buying that makes it rise, confirming the belief. The loop is self-fulfilling, and it runs on perception rather than fundamentals.
Fields of research and effort run their own version. The currency is not price but attention, and the loop turns like this. In a decentralized field with no central allocator, no one hands out priorities. Each actor has to decide for themselves what is worth working on or funding, and they need a signal. The cheapest available signal is what everyone else is already doing: what is being published, discussed, hired for, talked about. Attention becomes the proxy for importance, because activity is the most legible evidence of where value seems to be, unless there is a better way.
So talent and money follow attention. Output then skews toward the directions that already have the people and the funding. And that skew reads, to the next person scanning for a signal, as proof that this is where the action is, which pulls in the next wave of talent and money. Funding sits unevenly across the field not in spite of this dynamic but in large part because of it. Hot areas compound; quiet ones starve, regardless of how important they are.
Why decentralization hides the loop rather than preventing it
Here is the part that should unsettle anyone who treats decentralization as protective. The loop above requires no coordinator, no bad actor, and no concentration of formal power. It is the emergent result of many independent agents each making a locally reasonable choice with the cheapest signal available to them. That is what makes it invisible.
When a central body misallocates resources, there is a decision to criticize, a policy to change, a person to hold accountable. When a decentralized field misallocates, there is none of that. There is only the aggregate of free choices, which looks like the market working as intended. No one chose the skew, so there is no choice to contest. The distortion wears the costume of liberty very convincingly.
Good faith does not break the loop. It camouflages it.
The nonprofit, mission-driven framing adds a final twist that makes it worse, not better. Everyone in such a field believes they are optimizing for importance: for impact, for the mission, for the neglected problem. But the proxy they lean on in practice is attention, and attention diverges systematically from importance. It favors the recent over the enduring, the legible over the subtle, the well-connected over the isolated, and the already-crowded over the genuinely neglected. So a field full of people sincerely trying to work on what matters can be pulled, in aggregate, toward what is merely visible, while each participant feels they are being rigorously impact-driven. And all of this happens out of sight: each participant sees their own choices, which are reasonable, and never the skewed aggregate. In their defense, the blindness is structural, not willful. But so is the helplessness. Awareness is a remedy only when it has somewhere to go (a decision to contest, a policy to reverse, a person to confront), and here, again, there is none of that.
The fixer’s trap
Once you see the loop, the obvious response is to correct it directly: build a central map, a prioritization, a recommendation engine that tells the field where the real gaps are and points people toward the neglected work. Surely better guidance beats a blind proxy.
So the instinct to fix the loop by issuing better instructions makes you the loop’s new engine.
This is a trap, and it is worth being precise about why. A single, authoritative recommender does not stand outside the loop; it becomes a concentrated, fast version of it, bolted on top of the subtle one already turning at a lower pace. The existing loop runs through many channels and corrects slowly; an authoritative source that says “work on this” is a sharp, centralized forcing function that is harder to push back on precisely because it carries authority. And it is reflexive in the most direct way: it learns from the behavior it shapes. It recommends a direction, the field moves, the movement registers as signal, the recommendation is reinforced. The cure reintroduces exactly the centralized narrative-control that the field’s decentralized structure was supposed to rule out, now wearing the badge of refinement.
The resolution: a mirror with no one holding it
If recommendation is the wrong move and the status quo is a slow distortion, what is left? The answer turns on a distinction that is easy to miss: the difference between prescribing and describing.
The existing loop runs on attention-as-proxy because attention is the only cheap signal of where value lies. That proxy is exactly what is broken. And what stands the best chance of weakening a broken proxy is the visibility of the distortion itself rather than a louder counter-signal. A faithful, current description of the field (one that shows plainly where attention and money have piled up, and which important directions sit starved beside them) hands every independent actor the information to correct on their own judgment. It does not tell anyone what to do. It makes the imbalance legible enough that people can choose to push against it.
This looks like it walks back into the fixer’s trap, until you notice that, in the limit, there is no observer here at all. The map is not a separate party looking at the field and reporting back; it is the field’s own output, its publications and funding and citations, folded back into a form the field can read, a common substrate. No one stands outside holding it up, at least in principle; the system is at once observer and observed, or as close to that as anything real gets. A recommender amplifies the proxy; this exposes it, using nothing but the field’s own activity reflected to itself. The line that matters is between reflexivity that returns judgment to the reader (here is the skew, decide for yourself) and reflexivity that acts on behalf of the field (work on this, as it managed to get everyone else’s attention). The first disperses the power to choose across thousands of minds; the second does not give a choice.
None of this escapes the loop, and honesty requires saying so. The field already runs it, just blind; self-observation does not remove the loop but gives it a readout, a kind of proprioception it did not have before. That also relocates the discipline: with no external lens there is nothing to capture, and the only failure is distortion, and in all honesty, pure self-observation stays an asymptote rather than a claim. And if the reflection ever grew authoritative enough that showing the gaps began to function as funding them, that influence becomes something to watch and disclose, not assume away.
The map as a target, and the case for several
There is a second cost here that is not distortion at all, and a mirror can be perfectly faithful and still pay it. Once a description of the field is the thing people consult, work starts being shaped to read well on it. Forethought calls the general version epistemic misalignment, and the pressure it describes lands mostly on form rather than topic. Chasing whatever the map marks as neglected is roughly self-limiting: enough people go, the gap closes, the map says so. The harder version is that whatever shape of work the reflection can see (an output with a venue, an agenda it belongs to, a unit a funder can price) becomes the shape work takes, while work that fits no column never registers as absent, because nothing is counting it. Which is this essay’s own loop, one level up: the reflection reports where attention sits, work adjusts to be reportable, and the next reading is more confident for being backed by numbers.
That objection assumes the thing being consulted is a map. Underneath, it is not. It is a queryable substrate: the field’s own record of itself, held in a form anyone can interrogate. A map is one query over that substrate, frozen and published, and anyone willing to spend a little effort can ask their own instead (which agendas absorbed the last year of funding, what is being cited by whom and how recently, where the hiring went) and get an answer traceable to the same sources. So a bad map is not a durable object here. A slicing that flatters someone’s priors survives about as long as it takes one reader to check it. The classic failure of a single authoritative description, that it is wrong and no one can tell, has little room to operate against a substrate that answers questions directly.
What survives is not a wrong map but a single one. If one view is the view everybody reads, people write for it: work that fits its columns gets done and framed accordingly, work that fits nothing gets quietly dropped or never proposed. That is Goodhart in its familiar form, and it is a property of there being one target, not of there being a target at all. A field small enough to share one map is the easiest place for it to take hold, not the hardest: few enough readers that they end up consulting the same one, correlated enough that they read it the same way. The scale that makes the mirror cheap to build is the scale that makes it easy to make canonical.
So the precaution is not a better map but more of them, published side by side, drawn from the same substrate so that they disagree about what it means rather than about what happened. Versions that disagree cannot fail in unison. But the sharper point is what plurality does to the Goodharting itself. Against one target, work bending toward the target concentrates the field, which is the original disease. Against several that genuinely disagree about which directions are starved, bending toward any one of them is just picking an orientation and doing the work that orientation says is missing, while whoever reads the map beside it bends somewhere else. There is no single column to crowd into when the slicing is itself contested. At that point the optimizing stops being the failure and becomes the mechanism: it is how a reader turns a view of the field into work, and the plurality is what keeps the results from landing in the same arc. None of the maps will be true enough, and none needs to be. Their slants show against each other, the field’s flaw was never distortion but distortion it could not see, and with no settled image to obey and at last something to contest, judgment disperses again across thousands of minds.
What AI safety concretely gains from a live, actionable map
The abstract argument lands somewhere specific for technical AI safety, a field that runs the dynamics above. Here is what a live, source-traceable map of the field’s own internal state (its funding flows, its talent flows, its research output) actually buys.
- Neglectedness stops being blind guesswork. Grantmaking and career choice both depend on knowing what is underfunded relative to its importance. Today that judgment is made from stale, hand-assembled snapshots that are out of date even before they are posted, or are not relevant in the short time frame they are operating in. A live map means the field’s allocators are not reasoning blind about their own distribution.
- The herding becomes correctable. By showing where attention and money concentrate against where live problems sit unfunded, the map lets the field push back on its own loop, without a central allocator deciding anything. The correction comes from many independent readers, not from a controller.
- “What exists” becomes “what’s actionable now.” The field already has genuinely good descriptive maps. Shallow Review (shallowreview.ai) lays out the field’s problems and agendas in a beautifully organized taxonomy, and AISafety.info gives an engaging, accessible tour of the open questions. Both are excellent at what they do, and both share the same two limits: they are mostly static, and they tell you what exists rather than what is blocking each direction right now. A map organized around the current bottleneck in each line of work, and around what recent research addresses and what new directions it enables, turns a catalogue into a moving picture of where each thread is stuck and what just empowered it. That is the difference between a library and a dashboard.
- The newest work stays visible. Attention and citation counts systematically bury the most recent papers, because citations take years to accumulate, while overweighting the established canon. Reporting engagement by life-stage (frontier, consolidating, canonical) keeps the live edge of the field in view instead of hiding it behind a backward-looking metric.
- The field gets a shared, contestable substrate. When both sides of a disagreement about where the field should go can cite the same traceable starting point, and disagree openly about how it is sliced because the underlying sources are public, the quality of the field’s debate with itself rises. Argument improves when everyone is arguing over the same legible facts.
The general principle
The deepest mistake in responding to a hidden feedback loop in a decentralized field is to reach for a controller. The loop’s invisibility tempts you to install an authority that can finally direct things properly, and that authority becomes the loop’s most efficient engine, now with a halo of objectivity.
The better response is almost anticlimactic. You do not need someone to steer the field. You need the field to be able to see itself clearly enough to steer on its own. Against an invisible loop, the right instrument is not an allocator. It is the field’s own reflection, kept faithful and kept current, and watched closely for the day the reflection starts giving orders.
Appendix: sources and notes
The essay borrows several named concepts from other fields and compresses each into a clause. This appendix records what each source actually argued, and where the transfer to a research field holds or bends.
Reflexivity (Soros). From The Alchemy of Finance (1987), restated in “Fallibility, Reflexivity, and the Human Uncertainty Principle” (2014). The structure is two functions running in opposite directions: a cognitive function, the attempt to understand the world one lives in, and a manipulative function, the attempt to make an impact on that world and advance one’s own interests. Soros’s point is what the interference between them does to truth: with the manipulative function running, “the facts no longer serve as an independent criterion because the statement may be the product of the manipulative function.” The transfer to a research field is close but not exact, and the difference matters. In markets the manipulative function runs through price, and price eventually settles against something; a position is scored, and the belief that produced it is scored with it. Attention has no settlement. Nothing marks a direction to market, so the loop described here can run considerably longer than a market one before anything contradicts it.
The Matthew effect (Merton). “The Matthew Effect in Science” (Science, 5 January 1968, 159:3810, 56–63). The paper is usually remembered for its claim about reward: eminent scientists collect disproportionate credit for collaborative work and for discoveries made independently by several people. That is not the half this essay uses. Merton takes a second pass on the communication system, observing that the effect raises the visibility of contributions from scientists of established standing while lowering the visibility of contributions from the less well known. The claim here is not that famous researchers are over-credited but that visibility is itself doing the allocating, which is Merton’s communication-system version rather than his reward-system one.
Winner-take-all (Frank and Cook). The Winner-Take-All Society (1995): small differences in performance produce enormous differences in reward. The relevant argument is the one about waste rather than the one about inequality. Such markets draw too many contestants, in part because people systematically overestimate their odds of winning, and the result is a misallocation of talent, with able people crowding into a few high-visibility contests while other work goes understaffed. That is this essay’s skew restated as a labor-market fact, and it arrives, as here, without anyone having intended it.
Extremistan (Taleb). The Black Swan (2007). This one names a distribution rather than a mechanism: domains where a single observation can dominate the total, against Mediocristan, where no individual event moves the aggregate much. It earns its place because it describes the shape the loop produces, a few directions holding most of the field’s attention and most holding almost none, but it says nothing about how a field gets there. The terrain is Taleb’s; the mechanism is Soros’s.
Goodhart’s law. Charles Goodhart, at a Sydney conference in 1975: “Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes.” He later called it a throw-away line. The compressed version everyone quotes, that when a measure becomes a target it ceases to be a good measure, is Marilyn Strathern’s, from “‘Improving ratings’: audit in the British University system” (European Review 5:3, 1997, 305–321), an anthropological account of the audit explosion in British higher education. Two details bear directly on the argument. Goodhart’s original wording puts the failure in control: the regularity collapses when the measure is used to steer, which is the prescribe/describe line this essay draws, and the reason a description is not automatically subject to the law. And Strathern’s setting was academic audit, so the law’s sharpest formulation was written about a research field being measured. On this terrain it is not an analogy.
Epistemic misalignment (Forethought). AI for AI for Epistemics. Its risk register names two failure modes, both stated more generally there than this essay uses them. Epistemic misalignment is the worry that powerful tools steer thinking in directions that are not truth-tracking, in ways their users fail to detect, because there is no ground truth to check them against: Goodhart, with the compounding problem that the optimization is invisible from inside it. Trust lock-in, which the piece treats as its central concern, is the worry that people come to trust tools or ecosystems that do not deserve it, and that the trust becomes self-perpetuating because those tools keep recommending themselves; anything widely trusted to adjudicate what is true holds enormous influence and, if the trust is misplaced, is very hard to dislodge. Note that both risks are about a source that is the only one rather than about a source that is wrong, which is why the answer offered above is plurality rather than accuracy, and why the piece’s own prescription is that such infrastructure be open and auditable.
LLM policy
The appendix was written entirely by a language model, and then verified carefully: every claim in it checked against the source it cites. The essay itself was LLM-assisted and edited word by word, so if you do not find it tasteful, you can confidently blame my literary faculties and sensibilities.
Read next →
Open in Principle, Blind in Practice
Why not just fund the neglected directions? A funder willing to back overlooked work still needs the map, and without it reproduces the very loop they mean to escape, only running it in reverse.