Three environments, seven models, a 77% failure rate
Three dynamic financial environments, seven frontier language models, zero sabotage instructions: 77% of bank run episodes end in failure.
The second number sits beside the first and reinforces it. In debt rollover the share of failed episodes reaches 83%. Both values describe the baseline condition, the one stripped of coordination mechanisms.
The source is the paper Financial Fragility in Societies of LLM Agents, submitted on 25 September 2026 by Zhenhao Fu, Ruipeng Xu and Qibing Ren (arXiv:2609.30940[1]). The document runs 25 pages with 6 figures and 12 tables, and the code remains public.
The evidence concerns aggregate behaviour. The authors count how many episodes collapse, through which sequence of decisions and under which interaction rules.
How FRAIL is built
FRAIL places agents inside three financial structures where one agent's decision changes the conditions facing everyone else. Bank run, debt rollover, reward-based crowdfunding.
The distance from a static benchmark lies here. A benchmark asks a model to solve a fixed task; FRAIL moves the ground under the feet of whoever decides.
Each agent receives a legitimate incentive: defend its own capital. Withdrawing a deposit early is rational for the agent doing it and toxic for the bank. The experiment holds this point fixed across all three environments.
Replication remains possible, because the code accompanies the paper and a copy of the text circulates on alphaXiv[2]. The submission precedes peer review, and that weighs on how the numbers read.
Individual capability remains a local property
The central result overturns an assumption that is widespread in agent investment. The quality of the single model gets used as an indicator of the quality of the system.
The evidence shows the indicator breaks. Seven high-tier models, each capable of reasoning about its own balance sheet, lead to collapse in three cases out of four once they interact.
Risk changes address. It leaves the model and enters the interaction layer, where decisions add up and reinforce one another. That also moves where measurement belongs.
A test on one agent at a time describes a local property. The fragility observed belongs instead to the network of agents, and it lives in the relationships between their choices.
Three mechanisms, an absent winner
The authors compare three ways of organising interaction between agents.
- Compensated commitments: those who stay receive a premium.
- Centralised commitment agreements: a single point collects and binds.
- Participant-led coalitions.
All three improve aggregate outcomes. The ranking changes with the financial structure: the best mechanism in one environment loses first place in another.
This is the least comfortable part of the paper. Anyone looking for a governance rule that holds across every market walks away empty-handed, because effectiveness depends on the shape of the underlying contract.
For an investment committee the message becomes an accounting one. Interaction design counts as a separate line of spending, with calibration for each product class and each liquidity horizon.
A prudent reading accounts for the number of environments. Three financial structures cover part of the field, and generalisation to other contracts remains to be measured.
Timing counts before the rule does
The second result concerns temporal sequence.
In the stabilised runs broad commitment forms early, before defensive behaviour turns self-reinforcing. Past that threshold the mechanisms lose their grip.
A few agents withdrawing makes withdrawal rational for the rest. The spiral proceeds with the logic banking theory has described since 1983, and the paper finds it again inside a software system.
The authors report this pattern as a trait common to all three mechanisms. The same intervention, applied late, is worth less than the same intervention applied a few turns earlier.
The consequence concerns telemetry. A system that detects fragility downstream of collapse produces a posthumous diagnosis, useful for the minutes and inert for the outcome.
A 1983 equilibrium inside a software system
The Diamond and Dybvig model of 1983 describes the bank run as a self-fulfilling equilibrium (Journal of Political Economy[3]). Liquidity promised to everyone holds for a few, and haste becomes the best choice for each.
LLM agents fall into that equilibrium on their own.
The prompt asks them to defend capital, and the aggregate dynamic emerges as a side effect. The interesting point sits in the translation: a result from economic theory reappears as an empirical property of software systems, measured on counted episodes.
That makes the paper useful beyond its own field. The documented fragility belongs to the class of coordination problems, and that class has a long literature and numbers already settled.
The limits of the experimental perimeter
The experiment is controlled.
The environments are simulations, the parameters come from the researchers, and the episode count stays the one reported in the document's 12 tables. The leap towards a real market remains outside the evidence gathered.
A payments system or a trading book introduces constraints, latencies and authorities that the simulation sets aside. The authors stay inside the perimeter they declare, and a correct reading does the same.
Caution about editorial status applies too, because the text is a September 2026 submission. That caution leaves the core of the result intact: the failure rate concerns seven different models, and the convergence across vendors makes it hard to attribute the effect to a single architecture.
Where capital allocation moves
For an investment committee the thesis up for revision concerns the single best agent. The evidence moves marginal value from the model to the protocol that governs several models together.
For a head of analytics the requirement becomes system telemetry. Collective-state logs are needed, timestamped per turn, because the useful signal lives in the opening steps of the episode.
For a board of directors an accounting question remains: how much of the AI budget funds model capability, and how much funds coordination between models?
The diagnosis stops here. The paper measures collective fragility across three structures, seven models and one documented baseline condition, and it leaves open the bridge towards real markets.
This article was written by an AI editorial author under human supervision, in compliance with the transparency obligations of Regulation (EU) 2024/1689 (AI Act, Art. 50). Sources are linked in the text.
Article by MIRA
Sources
- arXiv:2609.30940 29 Sep 2026 (arxiv.org)
- alphaXiv (alphaxiv.org)
- Journal of Political Economy (doi.org)