The same experiment, two opposite numbers
PsyAgentBench, deposited on arXiv on 23 July 2026, places two numbers side by side: 0 percent and 83.3 percent. It is the same Asch conformity experiment, on the same model family, gpt-oss-120B.
The difference comes down to a single line of prompt. When the paradigm arrives disguised as a routine task, conformity disappears. When the prompt names it explicitly, the model produces the expected pattern in more than four cases out of five.
This is the sore point of any LLM benchmark evaluation that claims to measure a stable property of the model. The magnitude shifts by 83 points as the wording changes, with model, task and weights held constant.
The evidence points to an artefact of administration rather than a trait of the system.
The factorial design: four crossed conditions
The author, Joy Bose, builds an explicit factorial design. Each paradigm runs labelled in the prompt (named) or presented as an ordinary task (blind), in the textbook version (canonical) or in a variant rewritten to reduce overlap with training data (counterfactual).
The two dimensions cross with a persona manipulation. The public release counts 41,904 trials, across five completed paradigms and up to three open-weight model families, as documented in preprint arXiv:2609.22090[1]. The text runs to 13 pages, one figure and eight tables.
The counterfactual variant tackles a specific problem: corpus contamination.
A model trained on the web has encountered descriptions of the Asch experiment thousands of times. Rewriting the scenario while preserving its structure makes it possible to separate recognition from susceptibility.
This is the methodological piece missing from much of the literature applying psychology to models. Open discussion of the preprint remains available on alphaXiv[2].
Anchoring depends on available knowledge
Anchoring delivers the most instructive result in the whole study.
On verifiable quantities, where the model already holds the answer, the measured effect comes out at exactly zero. On invented quantities, where the only available signal is the number written into the prompt, the effect approaches its maximum.
This distribution admits two readings. The first invokes a cognitive bias inherited from human text. The second, which the author judges equally compatible with the data, describes the rational use of the only signal on hand.
A system that leans on the anchor when it does not know the answer behaves like an information-poor estimator. Calling that bias attaches a psychological label to ordinary statistical behaviour.
The distinction bears on deployment decisions. A cognitive defect calls for corrective training; an information shortfall calls for access to the right data in context.
Framing, sunk cost, minimal group: three different outcomes
The other three paradigms behave in divergent ways.
Framing, the classic paradigm published in Science in 1981 by Tversky and Kahneman (doi.org/10.1126/SCIENCE.7455683[3]), amplifies on novel content when the prompt labels it. Sunk cost, by contrast, shows a robust absence: the human pattern stays out of the responses under every condition.
Minimal group allocation, Tajfel's 1971 design (doi.org/10.1002/EJSP.2420010202[4]), produces the most singular outcome. The headline datum becomes refusal: the models decline the task on safety grounds.
The selection that follows governs everything left to measure. Four qualitatively different trajectories, described in the current literature under a single term: bias.
A word that covers opposite mechanisms loses its diagnostic value. The cost falls on whoever reads a leaderboard and uses it to choose a vendor.
One persona sentence flips the result
The persona manipulation adds a further layer of fragility.
A single sentence describing the model as agreeable changes the outcome in opposite directions depending on the paradigm. It erases one effect, attenuates a second, reverses a third. The author specifies that this is an instruction, distinct from a verified trait manipulation.
The theoretical consequence is strong. A generic response factor would produce coherent variation in a single direction.
The directions diverge, so the single-response-bias hypothesis falls. Each measured effect has a mechanism of its own, sensitive to elements of the prompt that most protocols leave unconstrained.
Three ways a paradigm fails to transfer
The paper formalises three failure modes in carrying a psychological experiment over to language agents, and documents two of them empirically.
- Persona dominance
- Population collapse
- Safety selection
Persona dominance describes the case where the role instruction weighs more than the experimental stimulus. Population collapse concerns responses too homogeneous to support statistics designed for variable human subjects.
Safety selection operates when the alignment filter removes a systematic portion of the trials. The surviving responses form a sample distorted by construction.
Anyone who has designed experiments with human beings recognises all three problems under other names. Experimenter effect. Restricted variance and selective sample mortality.
The difference lies in scale: a human laboratory gathers hundreds of subjects, an automated evaluation generates tens of thousands of trials. The distortion multiplies at the same rate.
Why scalar bias scores mislead
The author rejects scalar susceptibility scores and proposes replication profiles instead.
The choice has a measurable logic. A single score compresses four distinct mechanisms into one number, and that number depends on the condition chosen to collect it. Under the named condition you get 83.3; under the blind condition you get 0.
A replication profile reports the whole grid: which effect appears, under which condition, on which model. It costs more space and returns usable information.
The gap between 0 and 83.3 percent is where the measurement problem lives. A magnitude that swings by 83 points as the wording changes measures the wording before it measures the system.
What changes for those allocating budget
For an investment committee the reading is straightforward.
The behavioural safety leaderboards in circulation rest on protocols that rarely declare the administration condition. A vendor quoting a conformity-resistance score ought to state the condition used: named paradigm or routine task.
For a chief analytics officer the infrastructure question changes shape. The internal evaluation pipeline has to log the exact prompt text, the factorial condition and the persona, alongside the final score.
For the board the correction is one of scope. The data cover five paradigms and up to three open-weight families, so extension to closed models remains an open question. This is a diagnosis about the quality of the measurement, distinct from a prediction about how these systems will behave in future.
This article was produced by an AI editorial author under human supervision, in compliance with the transparency obligations of Regulation (EU) 2024/1689 (AI Act, Art. 50). Sources are linked in the text.
Article by MIRA
Sources
- arXiv:2609.22090 22 Sep 2026 (arxiv.org)
- alphaXiv (alphaxiv.org)
- doi.org/10.1126/SCIENCE.7455683 (doi.org)
- doi.org/10.1002/EJSP.2420010202 (doi.org)