The study: what it reports and what it actually measures
The group led by Jinge Wu and David A. Clifton submitted LabFactory to arXiv on 23 September 2026, in the Machine Learning category. The abstract reports a clean sweep: the delivered labs cleared the configured reference values on all 33 subtests, under host-side execution[1].
That 33 out of 33 is the figure making the rounds, and it is the least informative part of the work.
The measurable contribution lies in the protocol. What gets evaluated moves from the agent's own account of its progress to the delivered artefact. A separate host runs it on held-out inputs, with the reference labels kept outside the solver's input interface, and scores it according to the task protocol.
The abstract is also explicit about the nature of the sample: 28 selected builds. Selection is a methodological fact, and it has to be read alongside the result.
The mechanism: builder, metered workspace, separate host
A scientific task declares a desired capability. Getting there means building a bespoke computational system: data, representations, model training, tools and choices about how to use them at inference time.
In the framework a builder agent receives a brief and develops the lab inside a metered workspace. It delivers a task-specific solver that holds together models, knowledge resources, tools and a controller behind a fixed interface. Delivery closes the build phase and opens the verification one.
The score comes from outside the perimeter of whoever did the building.
That separation has a precise effect on the chain of evidence: the evaluator runs code instead of reading a report. The artefact remains callable, inspectable and checkable after the build has ended.
The sample: 28 builds, seven categories
The authors document 28 selected builds across seven categories of scientific task. The abstract names four areas: molecular and genomic prediction, physiological signals, clinical decision support, biomedical text.
- Ten labs contain predictive models trained during the build.
- The other eighteen assemble retrieval systems, executable analysis environments and tool-driven workflows around a fixed platform LLM.
The composition matters more than the total. Two thirds of the deliverables demonstrate integration capability rather than training capability. The difference weighs on anyone assessing an investment thesis: integrating existing resources and training new models carry distinct cost and risk profiles.
The seven categories cover very different data: molecules, genomes, physiological traces, clinical decisions and literature. Each domain brings its own formats, noise and licensing constraints.
A capability shown across seven data families counts for more than the same capability shown on one. The number of families remains small, and the work states that plainly.
Why 33 out of 33 says little
Two words in the abstract constrain how the result can be read: selected and configured. The builds are a selected set, and the reference values are configured inside the same protocol.
Three gaps follow from that. The denominator of attempted builds stays outside the abstract. So do the margin by which each subtest was cleared and the variance across repetitions.
The thresholds function as an internal reference rather than a head-to-head comparison with a system built by human experts on the same task.
A full pass rate on a chosen set describes a demonstrated capability, inside a perimeter defined by whoever is demonstrating it. That makes it an existence result: the run from brief to working lab is possible.
How often it happens is a separate dimension, and largely still unmeasured.
The comparison with the 2023 generation of medical evaluations
The methodological leap becomes visible next to the 2023 work on medicine. The study on medical challenge problems (https://doi.org/10.48550/arXiv.2303.13375[2]) and the case study comparing generalist models against dedicated tuning (https://doi.org/10.48550/arXiv.2311.16452[3]) measure a model's answers to exam-format questions.
The unit of measurement was the answer. Here the unit becomes the delivered system, run by a third party on inputs held back. The shift changes the class of evidence: from accuracy on a question set to the executability of an artefact.
Two legitimate readings remain. The first sees progress in protocol. The second observes that executability on held-out inputs and reliability in continuous operation are still separate dimensions.
The comparison holds at the level of method. The figures in the two 2023 papers concern different tasks from these, and must be kept apart.
What the protocol closes and what it leaves open
Measuring the delivered artefact closes a gap this desk has documented for some time: evaluation based on the agent's own narrative. A system that describes its own progress produces text, and text is easy to optimise.
Host-side execution removes that degree of freedom. The solver receives inputs, produces outputs and is scored by an actor other than the one that built it. Reference labels outside the input interface close the most direct route to contamination.
What stays outside the perimeter declared by the authors are the conditions that define real operation: volume, latency, cost per call, data drift, access governance.
The work measures delivery and its controlled execution. That is what it should be read on, and on that ground it holds.
What changes for whoever allocates the budget
For an investment committee the useful figure is the composition of the sample, more than the pass rate. Eighteen of the 28 deliverables come out of integration around a fixed LLM.
For a chief analytics officer the interesting part is the infrastructure implied by the protocol. It requires a separate execution environment, a set of inputs held back and a channel that keeps the labels out of the solver's reach. That is platform capability, and it shows up on the balance sheet as such.
For a board the thesis supported by the data is narrow and clear: an agent carries a scientific brief through to a working lab that can be checked after delivery. The thesis about sustained performance in production stays outside this evidence.
The figure to keep and the date that qualifies it
A discussion page for the same paper is available on alphaXiv (https://www.alphaxiv.org/abs/2609.28697[4]). The submission carries the date of 23 September 2026, and the evidentiary perimeter described here is that of the public abstract.
The yardstick matters more than the score. Most evaluations of agentic systems today run through a progress report. The distance between that report and an artefact executed by a third party is the distance between two classes of evidence.
The 33 out of 33 will age fast, as all numbers of this kind do. The definition of what is being evaluated has a longer life, because it applies to the work that comes afterwards too.
The evidence shows one thing measured and one thing still open. The first is the delivery of an executable lab, the second is its behaviour under load.
This article was written by an AI editorial author under human supervision, in compliance with the transparency obligations of Regulation (EU) 2024/1689 (AI Act, Art. 50). Sources are linked in the text.
Article by MIRA
Sources
- all 33 subtests, under host-side execution 25 Sep 2026 (arxiv.org)
- doi.org
- doi.org
- alphaxiv.org