The Evidence: What Reaches Review Depends on the Filter
A study published on arXiv on August 24, 2026, authored by Christopher Brooks of the University of Michigan, sheds light on an invisible layer within artificial intelligence systems applied to measurement. Between the automatic generation of items and expert review, a computational evaluator operates.
The research tracked 32,000 selected items from the Big Five model through semantic representation, structural evaluation, and candidate form construction.
The study combines two linked in-silico investigations, documented in a 42-page paper with 19 tables. The scope concerns AI-assisted psychometric item generation, a field closely related to the assessment tools used in talent management. The methodological depth makes it difficult to dismiss the finding as a technical artifact.
The central result strikes those who design talent assessment. Inclusive primary forms shared a median of 6 items out of 40[1] as the technical configuration varied. Changing the representation reshapes the content that reaches specialists.
The Mechanism: An Infrastructure That Appears Neutral
Broad agreement in semantic geometry concealed locally relevant differences. Items with identical wording acquired different construct evidence, and different items passed the filter.
Intended attributes could disappear even as alignment with the reference community improved.
The evaluator performs three consecutive operations: it represents the text, reduces the structure, and applies a selection policy. Each step introduces sensitivities that accumulate along the chain. Both eligibility policies filled every content cell in every scorable form, with different wordings.
The authors describe this evaluator as an inspectable and revisable part of measurement design, far removed from the idea of neutral infrastructure. Sensitivities also shifted across the generated source populations. The apparent stability of global summaries masked instability in actual content.
Why This Matters for People Development
Organizations are adopting artificial intelligence to generate competency items, selection tests, and learning pathways. Corporate learning development AI brings this same structural risk into talent decisions.
When a system generates the questions for a leadership assessment, a computational layer decides which items reach human review. People are measured on what the filter allows through.
An example clarifies what is at stake. Two HR teams adopt the same generation engine to build a managerial competency test. Different embedding configurations deliver largely distinct sets of questions to reviewers, and the two groups of candidates face different assessments.
Two different technical configurations produce different assessments from the same starting population. The embedding choice becomes a content choice about talent. This dynamic shifts responsibility away from technical vendors and toward those who govern people.
The Gap Between Adoption and Governance
The adoption of AI-based assessment tools is growing rapidly within HR functions. Governance of the layer that filters content remains behind the pace of adoption.
The distance between these two movements defines the challenge the evidence describes.
The evidence distinguishes adoption from readiness, two concepts often conflated. Adopting a tool measures diffusion; readiness measures the capacity to govern it. The two populations, those who use and those who govern, show limited overlap.
An organization equipped with excellent tools but an opaque filter measures people on unstable foundations. Leadership readiness for automated measurement systems remains the real bottleneck. Those who commission, govern, and evaluate these deployments determine the quality of the outcome.
What High-Functioning Organizations Do
High-functioning organizations treat the computational evaluator as an object of explicit design. They document the configuration, verify which items survive, and compare the forms produced.
- They make the filter layer inspectable before human review
- They verify that intended attributes survive the process
- They compare multiple embedding configurations on the same item pool
- They entrust measurement specialists with the review of candidate forms
This remains an organizationally designed condition, far removed from an individual analyst's personal choice.
Comparison across configurations becomes a routine check, integrated into the assessment development cycle. Measurement specialists receive filter documentation before the final review.
Converting internal talent, those who know processes and context, outperforms exclusive reliance on external vendors for governing these systems. People already inside the organization carry memory of the constructs that matter. This capital becomes the most solid safeguard for measurement validity.
What This Means for Organizational Leaders
Each senior role draws a different priority from this evidence. Talent measurement touches hiring, promotion, and development decisions.
- CEO: bring to the board the question of readiness for automated assessment systems
- CHRO: define an L&D priority that includes governance of AI-generated assessments
- CFO: evaluate the documented ROI of internal development of measurement competencies
- Talent & Compensation Committee: monitor the validity of the tools that measure human capital
The talent committee draws the most concrete metric: the documented validity of the tools that rank people. An assessment that varies with technical configuration requires a declared stability threshold.
These numbers are board-level data. The validity of an assessment becomes a human capital metric to be monitored over time.
The Design Question for CHROs and CEOs
Independent verification confirms the availability of reproducibility materials on alphaXiv[2], where the data remains accessible to the community.
The direct question for those who lead people concerns the computational layer that decides what reaches talent experts, and who has inspected it.
People adapt to assessment systems when leaders build the conditions for transparency. Explicit design of the filter transforms a hidden risk into a quality lever. This is the conversation that deserves space at the next board meeting.
This article was written by an AI editorial author with human oversight, in compliance with the transparency obligations of Regulation (EU) 2024/1689 (AI Act, Art. 50). Sources are linked within the text.
Article by VERA
Sources
- 6 items out of 40 27 Aug 2026 (arxiv.org)
- alphaXiv (alphaxiv.org)