A study from Princeton University and the University of Chicago found that large language models can develop new stereotypes during repeated hiring decision-making, even when the groups have no underlying differences. “In this paper, we argue that removing existing biases is only one aspect of the problem,” the authors wrote. “Like people, LLMs can also invent novel biases that influence human and agent behavior.”
The bias catch
Researchers tested GPT, Claude, Gemini and other models in a simulated hiring task adapted from psychology research. In each of 40 rounds, a model acting as a consultant selected one applicant from one of four fictional demographic groups for a job such as doctor, lawyer, janitor or childcare aide. Every group-job pairing had the same independently sampled 90% probability of success, meaning the model had no real performance differences to detect. Each model and prompt condition was tested across 30 runs.
Despite those equal odds, the models developed unequal patterns, repeatedly putting different fictional groups into different categories of jobs. The authors link the result to an early, random success or failure can lead a decision-maker to repeat apparently successful job pairings rather than test alternatives. The same dynamic has been documented in the human experiments that inspired the task
That finding discounts what most HR technology buyers have been told to expect. Standard fairness benchmarks, the kind vendors cite when they say a model has been tested for bias, generally show newer models performing better over time. This research measured whether a model builds a new stereotype from its own decision history.
Read more: AI works better when HR helps lead it, new research finds

What can be done?
The researchers tested several fixes. Prompting the model to reason step by step barely helped. Neither did increasing randomness in its outputs or shortening how much decision history it had to reference. Removing the game’s point system made no difference, either. What did work, at least in the researchers’ synthetic setting, was giving the model an explicit, measurable diversity objective instead of a general instruction to be fair. Models told to optimize for varied outcomes produced allocations close to random assignment.
When the researchers ran a version of the test where groups genuinely did perform differently, forcing a diversity objective into the mix lowered overall success rates. The intervention only helps when the underlying population truly is equivalent, and confirming that requires a judgment call.
The paper, presented at the International Conference on Machine Learning in 2026, tested a synthetic hiring scenario with invented demographic labels, not a live enterprise system. Its authors are careful to note the mechanism, not the specific numbers, as the transferable finding. Still, as more HR tools move from single-decision scoring toward agentic systems that learn and adapt across a series of choices, the question the study raises deserves a place in procurement conversations.
“This raises a central tension in alignment: How do we limit generalization in sensitive cases without suppressing reasoning as a whole?” wrote the authors. “The challenge ahead is to design interventions that selectively discourage harmful pattern-matching while preserving the constructive forms of abstraction that make LLMs powerful.”
The post AI hiring tools can invent their own bias, research finds appeared first on HR Executive.
This article was originally published on HR Executive. Click below to read the complete article.