← Back
SiTech Team⏱️ 3 წთ. საკითხავი

AI Hiring Bias: Study Shows LLMs Are More Likely to Form Biases Than Humans When Recruiting

AI Hiring Bias: Study Shows LLMs Are More Likely to Form Biases Than Humans When Recruiting

Princeton and UChicago researchers found that AI models exhibit 65% more segregation in candidate selection than humans — with OpenAI o3 scoring 1.83 on the segregation scale, approaching the theoretical maximum of 2.

A groundbreaking study from Princeton University and the University of Chicago reveals that large language models (LLMs) can develop more severe biases than humans when evaluating job candidates. The study, published at ICML in Seoul, sends a clear warning about the deployment of AI in human resources.

The Experiment: A Simulated Hiring Game

Researchers ran LLMs — including ChatGPT, Claude, and Gemini — through a simulated hiring game adapted from psychology. Each model was told it had been hired as a consultant for a fictional city and asked to help hire people for 20 jobs across different professions. Candidates came from four fictional ethnic groups: Tufa, Aima, Reku, and Weki.

Unbeknownst to the models, all candidates were equally likely to succeed at every job. Yet the models quickly started segregating candidates from different groups into different jobs based on early observations.

Shocking Results: AI 65% More Biased Than Humans

On the study's segregation scale (where 2 means complete segregation), human participants scored 0.84. The AI models scored roughly 65% higher. OpenAI's o3 scored 1.83 — close to the theoretical maximum of 2.

LLMs are "eager to create generalizations from limited data," says Ryan Liu, PhD student at Princeton and coauthor. "That's literally a lot of what they're optimized for."

The Exploration-Exploitation Dilemma

Every decision-maker faces a trade-off between sticking with what worked and trying something new — the exploration-exploitation dilemma. Because LLMs are trained on problems that reward generalization from few examples, they settle on hunches too early. The same instinct that helps them solve logic puzzles also makes them quick to stereotype.

Newer Models Are Worse

Newer reasoning models (o3, DeepSeek R1) showed even stronger biases. Higher reasoning capabilities didn't help — they made the bias worse. When LLMs rush to generalize in social settings, "that's when things tend to go wrong," says Liu.

The Fix: Incentivizing Diversity

Telling the model to be fair didn't help much. But promising the models a bonus for diverse hiring made them far less biased. The key insight: design goals that incorporate desirable social values.

Memory Makes It Worse

As chatbots gain memory and personalization features, they can "over-index on the same kinds of behaviors it's experienced before" and form deeper biases. Simply having chatbots remember less isn't a fix, because users want them to remember.

Implications for Georgian Businesses

For Georgian companies using AI in HR, this study raises critical questions. AI hiring tools must be tested for bias before deployment. The solution isn't to abandon AI but to design incentive structures that explicitly reward fairness and diversity.

Conclusion

This Princeton study is a wake-up call. AI models don't just inherit human biases — they invent new ones. The same optimization that makes them brilliant at logic makes them dangerous at social judgment.

📖 Source