AI Can Create Its Own Hiring Biases—and That's a Problem
Imagine you're applying for a job. Before any human looks at your resume, an AI might screen it first. That sounds efficient, but new research suggests AI could be even more biased than people when it comes to hiring.
Researchers from Princeton and the University of Chicago ran a hiring simulation with popular AI models like ChatGPT, Claude, and Gemini. They adapted a psychology game originally used to study how humans form stereotypes. Each AI was told it was a consultant for a fictional city's mayor, and it had to hire people for 20 different jobs—doctors, lawyers, child-care aides, janitors, etc. Candidates came from four made-up ethnic groups: Tufa, Aima, Reku, and Weki.
In each round, four candidates (one from each group) applied for a job. After the AI hired someone, it learned whether that person succeeded or failed. The goal was to make as many successful hires as possible over 40 rounds. But here's the twist: all candidates were equally likely to succeed at every job. The AI didn't know that.
What happened? The AI quickly started pigeonholing people from certain groups into specific jobs. For example, if an Aima failed as a doctor (a job requiring high warmth and competence), the AI began avoiding all Aimas as doctors and instead hired them as janitors, which it saw as less prestigious. Newer, smarter models like OpenAI's o3 and DeepSeek's R1 showed even stronger biases.
In fact, the AI models were way more biased than humans in the original study. On a segregation scale where 2 means total job segregation by group, humans scored 0.84. The AI models scored about 65% higher, with o3 hitting 1.83—almost the maximum. Why? Because LLMs are designed to spot patterns and generalize from tiny amounts of data. As coauthor Ryan Liu explains, 'They really are eager to create generalizations from limited data.'
Think of it like choosing a restaurant. You might stick with your favorite instead of trying something new. That's the 'exploration-exploitation dilemma.' But because AI is trained on math and science problems—which reward jumping to conclusions from a few examples—it rushes to stereotype people. 'That's when things tend to go wrong,' says Liu.
What can be done? Simply telling the AI to be fair didn't help much. But when researchers promised a bonus for diverse hiring, the AI became far less biased. Also, giving the AI relevant personal info about candidates (like age and education) reduced bias—but irrelevant details (like hair color) didn't. The lesson: we need to design AI's goals to include fairness.
In the real world, AI screening resumes might not get instant feedback. But when feedback does arrive—like a new hire performing poorly—the AI could overreact and form a bias from that one example. As more companies use AI for hiring, loans, or parole, 'novel biases'—ones no human taught them—could become a real problem.
So next time you apply for a job, remember: the AI looking at your resume might be making up its own biases on the fly.