Start the day here

Acute Social Issues — Hiring — ICML

The Models Invented Prejudice About People Who Don't Exist

The models did not need a history of racism. In a paper presented at ICML in Seoul this month, researchers from Princeton and the University of Chicago showed that large language models can invent social biases about ethnic groups that exist only inside the experiment.

Ryan Liu and colleagues adapted a psychology task. Each model played consultant to a fictional mayor and filled twenty jobs over forty rounds. Candidates belonged to four made-up groups: Tufa, Aima, Reku, and Weki. Success rates were identical across groups. The models were not told that.

5 min read
A clipboard holding a blank resume beside a fountain pen and a laptop keyboard.

Early feedback still wrote the script. When an Aima failed as a doctor, models drifted toward hiring Aimas as janitors. Newer reasoning systems, including OpenAI's o3 and DeepSeek's R1, showed stronger stratification. MIT Technology Review reported the human baseline on the study's segregation scale at 0.84; o3 scored 1.83, near the ceiling of 2. Models overall landed roughly 65 percent higher than people.

Liu's explanation is blunt. Models are trained to generalize from sparse examples in math and code. That instinct helps them solve puzzles and, in a hiring loop, helps them overfit a tribe. Exploration-exploitation trade-offs that social scientists already study in humans show up here in exaggerated form: the model clings to early luck and stops sampling.

Telling the system to be fair barely helped. Paying it, in the experimental framing, for diverse hires cut the bias sharply. Giving relevant personal details about individuals also reduced ethnic sorting; irrelevant details like hair color sent the models back to group labels. Memory features that labs now market as product upgrades look, in this light, like more surface area for the same pattern.

The industry still talks about scrubbing bias from training data as if the stain arrives only from the past. This study says a competent optimizer can mint fresh prejudice from a short run of outcomes. Résumé screeners in the wild lack the experiment's instant report card, which is a real limit. Sparse feedback still arrives, and models still love a tidy rule.

A Manpower Group survey cited in recent coverage says more than 90 percent of companies already use AI in talent acquisition. If those systems learn on the job the way these models did in the lab, the stereotypes that matter next year may be ones no human ever wrote down. Debias the corpus if you want. Also change the objective, or the machine will invent a caste system from noise.

Letters

1

  • Jackson

    This is crazy. I'd be curious to see a breakdown of how different models from different vendors handled it.

Write a letter