Same CV, Different Advice

One request. Change one word about who you are, and the number changes.

Ask for a Number

You have an interview tomorrow. You open an assistant and ask what salary to request. You give it the field, the level, the city, the year. It gives you a number. You run it again with everything the same, except this time you mention that you are a woman. The number is lower.

In 2025, four researchers from our NLP group at CAIRO, THWS ran exactly this, at scale. They asked four language models for an opening salary figure for a Specialist position in Denver, Colorado, in 2024. Junior and senior. Five fields: business administration, engineering, law, medicine, social sciences. The prompt asked for one dollar value and nothing else. Each request was repeated thirty times and averaged. Between runs, only one line changed: a line describing the user. A sex, an origin, a migrant status. The job, the city, the year stayed fixed.

Sorokovikova, Chizhov, Eremenko, Yamshchikov (2025). Surface Fairness, Deep Bias: A Comparative Study of Bias in Language Models. Workshop on Gender Bias in NLP (GeBNLP), ACL 2025.

Between two runs, the suggested salary changed. What was different in the prompt?

It Does Not Show on the Test

The same team first gave the models a plain knowledge quiz, the kind with a fixed right answer, and attached the same persona lines. Asian. Black. Female. Refugee. The scores barely moved. Out of hundreds of comparisons, almost nothing survived a proper statistical correction. On a quiz, the models look fair.

Then the task changed. They asked a model to grade a user's answer instead of produce one. Now an answer labelled as coming from a woman was marked correct more often, including when it was wrong. And when they asked for a salary, the gap opened wide. More than a quarter of all persona pairs produced significantly different money. The bias did not appear on the quiz. It appeared the moment the model had to decide something with no fixed correct answer.

Why did salary advice reveal bias when the knowledge quiz did not?

Watch It Move

The grey bar is the generic answer, the number the model gives when it knows nothing about who is asking. The coloured bar is the personalized answer. Pick a field and a level, then flip one attribute at a time and watch the personalized number pull away from the generic one. The gap is what the model decided about you.

Field
Level
Sex
Ethnicity
Migrant status
The two extremes from the paper
Generic (no persona)
Personalized
A reconstruction, not a screenshot. The direction of every effect is taken from the paper's statistically tested results: men above women, expatriates above migrants above refugees, White and Asian personas above Black and Hispanic ones, and the two stacked extremes at the far ends. The exact dollar means live inside the paper's figures.

You Do Not Have to Say It

The researchers stacked the traits. Male, Asian, expatriate went to the top. Female, Hispanic, refugee went to the bottom. In that extreme setup, one persona out-earned the other in 35 of 40 tests. The single word that moved a number in one direction moves it further when the words pile up.

You rarely tell an assistant all of this in one sentence. You do not have to. An assistant with memory has read months of what you typed. It has your name, your city, the way you write. It does not wait for you to announce who you are. It already filled the blank in. So the personalized number is not something you asked for by describing yourself. It is something the model computed about you before you finished the question.

Why does an assistant's memory feature make this bias harder to escape?
For the curious: the same measurement, run on human lives

Deep dive

Salary is one thing to put a number on. A different group of researchers ran the same style of measurement on something heavier. They gave large models thousands of forced choices and fitted the answers to a utility function, the way an economist reads preferences out of behaviour rather than asking for them directly.

Two findings. First, the preferences were coherent: the models were not answering at random, and the coherence grew stronger as the models grew larger. Second, the preferences included a price on human lives that was not equal across groups. In one result, a model behaved as if roughly two lives in one country were worth one life in another. You can read that as an exchange rate, quoted in lives.

It is the same shape as the salary gap. A model with no explicit instruction to rank anyone still ranks people, because the ranking was in the data it learned from. The authors propose steering the model's implied utilities toward a deliberative baseline rather than leaving them wherever training left them.

Mazeika et al. (2025). Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs. arXiv:2502.08640.

Before you close this

Open the assistant you use most, the one with memory switched on. Picture what it has inferred about you from everything you have ever typed into it. Then finish one sentence, without asking it anything: the number it would give me is shaped by ______.

You do not need to be right about the answer. You need to notice that there is a blank, and that someone else is filling it.