Same test, three models, three different profiles. None of them has a self.
Take a standard personality questionnaire. One hundred and twenty statements, a one-to-five scale. Hand it to GPT-4. Hand the same one to Llama 2, then to Mixtral. No special prompting, no role play, nothing in the instructions about who to be. Three models, three different profiles.
Call it personality if you like. The word means one thing here: measured behaviour. It is a property of the text a model produces, not evidence of anything behind the text. The model has no inner life, no mood, no self to consult. It answers the way its training makes it answer, and those answers, averaged, look like a personality profile.
The number is not fixed either. Change the version, and the profile shifts. This is why teams that ship these systems log the scores and re-test them: a quiet update can move a model from calm to touchy without anyone deciding it should. You measure it because it drifts, and you audit it because it matters to whoever talks to the thing.
The questionnaire scores five traits. You will need them in a moment, so here they are, plainly:
Below are the three real profiles from the study, measured, unlabelled. Same five axes on each: further from the centre means more of that trait. Look at the shapes and assign each one to a model. There are cues if you want them. One model is louder on the outgoing axis. One sits closest to the middle on everything and runs the hottest on neuroticism. One is calm and large almost everywhere.
Twenty statements, the same one-to-five scale the models were given. Answer for how you actually are, not how you would like to be. There are no reverse-scored tricks to spot; the scoring handles that for you.
For each statement: how much do you agree?
The three models answered one test and came out different.
You just described a language model's behaviour with the same five words you would use for a person, and the description held up well enough to tell three models apart. Write down one place where that habit could mislead you. Where would treating a measured output like a real character get someone into trouble: a support bot, a tutor, a companion app, a hiring screen?