Self-Understanding
Everybody Has a Score
The reading room of the public library closes at nine, and Odile has dragged the stacking chairs into a circle under the window that never shuts properly.
Wednesday, coming up on half past seven. Nine people, no food because the sign says no food, and a paperback with a cracked mirror on the cover. The quiz in the back is what everyone has come to discuss, and somebody has found the same one online: a free site, a couple of dozen statements, three colored bars at the end. Not a clinic. A web page with an ad for mattresses down one side.
The phone goes around. Terrence goes last and answers honestly, which turns out to be a social error. One of his bars comes back high. He reads it, laughs once — too loud for the room, and a man at the newspaper rack looks over — and turns the phone so the circle can see.
Nobody speaks. Then four people speak at once, at library volume, all of them kind and all of them slightly too warm — the register you use with somebody who has just mentioned a scan.
Terrence puts the phone face down on his knee. “Okay,” he says. “So what’s the line?”
Odile, who spent thirty-one years typing down what people said in courtrooms, looks up. “The line?”
“The number where it starts being a thing. Like blood pressure.” He turns the phone over and back. “What’s this one?”
Nobody has an answer. The paperback did not give them one, the website did not give them one, and there is not one to give.
There is no line, because there is nothing for a line to divide
These traits are dimensional. Plainly: everybody has some of each, and people differ in how much. Nobody has zero.
Picture eight people at a table eating the same plate of eggs. One hand reaches for the salt before tasting. One reaches twice. One never reaches at all, one hovers, and the rest are somewhere in between. Now draw the line. Say which four are Salt People. You cannot, and not because you lack data about their eating — because the question has nothing in it to answer. Height works the same way; nobody has ever seriously asked where tall stops and short begins.
So you have a score on all three. So do I, and so does everyone in that room, and everyone in your family, and the person you are worried about. There is no cutoff — not a soft one, not an unofficial one, not one clinicians know and are keeping quiet about. Blood pressure has a number because a body has an outcome the number predicts. This does not.
Therefore: a quiz cannot place anybody, and neither can a book. This one cannot. I cannot, and I have sat across from people for a living for a long time.
And since everybody has a score, recognizing something of yourself in a description is not a finding. It is what a description of a normal-range trait is built to do. If a small cold drop lands somewhere in the next section, it is not a verdict.
Three questionnaires and some undergraduates
So where did a phrase with no line in it come from?
In 2002, two psychologists named Delroy Paulhus and Kevin Williams published the paper that introduced the term. They gave three existing questionnaires to a few hundred undergraduate psychology students and looked at how the scores related to one another. That is the study. It is perfectly respectable. It did not find a new kind of human being.
The first sentence of their abstract is the part that gets sanded off in every retelling. They called these offensive yet non-pathological personalities, and labeled two of the three subclinical right there in the naming sentence — below the level at which a clinician would be looking at all. Their conclusion was that the three are overlapping but distinct as currently measured, which is a careful claim about what a set of scales was doing rather than a claim about kinds of persons.
The traits are duller than the cover art too. Narcissism, here, means grandiosity: thinking well of yourself past the point the evidence supports and needing others to agree. Picture the neighbor who rebuilt his own fence and is now in your driveway describing the post spacing, who will not be finished for eleven minutes and will not ask about your week. The fence is good. That is the part the cracked-mirror books cannot afford to say, because the whole product depends on the traits being alien.
What the instruments actually weigh
Two of the three keep failing to be two things. When researchers measure Machiavellianism and measure psychopathy, the scores behave so much alike that serious people argue the field is measuring one thing twice. Muris and colleagues, pooling a very large number of samples, found the overlap high enough to make the question serious, with narcissism sitting noticeably apart. Miller and colleagues published a paper whose title is also its argument — “Psychopathy and Machiavellianism: A Distinction Without a Difference?” — and found a model fusing the two fit the data as well as the three-part model everybody uses. Worse for the theory, the Machiavellianism scale correlates with impulsivity, when the whole idea of the Machiavellian is patience.
The other side is behavioral, and this is a live argument rather than a settled one. Jones and Paulhus set up situations and watched what people actually did. When cheating was nearly risk-free, all three traits predicted it. When there was serious risk of being caught, only psychopathy did. Under pressure, with something to lose, two traits that look identical on paper do different things.
The scale is not level, and it never was
Most of what you have read about this rests on one of two very short questionnaires: one carries nine statements per trait, the other four. The field’s own floor says that if a subscale measures one thing, at least half of what its items are doing should be that thing. When Knitter and colleagues tested both short scales against that floor, the subscales came in well beneath it.
Plainly: the tools are much weaker than the confidence of everything built on them.
Picture a kitchen scale you have never calibrated, on a counter that is not quite level. Weigh one bag of flour ten times and you learn something real. What you cannot do is compare it against a scale in somebody else’s kitchen, or tell a person their bag is over the line.
And nearly all of this research is self-report, one moment in one life, mostly from undergraduates and paid internet workers. So the field’s principal method for studying deception and self-flattery is to ask people who score high on deception and self-flattery to describe themselves accurately, once, on a form, for course credit. Reviewing the whole literature, Muris and colleagues found exactly one study that gathered information from anybody other than the participant. One.
I am not sneering; the researchers say all of this in their own abstracts. These are modest, real, badly instrumented findings, sold outside the building as a lens that can identify the person in your house.
A score, a judgment, and the instrument you already have
A number from a questionnaire is not a clinician’s judgment made against written criteria with a history in hand, and neither of those is the word somebody uses about an ex in a kitchen — though all three travel under the same handful of names. The middle one, the only one anybody can check, is also uncommon: gathering the community studies that used structured interviews, Dhawan and colleagues found narcissistic personality disorder running at roughly one in a hundred. What the confusion between the three does to an ordinary conversation is the whole subject of another piece in this set, and I will leave it there.
That is a real demolition and I am not taking any of it back. If something went flat in you reading it, that makes sense; people come to this material because a phrase out of a laboratory seemed to promise an outside authority at last.
Here is what the demolition does not touch. Your problem was never measurement. The question in a kitchen is not how much of this trait does he carry relative to a room of nineteen-year-olds who needed the course credit. It is what happened, how often, and what happened when it was raised.
Consider what you have, in the field’s own terms. It is longitudinal — a researcher measures once; you have watched one person across years, through two jobs, a funeral, and the month the boiler went. It is multi-informant — your observations, plus your sister’s, plus the thing the neighbor said in the driveway. And it was gathered where behavior actually cost something, which is the property this research most often apologizes for.
Your evidence is still about one person, it is not a diagnosis, and you were never a neutral observer. The correction is not to discount what you saw. It is to write it flat, and notice that the flat version survives being read back and the interpretation does not.
You have been collecting it for years without calling it data, because it did not look like data. It looked like Tuesday.
One thing you can do this week
Five things the quiz never asked. Ten minutes, alone, on paper.
Take one of the free quizzes — about yourself, not about anybody else. Answering on behalf of a person who is not in the room is the most common misuse of these forms, and it produces nothing but a mirror. Before you look at the result, write down what you expect each bar to say.
Then look, note where you were wrong, and do the actual exercise: write down five things the form never asked about. Start with these and add your own. What you do when you are wrong in front of somebody. How you behave toward a person who can do nothing for you. What happens in your house on an unremarkable Wednesday. Whether anybody has ever told you something hard about yourself and stayed.
Keep the second list. Throw the bars away.
The short version
- The Dark Triad came out of a 2002 questionnaire study of undergraduates. The researchers who named it called these traits offensive but non-pathological, and two of the three subclinical.
- The traits are dimensional. Everybody has a score on all three, there is no cutoff, and no quiz can place a specific person.
- Two of the three overlap so heavily on the standard scales that whether they are one thing is an open argument inside the field.
- The short instruments carry four to nine statements per trait and fall below the field’s own minimum standard. Nearly all of this work is one-time self-report from students and online panels.
- A score from a form is not a clinical judgment, and clinical judgments here are uncommon — roughly one in a hundred in community studies that used structured interviews.
- Recognizing yourself in a description of a normal-range trait is not a finding and not a verdict. It is what such a description is for.
Terrence and Odile are composites, assembled out of many rooms and belonging to nobody in particular.