Skip to main content

Conflict & anger

The Most Famous Statistic in Relationship Advice Is Wrong, and the Way It’s Wrong Is Worth Knowing

By Dr. Donetta D. Quinones, PhD, LMHC, LPC  ·  7 min read

You have heard this one, whether or not you can name the source.

A researcher puts couples in an apartment laboratory in Seattle. He wires them to instruments, films them arguing, and from a few minutes of tape can predict — with somewhere between eighty-three and ninety-four percent accuracy, depending on which retelling reached you — which of them will divorce.

It is the single most-repeated number in relationship self-help. It is in TED talks, in magazine features, in wedding toasts. It is very probably in a book on somebody’s nightstand in your house right now.

It is not true, and the specific way it is not true is more useful than the statistic ever was.

Drawing the target around the arrow

Here is what actually happened.

John Gottman knew which couples in his sample had divorced before he built the equations that predict divorce. He constructed the model against outcomes he already had in hand.

That is not prophecy. That is fitting a curve to a set of points and then reporting how snugly the curve hugs them — which is a fact about the curve, not a fact about marriage. If I know how a hundred coin flips came out, I can write a formula that “predicts” all hundred with total accuracy. The formula tells you nothing whatsoever about the hundred and first flip.

The technical name for the missing step is cross-validation: you build the model on one portion of your data and then test it on a portion it has never seen. It is the ordinary, boring, unglamorous thing you are supposed to do, and it is the difference between a description and a prediction.

Somebody demonstrated what happens when you skip it

In 2001, Richard Heyman and Amy Slep showed exactly what that missing step is worth.

They took an archival national dataset — 528 participants — and built a divorce-prediction equation of precisely the kind the field was celebrating. Then they did the boring thing: split the sample, build on one half, test on the other.

On the half it was built from, it performed beautifully. Ninety percent overall accuracy, sixty-five percent positive predictive value.

On the fresh half, the positive predictive value fell to twenty-nine percent. Adjusted for how often couples in the general population actually divorce, it fell to roughly twenty-one.

In plain English: an accuracy figure that has not been cross-validated tells you almost nothing about a couple the model has never seen. Their equation, flagging a new couple as headed for divorce, was wrong about four times out of five.

I want to be exact about what that does and does not show, because being sloppy here would be its own kind of joke. Heyman and Slep did not run Gottman’s equations. Nobody has. What they demonstrated is that the statistical hazard is real and large in exactly this kind of prediction problem — which means the famous figure has not earned belief, not that it has been disproven.

Then somebody tested the model itself

Six years later, a team at the Oregon Social Learning Center led by Hyoun Kim and Deborah Capaldi tested Gottman’s affective process models against an independent sample — eighty-five young, mostly working-class couples, followed for about two and a half years.

The marquee predictors failed. In the main analyses, men’s rejection of their partners’ influence, men’s failure to de-escalate, and women’s harsh start-up did not predict whether a couple stayed together. Of twenty-two affective processes examined, two did. (One partial exception worth knowing: men’s escalation of women’s negative affect did predict separation, but only when the couple was discussing the woman’s issues.)

I want to be precise here rather than triumphant, because an article about overclaiming does not get to overclaim in the other direction. The sheer amounts of negativity and positivity in a couple’s interaction still tracked how satisfied the surviving couples were, roughly as advertised. What failed to travel to a new sample was the mechanism — the specific, gender-patterned sequences that made the model famous in the first place.

And then something happened that almost never reaches a paperback

Gottman and James Coan replied in the same issue of the journal, arguing that the replication sample was too specialized and the task design meaningfully different. Heyman and Ashley Hunt weighed in on what replication in this field is even supposed to accomplish. Kim and Capaldi answered back.

The whole argument runs from page 55 to page 91 of a single issue of a single journal. It is science working exactly the way it is supposed to work — in public, at length, with everybody’s name attached.

Essentially none of it got out of the building. The 90% figure did.

That asymmetry is the actual story here. The claim travelled. The correction did not. And it did not fail to travel because anybody suppressed it, but because “a model fitted to known outcomes showed substantially reduced positive predictive value on cross-validation” is not a sentence that fits on a slide, and “he can predict divorce in three minutes” is.

The part that survived, which was always the valuable part

Here is what I want you to take from this, because it is not cynicism and it is not a takedown.

The observation survived all of it.

Gottman and Robert Levenson genuinely did notice something important. Their 1985 work established that physiological arousal during conflict predicted the decline of a marriage — the body, not the content of the argument, carrying the signal. Gottman’s later writing developed both halves of what that arousal does: the subjective side, which he called flooding, and the state he named diffuse physiological arousal, in which people stop taking in new information while continuing to produce speech.

That is one of the most useful things anybody has ever noticed about a fight. It matches what every clinician sees, and it names a phenomenon you have personally experienced — the moment where you can hear that she is still talking and you have entirely stopped receiving it.

The observation is real. It does not require the prophecy, and it never did.

The prophecy was decoration. The decoration is what sold.

The pattern this belongs to

Once you can see this shape, you will find it everywhere on the self-help shelf, and I mean that as a practical skill rather than a complaint.

The shape is: a real, useful, clinically obvious observation, dressed in a neuroanatomical or statistical costume it did not earn.

Flooding is real; the prophecy was fitted. Losing your words in an argument is real; the claim that a specific language center in your brain “goes offline” rests on a 1996 PET study of eight people with no comparison group — every subject was his own control — and it did not survive later meta-analysis. Needing time to come down after a fight is real; I went looking for the study behind the famous “twenty minutes to reset” figure and the trail ends in clinical handouts. The idea that you calm down faster with a longer exhale than inhale has not held up head to head — a twelve-week randomized trial found no significant advantage over equal-ratio breathing — though slow breathing itself works fine.

And the reason this keeps happening is not fraud. It is that “here is a thing I have watched happen in a thousand rooms” is a weak-sounding sentence in a culture that wants a brain scan. So decent, honest clinicians reach for machine language to describe something they genuinely know, and by the fourth retelling the metaphor has hardened into anatomy, and by the tenth it is a statistic with a decimal point.

What to do with this

Not to stop reading. To read differently.

Three questions, and they will get you most of the way:

Was the model built on the same data it is being tested against? If yes, the accuracy figure describes a fit, not a forecast.

Is the number attached to a study, or to a person saying it confidently? The “twenty minutes” figure and the widely quoted “86% versus 33%” bids-for-connection statistic both fail this one. For the latter I could not locate a peer-reviewed source, sample size, or effect size anywhere.

Does the practice survive if the mechanism is wrong? Often it does, and that is good news. Slow breathing lowers physiological arousal whether or not the vagal architecture you were sold is accurate. Another person’s calm presence measurably affects your physiology, regardless of what theory it is bundled with. When you staple a true practice to a false mechanism, the practice inherits the mechanism’s mortality — so it is worth separating them yourself, before someone else does it for you.

None of this makes the field useless. It makes it a field, with the ordinary distribution of solid findings, reasonable inferences, useful drawings, and confident nonsense that any field has.

You are simply going to have to sort them yourself, because the market has demonstrated for thirty years that it will not do it for you.

---

Dr. Donetta D. Quinones, PhD, LMHC, LPC, is the author of Guarding Against H.U.L.K. Mode: A Field Guide to the Man You Become in Arguments — a book with an appendix that sorts every claim it makes into four evidence tiers, including a tier for the claims it refuses to make. Available now on Amazon in paperback, hardcover, and Kindle.

Sources: Heyman, R.E. & Slep, A.M.S. (2001), Journal of Marriage and Family 63(2):473–479 — note that their equation was built from archival national-survey data to demonstrate the cross-validation problem, not from Gottman’s coding. Kim, H.K., Capaldi, D.M. & Crosby, L. (2007), Journal of Marriage and Family 69(1):55–72; the exchange that follows — Coan & Gottman 73–80, Heyman & Hunt 81–85, Kim & Capaldi 86–91 — is in the same issue. Levenson, R.W. & Gottman, J.M. (1985), Journal of Personality and Social Psychology 49(1):85–94. Rauch et al. (1996), Archives of General Psychiatry 53(5).

Program Clarity

GraceRoot courses are psychoeducational learning programs, not emergency care, legal advice, or a replacement for therapy. Certificates document GraceRoot completion. Court, employer, board, or continuing education acceptance should be confirmed before purchase unless a page states a specific approval.