We have read relationship research the way most people read it, one study at a time, and one study at a time is a flattering way to read anything. Each paper arrives with its own variable in the spotlight: attachment style, conflict style, how a couple talks about money, whether they laugh at the same things. Read in sequence, they all sound like the answer.
In 2020, eighty-six researchers put 43 of their own longitudinal datasets into one comparison to see which of those answers survived being ranked against the others. The material covered 11,196 couples, followed over months and in some cases years, and thousands of measures collected at the start. Models built on a person’s own answers about their relationship could account for a large share of how satisfied that person said they were at the moment of asking, and the single measure that showed up most reliably among them was whether they believed their partner wanted the relationship to last. Nothing anyone collected reliably predicted whose relationship was about to get better or worse.
Nobody could predict the second year
The datasets were not snapshots. Couples were followed across an average of four time points over roughly fourteen months, with some studies running as long as four years, and the analysis plan for each dataset was registered before anyone ran it. Nearly everything each partner reported at the start was fed into a machine-learning method built to sift many predictors at once without overfitting any single dataset, and the 43 results were then combined.
Predicting satisfaction at the start of a study worked well. A person’s own answers about their relationship accounted for about 45 percent of the variation in how satisfied they were, and their answers about themselves, their traits and moods and history, about 19 percent. At the end of each study the same predictors still worked, more weakly: 18 percent at best from the relationship answers, 12 percent at best from the answers about the self.
Those predictors were collected at the same sitting as the outcome, so part of the 45 percent is people describing their relationship in one place on a form and rating it in another. The figure already excludes trust, intimacy, love and passion, four measures that arguably describe relationship quality rather than predict it, and a stricter run that removed eight more variables did not change the conclusions. It remains a number about one afternoon.
Predicting change is where it stopped.
No model of who would rise and who would fall accounted for more than 5 percent of the variation, and the averages sat closer to 2 percent. Whether a relationship was on its way up or on its way down at the moment everyone filled in the forms was not legible in the forms.
Almost nobody wonders how happy they are today. They wonder where this is going, and that is the question the models could not touch. The authors’ own reading is that the answer lives in what they did not measure, in things that had not happened yet: stressful life events, the arrival of a child, the ordinary weather of a life, none of which can be read off a questionnaire filled in a year or two earlier.
The answers that kept earning their place
Five relationship measures earned their place across dataset after dataset. Believing a partner is committed came first, measured with lines as blunt as “My partner wants our relationship to last forever.” Then appreciation, the paper’s example being agreement with something like “I feel very lucky to have my partner in my life.” Then satisfaction with the couple’s sex life, then the belief that the partner is happy in the relationship, then how often the two of them fight.
On the individual side, the strongest were satisfaction with life in general, negative mood, depression, and the two sides of insecure attachment: the fear of being left and the reflex to keep distance. Gender, ethnicity and education mattered little, and so did the outwardly factual features of a couple’s life, whether they lived together, whether they had children, whether they were dating or married. Relationship length was the exception among those.
The most-quoted part of the paper is what happened when the partner’s own answers were added to a model. They barely moved it. Adding them, together with both people’s answers about themselves, shifted the result by somewhere between nothing and about two percentage points at the start of a study, and by no more than three and a half at the end. Whatever those things contribute appears to arrive through the way one person already sees the relationship, rather than sitting on top of it.
This is where the popular reading runs ahead of the paper. The headline version says your partner does not matter, or that compatibility is a myth. What was tested was one partner’s questionnaire answers as a predictor of the other partner’s rating, and the paper’s own account is that a partner’s influence is likely real but travels through the relationship the two of them build, which is what the person’s own answers already describe. A model that gains nothing from a variable has not shown that the variable does nothing. It has shown it had nothing left to add once the closer measure was in.
Every dataset came from the United States, Canada, Switzerland, New Zealand, the Netherlands or Israel, and nearly all of the relationship measures were rating scales people filled in about their own relationship rather than anything an observer watched happen. One thing works in the findings’ favor: more than half the couples in the project came from a marriage-support study that deliberately oversampled low-income households, with more than four in ten of them reporting incomes below the poverty line, and the pattern there matched the other 42 datasets.
That last point is where a later study pushes back. In 2021, James McNulty, Andrea Meltzer, Lisa Neff and Benjamin Karney pooled ten longitudinal studies of 1,104 married couples in the same journal, and instead of relying on questionnaires alone they added observers watching each couple work through a problem, plus repeated reports of stress from both spouses. On those measures both people’s qualities did predict how satisfaction changed, accounting for close to 16 percent of the variance in its trajectory, with the effect running through the observed behavior and shifting with how much stress each spouse was carrying. The unpredictability in the 43 datasets is a finding about self-report questionnaires, not a verdict on relationships.
Reading research is what we do; treating people is a different job entirely. A study is not a diagnosis, and none of this is an assessment of any particular marriage. It describes averages across thousands of couples in six countries. It does not tell a couple what to change, and it does not promise that attending to your own sense of being lucky will produce more of it. Couples and family therapists exist for relationships that have started to hurt.
So the next single study will arrive, as they do, with one variable in the spotlight. This one arrived with 2,413 measures at once, and the one that earned its place most often was a sentence a person could have written on a napkin: my partner wants our relationship to last forever.