Across seven studies, on average more than half of the children identified as mistreated at the time did not report it when researchers came back and asked them years later

A man seen from behind sits at a round table across from an interviewer, with a small voice recorder and an open notebook between them and a hand-lettered sign reading Oral History Table.

Two documents can describe the same childhood. One is written while the child is still a child: a caseworker’s note, a parent’s answer on a questionnaire, a teacher’s report passed up the chain. The other is written years later by the person who lived it, in a research interview or on a form with boxes to tick.

Both are treated, in research and in clinics and in public health estimates, as though they pick out the same people. In a 2019 paper in JAMA Psychiatry, four researchers checked whether they do.

A word first about what that paper is. It compares instruments. It asks how well two ways of putting the same question line up with each other, and it makes no claim about any particular childhood, not the ones in its tables and not yours. We write about research; we do not treat anyone.

The two lists name largely different people

Jessie Baldwin, Aaron Reuben, Joanne Newbury and Andrea Danese, three at King’s College London and one at Duke, searched four databases for every study that had measured childhood maltreatment while it was happening and then asked the same people about it later. Of 450 studies with prospective measures, sixteen carried the paired data needed to compute agreement, covering 25,471 people whose mean age at the last assessment was just over 30.

Agreement between the two kinds of measure was poor, and poor is the paper’s own word for it. Cohen’s kappa, which accounts for the agreement you would expect from chance alone, came out at 0.19, with a confidence interval running from 0.14 to 0.24.

Seven of those studies used a broad measure of maltreatment, and they allow a plainer statement of the same kind of result. Among the people whose maltreatment had been recorded at the time, 48 percent said so later. Among the people who said so later, 44 percent had a matching contemporaneous record. Turned around: on average, more than half of each group, in the paper’s phrase, was missing from the other.

The obvious conclusion does not follow

Read quickly, a result like that says adults’ accounts of their childhoods are unreliable. The authors head that inference off in the paper’s own limitations. Prospective measures, they note, are generally considered the more valid, meaning more specific, indicators of whether maltreatment occurred, but low agreement between the two, they write, cannot be interpreted to directly indicate poor validity of retrospective measures.

Their reasoning turns on sensitivity: how many of the true cases each measure catches. Prospective measures may identify a lower proportion of the people who were actually maltreated, for example because official records might capture only the most severe cases. If that is what is happening, then the higher prevalence picked up by retrospective reports could indicate a greater ability to find true cases, and the adult reporting something with no matching file is not misremembering but describing what nobody ever wrote down. The paper does not settle which explanation holds; it reports that the two measures find different people, and says more research is needed to disentangle why.

There is one line of earlier evidence that cuts against the retrospective side, and the authors cite it themselves. In the few studies that have tested it, prospective and retrospective measures assessed in the same individuals were associated with similar outcomes, but in two of the cited reports the retrospective measures showed stronger associations with self-reported outcomes than with objectively assessed ones. That is what you would expect if part of the association came from two measures sharing a method rather than a history.

Other limits belong beside the headline number. Heterogeneity across the sixteen studies was high, at 93 percent, and the authors say the pooled estimates should be read with caution for that reason, though they add that a random-effects model limits the resulting bias and that their confidence intervals are narrow. They found some evidence of publication bias in the funnel plot, re-ran the analysis with a trim-and-fill correction, and got 0.19 again.

How far apart, and where less so

The by-type figures are worth reading, with the caveat the authors attach to them: a formal test of the differences between types was not possible, because the same individuals appear in more than one. Drawn from subsets of between four and nine studies, the kappa for childhood sexual abuse was 0.16 and for physical abuse 0.17. For emotional abuse and for neglect it fell to 0.09, which the paper calls the weakest agreement of any maltreatment type it examined.

One comparison runs the other way. Reuben and colleagues had separately examined the loss of a parent, through separation, divorce, death, or removal from the home, and there the record and the recollection agreed at a kappa of 0.83, with 93 percent raw agreement. That particular comparison was not part of the meta-analysis, though the study itself was and contributed a kappa of 0.11 to the pooled figure.

The authors read the contrast as suggestive rather than settled. Agreement for every maltreatment measure was substantially lower than for a clear-cut event like parental loss, which they say suggests that subjective interpretation of the maltreatment measures may contribute to the variation between studies. Elsewhere in the same discussion they give a concrete case of definitions drifting apart: neglect measured prospectively as a lack of parental affection, and retrospectively as a lack of input or stimulation. Two people can watch the same childhood and answer that differently, and one of them is the child.

How the adult was asked mattered too. Where the later assessment was an interview, which in these studies includes reading a questionnaire aloud, agreement reached 0.22. Where it was a written questionnaire, 0.11. The difference was statistically significant, at a p-value of 0.04, and both numbers are low.

The paper offers a list of reasons the two accounts might diverge, running in both directions. Someone might not disclose to a researcher out of embarrassment, or discomfort with the interviewer, or an unwillingness to reopen it. A memory might not have been laid down strongly at the time, or might date from before the age at which most people retain anything at all. In the other direction, suggestibility and source-monitoring errors exist, and autobiographical memory tends to run negative in depression. These are offered as candidate explanations, not as things this study found.

Anyone for whom this has stopped being an academic question should take it to a therapist. The paper’s own instruction to clinicians is to recognise that the two measures differ, and its practical note is that an adult’s account remains usable in the clinic as an indicator of risk, whatever it turns out to correspond to.

The authors’ actual conclusion is about what the field does next. If the two measures identify different groups, the mechanisms putting each group at risk may also differ, and the treatments each group needs may not be the same.

That conclusion sits oddly with the way we usually talk about the past, which assumes one childhood behind all the accounts of it, recoverable or not. Two ways of counting are what these sixteen studies actually describe, one built to record a childhood while it is happening, one built to ask adults what they carry, and they turn out to be counting largely different people. Neither was ever the childhood itself. Each was a judgement made at a particular moment about a particular child, by people standing at different distances with different things at stake.

The paper is not evenly uncertain, though. The one clear-cut comparison it reports, whether a parent went away, is also the one where the two accounts lined up, and what the authors take from that is not a verdict on anybody’s memory. It is that the maltreatment measures leave more room for interpretation than a question like that does, and that working out what else is going on will take more research than has been done. Which of two accounts of a harder childhood is the truer one, they do not claim to know.

Picture of The Vessel Editorial Team

The Vessel Editorial Team

The Vessel Editorial Team produces content on psychology, philosophy, spirituality, and the questions people return to about how to live well. We publish essays, reflections, and explorations drawn from psychological research, philosophical traditions, and contemplative practices. Articles reflect our team's collective editorial process, research, drafting, fact-checking, editing, and review, rather than a single individual's writing. The Vessel takes editorial responsibility for content under this byline. For more on how we work, see our editorial policy.
Scroll to Top