The forces that drive behaviour are mostly invisible, even to the person doing the behaving. Which is why "why did you do that?" is one of the weakest questions in research.
Researchers split variables into two kinds. A manifest variable is something you can observe directly: age, the last brand someone bought, a click. A latent variable is a construct you can only infer from patterns across manifest indicators, things like brand affinity, political ideology, or need for cognition. The distinction is formal. Spearman (1904) inferred "general intelligence" (g) from the way test scores correlated, and modern latent-variable modelling (Bollen, 1989) is built on the idea. The catch for behavioural science is that almost everything that matters is latent, and the drivers of behaviour are typically not things people can introspect.
That second claim is the cornerstone, and it's well-evidenced. Nisbett & Wilson's (1977) review concluded that people have "little or no direct introspective access to higher-order cognitive processes." When asked why they chose, judged, or felt something, they don't read out a cause. They generate one from their implicit theories about what should have mattered. The most striking modern demonstration is choice blindness. In Johansson et al. (2005), people chose the more attractive of two faces, were then handed (by sleight of hand) the face they had rejected, and proceeded to explain, in detail, why they "preferred" it, all without noticing the switch. It replicates outside the lab, with shoppers tasting jam and tea (Hall et al., 2010). As Kahneman (2011) puts it, the mind acts first and narrates afterward, and the narration feels exactly like a reason.
The contested question is not whether people confabulate. That part is robust. It's how much, and when introspection can be trusted. Nisbett & Wilson themselves noted that reports can be accurate when the true cause is salient and a plausible candidate; the story only fails when the real driver is subtle or socially awkward. So "all self-report is worthless" overshoots the evidence. A second, separate dispute is about latent variables themselves: a factor extracted from a correlation matrix is a statistical summary, not a thing inside the head. Treating "Openness" or a "premium-seeker" segment as a cause is reification, and marketing is awash in barely-distinguishable latent constructs (engagement, love, equity, purpose) with shaky discriminant validity.
A latent construct is only as good as the theory and the indicators behind it; it has to be validated, not just extracted. And the confabulation paradigms, vivid as they are, are specific setups. The leap from "people sometimes invent reasons" to "introspection is always useless" is itself a misreading.
We still don't have a clean map of when introspection is trustworthy versus not (salience and plausibility are part of the answer, but not all of it), or how to recover latent drivers reliably without sliding into reification.
This is the foundation under every "why" question you've ever put in a survey. People are not lying when they tell you why they bought, voted, or quit. Or at least, they don't do it maliciously or in bad faith (the issue of spam respondents is indeed real, but that's not what we're discussing here). They're sincerely reporting a story their mind assembled after the fact. If you treat that story as the "cause" and act on it, you'll optimise the wrong thing. In other words, there is a hidden reason for people to answer they way they do, and it's not the one your survey probably wanted to discover.
"Why did you choose us?" collects stories, not causes. Attribution surveys ("what made you buy?", "what do you value most?") capture people's lay theories about themselves. They're useful as perception data and dangerous as causal data. Recover real drivers from behaviour instead: experiments, A/B tests, and comparisons across conditions. And when your analysis names a latent driver or segment, make it earn its keep. It should predict something a simpler, observable variable doesn't, or it isn't real enough to bet on.
Voters' stated reasons are rationalisations stacked on latent drivers: identity, group belonging, affect. "I voted on the economy" is a respectable story; the actual driver may be something the voter would never volunteer or even recognise. Don't build messaging off stated reasons. Build hypotheses from them and test them behaviourally.
Stated reasons for non-compliance are weak design inputs. Asking people "what would make you change?" returns their theory of themselves, which is often wrong. Pilot the intervention, watch what actually shifts behaviour, and let that drive the design rather than the focus-group explanation.
Demote the word "why." Treat every stated reason as a hypothesis to be checked against behaviour, comparison, or experiment. And discipline your latent constructs: name one only when it's been validated and earns its keep predictively. The honest posture is that you're inferring invisible drivers from visible traces, so triangulate, and stay suspicious of any explanation that arrived too neatly.