Almost everything we know about people's traits comes from asking them to rate themselves. That works better than you might fear, but it breaks in specific, predictable ways. Knowing those ways is the difference between using the data and being fooled by it.
Step back and notice what nearly every finding in this layer (L2 - Dispositions) is built on. When we say "people high in conscientiousness do better at work", or that "anxious attachers read silence as rejection", the underlying data almost always come from the same place: someone read a statement like "I am always prepared" and ticked a box from "strongly disagree" to "strongly agree". Self-report questionnaires are the workhorse of the field because they are cheap, fast, and, when the scale is well made, both reliable and genuinely predictive. None of what follows means the data are worthless. It means the workhorse has some predictable limps, and serious practicioners learn to spot them.
The most familiar is social desirability: people tilt their answers toward what looks good. The classic scale for catching it is decades old (Crowne and Marlowe, 1960), and later work split the tendency into two (Paulhus, 1984). One half is honest self-deception, you genuinely think you are more patient than you are. The other is deliberate impression management, you know the truth and tidy it up for the audience, and tellingly, impression-management scores jump the moment answers stop being anonymous. People shade harder when someone is watching.
A second family of distortions has nothing to do with the content at all. These are response styles: the habit of agreeing with whatever is asked (acquiescence), or of always reaching for the extreme ends, or of hugging the safe middle. Because they are content-free, they quietly bias not just scores but the apparent relationships between scales, and they differ systematically from group to group and country to country (Baumgartner and Steenkamp, 2001). Two people with identical views can land in different places purely from answering style.
Third, in any setting where the answer has stakes, there is faking. When a questionnaire is part of a job application, the incentive to present well is obvious, and people act on it. A striking paper assembled five former editors of the field's main journals, none with a stake in selling tests, to reconsider personality testing in hiring; they concluded that faking cannot really be prevented, and, more pointedly, that the validity of these tests for predicting job performance was low enough to question the whole practice (Morgeson et al., 2007).
The fourth is the most underappreciated and the most damaging to cross-country comparisons: the reference-group effect. People do not rate themselves against humanity in general, they rate themselves against the people around them. So a hard-working person in a hard-working culture, comparing to demanding local peers, may tick "moderately conscientious", while someone more relaxed in an easygoing culture rates themselves "very conscientious". The result is that self-report averages can fail to show real cultural differences, and sometimes reverse them: experts agree East Asian cultures are more collectivist, yet trait questionnaires did not show it, until researchers changed the comparison standard, at which point the expected difference reappeared (Heine et al., 2002).
The faking finding kicked off a live argument: does faking actually wreck a test's usefulness? People clearly do it, but several researchers counter that the tests' (already modest) ability to predict performance survives it reasonably well, since most candidates shade in similar directions. The sharper version of the critique sidesteps the question: the editors' panel argued the validity was low enough to begin with that faking is almost beside the point (Morgeson et al., 2007). So the real debate is less "do people lie on these" and more "are they accurate enough to make decisions with".
A deeper worry is whether the structures we measure even hold up outside the rich, Western samples they were built on. When a large project ran Big Five questionnaires face to face across 23 low- and middle-income countries, nearly 95,000 people, the familiar five traits frequently failed to appear at all, the items simply did not hang together the way they do in Western internet samples (Laajaj et al., 2019). Combined with the reference-group effect, this leaves cross-cultural comparison of self-reported traits on very thin ice.
If self-report is this leaky, why not measure another way? People have tried. You can ask people who know the person, use forced-choice formats that make impression management harder, watch behaviour directly, or read digital traces: one well-known study predicted personality and much else from Facebook "likes", with openness predicted about as accurately as a standard questionnaire predicts itself on a retest (Kosinski et al., 2013). Each alternative fixes some problems and brings others, from informants' own biases to serious questions of privacy and consent. There is no clean escape from self-report, only a menu of trade-offs.
Finally, the field has a famous cautionary tale: the Myers-Briggs Type Indicator. It is enormously popular and commercially huge, but psychometrically weak. It forces smooth, continuous variation into crisp types, so about half of people get sorted into a different type when they simply retake it weeks later, its underlying four-type theory does not hold up, and it predicts outcomes poorly (Pittenger, 2005; Stein and Swan, 2019). It is the standing example of how an intuitive, likeable instrument can thrive for decades while failing the basic tests of measurement.
Even the tools for catching these problems are imperfect, social-desirability scales catch only some of the shading, and every fix trades one bias for another. This literature is also itself mostly built on the same Western samples it warns about. And the honest headline is balance, not nihilism: self-report remains the most practical way to measure many traits, and well-built scales earn their keep. The point is to use the number while knowing what could have bent it.
Can disposition be measured well without leaning on a person's self-image at all? Is valid cross-cultural comparison of traits ever really possible, or only within a culture? And do digital traces genuinely free us from self-report's distortions, or just move the problem somewhere harder to see, at a steep cost to privacy?
The usable core: treat any self-report number as a measurement that motive, habit, and culture could have moved, lean on it most where stakes are low and comparisons are within one group, and reach for behaviour or multiple sources when a decision really matters.
Two practical warnings. First, be sceptical of personality questionnaires as hiring gates: candidates are motivated to fake, and the tests predicted performance weakly even before faking (Morgeson et al., 2007), so they pair badly with high-stakes decisions, exactly the lesson the hiring sections of the ability and dark-traits pieces reached from the other side (L2-06, L2-07). Use structured interviews, work samples, and several raters instead. Second, do not build customer segmentation on typologies like MBTI, and treat the neat trait averages in a survey deck with suspicion, especially when they compare groups or countries, because response styles and reference groups can manufacture differences that are not there.
Polling and message-testing run on self-report, and the sensitive questions are the ones most distorted. People under-report socially frowned-on views and over-report virtuous behaviour like voting, so raw answers on turnout, prejudice, or taboo positions are systematically off. The fixes are design choices: genuine anonymity, indirect questioning, and treating a confident topline on a sensitive issue as the floor of a range rather than the truth.
Official statistics and well-being surveys lean heavily on self-report, and the reference-group effect makes the popular international league tables genuinely treacherous: ranking countries on self-rated trust, life satisfaction, or conscientiousness can put them in the wrong order, because each population is grading itself against a different local standard (Heine et al., 2002). Self-reported health behaviours are similarly shaded. The safer path is to anchor questions to concrete, comparable events and to pair surveys with behavioural and administrative data.
Three habits. First, when you see a self-report number, ask what would make someone shade it: is it anonymous, are the stakes high, which way does "looking good" point? Second, distrust cross-cultural and cross-group comparisons of self-rated traits by default, the reference-group effect and shaky measurement invariance make them unreliable unless specifically validated. Third, for decisions that matter, triangulate, behaviour, forced-choice formats, and multiple informants beat a single self-report, and no serious choice should rest on an MBTI-style type.