Category Archives: WEIRD

It is OK to be a WEIRD science

It Is Fine to Be WEIRD

Americans love acronyms, and none has traveled further in social psychology than WEIRD — Western, Educated, Industrialized, Rich, and Democratic — coined to criticize a discipline for building a science of humanity out of American undergraduates. The critique was fair, but it points in the wrong direction. Being WEIRD is not the problem. It is perfectly legitimate for WEIRD researchers, paid by WEIRD institutions, to study WEIRD people in order to help WEIRD societies — to ask whether a therapy relieves depression in a Western clinic even if it would do nothing in another culture, or even if the disorder as we define it barely exists there. Local knowledge is not lesser knowledge.

The problem is not being WEIRD. It is being WEIRD while claiming to be universal — and psychology keeps making that claim because it has never quite decided whether it is a natural science of universal laws or a social science of particular societies. This essay is about that confusion, the two rituals that keep it in place — the apologetic limitations paragraph and the meta-analytic average that pools everyone into a number belonging to no one — and what becomes possible once we accept that it is fine to be WEIRD.

Philosophy with p-values

Psychology was born from philosophy but wanted to be physics. It took philosophy’s questions — how we perceive, learn, remember, decide — and set out to answer them with the methods of the natural sciences: experiment, measurement, and laws that hold for everyone. Call it philosophy with p-values. For the questions it started with, this worked reasonably well. Basic perception really is close to universal. A Weber fraction measured in Leipzig is likely to look much the same in Toronto; basic visual processes do not change dramatically from one culture to another. In the laboratory of basic processes, one human is often interchangeable enough with another that findings can generalize broadly.

That success was a trap. It made universality the price of admission to scientific psychology: if your findings held for everyone, everywhere, you were doing real science; if they did not, you were doing something softer and less scientific. The standard was manageable while psychology studied processes that are, in fact, close to universal. It became a problem when psychologists turned to everything else — love, prejudice, persuasion, the self — and kept the same standard.

There are two ways to apply a universal method to a variable subject. The first is to deny much of the subject matter. Behaviorism largely did that: mind, meaning, and emotion were pushed aside in favor of observable stimulus and response, which could be studied in a more law-like fashion. Whatever would not fit the method was ruled unscientific and shown the door.

The second way survived behaviorism and is the one we still practice. Instead of changing the questions, we changed the way we studied them. Social psychology adopted the model of the experimental laboratory — controlled experiments, manipulations, deception, participants reduced to “subjects” — so that social life could be studied as though it, too, obeyed general laws.

What psychology was reluctant to consider was another possibility: that the study of social behavior might be a different kind of science, with standards of its own and no need to discover universal laws.

Biology shows that this was a choice, not a necessity. It is a natural science with genuinely universal principles, and no one doubts its scientific credentials. Yet while reproduction is universal, how organisms reproduce varies enormously, and it would be absurd to insist on one detailed theory of sexual behavior spanning salmon, praying mantises, and swans. Biology does not apologize for this. It allows general evolutionary principles to coexist with the study of particular species and ecological contexts. The second does not have to dissolve into the first to count as science.

The False Promise of Experimental Social Psychology

The first challenge to universality came from studies within WEIRD cultures showing that people behave differently in the same situation. Walter Mischel’s Personality and Assessment (1968) argued that behavior could not be predicted very well from broad personality traits. Social psychologists pushed this argument further. Ross (1977) called the tendency to explain behavior in terms of personality while underestimating the situation the fundamental attribution error, and Ross and Nisbett (1991) went so far as to write that “one cannot predict with any accuracy how particular people will respond.” In this way, the search for universal laws of behavior could continue. Personality psychologists marshalled evidence that people really do differ from one another, but experimental social psychology did not respond by making those differences central to its theories. Social psychologists went on manipulating situations in the laboratory and largely treating differences between people as error variance.

The second challenge came from outside the Western laboratory. Beginning in the 1970s and gathering force through the 1980s, cross-cultural psychologists showed that many supposedly universal findings varied systematically from one culture to another — that the mind studied in Michigan was not simply the human mind. In principle, the point was conceded: virtually everyone now agrees that culture shapes cognition and behavior. In practice, much less changed, because studies kept being run in the same few places.

The complaint finally crystallized, decades later, in Henrich, Heine, and Norenzayan’s 2010 article on the “weirdest people in the world” — WEIRD, for Western, Educated, Industrialized, Rich, and Democratic. If anything, WEIRD may make psychology sound more diverse than it was. France is WEIRD, Germany is WEIRD, Sweden is WEIRD, but the standard participant was not drawn from a representative cross-section of Western societies. The empirical base was often much narrower — closer to a WASP monoculture of white, middle-class American college students.

The article was cited everywhere and changed surprisingly little, much like the cross-cultural critique before it. Papers still open with theories stripped of cultural context and test them on WEIRD samples; they simply close, now, with a limitations paragraph acknowledging that the sample was WEIRD and urging further research on other populations — by someone else. The acknowledgment goes in the one section of a paper that rarely changes the interpretation of the findings, and the work proceeds much as before.

And the asymmetry hides in plain sight, even in the names. Other countries mark their journals — the British Journal of this, the Iranian Journal of that — while the journals that set the field’s agenda do not: they are simply Psychological Science or the Journal of Personality and Social Psychology, as if nationality did not apply to them. So an American study in an American journal is readily read as a finding about people, while an Iranian study in an Iranian journal is read as a finding about Iranians. The universality we claim to have renounced still runs underneath — the default setting for one population, and the denial of that status to every other.

So the apology is the wrong response to the right observation. To acknowledge that a sample was WEIRD and then generalize anyway leaves the universalist assumption intact and merely adds a note of caution. The alternative is not to apologize but to specify — to name the population you studied and claim nothing beyond it. There is nothing unscientific about a finding that holds in one society and not another. Whether a result generalizes across cultures is a question to be answered, not a box to be ticked or an assumption to be smuggled in.

Everybody knows that people differ from one another; that does not make anyone weird. What is truly weird is a science of human behavior that ignores this diversity and imagines that behavior can be reduced to a single number — one universal effect of a situation on everyone. The remedy is not complicated: stop making global generalizations and ask instead a specific question about a specific population. Sometimes, in other words, it is perfectly OK to ask a WEIRD question.

Demographic Acronym “WEIRD” Overused in Psychology Research | Psychology Today

The Weirdest Statistical Method: Meta-Analysis

It is widely recognized that many psychological studies cannot provide conclusive evidence for an effect, let alone against one. Sample sizes are often too small to show that a result is more than a statistical fluke, or to pin down how large an effect actually is. The proposed solution is meta-analysis: find all the published studies on a topic and combine them to estimate the average effect size. That average is then presented as the true population effect size.

Pooling does solve the problem it was built for. Combine enough studies and the average becomes precise — the sampling error that plagued each small study shrinks toward zero. But a precise average is worth nothing if the quantity it averages over is not one quantity at all. If the true effect varies from study to study — across populations, procedures, and cultures — then pooling delivers an exact estimate of a number that describes no one. And heterogeneous effects are exactly what psychology’s meta-analyses pool.

To see why that is a mistake, and to see it clearly enough that no one can accuse me of being against averages, forget psychology for a moment and consider the average height of 8 billion human beings on this planet. Suppose it is 165 centimeters. There is nothing wrong with the number. It is correct, it is precisely estimated, and it is the honest answer to a well-defined question: what is the mean of the height distribution over all living people? Now try to use it. Design a doorframe? You do not build to the mean; you build to a high percentile, because the mean is silent about the tail and the tail is the entire point. Manufacture clothing? Here the average is not merely useless but actively misleading, because no one is average on every dimension at once, and a garment cut to the mean neck, mean arm, and mean torso fits no actual body.

So the average can be correct, precise, and useless, all at once, with no statistical error anywhere. The uselessness is not a flaw in the estimate. It is a property of the question. And the question carries a hidden assumption — one no one would defend if asked, and that the practice acts on regardless. Put the claim baldly and it is absurd: that there is a single true value and every individual simply equals it, the variation being error. No one believes that. but in a psychological meta-analysis the same assumption slips through, not as a belief anyone holds but as a convention everyone follows: the pooled effect is written up as the effect, cited as the effect, carried into the next study as the effect — as though the number applied to everybody.

So, to summarize, estimating a single average for a heterogeneous set of objects is a weird question that no one would consider meaningful to ask or to answer. But its weirdness is hidden by the universality assumption — that variation between people is mere error variance, and that the truth is a single number applying to everybody. A meta-analysis of mindfulness therapy illustrates that I am not attacking a strawman, but that the problem is real. I picked this meta-analysis because I am interested in the effectiveness of mindfulness therapy for WEIRD people in Canada. I don’t want to generalize to all other meta-analyses, but it is likely that this is not the only meta-analysis that failed to take cultural differences into account.

Does Mindfulness Therapy Work?

Goldberg and colleagues (2018) set out to answer exactly that question with a meta-analysis of mindfulness-based therapy. They collected studies spanning a range of disorders — depression, anxiety, substance use, and more — conducted in different populations and countries and using different mindfulness protocols, from standardized programs like MBSR and MBCT to local adaptations. They then pooled these studies to estimate average effects. For depression compared with passive controls, for example, the estimated effect was d =.6, with a 95% confidence interval from .5 to .7 — a moderate effect, the kind of number that easily becomes “mindfulness works” by the time it reaches a textbook or a clinician.

But look at the question that average answers — “the effect of mindfulness therapy” — and you will recognize the problem from the previous section. It treats “mindfulness therapy,” without further qualification, as though there were a meaningful effect to be estimated across all of these studies. Yet the studies were not one thing. Mindfulness for chronic pain in an American clinic and mindfulness for depression in a Chinese university are no more the same treatment of the same disorder in the same population than a newborn and an adult are the same height. The pool is heterogeneous by construction — across disorders, populations, and therapies at once. Asking for its single average effect is like asking for the average height of everyone on Earth.

I reanalyzed the open data with z-curve3 (Schimmack, 2026), which is designed to model heterogeneous evidence. The first step is to check for publication bias, and there was little evidence of it — unusual in psychological research, but less surprising in a meta-analysis that includes many nonsignificant results. This means that the full set of 214 positive effect-size estimates can be analyzed without a large correction for selective reporting. The studies were, on average, modestly powered: only about 45% reached significance, reflecting a literature that mixes a few well-powered studies with many underpowered ones.

But the decisive quantity is not the average power or even the average effect. It is the spread. z-curve3 estimates not only the mean true effect but also the distribution of true effects across studies. The estimated mean was .47, reassuringly close to Goldberg’s pooled estimate, with a standard deviation of .29. Those numbers imply a 95% prediction interval from about −.10 to 1.10. In other words, the true effect in another study drawn from this literature could plausibly range from a negligible negative effect to an enormous positive one.

That interval is the data refusing the question. If “the effect of mindfulness therapy” were a single useful quantity, the studies would cluster around it and the interval would be narrow. Instead, it spans almost the entire range of plausible effects.

And this is the point where the argument is often lost, so I want to be precise. The average of .47 is not meaningless. It is the correct answer to a narrow and legitimate question: if you drew another study at random from this same mixture of studies, .47 would be your best single guess for its effect. But patients are not looking to enter a lottery whose prize ranges from a small harm to a large benefit. They want to know whether a particular therapy has been shown to work for their problem and in a population like theirs.

There are, in fact, two problems stacked on top of each other. Even within a single population, an average conceals variation between individuals — some patients improve, some do not. I set that problem aside here because the meta-analytic average fails long before we reach it. It is already an average of different population averages, different treatments, and different disorders. It tells us little about how any particular therapy performs for any particular problem in any particular population. The within-population question is hard. This broader question may not even be well posed.

Goldberg and colleagues (2018) set out to answer exactly that question with a meta-analysis of mindfulness-based therapy. They collected studies spanning a range of disorders — depression, anxiety, substance use, and more — conducted on different populations in different countries, using different mindfulness protocols, from standardized programs like MBSR and MBCT to local adaptations, and pooled them into a single estimate. The headline was encouraging: a standardized mean difference somewhere between .46 and .73, a moderate-to-large effect, the kind of number that has become “mindfulness works” by the time it reaches a textbook or a clinician.

But look at the question that average answers — “the effect of mindfulness therapy” — and you will recognize the grammar from the previous section. It treats “mindfulness therapy,” unconditioned, as one homogeneous thing, an effect that exists and the meta-analysis merely measures. The studies it pooled were not one thing. Mindfulness for chronic pain in an American clinic and mindfulness for depression in a Chinese university are no more the same treatment of the same disorder in the same people than a newborn and an adult are the same height. The pool is heterogeneous by construction — across disorders, populations, and therapies at once. Asking for its single average effect is asking for the average height of everyone on Earth.

Hidden Moderators in Plain Sight

A meta-analyst can fairly say that I have described only half the job. Meta-analysts do not just compute an average; they also look for moderators — study features that predict when the effect is larger or smaller. Culture, dosage, type of control group, severity of the disorder: code each study on these characteristics, then test whether they track the effect sizes. This is the right instinct. If effects vary, find out what they vary with. Sometimes this works. But in psychology it often does not, and the reason is partly built into the way the search works.

A moderator analysis can only find variation associated with variables that were actually coded. You choose the study characteristics, code the studies on them, and ask which ones predict the results. That can reveal an explanation only if two things are true: you thought to measure it, and enough studies differ on it for a pattern to emerge. When the real source of variation is something no one thought to code — an unusual outcome measure, a quality problem, a researcher who strongly favored a particular result — the moderator analysis may come back empty.

Then there is variation that does not correspond to any broad study characteristic at all. Imagine making a smoothie with a dozen different fruits. Suppose it tastes off because of a single rotten blueberry — one study in forty with a broken measure, a p-hacked result, or invented data. There may be no useful moderator for that. “Rotten” is not a dimension along which the studies vary; it is a fact about one study. A moderator is a column in a spreadsheet, while a single unusual study is a row. To understand that study, eventually you have to look at the row.

Psychologists already know the shape of this problem from their own statistical tools. Factor analysis looks for variation that is shared across several measures. A strong relationship between just two variables does not ordinarily define a broad factor and may be treated as something specific to that pair. Cluster analysis asks a different question. If two variables correlate at .9, they can form a tight cluster whether or not they belong to any broader dimension. Moderator analysis resembles the factor approach: it looks for systematic variation along dimensions shared by multiple studies. It is less useful for a small pocket of studies that resemble one another for some idiosyncratic reason, and still less useful for a single unusual study. Those patterns become visible only when we stop looking exclusively at columns and start looking at rows.

In a meta-analysis of treatment effectiveness, however, the rotten blueberries are not the only studies we should be looking for. We also want the opposite — studies that provide especially strong and trustworthy evidence that the treatment works. But finding them is harder than sorting a forest plot by observed effect size. Large effects from small studies are especially vulnerable to sampling error, and extreme estimates are often extreme partly because of luck. Rank studies by the effects you happen to observe and you risk promoting the flukes.

This is where z-curve3 can help. It uses information from the distribution of results to shrink noisy study estimates toward more plausible values, correcting for regression to the mean and selective reporting. From the adjusted estimate it can compute a minimum effect size: a conservative lower-bound estimate of how large the effect could reasonably be after sampling error and uncertainty are taken into account. That makes a different kind of claim from the pooled average. The pooled mean asks for the center of the entire collection. The minimum effect size asks what can be said conservatively about one particular study.

And this brings us back to the blueberry. Moderator analysis asks which characteristics explain differences across studies. The corrected forest plot asks a different question: which individual studies provide the strongest evidence after noisy estimates have been pulled back toward more plausible values? The figure shows those studies, along with an estimate of how likely each result is to reach significance again in an exact replication of the same size. These are the promising fruits for a tasty smoothie. You find them not by blending everything together, and not only by coding broad dimensions, but by looking at the studies one at a time.

Do Western Patients Benefit from Eastern Mindfulness Therapy?

The figure shows a forest plot of the studies with the strongest evidence, sorted by their minimum effect size, from a high of 1.56 down to .41. Each study is identified by its first author and year. Look at which studies produced strong evidence of effectiveness on their own. Names like Majid, Zemestani, Kaviani, Omidi, Bakhshani, Panahi, Zhang, Chien, and Wang are Asian names, and closer inspection of the articles confirms it: these were studies of Asian participants. The strong evidence in this literature comes, overwhelmingly, from Iran and China. The pattern was sitting in the 2018 data; it took a 2021 umbrella review to note, across this body of work, that effects tend to run larger in Asian studies (Goldberg et al., 2021).

Given these results, a meta-analysis that pools all studies tells us nothing about the effectiveness of mindfulness therapy in WEIRD or in non-WEIRD samples. The average is too high for the Western patient, whose studies cluster low, and too low for the Iranian and Chinese patient, whose studies cluster high.

It may seem laudable that the meta-analysis included non-WEIRD samples. But dropping them in the blender is what created the heterogeneity that makes the average useless in the first place. The pooled number tells us nothing about either population on its own — it is an average across both that describes neither. And the fix is not complicated. Before you average a set of studies, you owe one check: do their results scatter by luck alone? If the only thing separating the estimates is sampling error, the studies were plausibly measuring one effect, and the average means something. If they scatter by more than luck — if real differences remain after chance is accounted for — then they were never one thing, and no single number should be reported for all of them.

So do Western patients benefit from Eastern mindfulness therapy? This meta-analysis cannot say. This is not a verdict on mindfulness therapy. It is a verdict on a method. There is nothing weird about studying WEIRD samples, if the question is whether mindfulness therapy helps WEIRD patients. What is weird is to mix populations, discover that the effects vary, and then report the average as if it applied to all of them.

Conclusion

Science is a process. While there are universal criteria that distinguish science from other belief systems, the universal aspect of science is to question itself and to learn from mistakes. This process can take time. Meta-analysis emerged in the 1970s to make sense of inconclusive and sometimes conflicting results in a growing literature of empirical studies. Over time, rules for meta-analyses were formulated. Nowadays, meta-analyses are often considered to be the gold standard to make sense of original studies and meta-analyses are highly cited as authoritative sources to make claims like “Mindfulness therapy works.”

Initial meta-analysis often assumed a single effect size. Over time, methods were developed to examine and quantify heterogeneity in population effect sizes. However, meta-analysts are still trying to figure out how to report heterogeneity and what to with it. This essay points out that heterogeneity in effect sizes cannot be ignored. Studies should be combined to reduce sampling error, but not to hide true variation across populations.

More broadly, psychologists need to become more comfortable to study specific populations rather than claiming that their study tests a universal hypothesis and then apologize for the fact that they studied only US Americans or another WEIRD population. Studies that do want to make universal claims (e.g., Ekman’s research on facial expression) do require cross-cultural data, but not all studies have to test universal hypotheses.

Further Readings

  • Ghai, S. (2021). “It’s time to reimagine sample diversity and retire the WEIRD dichotomy.” Nature Human Behaviour. This is probably the cleanest paper for your purpose. Ghai argues that dividing the world into WEIRD versus non-WEIRD collapses enormous heterogeneity into a binary classification. A sample from India, Nigeria, Chile, and rural China does not become meaningfully similar simply because all are “non-WEIRD.”
    Nature Human Behaviour article
  • Clancy, K. B. H., & Davis, J. L. (2019). “Soylent Is People, and WEIRD Is White: Biological Anthropology, Whiteness, and the Limits of the WEIRD.” Annual Review of Anthropology. This is a deeper conceptual critique. They argue that the individual components of WEIRD are poorly operationalized and that treating inhabitants of “WEIRD societies” as homogeneous erases substantial differences within those societies. Their broader argument is that the label can obscure the actual dimensions researchers need to measure.
    Annual Review article
  • Muthukrishna et al. (2020). “Beyond Western, Educated, Industrial, Rich, and Democratic (WEIRD) Psychology: Measuring and Mapping Scales of Cultural and Psychological Distance.” Psychological Science. This comes partly from the same intellectual tradition as the original WEIRD paper, but it implicitly identifies a major problem with the acronym: cultural variation is better conceived as multidimensional and continuous rather than as membership in two groups. They develop measures of psychological/cultural distance instead.
    Paper information and full-text links
  • Schimmelpfennig et al. (2024). “Methodological concerns underlying a lack of evidence for cultural heterogeneity in the replication of psychological effects.” Communications Psychology. This paper includes Henrich, Heine, and Norenzayan themselves. It explicitly warns against turning the letters of WEIRD into an empirical “WEIRDness” scale. Their point is important: WEIRD was originally a mnemonic/consciousness-raising device, not a theory of which cultural dimensions cause psychological variation. They criticize binary coding and mechanically decomposing countries according to the five letters because this produces classifications with poor theoretical and face validity. Open-access article
  • Jeffrey Sherman’s “There Is Nothing WEIRD About Basic Research: The Critical Role of Convenience Samples in Psychological Science” in American Psychologist (published online 2024; print 2025). Sherman accepts that psychology has a diversity problem, but challenges the inference that every study therefore requires culturally representative or highly diverse sampling. His argument is that the relevant question is what population a claim is intended to generalize to and what moderators the theory predicts. Convenience sampling can be entirely appropriate for basic research. He also stresses that “WEIRD sample” and “convenience sample” are not the same methodological problem.
  • Open manuscript copy