Category Archives: Ethics

Project Implicit: Insider and Outsider Perspectives

Here is an open access version of “Lessons from two decades of project implicit” by Kate A. Ratliff and Colin Tucker Smith. Microsoft Word – PI Chatper (Krosnick).docx

The chapter offers interesting insights into the history of Project Implicit by two insiders who worked for Project Implicit. This blog post provides comments on this history from the perspective of an outsider.

1. Big Sample Envy

“Nosek wanted to use the IAT in his research but was only allotted fifteen participant
hours through the Yale participant pool” (p. 98).

In most sciences, it is a blessing to be at a rich ivy league university with expensive equipment. Psychology is different because it relied mostly on undergraduate students as participants and classes at fancy ivy universities are small. This gave large state universities like Ohio State University or the University of Illinois at Urbana-Champaign. One might think, rich universities could just pay participants, but that did not appear to be the case. Thus, psychologists at the top universities often published studies with very small samples (Bargh et al., 1996), which led to the replication crisis in the 2010s (Doyen et al., 2012; Kahneman, 2012, 2017).

Project Implicit was born out of the desire to collect data with large samples.

“In the first version of the website, I set up the application to compute the scores within the app and just send a single line of data to the database– e.g., block means, errors. I could watch the file grow live with each person completing a test and their result being added to the database. It was truly mesmerizing. Watching a new line come in every few seconds compared to how laborious data collection had been before. It was some thing of a conversion experience to going all-in on on-line data collection.” (Brian Nosek, quoted in Ratliff & Smith, p. 98).

For an outsider, the statement is a clear admission that the primary purpose of Project Implicit was research and the use of online administration to get data from many people.

Ratliff and Smith further mention that the National Institute of Mental Health awarded a research grant ($2.5 million) to “further develop the virtual laboratory on the Internet” (p. 98).

False Feedback and Deception

The article also mentions the preconditions for research conducted with Project Implicit. (
(1) studies can be no longer than fifteen minutes (around ten minutes is the goal),
(2) study text should be no higher than an eighth-grade reading level
(3) studies may not include deception
(4) studies must include some kind of measure about which participants receive
feedback
(5) an appropriate debriefing that fulfills the educational mission of the
organization must be offered.

Several of these points are noteworthy from an outsider’s perspective. The short time frame makes it impossible to study causes or consequences of implicit biases experimentally. Even correlational studies that relate IAT scores to other measures may take longer. Thus, most studies are limited to the IAT scores themselves or correlations with demographic variables. This limits the usefulness of the virtual laboratory to study actual causes and consequences of implicit biases in real life. Not surprisingly, millions of people have completed an IAT, but sample sizes with actual measures of behavior are much smaller and often unable to reveal meaningful relationships (Kurdi et al., 2019).

The absence of deception and the requirement to provide feedback about IAT performance create a tension that is rarely acknowledged. One type of study in psychology deliberately gives people false feedback about a desirable trait. These studies use deception and require extensive debriefing to ensure that participants are not harmed by the false information. Project Implicit does not give blatantly false feedback, but many people will receive false feedback if a test has low validity. For example, an IQ test that correlates r = .6 with true intelligence (whatever that is) will give 20% of participants false feedback that they are below average (IQ below 100) if their true score is above average. IAT scores are much less valid than intelligence tests and even more people get false feedback. An ethical debriefing would require warning people that one possible explanation for a surprising result is measurement error, however Project Implicit has failed to provide this information. This resistance to debriefing participants properly about the low validity of IAT scores contradicts the claim that IAT research on Project Implicit should avoid deception and properly debrief participants.

The lack of proper debriefing can be explained by the insiders’ belief in implicit biases and the ability of IATs to measure them.

“When we started graduate school in 2003, few people outside of the field of social
psychology were talking about implicit bias. We earnestly explained to our friends and
family that people have attitudes and stereotypes that influence how they see and interpret
the world around them, and they might not even know it is happening. They were skep
tical. We told them about tests that help scientists uncover and quantify these biases.
They were notc onvinced. We told them to read Blink (Gladwell, 2005). A “real” author wrote
that; they started to get it. Now, of course, implicit bias is discussed everywhere– court
rooms, police departments, offices of human resources, corporate boardrooms, elementary
schools, and colleges. The idea that even “good people” may harbor unwanted attitudes and
stereotypes is commonplace, ordinary, perhaps even a bit insipid. We seem to have forgotten
that, just two decades ago, these ideas were quite radical.” (Ratliff & Smith, p. 97).

Research on the unconscious, however, shows how hard it is to study unconscious processes and that widespread beliefs in them do not mean that they exist. At one point in time, academic psychologists were attacked for questioning the validity of repressed memories and it is now widely accepted that some (not all!) of these memories were constructions of events that never happened.

Like some psychoanalysts who lashed out against scientific critics, Project Implicit insiders dismiss valid scientific criticism without engaging with the scientific arguments.

“we disagree with arguments that moderate correlations between IAT scores and self-report
suggest that the constructs are redundant (Schimmack, 2021), and thus implicit bias is
uninteresting. These and similar arguments are difficult to reconcile with many people’s surprise and even resistance when confronted with evidence of their own bias” (Ratliff & Smith, p. 112).

This response is almost comically similar to a cartoonish psychoanalyst who tells a patient that (a) “you unconsciously want to kill your father,” (b) you unconsciously want to sleep with your mother,” or (c) “you unconsciously want to have a penis.” When the patient responds that this is clearly not the case, the psychiatrists claims that they are just using defense mechanisms to deny the truth about their hidden motives.

According to Ratliff and Smith any denial of biases revealed by the IAT is a defensive response, when most of the time, it is much more likely that the IAT scores are biased. They also mischaracterize Schimmack’s evidence, which may reveal a defensive reaction of their own. Schimmack showed that a large portion of the variance in IAT scores is random and systematic measurement error. Once measurement error is statistically corrected, IAT scores and self-reports on the race IAT are highly correlated. Thus, there is no evidence that IAT scores reflect anything that could diverge from people’s self-perceptions. Moreover, their self-reported attitudes are often stronger predictors of behavior than the small amount of unique variance in IAT scores, even in studies done by IAT proponents (Axt et al., in press; Greenwald et al., 1998).

Accuracy and Ethics of Feedback

The section “Accuracy and Ethics in Providing IAT Feedback” promises to address these problems, but falls short of engaging with the low validity of IAT scores as measure of implicit biases.

“Research shows the IAT is an effective educational tool for raising awareness about implicit
bias, but the IAT cannot and should not be used for diagnostic or selection purposes (e.g., hiring or qualification decisions). For example, using the IAT to choose jurors is not justifiable, but it is appropriate to use the IAT to teach jurors about implicit bias” (Ratliff & Smith, p. 115).

What this statement leaves out is the reason why IATs should not be used for diagnostic purposes. The reason is that IAT scores have woefully inadequate validity; that is most of the variance in these scores is measurement error. So, how is it ethical to give people feedback about these scores if they are often invalid? The most revealing statement in the whole article is Ratliff and Smith’s answer to this question:

“This brings up an important question on which Project Implicit’s Scientific Advisory Board reflects frequently– is it ethical to pro vide participants feedback on their IAT performance? Thus far, the team has answered this question in the affirmative (a point to which we will return at the end of this section), but the team closely follows the literature on IAT reliability and malleability to make this decision and are open to reconsidering should the evidence suggest it is prudent to do so.”

The question is whether we can trust a team of researchers who are interested in collecting data in the virtual laboratory to make this ethical decision without conflict of interest. Maybe they should consult outsiders to avoid motivated biases that could harm people who receive false feedback without proper debriefing.

Aside from conflict of interest, a bigger problem is that the Project Implicit members have no formal training in developing, evaluating, and administering psychological tests, a discipline known as psychometrics and despite the similar name, largely removed from psychology. Even undergraduate students learn at some point that reliability is insufficient to evaluate test scores, but Ratliff and Smith never discuss validity and systematic measurement error in IAT scores.

They also confuse effect sizes for group means with scores of individuals. “The reasoning for these particular cut-offs is that, given that the standard deviations of IAT D-scores are rarely greater than 0.5 (Nosek et al., 2007), these IAT D-score cutoffs correspond approximately to Cohen’s d effect sizes of 0.3 (slight preference), 0.7 (moderate preference), and 1.3 (strong preference). These are above Cohen’s conventional cutoffs (i.e., 0.2, 0.5, 0.8), because the confidence interval around the estimate of a single score is likely to be greater than that of the confidence interval based on a sample mean. In other words, the feedback is somewhat conservative” (p. 101). This claim shows lack of knowledge about the scoring of test scores and the true amount of uncertainty around an individuals’ test score. Not surprisingly, they see no problem in providing invalid feedback based on their false assumption that the scoring is conservative.

The chapter does provide some interesting information about changes to the feedback that people are given. In the beginning, feedback claimed that IAT scores reveal unconscious biases. Ratliff and Smith emphasize that talks and educational materials no longer use the term unconscious (p. 112). Instead, “for several years now Project Implicit has used the term active awareness to reflect the fact that unawareness of implicit bias might be because one has
not reflected deeply about their biases rather than because one cannot” (p. 112).

However, there is no evidence for this claim. A search on the Project Implicit website did not retrieve any relevant hits that mention active awareness and evidence that IAT scores reflect biases that operate without active awareness. Instead, the website continues to claim that implicit biases exist without awareness.

Some outsiders might consider this double deception. The description of the way Project Implicit is presenting itself to the public is deceiving readers who do not fact check the claim and the claim “without awareness” deceives people who visit the website that the test can tell something about them that they do not already know.

Conclusion

In conclusion, Project Implicit was created as a research laboratory for short studies with the aim to get responses from a large number of people. Many other researches have surveys posted, but do not get millions of visitors to do their surveys. Project Implicit has benefited from an affiliation with Harvard that suggests to many Americans that it is solid science and from marketing the IAT as a “window into the unconscious” (Banaji & Greenwald, 2013). Criticism of the validity of the IAT has been brushed aside with the claim that “Project Implicit
gives feedback to participants about their IAT performance because of the perceived educational value in doing so.” The question remains who perceives this value. Many outsiders do not think that it is educational to give people false feedback about their unconscious. If the IAT is no different than a Rorschach test, why does it still get support from psychological science.

Fortunately, thanks to popular articles and blog posts the general public is learning more about the problems with the IAT and the concept of implicit biases (Schimmack, 2026; Singal, 2017). This blog post provides further evidence that the organization behind the online administration of the IAT lacks the scientific qualifications to do so and has put self-interest over ethics. Despite growing scientific evidence that IATs do not measure implicit biases, visitors are not given proper information about the accuracy of their feedback. Instead, resistance to the feedback is described as defensive. Ironically, the response by the scientific advisory board to criticism is a lot more defensive and less defensible than responses by people to do not believe the IAT.

Ethical Challenges for Psychological Scientists

Psychological scientists are human and like all humans they can be tempted to violate social norms (Fiske, 2015).  To help psychologists to conduct ethical research, professional organizations have developed codes of conduct (APA).  These rules are designed to help researchers to resist temptations to engage in unethical practices such as fabricate or falsify of data (Pain, Science, 2008).

Psychological science has ignored the problem of research integrity for a long time. The Association for Psychological Science (APS) still does not have formal guidelines about research misconduct (APS, 2016).

Two eminent psychologists recently edited a book with case studies that examine ethical dilemmas for psychological scientists (Sternberg & Fiske, 2015).  Unfortunately, this book lacks moral fiber and fails to discuss recent initiatives to address the lax ethical standards in psychology.

Many of the brief chapters in this book are concerned with unethical behaviors of students, in clinical settings, or ethics of conducting research with animals or human participants.  These chapters have no relevance for the current debates about improving psychological science.  Nevertheless, a few chapter do address these issues and these chapters show how little eminent psychologists are prepared to address an ethical crisis that threatens the foundation of psychological science.

Chapter 29
Desperate Data Analysis by a Desperate Job Candidate Jonathan Haidt

Pursuing a career in science is risky and getting an academic job is hard. After a two-year funded post-doc, I didn’t have a job for one year and I worked hard to get more publications.  Jonathan Haidt was in a similar situation.  He didn’t get an academic job after his first post-doc and was lucky to get a second post-doc,but he needed more publications.

He was interested in the link between feelings of disgust and moral judgments.  A common way to demonstrate causality in experimental social psychology is to use an incidental manipulation of the cause (disgust) and to show that the manipulation has an effect on a measure of the effect (moral judgments).

“I was looking for carry-over effects of disgust”

In the chapter, JH tells readers about the moral dilemma when he collected data and the data analysis showed the predicted pattern, but it was not statistically significant. This means the evidence was not strong enough to be publishable.  He carefully looked at the data and saw several outliers.  He came up with various reasons to exclude some. Many researchers have been in the same situation, but few have told their story in a book.

I knew I was doing this post hoc, and that it was wrong to do so. But I was so confident that the effect was real, and I had defensible justifications! I made a deal with myself: I would go ahead and write up the manuscript now, without the outliers, and while it was under review I would collect more data, which would allow me to get the result cleanly, including all outliers.

This account contradicts various assertions by psychological scientists that they did not know better or that questionable research practices just happen without intent. JH story is much more plausible. He needed publications to get a job. He had a promising dataset and all he was doing was eliminating a few outliers to bet an arbitrary criterion of statistical significance.  So what, if the p-value was .11 with the three cases included. The difference between p = .04 and p = .11 is not statistically significant.  Plus, he was not going to rely on these results. He would collect more data.  Surely, there was a good reason to bend the rules slightly or as Sternberg (2015) calls it going a couple of miles over the speed limit.  Everybody does it.  JH realized that his behavior was unethical, it just was not significantly unethical (Sternberg, 2015).

Decide That the Ethical Dimension Is Significant. If one observes a driver going one mile per hour over the speed limit on a highway, one is unlikely to become perturbed about the unethical behavior of the driver, especially if the driver is oneself.” (Sternberg, 2015). 

So what if JH was speeding a little bit to get an academic job. He wasn’t driving 80 miles in front of an elementary school like Diedrik Stapel, who just made up data.  But that is not how this chapter ends.  JH tells us that he never published the results of this study.

Fortunately, I ended up recruiting more participants before finishing the manuscript, and the new data showed no trend whatsoever. So I dropped the whole study and felt an enormous sense of relief. I also felt a mix of horror and shame that I had so blatantly massaged my data to make it comply with my hopes.

What vexes me about this story is that Jonathan Haidt is known for his work on morality and disgust and published a highly cited (> 2,000 citations in WebofScience) article that suggested disgust does influence moral judgments.

Wheatley and Haidt (2001) manipulated somatic markers even more directly. Highly hypnotizable participants were given the suggestion, under hypnosis, that they would feel a pang of disgust when they saw either the word take or the word often.  Participants were then asked to read and make moral judgments about six stories that were designed to elicit mild to moderate disgust, each of which contained either the word take or the word often. Participants made higher ratings of both disgust and moral condemnation about the stories containing their hypnotic disgust word. This study was designed to directly manipulate the intuitive judgment link (Link 1), and it demonstrates that artificially increasing the strength of a gut feeling increases the strength of the resulting moral judgment (Haidt, 2001, Psychological Review). 

A more detailed report of these studies was published in a few years later (Wheatley & Haidt, 2005).  Study 1 reported a significant difference between the disgust-hypnosis group and the control group, t(44) = 2.41, p = .020.  Study 2 produced a marginally significant result that was significant in a non-parametric test.

For the morality ratings, there were substantially more outliers (in both directions) than in Experiment 1 or for the other ratings in this experiment. As the paired-samples
t test loses power in the presence of outliers, we used its non-parametric analogue, the Wilcoxon signed-rank test, as well (Hollander&Wolfe, 1999). Participants judged the actions to be more morally wrong when their hypnotic word was present (M = 
73.4) than when it was absent (M = 69.6), t(62) = 1.74, p = .09, Wilcoxon Z = 2.18, p < .05. 

Although JH account of his failed study suggests he acted ethically, the same story also reveals that he did have at least one study that failed to provide support for the moral disgust hypothesis that was not mentioned in his Psychological Review article.  Disregarding an entire study that ultimately did not support a hypothesis is a questionable research practice, just as removing some outliers is (John et al., 2012; see also next section about Chapter 35).  However, JH seems to believe that he acted morally.

However, in 2015 social psychologists were well aware that hiding failed studies and other questionable practices undermine the credibility of published findings.  It is therefore particularly troubling that JH was a co-author of another article that failed to mention this study. Schnall, Haidt, Core, and Jordan (2015) responded to a meta-analysis that suggested the effect of incidental disgust on moral judgments is not reliable and that there was evidence for publication bias (e..g, not reporting the failed study JH mentions in his contribution to the book on ethical challenges).  This would have been a good opportunity to admit that some studies failed to show the effect and that these studies were not reported.  However, the response is rather different.

With failed replications on various topics getting published these days, we were pleased that Landy and Goodwin’s (2015) meta-analysis supported most of the findings we reported in Schnall, Haidt, Clore, and Jordan (2008). They focused on what Pizarro, Inbar 
and Helion (2011) had termed the amplification hypothesis of Haidt’s (2001) social intuitionist model of moral judgment, namely that “disgust amplifies moral evaluations—it makes wrong things seem even more wrong (Pizarro et al., 2011, p. 267, emphasis in original).” Like us, Landy and Goodwin (2015) found that the overall effect of incidental disgust on moral judgment is usually small or zero when ignoring relevant moderator variables.”   

Somebody needs to go back in time and correct JH’s Psychological Review article and the hypnosis studies that reported main effects with moderated effect sizes and no moderator effects.  Apparently, even JH doesn’t believe in these effects anymore in 2015 and so it was not important to mention failed studies. However, it might have been relevant to point out that the studies that did report main effects were false positives and what theoretical implications this would have.

More troubling is that the moderator effects are also not robust.  The moderator effects were shown in studies by Schnall and may be inflated by the use of questionable research practices.  In support of this interpretation of her results, a large replication study failed to replicate the results of Schnall et al.’s (2008) Study 3.  Neither the main effect of the disgust manipulation nor the interaction with the personality measure were significant (Johnson et al., 2016).

The fact that JH openly admits to hiding disconfirming evidence, while he would have considered selective deletion of outliers a moral violation, and was ashamed of even thinking about it, suggests that he does not consider hiding failed studies a violation of ethics (but see APA Guidelines, 6th edition, 2010).  This confirms Sternberg’s (2015) first observation about moral behavior.  A researcher needs to define an event as having an ethical dimension to act ethically.  As long as social psychologists do not consider hiding failed studies unethical, reported results cannot be trusted to be objective fact. Maybe it is time to teach social psychologists that hiding failed studies is a questionable research practice that violates scientific standards of research integrity.

Chapter 35
“Getting it Right” Can also be Wrong by Ronnie Janoff-Bulman 

This chapter provides the clearest introduction to the ethical dilemma that researchers face when they report the results of their research.  JB starts with a typical example that all empirical psychologists encountered.  A study showed a promising result, but a second study failed to show the desired and expected result (p > .10).  She then did what many researchers do. She changed the design of the study (a different outcome measure) and collected new data.  There is nothing wrong with trying again because there are many reasons why a study may produce an unexpected result.  However, JB also makes it clear that the article would not include the non-significant results.

“The null-result of the intermediary experiment will not be discussed or mentioned, but will be ignored and forgotten.” 

The suppression of the failed study is called a questionable research practice (John et al., 2012).  The Publication Manual of APA considers this unethical reporting of research results.

JP makes it clear that hiding failed studies undermines the credibility of published results.

“Running multiple versions of studies and ignoring the ones that “didn’t work” can have far-reaching negative effects by contributing to the false positives that pervade our field and now pass for psychological knowledge. I plead guilty.”

JP also explains why it is wrong to neglect failed studies. Running study after study to get a successful outcome, “is likely capitalize on chance, noise, or situational factors and increase the likelihood of finding a significant (but unreliable) effect.” 

This observation is by no means new. Sterling (1959) pointed out that publication bias (publishing only p-values below .05), essentially increases the risk of a false positive result from the nominal level of 5% to an actual level of 100%.  Even evidently false results will produce only significant results in the published literature if failures are not reported (Bem, 2011).

JP asked what can be done about this.  Apparently, JP is not aware of recent developments in psychological science that range from statistical tests that reveal missing studies (like an X-ray for looked file-drawers) to preregistration of studies that will be published without a significance filter.

Although utterly unlikely given current norms, reporting that we didn’t find the effect in a previous study (and describing the measures and manipulations used) would be broadly informative for the field and would benefit individual researchers conducting related studies. Certainly publication of replications by others would serve as a corrective as well.

It is not clear why publishing non-significant results is considered utterly unlikely in 2015, if the 2010 APA Publication Manual mandates publication of these studies.

Despite her pessimism about the future of Psychological Science, JP has a clear vision how psychologists could improve their science.

A major, needed shift in research and publication norms is likely to be greatly facilitated by an embrace of open access publishing, where immediate feedback, open evaluations and peer reviews, and greater communication among researchers (including replications and null results) hold the promise of opening debate and discussion of findings. Such changes would help preclude false-positive effects from becoming prematurely reified as facts; but such changes, if they are to occur, will clearly take time.

The main message of this chapter is that researchers in psychology have been trained to chase significance because obtaining statistical significance by all means was considered a form of creativity and good research (Sternberg, 2018).  Unfortunately, this is wrong. Statistical significance is only meaningful if it is obtained the right way and in an open and transparent manner.

33 Commentary to Part V Susan T. Fiske

It was surprising to read Fiske’s (2015) statement that “contrary to human nature, we as scientists should welcome humiliation, because it shows that the science is working.”

In marked contrast to this quote, Fiske has attacked psychologists who are trying to correct some of the errors in published articles as “method terrorists

I personally find both statements problematic. Nobody should welcome humiliation and nobody who points out errors in published articles is a terrorist.  Researchers should simply realize that publications in peer-reviewed journals can still contain errors and that it is part of the scientific process to correct these errors.  The biggest problem in the past seven years was not that psychologists made mistakes, but that they resisted efforts to correct them that arise from a flawed understanding of the scientific method.

36 Commentary to Part VI Susan T. Fiske

Social psychologists have justified not reporting failed study (cf. Jonathan Haidt example) by calling these studies pilot studies (Bem, 2011).  Bem pointed out that social psychologists have a lot of these pilot studies.  But a pilot study is not a study that tests the cause effect relationship. A pilot study tests either whether a manipulation is effective or whether a measure is reliable and valid.  It is simply wrong to treat studies that test the effect of a manipulation on an outcome a pilot study, if the study did not work.

“However, few of the current proposals for greater transparency recommend describingeach and every failed pilot study.”

The next statement makes it clear that Fiske conflates pilot studies with failed studies.

As noted, the reasons for failures to produce a given result are multiple, and supporting the null hypothesis is only one explanation. 

Yes, but it is one plausible explanation and not disclosing the failure renders the whole purpose of empirical hypothesis testing irrelevant (Sterling, 1959).

“Deciding when one has failed to replicate is a matter of persistence and judgment.”

No it is not. Preregister the study and if you are willing to use a significant result if you obtain it, you have to report the non-significant result if you do not. Everything else is not science and Susan Fiske seems to lack an understanding of the most basic reason for conducting an experiment.

What is an ethical scientist to do? One resolution is to treat a given result – even if it required fine-tuning to produce – as an existence proof: This result demonstrably can occur, at least under some circumstances. Over time, attempts to replicate will test generalizability.

This statement ignores that the observed pattern of results is heavily influenced by sampling error, especially in the typical between-subject design with small samples that is so popular in experimental social psychology.  A mean difference between two groups does not mean that anything happened in this study. It could just be sampling error.  But maybe the thought that most of the published results in experimental social psychology are just errors is too much to bear for somebody at the end of her career.

I have followed the replication crisis unfold over the past seven years since Bem (2011) published the eye-opening, ridiculous claims about feeling the future location of randomly displayed erotica. I cannot predict random events in the future, but I can notice trends and I do have a feeling that the future will not look kindly on those who tried to stand in the way of progress in psychological science. A new generation of psychologists is learning everyday about replication failures and how to conduct better studies.  For old people there are only two choices. Step aside or help them to learn from the mistakes of the older generation.

P.S. I think there is a connection between morality and disgust but it (mainly) goes from immoral behaviors to disgust.  So let me tell you, psychological science, Uranus stinks.