Alesina, A., Carlana, M., La Ferrara, E., & Pinotti, P. (2018). Revealing stereotypes: Evidence from immigrants in schools (NBER Working Paper No. 25333). National Bureau of Economic Research. https://doi.org/10.3386/w25333

Abstracts are short and cannot tell the whole story. A previous working paper and data in this article show that the IAT was also used to predict grading and that there was no evidence that IAT scores predicted lower grades for immigrant students.
The study
Alesina, Carlana, La Ferrara and Pinotti asked whether Italian middle-school teachers’ implicit stereotypes about immigrants, measured with an immigrant/Italian-name IAT, show up in how they grade immigrant students. They also asked whether telling teachers their IAT scores changes their grading.
The design has real strengths:
- 1,384 teachers in 102 schools took the IAT.
- Grades came from administrative records, so teachers didn’t know their grading was being observed.
- Blind-graded national tests (INVALSI) provide a performance benchmark.
- A randomized field experiment varied when teachers received their IAT feedback.
The paper circulated as NBER Working Paper 25333 (December 2018). It was published in the American Economic Review in July 2024. Between the two versions, the evidence on the central predictive-validity question changed little, but the presentation changed a lot.
The predictive-validity test
The key analysis regresses end-of-year grades on an Immigrant × Teacher-IAT interaction, controlling for INVALSI scores and teacher fixed effects. Grades are from the 2011–2016 cohorts. If the IAT measures something that drives discriminatory grading, teachers with higher scores should show a larger immigrant penalty at equal test performance.
Working paper (2018). The authors analyzed math and literature separately.
- For math, one SD higher on the IAT (0.26 D-units) was associated with a 0.033-point larger immigrant penalty (SE 0.019). This was significant only at p < .10 in every specification.
- For literature, the estimate was zero (−0.001, SE 0.017).
- The literature null was explained after the fact by lenience toward first-generation immigrants. The supporting interaction was itself not significant.
Published version (2024). The subjects are pooled, standard errors are clustered by teacher, and the IAT is entered in raw D-units.
- The interaction is −0.075 (SE 0.050), p ≈ .13. The authors state that it is not statistically significant.
- Per SD of IAT, that is about 0.02 grade points, or about 0.02 SD of grades.
- The 95% CI excludes any effect larger than about 0.04 SD per SD of IAT.
So with a smallest effect of interest of d = .10, the data are evidence of no meaningful predictive validity, not merely a lack of evidence for it.
What changed between versions
1. The main test became a footnote to subgroup analyses. After the pooled null, the AER version splits students by INVALSI score.
- For high-ability students: −0.139 (SE 0.080), p ≈ .08.
- For low-ability students: −0.031 (SE 0.060).
- No test of the difference between subgroups is reported.
Even the high-ability estimate implies at most about 0.07 SD per SD of IAT at the CI bound. This split does not appear in the working paper. Yet the abstract-level summary in the introduction now says that higher IAT scores are associated with lower grades for high-performing immigrants. The conclusion calls the pattern “strongly suggestive of bias.”
2. Measurement error was invoked, and a new measure was introduced. The authors attribute the null partly to noise in teacher-level bias estimates. In an online appendix, they show that a “naïve” teacher-level bias measure is not significantly related to the IAT. An empirical-Bayes-shrunken version of the same measure is. This alternative outcome appears only after the preregistered-looking specification fails, which is a classic forking path.
3. The measurement error that matters most is in the IAT itself. The AER version reports a new fact (footnote 6): the male-name and female-name immigrant IATs correlate only 0.28. That implies a reliability of roughly .44 for the averaged score. The authors present this as reassuring because 76% of teachers received consistent categorical feedback. But a test-retest correlation of .28 places a low ceiling on any criterion correlation. It also means many teachers were told they were “moderately” or “strongly” biased on the basis of noise.
4. The figures and text overstate the tables. The text describing Figure 5 says higher IAT is associated with significantly lower grades for immigrants. Table 3, which quantifies the same relationship, says it is not significant.
5. The framing moved from prediction to intervention. The working paper’s abstract led with the IAT–grading correlation. The AER abstract omits it and emphasizes the experiments. The working paper’s key moderator result—that the feedback effect was driven by teachers without explicit anti-immigrant views—has moved to an online appendix.
6. The IAT critique literature is acknowledged, then set aside. The AER version cites the predictive-validity critiques, including Blanton et al. (2009), Oswald et al. (2013), and Schimmack (2021). It responds that those studies used small samples and lacked real-world behavior. This study has a large sample and real-world behavior, and it reproduces the weak-to-null pattern those critiques describe.
The experiments
Field experiment. Receiving IAT feedback before grading narrowed the native–immigrant grade gap by 0.35 points (pooled ITT; SE 0.11). This is larger than the entire conditional gap the paper documents (about 0.10 points). It came from raising immigrants’ grades and lowering natives’ grades.
Did teachers with higher IAT scores respond more? The triple interaction is 0.214 (SE 0.302), which is null. In the field, the IAT predicts neither discrimination nor responsiveness to feedback. A reaction to being labeled biased explains the effect as well as correction of a bias does.
Online experiment. This was added for the AER version. Of 179 teachers who completed the survey, 146 are analyzed, and the attrition isn’t explained in the main text. Each teacher graded ten vignette tests with randomly assigned names.
- The average effect of personalized feedback versus a generic debiasing message is zero (0.017, SE 0.122).
- The headline result is a moderator effect: teachers with higher IAT scores responded more to their own feedback.
One result here does favor the IAT. In the active control group (about 74 teachers), higher IAT was associated with lower grades for immigrant-named tests (−0.426, SE 0.157). Three things should be weighed against it:
- It comes from a small sample, in an exploratory interaction model.
- The outcome is vignette grading, not real grades.
- The group had just received a debiasing message, and teachers with IAT = 0 graded immigrant names 0.42 points higher than native names.
It does not replicate in the field data with real grades and about 780 teachers.
The AEA registry IDs (AEARCTR-0003647 and -0013570) appear to postdate the respective data collections. If so, the moderator analyses were not preregistered. This is worth confirming against the registry.
Conclusion
The working paper’s main predictive-validity claim was a marginal (p < .10) association in math and a null in literature. In the published version, with subjects pooled and errors clustered by teacher, the IAT does not significantly predict discriminatory grading. The estimate is small enough to rule out even modest effects.
The publication process did not converge on this conclusion. Instead it added subgroup splits, an alternative bias measure, and a second, small online experiment. That gave the final paper a narrative of implicit bias that is “strongly suggestive” in the conclusion but not supported by the primary test.
Two of the paper’s own facts deserve more attention than they received. First, about 70% of teachers have “moderate to severe” IAT scores, yet grading discrimination is about 0.1 points. Second, the two IATs given to the same teachers on the same day correlate .28. Both are hard to reconcile with the view that the IAT measures a stable implicit bias that shapes behavior.
In real classroom grading of immigrant students, the IAT failed to predict discrimination.