Abstract
This report asks whether IMPACTER’s scorer gives a different result to the same reasoning depending on who wrote or said it: by race, gender, income, native language, including Spanish, and learning differences. On the 8,612-student sample of Report 2025-01, we ran four audits across 27 subgroups. Agreement: the machine–human standardized mean difference was below .10 in every subgroup (largest −.08, students with a speech or language impairment). Spanish speakers: students whose home language is Spanish, nearly a quarter of the sample, matched expert raters (SMD −.03), and transcription errors on their audio did not reach the score. Dialect: meaning-matched responses in African American English, Chicano English, Spanish-influenced learner English, and code-switched Spanish–English received the same level as the Mainstream American English version as often as a same-dialect paraphrase did (88.8–92.5% against a 92.9% noise floor), and no pair moved two levels. Differential functioning: 0 of 162 prompt-by-group tests reached Category B. Amplification: in 18 of 18 focal groups the voice-score gap was smaller than the teacher-rating gap for the same competency (mean −.09 SD).
- Evaluated for bias?
- Yes. Race and ethnicity, gender, income, native language (including Spanish) and learning differences: 27 subgroups, four audits.
- Who conducted it?
- Independent researchers at the University of Washington and Michigan State University (§4).
- How often?
- On every model version, before it reaches students. A failed check blocks the release. The next independent evaluation runs in Q1 2027, before Model 7.