Impacter PathwayTechnical Report Series · No. 2026-01 · v1.0

Heard The Same Way

A Bias Audit Of Impacter Pathway Across Race, Gender, Income, Language, And Learning Differences

Abstract

This report asks whether IMPACTER’s scorer gives a different result to the same reasoning depending on who wrote or said it: by race, gender, income, native language, including Spanish, and learning differences. On the 8,612-student sample of Report 2025-01, we ran four audits across 27 subgroups. Agreement: the machine–human standardized mean difference was below .10 in every subgroup (largest −.08, students with a speech or language impairment). Spanish speakers: students whose home language is Spanish, nearly a quarter of the sample, matched expert raters (SMD −.03), and transcription errors on their audio did not reach the score. Dialect: meaning-matched responses in African American English, Chicano English, Spanish-influenced learner English, and code-switched Spanish–English received the same level as the Mainstream American English version as often as a same-dialect paraphrase did (88.8–92.5% against a 92.9% noise floor), and no pair moved two levels. Differential functioning: 0 of 162 prompt-by-group tests reached Category B. Amplification: in 18 of 18 focal groups the voice-score gap was smaller than the teacher-rating gap for the same competency (mean −.09 SD).

Evaluated for bias?
Yes. Race and ethnicity, gender, income, native language (including Spanish) and learning differences: 27 subgroups, four audits.
Who conducted it?
Independent researchers at the University of Washington and Michigan State University (§4).
How often?
On every model version, before it reaches students. A failed check blocks the release. The next independent evaluation runs in Q1 2027, before Model 7.
−.08largest subgroup SMD
criterion |SMD| < .10
0dialect pairs moved
two or more levels
0 / 162prompt × group tests
at DIF Category B or C
18 / 18groups where voice shows
a smaller gap than teachers

1What we tested, and why

Bias in a language model does not always show up when you ask about it. It can sit in how a model responds to the way a person speaks, with the meaning held constant. A fair scorer has to be tested at that level, not only checked for overt bias.

Recent work established the method. By presenting the same content in African American English and in Mainstream American English, Hofmann et al. (2024) surfaced covert, dialect-triggered prejudice in language models that passed overt bias checks, and found that larger, preference-aligned models did not remove it. We used the same meaning-matched design on student responses, written and spoken (§3.2). Because nearly a quarter of the students in this sample speak Spanish at home, three of the four varieties we tested are ones Spanish-speaking students use: Chicano English, the English of students still learning it, and Spanish–English code-switching. We also tested individual grammatical features one at a time: the ones Pan et al. (2025) identify as driving model failure on dialect, and the Spanish-transfer constructions most common in our students’ writing (Table 4b).

Two further questions follow from the same literature. Spoken input raises the first: speech recognition makes more errors for Black speakers (Koenecke et al., 2020), so we measured whether those errors reach a student’s score (§3.3). The second is what bias should be measured against. People hold dialect prejudice too (Rickford & King, 2016), and teacher judgment is what voice scores supplement, so we asked whether voice widens or narrows the group gaps already present in teacher ratings of the same students (§3.6).

The core check is agreement with expert raters, group by group (§3.1). It is the comparison that surfaced subgroup departures in an operational essay scorer (Bridgeman, Trapani & Attali, 2012). It catches this kind of bias when it is there, which is why a clean result on it means something.

2Design

2.1 Sample and subgroups

The sample is the 8,612 students in grades 6–12 from 47 schools in 12 California districts studied in Report 2025-01, scored by the same release. Subgroup membership comes from district roster fields joined to de-identified records under each district’s data agreement. The model never sees a demographic field, and IMPACTER never infers one from a response. Home language comes from each district’s home language survey; Spanish is the largest group after English. Neurodivergence is taken from IEP primary disability category and Section 504 status. Groups under 100 students (†) are reported descriptively only. Non-binary students (‡, n = 102) enter the agreement analysis but are excluded from the multigroup models, as in Report 2025-01.

Table 1. Analytic sample. Full sample from Report 2025-01; double-scored subsample stratified to oversample smaller groups.
SubgroupStudents% of 8,612Double-scored
Race and ethnicity
Hispanic or Latino2,41228.0560
White2,23926.0540
Black or African American1,89522.0520
Asian1,72220.0461
Filipino1201.4100
Two or more races1551.8150
American Indian or Alaska Native†380.438
Native Hawaiian or Pacific Islander†310.431
Gender
Female4,20348.81,170
Male4,30750.01,128
Non-binary or another gender‡1021.2102
Socioeconomic status
Meal-eligible4,73755.01,250
Not meal-eligible3,87545.01,150
Language background
English only5,25361.01,000
English Learner1,46417.0700
Reclassified fluent (RFEP)1,55018.0500
Initially fluent (IFEP)3454.0200
Home language
Spanish2,06724.0640
English5,85668.01,400
Another language6898.0360
Neurodivergence
IEP · autism2072.4180
IEP · specific learning disability4224.9300
IEP · other health impairment (incl. ADHD)2152.5180
IEP · speech or language impairment1121.3112
IEP · all other categories1291.5129
Section 504 plan3624.2250
No IEP or 5047,16583.21,249

2.2 Four audits

Audit 1, human–machine agreement. Expert raters (inter-rater κ = .87; /iaa) scored the stratified subsample blind to demographics and to the machine score. For each subgroup we computed quadratic weighted kappa between raters and between rater and machine, and the standardized mean difference of machine minus human (Williamson, Xi & Breyer, 2012). This asks whether the machine departs from expert judgment more for some students than for others.

Audit 2, matched-guise dialect test. Following Hofmann et al. (2024), we drew 240 responses spanning every rubric level. Linguists who are fluent speakers of each variety rendered each one into African American English, Chicano English (Fought, 2003), the Spanish-influenced English of students learning English, and Spanish–English code-switched English, keeping every rubric-relevant idea. A second linguist confirmed meaning equivalence. Because any rewrite changes some wording, we also built a control: a meaning-preserving paraphrase in the same dialect. The control sets the noise floor, so a dialect effect is whatever the guises show beyond it. A spoken subset of 60 pairs per guise was recorded by fluent speakers and run through the full audio pipeline.

Audit 3, transcription. On 1,170 audio responses with human verbatim transcripts, we measured word error rate by subgroup. We then scored each response twice, from the ASR transcript and from the human transcript, so any transcription disparity that reached a student’s level would appear as an SMD (Koenecke et al., 2020).

Audit 4, differential functioning and invariance. DIF asks whether students with the same standing on a competency get different scores on a prompt because of group membership. The matching variable decides whether the test can see anything. Matching on the voice score would hide any bias common to every voice prompt, so we matched on a latent composite of the teacher, scenario, and self-report measures from Report 2025-01, which share no instrument with voice. We used ordinal logistic regression (Zumbo, 1999) on the 9 prompts for 18 focal groups, with Benjamini–Hochberg control (1995). We then extended the multigroup CFA of Report 2025-01 to the new dimensions.

Table 2. Criteria fixed before analysis.
CheckStatisticCriterionSource
Human–machine, each subgroupSMD, machine − human|SMD| < .10Williamson et al., 2012
Human–machine, overallSMD|SMD| < .15Williamson et al., 2012
Agreement floorQWKh–m≥ .70Williamson et al., 2012
Agreement degradationQWKh–h − QWKh–m< .10Williamson et al., 2012
Dialect guisestandardized shift vs MAE|Δ| < .10, and no pair ≥ 2 levelsHofmann et al., 2024 design
Transcription to scoreSMD, ASR vs human transcript|SMD| < .10Koenecke et al., 2020
Differential functioningordinal logistic ΔR²< .035 (Category A)Jodoin & Gierl, 2001
InvarianceΔCFI / ΔRMSEA≥ −.01 / ≤ .015Cheung & Rensvold, 2002; Chen, 2007

2.3 What the model is asked to do

Covert dialect prejudice has been documented in general-purpose language models asked to judge the speaker: what they are like, what job they should hold, whether they are guilty (Hofmann et al., 2024). IMPACTER is not generative AI. It writes nothing, and it is never asked about the speaker. It is supervised machine learning: a DeBERTa-v3-base encoder (He et al., 2021) with an ordinal head, trained on responses that expert human raters scored against the rubric and on synthetic examples spanning every rubric level. It reads a response and returns a rubric level for the reasoning in it: what happened, what the student did, and what they learned.

That makes human agreement the definition of accuracy, not a side metric. Raters are trained until their own agreement reaches κ = .87, and the model is judged by quadratic weighted kappa against them, a measure of how closely its levels track expert human judgment. A model built this way can be no fairer than the raters it learns from, and it can only become less fair by matching them better for some students than for others. Audit 1 tests exactly that, and the dialect test checks whether the raters’ own judgments, as learned, carry a dialect penalty.

The scorer is a 184M-parameter encoder with no human-preference alignment layer. That matters because alignment has been shown to mask covert prejudice rather than remove it (Hofmann et al., 2024); here there is no such layer, so what the audits measure is what the model does. The narrower task is a reason to expect less bias, not proof of it. Pretraining associations live in the encoder whatever head sits on top. The audits test the expectation.

3Results

3.1 Agreement with expert raters

Overall, QWKh–m = .85 against QWKh–h = .87, and overall SMD = −.01. Every subgroup met all three criteria (Table 3). The largest departure was for students with a speech or language impairment (−.08), followed by English Learners and students with autism (−.06 each). All three are negative, and all three are groups whose spoken narration differs from the norm. The direction matters: where the machine departs from raters, it scores those students slightly below what raters gave them. We report it as the place to watch, and §5 makes it a monitored threshold.

Figure 1. Machine minus human SMD by subgroup. Dashed lines mark the ±.10 criterion.
−.100+.10within criterion, |SMD| < .10RACE AND ETHNICITYHispanic or LatinoHispanic or Latino: SMD −.03WhiteWhite: SMD +.02Black or African AmericanBlack or African American: SMD −.05AsianAsian: SMD +.03FilipinoFilipino: SMD +.01Two or more racesTwo or more races: SMD −.01American Indian or Alaska NativeAmerican Indian or Alaska Native: SMD −.07Native Hawaiian or Pacific IslanderNative Hawaiian or Pacific Islander: SMD −.04GENDERFemaleFemale: SMD +.02MaleMale: SMD −.03Non-binary or another genderNon-binary or another gender: SMD +.06SOCIOECONOMIC STATUSMeal-eligibleMeal-eligible: SMD −.02Not meal-eligibleNot meal-eligible: SMD +.02LANGUAGE BACKGROUNDEnglish onlyEnglish only: SMD +.01English LearnerEnglish Learner: SMD −.06Reclassified fluent (RFEP)Reclassified fluent (RFEP): SMD −.01Initially fluent (IFEP)Initially fluent (IFEP): SMD +.02HOME LANGUAGESpanishSpanish: SMD −.03EnglishEnglish: SMD +.01Another languageAnother language: SMD −.04NEURODIVERGENCEIEP · autismIEP · autism: SMD −.06IEP · specific learning disabilityIEP · specific learning disability: SMD −.04IEP · other health impairment (incl. ADHD)IEP · other health impairment (incl. ADHD): SMD −.05IEP · speech or language impairmentIEP · speech or language impairment: SMD −.08IEP · all other categoriesIEP · all other categories: SMD −.05Section 504 planSection 504 plan: SMD −.02No IEP or 504No IEP or 504: SMD +.01
Table 3. Human–machine agreement by subgroup. SMD is machine minus human in pooled-SD units; positive means the machine scored the group higher than expert raters did.
SubgroupnQWK h–hQWK h–mDegradationSMDMeets all
Race and ethnicity
Hispanic or Latino560.87.85.02−.03✓
White540.88.86.02+.02✓
Black or African American520.86.83.03−.05✓
Asian461.87.86.01+.03✓
Filipino100.86.84.02+.01✓
Two or more races150.87.85.02−.01✓
American Indian or Alaska Native†38.84.80.04−.07✓
Native Hawaiian or Pacific Islander†31.85.82.03−.04✓
Gender
Female1,170.87.86.01+.02✓
Male1,128.87.85.02−.03✓
Non-binary or another gender‡102.85.82.03+.06✓
Socioeconomic status
Meal-eligible1,250.87.85.02−.02✓
Not meal-eligible1,150.88.86.02+.02✓
Language background
English only1,000.88.86.02+.01✓
English Learner700.85.82.03−.06✓
Reclassified fluent (RFEP)500.87.85.02−.01✓
Initially fluent (IFEP)200.87.86.01+.02✓
Home language
Spanish640.87.85.02−.03✓
English1,400.88.86.02+.01✓
Another language360.86.84.02−.04✓
Neurodivergence
IEP · autism180.85.82.03−.06✓
IEP · specific learning disability300.86.83.03−.04✓
IEP · other health impairment (incl. ADHD)180.85.82.03−.05✓
IEP · speech or language impairment112.84.80.04−.08✓
IEP · all other categories129.85.82.03−.05✓
Section 504 plan250.87.85.02−.02✓
No IEP or 5041,249.88.86.02+.01✓

3.2 The dialect test

If the encoder carried covert dialect associations (Hofmann et al., 2024) into rubric judgments, the African American English, Chicano English and Spanish-influenced guises would score lower than the MAE version of the same reasoning. Against the control’s 92.9% same-level rate, the written guises held the same level 91.2%, 92.5%, 90.4% and 88.8% of the time. The guise closest to a Spanish-speaking English learner’s own writing moved no more than the others, and no Spanish-transfer feature on its own moved scores beyond the noise floor (Table 4b). Downward moves outnumbered upward ones in every guise, which is the pattern to report rather than average away, but every standardized shift stayed inside |Δ| < .10 and no pair moved two levels. Spoken guises, which pass through transcription first, moved slightly more, as Audit 3 predicts.

Table 4. Matched-guise dialect test. Each source response was rendered in Mainstream American English (MAE) and in each guise with meaning held constant; the outcome is the change in level from the MAE version.
GuisePairsSame level %1 level lower1 level higher≥ 2 levelsStandardized shift
Control · MAE re-paraphrase noise floor24092.9890.00
African American English24091.21470−.03
Chicano English24092.51170−.02
Spanish-influenced learner English24090.41670−.04
Spanish–English code-switching24088.81980−.05
Spoken · African American English through ASR6088.3520−.05
Spoken · Chicano English through ASR6088.3520−.05
Spoken · Spanish-influenced learner English through ASR6086.7620−.05
Spoken · code-switching through ASR6085.0630−.05
Table 4b. Feature-level guises. One feature inserted at a time into otherwise unchanged MAE responses: the African American English constructions Pan et al. (2025) identify as driving model failure on dialect, and three Spanish-transfer constructions common in the English of Spanish-speaking students.
Feature insertedPairsSame level %1 lower1 higherStandardized shift
Control · same-dialect paraphrase12092.545+.01
Zero copula (“she my friend”)12092.554−.01
Existential it (“it’s a lot of people there”)12093.344.00
Y’all12094.234+.01
Habitual be (“he be helping”)12090.874−.03
Negative concord (“didn’t nobody help”)12090.084−.04
Dropped subject (“is very fun”)12092.554−.01
Possessive with of (“the house of my friend”)12093.344.00
Adjective after noun (“the car red”)12091.764−.02

3.3 What reaches the scorer

A spoken response is converted to text once, at the start. From then on the scorer has no audio. It has no acoustic model, no voiceprint, and no measure of pronunciation, accent, fluency or tone. It does not grade vocabulary or grammar either. It reads the architecture of the reasoning: temporal sequence, causal links, and connection to others. This is the relevant point in Koenecke et al. (2020), who found commercial speech recognition made nearly twice as many word errors for Black speakers and traced the gap to acoustic models. The scorer contains no acoustic model.

The conversion step itself is not exempt, so we measured it. Word error is higher for Black students (10.4% against 6.1% for White students), and higher for students whose home language is Spanish (9.1%), higher still for English Learners and for students with a speech or language impairment. The question that matters is whether those errors changed anyone’s level. Scoring the same audio from the machine transcript and from a human verbatim transcript, the difference stayed at or under −.06 SD in every group. Misheard words rarely break the structure being read. Typed responses remain available to any student.

Table 5. Transcription accuracy and whether it reaches the score. Last column: SMD between scores from the ASR transcript and from a human verbatim transcript of the same audio.
SubgroupAudio nWER %Gap vs White (pts)Score SMD, ASR vs human transcript
White1806.1ref−.01
Black or African American15010.4+4.3−.03
Hispanic or Latino2208.2+2.1−.02
Asian1207.0+0.9−.01
Filipino1007.4+1.3−.01
Home language Spanish2109.1+3.0−.03
English Learner20011.9+5.8−.04
Reclassified fluent (RFEP)1208.0+1.9−.02
IEP · speech or language impairment8014.6+8.5−.06

3.4 Differential functioning

Matched on the external criterion, no prompt showed moderate or large DIF for any focal group. The largest effect, .027 for students with a speech or language impairment, is still Category A. The ordering repeats Table 3: the groups whose narration differs most sit closest to the line, while staying under it.

Table 6. Differential functioning by focal group: ordinal logistic regression, matched on the external non-voice criterion. Category A < .035; B .035–.070; C > .070.
GroupPromptsMax ΔR²B flagsC flags
Race and ethnicity
Hispanic or Latino9.00600
Whitereference group
Black or African American9.00900
Asian9.00400
Filipino9.00500
Two or more races9.00400
American Indian or Alaska Native†below inferential floor; not tested
Native Hawaiian or Pacific Islander†below inferential floor; not tested
Gender
Female9.00700
Malereference group
Non-binary or another gender‡below inferential floor; not tested
Socioeconomic status
Meal-eligible9.00500
Not meal-eligiblereference group
Language background
English onlyreference group
English Learner9.01400
Reclassified fluent (RFEP)9.00600
Initially fluent (IFEP)9.00300
Home language
Spanish9.00800
Englishreference group
Another language9.01000
Neurodivergence
IEP · autism9.01900
IEP · specific learning disability9.01200
IEP · other health impairment (incl. ADHD)9.01600
IEP · speech or language impairment9.02700
IEP · all other categories9.01300
Section 504 plan9.00600
No IEP or 504reference group

3.5 Measurement invariance

Scalar invariance, already shown for grade band, gender and English Learner status, held across race and ethnicity, home language (Spanish vs English), meal eligibility, and IEP/504 status. The same structure, loadings and intercepts operate in every group, which is what licenses comparing group means at all.

Table 7. Measurement invariance of the retained CT-CM model, extending Report 2025-01 Table 9. Criterion ΔCFI ≥ −.01.
Levelχ²dfCFIRMSEAΔCFIHeld
Race and ethnicity · six groups
Configural1,321.4270.956.031refref
Metric1,402.8315.954.030−.002Yes
Scalar1,511.6360.951.030−.003Yes
Meal eligibility · eligible vs not
Configural429.690.960.029refref
Metric446.199.958.028−.002Yes
Scalar470.9108.956.028−.002Yes
Home language · Spanish vs English
Configural438.290.958.029refref
Metric455.999.956.029−.002Yes
Scalar481.3108.953.028−.003Yes
Neurodivergence · IEP or 504 vs neither
Configural455.390.955.031refref
Metric476.099.952.030−.003Yes
Scalar505.2108.948.030−.004Yes

3.6 Does voice amplify existing gaps?

Group differences in a measured skill can be real, and the Standards separate such impact from bias (AERA, APA & NCME, 2014, ch. 3). The test that answers the word amplify is comparative. Human judgment is the baseline, and it is not neutral: listeners penalize dialect too (Rickford & King, 2016), and in Report 2025-01 teacher ratings carried five times the method variance of voice. If a gap is larger on voice scores than on teacher ratings of the same competency, voice is adding it. In every focal group the voice gap was the smaller one, by a mean of −.09 SD. For students whose home language is Spanish, teacher ratings showed a gap of −.18 and voice −.07. The largest reductions were for students with IEPs and for English Learners, the students whose teachers’ impressions are most likely to be shaped by how they talk.

Figure 2. For each focal group, the teacher-rating gap (blue) and the voice-score gap (coral) for the same students and competencies.
teacher rating gapvoice gap−.30−.20−.100+.10+.20Hispanic or LatinoHispanic or Latino: teacher −.14Hispanic or Latino: voice −.06AsianAsian: teacher +.12Asian: voice +.09FilipinoFilipino: teacher +.05Filipino: voice +.03Black or African AmericanBlack or African American: teacher −.21Black or African American: voice −.08Two or more racesTwo or more races: teacher −.05Two or more races: voice −.02FemaleFemale: teacher +.24Female: voice +.11Meal-eligibleMeal-eligible: teacher −.19Meal-eligible: voice −.10English LearnerEnglish Learner: teacher −.27English Learner: voice −.14Reclassified fluent (RFEP)Reclassified fluent (RFEP): teacher −.03Reclassified fluent (RFEP): voice +.02Initially fluent (IFEP)Initially fluent (IFEP): teacher +.06Initially fluent (IFEP): voice +.05SpanishSpanish: teacher −.18Spanish: voice −.07Another languageAnother language: teacher −.12Another language: voice −.04IEP · autismIEP · autism: teacher −.31IEP · autism: voice −.18IEP · specific learning disabilityIEP · specific learning disability: teacher −.29IEP · specific learning disability: voice −.12IEP · other health impairment (incl. ADHD)IEP · other health impairment (incl. ADHD): teacher −.33IEP · other health impairment (incl. ADHD): voice −.15IEP · speech or language impairmentIEP · speech or language impairment: teacher −.22IEP · speech or language impairment: voice −.11IEP · all other categoriesIEP · all other categories: teacher −.25IEP · all other categories: voice −.13Section 504 planSection 504 plan: teacher −.10Section 504 plan: voice −.04
Table 8. Group gaps by method, in latent-SD units (focal minus reference), averaged over the three competencies. Amplification = |voice gap| − |teacher gap|; negative means voice shows the smaller gap.
Focal groupVoiceTeacher ratingSelf-reportAmplification
Race and ethnicity
Hispanic or Latino−.06−.14−.03−.08
Black or African American−.08−.21−.02−.13
Asian+.09+.12+.04−.03
Filipino+.03+.05+.01−.02
Two or more races−.02−.05.00−.03
Gender
Female+.11+.24+.06−.13
Socioeconomic status
Meal-eligible−.10−.19−.04−.09
Language background
English Learner−.14−.27−.05−.13
Reclassified fluent (RFEP)+.02−.03+.01−.01
Initially fluent (IFEP)+.05+.06+.02−.01
Home language
Spanish−.07−.18−.03−.11
Another language−.04−.12−.02−.08
Neurodivergence
IEP · autism−.18−.31−.09−.13
IEP · specific learning disability−.12−.29−.06−.17
IEP · other health impairment (incl. ADHD)−.15−.33−.08−.18
IEP · speech or language impairment−.11−.22−.03−.11
IEP · all other categories−.13−.25−.05−.12
Section 504 plan−.04−.10−.03−.06

4Who ran it, and how often

  • Independent evaluation. The study was conducted independently by researchers at the University of Washington (Jerrod Robert, Ph.D.; Sungmin Pak, Ph.D.) and Michigan State University (Derek Seji, Ph.D.; Michelle Greenblat, Ph.D.). IMPACTER provided de-identified data and model documentation.
  • Every model version. The full evaluation, including the dialect test, repeats on every model version before it reaches students. A failed criterion in Table 2 blocks the release. The next independent evaluation runs in Q1 2027, before Model 7 reaches students, and this report will be reissued with its results.
  • On request. Any partner district can have Tables 3, 4 and 8 computed for its own students, including its Spanish-speaking students, reported only to the district.

5Limits

Grades. The sample is grades 6–12. Elementary students, including districts serving TK–8, are not represented. Neurodivergence is a records proxy that misses undiagnosed students. Dialects: four varieties were tested, three of them used by Spanish-speaking students; Vietnamese- and Tagalog-influenced English, Hawaiian Creole, and Indigenous Mexican languages were not. Responses written entirely in Spanish are outside this report. Small groups: American Indian or Alaska Native and Native Hawaiian or Pacific Islander students fell below the inferential floor. For them, no flag means too little data to test, not evidence of fairness. Watch list: speech or language impairment, English Learners and autism sit closest to every criterion. They get a tighter internal threshold (|SMD| < .08) at each release.

·Acknowledgments

D. Hwang, A. Michelson, A. Le and D. Siddiqi of Impacter Education Corporation provided de-identified data and model documentation for this evaluation.

·References

  1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. AERA.
  2. Benjamini, Y., & Hochberg, Y. (1995). Controlling the false discovery rate. Journal of the Royal Statistical Society B, 57(1), 289–300. doi:10.1111/j.2517-6161.1995.tb02031.x
  3. Bridgeman, B., Trapani, C., & Attali, Y. (2012). Comparison of human and machine scoring of essays: Differences by gender, ethnicity, and country. Applied Measurement in Education, 25(1), 27–40. doi:10.1080/08957347.2012.635502
  4. Chen, F. F. (2007). Sensitivity of goodness of fit indexes to lack of measurement invariance. Structural Equation Modeling, 14(3), 464–504. doi:10.1080/10705510701301834
  5. Cheung, G. W., & Rensvold, R. B. (2002). Evaluating goodness-of-fit indexes for testing measurement invariance. Structural Equation Modeling, 9(2), 233–255. doi:10.1207/S15328007SEM0902_5
  6. Fought, C. (2003). Chicano English in context. Palgrave Macmillan.
  7. He, P., Gao, J., & Chen, W. (2021). DeBERTaV3: Improving DeBERTa using ELECTRA-style pre-training with gradient-disentangled embedding sharing. arXiv:2111.09543.
  8. Hofmann, V., Kalluri, P. R., Jurafsky, D., & King, S. (2024). AI generates covertly racist decisions about people based on their dialect. Nature, 633, 147–154. doi:10.1038/s41586-024-07856-5
  9. Hwang, D., Michelson, A., Le, A., & Siddiqi, D. (2025). Multitrait–multimethod validation of human skills measurement via authentic student voice (Technical Report 2025-01). IMPACTER Education Corporation.
  10. Jodoin, M. G., & Gierl, M. J. (2001). Evaluating Type I error and power rates using an effect size measure with the logistic regression procedure for DIF detection. Applied Measurement in Education, 14(4), 329–349. doi:10.1207/S15324818AME1404_2
  11. Johnson, M. S., Liu, X., & McCaffrey, D. F. (2022). Psychometric methods to evaluate measurement and algorithmic bias in automated scoring. Journal of Educational Measurement, 59(3), 338–361. doi:10.1111/jedm.12335
  12. Koenecke, A., Nam, A., Lake, E., Nudell, J., Quartey, M., Mengesha, Z., Toups, C., Rickford, J. R., Jurafsky, D., & Goel, S. (2020). Racial disparities in automated speech recognition. PNAS, 117(14), 7684–7689. doi:10.1073/pnas.1915768117
  13. Loukina, A., Madnani, N., & Zechner, K. (2019). The many dimensions of algorithmic fairness in educational applications. Proceedings of the 14th Workshop on Innovative Use of NLP for Building Educational Applications, 1–10. doi:10.18653/v1/W19-4401
  14. Pan, E., Choi, A. S. G., ter Hoeve, M., Seto, S., & Koenecke, A. (2025). Analyzing dialectical biases in LLMs for knowledge and reasoning benchmarks. Findings of EMNLP 2025. arXiv:2510.00962.
  15. Rickford, J. R., & King, S. (2016). Language and linguistics on trial: Hearing Rachel Jeantel (and other vernacular speakers) in the courtroom and beyond. Language, 92(4), 948–988. doi:10.1353/lan.2016.0078
  16. Williamson, D. M., Xi, X., & Breyer, F. J. (2012). A framework for evaluation and use of automated scoring. Educational Measurement: Issues and Practice, 31(1), 2–13. doi:10.1111/j.1745-3992.2011.00223.x
  17. Zumbo, B. D. (1999). A handbook on the theory and methods of differential item functioning (DIF). Department of National Defense, Ottawa.

Report 2026-01 · v1.0 · Sample of Report 2025-01 · Scoring release DeBERTa-v3-base, CORAL ordinal head · Independent evaluation by researchers at the University of Washington and Michigan State University · Companion: Technical Report 2025-01, Model Card v6.0 · info@impacterpathway.com