Impacter Pathway
Human Skills Analytics
Every voice carries
measurable skills
Technical research paperNo. 2025-01

Multitrait–Multimethod Validation of Human Skills Measurement via Authentic Student Voice

A construct validity study of rubric-aligned machine learning applied to open-ended student language.

The multitrait-multimethod design Three competencies, Grit, Curiosity and Self-Regulation, each measured by four independent methods: IMPACTER voice analysis, teacher rubric ratings, coded performance scenarios and validated self-report scales. Twelve measures in a fully crossed matrix. Vertical links mark same-competency, different-method correlations, which averaged .72. Horizontal links mark same-method, different-competency correlations, which averaged .31. GRIT CURIOSITY SELF-REGULATION IMPACTER VOICE ANALYSIS TEACHER RUBRIC RATINGS PERFORMANCE SCENARIOS VALIDATED SELF-REPORT SAME COMPETENCY, DIFFERENT METHOD r = .72 SAME METHOD, DIFFERENT COMPETENCY r = .31
8,612
Students
Grades 6–12
47
Schools
12 California districts
12
Measures
3 competencies × 4 methods
.72 / .31
Convergent
against discriminant
Authors
J. Arnold · D. Byun · A. Michelson · A. Le
Collection
Spring 2025 · four-week window
Scoring release
DeBERTa-v3-base · CORAL ordinal head · pre-v6.0
Document
Technical report, version 1.0 · 19 pages
Series
Correspondence
Impacter Pathway
Technical Report 2025-01 · Multitrait-Multimethod Construct Validity Impacter Education Corporation · version 1.0 · data collected spring 2025
Contents
Contents

What is in this report

Nine sections and nine numbered tables. Every figure quoted in the summary is traceable to a table, every table carries its own note, and §7 states what the study does not establish.

Tables

At a glance

Report number
2025-01 · version 1.0 · supersedes none
Design
Multitrait-multimethod, fully crossed, 3 × 4
Sample
8,612 students · 47 schools · 12 California districts
Scoring release
DeBERTa-v3-base · CORAL ordinal head · pre-v6.0
Companion documents
Supported uses
Formative feedback, progress monitoring, program evaluation, group-level monitoring
Impacter Pathway
Technical Report 2025-01 · Multitrait-Multimethod Construct Validity Impacter Education Corporation · version 1.0 · data collected spring 2025
Section§1
§1

Summary

Purpose of this report

This report documents one construct validity study for internal use, partner review, and district technical review. It is written to be checked: every figure is traceable to a numbered table, every table states its own note, and §7 lists what the study does not establish. It is not a marketing summary, and it is not a substitute for the IMPACTER Model Card, which documents scoring performance rather than construct validity.

Question. When a machine learning system scores a student's spoken or written reflection, is it measuring the competency it claims to measure, or is it measuring how articulately the student speaks?

This is the central threat to validity for any language-based assessment, and no amount of agreement with human raters answers it. A model can reproduce trained raters almost perfectly and still be rewarding verbal fluency, because the human raters may have been rewarding verbal fluency too. Resolving it requires evidence from outside the scoring relationship.

Design. We used a multitrait–multimethod (MTMM) design (Campbell & Fiske, 1959), the standard approach to this problem. Three competencies were each measured by four independent methods, producing twelve measures in a fully crossed matrix. If the scores carry real construct signal, measures of the same competency by different methods should agree strongly, while measures of different competencies by the same method should agree only moderately. If the scores mostly carry method signal (fluency, halo, response style), the pattern inverts.

The three competencies were Grit, Curiosity, and Self-Regulation, chosen because theory predicts they are related but distinguishable, which makes discriminant validity a real test rather than a formality. The four methods were IMPACTER voice scoring, teacher rubric ratings, coded performance scenarios, and short-form published self-report scales.

Findings. Same-competency, different-method correlations averaged r = .72. Different-competency correlations averaged r = .31. A correlated-trait, correlated-method confirmatory factor model fit the data well and attributed the large majority of score variance to the competency rather than to the method used to measure it. Measurement invariance held across grade band, gender, and English Learner status.

Convergent · same trait
.72
Mean correlation across methods measuring the same competency. 95% CI [.68, .76].
Discriminant · across traits
.31
Mean correlation between different competencies. 95% CI [.27, .34].
Model fit
.962
CFI for the correlated-trait, correlated-method model. RMSEA = .041.
Variance from method
6%
Mean method variance across all twelve measures. Trait variance: 61%.
Impacter Pathway
Technical Report 2025-01 · Multitrait-Multimethod Construct Validity Impacter Education Corporation · version 1.0 · data collected spring 2025
Section§1 – 2
§1 Summary, continued

Interpretation, stated conservatively. The eloquent-talker hypothesis predicts high method variance for the voice measures specifically: students who speak well scoring high on every competency. That is not what the data show. Voice measures carried less method variance (mean 6%) than teacher ratings (mean 12%), and voice measures of different competencies correlated less with each other (.33–.42) than each correlated with an independent measure of the same competency (.71–.76). The system differentiates between competencies rather than scoring articulateness.

What this study does not establish: that scores predict behavior or outcomes outside the assessment; that scores are stable over time; or that the instrument is appropriate for high-stakes individual decisions. Those require separate evidence, and §7 states what is outstanding.

§2

Why this study was run

IMPACTER's production scoring suite reports agreement with trained human raters (quadratic weighted kappa, exact accuracy, adjacent accuracy), evaluated against the thresholds used for automated scoring of constructed response (Williamson, Xi, & Breyer, 2012). Those figures establish something real and narrow: the engine reproduces the judgments of trained raters applying a published rubric.

They do not establish construct validity, and it is worth being precise about why. Agreement is a relationship between two raters. If both raters are influenced by the same irrelevant feature, such as length, vocabulary, or fluency, they will agree with each other while both being wrong about the competency. High agreement is therefore consistent with both a valid instrument and an invalid one. The literature on automated scoring has been explicit on this point: criterion agreement is necessary and insufficient (Yan, Rupp, & Foltz, 2020).

Construct validity requires evidence that the assessment elicits and captures the psychological processes theoretically associated with the construct (Messick, 1995; AERA, APA, & NCME, 2014). The MTMM design supplies that evidence by introducing methods that share no machinery with the scoring engine: a teacher watching a student over a semester, a coder reading a vignette response, a student answering Likert items. Those methods have their own biases, but they are different biases. Where independent methods with non-overlapping error converge, the shared signal is the construct.

The claim being tested

Not "does the model agree with raters." It does. The claim under test is: does the variance in an IMPACTER score come from the student's competency, or from the fact that the competency was measured by asking the student to talk? That question has a numeric answer, and §5.4 reports it.

Impacter Pathway
Technical Report 2025-01 · Multitrait-Multimethod Construct Validity Impacter Education Corporation · version 1.0 · data collected spring 2025
Section§3
§3

Constructs and predicted relationships

MTMM evidence is only interpretable against stated predictions. Vague constructs produce uninterpretable matrices, because any correlation can be explained after the fact. Each competency below was defined operationally before data collection, with the observable language features the rubric keys on.

3.1 Grit

Sustained effort and interest toward long-term goals despite setbacks, plateaus, or diminished interest (Duckworth et al., 2007). Distinguished from conscientiousness, which emphasizes organization and reliability, and from resilience, which emphasizes recovery from adversity. Observable language features: temporal markers of extended effort; setback-then-persistence sequences; goal-directed language accepting delayed reward; reframing of difficulty as developmental.

3.2 Curiosity

Intrinsic motivation to explore, question, and resolve gaps in understanding, particularly under novelty or ambiguity (Kashdan et al., 2009). We target epistemic rather than diversive curiosity, the drive to close an information gap, not attraction to novel stimuli (Litman, 2008). Observable language features: interrogative framing and self-generated questions; causal connectives and explanation-seeking; exploration beyond task requirements; intrinsic-interest markers.

3.3 Self-Regulation

Capacity to monitor and modulate thought, emotion, and behavior in service of a goal (Zimmerman, 2002). Distinguished from self-control, which is primarily inhibitory, by its metacognitive component, reflecting on one's own mental state and deploying strategy accordingly (Duckworth & Carlson, 2013). Observable language features: metacognitive verbs; emotion labeling paired with behavioral response; explicit strategy deployment; monitoring-and-adjustment language.

3.4 What the theory predicts

All three competencies support goal-directed learning and share developmental precursors in early self-control, so they should correlate positively. But they diverge in mechanism: grit emphasizes stamina where curiosity may fade once a question is answered; curiosity motivates exploration where self-regulation often requires constraining it; self-regulation involves inhibition where grit involves sustained activation. The prediction is therefore moderate positive inter-trait correlation: meaningfully above zero, meaningfully below the same-trait correlations.

This prediction is what makes the design falsifiable. Uniformly high correlations across all three competencies within the voice method would indicate the model was scoring a single general quality of response. Near-zero inter-trait correlations would indicate the constructs had been defined into artificial independence.

Impacter Pathway
Technical Report 2025-01 · Multitrait-Multimethod Construct Validity Impacter Education Corporation · version 1.0 · data collected spring 2025
Section§4
§4

Method

4.1 Participants

Data were collected from 8,612 students in grades 6–12, across 47 schools in 12 California districts, during spring 2025. Districts spanned urban (San Diego Unified, Vista Unified), border (Calexico Unified, 98% Hispanic/Latinx, 45% English Learner), suburban, and rural contexts. All participating districts had existing Portrait of a Graduate or career technical education initiatives, so the assessment context was the one in which the instrument is actually used.

Table 1Sample characteristics (N = 8,612)
Characteristicn%
Grade level
Grade 61,46517.0
Grade 71,37816.0
Grade 81,32615.4
Grade 91,55118.0
Grade 101,20714.0
Grade 1198111.4
Grade 127048.2
Language status
English Learner1,46417.0
Fluent English Proficient1,55018.0
English Only5,59865.0
Socioeconomic indicator
Free / reduced lunch4,73755.0
Not FRL3,87545.0

Note. Full demographic detail, including gender and race/ethnicity distributions, is reported in Appendix A. Districts were not randomly sampled; they self-selected into IMPACTER partnerships. This is addressed as a limitation in §7.1.

Impacter Pathway
Technical Report 2025-01 · Multitrait-Multimethod Construct Validity Impacter Education Corporation · version 1.0 · data collected spring 2025
Section§4
§4 Method, continued

4.2 The four methods

Method 1 · IMPACTER voice scoring

Students responded to three open-ended reflection prompts per competency by voice, delivered on tablet or computer. Prompts were designed by expert panel to elicit construct-relevant cognitive processes without requiring particular vocabulary or response structure, and were pilot tested with 200 students across grade levels; cognitive interviews confirmed students engaged the intended processes.

Example prompts. Grit: "Tell me about a time you worked on something that took a long time to finish. What kept you going when it got hard?" Curiosity: "Describe a time when you didn't understand something and wanted to figure it out. What did you do?" Self-Regulation: "Tell me about a situation where you had to control your emotions or focus even though it was difficult."

Audio was transcribed by automated speech recognition and scored by a fine-tuned DeBERTa-v3-base encoder with a CORAL ordinal regression head, against construct-specific rubrics on a 0–4 scale (0–1 Foundational, 2 Emerging, 3 Developing, 4 Competent/Advanced). Scores were normalized for response length and calibrated against rotating anchor sets of expert-validated exemplars.

Scoring model and version

Responses in this study were scored by a fine-tuned DeBERTa-v3-base encoder with a CORAL ordinal regression head, the same architecture as IMPACTER's current production suite. The release in production at the time of collection reported 97% mean adjacent agreement with expert consensus on the anchor sets; the current production release (v6.0) reports 98.0%.

The evidence in this report attaches to the release that scored the data. Because the architecture, loss function, and rubric anchoring are unchanged across the two releases, the findings are expected to hold for v6.0, but re-estimation is listed in §7.3 rather than assumed. Validity claims travel with model versions, and IMPACTER versions and freezes production models specifically so that a claim can be tied to the release it was established on.

Impacter Pathway
Technical Report 2025-01 · Multitrait-Multimethod Construct Validity Impacter Education Corporation · version 1.0 · data collected spring 2025
Section§4
§4 Method, continued

Method 2 · Teacher rubric ratings

Classroom teachers rated a random subsample of their students (typically 5–8 per teacher) on 4-point behavioral rubrics mirroring the scoring structure. Teachers completed 60 minutes of training on construct definitions and rubric anchors with practice coding and calibration, and were paid a stipend. Ratings were completed within two weeks of the student assessment to limit developmental drift. Inter-rater agreement from dual-coding 10% of cases: Cohen's κ = .71–.78.

Method 3 · Coded performance scenarios

Students responded to brief vignettes eliciting construct-relevant reasoning about hypothetical situations rather than autobiographical recall, an important independence property, since it does not ask students to narrate their own past. Responses were coded by trained research assistants after 8 hours of training and demonstrated ≥80% agreement with gold-standard codes. Twenty percent double-coded; ICC(2,1) = .82–.87.

Method 4 · Published self-report scales

Short-form established instruments: Short Grit Scale (Grit-S; Duckworth & Quinn, 2009), 8 items, α = .82; Epistemic Curiosity subscale (Litman, 2008), 5 items, α = .79; Brief Self-Control Scale (Tangney et al., 2004), 6 items, α = .81. Test-retest over two weeks in a subsample (n = 248) ranged r = .76–.83.

4.3 Administration and fidelity

All four methods were administered within a 4-week window. Half of schools completed self-reports before scenarios to test order effects; no significant differences emerged (all F < 1.5, p > .20), so data were combined. Completion exceeded 90% for every method. Three schools (287 students) were excluded for completion below 85% or protocol deviation.

4.4 Analysis

All measures were standardized within method. Outliers beyond 3.5 SD were winsorized (0.4% of values). Missing data (6.2% overall) were handled by full information maximum likelihood in the structural models and pairwise deletion in the correlation matrix; multiple imputation (20 imputations) produced substantively identical results.

We estimated a correlated-trait, correlated-method (CT-CM) confirmatory factor model (Eid et al., 2008): three correlated trait factors, four correlated method factors, twelve observed indicators, cross-loadings fixed at zero. This specification permits direct estimation of trait, method, and residual variance components. Fit was evaluated against conventional thresholds (CFI and TLI ≥ .95 excellent; RMSEA ≤ .05; SRMR ≤ .05). Beyond classical MTMM patterns we computed average variance extracted (Fornell & Larcker, 1981) and heterotrait–monotrait ratios (Henseler et al., 2015), which are stricter tests of discriminant validity. Measurement invariance was tested in the configural → metric → scalar sequence, with ΔCFI < .01 as the criterion (Cheung & Rensvold, 2002). Models were estimated in Mplus 8.8 with MLR estimation; descriptive and sensitivity analyses in R 4.3.1.

Impacter Pathway
Technical Report 2025-01 · Multitrait-Multimethod Construct Validity Impacter Education Corporation · version 1.0 · data collected spring 2025
Section§5
§5

Results

5.1 Reliability of each method

Before examining cross-method relationships, each method must be shown to produce consistent scores; an unreliable measure cannot converge with anything.

Table 2Reliability by method and construct
Method / constructStatisticEstimate95% CI
IMPACTER voice
GritAnchor-set consistency.91[.89, .93]
CuriosityAnchor-set consistency.89[.87, .91]
Self-RegulationAnchor-set consistency.90[.88, .92]
Teacher ratings
GritCohen's κ.76[.69, .83]
CuriosityCohen's κ.71[.64, .78]
Self-RegulationCohen's κ.78[.71, .85]
Performance scenarios
GritICC(2,1).84[.80, .88]
CuriosityICC(2,1).82[.78, .86]
Self-RegulationICC(2,1).87[.83, .91]
Self-report scales
Grit (Grit-S)Cronbach's α.82[.80, .84]
Curiosity (EC)Cronbach's α.79[.77, .81]
Self-Regulation (BSCS)Cronbach's α.81[.79, .83]

Note. The four statistics are not interchangeable and should not be compared directly across methods. α measures internal consistency of items, κ and ICC measure agreement between raters, and anchor-set consistency measures scoring stability of the model across administrations. Each clears the conventional threshold for its own type. The voice figures are a scoring-stability property and are not evidence of construct validity; that is the subject of §5.2 onward.

Impacter Pathway
Technical Report 2025-01 · Multitrait-Multimethod Construct Validity Impacter Education Corporation · version 1.0 · data collected spring 2025
Section§5
§5 Results, continued

5.2 The MTMM correlation matrix

The matrix below is the core evidence. Read it as four blocks. The shaded cells are same-competency, different-method correlations, which are the convergent evidence. The cells within each method block are different-competency, same-method correlations, which are the discriminant evidence.

Table 3Multitrait–multimethod correlation matrix (N = 8,612)
Measure V·GV·CV·S T·GT·CT·S P·GP·CP·S R·GR·CR·S
Voice · Grit1.00
Voice · Curiosity.421.00
Voice · Self-Reg.33.381.00
Teacher · Grit.74.39.281.00
Teacher · Curiosity.41.72.34.471.00
Teacher · Self-Reg.29.36.70.45.511.00
Scenario · Grit.71.36.27.68.38.311.00
Scenario · Curiosity.39.69.31.42.66.35.441.00
Scenario · Self-Reg.27.32.67.33.39.63.36.411.00
Scale · Grit.76.35.25.71.37.28.73.38.291.00
Scale · Curiosity.37.73.29.40.69.32.41.71.34.461.00
Scale · Self-Reg.25.31.71.31.38.68.32.37.69.38.441.00

Note. V = IMPACTER voice; T = teacher rating; P = performance scenario; R = self-report scale. G = Grit; C = Curiosity; S = Self-Regulation. Shaded cells are monotrait–heteromethod (convergent) correlations. All r > .10 significant at p < .001, two-tailed. Cell n varies 8,147–8,612 due to pairwise deletion.

Impacter Pathway
Technical Report 2025-01 · Multitrait-Multimethod Construct Validity Impacter Education Corporation · version 1.0 · data collected spring 2025
Section§5
§5.2 The MTMM matrix, read against the criteria

Convergent evidence

Same-competency correlations across methods: Grit mean r = .73 (range .71–.76); Curiosity mean r = .71 (.69–.73); Self-Regulation mean r = .69 (.67–.71). All substantially exceed the .50 benchmark conventionally treated as convergent evidence. Notably, the voice measure achieved the highest or second-highest correlation with each alternate method for every competency; it converges with independent methods at least as well as teacher ratings or published scales do.

Discriminant evidence

Different-competency correlations within method: voice mean r = .38 (.33–.42); teacher mean r = .48 (.45–.51); scenario mean r = .40 (.36–.44); scale mean r = .43 (.38–.46). Every one of these falls below the convergent correlations.

The comparison that matters most: Voice·Grit correlates .74 with Teacher·Grit, but only .42 with Voice·Curiosity. Same student, same instrument, same session, different competency, and the correlation drops 32 points. Steiger's Z for dependent correlations, Z = 41.2, p < .001. If the model were scoring articulateness, that gap could not exist.

Campbell–Fiske criteria

Overall convergent mean (.72) versus discriminant mean (.31), Steiger's Z = 87.3, p < .001.

Impacter Pathway
Technical Report 2025-01 · Multitrait-Multimethod Construct Validity Impacter Education Corporation · version 1.0 · data collected spring 2025
Section§5
§5 Results, continued

5.3 Confirmatory factor analysis

The classical matrix is descriptive. Modeling traits and methods as simultaneous latent factors tests whether the predicted structure actually reproduces the observed correlations, and quantifies how much of each score comes from where.

Table 4Nested model comparison
Modelχ²dfCFITLIRMSEA [90% CI]SRMR
1 · Traits only, no method factors1987.451.814.767.092 [.088, .096].071
2 · Correlated traits, correlated methods412.848.962.951.041 [.037, .045].032
3 · Correlated traits, uncorrelated methods523.154.948.935.048 [.044, .052].039

Note. MLR estimation, robust standard errors. Model 2 improves on Model 1 by Δχ² = 1574.6, p < .001 (Satorra–Bentler corrected), ΔCFI = .148, ΔRMSEA = .051. Model 2 is retained. The poor fit of Model 1 is itself informative: method effects are real and must be modeled, not assumed away.

Table 5Standardized loadings from the retained CT-CM model
IndicatorTrait λSEMethod λSE
IMPACTER voice
Voice · Grit.79.02.24.03.68
Voice · Curiosity.82.02.19.03.71
Voice · Self-Reg.77.02.27.03.66
Teacher ratings
Teacher · Grit.73.02.35.03.65
Teacher · Curiosity.76.02.31.03.68
Teacher · Self-Reg.71.02.38.03.64
Performance scenarios
Scenario · Grit.78.02.22.03.66
Scenario · Curiosity.80.02.18.03.67
Scenario · Self-Reg.75.02.26.03.63
Self-report scales
Scale · Grit.81.02.17.03.69
Scale · Curiosity.84.02.14.03.73
Scale · Self-Reg.79.02.21.03.67

Note. All loadings p < .001. R² is the proportion of indicator variance explained by trait and method factors jointly. Trait loadings range .71–.84 (mean .78); method loadings range .14–.38 (mean .24). Trait loadings exceed method loadings for every one of the twelve indicators without exception.

Impacter Pathway
Technical Report 2025-01 · Multitrait-Multimethod Construct Validity Impacter Education Corporation · version 1.0 · data collected spring 2025
Section§5
§5 Results, continued

5.4 Variance decomposition: the eloquent-talker test

Squaring the standardized loadings partitions each observed score into the portion attributable to the competency, the portion attributable to the measurement method, and residual.

Table 6Variance components by indicator
IndicatorTraitMethodResidualTotal
IMPACTER voice
Voice · Grit.62.06.321.00
Voice · Curiosity.67.04.291.00
Voice · Self-Reg.59.07.341.00
Teacher ratings
Teacher · Grit.53.12.351.00
Teacher · Curiosity.58.10.321.00
Teacher · Self-Reg.50.14.361.00
Performance scenarios
Scenario · Grit.61.05.341.00
Scenario · Curiosity.64.03.331.00
Scenario · Self-Reg.56.07.371.00
Self-report scales
Scale · Grit.66.03.311.00
Scale · Curiosity.71.02.271.00
Scale · Self-Reg.62.04.341.00
Mean across indicators.61.06.331.00

Note. Components are squared standardized loadings from Table 5; residual = 1 − (trait + method). The mean row is the arithmetic mean of the twelve indicator values (trait .608, method .064, residual .328), reported to two decimals. Mean method variance by method: voice .06, teacher .12, scenario .05, scale .03.

What this rules out. The eloquent-talker hypothesis makes a specific, testable prediction: if the voice measure primarily captures verbal facility, then method variance for the voice indicators should be high (a fluent student scores high on everything the instrument asks) and trait-specific variance should be low. The observed pattern is the reverse. Voice indicators carry mean method variance of .06 against mean trait variance of .62, a ratio of roughly 10:1, and their method variance is half that of teacher ratings (.12).

This is the finding that does the work in this report, and it deserves a conservative statement. It does not show the instrument is free of verbal-ability influence; language-based measurement will always carry some. It shows that the dominant source of variance in a score is which competency is being measured, not that the measurement happened through language. Method variance of 6% also sits well below the 25–50% commonly observed across social-science measurement (Podsakoff et al., 2003), which is a point in favor of standardized linguistic analysis relative to human observational rating rather than a claim of perfection.

Impacter Pathway
Technical Report 2025-01 · Multitrait-Multimethod Construct Validity Impacter Education Corporation · version 1.0 · data collected spring 2025
Section§5
§5 Results, continued

5.5 Stricter discriminant tests

Campbell–Fiske criteria are necessary but permissive. Average variance extracted and heterotrait–monotrait ratios are the stricter contemporary tests.

Table 7Average variance extracted and Fornell–Larcker check
ConstructAVE√AVELargest inter-construct rCriterion met
Grit.61.78.48 (Curiosity)Yes
Curiosity.59.77.48 (Grit)Yes
Self-Regulation.58.76.44 (Curiosity)Yes

Note. AVE above .50 indicates the construct explains more than half of its indicators' variance. The Fornell–Larcker criterion requires √AVE to exceed every correlation with another construct; all three clear it with margin.

Table 8Heterotrait–monotrait ratios
Construct pairHTMT95% CIBelow .85
Grit ↔ Curiosity.67[.63, .71]Yes
Grit ↔ Self-Regulation.59[.55, .63]Yes
Curiosity ↔ Self-Regulation.73[.69, .77]Yes

Note. HTMT below .85 supports discriminant validity (Henseler et al., 2015). All confidence intervals exclude .85. The three competencies are distinguishable rather than three labels for a single "good student" factor.

Latent trait factor correlations were .48 (Grit–Curiosity), .44 (Curiosity–Self-Regulation), and .39 (Grit–Self-Regulation), moderate and positive, which is what §3.4 predicted. The predictions were stated in advance and the data matched them; that correspondence is itself part of the evidence.

Impacter Pathway
Technical Report 2025-01 · Multitrait-Multimethod Construct Validity Impacter Education Corporation · version 1.0 · data collected spring 2025
Section§5
§5 Results, continued

5.6 Measurement invariance

A validity claim that holds only for some students is not a validity claim. We tested whether the measurement structure is equivalent across the subgroups most at issue for a language-based instrument.

Table 9Multigroup measurement invariance
ComparisonLevelχ²dfCFIRMSEAΔCFIHeld
Grade band · middle (6–8) vs high (9–12)
ConfiguralSame structure421.796.959.043refref
MetricSame loadings438.2105.957.042−.002Yes
ScalarSame intercepts463.5114.954.043−.003Yes
Gender · female vs male
ConfiguralSame structure436.896.961.044refref
MetricSame loadings451.3105.959.043−.002Yes
ScalarSame intercepts472.1114.957.044−.002Yes
English Learner status · EL vs non-EL
ConfiguralSame structure447.296.956.045refref
MetricSame loadings465.8105.954.045−.002Yes
ScalarSame intercepts488.3114.951.046−.003Yes

Note. Criterion ΔCFI < .01 (Cheung & Rensvold, 2002). Scalar invariance was achieved in all three comparisons, which licenses comparison of group means. The non-binary/other gender category (n = 102) was too small for stable multigroup estimation and was excluded from the gender comparison; this is a limitation, not a finding.

On the English Learner result. This is the invariance test that matters most for an instrument built on student language, and it warrants care rather than celebration. Scalar invariance means the measurement model, including structure, loadings, and intercepts, operates equivalently for English Learners and non-English Learners in this sample. It does not mean the instrument is free of language-related advantage. It means that whatever advantage exists does not distort the relationship between the construct and its indicators. Item-level differential functioning analysis, which asks a narrower and harder question, has not been conducted and is listed in §7.3.

Impacter Pathway
Technical Report 2025-01 · Multitrait-Multimethod Construct Validity Impacter Education Corporation · version 1.0 · data collected spring 2025
Section§6
§6

Evidence coverage against the Standards

The Standards for Educational and Psychological Testing (AERA, APA, & NCME, 2014) organize validity evidence into five sources. Stating where this study lands on each is more useful than a summary verdict.

Scope of use

This report supports the interpretation of IMPACTER scores for formative feedback, progress monitoring, program evaluation, and group-level standards monitoring. It does not support use for high-stakes accountability decisions about individual students, educator evaluation, or special education eligibility or placement. Those uses require consequential and criterion evidence this study does not provide, and the published IMPACTER Model Card lists them as out of scope.

Impacter Pathway
Technical Report 2025-01 · Multitrait-Multimethod Construct Validity Impacter Education Corporation · version 1.0 · data collected spring 2025
Section§7
§7

Limitations and outstanding work

7.1 Sampling

All data are from California districts, limiting generalization to states with different frameworks and accountability contexts. More importantly, participating districts were not randomly sampled; they self-selected into IMPACTER partnerships and therefore likely over-represent schools already oriented toward measuring these competencies. Replication in districts with no prior relationship would be a stronger test.

7.2 Design

Three limitations bear directly on interpretation. First, all four methods require either language production or language interpretation, so they are independent in machinery but not fully independent in modality; adding parent report, peer nomination, or direct behavioral observation such as time-on-task coding would reduce shared method variance further. Second, the four-week administration window limits developmental drift but introduces temporal proximity as a possible source of shared variance; an 8–12 week spacing would be cleaner. Third, only three of the eight competencies in the IMPACTER architecture were examined, and they were selected partly for theoretical distinctiveness. Pairs expected to overlap more, such as Compassion and Perspective-Taking, would be a more demanding discriminant test, and that test has not been run.

7.3 Outstanding evidence

  • Re-estimation on the current production release. This study scored under the pre-v6.0 release of the DeBERTa-v3-base suite. The architecture is unchanged in v6.0, so the findings are expected to replicate, but they should be re-estimated rather than assumed to carry forward.
  • Criterion and predictive validity. Whether scores predict grade point average trajectory, course completion, attendance, disciplinary referral, or postsecondary enrollment. Not examined here.
  • Test–retest stability. Deterministic scoring means an identical response receives an identical score. That is a property of the engine, not a stability estimate for the construct in students, and the two should not be conflated.
  • Item-level differential functioning. Prompt-level DIF analysis across English Learner status, race/ethnicity, socioeconomic status, and disability status.
  • Intervention sensitivity. Whether scores move in the expected direction following instruction, which is the evidence that makes these competencies assessable as malleable rather than fixed.
  • Full discriminant coverage. Extension to the remaining five competencies, including the conceptually adjacent pairs.
Impacter Pathway
Technical Report 2025-01 · Multitrait-Multimethod Construct Validity Impacter Education Corporation · version 1.0 · data collected spring 2025
Section§8 – 9
§8

Bottom line

Agreement with human raters establishes that the scoring engine reproduces trained raters. It does not establish that the instrument measures the competency it names. This study addressed the second question with the design built for it.

Result: variance in an IMPACTER score is dominated by which competency is being measured, not by the fact that the measurement was made through student language. Same-competency correlations across four independent methods averaged .72 against .31 across competencies; trait variance exceeded method variance by roughly ten to one; and the measurement model held equivalently across grade band, gender, and English Learner status.

Bounded to the release scored, the three competencies examined, and the California sample described in §4.1. Not a claim of readiness for high-stakes individual decisions. §7.3 is the complete list of what remains.

§9

References

AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. testingstandards.net

Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81–105. doi:10.1037/h0046016

Cheung, G. W., & Rensvold, R. B. (2002). Evaluating goodness-of-fit indexes for testing measurement invariance. Structural Equation Modeling, 9(2), 233–255. doi:10.1207/S15328007SEM0902_5

Duckworth, A. L., & Carlson, S. M. (2013). Self-regulation and school success. In B. W. Sokol et al. (Eds.), Self-regulation and autonomy. Cambridge University Press. doi:10.1017/CBO9781139152198.015

Duckworth, A. L., Peterson, C., Matthews, M. D., & Kelly, D. R. (2007). Grit: Perseverance and passion for long-term goals. Journal of Personality and Social Psychology, 92(6), 1087–1101. doi:10.1037/0022-3514.92.6.1087

Duckworth, A. L., & Quinn, P. D. (2009). Development and validation of the Short Grit Scale (Grit-S). Journal of Personality Assessment, 91(2), 166–174. doi:10.1080/00223890802634290

Duckworth, A. L., & Yeager, D. S. (2015). Measurement matters: Assessing personal qualities other than cognitive ability for educational purposes. Educational Researcher, 44(4), 237–251. doi:10.3102/0013189X15584327

Eid, M., Nussbeck, F. W., Geiser, C., Cole, D. A., Gollwitzer, M., & Lischetzke, T. (2008). Structural equation modeling of multitrait-multimethod data. Psychological Methods, 13(3), 230–253. doi:10.1037/a0013219

Fornell, C., & Larcker, D. F. (1981). Evaluating structural equation models with unobservable variables and measurement error. Journal of Marketing Research, 18(1), 39–50. doi:10.1177/002224378101800104

Henseler, J., Ringle, C. M., & Sarstedt, M. (2015). A new criterion for assessing discriminant validity in variance-based structural equation modeling. Journal of the Academy of Marketing Science, 43(1), 115–135. doi:10.1007/s11747-014-0403-8

Kashdan, T. B., Gallagher, M. W., Silvia, P. J., et al. (2009). The Curiosity and Exploration Inventory-II. Journal of Research in Personality, 43(6), 987–998. doi:10.1016/j.jrp.2009.04.011

Litman, J. A. (2008). Interest and deprivation factors of epistemic curiosity. Personality and Individual Differences, 44(7), 1585–1595. doi:10.1016/j.paid.2008.01.014

Messick, S. (1995). Validity of psychological assessment. American Psychologist, 50(9), 741–749. doi:10.1037/0003-066X.50.9.741

Podsakoff, P. M., MacKenzie, S. B., Lee, J.-Y., & Podsakoff, N. P. (2003). Common method biases in behavioral research. Journal of Applied Psychology, 88(5), 879–903. doi:10.1037/0021-9010.88.5.879

Tangney, J. P., Baumeister, R. F., & Boone, A. L. (2004). High self-control predicts good adjustment. Journal of Personality, 72(2), 271–324. doi:10.1111/j.0022-3506.2004.00263.x

Williamson, D. M., Xi, X., & Breyer, F. J. (2012). A framework for evaluation and use of automated scoring. Educational Measurement: Issues and Practice, 31(1), 2–13. doi:10.1111/j.1745-3992.2011.00223.x

Yan, D., Rupp, A. A., & Foltz, P. W. (Eds.). (2020). Handbook of automated scoring. CRC Press. doi:10.1201/9781351264808

Zimmerman, B. J. (2002). Becoming a self-regulated learner. Theory Into Practice, 41(2), 64–70. doi:10.1207/s15430421tip4102_2

Impacter Pathway
Technical Report 2025-01 · Multitrait-Multimethod Construct Validity Impacter Education Corporation · version 1.0 · data collected spring 2025
Back matter
Disclosure

Status, funding and document control

Disclosure and status

Conducted and funded by IMPACTER Education Corporation. The design, analysis, and interpretation are IMPACTER's own; no external evaluator conducted the analysis, and readers should weigh it accordingly. It is published in full, including design, figures, and limitations, so that it can be independently evaluated and argued with.

Human subjects and data handling

Institutional review board approval and parental consent procedures are documented in Appendix B, following FERPA and California Education Code §49073. Participation used standard opt-out procedures with student assent at the time of assessment. All data were de-identified before analysis. Audio was transcribed and deleted within 30 days under district data retention agreements.

Data availability

Aggregated data and analysis code are available to qualified researchers on request under a data use agreement. Student-level data cannot be shared under FERPA.

Appendices

A. Full demographics. B. IRB and consent documentation. C. Prompts and scoring rubrics. D. Expert panel and cognitive interview protocols. E. Sensitivity analyses including multiple imputation. F. Full multigroup invariance output. Available on request.

Document control

Report No. 2025-01  ·  Version 1.0  ·  Supersedes none
Data collection Spring 2025  ·  Scoring release DeBERTa-v3-base with CORAL ordinal head, pre-v6.0
Companion documents IMPACTER Model Card v6.0  ·  Ordinal Scoring Pipeline Technical Note
Review revise on re-estimation under v6.0, or on extension to additional competencies
Address impacterpathway.com/assets/docs/research/multitrait-multimethod-construct-validity
Correspondence josh@impacterpathway.com  ·  Series impacterpathway.com/research