Under The Hood
In Your Hands
The Big Picture
The Solutions
The Methods
The Foundations
The Case
The Science
The Proof
The Arc
The People
The Promise
Transparency · Model Card v6.0

How Our ML Scores
Student Voice

Every IMPACTER model release ships with a public Model Card: architecture, training data, human-agreement benchmarks, and fairness checks. Because a score you can't audit is a score you can't trust.

0.000
Avg QWK
Human–ML agreement (Quadratic Weighted Kappa). Industry threshold: ≥ 0.70.
0.0%
Avg Exact Accuracy
ML scores matching trained human raters.
0.0%
Avg Adjacent Accuracy
Within ±1 point of expert scores —
the human–human benchmark.
01 · The Science

Validated Rubrics
Not Black-Box LLMs

Eight competencies, each benchmarked against trained human raters.
Click any competency to see the linguistic markers the model actually reads.

Agreement With Human Raters,
Trait by Trait.

1.0 = perfect agreement · 0.70 = production threshold

Architected by
Dr. Andrew Arnold
Senior Scientific Advisor
  • Principal ML Engineer, Shopify
  • Chief Scientist, AWS CodeWhisperer
  • PhD in Machine Learning, Carnegie Mellon
  • Architect of IMPACTER's scoring pipeline
Dr. Andrew Arnold
Dashed line = 0.70 industry threshold
≥ 0.70
00.250.500.751.0

    SOURCE: HARVARD MCC DUAL-AXIS RUBRIC · IMPACTER MODEL CARD v6.0
    02 · Overview & Scope

    What These Models Do —
    And What They Don't

    The v6.0 suite scores K-12 open-ended responses on eight human-skills competencies, calibrated to human expert scoring of student reasoning, reflection, and narrative evidence.

    Intended Use

    • Scoring student reflections for Portrait of a Graduate durable-skills competencies
    • Behavioral health and social-emotional reflection scoring
    • Workforce and career readiness / CTE assessments
    • Short-form constructed-response tasks

    Out of Scope

    • High-stakes summative academic content scoring
    • Individual diagnostic or clinical decisions without human review
    • Any use outside rubric-anchored prompts and approved task types
    03 · Training & Governance

    Trained On Authentic Voice
    Frozen For Audit

    Models train on a deidentified corpus: the words of a response and the level an expert rater assigned it on a 0–4 scale, from partners whose agreement covers it, and never from a school that has said no.
    No identified student data trains any model. Each release is evaluated on held-out sets of newly collected responses.

    1

    Registered

    Every new version is registered in MLflow with full lineage, from training run to evaluation record. Nothing reaches production unrecorded.

    2

    Validated

    Promoted to Production only if QWK improves over the current model, with SMD and subgroup fairness checks on every release.

    3

    Frozen

    Production models are frozen for deterministic scoring, identical responses receive identical scores, every administration.

    Model Configuration

    Architecture
    DeBERTa-v3-base · 184M params
    Version
    MODEL SUITE v6.0
    Loss Function
    CORAL ORDINAL REGRESSION
    Platform
    DATABRICKS + MLFLOW REGISTRY
    Score Scale
    0–4 RUBRIC-ALIGNED ORDINAL
    Promotion Gates
    QWK ≥ 0.70 · DEGRADATION ≤ 0.10 · |SMD| < 0.10
    04 · Fairness

    Checked For Drift
    Before Every Promotion

    Distributional SMD checks, by score band and response length, show no systematic drift.
    Scalar invariance holds for grade band, gender, and English learner status (Technical Report 2025-01). Race/ethnicity, socioeconomic status, and disability are listed as outstanding. A district's own subgroup analysis runs only at its direction and is reported to it.

    Structural SMD — all bands well under threshold

    THRESHOLD: |SMD| < 0.10 · OBSERVED: 0.03–0.08
    |SMD| < 0.10

    Version note: All statistics report Model Suite v6.0 performance on held-out validation data. Performance varies by assessment, prompt design, and cohort; each release publishes updated figures. Emotional inference and biometric capture are disabled for K-12 deployments.

    Measurement You Can Interrogate

    Get the full Model Card, scoring rubric, and technical architecture documentation, or walk through it live with our data science team.