How Our ML Scores
Student Voice
Every IMPACTER model release ships with a public Model Card: architecture, training data, human-agreement benchmarks, and fairness checks. Because a score you can't audit is a score you can't trust.
the human–human benchmark.
Validated Rubrics
Not Black-Box LLMs
Eight competencies, each benchmarked against trained human raters.
Click any competency to see the linguistic markers the model actually reads.
Agreement With Human Raters,
Trait by Trait.
1.0 = perfect agreement · 0.70 = production threshold
- Principal ML Engineer, Shopify
- Chief Scientist, AWS CodeWhisperer
- PhD in Machine Learning, Carnegie Mellon
- Architect of IMPACTER's scoring pipeline
What These Models Do —
And What They Don't
The v6.0 suite scores K-12 open-ended responses on eight human-skills competencies, calibrated to human expert scoring of student reasoning, reflection, and narrative evidence.
Intended Use
- Scoring student reflections for Portrait of a Graduate durable-skills competencies
- Behavioral health and social-emotional reflection scoring
- Workforce and career readiness / CTE assessments
- Short-form constructed-response tasks
Out of Scope
- High-stakes summative academic content scoring
- Individual diagnostic or clinical decisions without human review
- Any use outside rubric-anchored prompts and approved task types
Trained On Authentic Voice
Frozen For Audit
Models train on district-provided responses, hand-scored by human experts on a 0–4 scale,
and evaluated on held-out sets of newly collected responses.
Registered
Every new version is registered in MLflow with full lineage, from training run to evaluation record. Nothing reaches production unrecorded.
Validated
Promoted to Production only if QWK improves over the current model, with SMD and subgroup fairness checks on every release.
Frozen
Production models are frozen for deterministic scoring, identical responses receive identical scores, every administration.
Model Configuration
- Architecture
- DeBERTa-v3-base · 184M params
- Version
- MODEL SUITE v6.0
- Loss Function
- CORAL ORDINAL REGRESSION
- Platform
- DATABRICKS + MLFLOW REGISTRY
- Score Scale
- 0–4 RUBRIC-ALIGNED ORDINAL
- Promotion Gates
- QWK ≥ 0.70 · DEGRADATION ≤ 0.10 · |SMD| < 0.10
Checked For Drift
Before Every Promotion
Distributional SMD checks, by score band and response length, show no systematic drift.
Full subgroup analysis (EL status, race/ethnicity) runs on partner-specified validation sets.
Structural SMD — all bands well under threshold
THRESHOLD: |SMD| < 0.10 · OBSERVED: 0.03–0.08Version note: All statistics report Model Suite v6.0 performance on held-out validation data. Performance varies by assessment, prompt design, and cohort; each release publishes updated figures. Emotional inference and biometric capture are disabled for K-12 deployments.
Measurement You Can Interrogate
Get the full Model Card, scoring rubric, and technical architecture documentation, or walk through it live with our data science team.