Every score we publish is checked by independent judges, trained humans and a model held to their standard. When they disagree, the human wins and the model learns.




Every score beginsSustained effort toward a goal, held through setbacks. One of the eight durable skills we measure. Below: what Grit sounds like at each level. Click a level to hear it in a real sentence.

“At first I kept messing up the routine and I wanted to quit, but I kept practicing it every night until it finally clicked. Because I stuck with it, I made the team, and now when something is hard I remind myself it just takes reps.”
The working interface our raters use. This is science.
Set the Human Score on each row, and where it differs, yours will win.
You just ran the scoring room. Now slow one disagreement down. Take the call yourself: do you notice L2, or L3?

“My auntie always says quitters never win, so I guess I just don’t quit. Like last semester with algebra, everybody said drop it, and I stayed.”
Detected a persistence verb, but classified “quitters never win” as a borrowed phrase, not first-person evidence.

“My grandpa always tells me slow is smooth, and honestly he’s right. When my science fair project kept failing, I rebuilt it three times until it worked.”
Same pattern — a borrowed phrase carrying lived persistence — now credited correctly, because a human corrected it.
A system that never disagrees with itself is a system nobody is checking. Every divergence is surfaced and turned into training data.
Rubrics define the skill. Humans apply the rubric. The model is trained on their judgment and overruled by it. Every time, by design.
Experts agree with each other at κ 0.87 — the model sits inside the human range.
Within one rubric level of the expert, on responses the model has never seen.
Identical to the expert. Everything further apart is overruled by a human.