“This shows that student volunteers can help the community.”
Find feedback drift
before students do.
Surface similar student evidence that received meaningfully different scores or feedback—then let the teacher make the call.
Similar evidence, different scores
Two submissions show the same unsupported-claim pattern at similar severity.
“The food bank example is proof that required service…”
Inspect the completed analysis—without supplying an API key.
Six full fictional essays were analyzed in Codex, passed through the same verbatim-evidence gate used by the server route, and saved for deterministic R1–R4 review on this static site.
Model-assisted extraction.
Rule-based review.
The model structures evidence. Versioned code decides which patterns deserve review.
Bring the rubric and feedback
Use anonymous student labels, rubric scores, and the feedback already written by the teacher.
Extract evidence signals
GPT-5.6 returns constrained signals with verbatim evidence spans—not final grades.
Compare with visible rules
TypeScript rules find score gaps, uneven coverage, and score-feedback mismatches.
Teacher makes the call
Confirm, dismiss, resolve, and export a transparent decision log.
A calibration mirror,
not a grading authority.
Never changes a score automatically
Never labels a teacher as unfair or biased
Never ranks students by risk or ability
Keeps every signal tied to source evidence