Evidence-first teacher review

Find feedback drift
before students do.

Surface similar student evidence that received meaningfully different scores or feedback—then let the teacher make the call.

See the method
No account or API key Synthetic student data No automated grade changes
Calibration reviewSynthetic demo
High priority

Similar evidence, different scores

R1 · v1.0

Two submissions show the same unsupported-claim pattern at similar severity.

Student A3/4

“This shows that student volunteers can help the community.”

Evidence use
vs
Student B1/4

“The food bank example is proof that required service…”

Evidence use
Would the same scoring interpretation apply to both?
6synthetic submissions
4deterministic audit rules
100%evidence-linked findings
0automatic grade changes
GPT-5.6 Codex sample run

Inspect the completed analysis—without supplying an API key.

Six full fictional essays were analyzed in Codex, passed through the same verbatim-evidence gate used by the server route, and saved for deterministic R1–R4 review on this static site.

32verified excerpts
13review questions
0API calls on this site
Completely fictional data · scores never changed
How it works

Model-assisted extraction.
Rule-based review.

The model structures evidence. Versioned code decides which patterns deserve review.

01

Bring the rubric and feedback

Use anonymous student labels, rubric scores, and the feedback already written by the teacher.

02

Extract evidence signals

GPT-5.6 returns constrained signals with verbatim evidence spans—not final grades.

03

Compare with visible rules

TypeScript rules find score gaps, uneven coverage, and score-feedback mismatches.

04

Teacher makes the call

Confirm, dismiss, resolve, and export a transparent decision log.

Product boundary

A calibration mirror,
not a grading authority.

Never changes a score automatically

Never labels a teacher as unfair or biased

Never ranks students by risk or ability

Keeps every signal tied to source evidence