Assessment Desk

AI contributes a lens. Teachers retain the judgement.

A curriculum-aware assessment workspace shaped by classroom experience and grounded in research on feedback, validity and human judgement.

The premise

We are right not to trust AI with educational judgement. Language models can be inconsistent, persuasive when wrong and capable of reproducing bias at scale. But distrust is not a reason to ignore what they can contribute.

Human assessment has limitations too. Research on feedback, rater behaviour and construct validity describes the effects of workload, fatigue, order, halo and qualities of presentation that are irrelevant to what an assessment intends to measure. The answer is not to choose the supposedly better judge. It is to make the criteria, evidence and reasoning more visible so that each can be challenged.

A useful second lens

I began building Assessment Desk while teaching Language Acquisition in an IB Middle Years Programme school in Tokyo. The working command-line tool connects assignments to curriculum criteria, ingests written documents, handwriting and transcribed speech, and prepares structured analysis for a teacher to review.

The model’s role is sustained attention, not professional authority. A suggested judgement should appear beside its rationale and the evidence in the student’s work. The teacher can inspect it, disagree with it and decide what it means.

Why the criteria matter

A language model can often notice weak reasoning behind polished language without being handed a school rubric. Criterion-binding serves a different purpose: governance. It aligns the analysis with what the school says it measures and gives teachers and moderators a basis for disagreement.

The project has also exposed a deeper issue. A rubric may combine published descriptors, school guidance and local interpretations without recording where each statement came from or what authority it carries. Encoding criteria for machine use forces those distinctions into the open, improving the assessment framework for human markers too.

From classroom pipeline to research questions

Assessment Desk began as a working command-line system rather than a controlled study. It connected real assignments to curriculum criteria, processed written documents, handwriting and recorded speech, stored the model’s structured responses, and carried scores into teacher-facing reports. That history provides something valuable but narrower than validation: a real pipeline that can now be audited.

The audit asks whether each suggested judgement can be traced through the assignment, the evidence the system actually assessed, any OCR or transcription repair, the relevant criterion, and the response returned to the teacher. It also asks a practical question: would a teacher regard the result as reasonable and useful, including in cases where the teacher ultimately disagrees?

That review has already shown why context matters. A recording may be rehearsed, short-notice or spontaneous. An OCR-cleaned text may be the grading input, while a much more polished version stored beside it may have been created only as a constructive example for the student. If those roles are collapsed, an apparently rigorous analysis can answer the wrong question.

This is not evidence that AI grades accurately, eliminates bias or outperforms teachers. The literature on workload, halo effects, construct-irrelevant variance and automation bias motivates the need for visible evidence and accountable review; the existing classroom corpus was not designed to test those effects.

It does create a basis for better research. A future study could deliberately control what each rater sees, distinguish prepared from time-bounded speaking, compare raw and repaired transcripts, and record independent human adjudication of scores and feedback. The aim would not be to make the teacher disappear. It would be to discover where an auditable second lens helps, where it distorts, and how disagreement can improve the final judgement.

The public Assessment Desk demo uses fictional data and currently shows the teacher workspace rather than live assessment results. The next interface work is to make the full evidence chain—and the teacher’s authority over it—visible.

Interested?

Get in touch