Every centre is drawn to automated marking: fast, independent of teacher availability, consistent across classes. The short answer: it works at scale provided a gate decides which responses are released and which go to a human.
Three characteristic failure modes
- Invented evidence. Feedback quotes a sentence "from your response" that does not exist, or is trimmed until it means something else. The learner cannot find it in their own work.
- Measuring the wrong thing. The clearest case is pronunciation: a system that scores it from a transcript is marking vocabulary and grammar a second time. Evidence for that criterion must come from measurement of the audio signal itself.
- Broken input, confident output. A Writing response whose Task 1 chart image is missing, a Speaking recording containing one of three parts, a silent audio file. If the system still returns a number, that number is meaningless and looks entirely normal.
The third is the dangerous one, because it fails silently.
The minimum release gate
Before an automated score reaches a learner, require every one of these — and treat missing evidence as a failure, because passing something through in silence is the worst failure mode here:
- The marking was not rejected by the validator.
- Every quotation in the feedback traces back to the exact wording of the submitted work.
- No criterion is flagged as unmarkable.
- Input data is valid: prompt images present, all recording parts present, the response is not a restatement of the prompt.
- The confidence score meets the threshold the centre sets.
This is how LinguaGrade operates: responses clearing every condition are released automatically, and anything failing one waits in a teacher's queue. Centres switch auto-release on or off and set their own threshold.
Calibrate before trusting at scale
Do not pick a threshold by feel. Take a few dozen responses already marked by teachers, have the system mark them again, and compare. Watch two numbers: how often the two differ by more than half a band, and how often the system was confident and wrong. The second decides your threshold.
Only after that should auto-release be enabled for real learners. Until then, route everything through teachers and treat machine marking as a draft that makes them faster.
What to tell learners
Transparency is what keeps trust: say that work is marked first by the system with teacher review where required, give an appeals route, and name who is finally accountable. A centre that states this plainly is more credible than one that leaves learners guessing.
Frequently asked questions
Can an AI band replace an official result?
No. It is internal practice feedback, not a recognised examination result.
What share of responses should reach a teacher?
It depends on your threshold and input quality. Watch whether the share is stable or rising — rising means your input quality is degrading.
Should learners be told a machine marked their work?
Yes. It does not reduce the value of the feedback and it protects the centre when an appeal arrives.