Would your AI oversight survive a test taker’s appeal?
This article explains why human oversight is critical when using AI in assessment.
AI, particularly generative AI, is increasingly being used throughout the assessment lifecycle. There are hundreds of potential use cases, some common ones are:
- Using AI to create questions or put them together into tests, including with Questionmark’s AI authoring tool, Author Aide;
- Helping find test fraud by monitoring test takers and flagging potential issues;
- Using AI to help score essays and other unstructured test answers, which human graders can review, including with Questionmark’s AI scoring and feedback tool, AI Scoring.

AI is often very capable. But it sometimes makes mistakes, sometimes can be biased, and often is very confident about its output. It’s therefore critical to have human review of AI to mitigate the risk of negatively impacting test takers.
Thus, humans need to review AI-authored questions before they are used in tests. Any accusation of cheating from an AI source needs validating by a human. And scores produced by AI need to have careful human review before relying on them. Most other uses of AI in assessment need human oversight.
How to ensure meaningful human oversight of AI in assessment
There are many ways of structuring how humans interact with and review the output of AI. A common approach is “Human In The Loop” (HITL) which means that a human reviews every AI decision, but there are other approaches including “Human On The Loop” (HOTL). Whichever way human review is structured, the key is that human review or oversight is focused on mitigating risks and is meaningful.
Oversight that is cursory or by people without the right skills has no value. For example, if AI scores 400 essays and a human spends an hour or two scanning them and reviewing the scores, it’s likely just lip service or rubber stamping. Meaningful human oversight doesn’t mean that humans have to be involved in every detail, but they have to be involved in the right places in the right ways and the right times, based on the risks involved.
As the ATP’s “Human Oversight of AI in Assessment: Considerations for Responsible Use” says: “when planning human oversight of AI, the key is to ensure appropriate implementation of meaningful human oversight”. It goes on to say that “In an assessment context, this human oversight involves reviewing and validating how AI systems operate with respect to decisions and outcomes that impact the validity, quality, reliability, and fairness of assessments.”
Some of the things that meaningful oversight needs:
- Competence of the overseers
- Criteria and guidelines on when and how to intervene
- Authority to override the AI
- Appropriate incentives for overseers
- Time to consider the AI outputs properly
- Measures to deal with the risk of automation bias (that one believes the AI because it appears confident)
- Quality control and review of the oversight process
These are the ingredients of meaningful oversight. Later in this article, we turn them into a practical checklist, the evidence you’d actually want on file to prove that oversight happened.
The role of validity and defensibility for AI in assessment
A critical reason for human oversight of AI in assessment is the need to ensure the validity of the assessment results, that is to say that it is appropriate to use the results of the assessments for the intended purpose.
If someone fails an assessment when they should have passed, because AI has made a mistake, then that may have serious consequences for that person. They could miss out on an educational or work opportunity.
And if someone passes an assessment when they are not competent and so should have failed, that can have serious consequences too. Someone could be placed in a role where they impact health and safety, or it can displace an opportunity for someone else.
A corollary of validity is defensibility. Defensibility means that a testing organization should be able to justify an assessment decision if it is disputed, for example by a test taker, an appeal panel, a court or a regulator. Critical to defensibility is that documentation and evidence are assembled as the test is authored and delivered. It’s important that any use of AI has appropriate oversight to ensure that any assessment result is defensible.
Without effective oversight, a test taker or other stakeholder disputing a test result could argue that the AI has made an error (or hallucination) and/or been biased, and that it’s not fair that the result determination relies on AI. Depending on the context of the assessment, this could lead to a legal case, regulator action and/or dissatisfaction in the community of test takers.
To deal with this, it’s important not only to have effective and meaningful oversight but also to ensure that oversight is well documented. Documentation is what makes real oversight defensible.
Legal requirements for effective AI oversight in Europe
Particularly in Europe, but also in other jurisdictions, there are legal requirements or constraints on use of AI in assessment, and these can require human oversight.
One area of legal constraint is the EU AI Act. For high-risk use cases, which include using AI to score many summative assessments, these take effect from December 2027. The EU AI Act sets out some high-level requirements on human oversight; this includes “Deployers shall assign human oversight to natural persons who have the necessary competence, training and authority, as well as the necessary support.” The EU is expected to provide further guidance on oversight before the law goes live.
However, the GDPR has been in effect for several years and also has some oversight requirements. The GDPR gives test takers in Europe the right (with some exceptions) “not to be subject to a decision based solely on automated processing … which significantly affects him or her”. This potentially impacts use of AI in assessment, which has significant impact on people, including AI scoring and AI fraud detection.
European guidance says “To qualify as human involvement, the controller must ensure that any oversight of the decision is meaningful, rather than just a token gesture. It should be carried out by someone who has the authority and competence to change the decision.”
There have been various EU regulatory actions and court rulings in relation to use of AI with insufficient oversight:
- A large ride-hailing, transport and delivery company used automation to deactivate accounts of drivers due to fraud concerns. In 2023, a Dutch appeal court found that the company’s human review was little more than a symbolic act; for example, the company could not show what qualifications or knowledge its reviewers had. The company was given a heavy fine in 2026 related to automated decisions without human involvement.
- In Italy, the data protection regulator (the Garante) fined a food delivery company in 2021 for failures in how it used algorithms to manage its riders. It fined the company again in 2024, this time €5 million. The 2024 decision required the company to ensure that decisions made by algorithms are verified by adequately trained operators.
- There has also been a regulatory case in the assessment industry. The Portuguese DPA found many issues with a proctoring application in 2021. Part of this involved human oversight concerns in relation to how teachers should interpret cheating flags, with an “absence of specific guidelines on the interpretation”, and the risk that it might “lead teachers to validate the systems’ decisions as a rule”.
UK law is keener to encourage automation, but there are still rules on automated decisions, including the requirement to offer human review on request.
Legal requirements for effective AI oversight in the U.S
In general, there is less regulatory constraint on the use of AI in assessments in the U.S., so human oversight is primarily driven by the desire of testing organizations to produce valid and trustable results and the risk of test taker legal action if they do not.
General fairness and non-discrimination laws could impact use of AI in assessment without sufficient oversight. And some states have automated decision-making technology (“ADMT”) laws which place requirements on using AI or other automation to make decisions. For example, California, Colorado and Connecticut have ADMT laws set to come into effect that may impact assessment by employers.
AI regulation is a fast-moving area and testing organizations are encouraged to consult lawyers for precise impacts on them.
Comparing legal requirements for AI oversight by region
| Region | Law | Oversight requirements | Timing |
|---|---|---|---|
| EU | EU AI Act | High-risk uses (e.g. AI scoring of summative assessments) require oversight assigned to competent, trained, authorized people | Applies from Dec 2027 |
| EU | GDPR (Article 22) | Right not to be subject to a decision based solely on automated processing with significant effect; oversight must be “meaningful,” not a token gesture | In effect since 2018 |
| UK | UK Data protection Law | Less restrictive than the EU, but still requires human review of automated decisions on request | Revised in 2026 |
| US | ADMT laws (CA, CO, CT) | Place requirements on using AI/automation to make employment-related decisions | Coming into effect, dates vary by state |
| US (federal) | General fairness/non-discrimination law | No AI-specific oversight mandate, but discrimination claims possible if oversight is insufficient | Current |
AI and accessibility
Another issue to be aware of in relation to human oversight and assessment is accessibility.
There are strict accessibility laws in the U.S., Europe and in many other jurisdictions. If AI does something which crosses an accessibility threshold, then there could be negative consequences for a testing organization.
For example, if an AI constructed item doesn’t work with a screen reader or fails some other accessibility rule, that would be a significant negative consequence.
Also if AI proctoring raises a flag against a test taker due to their disability, that could also have a consequence.
What review is meaningful enough to be defensible?
So what do you need to do to have meaningful enough human oversight? This depends critically on the risks if AI gets things wrong. Using AI to generate formative feedback in low stakes quizzes is very different from using AI in high-stakes assessments.
What is needed will depend on an organization’s circumstances and the use cases and risk involved. For each ingredient above, here’s the kind of evidence a regulator, appeal panel or auditor would look for:
- Policies and procedures around human oversight, and good governance around these
- Evidence that reviewers have the appropriate competence and domain knowledge to be able to review the AI
- Reviewers should have guidance to follow, including criteria on when to override the AI
- Reviewers should have documented training on the system and on the risks of automation bias
- An effective user interface for reviewers to be able to see the input and output of the AI so that reviewers can judge the AI’s decision
- Evidence that reviewers have time to do the review
- Evidence that reviewers can and do catch AI errors
- Quality control and bias monitoring of the oversight
Detailed guidance on human oversight in assessment is available from the ATP document: Human Oversight of AI in Assessment: Considerations for Responsible Use published in June 2026.
Summary
We hope this information is helpful to those considering how to deploy human oversight on AI and assessment. You may also find our parallel article helpful: Why human oversight in AI-based assessments matters for bias, trust, and accuracy, and AI & Assessments: When the Stakes Are High.
FAQs
No, there are not universally agreed rules. The ATP document mentioned in this article is a good source of best practice.
Likely yes if the stakes of the assessment are high and if test takers are in Europe. Not required for low stakes AI scoring that does not have significant effects on test takers.
Yes for higher stakes use cases that fit into the high risk rules under the EU AI Act.
Potentially yes. They have extraterritorial effect in relation to test takers in Europe. But please check with your lawyer.
At the time of writing, there are two standards under development. ISO are developing ISO/IEC 42105 Information technology — Artificial intelligence — Guidance for human oversight of AI systems. And CEN are developing prEN 18229-3 which is the AI trustworthiness framework – Part 3: Human oversight. Both may be useful once finalized but are focused across sectors and will only have some value for use in assessment.
Disclaimer: This document provides general information about AI and does not contain legal advice for any individual organization’s specific circumstances. Readers should consult their own legal experts for legal advice on the matters addressed in this document.
John Kleeman is the Founder of Questionmark, Co-Chair of the ATP’s Responsible AI in Assessment Subcommittee and one of the authors of the ATP’s Human Oversight of AI in Assessment guidance (June 2026).