Most institutions have spent the past year discovering that AI detection does not give them a straight answer. If that is where you find yourself, this piece is for you. It sets out why the familiar responses to AI in student work fall short, and what a defensible path forward looks like when the goal is to see understanding rather than to police it.
If you lead assessment or integrity at an institution, you have almost certainly had a version of this conversation in the past year. Students are using generative AI, some of it in ways your policies never anticipated, and the tools you bought to catch it are not giving you a straight answer. The natural next question is what to do instead, and the choices on the table can feel like three bad options. This article makes the case for a fourth one, and it starts by being honest about why the first three fall short.
The appeal of AI detection is obvious. It promises a simple score that tells you whether a piece of writing was produced by a machine. The problem is that the score is not reliable enough to act on, and the evidence for that is no longer in dispute.
The clearest signal came from OpenAI, the company behind ChatGPT. It built its own AI text classifier and then retired it in July 2023, citing a low rate of accuracy. In its own evaluation, the tool correctly identified only around a quarter of AI-written text while incorrectly flagging roughly one in ten pieces of genuinely human writing. When the organisation that builds the AI cannot reliably detect the AI, that tells you something important about the whole category.
The reliability problem gets worse, not better, for the students you most want to protect. A Stanford study published in Patterns tested seven detectors on essays written by non-native English speakers, all written entirely by humans, and found that the tools flagged more than sixty per cent of them as AI-generated. For essays by native English speakers, the same detectors made almost no such mistakes. The reason is structural rather than a bug that will be patched away. Careful second-language writing tends to use predictable vocabulary and formal sentence structure, which is exactly the pattern these tools read as machine-like. A tool that systematically misreads international students is not a foundation you can build an integrity process on.
The consequences are not hypothetical. Vanderbilt University disabled its detection tool in 2023 after working through the numbers, noting that even a 1% false positive rate, applied to the roughly 75,000 papers it processes each year, would mean around 750 students wrongly flagged. A growing number of universities have since switched their tools off for similar reasons.
The most serious case so far is Australian Catholic University, which logged close to 6,000 academic misconduct referrals in 2024, around 90% of them linked to suspected AI use. The university has said that around a quarter of all referrals were dismissed after investigation, and that any case resting solely on the detection report was dismissed immediately. It stopped using the tool in March 2025. Even on the university's own account, that is thousands of students pulled into a misconduct process, some of them waiting months for a decision while their results were withheld. This is the picture behind the question so many people are now asking, which is whether AI detection is accurate enough to rely on. The evidence says no.
This is the most common objection, and it deserves a direct answer rather than a sidestep. Having a detection tool in place is not the same as having a defensible integrity process, because the tool's output cannot carry the weight an integrity case requires.
A detection score is not verifiable evidence in the way a plagiarism match is. A plagiarism report points to a specific source document that anyone can read and check. An AI score points to no source anyone can verify, which is why so many institutions, including ones that still run these tools, have ruled that a detection flag cannot be the sole basis for a misconduct finding.
The clearest sign of the limits comes from the vendors themselves. Turnitin now says that because its false positive rate is not zero, instructors need to apply their own judgement, their knowledge of the student, and the context of the assignment before drawing any conclusion. The product has changed to match. Results in the low range are no longer shown as a number at all, because the company judged them too unreliable to display. When a vendor suppresses part of its own output, it is telling you something about how much weight the rest of it can bear.
So if you already use Turnitin, the honest position is that it can be one signal among several, but it cannot be the answer to AI in assessment. You still need an approach that actually shows you whether a student understands their own work.
Detection is the first of those options, and the case against it is above. When it fails, most institutions reach for one of two familiar fallbacks, and both come at a real cost.
The second option is heavier surveillance, usually in the form of proctoring. It treats every student as a potential cheat, it raises genuine privacy concerns, and it damages the trust between students and the institution that healthy learning depends on. It also does nothing to tell you whether the student understood the material, only whether they behaved during a fixed window.
The third option is a retreat to pen and paper. Locking assessment into a supervised hall does make AI harder to use in the moment, but it gives up a great deal in the process. It narrows what you can assess, it pushes you back toward memory and time pressure rather than genuine understanding, and it quietly abandons the more authentic, real-world assessment work that programmes have spent years developing. Choosing security by giving up on quality is not much of a choice.
That leaves the fourth option, and it is the one grounded in pedagogy rather than enforcement. Instead of trying to prove a negative, that a student did not use AI, you design assessment that surfaces a positive, which is whether the student genuinely understands their work. The most direct way to do that is to have the student explain their reasoning out loud and respond to questions in real time. This is what oral assessment does, and it changes the problem entirely. A generative model can write a flawless essay, and it cannot sit in the room and think in the student's place. When understanding has to be demonstrated in conversation, the question of which tools a student used at home stops being the thing you have to police.
Oral assessment is not a workaround. It is a well-evidenced assessment method in its own right, which is what makes it defensible in a way detection never was.
Research on oral formats is consistent about one thing in particular, which is that reliability comes from structure. When every student answers a comparable set of questions, when marking runs against a clear rubric, and when assessors have been briefed on how to apply it, oral assessment holds up well and is especially strong at measuring depth of understanding and the ability to apply theory to a real situation. An unstructured chat does not give you that, so the design work matters.
What follows from that structure is a process you can actually defend. The judgment stays with the educator, based on evidence anyone can review, which is the opposite of an opaque score you have to take on faith. It protects academic integrity by design rather than by suspicion, because there is nothing to fake when a student has to reason on the spot. And it treats students as capable people demonstrating what they know, rather than suspects to be screened, which is far better for the relationship between a student and their institution.
The fair challenge to all of this is workload. Speaking to every student individually sounds like something that works for a seminar of fifteen and collapses at three hundred, and if oral assessment meant scheduling and sitting through every conversation by hand, that would be true.
This is where the design of the assessment does most of the work. Orals do not have to be long to be revealing, because a focused set of questions about a student's own submitted work will tell you very quickly whether they can account for it. They do not all have to happen live, since asynchronous responses recorded against the same rubric can be reviewed when it suits the marker. They also do not have to apply to every task in a module, because using orals at the points where the stakes are highest gives you most of the assurance for a fraction of the effort. The practical question is not whether you can afford oral assessment, but where in the programme it earns its place.
The choice is not really between three imperfect enforcement tactics. It is between chasing an unreliable score and building assessment that shows you what a student actually understands. Detection asks technology to answer a question it cannot answer. Oral assessment asks the student, which is where the answer has always lived.
At FeedbackFruits, our Oral Assessment & Practice solution gives institutions a defensible, rubric-based way to do exactly this, at a scale that works for real cohorts and inside the LMS your teachers already use. Because judgment stays with the educator and every grade rests on evidence you can review, it stands up to scrutiny in a way a detection flag never could. Pairing it with Oral Practice lets students build confidence with the format first, so the assessment is fair as well as rigorous.
If you want the wider case for why orals show genuine understanding, start with our guide to oral assessment in higher education. If scale is your main concern, our free ebook Oral Assessment at Scale works through the full approach.