top of page

Academic Integrity in the Age of AI: Why Detection Is the Wrong Battle


A Higher Education Policy Institute survey from early 2025 found that 88% of university students reported using generative AI for assessments (HEPI, 2025). That's up from 53% the year before.

The institutional response, in many cases, has been to deploy AI detection tools. Run the submission through the detector. Flag anything above a threshold. Investigate.

The problem is that this strategy is failing, and in some cases, actively backfiring.


The detection trap


AI detection tools work probabilistically. They don't identify AI-generated text with certainty; they identify text that resembles AI-generated text. The difference matters enormously.

Australian Catholic University reported nearly 6,000 AI cheating allegations in 2024, roughly 90% of all its academic integrity cases (Futurism, 2025). Around 25% of those referrals were dismissed after investigation. Students were accused, put through formal processes, and cleared because the detector was wrong.

False positives don't just create legal and pastoral problems. They erode trust. Students who are falsely accused become understandably hostile to institutions they feel surveilled by. Educators who flag a student on algorithmic grounds and are then proven wrong become reluctant to raise integrity concerns at all.

Several universities that adopted detection tools early have since quietly reversed course. The arms race is one they can't win: as detection tools improve, AI writing tools simply adjust to avoid detection signatures. This is a game without an end state.


The better question


Here's the question that actually matters: not "did this student use AI?" but "does this assessment still test what we think it tests?"

If a student can complete your assessment to a passing standard using an AI tool without understanding the material, that's an assessment design problem. The AI didn't create the problem. It revealed it.

The most useful thing detection-era pressure has done is force educators to ask whether their assessments were ever really measuring what they intended to measure. Many weren't. An essay that asked students to summarise and synthesise secondary sources was always somewhat gameable: by patchwork plagiarism, by outsourcing to a tutor, or now by an AI. The assessment design, not the technology, is the root issue.


What redesigned assessment looks like

Assessments that are hard to fake with AI tend to share a few characteristics: they require specific, contextualised knowledge; they involve a process, not just an output; or they happen in a setting where the thinking is visible.

Oral components. A student who submitted an AI-generated essay but understands the material will be able to discuss it. A student who doesn't will struggle when asked to expand on a point, defend a claim, or respond to a counter-argument. Oral components don't need to replace written work. Even a short follow-up conversation can dramatically raise confidence in authorship.

Staged submissions. Requiring a problem statement, a draft, a peer review response, and a final version means the process of thinking is documented. AI can contribute at stages, but stitching together a consistent intellectual thread across multiple drafts is harder to fake than producing a single polished output.

Reflection journals and process notes. Asking students to document their thinking as they work (what they tried, what didn't work, what they changed and why) creates a record of reasoning that AI can't generate retroactively.

In-person assessments. Some programmes are simply returning to supervised, in-person work for high-stakes assessment. It's not always practical, but for terminal assessments it's reliable.

On the AI tool side, platforms designed with integrity in mind approach the problem differently from detection tools. Noodle Factory's Walter, for instance, keeps student interactions grounded in the educator's approved course content — not the open internet. Students can use AI support, but the support is bounded by what the course actually covers. That's a different approach to the integrity question: not catching misuse after the fact, but designing the AI experience so that boundaries are built in from the start.


What this requires of institutions

Assessment redesign at scale is not a quick fix. It requires time, professional development, and crucially, permission for educators to experiment and iterate. Policies that penalise any AI use without distinguishing between misuse and legitimate engagement make that harder.

The institutions navigating this well are the ones treating it as a curriculum design opportunity rather than a disciplinary crisis. They're asking: if AI is now part of how students work, what does rigorous assessment look like in that context? That's a harder question than "did they cheat?" but it's the one that actually leads somewhere useful.


Sources


Noodle Factory builds AI tutoring tools for universities, polytechnics, and K-12 schools across Southeast Asia. Learn more at noodlefactory.ai.

bottom of page
Chat with Walter