top of page

Assessment Redesign 101: How to Make Exams AI Can't Just Answer

Updated: Jul 21

There is a game being played in higher education right now, and detection is losing.


AI-related academic misconduct now accounts for 60–64% of all cheating cases in higher education globally — a figure that has shifted dramatically in just two academic years (Feedough, 2025). Detection tools flag false positives, penalise non-native English speakers, and generate legal risk for institutions. Meanwhile, students iterate faster than any policy committee can meet.


The response from most institutions has been reactive: tighten the policy, buy detection software, update the handbook. But if the tool a student is using is already baked into their operating system, their browser, and their phone keyboard, detection alone will never be enough.


The more durable answer is not better surveillance. It is better design.


In Southeast Asia, the pressure is particularly acute. Across Indonesia, Malaysia, the Philippines, Thailand, and Vietnam, 90% of university students now use generative AI — saving an average of 5.3 hours per week on academic tasks (Deloitte Asia Pacific, 2024, n=2,903 students). In Singapore, 93% of university students report having used AI for assignments or study tasks (Studiosity/British Council, 2025). AI-related searches across the region have grown 11× since 2020 (e-Conomy SEA 2024, Google/Bain/Temasek). The question is no longer whether your students are using AI. It is whether your assessments are designed to surface genuine learning regardless.


Why Detection Is a Losing Game


AI detection is probabilistic, not definitive. Turnitin's own research acknowledges significant false positive rates, and courts in several countries have already ruled against institutions that penalised students on the basis of AI detection scores alone.


The problem runs deeper than false positives: students systematically underreport AI use. Research at Thai Nguyen University in Vietnam found that actual AI-assisted cheating prevalence is nearly three times higher than what students admit in direct surveys — measured via anonymous list experiments that remove social desirability bias (Springer Education and Information Technologies, 2024). In Malaysia, only 20% of undergraduates disagreed that using ChatGPT for schoolwork constitutes cheating; the other 80% acknowledge the integrity issue, yet usage accelerates (NCBI/PMC, 2024, n=443). This means any institution calibrating its response to self-reported data is likely working with a significant undercount.


More fundamentally, detection assumes the problem is the tool. It is not. The real problem is assessment design that never required a student to demonstrate their own thinking in the first place. An essay prompt that asks "Discuss the causes of the 2008 financial crisis" was already answerable without much original thought — AI has simply made that more visible.


The Australian government's Tertiary Education Quality and Standards Agency (TEQSA) puts it plainly: institutions need to "redesign assessments toward process-based, oral, and applied evaluations that are harder to outsource to language models" (TEQSA, 2025). The University of Sydney has already embedded a two-lane assessment framework distinguishing between supervised secure tasks and open tasks requiring real-world reasoning — and is mapping every degree programme to this structure for 2026 (Teaching@Sydney, 2025).


The direction is clear. Here is how to move in it.


Five Assessment Strategies That Require a Human


1. Staged Submissions with Process Evidence


Instead of collecting one final assignment, break the task into checkpoints: a topic proposal, an annotated source list, a rough draft with reflective commentary, and a final version. Grade each stage.


This works because AI produces polished outputs, not learning trails. A student who genuinely engaged with the material will show evolution — ideas developing, arguments shifting, sources being reconsidered. A student who fed a prompt into a language model will show a suspiciously complete draft appearing fully formed at checkpoint three.


At the AUN-QA network level, process-oriented assessment has been identified as one of the core responses to AI in ASEAN higher education (AUN-SEC, 2024). Polytechnics in Singapore have begun requiring students to submit timestamped drafts alongside final project reports — a low-tech solution that costs nothing to implement.


The APRU's 2025 Generative AI in Higher Education Whitepaper — drawing on 70+ participants across Asia Pacific, including institutions from Indonesia and the Philippines — reinforces this direction through its CRAFT Framework: Culture, Rules, Access, Familiarity, and Trust. Its central finding: effective assessment reform requires students as co-design partners from the outset, not as passive recipients of updated policies (APRU, January 2025). A parallel systematic review of ChatGPT use in ASEAN higher education (Wiley, 2025) found that most existing institutional policies focus on prohibition rather than design — and recommends shifting toward accountability protocols that build student agency.


Try this: For a 2,000-word business report, require a 200-word proposal (Week 2), an annotated bibliography (Week 5), a 500-word section draft with tutor feedback (Week 8), and the final submission (Week 12). Weight the stages at 30% of the overall grade.


2. Oral Defences and Viva Components


An oral defence requires a student to explain, justify, and respond in real time. It cannot be outsourced to any tool that is not in the room with them.


Stanford's Academic Integrity Working Group, formed in 2024, specifically recommends oral exams and in-class formats for high-stakes assessments where demonstrating genuine understanding matters (Stanford, 2025). Research published in Frontiers in Education in 2024 found that viva voce examinations are among the most effective AI-resistant formats precisely because they evaluate comprehension in the moment, not polished text produced before the moment (Frontiers in Education, 2024).


Oral defences do not have to replace written work — they can complement it. A 5–10 minute conversation following a written submission, where the student is asked to explain one section or respond to a challenge, changes the entire dynamic of the assignment.


Try this: After a group project submission, hold individual 8-minute oral Q&A sessions. Ask each student to explain one decision made in the project, what they would change, and why.


3. Reflection Journals Tied to Lived Experience


Reflection is difficult to fake because genuine reflection references specific, personal, situated experience — the kind of granular detail that AI cannot generate without access to the student's actual life.


A weekly reflection journal tied to internship tasks, lab sessions, community service learning, or a workplace visit requires students to write about what they specifically observed, felt, and concluded. The more localised and experiential the prompt, the harder it is to outsource.


Research in the Springer AI and Ethics journal (2025) identifies reflective writing as one of the assessment forms most resistant to AI substitution, noting that "personal, subjective experiences are challenging for AI to replicate" (Springer, 2025).


Try this: After each practical lab session, have students submit a 300-word reflection answering three questions: What specifically surprised you today? What did you do when something went wrong? What would you do differently next time? Grade for specificity, not correctness.


4. Performance-Based and Scenario Tasks


Ask students to do something, not just write about it. Design a lesson plan and teach a 10-minute segment. Conduct a client interview and record it. Build a working prototype. Troubleshoot a real system fault. Present findings to a panel that includes an industry practitioner.


Performance tasks assess competence, not knowledge recall. They require students to integrate understanding in a live, observable context. They also produce evidence that is inherently personal — a video recording of a teaching demonstration, for instance, cannot be generated by any AI tool currently available.


This matters in the SEA context because the skills gap is accelerating. The e-Conomy SEA 2025 report projects that job skills across Southeast Asia will change by 72% between 2016 and 2030 — compared with only 40% change in the preceding decade — placing enormous pressure on institutions to produce graduates who can demonstrate applied competence, not credential-adjacent text (Google/Bain/Temasek, 2025). Performance tasks generate the kind of evidence that proves that competence.


Working Futures' 2025 report on authentic assessment in the AI era documents institutions using portfolio-based evidence, industry-linked projects, and observable performances as the primary shift away from AI-vulnerable traditional assessments (Working Futures, 2025).


Try this: Replace the final essay in a communications module with a 3-minute recorded pitch to a simulated client, followed by a written brief explaining the strategic choices made. Assess the pitch and the reasoning separately.


5. Constrained In-Class Tasks with AI Permitted


Rather than banning AI entirely, design in-class tasks where AI use is explicit, visible, and evaluated. Give students a 90-minute in-class session, provide access to a language model, and ask them to complete a task that requires critical judgment at every step — evaluating AI outputs, correcting errors, and justifying their final decisions in writing.


This approach, part of what Perkins et al. (2024) call the AI Assessment Scale, grades the quality of human reasoning about AI output rather than AI output itself. Students who understand the subject will catch errors, refine arguments, and produce better work. Students who do not will accept everything the model generates — and this becomes visible immediately.


Try this: Give students a 600-word AI-generated case study with three deliberate factual or logical errors embedded. Ask them to identify the errors, explain why each is wrong, and rewrite the relevant sections. Time-limited. In class.


What SEA Governments Are Now Requiring


The pace of regulatory change across Southeast Asia makes a reactive institutional posture increasingly untenable.


In February 2026, the Philippines Department of Education issued DepEd Order No. 003 — 49 pages of foundational guidelines requiring students to disclose AI use in all assignments and prohibiting manipulative AI tools targeting minors. In March 2026, Indonesia formalised its response through a Joint Ministerial Decree signed by seven cabinet ministers, setting age-appropriate rules on permitted AI tools from early childhood through university — one of the most comprehensive government-level AI-in-education frameworks in the region.


Malaysia's Budget 2025 earmarked RM 50 million for AI education across all research universities, up from RM 20 million the year prior. Thailand's Ministry of Education partnered with Microsoft (June 2025) to train 4,500 teachers reaching more than 400,000 students, targeting 1 million AI-literate citizens by end of 2025. Singapore's MOE EdTech Masterplan 2030 embeds AI into four formal learning modes — including learn to use AI and learn beyond AI — backed by a commitment of over SGD $1 billion to national AI capability development.


UNESCO data reinforces why institutional response cannot wait: nearly two-thirds of higher education institutions surveyed globally already have AI guidance or are developing it (UNESCO, 2025), up from less than 10% just three years prior. The institutions that move now are setting the floor that late movers will eventually be required to meet.


Getting Started Without Overhauling Everything


You do not need to redesign your entire module in one semester. Here is a practical sequence:


Start with one assessment. Pick the one most obviously vulnerable to AI (usually: open-ended essay, reflection without specificity, generic case study). Redesign it using one of the strategies above.


Add a single process checkpoint. Even one intermediate submission requirement changes student behaviour significantly.


Be transparent with students. Tell them why you are redesigning assessments. Most students welcome clarity. Explaining that the goal is to ensure their learning is visible — not to catch them — shifts the dynamic.


Use rubrics that reward process. If your marking criteria only reward the final product, you cannot credibly grade process. Build criteria for drafts, reflections, and oral responses that stand on their own.


Coordinate with your department. Isolated assessment redesign is less effective than a coordinated programme-level approach. Even a short department conversation about which assessments are highest risk is a useful starting point.


The Bigger Picture


AI is not going away, and the students entering universities in 2026 have grown up using it. The question is not whether they will use it — most already do (Digital Education Council, 2024). The question is whether your assessments are designed to reveal genuine learning regardless.


Detection chases a moving target. Design creates a stable one.


The educators who will navigate this period best are not the ones who build higher surveillance walls. They are the ones who design assessments worth doing — tasks that require a student to be present, to think on their feet, to draw on experience that is genuinely theirs.


That is not a concession to AI. That is what good assessment has always looked like.


Want to see what AI-ready assessment design looks like in practice? Noodle Factory's Walter platform supports educators in building adaptive, process-linked learning activities that generate evidence of student engagement at every stage. Learn more here.


Sources

bottom of page
Chat with Walter