Does generative AI harm learning or support it? The honest answer is that it depends on the task, the stage of learning and what the assessment is actually meant to demonstrate. That can sound like an evasion until someone turns it into a framework that educators can use. The AI Assessment Scale (AIAS), developed by Mike Perkins, Leon Furze, Jasper Roe and Jason MacVaugh, is one such framework. Its latest revision is worth a closer look because it changes how the scale itself should be understood.
Table of Contents
ToggleWhat does the AI Assessment Scale look like now?
The current version, AIAS 2.1, keeps five levels: No AI, AI Planning, AI Collaboration, Full AI and AI Exploration. The AIAS website reports use in more than 350 institutions worldwide and translations into more than 30 languages. TEQSA has also identified the AI Assessment Scale as one option institutions can use when implementing generative AI in assessment.
The important change is the framing. Each level describes a different kind of assessment task, not a higher or lower degree of permission. No level is supposed to be better than another. Only Level 1 carries a padlock because it is the level where AI use is excluded and secure assessment conditions are required.
Why did the AI Assessment Scale change?
The first version used traffic lights: red, amber and green. The problem was not the colours themselves. It was the hierarchy they suggested. Red looked like the undesirable option and green like the desirable one, even though a well-designed Level 4 assessment is not inherently better than a well-designed Level 1 assessment.
The deeper issue appears in the authors' 2025 revision paper. They explain that broad permission to use AI can become difficult to control because educators cannot reliably verify where a student's own contribution ends. The revised framework also removes the previous requirement for students to submit an appendix of original work, reflecting the limits of trying to verify authorship through detection alone.
That is a meaningful change in direction. The AI Assessment Scale is no longer best understood as a permission ladder. It has become a framework for assessment redesign: deciding what students need to demonstrate, what role AI can play in that task and how the resulting assessment remains valid.
The gap between AI rules and what educators can actually verify
Corbin and colleagues described a similar problem in Assessment & Evaluation in Higher Education in 2025. Their study looked at how students and teachers understand acceptable and unacceptable AI use in assessment and found that simple boundaries often become difficult to apply in practice. They argue that AI assessment policy needs to consider not only what is permitted, but also whether the rules can actually be enforced and whether they preserve meaningful learning.
That connects closely to what we have been writing about throughout this series. Bavaria's draft law takes a legal route: universities may not simply exclude AI from unsupervised written examinations and would have to define how AI use is documented. TEQSA points towards learning process evidence. The AIAS takes a pedagogical route. It allows educators to define different kinds of AI use, but the assessment still has to demonstrate the learning it claims to measure.
Three different routes therefore meet at a similar practical problem. A pedagogical framework, a German legislative proposal and an Australian higher education regulator all leave educators with the same task: the stated rules around AI in assessment only matter if the assessment design and evidence make those rules meaningful.
Which assessment tasks should stay AI-free?
The familiar comparison is the calculator. We restrict calculators when students are still building certain mathematical skills, the argument goes, so perhaps AI should be restricted in the same way. The research gives us a more useful answer. Hembree and Dessart's meta-analysis found that calculator use generally improved pencil-and-paper skills, with an important exception in Grade 4, where sustained use appeared to hinder some basic skills. Ellington's later review of 54 studies found no evidence that calculator use generally prevented skill development.
The exception is actually the useful part. A tool should be restricted when using it gets in the way of the specific capability the assessment is meant to build. That is a much stronger principle than simply declaring a technology good or bad.
The same logic applies to AI in education. In a large field experiment involving nearly 1,000 high-school mathematics students, Bastani and colleagues found that unrestricted GPT-4 access improved performance during practice but was followed by a 17% reduction in performance on a later unassisted exam. A different AI tutor, designed to provide hints rather than simply give solutions, avoided that penalty. The technology was not the only variable. The way it was incorporated into the learning task mattered.
That is the AI Assessment Scale argument in one experiment.
The decision belongs in the classroom
Which AIAS level fits a particular assessment is not really a decision for a ministry or regulator. It belongs much closer to the classroom, with the educator who knows the cohort, the discipline and the purpose of the task. A first-semester essay designed to build basic argumentation skills may reasonably sit at Level 1. A capstone project that asks students to analyse a complex dataset may be less meaningful if AI support is artificially excluded.
That requires two things: a clear framework and a way to make the chosen level credible. AIAS now provides the framework with considerably more clarity. The harder part comes next. A Level 1 assessment needs secure conditions. The levels above it need enough evidence to show how the work was produced and where AI contributed. Neither outcome happens simply by printing an AIAS level on the assignment brief.
AI, cheating and academic integrity are not solved by a label
This is where the distinction between assessment design and academic integrity becomes important. Clear AI rules can reduce uncertainty, but they do not by themselves prevent cheating, plagiarism or inappropriate delegation of the work to AI. A prohibition is only meaningful when an institution has a reasonable way to uphold it. Likewise, allowing AI does not remove the need to establish who did the work and whether the intended learning actually took place.
That is why the strongest assessment designs increasingly combine clear expectations with evidence. The question is not simply whether AI was available. It is whether the assessment can still support a defensible judgment about the student's learning and contribution.
Sources and further reading
This article draws on the following research and guidance:
- Perkins, Roe & Furze (2025): Reimagining the Artificial Intelligence Assessment Scale — the paper behind AI Assessment Scale 2.1 and its shift towards assessment redesign.
- Bastani et al. (2025): Generative AI without guardrails can harm learning — a large field experiment examining how different forms of GPT-4 support affected learning and later unassisted performance.
- Corbin et al. (2026): The wicked problem of AI and assessment — an analysis of why generative AI creates a broader assessment-design problem rather than one that can be solved through a single policy or detection tool.
Making a chosen assessment level real
That second part is what we work on at Mentafy: documenting the writing process so that personal contribution, AI use and the development of the work become visible. The aim is simple. When an educator chooses an assessment level, the evidence should support it. See our toolset at a glance, or get in touch.





No comment yet, add your voice below!