For three years, the standard line on AI detectors has been that they do not work reliably enough to base decisions on. That needs updating. But the update is narrower than the headlines suggest, and it does not rescue detection.

What has genuinely improved

In June 2026, a peer-reviewed study in the International Journal for Educational Integrity tested four commercial detectors on 160 documents of known origin. All four correctly identified every fully human text. Pangram performed best. A University of Chicago working paper reached a similar result a year earlier: on medium and long passages, Pangram's false-positive rate was effectively zero.

Compare that with the still widely cited 2023 evaluation, which found all 14 tested tools below 80% accuracy. The false-positive problem behind the horror stories – honest students accused based on a score – is largely solved for long texts by the best tools. That is a real achievement and deserves to be acknowledged.

What those tests do not measure

Now consider what is actually being tested: raw model output, or output prompted to sound more human.

That is not what reaches an examination office. A student outsourcing a paper does not paste raw ChatGPT into the submission box. They run it through a humaniser, check the result with a free public detector, and iterate until it comes back clean. As one academic told The Atlantic: "Basically it's an arms race."

For that scenario, the evidence is mixed. A 2026 University of Notre Dame study found that after a commercial humaniser, fewer than 4% of fully AI-written rewrites were still flagged. The Chicago paper, using a different humaniser, found one tool holding up well. Neither study is peer-reviewed. The peer-reviewed study, meanwhile, humanised its texts through prompting rather than a dedicated tool.

So the more accurate summary is not "detectors now work". They have become considerably better at a test that only partly resembles the cases examiners face. The most striking figures, including the widely repeated one false positive in 25,000, come from vendor testing.

Even perfect AI detectors would answer the wrong question

Suppose the arms race were settled tomorrow in the detectors' favour. You would still be left with this: a detector answers only one question – whether AI was involved in producing a text.

That question is steadily losing its force. Bavaria's government has proposed that universities may no longer simply prohibit AI in unsupervised written examinations, but instead must set rules for documenting its use. Australia's regulator advises institutions to allow AI within defined parameters. The AI Assessment Scale builds four of its five levels around permitted use.

Where AI is allowed, "AI was involved" is not a finding. It is a description of normal practice.

And it catches the wrong people

The Notre Dame study also examined what accurate detection does to honest students. Light, policy-compliant AI editing, the kind many policies now explicitly allow, was flagged 38% to 80% of the time.

Put that beside the humaniser result and the picture is clear. The authors state it plainly: "Honest AI-editing results in a higher sanction risk than humanizer-assisted evasion."

Anthropic says much the same about its own text watermark. It shows that Claude was involved, but "cannot distinguish 'Claude wrote this' from 'Claude heavily edited this'." Involvement is detectable. Contribution is not.

What a decision actually requires

Courts have already drawn this line. In January 2026, a New York court overturned a misconduct finding based on a 100% AI score. The issue was not whether the detector was right, but the procedure: the student was denied an advisor, his contrary evidence was not considered, and the same official decided both the case and the appeal. The finding was "without valid basis and devoid of reason."

German administrative courts have reached compatible conclusions. A detector result can be, at most, one indication among others; the burden of proof remains with the examination authority.

The vendors of AI detection do even agree. The general consensus is, that its score is the start of a conversation, not evidence. Even the 2026 study most favourable to detection concludes that these tools "should not be used as sole evidence in high-stakes decision-making."

The question that replaces it

What did this student contribute? That cannot be read from a finished text, however good the classifier, because the finished text is precisely what a humaniser rewrites. Answering it requires evidence of how the work came together: drafts, revisions, sources consulted, and the order in which decisions were made.

Australia's TEQSA calls this learning-process evidence and asks institutions to build the infrastructure to collect it, while honestly describing such data as "partial windows" into a student's activity.

We should be equally honest about its limits. A 2026 preprint showed that keystroke-based checks can be defeated more than 99% of the time by simply retyping AI output. Process data reliably shows that a human typed the text. On its own, it does not show that the human originated it.

That is not an argument against it. It is an argument for treating it as what it is: substantially better evidence than an output score, still evidence rather than proof, and something for a person to weigh alongside the work itself and, where relevant, a short conversation about it.

Detection still has a place in that stack. Where AI is genuinely prohibited and the text is long, a modern detector is a cheap and reasonable first screen. But it is one signal, it is part of an arms race it may not win, and it answers a question that examination rules are steadily making less relevant.


Making contribution visible

That shift is what we build for at Mentafy: documentation of how work actually takes shape, in the tools students already use, so that both personal contribution and AI use can be seen. See our toolset at a glance, or get in touch.

Sources

  • Van Vlasselaer, E., Van Droogenbroeck, F. & Spruyt, B. (2026). Who wrote this? Evaluating the reliability of AI detection tools in higher education. International Journal for Educational Integrity, 22, 16. DOI 10.1007/s40979-026-00226-w
  • Jabarian, B. & Imas, A. (2025). Artificial Writing and Automated Detection. Becker Friedman Institute Working Paper 2025-116, University of Chicago. (Working paper, not peer reviewed.)
  • Karr, C., Khvatskii, G., Hua, Y. & Chawla, N. (2026). Why AI Detection Fails for Academic Integrity. arXiv:2608.11256 (Preprint.)
  • Weber-Wulff, D. et al. (2023). Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19, 26. DOI 10.1007/s40979-023-00146-z
  • Anthropic (2026). How Claude's text watermark works.
  • Lodge, J. M. et al. (2025). Enacting assessment reform in a time of artificial intelligence. TEQSA.
  • Lodge, J. M. et al. (2026). Assuring quality learning in a gen AI-integrated future: The role of adaptive capabilities. TEQSA.
  • Perkins, M., Furze, L., Roe, J. & MacVaugh, J. (2024). The AI Assessment Scale. Journal of University Teaching and Learning Practice, 21(06).
  • Matter of Newby v Adelphi Univ., 2026 NY Slip Op 26021 (Sup. Ct. Nassau County, 28 January 2026).
  • Condrey, J. (2026). On the Insecurity of Keystroke-Based AI Authorship Detection. arXiv:2601.17280 (Preprint.)
  • Bayerisches Staatsministerium für Wissenschaft und Kunst (2026). Gesetzentwurf zur Änderung des Bayerischen Hochschulinnovationsgesetzes, cabinet decision of 23 June 2026. (Draft; not yet enacted.)
  • Oremus, W. (2026). The Pangram Backlash Unfolding on College Campuses. The Atlantic, 21 September 2026.

Recommended Posts

No comment yet, add your voice below!


Add a Comment

Your email address will not be published. Required fields are marked *