The Massachusetts Institute of Technology advised its faculty not to rely on AI text detectors, and Stanford and Harvard are on the same path. The evidence shows these tools fail and their mistakes fall on the most vulnerable.
Underneath lies a striking academic hypocrisy: using an AI tool to police the use of AI tools. Meanwhile, humans, exposed daily to the prose of the machines, are starting to write like them.
When the algorithm accused Cervantes
By: Gabriel E. Levy B.
Since the arrival of ChatGPT as the first massive generative AI model in November 2022, practically every university on the planet and the newsrooms of traditional media decided, hastily and without solid evidence or experimental field work, to deploy software and algorithms that promised, and worse still keep promising, to distinguish between human prose and texts created by algorithmic Artificial Intelligence systems built on LLM technology, or large language models, such as ChatGPT, Claude, Grok, DeepSeek or Gemini.
Generally, although they do not always work this way, these programs return a percentage of AI use in a text, and that number, for a great many professors around the world, is what defines suspicion for or against the student. There are dozens of documented cases of students who were sanctioned or expelled for using it and who later proved, in court or on appeal, the imprecision of these systems; many of them even managed to prove that they themselves had written the text.
Possibly the most talked-about case among the many errors these systems make occurred in December 2024, when the Spanish writer Pedro Torrijos submitted the first paragraph of Don Quixote to a detector and got a startling verdict: an 86% probability of being the work of artificial intelligence. In other words, according to detection programs of this kind, such as Pangram or Copyleaks, Cervantes would have used ChatGPT or Gemini more than 420 years ago to write his masterpiece. The paradoxical and at the same time ridiculous part is not the incident itself, but the fact that professors use that tool to decide whether a student used AI in the classroom.
But this is not the only case: a few months later, science communicator Fernando Bujedo repeated the exercise with the tool JustDone and got 78% for Cervantes and 89% for his own master’s thesis, written in 2017, when generative AI did not exist.
Yet the most significant fact is that all the detection systems, such as Pangram, Copyleaks, ZeroGPT or JustDone, and the platforms themselves, such as ChatGPT or Claude, warn in their terms of use that they do not guarantee their results, that there is a high probability of error in their detections and that they must not be used as an objective criterion to determine the use of AI. In other words, they define themselves as support tools, not instruments of judgment.
All of this throws an unavoidable question at us:
Can a tool capable of mistaking Cervantes for a chatbot decide a student’s academic fate?
Two warnings from Cambridge
In August 2023, the educational technology unit at MIT Sloan published a guide whose title leaves no room for doubt: “AI detectors don’t work and it is better to do something else.”
MIT Sloan scholar Anna Wright documented at length, with clear evidence and academic rigor, the high error rates of these programs, and proposed a different path with policies agreed in each course, dialogue with students, redesigned assignments and written accounts of the writing process for each paper.
But the most important thing, and what works best, is asking the student to explain in their own words what is written in the text. This method, which contradictorily requires none of the very technology it seeks to punish, exposes in most cases when the student was lazy and used AI to write; and when the student has properly absorbed the knowledge, it means the academic purpose of learning was indeed fulfilled.
Three years later, in August 2026, the institute went further.
A special committee led by professors Eric Klopfer and Sam Madden, convened by president Sally Kornbluth, recommended against relying on these systems because they feed an arms race with so-called “humanizers”, services that erase the traces of AI.
The committee warned that these programs often mistake the writing of non-native English speakers and neurodivergent students, and MIT’s own disciplinary body stopped accepting a detector score as sufficient proof of misconduct.
MIT is not alone and this is no isolated case: Stanford’s institute for human-centered AI, one of the world’s most advanced in the complex and deep use of LLM models, urged people not to trust detectors it describes as unreliable and easy to fool, and Harvard’s Faculty of Arts and Sciences called it inadvisable to rely on automated detection and declined to license such a tool for its courses.
There is, besides, a deeper contradiction that no software update can fix:
we use an automation tool to detect the automation of a text. We delegate to an algorithm the task of deciding whether another algorithm wrote something, and we turn suspicion into an industrial procedure that dispenses with careful reading and with the voice of the accused.
The very logic we wanted to police ended up running the policing, and human judgment was reduced to validating the percentage on a screen.
The inconsistency reaches the disciplinary field itself: if a student deserves punishment for delegating an essay to AI, the professor who delegates grading to a detector commits the same offense, since they too used a tool to replace their academic duty, which is to read and judge their students’ work.
The perplexity trap
Detectors measure two statistical properties of language.
The first is perplexity, that is, how predictable each word turns out to be. The second is the variation in sentence length and structure, known as burstiness.
Formal, codified prose scores low on perplexity, and for that reason it looks, in the eyes of the algorithm, like the prose of an LLM or generative AI model.
Journalist Benj Edwards demonstrated the phenomenon in Ars Technica in July 2023, when GPTZero attributed fragments of the United States Constitution to artificial intelligence.
Edward Tian, founder of the company, explained that the document appears so many times in the training data of the models that they learned to imitate it.
Another detector, ZeroGPT, gave the Declaration of Independence a 97.93% probability of being artificial.
The numbers confirm the pattern.
OpenAI shut down its own classifier in July 2023 after admitting that it recognized only 26% of AI texts and falsely accused 9% of human ones, and a team led by researcher Debora Weber-Wulff tested 14 tools without any of them reaching 80% accuracy.
An imprecision that discriminates
The cost of those errors is not evenly shared.
A Stanford University team, headed by professor James Zou, published an experiment in the journal Patterns in July 2023: it ran 91 TOEFL essays, written by non-native English students, through seven well-known detectors. On average the programs falsely flagged 61.3% of the essays, and one of the tools flagged almost 98%.
Those same systems got it right on more than 90% of texts by American eighth graders. The statistical trap repeats itself: someone writing in a second language uses more common vocabulary and simpler sentences, and that plainness sets off the alarm.
The detector mistakes command of the language for authorship.
Vanderbilt University disabled Turnitin’s detector in August 2023 after calculating that, with 75,000 papers a year and a 1% error rate, the program could wrongly accuse some 750 students. Turnitin and GPTZero defend their improvements and insist that their scores do not prove misconduct, a nuance that does not always survive the trip to the classroom.
The inverted mirror
The other half of the story happens on our own keyboards. The more AI text we read, the more its style seeps into ours.
The Max Planck Institute for Human Development analyzed more than 730,000 hours of academic talks and podcast episodes and found that, since the chatbot appeared, the model’s favorite words gained ground in spontaneous speech: “delve” grew by 48% and “meticulous” by 40%. The footprint is deeper in scientific writing.
Dmitry Kobak and his team reviewed more than 15 million biomedical abstracts and concluded in Science Advances that at least 13.5% of those published in 2024 passed through a language model, with peaks of 40% in some fields.
A Cornell University study also showed that writing assistants push authors from India toward Western styles and erase cultural nuance. Even punctuation joined the dispute: the em dash, beloved by Dickinson and by Woolf, became suspected of betraying AI, and many writers now avoid it.
The circle closes: the models learned from human prose and today they shape it, while the next generation of systems will train on that standardized writing.
In short, MIT reached a blunt conclusion: AI detectors fail too often to sustain an accusation, and their mistakes punish the most vulnerable. Cervantes flagged as a chatbot captures that technical impotence better than any statistic. The real challenge is cultural, because machine prose is already rubbing off on ours, and an uncomfortable question hangs in the air: how many of us are already writing, without noticing, like the machine.
References:
- Wright, A. (2023). AI detectors don’t work. Here’s what to do instead. MIT Sloan Teaching & Learning Technologies. https://mitsloanedtech.mit.edu/ai/teach/ai-detectors-dont-work/
- MIT Ad Hoc Committee on AI Use in Teaching, Learning, and Research Training. (2026). Report. Massachusetts Institute of Technology. https://aiandeducation.mit.edu/report/
- Edwards, B. (2023). Why AI detectors think the US Constitution was written by AI. Ars Technica. https://arstechnica.com/information-technology/2023/07/why-ai-detectors-think-the-us-constitution-was-written-by-ai/
- Liang, W., Yuksekgonul, M., Mao, Y., Wu, E. & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), 100779. https://doi.org/10.1016/j.patter.2023.100779
- Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., Foltýnek, T., Guerrero-Dib, J., Popoola, O., Šigut, P. & Waddington, L. (2023). Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19(26). https://doi.org/10.1007/s40979-023-00146-z
- OpenAI. (2023). New AI classifier for indicating AI-written text. https://openai.com/index/new-ai-classifier-for-indicating-ai-written-text/
- Vanderbilt University. (2023). Guidance on AI detection and why we’re disabling Turnitin’s AI detector. https://www.vanderbilt.edu/brightspace/2023/08/16/guidance-on-ai-detection-and-why-were-disabling-turnitins-ai-detector/
- Yakura, H., Lopez-Lopez, E., Brinkmann, L., Serna, I., Gupta, P. & Rahwan, I. (2024). Empirical evidence of Large Language Model’s influence on human spoken communication. arXiv. https://arxiv.org/abs/2409.01754
- Kobak, D., González-Márquez, R., Horvát, E.-Á. & Lause, J. (2025). Delving into LLM-assisted writing in biomedical publications through excess vocabulary. Science Advances, 11(27). https://www.science.org/doi/10.1126/sciadv.adt3813
- Agarwal, D., Naaman, M. & Vashistha, A. (2025). AI suggestions homogenize writing toward Western styles and diminish cultural nuances. CHI 2025. https://dl.acm.org/doi/10.1145/3706598.3713564
- MuyComputer. (2024). Cervantes era una inteligencia artificial. https://www.muycomputer.com/2024/12/02/cervantes-era-una-inteligencia-artificial/
- Omnia. (2025). Pone a prueba un detector de plagios de IA y el resultado expone hasta a Cervantes. https://www.omnia.com.mx/noticia/363919
- Stanford HAI (Myers, A.). (2023). AI-detectors biased against non-native English writers. Stanford Institute for Human-Centered Artificial Intelligence. https://hai.stanford.edu/news/ai-detectors-biased-against-non-native-english-writers
- Harvard University, Office of Undergraduate Education. (2023). Generative AI guidance. Faculty of Arts and Sciences. https://oue.fas.harvard.edu/faculty-resources/generative-ai-guidance/



