For years the worry in classrooms was plagiarism. Now it's almost the reverse: proving that a person, not a machine, did the writing. Schools and publishers have reached for a new class of software — AI detectors — that promise to spot text made by tools like ChatGPT. The catch is that these programs are shaky, their own makers admit it, and people keep leaning on them anyway.

Start with where this came from. Anti-plagiarism software predates ChatGPT by a long stretch. Programs such as Turnitin check a submission against a huge library of web pages, journal articles and past work, hunting for sentences and phrases that line up too neatly. Turnitin even hands back a percentage meant to show how much of a student's text matches other sources. That figure was always slippery — a high score could mean deliberate copying or an innocent coincidence — and the false alarms pushed some teachers to quietly drop it.

The newer tools try to answer a harder question. Not "was this copied," but "was this written by a person at all." GPTZero, Pangram and Turnitin's own detector don't line your words up against a database. They feed your writing to AI models of their own and estimate the odds that no person wrote it. GPTZero explains that its software weighs a passage's word choice, rhythm and sentence shape, and it hunts for cues in a passage's length and overall tone that surface more often in machine-made writing. That's a far softer kind of evidence than a matched sentence you can point to online, and it trips over people who didn't grow up speaking English.

Even so, uptake has been quick. The nonprofit Center for Democracy & Technology ran the numbers: across 2024 and 2025, 43 percent of US teachers in grades six through 12 reached for AI detectors on a regular basis. Some schools didn't even choose to: universities already running Turnitin discovered the company had switched on AI detection automatically when it launched in 2023.

The vendors publish reassuring numbers. Turnitin says under 1 percent of the human writing it checks gets wrongly labeled as AI. Pangram pegs its own false-positive rate at 1 in 10,000, and GPTZero reports something similarly small. Yet the same companies hedge hard. Turnitin allows that its tool "may not always be accurate" and says it shouldn't be used to punish a student. Grammarly tells users they "should never rely on the results of an AI detector alone." GPTZero concedes that "no AI detector can ever truly be 100% perfect." OpenAI went furthest: in 2023 it pulled its own AI-writing detector because it simply wasn't accurate enough.

None of that has cooled the accusations. Online, people now trade charges of "sounding like AI," and cheap detection tools pour fuel on the pile. Some of the fallout has been severe. Last month the publisher Minotaur walked away from a book deal worth $2 million over fears that its writer, Jerry Falade, had turned to AI — a claim he flatly rejects.

The courtroom cases sharpen the stakes. Thierry Rignol, who is French, took Yale to court last year. A professor had put his final exam through GPTZero, concluded chunks of it were machine-made, and hit him with a failing mark and a suspension lasting a year. His lawsuit argues these tools have a documented habit of misfiring on writers who aren't native English speakers. February brought another win for a student: an Adelphi University undergraduate beat the school in court after a professor leveled the same charge. The filing never says which detector the professor relied on, though Adelphi does license Turnitin.

The bias worry isn't anecdotal. A Stanford study from 2023 reached the same conclusion: essays by writers whose first language isn't English got flagged as AI far more often than work from native speakers. Researchers warn the tools may also misjudge neurodivergent writers. So what are these systems actually keying on? By UCLA's account, they hunt for repeated words and phrases, prose that reads too formal or too loose, and sentences that don't quite make sense. QuillBot adds a measure of a text's "unpredictability," on the logic that machines gravitate to the most common, most expected phrasing. A uniform sentence structure counts against you too. The catch is obvious: plenty of humans just write that way.

The accusations keep flying regardless. Last week Jack Osbourne — Ozzy Osbourne's son — told his 3.5-million-plus social following that Kat Tenbarge, a reporter who contributes to The Verge, had used AI to draft a Rolling Stone piece, waving output from a detector called Getsolved as "proof." Tenbarge rebutted the accusation, both on video and on her website. Osbourne has neither retracted the claim nor taken the video down, leaving her to field the trolls. Stories like hers are piling up, and the accusers routinely ignore the fine print the detectors themselves attach.

So where does this leave things? A growing list of universities has decided the uncertainty isn't worth it. Yale, Johns Hopkins, Vanderbilt and Georgetown are among those that have switched off or fenced in AI detection. MIT puts it bluntly: "AI detectors don't work."

Rather than police the writing after the fact, many schools are rebuilding the assignments instead. The University of Chicago suggests asking students to read more slowly, chopping big essays into stages, and building in moments of reflection. Stanford points professors toward in-class assessments. MIT tells instructors to leave space for students to say, without penalty, that they leaned on AI for help.

The suspicion is spreading past the classroom, too. Substack has wired Pangram into the app so readers can sweep a blog for possible AI, and LinkedIn slapped a "seems like AI slop" button onto posts. Writers are pushing the other way: the Authors Guild now issues "Human Authored" certifications, and writers can tack on badges reading Not by AI or Written by Human. Wikipedia has published its own guide to spotting machine text — watch for writing that "puffs up" a topic or offers only "superficial analysis of information" — and banned AI-generated articles outright.

Here's the practical takeaway. A detector's verdict is a guess, not a finding, and the firms that sell them say as much in their own disclaimers. If you write — especially in a second language, or in a plain, orderly style — a false flag is a real risk worth guarding against, so keep your drafts, version history and notes. And if you're on the accusing end, a detector score is nowhere near proof. The reliable signal was never the software. It's the process behind the work.

This article was rewritten from reporting by The Verge AI. Source: The Verge AI