What Actually Happened
Here's what actually happened over the last two years. AI got good at writing. Not Hemingway good. Not even good-blogger good. But good enough that the volume dial on content went to eleven and stayed there. Marketers noticed. Educators panicked. An entire industry of AI detection tools materialized overnight, promising to separate the human from the machine.
The problem is they can't. Not reliably. Not consistently. Not even close.
Three Detectors, Two Verdicts
I ran the same piece of writing through three major detectors last week. GPTZero called it 100% AI. Quillbot called it 0% AI. ZeroGPT agreed with Quillbot. Same text. Three tools. Two completely opposite verdicts. If you're using these tools to make consequential decisions — failing a student, rejecting a job applicant, pulling a piece of content — you're making those decisions on a coin flip dressed up in a confidence score.
The Math Behind It
The math behind this is actually well understood. Researchers call it the Sadasivan Impossibility Theorem. As AI writing gets better, the statistical distance between human text and AI text shrinks toward zero. The best possible detector eventually performs no better than random guessing. We're not there yet — but we're closer than the companies selling detection software would like you to know.
The False Positive Problem
What makes it worse is the false positive problem. Stanford researchers tested leading GPT detectors against essays written by non-native English speakers. The false positive rate hit 61.3%. Over half of genuine human writing flagged as AI — simply because the writers used cleaner grammar and simpler sentence structure. Disciplined writing looks like AI to a tool that was trained to treat complexity as humanity.
GPTZero — the most sophisticated and most widely used detector — has a specific bias worth understanding. It measures perplexity and burstiness: how surprising each word choice is, and how much sentence length varies. Short, punchy, declarative sentences score as low perplexity. Fragments score as low perplexity. Simple direct prose scores as low perplexity. Hemingway would fail GPTZero. So would most good copywriters.
Meanwhile the sentences that score as most human? Familiar quoted phrases. Common transitions. The exact constructions that every writing teacher tells you to cut.
The detector has it backwards. It's penalizing good writing and rewarding mediocre writing, because mediocre writing is what most of the human training data looked like.
AI Slop Is Real. That's Not the Point.
None of this means AI writing is good. Most of it isn't. The slop is real — generic, hedged, comprehensively unhelpful content that covers all the expected angles and illuminates exactly nothing. You know it when you read it. The problem isn't that AI can write. The problem is that AI can write the way a committee would approve, and committees have been approving bad content for decades.
The Actual Question
The actual question worth asking is simpler and harder to automate: is this worth reading?
Does it have a point of view? Does it say something specific or does it hide behind abstractions? Does it give you one real thing you can use, or does it give you six things that are technically true and practically useless? Is there a person behind it — someone who made a decision about what mattered and what didn't — or is it the written equivalent of a loading screen?
Those questions don't care whether a human or a machine produced the first draft. They care about the output. And they're the questions editors, readers, and anyone who actually values their time should be asking.
The tools will keep getting better at detecting AI. The AI will keep getting better at evading detection. That arms race has a known mathematical endpoint and it isn't detection winning. Meanwhile the content that actually matters — specific, opinionated, useful, honest about what it doesn't know — will keep being written by whoever is willing to do the work of thinking clearly, human or otherwise.
Stop asking who wrote it. Start asking whether it earned your attention.
That's the only metric that ever mattered.