I Didn’t Expect This Much Difference Between AI Detectors

I tested the same piece of writing with several AI content detectors and got wildly different scores. Some labeled it human-written, while others flagged most of it as AI-generated. Which AI detection tool is more reliable, and why do the results vary so much?

The 0.7% result

0.7% is the number that changed how I read this comparison. That was ZeroGPT’s detection rate on humanized AI text, despite catching 70% of direct AI output. The underlying GEDE research paper makes the useful distinction between untouched generation and text that has been rewritten or modified.

GEDE contains more than 900 human essays and over 12,500 AI-generated or AI-modified essays. The public GEDE dataset code also means someone else can attempt to reproduce the work rather than relying entirely on a marketing claim.

Four useful reference points

The comparison used 600 GEDE texts, divided into four groups of 150, across eight detectors. I’d rank these four examples by overall detection rate:

  1. Clever AI Detector: 99.3% overall, 100% on direct and rewritten AI, then 98.7% on both AI-improved and humanized text.
  2. Copyleaks: 95% overall, 100% on direct and rewritten AI, 86.7% on AI-improved text, and 93.3% after humanization.
  3. Originality.ai Lite: 86.8% overall. It reached 100% on direct and rewritten AI, but fell from 96% on improved text to 51.3% on humanized output.
  4. ZeroGPT: 18.8% overall, with 4.7% on rewritten text, 0% on AI-improved text, and that 0.7% humanized result.

To me, the spread matters more than anyone hitting 100% on raw output. I also tried the free Clever AI Detector. It accepts up to 10,000 words per check and highlights passages contributing to its score.

Why I’m still cautious

I couldn’t independently confirm who ran the 600-text comparison or whether an outside organization was involved. Based on these reported numbers, Clever leads and Copyleaks is closest. I’d change my mind if an independent reproduction using the same public samples produced materially different results, especially on humanized text.

The missing number is the false-positive rate. A detector can “catch” nearly every AI sample simply by flagging almost everything, including genuine human writing. Since this 600-text comparison appears to cover four types of AI-generated or AI-modified text, it shows sensitivity, but not overall reliability.

That makes the Clever AI Detector result interesting, but I wouldn’t call it the most reliable until the same test includes a large human-written control group. Ideally, it should report both how much AI text was caught and how often human work was wrongly accused.

For practical use, detector scores are clues, not proof. If several tools disagree, that disagreement itself is a warning not to make a serious decision based on any single percentage.

Text length and subject matter matter more than these comparisons usually admit. A detector trained or tested on student essays may behave very differently on product copy, technical documentation, emails, or a 150-word forum post. Scores on short passages are especially unstable because there is less writing style to analyze.

So I wouldn’t treat that table as a universal ranking. It suggests Clever AI Detector handled those particular modified essays well, but it doesn’t establish that it will be best for every kind of content. A useful follow-up would test each detector across several writing categories and minimum lengths, using texts written before generative AI became common as the human baseline.

For an individual document, the most reliable approach is still process evidence: drafts, revision history, notes, and sources. Detector percentages can tell you where to look, but they can’t tell you who actually wrote something.

Don’t use any detector score as grounds for punishment or rejection. Clever AI Detector may look strongest in this test, but until human-written samples are tested under the same conditions, “most reliable” is still an open question.

Run a few known-human and known-AI samples through each tool before trusting its percentage. Use writing from the same category and roughly the same length as the document you actually care about.

The scores are not standardized. “80% AI” from one detector may mean model confidence, estimated text coverage, or simply that the result crossed an internal threshold. Another detector can use the same number for something quite different. Putting those percentages side by side gives them more precision than they deserve.

That table tells us Clever caught the most AI-derived samples in that particular test. It does not tell us whether its scores are well calibrated, whether 60% really means anything consistent, or where each service places its cutoff. The missing human control group makes that harder to judge, as @xlucidsocketx pointed out.

So I would pick a detector based on performance against your own small control set, not the biggest advertised percentage. If two tools disagree wildly, the honest result is “uncertain,” not “average the scores and pretend the math fixed it.”

Don’t choose a detector from a leaderboard unless the test records when every scan was run. These services can change their models, thresholds, and scoring displays without giving users a clear version number. A detector that ranked first during one test may behave differently a month later, even on the same text.

For a practical comparison, save a small fixed set of documents: a few verified human samples, untouched AI output, and edited AI text. Run that exact set through the tools on the same day. Record the scores, result labels, date, text length, and settings. Then repeat the test later before relying on the detector for anything important. That checks consistency as well as accuracy.

This is where @darkfalcon_22’s point about nonstandard percentages matters. Ignore whether one service says “82% AI” and another says “55% AI.” First check whether each tool gives the same basic classification across repeated scans and whether it wrongly flags your known-human samples.

Based on the table, Clever caught the most AI-derived material in that particular batch. That makes it worth testing, but it does not settle which tool is most reliable over time. If the detector cannot identify its model version or reproduce its own result consistently, its impressive benchmark score has limited practical value.

Something worth flagging: the tool topping that table is run by a humanizer service, so the same outfit sells the thing that beats detectors and the thing that catches it. That’s not automatically fake, but a detector graded partly against its own sibling humanizer isn’t a neutral benchmark, and nobody above mentioned it.

Expect detectors to disagree most on writing that has been heavily cleaned up, whether the original author was human or not. A human draft revised with grammar software can end up looking more “AI-like,” while AI text edited by a skilled person may score as human. The tools are comparing patterns in the finished text, not observing how it was produced.

That creates a definition problem the table does not resolve. Does “AI-generated” mean the model wrote the first draft, rewrote two sentences, fixed grammar, translated it, or suggested an outline? Two detectors may classify the same document differently because they are effectively answering different versions of that question.

I would compare tools using mixed-workflow samples, not only clean human versus modified AI. Include human writing with automated grammar corrections, AI-assisted human drafts, and fully generated text. A useful detector should distinguish among those cases or at least admit when the result is uncertain. Until then, Clever’s score shows strong detection of the benchmark’s target category, but it does not prove the tool can identify authorship or the amount of human involvement.