False positives: when a detector is wrong
Detectors get human writing wrong in predictable ways. If you've been flagged unfairly, the failure mode probably has a name — and a defence.
he unfair flag is not a rare bug. It is a known, predictable failure mode of every AI detector ever built, and any tool that pretends otherwise is selling you something. If your writing has been called AI when it wasn't, you are not paranoid — and the failure has a name.
Who gets flagged unfairly
Three groups carry most of the false positives. The first is non-native English speakers — a Stanford-led study published in Patterns (2023)00130-7) found that seven leading detectors misclassified more than half of TOEFL essays by non-native writers as AI-generated, while near-perfectly identifying essays by U.S.-born students. The careful grammar and the even rhythm that come from formal language study can mimic, on the surface, what unedited models produce. The second is strong essayists and technical writers: prose that has been edited until every sentence is tight loses some burstiness, and a detector that conflates polish with machine output will mistake one for the other. The third is anyone writing in a genre with a fixed structure — legal briefs, academic abstracts, lab reports — where the form itself forces an even register.
Calibration is the whole game
An honest detector publishes its calibration: out of every 100 texts it calls 80% AI, how many really are AI? A poorly calibrated detector can be confidently and consistently wrong about the same kind of writing — and if that writing happens to be yours, the score will hurt you again and again. This is why a confidence interval, a recall figure, and a stated false-positive rate matter more than a sharp three-digit percentage. A detector that won't tell you how often it gets it wrong is not a detector you can defend yourself against.
A confident wrong score is worse than a hedged right one. The bravery of a detector should be calibration, not certainty.
If you've been flagged unfairly
Several things help. Keep your drafting evidence — version history, scratch notes, a timeline of your sources. Most institutions take that seriously when a detector score is contested. Ask which detector was used and what its published false-positive rate is; if the tool can't answer, that itself is part of your defence. Run the same text through a second detector with different calibration, and read the spread, not just one number. And if a teacher or HR process is leaning entirely on a single percentage, point out, calmly, that we have not seen a peer-reviewed evaluation that supports treating any one detector score as proof on its own.
The other thing worth saying: the best protection over time is voice. Generic, even, register-flat prose is easier to mis-classify in either direction. The more your writing reads like you in particular, the harder it is for a detector — or a person — to mistake it for anyone else.
Up next in Writing Papers in the AI Era: Is using an AI humanizer cheating?.
What's the false-positive rate of a good detector?+
On long-form text, a well-calibrated detector should be in the low single digits. Anything that claims zero is either bluffing or hasn't tested on hard cases like non-native English and formal genres.
I'm a non-native speaker — what should I do?+
Keep drafting evidence, request a second opinion from a differently calibrated tool, and trust your own process. Honest detectors disclose this exact failure mode; the score is one input, not the verdict.
Can I appeal an unfair AI flag?+
Yes — and version history plus a calibrated second read are your strongest case. A single detector score is not, on its own, evidence that meets a fair standard of proof.