A detector that says your essay is "74% AI" has not measured anything about your work. It has measured how predictable the surface of the text is, run that through a model, and reported the result. Whether that 74% is meaningful depends entirely on the detector's calibration curve.
What a calibration curve is
It's a plot, built on thousands of labelled examples, comparing the detector's stated confidence to the true rate of correctness. A well-calibrated detector hugs the diagonal — when it says 70%, it's right about 70% of the time. A curve that bows away from the diagonal is the visual signature of a detector that bluffs: it states high confidence on examples it is actually unsure about.
Three things to do with your score
- Treat it as evidence, not a verdict. A confident essayist reads the score the way they'd read a peer-review comment — as a signal to look at the passage more carefully.
- Look at the per-sentence breakdown, not just the overall number. snizzly's detector tells you which specific sentences pulled the score up. The overall number averages those — the sentence-level view is where the editing happens.
- Re-score after edits, and watch the direction of travel. A single number means little; a number that moved from 81 to 22 after you replaced the generic openings with specific ones tells you something real.
Before
[Detector score: 81% AI] In conclusion, the literature presents compelling evidence that suggests a significant correlation. Furthermore, additional research is needed to fully explore the implications of these findings.
After
[Detector score: 22% AI] The pattern holds across all three studies, though the effect size in Andersson (2019) is half what the earlier papers reported — possibly because the post-2015 sample includes the policy change. Worth following up.