“The detector said 87%” is not the closing scene of a courtroom drama.
A peer-reviewed Stanford-led study tested seven widely used detectors on human-written essays. The tools frequently misclassified writing by non-native English speakers as AI-generated; the researchers reported that more than half of the TOEFL essays were falsely flagged, while essays from U.S.-born eighth graders were handled far more accurately. They also showed that prompting changes could help generated text evade detection.
That is not proof that every detector always fails. It is proof that the output depends on the tool, sample, language patterns, model family, editing, and decision threshold. The number is a model’s estimate under particular conditions. It is not a chain of custody.
The false-positive cost is not theoretical.
Writers who use formal prose, predictable sentence structures, editing tools, translation assistance, or English as an additional language may be disproportionately exposed to suspicion. Publicly naming a person based on a detector score can cause reputational harm that no later correction fully reverses.
If provenance matters, seek direct evidence: disclosed workflows, document history, credible reporting, contractual records, or platform-required declarations. Use detector output cautiously, privately, and as one imperfect signal. If your method begins and ends with pasting three paragraphs into a website, congratulations on completing an online quiz.