Turnitin's AI detector draws false-positive accusations against students
Turnitin had claimed a false-positive rate under 1%, but its own later disclosure put sentence-level false positives near 4%; Vanderbilt disabled the tool that August.
- Culture & impact
- Notable
Turnitin’s AI-writing detector, deployed automatically across the plagiarism-checking software used by thousands of universities since April 2023, drew a growing run of reports through mid-2023 of students falsely accused of using AI tools they had not used. Because the feature was switched on for existing Turnitin customers with under 24 hours’ notice and no opt-out, instructors and students encountered it before institutions had settled on how to treat its results.
Turnitin had marketed the tool with a claimed false-positive rate under 1%. Its chief product officer subsequently acknowledged that real-world performance fell short of that laboratory figure and disclosed a sentence-level false-positive rate of roughly 4%, with the tool’s most unreliable outputs concentrated in low-confidence results and in text mixing human and AI-generated sentences — precisely the kind of writing, produced by a student who used an AI tool to revise rather than generate a whole essay, that the detector was ostensibly built to catch.
The gap between the marketed accuracy and the disclosed rate had institutional consequences. Vanderbilt University disabled Turnitin’s AI detector that August, citing the disclosed false-positive rate against its roughly 75,000 annual student submissions — a scale at which even a low single-digit error rate implied hundreds of students potentially mislabelled — alongside concerns that Turnitin provided no detailed account of how the detector worked and that AI detectors more broadly were shown to misclassify text by non-native English speakers at elevated rates. Other institutions followed with their own decisions to disable or restrict the feature, and the episode became an early, frequently cited case of an AI-detection product deployed faster than its accuracy could be verified.