Turnitin AI Detector Accuracy in 2026: Scores and False Positives

What Turnitin's AI score actually measures, what its under-1% false positive claim covers, what independent tests found, and what to do at each score range.

By 13 min read
A laptop showing a document with some paragraphs highlighted in blue, beside a printed essay and a pen on a desk

Turnitin's AI detector is reasonably good at catching unedited chatbot output and is tuned to flag human writing rarely, but it is not accurate enough to prove anything on its own. Turnitin's own figures are a document-level false positive rate under 1% (for papers scored 20% or higher), a sentence-level false positive rate of about 4%, and a willingness to miss roughly 15% of AI text to keep false positives low. Independent peer-reviewed tests rank it at or near the top of the detectors they compared and still conclude that no detector is reliable enough to decide a misconduct case, especially once text has been edited. The percentage is a share of flagged prose, not a confidence level, and scores under 20% are hidden behind an asterisk because Turnitin found more false positives there.

What the Turnitin AI score actually measures

The number in the AI writing indicator is the percentage of qualifying text that the model thinks was likely AI-generated, or AI-generated and then run through a paraphraser or "bypasser" tool. Turnitin's AI writing detection FAQ defines qualifying text as prose sentences in long-form writing. Bullet points, short list items, tables, code, poetry, scripts, annotated bibliographies and the reference list are not counted.

So a 40% score on a 1,500-word paper does not mean "40% chance this used AI," and it does not mean 600 words were flagged. If 1,000 of those words were prose sentences and the model flagged about 400 of them, you get 40%. That is why Turnitin warns that the percentage "does not necessarily correlate to the amount of text" highlighted when a document mixes prose with lists or tables.

How the model produces the number

Per the FAQ, sentences are split into overlapping segments, a transformer-based classifier gives each segment a value between 0 and 1, each sentence inherits the scores of the segments it sits in, and the pooled sentence scores are aggregated into the document percentage. Turnitin says the model is not programmed to look at named features like perplexity or burstiness; it learns patterns from a corpus of academic writing and AI output, which is why individual predictions can't be explained sentence by sentence. (Our explainer on perplexity and burstiness covers the general mechanics.)

Two consequences follow. Short papers are close to all-or-nothing: with a few hundred words there is a single segment and no overlap, so a mixed text can be flagged as entirely AI. And transitions matter: Turnitin's chief product officer reported in a May 2023 update that 54% of false-positive sentences sat right next to genuine AI writing, which is why a human-written conclusion after an AI-written body often gets highlighted too.

What the indicator can show

IndicatorWhat Turnitin says it meansWhat to do
- - (gray, no number)Not processed: under 300 words of prose, over 30,000 words, unsupported language or file type, over 100 MB, or submitted before AI detection was enabledCheck the file requirements and resubmit; the writing was not judged
!Processing error on Turnitin's sideResubmit; contact support if it persists
0%No qualifying prose was flaggedNothing
*% (1 to 19%)Something was detected below the 20% threshold; Turnitin stopped showing these numbers in July 2024 because false positives are more common hereTurnitin calls this range less reliable; it is not an accusation
20% to 100%That share of qualifying prose was flagged as likely AI-generated or AI-paraphrased; Turnitin's under-1% document false positive claim applies hereInstructors: read the highlights, note where they start and stop, talk to the student. Students: gather drafts, version history and notes

The 300-word floor, the 30,000-word ceiling and the accepted file types (.docx, .pdf, .txt, .rtf) are listed in Turnitin's Using the AI Writing Report guide. The minimum was 150 words at launch; Turnitin raised it in May 2023 after finding more false positives on short submissions.

Why the 20% asterisk exists

The asterisk is the clearest admission in Turnitin's documentation that the tool is less trustworthy at the low end. Its release notes for July 16, 2024 say that "no score or highlights are attributed for AI detection scores in the 1% to 19% range" to avoid false positives, and the indicator shows *% instead. Before that, from May 2023, those scores were shown with an asterisk attached, after Turnitin's testing on 800,000 pre-ChatGPT papers found "a higher incidence of false positives" below 20%. It has never published a separate false positive rate for that range.

So if an instructor says your paper came back with an asterisk, the detector itself is saying the signal is too weak to report. Reports generated before July 8, 2024 may still show a number under 20%, with the same caveat.

What Turnitin's "less than 1% false positives" claim covers

Turnitin has repeated the under-1% figure since its March 2023 false positives post. It is narrower than it sounds.

  • It is a document-level rate for documents scored 20% or more. The FAQ states the goal is to keep false positives "under 1% for documents with over 20% of AI writing," validated before each model release on more than 700,000 pre-ChatGPT papers.
  • The sentence-level rate is about 4%. In a 2023 post, Turnitin said a highlighted sentence has roughly a 4% chance of being human-written, most often at the boundary between human and AI text.
  • It is bought with false negatives. The FAQ's example is that a paper scored at 50% might contain as much as 65% AI writing, and Purdue's instructor guidance relays Turnitin's estimate that it misses around 15% of AI-generated content.
  • It was measured by Turnitin, on pre-2022 papers. A sensible test for false positives, but it says nothing about writing drafted by a person and polished by a grammar tool, or produced by a multilingual student in 2026.

At scale, even 1% is a lot of people. Vanderbilt's August 2023 decision to turn the feature off did the arithmetic: at Turnitin's stated rate, around 750 of the 75,000 papers Vanderbilt submitted in 2022 could have been wrongly labeled. For volume, Turnitin's March 2024 announcement put it at over 200 million papers processed, about 11% scored at 20% or more and about 3% at 80% or more.

What independent studies found

Three peer-reviewed evaluations are the ones universities cite when deciding whether to trust the score.

Weber-Wulff et al. (2023), in the International Journal for Educational Integrity (also a preprint), tested 14 detectors including Turnitin on 54 documents: human-written, machine-translated, ChatGPT-generated, manually edited ChatGPT output, and ChatGPT output paraphrased with Quillbot. Turnitin ranked first overall and produced no false accusations on the human-written set, but the headline was that every tool "scored below 80% of accuracy and only 5 over 70%." On machine-paraphrased AI text, accuracy across tools fell to 26%, and Turnitin missed most of the Quillbot samples. That is the weakness Turnitin has worked hardest on since; see whether Turnitin catches Quillbot paraphrasing.

Liang et al. (2023), published in Patterns and on arXiv, is the study behind the "biased against non-native speakers" headlines. Seven detectors flagged an average of 61% of 91 TOEFL essays by non-native writers as AI-generated, 97.8% of the essays were flagged by at least one tool, and the same tools were near-perfect on essays by US eighth graders. Turnitin was not one of the seven, and its principal machine learning scientist responded that all 91 essays were under 150 words, below its minimum, and that internal testing showed no statistically significant bias against English language learners. That claim is unaudited, and Turnitin's own FAQ lists "content without a lot of structural variation" and repetitive text as common false-positive patterns, which describes plenty of honest second-language writing.

Perkins et al. (2024), in the International Journal of Educational Technology in Higher Education and on arXiv, tested detectors including Turnitin, Copyleaks, GPTZero and ZeroGPT on AI text, versions altered with evasion techniques (misspellings, added burstiness, Quillbot, prompting for a non-native style) and human controls. Average accuracy on unmodified AI text was 39.5%; on manipulated text it fell to 22.2%. Turnitin was second-best on clean AI text but, under the study's scoring, showed the largest drop of any tool on manipulated text, from 50% to 7.9%. The authors concluded detectors "cannot currently be recommended" for determining integrity violations.

The pattern: Turnitin is among the better detectors, it rarely flags clean human prose outright, and it can be evaded by editing, which is why it can't carry a case by itself. All three studies predate Turnitin's 2025 and 2026 model updates, so the numbers are a snapshot, not a verdict on today's model.

What universities say an acceptable score is

Nothing, and Turnitin agrees: its sentence-level post says there is "no 'right' or 'target' score" with the AI indicator. What counts as acceptable is whatever your course's AI policy allows, which might be "none," "for brainstorming only," or "anything, with disclosure." Institutional guidance is consistent on the point that matters:

  • Purdue tells instructors the percentage "should not be used as the sole basis for action or a definitive grading measure" (Purdue Online); Illinois Tech repeats the same line.
  • University of Kentucky kept the feature on but told faculty to handle scores "with healthy skepticism," since unlike a similarity match there is no source document to corroborate an AI flag (UK Provost).
  • University of Michigan does not recommend AI detection at all "given their high error rate" (U-M ITS). Vanderbilt disabled it in August 2023, citing the false positive math, the non-native speaker research and the lack of detail about how the model works; ANU disabled it on 1 January 2024.

Instructors: put Turnitin's own sentence in your syllabus, the AI score is a reason to look closer, not a finding. Students: read your course's AI policy now, because it, not the detector, defines what is allowed. Our guide on what to do when an essay is flagged as AI but you wrote it covers the conversation itself.

Does Turnitin detect ChatGPT, Claude and Gemini, or just plagiarism?

Both, separately. The similarity score matches your text against Turnitin's database of web pages, publications and past student papers; the AI score is a classifier's prediction with no source to match. Turnitin's FAQ says the two are "completely independent," so a paper can be 2% similar and 90% AI, or the reverse.

Turnitin's current English model lists GPT-4o through GPT-5.4-pro, o1-mini, Gemini 1.0 through the Gemini 3.1 Pro preview, Claude 3 Haiku through Claude Sonnet 4.6 and Opus 4.5, Llama 3.3 and 4, Mistral Large 3, DeepSeek v3.2, Grok 4.1 and Nova 2 Lite, "and tools based on these LLMs as well." So it was trained on Claude output, rather than on ChatGPT alone. The release notes record recall improvements in October 2025 and February 2026, and in July 2026 a consolidation of the previous multi-model ensemble into a single model. None of these updates re-score old submissions; a paper has to be resubmitted to get a new number.

Paraphrasers, humanizers and the "AI bypasser" indicator

Turnitin added detection of AI text spun through a paraphraser in December 2023, gave it a separate report category (purple highlights) in July 2024, and added detection of "AI bypasser" tools, meaning humanizers, in August 2025. On 4 August 2026 it merged the categories back into a single blue "AI-generated" highlight while keeping the underlying detection. Two limits: paraphrase and bypasser detection exist only for English, and Turnitin won't name the tools it trained against. The honest reading is that a humanizer pass is no longer a reliable way to hide AI drafting.

Which languages are supported

English at launch (April 2023), Spanish since September 2024, Japanese since April 2025, and Modern Standard Arabic since 18 August 2026, each a separate model trained on fewer LLMs than the English one. Any other language returns the empty - - state.

Can students check their own Turnitin AI score?

No. Turnitin's FAQ says the indicator and report "are not visible to students," though an instructor can download the report as a PDF and share it. Draft Coach, the student add-on for Google Docs and Word Online, runs similarity, citation and grammar checks only; its FAQ mentions nothing about AI detection. There is no public Turnitin AI checker, and any website promising "your Turnitin AI score" is not running Turnitin's model.

What you can get is a proxy. Running a draft through Rewritica's free AI checker or GPTZero shows you what a detector sees: the passages it finds most AI-like, usually generic transitions, repeated sentence shapes and summary paragraphs. It won't produce Turnitin's number, because it is a different classifier with a different threshold. Use it to find the flat stretches and rewrite them with specifics, not to chase a score.

A worked example

Say a 1,400-word essay comes back at 38%. Roughly 1,100 words are prose sentences (the heading, a bulleted list and the reference list don't count), so the model flagged about 420 words' worth. In the report, the highlights cover the introduction and most of the second section, then stop cleanly at a paragraph that cites a class discussion by date. That shape is more informative than the number: a block of flagged text with a hard edge looks more like pasted output than a sprinkling of flagged sentences across the paper, which is more typical of a false positive on formulaic writing. Either way, the next step is the one Turnitin recommends: ask the student how the paper was written.

If you are the student in that meeting and you wrote it, bring the evidence before you're asked:

Hi Professor Adeyemi, thanks for letting me know about the AI report on my essay. I wrote it myself and I'd like to show you how. I've attached my outline from 3 March, the Google Docs version history (about 40 revisions over nine days), and my notes on the two readings. I'm happy to talk through any paragraph or write a short sample in your office. Could we meet this week?

Version history, earlier drafts, research notes and anything with timestamps carry the most weight. Our guide on how to prove you didn't use AI covers what to collect and how to present it.

Frequently asked questions

What does a Turnitin AI score of 40% mean?

It means the model thinks roughly 40% of the qualifying prose sentences in the file were likely AI-generated (or AI-generated and then paraphrased). It is a share of text, not a 40% probability that the paper used AI, and Turnitin says it should not be the sole basis for any action.

What is Turnitin's false positive rate?

Turnitin says under 1% of fully human-written documents get flagged, but that figure applies to documents scored at 20% or above, and it has separately reported a sentence-level false positive rate of around 4%. Independent academic tests have found higher error rates in specific conditions.

Why does my Turnitin AI score show an asterisk or two dashes?

An asterisk (*%) means the model found something between 1% and 19%, a range Turnitin no longer reports because false positives are more common there. Two dashes (- -) mean the file was not processed, usually because it has fewer than 300 words of prose, is over 30,000 words, is in an unsupported language or file type, or was submitted before AI detection was enabled.

Can students check their own Turnitin AI score before submitting?

No. The AI indicator is visible only to instructors and administrators, and Draft Coach (the student add-on for Google Docs and Word Online) checks similarity, citations and grammar, not AI. A free checker such as Rewritica's or GPTZero gives you a rough proxy, but it is a different model and will not match Turnitin's number.

Does Turnitin detect ChatGPT, Claude and Gemini?

Yes. Turnitin's current English model lists GPT-4o through GPT-5.4, Gemini 1.0 through 3.1, Claude 3 Haiku through Sonnet 4.6 and Opus 4.5, Llama 3.3 and 4, Mistral, DeepSeek and Grok among the models it was trained to detect, plus tools built on them.

What Turnitin AI score is acceptable?

There is no universal number. Turnitin says there is no right or target score, and every university we checked tells instructors not to act on the percentage alone. What counts as acceptable is set by your course's AI policy, not by the detector.

Bottom line

Turnitin's AI detector is a decent signal and a poor judge. Its under-1% false positive claim is real but narrow (documents at 20% or above, measured on pre-ChatGPT papers), its sentence-level errors run around 4%, and independent studies show it can be both fooled by editing and tripped by flat human prose. Instructors should treat the score as a prompt to read the highlights and talk to the student, which is what Turnitin itself asks. Students should read the course AI policy, keep version history on everything, and use a free checker to find and fix generic passages rather than to chase a number.

Written by the Rewritica Editorial Team. We research every guide against primary sources (vendor documentation, university policies, and peer-reviewed studies) and update it when the facts change. Spotted an error? Tell us.

  • turnitin
  • ai detection
  • false positives
  • academic integrity
  • ai score

Related articles