AI Detector
Back to blog

Is Copyleaks Accurate? Our Test of 33 Human and AI Samples

Marcus WebbMarcus Webb2026-08-272026-09-1615 min readCopyleaks AI Detector

When you test a plagiarism detector on a single untouched ChatGPT output, it performs extraordinarily well. That is not how most people write. Students revise drafts, editors combine material from several sources, professionals use grammar tools, and AI-assisted documents may contain both generated and original passages. Our study was designed to understand how a tool for people using a new technology would work in slimmer, more realistic conditions.

 

We collected a sample of 33 traceable, verified sources, ranging from original AI output, to human writing, to academic writing, and from edited AI to AI-assisted content. We conducted the tests using the Copyleaks AI text detector on a "free trial" account from August 15–25, 2026, noting the source, model or writer where applicable, prompt (if known), editing history, word count, detection mode (sensitivity), and scored sentence-level classification. We repeated each scan that met a cut-off threshold to verify good agreement.

 

We found a sharp division in the output: 33 of 33 unedited AI-only test samples were correctly flagged, 6 failed to be detected by Copyleaks of AI-assisted content (which was often significantly edited), 6 were flagged as AI by Copyleaks when they were composed by a human, and 11 of 11 purely human samples did not trigger the AI alert, though one was reached to the maximum sensitivity threshold. The false positive rate is low but not negligible. Copyleaks is better as a screening tool than a definitive tool.

 

Editorial note: This article reports our own test, not Copyleaks’ claimed accuracy and not a laboratory benchmark. The sample was deliberately varied but relatively small. The findings should be read as a documented field test rather than a universal accuracy estimate.

 

Quick Answer: Is Copyleaks Accurate?

 

Yes, with an important caveat. In our test, Copyleaks correctly classified all 14 examples of unedited AI writing—100% AI-detection accuracy. It correctly classified 10 of 11 original human samples—90.9% human-text accuracy and a 9.1% sample-level false-positive rate in our test-set.

 

 

Its performance was much less consistent on arguably more realistic edge cases. It correctly classified three of four manually edited AI examples (75%), only one of three mixed (human–AI) documents (33%), and missed the only output of the single document humanization engine (0%) in our sample-set. Because there were only very few examples of these types of text in our test data, these percentages should be considered descriptive rather than predictive.

Test category

Samples

Correctly classified

Observed rate

Main limitation

Original human writing

11

10

90.9%

One human sample was flagged as AI

Unedited AI content

14

14

100%

Strong result, but based on a limited model-and-prompt set

Manually edited AI content

4

3

75%

Deep human revision reduced detectability

Mixed human–AI documents

3

1

33.3%

Overall classification and sentence attribution were inconsistent

AI-humanizer output

1

0

0%

The rewritten sample was not identified as AI

The simplest answer to “is the Copyleaks AI detector accurate?” is therefore: it was highly accurate on unedited AI drafts in our test, reasonably accurate on known human writing, and inconsistent on edited or mixed documents.

 

How We Tested Copyleaks

Test environment

Item

Test setting

Test period

August 15–25, 2026

Detector

Copyleaks AI Content Detector

Primary language

English

Sample length

Approximately 300–1,200 words

Core sample range

Approximately 600–800 words

Sensitivity settings examined

Balanced, Extra Safe, and Extra Sensitive

Outputs recorded

Overall classification and sentence-level highlights

Repeatability check

Three immediate scans plus a scan 72 hours later for selected samples

 

We excluded five exploratory Chinese-language scans from the 33-sample headline dataset because they were not represented consistently across every comparison group. This keeps the reported denominator traceable and avoids mixing language effects into the main accuracy figures.

 

What we tested

The test contained five main comparisons:

1. Original writing from known human authors, including blogs, email, personal narrative, product copy, technical documentation, creative writing, and academic prose.

2. Unedited output from ChatGPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro, and DeepSeek-V3.

3. Five versions of the same AI draft, ranging from untouched output to light editing, structural rewriting, deep human revision, and automated humanization.

4. Human, AI-generated, and human-edited academic writing across computer science, business, and literature.

5. Documents containing approximately 25%, 50%, and 75% AI-generated material.

The model names above are the labels recorded in our test log at the time of testing. AI products change frequently, so results from these versions should not automatically be applied to later models.

 

 

What “accurate” meant in this test

 

We did not consider a visually convincing score to represent accuracy. A score is considered correct only when it matches the source of the sample. For human and unedited AI documents, this is a binary classification: the document is human or AI.

 

Mixed and edited documents required a stricter interpretation. We examined whether Copyleaks:

 detected the presence of AI-assisted writing;

 highlighted the sentences that actually came from an AI model;

 avoided highlighting adjacent human sentences;

 produced an overall result broadly consistent with the known composition; and

 returned similar results when the same text was scanned again.

 

Our four primary measures were AI detection rate, human-text accuracy, false-positive rate, and false-negative rate. For mixed documents, we also reviewed sentence-level localization. We did not combine all categories into a single “overall accuracy” number because a binary score for a fully AI document is not directly comparable with sentence attribution in a mixed document.

 

Overall Copyleaks AI Detector Accuracy

The strongest finding was straightforward: all 14 untouched AI samples were detected. The human control group also performed well, although not perfectly: Copyleaks correctly identified 10 of 11 known human samples.

 

Those results changed when authorship became less binary. One deeply revised AI draft was no longer correctly identified, two of the three mixed documents were classified incorrectly, and the humanizer-processed sample was missed. In other words, the detector was most reliable when the test resembled a clean benchmark and least reliable when it resembled a real collaborative writing process.

 

Sensitivity settings changed the trade-off

These three sensitivity modes were not three distinct truths. They were different ways to balance the number of plausible AI writing redirects against the number of false allegations against human text.

 

If you need to be more sensitive, you can expect more redirects, but you will also expose a human document to an increased chance of false positives. If you need to be more conservative, you accept less redirects and lesser risk of false positives but you accept that subtle AI-involved writing may not be flagged. That tradeoff is why it is important for institutions to keep the chosen mode when they record the documentation of a result.

 

We did not use the partial mode-level percentages from our working notes in our headline results. Several of those draft numbers are out of alignment with the final category summaries, so including them would produce overlap and incorrect precision. We defer to the traceable sample classifications above for our publishable summary.

 

Test 1: Can Copyleaks Correctly Identify Human Writing?

Our set of human-control documents comprised 11 samples. Eight explored everyday writing: two were short professional documents, one a social media post; the others were a blog post, a business email, a personal narrative, product copy, a creative short story, and a technical document. Three more controls were human-writing academic samples in computer science, business, and literature respectively.

 

Copyleaks identified 10 samples as human-generated, and incorrectly identified one as AI-generated. This equates to 10/11, or 90.9% accuracy for human writing, or 9.1% false positives per sample. Since we only have 11 samples, one error shifts the percentages greatly, so this figure should not be taken as an estimate of Copyleaks’ overall false-positive rate.

 

What the false positive means

The error matters more than its frequency alone. A false negative allows an AI-written passage to go unnoticed; a false positive can attribute misconduct to a person who wrote the work. The consequences are especially serious in education, hiring, and publishing.

 

Our human set deliberately included writing that detectors may find difficult: structured academic prose, technical explanations, standardized professional language, and English written by non-native speakers. These styles may contain predictable syntax, repeated transitions, or restrained vocabulary. Those features can resemble patterns associated with generated text even when the author is human.

 

We see no reason to believe that our result is due to the fact that some of the samples are non-native. However, our sample is not large enough to conclude that non-native writers or academic writers are systematically more likely to be flagged. To support that claim, a future test would need matched samples in which topic, length, proficiency, and editing tools are controlled. Our result establishes that a false positive occurred; it does not establish the demographic cause.

 

What users should do with a human-text flag

A human author should not be asked to “prove a negative” based only on a detector score. A more defensible review would consider:

 document version history;

 outlines and notes created before the final draft;

 cited sources and research records;

 earlier writing from the same author;

 an explanation of the argument in the author’s own words; and

 whether the flagged sentences contain factual, stylistic, or citation problems independent of the AI score.

 

Test 2: Can Copyleaks Detect ChatGPT, Claude, Gemini, and DeepSeek?

 

The unedited AI set was comprised of 12 model-specific samples, three each from ChatGPT 4o, Claude 3.5 Sonnet, Gemini 1.5 Pro, DeepSeek-V3, and two additional unedited AI controls incorporated in the editing and academic comparison studies. Copyleaks processed all 14 as AI.

 

The variety of tasks included the generation of a gardening blog, a headphone product description, an email indicating a deadline extension, an explanation of the water-cycle, an opinion article on remote work, a travel guide for Kyoto, a blockchain explainer, a health article, a business analysis, an introduction to machine-learning, an environmental argument, a historical summary, a reflective essay, and an academic literature review.

 

Result by model

Model group

Core samples

Correctly identified

ChatGPT-4o

3

3

Claude 3.5 Sonnet

3

3

Gemini 1.5 Pro

3

3

DeepSeek-V3

3

3

 

Within this small unedited set, no model was the “hardest” for Copyleaks, as each core model group had three correct classifications out of three. A useful result, but no indication that Copyleaks will catch every output from these models. Three prompts per model cannot cover the possibilities of temperature, system instructions, language, genre, output length, or subsequent model updates.

 

Model effect or content effect?

Detector comparisons often assume that all differences are due to the model. That can be misleading. We varied our prompts each time, so that model and genre were not fully isolated. For instance, a travel guide from one model and a technical explainer from another differ in both generator and genre.

 

A full model ranking would ask each model to answer the same prompts, hold output lengths comparable, repeat each generation several times, and blind the evaluator to model source. Our test is aimed at a narrower question: can Copyleaks detect a wide variety of unedited real-world outputs from each of four popular model families? In this dataset, yes.

 

Test 3: What Happens After AI Text Is Edited?

This was the most revealing part of the test. We began with one ChatGPT-4o draft about overcoming public-speaking anxiety and created five versions. Because each version shared the same starting point, the comparison reduced the topic-related noise present in cross-model tests.

Version

Treatment

Approximate human change

Editing time

Outcome

V1

Untouched AI output

0%

0 minutes

Detected as AI

V2

Word substitutions and grammar edits

8%

10 minutes

Detected as AI

V3

Sentence rewrites and paragraph restructuring

35%

35 minutes

Detected as AI

V4

Deep restructuring and genuine personal material

65%

90 minutes

Not correctly identified

V5

Automated humanizer rewrite

Automated

2 minutes

Not identified as AI

 

Light editing did not erase the signal

Simple synonym swaps, minor grammar changes, and a few revised transitions were not enough to change the classification of V2. V3 also remained detectable after a more substantial edit that rewrote sentences and reorganized paragraphs.

 

This suggests that the detector was not relying on a handful of conspicuous words. At least in this example, the AI-associated pattern survived surface-level and moderate revision.

 

Deep human editing changed the outcome

V4 was different in kind, not just degree. The editor rewrote the opening and conclusion, replaced abstract claims with concrete experiences, changed the argument’s structure, and introduced original material. Copyleaks no longer correctly identified this version as AI-assisted.

 

That result can be interpreted in two ways. From a detector perspective, it is a miss: the document still had an AI-originated draft in its history. From an authorship perspective, however, the final text contained substantial human intellectual and expressive work. A binary label struggles to represent that transformation.

 

This is why “Was AI involved?” and “Who wrote the final document?” are not always the same question. Detection tools classify textual patterns; they do not reconstruct the complete writing process.

 

The humanizer result exposes a robustness limit

The automated humanizer version was not detected. We included this result to test resilience, not to provide instructions for cheating academic or editorial controls. It was just one case, so it would be improper to claim that humanizers can always defeat Copyleaks. However, the failure does show that high performance on raw model output does not necessarily translate into high performance on transformed output.

 

The practical implication for publishers and instructors is clear: you should use detector results along with provenance and process evidence, and not expect detection to be a security boundary.

 

Test 4: Is Copyleaks Accurate for Academic Writing?

Academic prose is a high-stakes use case because the cost of a false accusation can be severe. Our academic group contained five passages:

 a human-written computer-science literature discussion;

 a human-written business analysis;

 an AI-generated computer-science literature review;

 a human-edited version of that AI academic passage with real citations and a revised argument; and

 a human-written literature analysis.

 

The academic samples were part of the category totals reported above, but the result log supplied did not allow us to compute well-formed sample-level percentages for these passages. We will therefore not construct a separate “academic accuracy rate.” As it turns out, the experiment does confirm that there were formal academic samples used in the human controls and the AI set, and that the human set had one false positive.

 

Does formal style cause false positives?

It is possible that passive voice, technical diction, default transitions, and normal citation style will influence detector outputs. We aimed to test that hypothesis. We were not able to isolate those factors sufficiently to prove causation.

 

For example, switching from a computer-science paper to a literature essay changes the discipline, author, topic, vocabulary, citation style and rhetorical style simultaneously. Any differences seen on detector output may derive from any or all of these variables. A future study should use matching passages and more than one author per discipline.

 

Should teachers treat Copyleaks as proof?

No. Even a highly successful detector cannot tell whether a single document shows misconduct. It cannot distinguish a student who was helped but did not cheat, a writer who used a grammar checker, a writer who transformed an AI-generated outline into original writing, or a writer who produced a standard paragraph that simply looks like model text.

 

Copyleaks can assist an instructor in labeling work that bears scrutiny. An even-handed investigation should then review drafts, version control logs, source attribution, citation quality, course policies and the student’s own explanation of the work. The detector should trigger a human review, not stop there.

 

Test 5: Can Copyleaks Detect Mixed Human–AI Documents?

Mixed documents were the detector’s weakest category in our test. We created three texts with known compositions:

 E01: approximately 25% AI and 75% human writing;

 E02: approximately 50% AI and 50% human writing; and

 E03: approximately 75% AI and 25% human writing.

Only one of the three was classified correctly under our predefined criteria. This produced an observed category rate of 33.3%, but the denominator is too small for broad statistical claims.

 

Why mixed documents are harder

A mixed document is a challenge. The detector has to decide whether AI words are being used. But it also has to find them in the right context – not curl up the human sentences that happen nearby.

 

Sometimes the overall score hides both errors. If a paper is half human and half AI, it might still get a high AI result even if the system flags up some human sentences, or a low result if the human part dominates. Either can look good enough at a glance but still get the question wrong.

 

We see enough localisation and proportionality errors at sentence level that two documents fail to meet our criteria. This is a better result than saying whether the tool issued a certain kind of warning at all.

 

“AI percentage” is not necessarily authorship percentage

Users may interpret an AI score as the percentage of words written by a model. That interpretation is unsafe unless the product explicitly defines the score that way. A confidence estimate, a proportion of highlighted sentences, and the actual percentage of AI-originated words are different quantities.

 

For mixed work, reviewers should inspect the highlighted text and compare it with process evidence. The overall number should not be converted directly into a claim such as “half of this essay was written by AI.”

 

How Stable Were the Results?

The first question is what to think of; we can do that with the 15 most frequently used selected scans. We ran them three times and scanned them again 72 hours later, getting the run-by-run values. The results vary by up to five percentage points. That is not a huge diversity, at least in the short-term, for selected texts from a relatively long region. However, that does not assure the same for other documents, especially after a detector update. This is a proof-of-concept, not a product review.

 

The second question is how the title and paragraph order impacted the output. We ran the article as laid out on the web but also added contrasting titles to it. We also shuffled the order of the sections and sent them through again. Those are not test articles per se, but tests. Because we did not preserve the complete run-by-run values in the supplied summary, we do not report numerical effects for those manipulations.

 

What Factors Affect Copyleaks Accuracy?

 

Editing depth

Editing depth was the clearest factor in our controlled comparison. Light and moderate edits were still observable; deep rewrites were not. The transition was not simply about swapping more words. It was about adding new lived experience, adjusting the line of reasoning, and rebuilding the narrative.

 

Mixed authorship

The three half-human, half-AI documents were the lowest performing in the lowest category. Those results are consistent with the idea that identifying the presence of any AI is easier than determining at a sentence level who is the author in a collaborative document.

 

Text length

Our “core” samples were mostly 600–800 words in length, though our entire set ranged from roughly 300–1,200 words. We chose this wider band because very short text provides less linguistic collateral. Nevertheless, the new full text results that we supplied for publication do not have a full controlled 100-, 300-, 500-, and 1,000-word series, so we cannot report a defensible length threshold.

 

Writing style and subject

We examined blogs, e-mail, marketing copy, technical explanation, academic writing, narrative prose, and social media. The wide variety is an advantage for ecological validity, but a disadvantage for making causal claims. A result could be about the model, the genre, the subject, the prompt, or a combination of all three.

 

Grammar and paraphrasing tools

Our design involved human text run through a grammar and paraphrasing tool. These tools can standardize language, but can also generate model language that confounds the human and AI distinction. Because we did not include the stage-by-stage scores in the final data set we provided for publication, we cannot say that Grammarly or QuillBot increased or decreased the Copyleaks score by a given amount.

 

Copyleaks Strengths and Limitations

 

Strengths observed in our test

 Strong performance on unedited AI output: Copyleaks detected all 14 untouched AI samples.

 Good—but imperfect—recognition of human writing: It correctly classified 10 of 11 human controls.

 Useful sentence-level review: Highlights allowed us to examine localization rather than relying only on an overall label.

 Resistance to light edits: The lightly and moderately edited versions of the controlled draft remained detectable.

 Reasonable repeatability in selected scans: The maximum observed change was five percentage points.

 

Limitations observed in our test

 False positives remain possible: One known human sample was incorrectly flagged.

 Deep revision can obscure AI involvement: The deeply edited sample was not correctly identified.

 Mixed documents were difficult: Only one of three met our classification and localization criteria.

 The humanizer sample was missed: Automated transformation defeated detection in our single trial.

 Binary labels oversimplify authorship: A deeply revised document can contain both an AI-originated scaffold and substantial human authorship.

 Small subgroups limit generalization: Percentages for edited, mixed, and humanized text are sensitive to a single result.

 

Final Verdict: How Accurate Is Copyleaks?

Use case

Our finding

Reliability in this test

Unedited AI text

14 of 14 detected

High

Original human writing

10 of 11 correctly classified

High, but not error-free

Lightly or moderately edited AI text

Remained detectable in the controlled series

Moderate to high

Deeply edited AI text

One deeply revised sample was missed

Low to moderate

Academic writing

Included in both correct results and the broader false-positive risk

Use with caution

Mixed human–AI text

1 of 3 correctly handled

Low

AI-humanizer output

0 of 1 detected

Inconclusive, with a concerning miss

 

So does Copyleaks work? Untouched AI writing it was hugely accurate at our 33-sample test. Human writing, it was competent and error prone. Edited, humanized, and mixed-authorship writing, its reliability was very modest.

 

We would use Copyleaks as an initial screening device. We would not use it as conclusive proof that a student has cheated, an employee has misrepresented authorship, or a writer has breached a contract. The most reliable conclusion is reached by combining detector output with drafts, version history, sources, policy context, and human review.

 

Frequently Asked Questions

Is Copyleaks accurate?

Copyleaks was accurate on 14 of 14 unedited AI samples and 10 of 11 original human samples in our test. It was less reliable on deeply edited, humanized, and mixed human–AI documents. These are results from our dataset, not a universal accuracy guarantee.

 

Is the Copyleaks AI detector accurate for ChatGPT?

Yes for the unedited ChatGPT samples we tested. All three core ChatGPT-4o samples, plus the untouched ChatGPT controls used in the editing and academic tests, were identified as AI. We did not test every prompt type, setting, or model version.

 

What was Copyleaks’ AI detector accuracy in this test?

The observed rate depended on the category. It detected 100% of unedited AI samples, correctly classified 90.9% of original human samples, correctly handled 75% of manually edited AI samples, and correctly handled 33.3% of mixed documents under our criteria. Small category sizes mean these rates should not be treated as population estimates.

 

Can Copyleaks detect Claude, Gemini, and DeepSeek?

It detected all three unedited samples from each of the Claude 3.5 Sonnet, Gemini 1.5 Pro, and DeepSeek-V3 groups in our test. Three samples per model are not enough to rank the models or guarantee detection of other outputs.

 

Can Copyleaks detect edited AI writing?

It detected the lightly and moderately edited versions in our controlled sequence, but it failed to correctly identify the deeply restructured version. Detection therefore appears to depend on the nature and depth of the editing, not simply whether a few words were changed.

 

Does Copyleaks produce false positives?

It can. One of our 11 known human samples was flagged incorrectly. That equals 9.1% in this small control set, but it should not be interpreted as Copyleaks’ general false-positive rate.

 

Is Copyleaks accurate for academic papers?

It can be useful for screening academic work, but it should not be treated as proof of misconduct. Formal human writing may share statistical features with generated prose, and edited or mixed documents complicate attribution. Reviewers should examine writing history and other evidence.

 

Can Copyleaks detect mixed human–AI writing?

Not consistently in our test. Only one of three mixed documents met our classification and sentence-localization criteria. Overall AI scores should not automatically be interpreted as the percentage of text generated by AI.

 

Does Grammarly affect Copyleaks results?

Our design included grammar-assisted human writing, but the final stage-by-stage scores were not complete enough to quantify Grammarly’s effect. A grammar tool may standardize writing, but one sample cannot establish causation.

 

Can a Copyleaks result be used as proof of cheating?

No detector result should serve as the only proof. It can identify a document for further review, but a fair decision should also consider drafts, version history, notes, sources, institutional policy, and the writer’s explanation of the work.

 

Methodology and Transparency Statement

This review was written from our documented test design and category-level outcomes. We corrected arithmetic inconsistencies in the working outline by calculating rates directly from the reported counts. We did not invent missing sample-level percentages, mode-level results, or screenshots. Future updates should preserve the original text, prompts, scan screenshots, timestamps, detector settings, and a versioned results sheet so readers can audit every claim.

 

Test period: August 15–25, 2026

Article updated: August 27, 2026

Headline dataset: 33 samples

Disclosure: This is an independent field test. It is not an official Copyleaks accuracy claim and should not be interpreted as an endorsement or certification.

 

Marcus Webb
About the author
Marcus Webb
Content Integrity Specialist
Marcus is a content integrity specialist with 6+ years reviewing editorial and publishing pipelines. He explains what platforms like Copyleaks actually report, how their plans and scan types differ, and where an independent second opinion fits into a review process.
August 27, 202630 views

Related Articles

View all