Home / Reviews / AI Detector / GPTZero AI
AI Detector Content Verification

GPTZero Review

gptzero.me · AI text detector, now part of Superhuman

GPTZero is the AI detector that shows its work. Its Advanced Scan ranks a document line by line and gives a short reason for each flag instead of handing back a bare percentage. In our session it caught unedited ChatGPT output at 100% and held that score on shorter extracts, yet it also marked 266 words of plain, hand-written English as 100% AI. It suits educators and editors who want an explainable first read on authenticity, so long as they treat the result as a starting point for a conversation rather than proof.

Reviewer Score
6.8
/10
Mixed Results
Sharp on obvious AI
Shaky on real writing
AI Detection8.5
Result Explanations8.5
False Positive Control4.5
Free Tier8.0
Support and Billing5.0
Platform
Web + Chrome extension
Free Plan
10,000 words / month
Models Detected
ChatGPT, Claude, Gemini + more
Full-Support Languages
5 incl. EN, DE, ES
Max Batch
250 files, Professional
Premium
$9.09 /mo billed annually

What GPTZero Does

01 . Overview

GPTZero reads a block of text and estimates how much of it was written by a large language model. Edward Tian built the first version over a Princeton winter break in early 2023, and the product has since grown into a detection suite used across education, publishing, hiring and legal review. The company says it has served more than 17 million users and now sits inside Superhuman.

The pitch rests on two claims. First, accuracy: GPTZero markets a 99% detection rate and says it is the only detector de-biased for English-as-a-second-language writers. Second, interpretability, which is where it tries to separate itself from the pack. Rather than returning a single number, the Advanced Scan highlights individual sentences and explains why each was flagged. Access starts with a permanent free plan capped at 10,000 words a month, then climbs through paid Essential, Premium and Professional tiers, with Premium shown at $9.09 a month on annual billing at the time of writing.

What Happened When We Tested It

02 . 6 tests spanning raw detection, false positives and the accuracy gap

We worked from a free account plus the paid dashboard, feeding the detector a mix of machine text, historical human text and writing produced by hand. The goal was not to confirm the marketing figure but to find the edges where a detector like this helps a teacher and where it could get a student wrongly accused. The session moved from the easy cases to the ones that expose the model's weak spots.

Unedited ChatGPT Output Was Caught at Once

Test 01 . Raw AI
★ Prompt Used

"Write a 490-word essay on the causes of the First World War."

GPTZero result screen showing an unedited ChatGPT essay scored 100 percent AI with zero percent human and mixed
Unedited ChatGPT output returned a 100% AI verdict on the basic scan.
What we observed: The essay came back at 100% AI, with 0% assigned to both the human and mixed buckets. Re-scanning shorter extracts of the same text held the score each time, so the instability that critics often pin on short samples did not show up here.

The Sentence-Level View Carried the Session

Test 02 . Advanced Scan
GPTZero Advanced Scan tab ranking individual sentences by how strongly each pushed the verdict toward AI
The Advanced Scan ranks each sentence and offers a reason for every flag.
The signal here: Expanding a flagged line reveals a short explanation of why it scored the way it did. For a teacher looking at a flagged assignment, seeing which sentences drove the number beats being handed a percentage with no reasoning attached. Most rival detectors stop at the headline figure.

Eighteenth-Century Human Text Passed Cleanly

Test 03 . Historical prose
GPTZero scoring 248 words of the Bill of Rights as 93 percent human with a highly confident human verdict
The Bill of Rights was read as human at 93%.
Worth noting: We pasted 248 words of the Bill of Rights from the National Archives transcription. That text predates language models by more than two centuries, so any AI score would be a false positive by definition. GPTZero cleared it at 93% human. Formal, archaic legal prose is exactly the writing critics expect a detector to misread, and it did not.

The Hallucination Detector Waved Through Fake Sources

Test 04 . Citation check
GPTZero Hallucination tab reporting zero unsupported claims after scanning a passage built on three invented citations
Three fabricated citations produced one soft warning and no unsupported-claim flags.
The key finding: We wrote a passage around three invented citations, making up the authors, journals, sample sizes and reported figures. The Hallucination tab returned zero unsupported claims and marked a single claim as partially supported. A student inventing a bibliography would clear this check. The tool also refuses to scan anything under 250 characters, so a lone suspicious sentence cannot be tested on its own.

Plain Hand-Written English Was Flagged 100% AI

Test 05 . False positive
GPTZero scoring 266 words of simple hand-written text about exercise as 100 percent AI generated
Text written by hand in plain English, scored 100% AI.
Copyleaks detector scoring the same 266-word passage as zero percent AI and 100 percent human
Copyleaks read the same passage as 0% AI.
QuillBot detector marking the same passage as 100 percent human alongside a warning against relying on AI detection alone
QuillBot cleared it too, with a warning against trusting detection alone.
The real story: We typed 266 words about the benefits of exercise in short sentences and everyday vocabulary, the register a school student would use. GPTZero returned 100% AI and explained that the text resembled other AI documents it had seen. Copyleaks and QuillBot both scored the identical passage as fully human. Low-perplexity writing, meaning text where the next word is easy to guess, produces the same statistical signature a model does, and the tool cannot tell predictable human writing apart from generated writing. It is the mechanism behind reports that detectors flag non-native English at far higher rates.

The Vendor's Own Numbers Sit Awkwardly Beside the Result

Test 06 . Accuracy claims
GPTZero FAQ statement that no AI detector is fully accurate and that results should not be used to punish anyone
GPTZero's published note on the limits of AI detection.
Where it lands: GPTZero's FAQ says no detector is fully accurate and that results should not be used to punish, then adds that repeated retraining has cut its ESL false-positive rate to 1%. An independent Pangram evaluation of the updated model on the 91-essay TOEFL benchmark reported 7.7%, dropping to 1.1% only when borderline results count as passes. Any institution weighing a purchase should look hard at that gap.
Catching obvious AI and showing its reasoning

The strongest thing we confirmed was the pairing of a clean 100% verdict on raw ChatGPT text with a per-sentence breakdown that names which lines drove the score. That combination of a firm call plus a visible rationale is rare among detectors and is the honest reason to reach for GPTZero over a bare-percentage tool.

!
A false 100% AI verdict on genuine writing

The plain-English passage we wrote by hand scored 100% AI while two other detectors cleared it at zero. Anyone planning to use a GPTZero score to challenge a student needs to sit with that: the tool can be completely wrong, completely confidently, on writing a person actually wrote.

How this review was put together. Testing ran in September 2026 on a free account and the paid dashboard, using machine text from ChatGPT, historical human text from the National Archives, a fabricated-citation passage, and original hand-written prose cross-checked against Copyleaks and QuillBot. Pricing was verified directly on the official pricing page at gptzero.me/pricing, with tier limits read from the site's own feature configuration. Independent accuracy figures come from a published Pangram TOEFL evaluation.

10-Point Feature Review

03 . Scored against observed behaviour
Feature Scores at a Glance
Raw AI Text Detection
8.5
Sentence-Level Explanations
8.5
False Positive Control
4.5
Hallucination Detection
5.0
Free Plan and Access
8.0
Model and Language Coverage
8.0
Pricing and Value
6.5
Feature Breadth
7.5
Accuracy Claims vs Evidence
5.5
Support and Billing
5.0
Feature-by-feature breakdownWeighted toward detection accuracy and false-positive control
01
Raw AI Text Detection Strength
Unedited ChatGPT output scored 100% AI and held that number when we re-scanned shorter extracts, so no short-sample wobble appeared. It stops short of 9.0 because one model and one prompt is a narrow proof, and it earns more than 8.0 because the result was clean and repeatable.
Verdict: 100% AIShort samples: Held
8.5
Strong
02
Sentence-Level Explanations Strength
The Advanced Scan ranks each line by contribution and gives a reason per flag, which almost no rival offers. It does not reach 9.0 because the reasons are brief and not always actionable, and it clears 8.0 because per-sentence rationale is uncommon and useful in daily practice.
Per-sentence reasons: YesLine ranking: Yes
8.5
Strong
03
False Positive Control Weak spot
Plain hand-written English scored 100% AI while Copyleaks and QuillBot cleared the same text at zero. It sits above 4.0 because archaic human prose did pass at 93%, and it cannot reach 5.0 because a full false hit on ordinary writing is the worst possible outcome for the tool's main job.
Plain human text: 100% AICross-checks: 0% AI
4.5
Poor
04
Hallucination Detection Weak spot
Two of three fabricated citations passed with no unsupported-claim flag, and only one soft warning appeared. It holds above 4.5 because the tool did surface one nuance flag, and it stays below 5.5 because a fabricated bibliography clearing the check defeats the feature's stated purpose.
Fake sources flagged: 0 of 3Verified counter: 0 of 2
5.0
Weak
05
Free Plan and Access Strength
A permanent 10,000-word monthly plan needs no card, confirmed on the official page. It falls short of 8.5 because the free tier locks vocabulary scanning along with plagiarism and reports, and it clears 7.5 because a non-expiring free allowance is more generous than most rivals give.
Free words/mo: 10,000Card required: No
8.0
Strong
06
Model and Language Coverage Strength
The detector targets ChatGPT, GPT-5, Gemini, Claude, Llama and Deepseek, with five languages fully supported. It stops below 8.5 because non-English coverage is thinner, and it earns above 7.5 because the model list is current and wide.
Models: 6+Full languages: 5
8.0
Strong
07
Pricing and Value Fair
Entry pricing reads reasonably and the free tier lowers the risk of committing, but the headline rate rides a 30% promotion over the standard annual price, so a renewal will cost more than the sticker. It misses 7.0 because the promo hides the durable rate, and it beats 6.0 because the underlying annual pricing is competitive for the category.
Premium shown: $9.09/moStandard annual: $12.99/mo
6.5
Fair
08
Feature Breadth Strength
One workspace covers AI scanning, plagiarism, a writing report, grammar help, a Chrome extension and LMS hooks. It stays under 8.0 because several of these run shallow next to specialists, and it clears 7.0 because having the range in a single tool has real day-to-day value.
Core tools: 6API tier: Professional
7.5
Solid
09
Accuracy Claims vs Evidence Weak spot
GPTZero markets 99% accuracy and a 1% ESL false-positive rate, but an independent Pangram test of the updated model on the TOEFL set reported 7.7%, and our own plain-English miss lines up with that gap. It holds above 5.0 because recall on raw AI runs high, and it stays below 6.0 because the marketing figure and the independent figure diverge sharply.
Vendor ESL FPR: 1%Independent TOEFL: 7.7%
5.5
Weak
10
Support and Billing Weak spot
Public sentiment skews negative, with a 2.4 out of 5 Trustpilot score and recurring cancellation and unreachable-support complaints. It rises above 4.5 because G2 enterprise reviewers report smoother handling, and it stays below 5.5 because billing and support friction turns up again and again.
Trustpilot: 2.4/5Common gripe: Cancellation
5.0
Weak

Pros and Cons

04 . What stood out, good and bad
+What works
  • Reliable on unedited AI: raw ChatGPT text was caught and stayed caught across re-scans
  • Reasons, not just a number: each flagged sentence comes with a short why
  • Handles formal human prose: archaic legal writing cleared without misfiring
  • Wide model coverage: spots ChatGPT, Claude, Gemini, Llama and Deepseek output
  • Everything in one place: detection, plagiarism, a writing report and grammar under one login
  • Free to start and stay: a standing monthly allowance with no card up front
  • Names the ESL problem: the model is retrained specifically to reduce bias against second-language writers
What frustrates
  • False positives on plain writing: simple hand-written English can be called 100% AI
  • Fabricated sources slip past: invented citations cleared the hallucination check
  • 250-character floor: a single suspect sentence cannot be tested alone
  • Marketing outruns the evidence: independent false-positive figures are far higher than the published one
  • Support and cancellation gripes: reviewers report trouble reaching anyone and trouble stopping billing
  • Confusing on humanized text: output can highlight most of a passage yet report a low overall score

Pricing Breakdown

05 . Paid tiers on annual billing
Feature Premium$9.09/mo billed annuallyMost Popular Professional$17.49/mo billed annually
Words per month 500,000
Advanced AI Scan, plagiarism and downloadable reports
Batch files scanned at once 250
Page-by-page scanning
LMS integration and enterprise-grade security

What you're actually paying: The figures above ride a live 30% promotion. The standard annual rates are $12.99 a month for Premium and $24.99 for Professional, and the discount is what pulls them down to $9.09 and $17.49, so a renewal after the promo lapses will cost more than the price you first see. The page's default annual saving is 45% against monthly billing, and the promo stacks on top to read as 62%. A cheaper Essential tier sits below Premium at 150,000 words a month, but it drops plagiarism checking and downloadable reports, both gated to Premium and up. Below every paid plan, the free tier still covers 10,000 words a month for good.

GPTZero Premium and Professional plan cards on the annual billing view, each showing a 30 percent off badge over a struck-through price
GPTZero's Premium and Professional cards on the annual view, with the promotion applied.

Pricing verified September 7, 2026. Promotions and tiers shift without notice, so confirm the current numbers on the official pricing page before you subscribe.

GPTZero vs the Top 3 Alternatives

06 . How it sits against the field
GGPTZero CCopyleaks OOriginality.ai TTurnitin
Free planYes, 10,000 words/moLimited free scanNone, credit-basedInstitution only
Per-sentence reasonsNatural-language, each linePhrase highlightsHighlights, thin reasonsReport to instructor
ESL de-biasingStated, retrained for itNot emphasisedNot emphasisedNot published
Best ForEducators wanting explainable readsPlagiarism and AI in one reportPublishers and SEO teamsUniversities running an LMS

How to choose: GPTZero's edge is the pairing of a real free tier with results a non-technical reviewer can actually read. Copyleaks earns its keep when you want plagiarism and AI detection folded into a single report, and Originality.ai fits publishers who already think in credits and care about paraphrase detection more than a free front door. Turnitin only enters the conversation if your institution already runs it inside its learning platform, since there is no self-serve way in.

What Users Are Saying

07 . Voices from Trustpilot and G2
The actual program works great, however they will not let you cancel. I have multiple emails from them confirming they cancelled my account.
Trustpilot reviewer
On billing and cancellation
★★☆☆☆
trustpilot.com
It was useful until all the humanizing models came. Now it highlights more than half the text as AI and tells you 2% AI. Also there is no one to reach for customer service.
Trustpilot reviewer
On humanized text and support
★★☆☆☆
trustpilot.com
Versatility not found in other AI detectors. You can check your writing, check for AI vocabulary, run an AI scan and check for plagiarism. It's a one-stop shop.
Trustpilot reviewer
On the range of tools
★★★★★
trustpilot.com
We double-check every contract through it to confirm whether it was generated by an AI model or written by the person, so we can tell if someone is submitting fake contracts.
Verified G2 user
On enterprise contract review
★★★★☆
g2.com
Nothing to complain about, everything is fine. The only issue for me is the pricing; it feels a bit high.
Verified G2 user
On value for money
★★★★☆
g2.com
· The Verdict ·
6.8/10
A sharp reader of obvious AI that you cannot hand a verdict to

Use GPTZero if you want a fast, explainable first read on machine text, value the per-sentence reasoning as a prompt for a conversation, and want a genuine free tier plus wide model coverage while you decide.

Skip GPTZero if you plan to turn a score into an accusation, because plain human writing can come back 100% AI. For a lighter free check look at Scribbr at 7.2, or Grammarly at 7.2 if you want detection sitting next to writing help.

Category RankMixed Results
Compared To3 Alternatives
Entry Price$9.09/mo
Free TierYes

Discussion

Join the discussion and share your perspective.