gptzero.me · AI text detector, now part of Superhuman
GPTZero is the AI detector that shows its work. Its Advanced Scan ranks a document line by line and gives a short reason for each flag instead of handing back a bare percentage. In our session it caught unedited ChatGPT output at 100% and held that score on shorter extracts, yet it also marked 266 words of plain, hand-written English as 100% AI. It suits educators and editors who want an explainable first read on authenticity, so long as they treat the result as a starting point for a conversation rather than proof.
GPTZero reads a block of text and estimates how much of it was written by a large language model. Edward Tian built the first version over a Princeton winter break in early 2023, and the product has since grown into a detection suite used across education, publishing, hiring and legal review. The company says it has served more than 17 million users and now sits inside Superhuman.
The pitch rests on two claims. First, accuracy: GPTZero markets a 99% detection rate and says it is the only detector de-biased for English-as-a-second-language writers. Second, interpretability, which is where it tries to separate itself from the pack. Rather than returning a single number, the Advanced Scan highlights individual sentences and explains why each was flagged. Access starts with a permanent free plan capped at 10,000 words a month, then climbs through paid Essential, Premium and Professional tiers, with Premium shown at $9.09 a month on annual billing at the time of writing.
What Happened When We Tested It
02 . 6 tests spanning raw detection, false positives and the accuracy gap
We worked from a free account plus the paid dashboard, feeding the detector a mix of machine text, historical human text and writing produced by hand. The goal was not to confirm the marketing figure but to find the edges where a detector like this helps a teacher and where it could get a student wrongly accused. The session moved from the easy cases to the ones that expose the model's weak spots.
Unedited ChatGPT Output Was Caught at Once
Test 01 . Raw AI
★ Prompt Used
"Write a 490-word essay on the causes of the First World War."
Unedited ChatGPT output returned a 100% AI verdict on the basic scan.
What we observed: The essay came back at 100% AI, with 0% assigned to both the human and mixed buckets. Re-scanning shorter extracts of the same text held the score each time, so the instability that critics often pin on short samples did not show up here.
The Sentence-Level View Carried the Session
Test 02 . Advanced Scan
The Advanced Scan ranks each sentence and offers a reason for every flag.
The signal here: Expanding a flagged line reveals a short explanation of why it scored the way it did. For a teacher looking at a flagged assignment, seeing which sentences drove the number beats being handed a percentage with no reasoning attached. Most rival detectors stop at the headline figure.
Eighteenth-Century Human Text Passed Cleanly
Test 03 . Historical prose
The Bill of Rights was read as human at 93%.
Worth noting: We pasted 248 words of the Bill of Rights from the National Archives transcription. That text predates language models by more than two centuries, so any AI score would be a false positive by definition. GPTZero cleared it at 93% human. Formal, archaic legal prose is exactly the writing critics expect a detector to misread, and it did not.
The Hallucination Detector Waved Through Fake Sources
Test 04 . Citation check
Three fabricated citations produced one soft warning and no unsupported-claim flags.
The key finding: We wrote a passage around three invented citations, making up the authors, journals, sample sizes and reported figures. The Hallucination tab returned zero unsupported claims and marked a single claim as partially supported. A student inventing a bibliography would clear this check. The tool also refuses to scan anything under 250 characters, so a lone suspicious sentence cannot be tested on its own.
Plain Hand-Written English Was Flagged 100% AI
Test 05 . False positive
Text written by hand in plain English, scored 100% AI.Copyleaks read the same passage as 0% AI.QuillBot cleared it too, with a warning against trusting detection alone.
The real story: We typed 266 words about the benefits of exercise in short sentences and everyday vocabulary, the register a school student would use. GPTZero returned 100% AI and explained that the text resembled other AI documents it had seen. Copyleaks and QuillBot both scored the identical passage as fully human. Low-perplexity writing, meaning text where the next word is easy to guess, produces the same statistical signature a model does, and the tool cannot tell predictable human writing apart from generated writing. It is the mechanism behind reports that detectors flag non-native English at far higher rates.
The Vendor's Own Numbers Sit Awkwardly Beside the Result
Test 06 . Accuracy claims
GPTZero's published note on the limits of AI detection.
Where it lands: GPTZero's FAQ says no detector is fully accurate and that results should not be used to punish, then adds that repeated retraining has cut its ESL false-positive rate to 1%. An independent Pangram evaluation of the updated model on the 91-essay TOEFL benchmark reported 7.7%, dropping to 1.1% only when borderline results count as passes. Any institution weighing a purchase should look hard at that gap.
★
Catching obvious AI and showing its reasoning
The strongest thing we confirmed was the pairing of a clean 100% verdict on raw ChatGPT text with a per-sentence breakdown that names which lines drove the score. That combination of a firm call plus a visible rationale is rare among detectors and is the honest reason to reach for GPTZero over a bare-percentage tool.
!
A false 100% AI verdict on genuine writing
The plain-English passage we wrote by hand scored 100% AI while two other detectors cleared it at zero. Anyone planning to use a GPTZero score to challenge a student needs to sit with that: the tool can be completely wrong, completely confidently, on writing a person actually wrote.
How this review was put together. Testing ran in September 2026 on a free account and the paid dashboard, using machine text from ChatGPT, historical human text from the National Archives, a fabricated-citation passage, and original hand-written prose cross-checked against Copyleaks and QuillBot. Pricing was verified directly on the official pricing page at gptzero.me/pricing, with tier limits read from the site's own feature configuration. Independent accuracy figures come from a published Pangram TOEFL evaluation.
10-Point Feature Review
03 . Scored against observed behaviour
Feature Scores at a Glance
Raw AI Text Detection
8.5
Sentence-Level Explanations
8.5
False Positive Control
4.5
Hallucination Detection
5.0
Free Plan and Access
8.0
Model and Language Coverage
8.0
Pricing and Value
6.5
Feature Breadth
7.5
Accuracy Claims vs Evidence
5.5
Support and Billing
5.0
Feature-by-feature breakdownWeighted toward detection accuracy and false-positive control
01
Raw AI Text Detection Strength
Unedited ChatGPT output scored 100% AI and held that number when we re-scanned shorter extracts, so no short-sample wobble appeared. It stops short of 9.0 because one model and one prompt is a narrow proof, and it earns more than 8.0 because the result was clean and repeatable.
Verdict: 100% AIShort samples: Held
8.5
Strong
02
Sentence-Level Explanations Strength
The Advanced Scan ranks each line by contribution and gives a reason per flag, which almost no rival offers. It does not reach 9.0 because the reasons are brief and not always actionable, and it clears 8.0 because per-sentence rationale is uncommon and useful in daily practice.
Per-sentence reasons: YesLine ranking: Yes
8.5
Strong
03
False Positive Control Weak spot
Plain hand-written English scored 100% AI while Copyleaks and QuillBot cleared the same text at zero. It sits above 4.0 because archaic human prose did pass at 93%, and it cannot reach 5.0 because a full false hit on ordinary writing is the worst possible outcome for the tool's main job.
Plain human text: 100% AICross-checks: 0% AI
4.5
Poor
04
Hallucination Detection Weak spot
Two of three fabricated citations passed with no unsupported-claim flag, and only one soft warning appeared. It holds above 4.5 because the tool did surface one nuance flag, and it stays below 5.5 because a fabricated bibliography clearing the check defeats the feature's stated purpose.
Fake sources flagged: 0 of 3Verified counter: 0 of 2
5.0
Weak
05
Free Plan and Access Strength
A permanent 10,000-word monthly plan needs no card, confirmed on the official page. It falls short of 8.5 because the free tier locks vocabulary scanning along with plagiarism and reports, and it clears 7.5 because a non-expiring free allowance is more generous than most rivals give.
Free words/mo: 10,000Card required: No
8.0
Strong
06
Model and Language Coverage Strength
The detector targets ChatGPT, GPT-5, Gemini, Claude, Llama and Deepseek, with five languages fully supported. It stops below 8.5 because non-English coverage is thinner, and it earns above 7.5 because the model list is current and wide.
Models: 6+Full languages: 5
8.0
Strong
07
Pricing and Value Fair
Entry pricing reads reasonably and the free tier lowers the risk of committing, but the headline rate rides a 30% promotion over the standard annual price, so a renewal will cost more than the sticker. It misses 7.0 because the promo hides the durable rate, and it beats 6.0 because the underlying annual pricing is competitive for the category.
Premium shown: $9.09/moStandard annual: $12.99/mo
6.5
Fair
08
Feature Breadth Strength
One workspace covers AI scanning, plagiarism, a writing report, grammar help, a Chrome extension and LMS hooks. It stays under 8.0 because several of these run shallow next to specialists, and it clears 7.0 because having the range in a single tool has real day-to-day value.
Core tools: 6API tier: Professional
7.5
Solid
09
Accuracy Claims vs Evidence Weak spot
GPTZero markets 99% accuracy and a 1% ESL false-positive rate, but an independent Pangram test of the updated model on the TOEFL set reported 7.7%, and our own plain-English miss lines up with that gap. It holds above 5.0 because recall on raw AI runs high, and it stays below 6.0 because the marketing figure and the independent figure diverge sharply.
Vendor ESL FPR: 1%Independent TOEFL: 7.7%
5.5
Weak
10
Support and Billing Weak spot
Public sentiment skews negative, with a 2.4 out of 5 Trustpilot score and recurring cancellation and unreachable-support complaints. It rises above 4.5 because G2 enterprise reviewers report smoother handling, and it stays below 5.5 because billing and support friction turns up again and again.
Trustpilot: 2.4/5Common gripe: Cancellation
5.0
Weak
Pros and Cons
04 . What stood out, good and bad
+What works
Reliable on unedited AI: raw ChatGPT text was caught and stayed caught across re-scans
Reasons, not just a number: each flagged sentence comes with a short why
Handles formal human prose: archaic legal writing cleared without misfiring
Wide model coverage: spots ChatGPT, Claude, Gemini, Llama and Deepseek output
Everything in one place: detection, plagiarism, a writing report and grammar under one login
Free to start and stay: a standing monthly allowance with no card up front
Names the ESL problem: the model is retrained specifically to reduce bias against second-language writers
−What frustrates
False positives on plain writing: simple hand-written English can be called 100% AI
Fabricated sources slip past: invented citations cleared the hallucination check
250-character floor: a single suspect sentence cannot be tested alone
Marketing outruns the evidence: independent false-positive figures are far higher than the published one
Support and cancellation gripes: reviewers report trouble reaching anyone and trouble stopping billing
Confusing on humanized text: output can highlight most of a passage yet report a low overall score
Pricing Breakdown
05 . Paid tiers on annual billing
Feature
Premium$9.09/mo billed annuallyMost Popular
Professional$17.49/mo billed annually
Words per month
300,000
500,000
Advanced AI Scan, plagiarism and downloadable reports
✓
✓
Batch files scanned at once
50
250
Page-by-page scanning
✗
✓
LMS integration and enterprise-grade security
✗
✓
What you're actually paying: The figures above ride a live 30% promotion. The standard annual rates are $12.99 a month for Premium and $24.99 for Professional, and the discount is what pulls them down to $9.09 and $17.49, so a renewal after the promo lapses will cost more than the price you first see. The page's default annual saving is 45% against monthly billing, and the promo stacks on top to read as 62%. A cheaper Essential tier sits below Premium at 150,000 words a month, but it drops plagiarism checking and downloadable reports, both gated to Premium and up. Below every paid plan, the free tier still covers 10,000 words a month for good.
GPTZero's Premium and Professional cards on the annual view, with the promotion applied.
Pricing verified September 7, 2026. Promotions and tiers shift without notice, so confirm the current numbers on the official pricing page before you subscribe.
GPTZero vs the Top 3 Alternatives
06 . How it sits against the field
GGPTZero
CCopyleaks
OOriginality.ai
TTurnitin
Free plan
Yes, 10,000 words/mo
Limited free scan
None, credit-based
Institution only
Per-sentence reasons
Natural-language, each line
Phrase highlights
Highlights, thin reasons
Report to instructor
ESL de-biasing
Stated, retrained for it
Not emphasised
Not emphasised
Not published
Best For
Educators wanting explainable reads
Plagiarism and AI in one report
Publishers and SEO teams
Universities running an LMS
How to choose: GPTZero's edge is the pairing of a real free tier with results a non-technical reviewer can actually read. Copyleaks earns its keep when you want plagiarism and AI detection folded into a single report, and Originality.ai fits publishers who already think in credits and care about paraphrase detection more than a free front door. Turnitin only enters the conversation if your institution already runs it inside its learning platform, since there is no self-serve way in.
What Users Are Saying
07 . Voices from Trustpilot and G2
The actual program works great, however they will not let you cancel. I have multiple emails from them confirming they cancelled my account.
It was useful until all the humanizing models came. Now it highlights more than half the text as AI and tells you 2% AI. Also there is no one to reach for customer service.
Versatility not found in other AI detectors. You can check your writing, check for AI vocabulary, run an AI scan and check for plagiarism. It's a one-stop shop.
We double-check every contract through it to confirm whether it was generated by an AI model or written by the person, so we can tell if someone is submitting fake contracts.
A sharp reader of obvious AI that you cannot hand a verdict to
Use GPTZero if you want a fast, explainable first read on machine text, value the per-sentence reasoning as a prompt for a conversation, and want a genuine free tier plus wide model coverage while you decide.
Skip GPTZero if you plan to turn a score into an accusation, because plain human writing can come back 100% AI. For a lighter free check look at Scribbr at 7.2, or Grammarly at 7.2 if you want detection sitting next to writing help.
Discussion
Join the discussion and share your perspective.