The six AI detectors most people end up choosing between all advertise accuracy percentages within a point of each other. Not one leads with the thing that decides whether you can safely use it on the document in front of you: a client report under an NDA, a personal essay, an unannounced product page. Every one of them wants the same thing first, which is the whole document pasted into a box and a few seconds of trust that whatever happens on the other end matches the privacy paragraph in the footer.
That paste is the part nobody compares. This post lines up the six on the four terms that separate them, explains why free tiers stop you where they do, shows how fast a 99% accuracy claim falls apart under paraphrasing, and leaves you choosing based on what you are checking rather than on a marketing page.
The Question Vendor Pages Leave Out
Open any AI detector’s homepage and count the numbers it shows you. Accuracy percentages, supported languages, words per scan, dollars per month. Now look for a sentence explaining where your text goes after you click the button. On most of these pages that sentence lives in a privacy policy three clicks away, even though it settles something the accuracy figure never touches: whether a document you were careful with ends up on hardware you do not control.
Free and no-signup are not the same claim as private, and ZeroGPT is the cleanest illustration of that gap. It costs nothing, never asks for an account, and its own FAQ is unusually direct about the rest, stating that submissions get processed to generate a result and then discarded, on infrastructure hosted in Germany under GDPR. That is more specific than most free detectors bother publishing, and the gap between a discard policy and never transmitting the text at all is worth understanding before you decide it does not matter for your document.
A discard policy is a promise about what someone else’s servers do after they receive your words. No transmission is a property you can verify yourself by watching your own network activity. Both are honest positions, but they protect a confidential draft differently, and processing location is the one axis here that a competitor cannot revise without rebuilding the product, while every price and cap below can change with a single blog post.
Six Detectors, Side by Side
Accuracy leads every vendor page because it sounds the most like a product difference. In practice, four things decide whether a detector fits the document in front of you: whether you have to create an account, how much text you get before hitting a wall, what it costs once you outgrow the free tier, and whether your writing leaves the machine you are sitting at. The table below lines up all six on those terms, using each vendor’s own pricing and product pages.
| Detector | Signup required | Free-tier cap | Price | Text leaves your browser |
|---|---|---|---|---|
| CapyToolkit Offline & Private AI Text Detector | Never | None below 2,000,000 characters | Free, always | Never |
| GPTZero | Not for a base scan | ~10,000 characters per scan | Around $14.99/mo | Yes |
| Originality.ai | Yes, paid only | No ongoing free plan | $14.95/mo | Yes |
| Copyleaks | Not for a base scan | 25,000 characters per scan | $13.99/mo | Yes |
| QuillBot AI Detector | Effectively, yes | 1,200 words per scan, 6 scans/day | Free tier plus paid Premium | Yes |
| ZeroGPT | No | 15,000 characters per scan | Free | Yes, then discarded |
Five of those rows tell a pricing story. One tells a different story entirely. GPTZero, Originality.ai, Copyleaks, QuillBot, and ZeroGPT all send your text to a server before scoring it, and that stays true across every price point on the list, from ZeroGPT’s free scan to Originality.ai’s subscription. Paying more never changes that, because a server-based detector cannot stop being server-based by adjusting a plan. That last column describes architecture rather than packaging.
Reading the caps in real document terms
Character limits stay abstract until you translate them into pages. Roughly 10,000 characters, where GPTZero’s base scan tops out, is two or three pages of double-spaced text, enough for a student essay and nowhere near a dissertation chapter. QuillBot’s 1,200-word ceiling covers a section rather than a finished article. Copyleaks’ 25,000 anonymous characters swallow most blog posts and strain on a long report, while ZeroGPT’s 15,000 sits between.
Where a cap actually bites depends less on your longest document than on how you work. A long report hits the ceiling on the first paste, but multi-pass editing hits it in a quieter way, since a writer who checks a draft, revises, and checks again puts the same document through detection three or four times before publishing. By contrast, the 2,000,000-character ceiling on the local tool exists because a browser tab has finite memory, not because a billing tier needs a wall to sell you past.
Where the free tiers quietly stop
Copyleaks deserves credit for the most generous anonymous allowance here: 25,000 characters with no login. That free scan withholds the reasoning behind the number. Sentence-by-sentence highlighting, the phrase-level AI Logic breakdown, and saved scan history all sit behind a free account, and sustained use past the cap runs through a plan starting at $13.99 a month. You get the score for nothing; understanding the score usually costs something.
Originality.ai draws the line even earlier by running no ongoing free plan whatsoever. Limited promotional scans surface from time to time through signup offers and its browser extension, but regular use starts at a paid subscription. QuillBot splits the difference, handing free accounts a single rewrite card explaining one flagged section while reserving detailed explainers across the whole text for Premium. Across all three the pattern repeats: the headline percentage is free, and the explanation of it is the product.
What Each Paid Tier Actually Buys You
Comparing a free tool against paid ones is easy to do dishonestly, and the dishonest version treats every subscription as a removed cap with a price tag glued on. These plans buy features that exist, work, and solve problems a browser-only tool does not attempt to solve. Skipping past them would make the rest of this comparison worthless.
Hold onto one distinction here. Paying for capability you will genuinely use is a straightforward trade, while paying to lift a cap you would never have reached is a subscription bought out of vague anxiety. Which of those you are doing depends entirely on the job sitting in front of you right now.
Classroom and institutional features
GPTZero’s strongest feature has nothing to do with scoring finished text. Writing Replay records the keystroke history behind a document and plays back how it came together, and direct integrations wire the tool into Canvas and Google Classroom, which is why the classroom-oriented case for GPTZero over a browser-only detector holds up even when everything else favors local processing. Watching a document get written is a fundamentally different evidence problem from scoring the finished draft, and no client-side tool can reconstruct a history it never observed.
Copyleaks reaches further into institutions than that. Its LMS integrations cover Canvas, D2L, Moodle, Blackboard, Schoology, Edsby, and Sakai, an enterprise API exposes detection to your own systems, coverage spans 30-plus languages, and it publishes per-language accuracy figures instead of one blended number, all of which makes the institutional deployment case for Copyleaks a genuinely separate product category rather than an expensive version of the same thing. An administrator rolling detection out across hundreds of accounts needs provisioning, reporting, and an integration point, and a single-page browser tool has none of those by design.
Bundles versus a single job
Originality.ai sells consolidation more than it sells detection. One subscription covers plagiarism detection, readability scoring, grammar checking, and fact-checking, all drawing against a single credit balance, and for an editorial team screening freelance submissions that consolidation is exactly where the credit-based bundle earns its monthly cost. Juggling four tools with four logins costs a content operation real time, and one dashboard removes that friction in a way a standalone checker never will.
QuillBot builds the same argument around writing rather than editing, wrapping its detector in a paraphraser, a grammar checker, and a humanizer that share one account and one context, which gives a writer already polishing a draft there real continuity. The counterweight is straightforward: if AI likelihood is the only thing you need to know, a bundle means paying monthly for tools you will never open. Some of that ground is covered by single-purpose local pages anyway, since readability scoring and prompt token counting exist as their own free CapyToolkit tools rather than as features of a suite you subscribe to.
Why Caps and Credits Exist at All
Free-tier limits get read as greed far more often than they deserve. A server-based scan carries a real per-request cost that scales directly with how much people use it, and a cap is simply the instrument that keeps that cost predictable. Every vendor in the table above is paying for the same underlying line items every time you click the scan button:
- Inference time on a GPU or CPU, which grows with document length
- Bandwidth in both directions, since your text travels out and a result travels back
- Storage and database capacity for scan history, saved reports, and account state
- The engineering payroll that keeps it available, patched, and fast enough to feel instant
QuillBot shows the shape of this most plainly, capping each free scan at 1,200 words and limiting free accounts to six scans a day, and a single editing session covering a résumé, a cover letter, and a couple of social posts can burn through that six-scan daily allowance before lunch. Six sounds generous in the abstract. It stops sounding generous the moment you recheck anything.
Originality.ai’s credit system is the cleanest illustration, because the arithmetic is fully exposed. One credit covers 100 words of AI detection on its own, and a combined AI plus plagiarism scan costs two credits per 100 words, so a 2,000-word post costs 20 credits checked one way and 40 checked the other. Subscription credits also expire at the end of each billing cycle rather than rolling forward. When the computation happens on your own device, none of those line items exist, and nothing is left for a cap to ration, which is why every tool in this catalog of CapyToolkit’s browser-based utilities that process everything locally stays free without a plan behind it.
How Local Scoring Works Without a Server
“Runs in your browser” gets used as a slogan often enough that it stops meaning anything, so it is worth describing exactly what executes. Click Check for AI, and four scoring functions run instantly over the text sitting in the textarea. No request goes out. Nothing is queued for upload. The words never leave the tab they were pasted into.
CapyToolkit’s detector splits that work across two layers. Four explainable statistical signals score whatever text is in the box the moment you click, and an optional trained classifier adds a second, more accurate opinion once your text passes 100 words. Each layer answers a slightly different question, and the two disagree often enough that seeing both beats a single blended percentage with no reasoning attached.
The four instant signals
Sentence-length burstiness carries the most weight of the four, at 30% of the ensemble, and it measures the standard deviation of words per sentence across your document. Human writing mixes short punchy sentences with long winding ones; model output trends toward a narrower band. Transition density counts formal connectors like “furthermore,” “in conclusion,” and “it is important to note” against total length, then scores that rate against a human baseline of roughly one per hundred words and an AI-typical rate of about four. Word predictability checks your two-word phrases against a table of the 15,000 most common English word pairs, and vocabulary variety tracks lexical diversity across a moving 50-word window.1
All four are plain JavaScript arithmetic with no model behind them, which is why the score appears before you finish rereading what you pasted. When one crosses the threshold, it surfaces below your score as a named reason you can click to highlight the exact sentences behind it, so you can see which specific signal pushed a draft toward AI-typical instead of arguing with a bare number. Simple statistics earn their keep there, since every input behind the number stays inspectable.
The classifier and the 80/20 blend
The second layer is a RoBERTa-based classifier trained to separate human from machine text, quantized to int8 and weighing in around 120MB, which downloads once when your text reaches 100 words and then runs through WebAssembly on your own processor.2 Your browser caches the model afterward, so later visits skip the network entirely, and because a service worker caches the page and its scripts too, scoring keeps working with no connection at all once you have loaded the page a single time. No account gates any of that. Advanced mode isn’t a paid tier hiding behind Basic, so the classifier, the four signals, and the sentence highlighting all work on your first visit with no email address and no usage cap.
Blending the two layers at 80% classifier and 20% signal ensemble is not a round number chosen for tidiness. The classifier reports 99.28% AUROC on the RAID benchmark, while GLTR, the closest published analogue to a rank-and-frequency signal ensemble, scores 62.6% accuracy on RAID itself once every detector’s threshold is tuned to a 5% false positive rate, and weighting each method by how far it sits above random guessing lands close to 80/20 in the classifier’s favor.23 Worth stating plainly rather than burying: AUROC and accuracy are different metrics, and no published study has scored both methods head-to-head under identical conditions.
Reading Accuracy Claims Without Getting Fooled
Nearly every vendor on this list advertises a figure somewhere near 99%. When six competitors converge on the same number, that number has stopped carrying comparative information and started functioning as table stakes, which means picking between them on accuracy alone is picking between identical claims measured under conditions nobody published.
Three questions turn a headline percentage back into something you can reason about. Measured on which benchmark? Against output from which generators? And with or without adversarial paraphrasing applied first? A detector tested only against GPT-family output with no evasion attempts is answering a much easier question than one tested across eleven generators and eleven attack types.
What the benchmarks actually show
RAID is the largest published benchmark of machine-generated text. Its corpus spans more than 6 million generations across 11 generative models, 8 domains, 11 adversarial attacks, and 4 decoding strategies, and the authors benchmarked 12 detectors against it. Reading the full RAID evaluation of machine-generated text detectors is uncomfortable if you have been treating vendor numbers as settled, because its central finding is that detectors are far more fragile against unfamiliar models and adversarial phrasing than their own published figures suggest.
Independent evaluations outside the benchmark community reach the same place by a different route. A widely cited academic assessment of 14 commercial and open-source detection tools found that every one of them scored below 80% accuracy, and only 5 cleared 70%, with performance dropping consistently once text had been paraphrased.4 Those figures sit well below the marketing consensus, and researchers with nothing to sell produced them.
How fast a 99% claim collapses
One documented example makes that fragility concrete. Against unmodified machine text, RAID measured Originality at 85.0% accuracy, a genuinely strong result that collapsed to 9.3% once the same passages were rewritten with homoglyphs, meaning Cyrillic characters that render identically to their Latin counterparts.3 A detector catching six of every seven samples before the attack caught fewer than one in ten after it, with nothing changed that a human reader would notice.
Nothing about that failure is specific to Originality, and the local tool above is no more immune than the commercial ones. The same benchmark shows that no single technique breaks every detector the same way: paraphrasing actually raised Originality’s accuracy while cutting GLTR’s by more than 15 points, and the character-level tricks did the widest damage. Treat any headline figure as a claim about clean, unmodified prose, and expect a different number once someone has edited, reworded, or obfuscated what you are checking.
The False Positive Problem Every Detector Shares
The failure mode that actually damages people has nothing to do with a detector missing AI text. Real harm starts when a detector flags someone who wrote every word themselves, and that flag reaches a teacher or a hiring manager who treats a percentage as a finding. Missed detections cost a vendor a little credibility. False positives cost individual writers grades, contracts, and reputations, an asymmetry that never shows up in a marketing comparison.
The most important evidence here comes from a Stanford-affiliated study that ran 91 TOEFL essays written by non-native English speakers through seven widely used detectors. As documented in the study on detector bias against non-native English writers, those detectors misclassified the non-native essays as AI-generated at an average rate of 61.3%, while essays written by native speakers were almost never misflagged. Those detectors labeled more than half of a group of real human writers as machines, and the label tracked who they were rather than what they did.5
Mechanically, this is not a bug in any single vendor’s model. Simpler and more standard sentence structures score as statistically predictable, and predictability is precisely what these signals hunt for, so writing that leans on conventional phrasing scores higher regardless of who or what produced it. Several categories of entirely human writing trip detectors for the same underlying reason:
- Academic prose, which rewards formal connectors and consistent sentence construction
- Technical documentation, where uniform phrasing is a style requirement rather than an accident
- Heavily edited drafts, since revision tends to smooth out the variation that reads as human
- Second-language writing, for the reasons the Stanford results document directly
- Anything following a rigid template, including reports, disclosures, and structured summaries
OpenAI’s own retreat is the strongest institutional admission available on this point. The company shut down its AI Classifier on July 20, 2023, citing the tool’s low rate of accuracy, doing so as the organization with the deepest knowledge of how that text gets generated.6 If the model’s own author could not make detection reliable enough to keep shipping, no vendor’s 99% deserves to be read as a guarantee about your specific paragraph.
All of this narrows what a score is actually good for. A percentage is directional information about patterns in a piece of text, never a forensic record of authorship, and presenting it as proof of anything misuses the number regardless of which tool produced it. Explainability matters most precisely here, because a tool that names which signal fired lets you look at the sentences responsible and judge them yourself, while a bare percentage leaves you nothing to examine.
Choosing Based on What You’re Actually Checking
The useful question has never been which detector wins. What you are checking, and what you are willing to transmit, are the two answers that together eliminate most of the list before accuracy ever enters the conversation. A teacher screening thirty essays has a different problem from a writer checking a personal essay nobody has seen yet. Work down this path in order and the choice usually settles itself:
- If the text is confidential, client-owned, under NDA, or simply unannounced, keep it local and stop there. No accuracy advantage justifies transmitting a document you were not supposed to share.
- If you are deploying detection across a classroom or an institution, GPTZero and Copyleaks are the two that build for it, and the LMS integration you need decides between them.
- If your writing already happens inside a suite, use the detector bundled with it. The continuity is worth more than a marginal accuracy difference you cannot verify anyway.
- If it is a one-off check on text you would not mind emailing to a stranger, any free option in the table works, and ZeroGPT asks the least of you.
For confidential drafts specifically, the local habit tends to generalize. If you would rather not paste a client report into a detector, you probably also want to strip names, emails, and account numbers out of a document before it reaches any AI service, which is the same instinct applied one step earlier in the workflow. Both come down to deciding what leaves your machine, rather than reading a privacy policy afterward and hoping.
Pricing shifts, caps get raised, and accuracy claims get restated with every release, so a comparison built on those three alone expires quickly. Where a tool computes your score is an architectural decision its builders made at the start. Worth ending on the honest note, though: not one detector here, paid or free, can prove who actually wrote a piece of text. The question worth asking is which one tells you enough about its own reasoning to be worth consulting at all.
- 1.
Peter Norvig, “Natural Language Corpus Data: Beautiful Data,” norvig.com, accessed August 2026. https://norvig.com/ngrams/
- 2.
ONNX Community, “TMR: Target Mining RoBERTa AI Text Detector,” huggingface.co, accessed August 2026. https://huggingface.co/onnx-community/tmr-ai-text-detector-ONNX
- 3.
Liam Dugan, Alyssa Hwang, Filip Trhlik, Josh Magnus Ludan, Andrew Zhu, Hainiu Xu, Daphne Ippolito, and Chris Callison-Burch, “RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors,” Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, 2024, pp. 12463–12492.
- 4.
Debora Weber-Wulff et al., “Testing of Detection Tools for AI-Generated Text,” link.springer.com, December 2023. https://link.springer.com/article/10.1007/s40979-023-00146-z
- 5.
Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu, and James Zou, “GPT Detectors Are Biased Against Non-Native English Writers,” Patterns, vol. 4, no. 7, July 2023.
- 6.
OpenAI, “New AI Classifier for Indicating AI-Written Text,” openai.com, January 2023. https://openai.com/index/new-ai-classifier-for-indicating-ai-written-text/