AI Humanizer Benchmark for Pangram
Table of Contents
What Current AI Humanizer Benchmarks Test
Why Pangram Is the Detector That Matters in 2026
How Leaving Out Pangram Hurts a Leaderboard's Credibility
What Happened When Pangram 4 Caught Up With Us
Introducing HumanizerEval
How to Read Any AI Humanizer Benchmark
Test Super Against Pangram Yourself
FAQ
You found an AI humanizer benchmark with a clean leaderboard, weighted scores and a GitHub repo full of raw data. It looks like exactly the research you wanted before paying for a tool. Then you scroll to the bottom and read who built it.
On both of the best known humanizer leaderboards, the company running the site also makes the humanizer ranked first. And neither one tests against Pangram, the AI detector that independent researchers keep rating as the hardest to fool.
That second problem is the bigger one. A humanizer ranking that leaves out the toughest detector on the market is measuring how well tools beat the easy tests. Here's what those benchmarks check, why Pangram belongs in any serious evaluation, and why we built HumanizerEval to put it front and center.
What Current AI Humanizer Benchmarks Test
Credit where it's due: both sites publish their inputs, outputs, and detector verdicts publicly. That's more transparency than most "best humanizer" listicles offer. But transparency about a test doesn't fix what the test leaves out.
HumanizerBench is operated by WriteHuman, and says so on the site. As of its September 1, 2026 update:
14 humanizers, 33 prompts each, across eight writing categories
Five detectors: GPTZero, Originality.ai, Copyleaks, Winston AI and ZeroGPT
Scoring: bypass rate (42%), meaning preservation (32%), readability (16%), consistency (10%)
First place: WriteHuman, at 78.29
AI Humanizer Benchmark is run by the team behind UndetectedGPT, which it also discloses. As of its September 3, 2026 update:
11 humanizers tested on 33 texts across seven categories
Seven detectors: GPTZero, Originality.ai, Copyleaks, Winston AI, ZeroGPT, QuillBot and Grammarly
The same 42/32/16/10 weighting as HumanizerBench
First place: UndetectedGPT, at 84.9
Look at the detector lists again. Pangram isn't on either one.
Why Pangram Is the Detector That Matters in 2026
Pangram is the detector that was specifically trained to catch humanized text, and the one academic researchers have singled out most often.
When Pangram shipped its fourth model on July 29, 2026, the company claimed that Pangram 4 detects humanized AI output 98.83% of the time across 13 commercial humanizers, with a false positive rate of 0.0041%. Those are vendor numbers, so treat them the way you'd treat any vendor number (including ours). The independent research is what makes Pangram hard to ignore:
A University of Chicago working paper by Brian Jabarian and Alex Imas, comparing Pangram, Originality.ai, GPTZero and an open-source RoBERTa model, found Pangram's false positive and false negative rates were close to zero on medium and long passages. When the researchers ran text through a humanizer, GPTZero's false negative rate climbed to around 0.50 or higher; Pangram's stayed low.
A June 2026 peer-reviewed study in the International Journal for Educational Integrity, evaluating GPTZero, Pangram, Copyleaks and Turnitin on student papers, found Pangram reached 95% inclusive accuracy on hybrid and humanized texts. GPTZero scored 0% on the hybrid papers.
Adoption matters too. Pangram 4 launched alongside a $9 million round led by Menlo Ventures, and VentureBeat reported that Substack and Quora use Pangram along with academic institutions and media organizations. If your writing gets checked somewhere that takes detection seriously, Pangram is increasingly the tool doing the checking.
How Leaving Out Pangram Hurts a Leaderboard's Credibility
We aren't saying the other benchmarks faked anything. Their data is public, and anyone can check it. The problem is the design of the test, and it shows up in four ways.
1. They grade humanizers against the detectors that are easiest to beat
GPTZero appears on both leaderboards. It's also the detector that, in the University of Chicago study, missed roughly half of humanized AI text. A humanizer that scores 85 against detectors like that has proven it can clear a bar that research already shows is low. It hasn't proven anything about the bar that matters.
2. They don't reflect where your text actually gets checked
A freelance writer submitting to a Substack publication, or a student whose university has adopted Pangram, doesn't care how a tool performs on ZeroGPT. Readers use these leaderboards to predict real outcomes. When the strictest detector in real-world use is missing, the ranking can't make that prediction.
3. The conflict of interest has nowhere to hide
Both operators disclose that they sell humanizers, and both of their tools rank first. That alone doesn't mean the results are wrong. But when the hardest available test is absent, readers are left guessing whether the first-place tool would still hold its spot against Pangram. A leaderboard run by a competitor should work harder than a neutral one to close off that question, not leave it open.
4. The results are already out of date
As of this writing, HumanizerBench was last updated September 1 and AI Humanizer Benchmark on September 3. StealthGPT released Super, our new humanization model, on September 8. Whatever StealthGPT score those sites show (9th and 7th, respectively) reflects the model Super replaced.
What Happened When Pangram 4 Caught Up With Us
Here's the part most humanizer companies wouldn't write about.
Pangram has tested StealthGPT for a while. In its August 2025 humanizer benchmark, Pangram reported catching our older model's output 95.6% of the time, and the University of Chicago researchers used StealthGPT as the humanizer in their study. When Pangram 4 launched, our founder Jozef Gherman put it bluntly, "it broke our previous models."
So we rebuilt. Super was trained with Pangram 4 as the target, not as an afterthought. In our internal benchmark run on September 6, 2026, Super passed Pangram 4 on 89% of samples; the same run showed 97% on GPTZero, 94% on Originality.ai and 92% on Winston AI. We walk through that test, and the gaps we found in Pangram's own accuracy claims, in our breakdown of whether Pangram's AI detector is really as accurate as it says.
Those are our numbers from our dataset. You shouldn't take them on faith, and that's exactly the point of the next section.
Introducing HumanizerEval
HumanizerEval is our AI humanizer leaderboard, built around the test the other two skip. Every humanizer on it is run against Pangram, alongside GPTZero, Originality.ai, Winston AI, & ZeroGPT.
What you'll find there:
10 humanizers tested on identical inputs using default settings
Pangram results reported for every tool, not folded into an average where one strict detector can get drowned out
Meaning preservation and readability scored alongside bypass rate, because a humanized essay that no longer says what you wrote isn't worth submitting
Every input, output and detector verdict published on GitHub and HumanizerEval.
If a competitor beats Super on Pangram, the data will show it, and anyone can pull the repo and check. We think Super holds up. HumanizerEval is how we let you verify that instead of asking you to believe it.
How to Read Any AI Humanizer Benchmark
Whether you're looking at our leaderboard or someone else's, run it through these questions before you trust a ranking:
Who operates it? If the first-place tool belongs to the site owner, look harder at everything below.
Which detectors are included? No Pangram means no evidence against the detector most likely to catch you. Turnitin matters for students, too.
When was it last updated? Humanizers and detectors both ship new models every few months. A ranking from before a major release is describing tools that no longer exist.
Are the full outputs public, or just scores? You want to be able to read what the humanizer actually produced.
Is meaning preservation scored? Bypass rates are easy to inflate by rewriting text into something the author never said.
How long are the test passages? The University of Chicago data showed detectors behave differently on short text, so a benchmark built only on short samples can flatter weak tools.
If you're newer to how humanizers work in the first place, our guide on how to humanize AI text and bypass every AI detector for free covers the basics before you start comparing leaderboards.
Test Super Against Pangram Yourself
The fastest way to settle the question is to run your own text. Paste a draft into StealthGPT's AI Humanizer with the Super model, then check the result on Pangram directly. If you need Super at volume, it's available through our API and MCP integration; see plans and API pricing to get started, no credit card required.
FAQ
Does HumanizerBench test against Pangram?
No. As of its September 1, 2026 update, HumanizerBench tests against GPTZero, Originality.ai, Copyleaks, Winston AI and ZeroGPT. AI Humanizer Benchmark adds QuillBot and Grammarly, but it doesn't include Pangram either.
Who runs the major AI humanizer leaderboards?
HumanizerBench is operated by WriteHuman, and AI Humanizer Benchmark is run by the team behind UndetectedGPT. Both disclose this, and both rank their own tool first.
Is Pangram more accurate than GPTZero?
Independent research says yes. A University of Chicago study found GPTZero missed roughly half of humanized AI text while Pangram's miss rate stayed low, and a June 2026 peer-reviewed study found Pangram far ahead of GPTZero, Copyleaks and Turnitin on hybrid and humanized papers.
Can any AI humanizer bypass Pangram?
Pangram 4 is the hardest detector we've tested against. In our internal September 2026 benchmark, StealthGPT's Super model passed Pangram 4 on 89% of samples. HumanizerEval publishes Pangram results for every tool it ranks, so you can compare humanizers on the same test instead of relying on each company's own claims.
Why don't other humanizer benchmarks include Pangram?
Neither site explains the omission. Whatever the reason, the result is the same: their rankings can't tell you how a humanizer performs against the detector that was built specifically to catch humanized text.