AI humanizer benchmark: the best tools of August 2026
The August 2026 cycle of HumanizerBench tested 12 AI humanizers against 33 prompts and five commercial AI detectors. WriteHuman finished first, ahead of Humanize AI Pro and Stealth Writer.
Ten of the twelve finished in a different position than they held in July. That looks like a turbulent month, and underneath the rankings it mostly wasn’t. Setting aside StealthGPT, whose scores moved a long way and are covered below, the other ten tools that attempt evasion averaged 0.738 detector bypass in July and 0.742 in August. The field did roughly what it did last month. What moved the board was the quality-penalty layer and a single category that almost nobody can beat.
Business email is the category the whole board fails
Averaged across the eleven tools that attempt evasion, business email cleared detectors 17.5 percent of the time. Every other category landed between 60 and 83 percent. Four tools scored under 0.02, meaning essentially every business email they produced was caught.
The obvious explanations don’t hold up. It isn’t length: those samples average 250 words, almost exactly the same as discussion board posts, which the field cleared at 80 percent. It isn’t that the tools left the text alone either, since verbatim overlap between input and output on business email is about what it is on blog posts.
What the detector-level data shows is narrower. On August’s business emails, Copyleaks scored the field at 0.061 and Winston at 0.075, against 0.72 to 0.85 on the categories they found normal. GPTZero sat at 0.246. Because each test’s bypass score is the median of five detectors, three low votes decide it, and Originality.ai at 0.502 and ZeroGPT at 0.794 cannot pull the median back up. Why those three detectors treat this register so differently is not something the published files can answer. We can show it happens, consistently, to every tool.
Two caveats belong with that finding. Business email is three samples per cycle, and it has only existed since July, so this is the second month of evidence, not a long trend. The prompt topics also rotate every cycle by design, so a category average is never a clean month-over-month comparison. What survives both caveats is that two different business-email topics, in two consecutive cycles, produced the worst category result on the board by a wide margin.
Why WriteHuman finished first in August 2026
WriteHuman posted the highest bypass rate in the field at 0.891 and took a single penalty point, for one length-inflated output. It is the only tool this cycle that paired top-of-field evasion with an essentially clean fidelity record, which is what the composite is built to reward.
It also led on the pre-penalty score, at 77.69 against Undetectable AI’s 74.48. That is a change from July, when the order at the top was decided by the penalty layer: before penalties were applied that month, Undetectable AI was ahead by 8.10 points.
Where the margin comes from is worth knowing if you are choosing a tool. WriteHuman’s composite of 76.69 sits 5.29 points above Humanize AI Pro, and most of that separation is business email, the category covered above, where WriteHuman scored 0.536 against 0.010. Recompute the cycle without those three samples and the gap is 0.94 points, with Undetectable AI moving from eighth to sixth. That recomputation is ours, run with the cycle’s own published scoring script rather than published as a result. If business email is not the kind of writing you do, first and second place are closer than the composite suggests.
HumanizerBench is operated by WriteHuman, which is why the scoring works the way it does: the weights and penalty rules are published in advance, the script that applies them ships alongside the raw data for each cycle, and anyone can re-run it. The why we built this page covers that arrangement in full.
Undetectable AI had the second-best bypass rate and finished eighth
Thirteen points separate Undetectable AI’s raw score from its published one. Ten came from length inflation, which it triggered on 24 of 33 outputs, enough to hit that penalty’s cap. Three more came from meaning drift on three outputs that fell below the 0.85 threshold. Its bypass rate of 0.860 was second only to WriteHuman’s 0.891, and 2.5 points clear of third.
Without those penalties it would have finished second at 74.48. That counterfactual is ours, not a published figure, and the cap works in its favor as well as against it: 27 penalty occurrences were clipped to 13 points applied.
Its median output ran 1.63 times the length of its input in August, 1.59 in July, and 1.53 in June. Whether that is a drift or a fixed property of the product, three cycles cannot tell you. What it does mean for someone using the tool is concrete. You paste in 250 words and get back roughly 400, which you then have to cut. Padding looks like a cheap route to a detector score, and the length penalty exists to measure the gap between a bypass number obtained that way and the rewrite the user actually asked for.
Undetectable AI’s six-place fall came from three roughly equal causes, not one: bypass cost it 4.09 points, category consistency 3.21, and penalties 3.00. The consistency loss is almost entirely the business-email cliff, where it went from 0.976 in July to 0.018 in August while its six other categories all stayed between 0.86 and 1.00. It has not stopped beating detectors. On the 30 non-business-email samples it posts the highest bypass rate of any tool in the cycle.
StealthGPT’s detection scores moved against it on all five detectors
StealthGPT’s bypass rate went from 0.816 to 0.275, and it moved from eighth to twelfth. The change is spread evenly across the panel: Originality.ai 0.784 to 0.202, Copyleaks 0.828 to 0.242, Winston 0.780 to 0.252, ZeroGPT 0.801 to 0.460, GPTZero 0.805 to 0.526.
The rest of its profile moved the other way. All 33 tests returned complete output, its readability rose from 0.606 to 0.709, its meaning preservation was flat, and its penalties halved from 10 points to 5. On the published quality measures it is a better rewrite than it was in July.
Five detectors from five vendors shifting against one tool in the same month, in a cycle where the rest of the field was flat, points to something changing on StealthGPT’s side rather than five independent detector updates. The published data cannot establish that, and we have no visibility into its pipeline, so we note it as the reading the numbers support rather than a conclusion.
A four-place climb that was not an improvement
ai-humanize-io rose from tenth to sixth, the largest gain of the cycle. Its bypass rate fell, its meaning preservation fell, and its readability fell. It climbed because it stopped padding: length-inflation penalties went from eight occurrences to zero, and its total penalty went from 10 points to 2. Its pre-penalty score actually dropped 3.09 points.
Rank movement and capability movement are different things, which is a reason to read the sub-score columns rather than the position.
GPTZero disagrees with the other four detectors
GPTZero has the widest spread of any detector in the cycle, from 0.131 to 0.884, and it splits the field into two groups rather than ranking it. In all three published cycles, the same four tools have come last against GPTZero: HIX Bypass, Humbot, Super Humanizer, and Stealth Writer. Three of them are far weaker against GPTZero than against everything else. HIX Bypass scores 0.131 on GPTZero and averages 0.725 across the other four. Humbot scores 0.218 against 0.682. Stealth Writer scores 0.308 against 0.735.
Those tools are not weak. HIX Bypass finished fourth overall. They appear to have been built against a detector signature that GPTZero does not share, and because the composite uses a median of five detectors, one dissenting detector rarely changes a tool’s overall number. If GPTZero is the detector you actually care about, the per-detector pages will serve you better than the main ranking.
Across the field, Originality.ai remained the hardest detector to beat in all three cycles, at a field average of 0.495 in August. ZeroGPT remained the easiest at 0.786.
Editing is not evasion: what Grammarly shows
Grammarly functions as the control in every cycle. Its bypass rate is exactly zero in all seven categories. Every one of its 33 outputs was identified. It also posted the highest meaning preservation in the field at 0.943 and took no penalties in any of the three cycles, because it barely changes length or meaning. It is an editor, not a humanizer, and the distinction gets blurred constantly in search results.
Grammarly finishing eleventh rather than last is worth explaining, because it exposes something about the formula. Category consistency is scored as the spread between a tool’s category results, so a tool that is caught uniformly scores a perfect 1.0 and collects the full 10 points. Grammarly earns those points for being consistently detected. That is a real quirk of the composite, and it is part of why a tool with zero evasion finishes above StealthGPT.
How to reproduce the August 2026 results
Every figure above comes from the published August cycle: the prompts, the 33
input samples, all 396 humanized outputs from the twelve tools, every
response from all five detectors, and the frozen scoring script that produced the
composites. Clone the
public repository and run
npm run verify to recompute the leaderboard from the raw files. The weights and
penalty rules are on the methodology page.
Where this analysis goes beyond the published leaderboard, it says so. The pre-penalty scores and the recomputation without business email are ours, produced with the cycle’s own scoring script over the cycle’s own data. Anyone can run the same checks against the same files and get the same answers.