Skip to content
The September 2026 results are in: WriteHuman holds #1, and newcomer SupWriter debuts at #5.Read the full analysis →
HumanizerBench
HumanizerBench

September 2026 AI Humanizer Rankings

WriteHuman finished first at 78.29, ahead of Stealth Writer at 75.39 and HIX Bypass at 73.97. Fourteen tools are ranked. SupWriter is the strongest debut, landing fifth at 72.80 on its first cycle; the other two newcomers, StealthBypass and Penlify, took the bottom two places.

September 2026 leaderboard Full interactive table →
#HumanizerOverallBypassMeaningRead.Pen.
1 -
78.2990.172.356.9
None
2 ▲ 1
75.3986.476.347.2
−1.0Penalties applied
  • Meaning drift−1.0
    Applied 1× max −10.0

    The output's meaning drifted significantly from the original input.

3 ▲ 1
73.9786.376.934.9
None
4 ▲ 2
72.9481.573.258.4
−1.0Penalties applied
  • Length inflation−1.0
    Applied 1× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

5 -
72.8080.379.045.7
None
6 ▲ 2
71.7582.182.161.7
−6.0Penalties applied
  • Length inflation−6.0
    Applied 6× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

7 ▼ 5
70.0770.874.761.8
None
8 ▼ 1
67.5882.166.165.4
−6.0×2Penalties applied
  • Meaning drift−1.0
    Applied 1× max −10.0

    The output's meaning drifted significantly from the original input.

  • Length inflation−5.0
    Applied 5× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

9 ▲ 3
65.1079.659.470.9
−5.0Penalties applied
  • Meaning drift−5.0
    Applied 5× max −10.0

    The output's meaning drifted significantly from the original input.

10 ▼ 1
61.7083.966.660.4
−12.0×2Penalties applied
  • Meaning drift−2.0
    Applied 2× max −10.0

    The output's meaning drifted significantly from the original input.

  • Length inflation−10.0
    Applied 12× max −10.0at max

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

11 ▼ 1
60.9764.474.252.1
−5.0Penalties applied
  • Length inflation−5.0
    Applied 5× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

12 ▼ 1
60.3541.186.466.1
−3.0Penalties applied
  • Length inflation−3.0
    Applied 3× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

13 -
39.3878.639.554.1
−23.0×4Penalties applied
  • Identical to input−2.0
    Applied 1× max −10.0

    The tool returned the input mostly unchanged. No real humanization happened.

  • Meaning drift−10.0
    Applied 17× max −10.0at max

    The output's meaning drifted significantly from the original input.

  • Length inflation−1.0
    Applied 1× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

  • Length deflation−10.0
    Applied 21× max −10.0at max

    The output came back much shorter than the input. The rewrite dropped content instead of paraphrasing it.

14 -
36.570.062.771.8
−5.0Penalties applied
  • Meaning drift−5.0
    Applied 5× max −10.0

    The output's meaning drifted significantly from the original input.

Overall is 0 to 100; bypass, meaning, and readability are 0 to 100 (higher is better). Penalties are points deducted for output-quality issues.

Almost everyone scored higher, and one tool fell

Ten of the eleven tools tested in both August and September scored higher this month. Those eleven are compared against each other throughout this section.

StealthGPT improved most, gaining 19.55 points and climbing from last of the eleven to eighth. It is almost entirely one thing: its bypass rate went from 0.2753 to 0.7961, which is worth about 22 points by itself. Slight losses on meaning and consistency take it back to the net 19.55.

Humanize AI Pro is the one tool that fell. August’s runner-up lost 1.33 points, falling from second of the eleven to sixth. It got better at everything except evading detectors, with meaning and readability both up, but its bypass rate fell from 0.8354 to 0.7083, and bypass is 42 of the 100 available points. Nothing else on the scorecard is big enough to cover a drop that size.

Undetectable.ai made the second-largest gain and passed two rivals, for reasons worth their own section.

Undetectable climbed by padding less

Undetectable gained 10.26 points, and most of that is penalty relief rather than better writing. The methodology penalises any output longer than 1.4 times the input. In August, 24 of its 33 outputs crossed that line; in September, six did. Its penalties fell from 13 points to six, and its outputs shrank from an average of 562 words to 433. Its bypass rate actually slipped slightly.

Walter Writes pulled the same lever the other way. Its outputs over the 1.4x line went from five to twelve. Length inflation alone hit its ten-point ceiling, and two further points for meaning drift took its total penalty to twelve. It got substantially better at evading detectors, gaining 6.68 points of bypass, and still ended up almost exactly where it started, one place lower.

Both moves point the same way, and the per-test data backs it up. Inside a single tool, longer output tends to score better against detectors: the correlation between how much a tool expanded a text and how well that text evaded detection is positive for every tool that pads, running from 0.26 for Super Humanizer to 0.66 for Grammarly, with Walter Writes at 0.31 and Undetectable at 0.33.

Pool every test together and that relationship vanishes, because tools differ from each other far more than their own outputs vary. It only shows up when you compare a tool against itself. It is a correlation rather than a demonstrated cause, but it is consistent enough to be worth stating: padding appears to buy bypass.

Which is why the length penalty exists. An output 40 percent longer than what you pasted in is not a rewrite of your text, and a score that rewarded it would be measuring the wrong thing.

StealthBypass returns fragments, not rewrites

On 21 of StealthBypass’s 32 scored tests, the tool returned a fragment rather than a rewrite, an average of 32 words back from inputs several hundred words long. Its other 11 outputs average 414 words and look like real rewrites.

Penalties took 23 points off its composite, the heaviest on the board, and dropped it to thirteenth at 39.38. Its bypass rate of 0.7864 looks respectable, but a 32-word fragment is easy for a detector to pass and useless to a writer. On the 11 tests where it returned full-length text, its bypass rate was 0.4920.

Penlify writes well and evades nothing

Penlify posted the best readability score on the board, 0.7181, and a bypass rate of 0.0001. Four of the five detectors caught it on every single test: Copyleaks, GPTZero, Originality and Winston never once rated its output as human. ZeroGPT disagreed, passing 12 of its 32 outputs, but per-test bypass is the median of five detectors, so a lone dissenter cannot move the score. Best writing on the board, last place at 36.57.

Humbot went dark mid-cycle

Humbot is not ranked this month. Its first eight tests completed, but slowly: they took between two and nine minutes each, against under a minute for most of the field. After that the site stopped returning output at all, so we skipped the remaining 25 tests and moved on. It looks like an outage on their end. We list a tool as unavailable rather than scoring it as a failure when we cannot get output, so Humbot has no September score.

Business emails still defeat almost everything

Not all writing is equally hard to disguise. Across all 14 ranked tools, business email is by far the hardest prompt: a mean bypass rate of 0.2626 over 42 tests, with only 28.6 percent of outputs scoring 0.5 or better. Cover letters are next at 0.3548. The third-hardest prompt, landing page copy, sits at 0.7208, a different world.

Short, formulaic, conventional business writing appears to leave humanizers very little room to work with. We will add more prompts of this kind next cycle to test that.

About a third of the field’s gain is one detector

The whole field looks better this month, and a large part of that is a single detector. Across the eleven tools tested in both cycles, mean bypass rose from 0.6344 in August to 0.7713 in September. Recompute the same tests using only the other four detectors and the improvement shrinks by roughly a third, from 13.7 points to 9.5. That missing third is attached to Originality alone.

Originality’s own mean human score over those tests went from 0.4659 to 0.8450. Every detector loosened somewhat; the next largest move was Copyleaks, at less than half as far. We cannot tell from our data whether Originality changed or the texts did, since every cycle generates new source material. Anyone reading a big month-over-month improvement should keep that in mind.