Skip to content
October 2026 results: WriteHuman holds #1, StealthGPT jumps to #2. Read the analysis →
HumanizerBench
HumanizerBench

October 2026 AI Humanizer Rankings

As of October 2026, the top-ranked AI humanizer on HumanizerBench is WriteHuman, scoring 74.89 out of 100.

WriteHuman finished first for the third month running, at 74.89. StealthGPT was 0.39 behind at 74.50, up from ninth in September, and Undetectable.ai was third at 71.18. Fourteen tools are ranked, and every one of them completed all 30 tests. Clever AI Humanizer debuts fourth at 69.02, and Humbot is back on the board after September’s outage.

October 2026 leaderboard Full interactive table →
#HumanizerOverallBypassMeaningRead.Pen.
1 -
74.8988.860.055.9
None
2 ▲ 7
74.5088.161.872.6
−3.0Penalties applied
  • Meaning drift−3.0
    Applied 3× max −10.0

    The output's meaning drifted significantly from the original input.

3 ▲ 3
71.1874.679.263.9
−5.0Penalties applied
  • Length inflation−5.0
    Applied 5× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

4 -
69.0280.159.568.6
−4.0Penalties applied
  • Meaning drift−4.0
    Applied 4× max −10.0

    The output's meaning drifted significantly from the original input.

5 ▲ 2
68.8763.475.062.2
−1.0Penalties applied
  • Meaning drift−1.0
    Applied 1× max −10.0

    The output's meaning drifted significantly from the original input.

6 ▼ 2
68.2466.568.656.9
None
7 ▲ 4
67.5861.876.656.6
−1.0Penalties applied
  • Meaning drift−1.0
    Applied 1× max −10.0

    The output's meaning drifted significantly from the original input.

8 ▼ 6
65.9871.371.641.8
−3.0×2Penalties applied
  • Meaning drift−2.0
    Applied 2× max −10.0

    The output's meaning drifted significantly from the original input.

  • Length inflation−1.0
    Applied 1× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

9 ▼ 6
65.1456.878.942.2
None
10 -
63.5372.870.463.5
−9.0Penalties applied
  • Length inflation−9.0
    Applied 9× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

11 ▼ 6
62.7550.077.549.2
None
12 ▼ 4
61.7995.349.944.2
−11.0×2Penalties applied
  • Meaning drift−9.0
    Applied 9× max −10.0

    The output's meaning drifted significantly from the original input.

  • Length inflation−2.0
    Applied 2× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

13 -
61.7471.468.665.0
−10.0Penalties applied
  • Length inflation−10.0
    Applied 19× max −10.0at max

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

14 ▼ 2
61.2938.185.572.2
−3.0Penalties applied
  • Length inflation−3.0
    Applied 3× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

Overall is 0 to 100; bypass, meaning, and readability are 0 to 100 (higher is better). Penalties are points deducted for output-quality issues.

This month is scored differently

October is the first cycle under methodology 1.3.0, and two changes matter for reading the board.

First, Pangram joins the detector panel, so every output is now checked by six detectors instead of five. Second, each output’s bypass score is now the average of all six detectors rather than the median. With six detectors the median only reflects the middle two, so a tool could be caught outright by two of them and still score near 100. The average counts every detector equally.

The source texts also come from newer models this month (GPT-6 Sol, Gemini 3.8 Flash and Claude Sonnet 5.5), and the prompt draw is different, as it is every cycle. Composite scores are not directly comparable with September’s. Eight of the twelve tools ranked in both months scored lower.

The rule change moves the order as well as the scores. Rescored the September way, with the median of the five original detectors, Stealth Writer would be fourth rather than eighth. The next section shows why the rule changed.

Pangram caught almost everything

Pangram’s average human score across all 420 outputs was 0.1687. Nine of the fourteen tools scored exactly zero from Pangram on every one of their 30 outputs.

Only four tools got meaningful passes. Phrasly averaged 0.8556 and passed 26 of its 30 tests, StealthGPT averaged 0.6994 with 22 passes, WriteHuman 0.5032 with 15, and Clever AI Humanizer 0.2647 with 8. Undetectable.ai passed once.

Under the average, a detector that rates every output as AI takes a sixth of the bypass score with it. Bypass is worth 42 of the 100 points, so a clean sweep by Pangram costs a tool 7 points on its own, and caps its bypass rate at 0.8333 even if the other five detectors pass everything. Under a median, one detector dissenting this consistently barely moves the score.

Stealth Writer and HIX Bypass fell on two detectors

Stealth Writer dropped from second to eighth. Against four of the detectors it is close to perfect: Originality 0.999, Winston 0.994, Copyleaks 0.967 and ZeroGPT 0.907. But Pangram gave it zero on every test and GPTZero averaged 0.409. Its median-of-five bypass would have been 0.9922; its actual bypass rate is 0.7127.

HIX Bypass fell from third to ninth the same way, with Pangram at zero and GPTZero at 0.114. SupWriter, fifth on its debut last month, fell to eleventh with GPTZero at 0.089. In both cases two detectors catching the tool used to be absorbed by the median, and now it is not.

GPTZero moved in two directions this month. The tools at the top of the board scored higher against it than in September (WriteHuman went from 0.718 to 0.995, Phrasly from 0.814 to 0.986), while HIX Bypass, SupWriter and Grammarly all fell sharply. Every cycle uses new source texts, so we cannot tell from our data how much of that is the detector and how much is the texts.

StealthGPT came within 0.39 points

StealthGPT climbed seven places, from ninth to second. Its bypass rate rose from 0.7961 to 0.8812, and it was one of only two tools to pass Pangram more often than not. It also posted the best readability on the board, 0.7260.

What kept it out of first was meaning drift. Three of its outputs fell below the meaning-preservation floor, at a point each. Without those three points it would have finished first at 77.50.

Phrasly evades best and keeps the least

Phrasly posted the highest bypass rate on the board, 0.9534, and was the tool Pangram passed most often. It also kept the least of what it was given. Its meaning-preservation score fell from 0.6607 in September to 0.4994, the lowest of the fourteen, and nine of its 30 outputs drifted far enough from the source to draw a penalty. With two more for length inflation, penalties took 11 points off its composite, and it finished twelfth at 61.79.

That is the trade the composite is built to price. A rewrite that evades every detector but no longer says what you wrote is not much use to a writer.

Clever AI Humanizer debuts fourth on a free plan

Clever AI Humanizer was tested on its free tier and finished fourth at 69.02, the best result on the board from a free plan. Its bypass rate was 0.8007, fourth-highest, and it was the fourth tool to get any real traction against Pangram. Four meaning-drift penalties cost it 4 points.

Humbot is back, and it pads

Humbot completed all 30 tests this month after going dark partway through September. It finished thirteenth at 61.74, almost entirely because of length. The methodology penalises any output longer than 1.4 times its input, and 19 of Humbot’s 30 outputs crossed that line. Its outputs averaged 527 words from inputs averaging 360. The length penalty is capped at 10 points, and it hit the cap. Without it, Humbot would have placed third.

Walter Writes had the next-largest length penalty, with nine outputs over the line, and Undetectable.ai had five.

Cover letters are still the hardest

Cover letters were the hardest prompt this month, with a mean bypass of 0.6007 across all fourteen tools, and the only prompt where fewer than three quarters of outputs scored 0.5 or better. They were also the second-hardest in September. News articles were the easiest at 0.7874.