Skip to content
October 2026 results: WriteHuman holds #1, StealthGPT jumps to #2. Read the analysis →
HumanizerBench

← Latest leaderboard

AI Humanizer Leaderboard

As of October 2026, the top-ranked AI humanizer on HumanizerBench is WriteHuman, scoring 74.89 out of 100. Since September 2026, StealthGPT climbed 7 places, the biggest gain. Stealth Writer, HIX Bypass and SupWriter each fell 6 places, the biggest drops. Clever AI Humanizer was ranked for the first time.

14 humanizers · 30 prompts each · 6 detectors · Last tested · Methodology v1.3.0

Rank

Position in this cycle by overall score.

Learn more →
Humanizer Overall

Weighted blend of bypass rate (42%), meaning preservation (32%), readability (16%), and consistency (10%). Penalties may reduce this score.

Learn more →
↓
Bypass

Fraction of detector tests where the humanized output was classified as human, across all 6 detectors. Click a score to read the raw outputs and verdicts behind it.

Learn more →
Meaning

Semantic similarity (embedding cosine) between input and humanized output. Higher = output preserves the input's meaning.

Learn more →
Readability

Writing quality of the output: clarity, fluency and naturalness, rated by a language model. Higher = better.

Learn more →
Penalties

Total points deducted from the overall score when output-quality issues were detected. Hover a row's chip to see why.

Learn more →
Trend

Leaderboard position across recent cycles. The line rises when the tool moves up the ranking, independent of absolute score.

Learn more →
Video
1 -
74.89 88.8 60.0 55.9 None
2 ▲ 7
74.50 88.1 61.8 72.6 −3.0Penalties applied
  • Meaning drift−3.0
    Applied 3× max −10.0

    The output's meaning drifted significantly from the original input.

3 ▲ 3
71.18 74.6 79.2 63.9 −5.0Penalties applied
  • Length inflation−5.0
    Applied 5× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

4 -
69.02 80.1 59.5 68.6 −4.0Penalties applied
  • Meaning drift−4.0
    Applied 4× max −10.0

    The output's meaning drifted significantly from the original input.

5 ▲ 2
68.87 63.4 75.0 62.2 −1.0Penalties applied
  • Meaning drift−1.0
    Applied 1× max −10.0

    The output's meaning drifted significantly from the original input.

6 ▼ 2
68.24 66.5 68.6 56.9 None
7 ▲ 4
67.58 61.8 76.6 56.6 −1.0Penalties applied
  • Meaning drift−1.0
    Applied 1× max −10.0

    The output's meaning drifted significantly from the original input.

8 ▼ 6
65.98 71.3 71.6 41.8 −3.0×2Penalties applied
  • Meaning drift−2.0
    Applied 2× max −10.0

    The output's meaning drifted significantly from the original input.

  • Length inflation−1.0
    Applied 1× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

9 ▼ 6
65.14 56.8 78.9 42.2 None
10 -
63.53 72.8 70.4 63.5 −9.0Penalties applied
  • Length inflation−9.0
    Applied 9× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

11 ▼ 6
62.75 50.0 77.5 49.2 None
12 ▼ 4
61.79 95.3 49.9 44.2 −11.0×2Penalties applied
  • Meaning drift−9.0
    Applied 9× max −10.0

    The output's meaning drifted significantly from the original input.

  • Length inflation−2.0
    Applied 2× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

13 -
61.74 71.4 68.6 65.0 −10.0Penalties applied
  • Length inflation−10.0
    Applied 19× max −10.0at max

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

14 ▼ 2
61.29 38.1 85.5 72.2 −3.0Penalties applied
  • Length inflation−3.0
    Applied 3× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

Swipe the table sideways to see every column.

Gold, silver and bronze mark the top three scores in each column. Tied scores share a color.

Screen captures of the runs behind the October 2026 board: 14 recordings, one per tool, 3h 20m of testing, 420 runs. Pick a tool to watch it being tested, then jump to any run. Nothing downloads until you press play. Hosted links may expire over time; the published JSON stays the record.

WriteHuman

11:46 · 30 runs

0:00 of 11:46

WriteHuman

2026-10-01 23:03 UTC · 11:46 · 30 runs

Visit WriteHuman

Runs in this recording

Click a time to jump.

Recordings

How scores are computed

Full methodology
Bypass
Share of detector tests where the output was classified as human.
Meaning
How well the output preserves the original meaning.
Readability
How clear, fluent, and natural the writing reads, rated by a language model.
Consistency
How evenly the humanizer performs across writing categories.

Penalties

Taken off the total for each flagged test. Each rule stops at −10.0, −50.0 in total, and a score cannot go below 0.

Identical to input −2.0
Refusal −1.0
Meaning drift −1.0
Length inflation −1.0
Length deflation −1.0