Skip to content
October 2026 results: WriteHuman holds #1, StealthGPT jumps to #2. Read the analysis →
HumanizerBench

← Latest leaderboard

AI Humanizer Leaderboard

In August 2026, the top-ranked AI humanizer on HumanizerBench was WriteHuman, scoring 76.69 out of 100. Since July 2026, AI Humanize io climbed 4 places, the biggest gain. Undetectable.ai fell 6 places, the biggest drop.

12 humanizers · 33 prompts each · 5 detectors · Last tested · Methodology v1.2.0

Rank

Position in this cycle by overall score.

Learn more →
Humanizer Overall

Weighted blend of bypass rate (42%), meaning preservation (32%), readability (16%), and consistency (10%). Penalties may reduce this score.

Learn more →
↓
Bypass

Fraction of detector tests where the humanized output was classified as human, across all 5 detectors. Click a score to read the raw outputs and verdicts behind it.

Learn more →
Meaning

Semantic similarity (embedding cosine) between input and humanized output. Higher = output preserves the input's meaning.

Learn more →
Readability

Writing quality of the output: clarity, fluency and naturalness, rated by a language model. Higher = better.

Learn more →
Penalties

Total points deducted from the overall score when output-quality issues were detected. Hover a row's chip to see why.

Learn more →
Trend

Leaderboard position across recent cycles. The line rises when the tool moves up the ranking, independent of absolute score.

Learn more →
Video
1 -
76.69 89.1 71.7 57.8 −1.0Penalties applied
  • Length inflation−1.0
    Applied 1× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

2 ▲ 1
71.40 83.5 67.1 52.3 None
3 ▲ 1
69.83 79.7 72.0 43.7 −1.0Penalties applied
  • Length inflation−1.0
    Applied 1× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

4 ▲ 2
67.93 75.2 76.7 38.3 −1.0Penalties applied
  • Length inflation−1.0
    Applied 1× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

5 -
66.95 71.7 75.4 43.3 −1.0Penalties applied
  • Length inflation−1.0
    Applied 1× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

6 ▲ 4
64.57 70.8 64.4 54.6 −2.0Penalties applied
  • Meaning drift−2.0
    Applied 2× max −10.0

    The output's meaning drifted significantly from the original input.

7 ▲ 2
63.30 76.4 62.9 67.2 −8.0×2Penalties applied
  • Meaning drift−4.0
    Applied 4× max −10.0

    The output's meaning drifted significantly from the original input.

  • Length inflation−4.0
    Applied 4× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

8 ▼ 6
61.49 86.0 72.1 55.2 −13.0×2Penalties applied
  • Meaning drift−3.0
    Applied 3× max −10.0

    The output's meaning drifted significantly from the original input.

  • Length inflation−10.0
    Applied 24× max −10.0at max

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

9 ▼ 2
60.83 68.0 68.5 65.3 −8.0×2Penalties applied
  • Meaning drift−3.0
    Applied 3× max −10.0

    The output's meaning drifted significantly from the original input.

  • Length inflation−5.0
    Applied 5× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

10 ▲ 1
55.86 41.5 73.1 58.2 −1.0Penalties applied
  • Meaning drift−1.0
    Applied 1× max −10.0

    The output's meaning drifted significantly from the original input.

11 ▲ 1
52.91 0.0 94.3 79.5 None
12 ▼ 4
45.55 27.5 62.9 70.9 −5.0×3Penalties applied
  • Meaning drift−1.0
    Applied 1× max −10.0

    The output's meaning drifted significantly from the original input.

  • Length inflation−2.0
    Applied 2× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

  • Length deflation−2.0
    Applied 2× max −10.0

    The output came back much shorter than the input. The rewrite dropped content instead of paraphrasing it.

Swipe the table sideways to see every column.

Gold, silver and bronze mark the top three scores in each column. Tied scores share a color.

Screen captures of the runs behind the August 2026 board: 13 recordings, one per tool, 2h 33m of testing, 429 runs. Pick a tool to watch it being tested, then jump to any run. Nothing downloads until you press play. Hosted links may expire over time; the published JSON stays the record.

WriteHuman

11:56 · 33 runs

0:00 of 11:56

WriteHuman

2026-08-01 21:22 UTC · 11:56 · 33 runs

Visit WriteHuman

Runs in this recording

Click a time to jump.

Recordings

How scores are computed

Full methodology
Bypass
Share of detector tests where the output was classified as human.
Meaning
How well the output preserves the original meaning.
Readability
How clear, fluent, and natural the writing reads, rated by a language model.
Consistency
How evenly the humanizer performs across writing categories.

Penalties

Taken off the total for each flagged test. Each rule stops at −10.0, −50.0 in total, and a score cannot go below 0.

Identical to input −2.0
Refusal −1.0
Meaning drift −1.0
Length inflation −1.0
Length deflation −1.0