Skip to content
The September 2026 results are in: WriteHuman holds #1, and newcomer SupWriter debuts at #5.Read the full analysis →
HumanizerBench

← Latest leaderboard

AI Humanizer Leaderboard

Cycle: September 2026 • 33 prompts per humanizer

Last tested
Prompts
33
Methodology
v1.2.0
14 humanizers
Rank

Position in this cycle by overall score.

Learn more →
Humanizer Overall

Weighted blend of bypass rate (42%), meaning preservation (32%), readability (16%), and consistency (10%). Penalties may reduce this score.

Learn more →
Bypass Rate

Fraction of detector tests where the humanized output was classified as human, across all 5 detectors. Click a score to read the raw outputs and verdicts behind it.

Learn more →
Meaning

Semantic similarity (embedding cosine) between input and humanized output. Higher = output preserves the input's meaning.

Learn more →
Readability

Writing quality of the output — clarity, fluency, and naturalness, rated by a language model. Higher = better.

Learn more →
Penalties

Total points deducted from the overall score when output-quality issues were detected. Hover a row's chip to see why.

Learn more →
Trend

Leaderboard position across recent cycles. The line rises when the tool moves up the ranking, independent of absolute score.

Learn more →
Last Tested
1 -
78.29 90.1 72.3 56.9 None
2 ▲ 1
75.39 86.4 76.3 47.2 −1.0Penalties applied
  • Meaning drift−1.0
    Applied 1× max −10.0

    The output's meaning drifted significantly from the original input.

3 ▲ 1
73.97 86.3 76.9 34.9 None
4 ▲ 2
72.94 81.5 73.2 58.4 −1.0Penalties applied
  • Length inflation−1.0
    Applied 1× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

5 -
72.80 80.3 79.0 45.7 None
6 ▲ 2
71.75 82.1 82.1 61.7 −6.0Penalties applied
  • Length inflation−6.0
    Applied 6× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

7 ▼ 5
70.07 70.8 74.7 61.8 None
8 ▼ 1
67.58 82.1 66.1 65.4 −6.0×2Penalties applied
  • Meaning drift−1.0
    Applied 1× max −10.0

    The output's meaning drifted significantly from the original input.

  • Length inflation−5.0
    Applied 5× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

9 ▲ 3
65.10 79.6 59.4 70.9 −5.0Penalties applied
  • Meaning drift−5.0
    Applied 5× max −10.0

    The output's meaning drifted significantly from the original input.

10 ▼ 1
61.70 83.9 66.6 60.4 −12.0×2Penalties applied
  • Meaning drift−2.0
    Applied 2× max −10.0

    The output's meaning drifted significantly from the original input.

  • Length inflation−10.0
    Applied 12× max −10.0at max

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

11 ▼ 1
60.97 64.4 74.2 52.1 −5.0Penalties applied
  • Length inflation−5.0
    Applied 5× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

12 ▼ 1
60.35 41.1 86.4 66.1 −3.0Penalties applied
  • Length inflation−3.0
    Applied 3× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

13 -
39.38 78.6 39.5 54.1 −23.0×4Penalties applied
  • Identical to input−2.0
    Applied 1× max −10.0

    The tool returned the input mostly unchanged. No real humanization happened.

  • Meaning drift−10.0
    Applied 17× max −10.0at max

    The output's meaning drifted significantly from the original input.

  • Length inflation−1.0
    Applied 1× max −10.0

    The output ran much longer than the input, a common trick that pads text to dilute the AI signal.

  • Length deflation−10.0
    Applied 21× max −10.0at max

    The output came back much shorter than the input. The rewrite dropped content instead of paraphrasing it.

14 -
36.57 0.0 62.7 71.8 −5.0Penalties applied
  • Meaning drift−5.0
    Applied 5× max −10.0

    The output's meaning drifted significantly from the original input.

Screen captures from the benchmark runs, kept for auditability. Hosted links may expire over time; the JSON in this repo is the source of truth.

How scores are computed

See the full methodology →

Overall score formula

  • Bypass 42% Share of detector tests where the output was classified as human.
  • Meaning 32% How well the output preserves the original meaning.
  • Readability 16% How clear, fluent, and natural the writing reads, rated by a language model.
  • Consistency 10% How evenly the humanizer performs across writing categories.

Penalties (right) are deducted from the overall score. Final score is clamped at 0.

Possible penalties

Code Penalty Max
Identical to input
The tool returned the input mostly unchanged. No real humanization happened.
−2.0 −10.0
Refusal
The tool refused to humanize the input, often due to a content-policy block.
−1.0 −10.0
Meaning drift
The output's meaning drifted significantly from the original input.
−1.0 −10.0
Length inflation
The output ran much longer than the input, a common trick that pads text to dilute the AI signal.
−1.0 −10.0

Maximum possible total penalty: −50.0 (off 100).