AI Humanizer Leaderboard
As of October 2026, the top-ranked AI humanizer on HumanizerBench is WriteHuman, scoring 74.89 out of 100. Since September 2026, StealthGPT climbed 7 places, the biggest gain. Stealth Writer, HIX Bypass and SupWriter each fell 6 places, the biggest drops. Clever AI Humanizer was ranked for the first time.
14 humanizers · 30 prompts each · 6 detectors · Last tested · Methodology v1.3.0
| Rank Position in this cycle by overall score. Learn more → | Humanizer | OverallOverall Weighted blend of bypass rate (42%), meaning preservation (32%), readability (16%), and consistency (10%). Penalties may reduce this score. Learn more → ↓ | BypassBypass Rate Fraction of detector tests where the humanized output was classified as human, across all 6 detectors. Click a score to read the raw outputs and verdicts behind it. Learn more → | Meaning Semantic similarity (embedding cosine) between input and humanized output. Higher = output preserves the input's meaning. Learn more → | Readability Writing quality of the output: clarity, fluency and naturalness, rated by a language model. Higher = better. Learn more → | Penalties Total points deducted from the overall score when output-quality issues were detected. Hover a row's chip to see why. Learn more → | Trend Leaderboard position across recent cycles. The line rises when the tool moves up the ranking, independent of absolute score. Learn more → | Video |
|---|---|---|---|---|---|---|---|---|
| 1
-
| 74.89 | 88.8 | 60.0 | 55.9 | None | | ||
| 2
▲ 7 | 74.50 | 88.1 | 61.8 | 72.6 | −3.0Penalties applied
| | ||
| 3
▲ 3 | 71.18 | 74.6 | 79.2 | 63.9 | −5.0Penalties applied
| | ||
| 4
-
| 69.02 | 80.1 | 59.5 | 68.6 | −4.0Penalties applied
| | ||
| 5
▲ 2 | 68.87 | 63.4 | 75.0 | 62.2 | −1.0Penalties applied
| | ||
| 6
▼ 2 | 68.24 | 66.5 | 68.6 | 56.9 | None | | ||
| 7
▲ 4 | 67.58 | 61.8 | 76.6 | 56.6 | −1.0Penalties applied
| | ||
| 8
▼ 6 | 65.98 | 71.3 | 71.6 | 41.8 | −3.0×2Penalties applied
| | ||
| 9
▼ 6 | 65.14 | 56.8 | 78.9 | 42.2 | None | | ||
| 10
-
| 63.53 | 72.8 | 70.4 | 63.5 | −9.0Penalties applied
| | ||
| 11
▼ 6 | 62.75 | 50.0 | 77.5 | 49.2 | None | | ||
| 12
▼ 4 | 61.79 | 95.3 | 49.9 | 44.2 | −11.0×2Penalties applied
| | ||
| 13
-
| 61.74 | 71.4 | 68.6 | 65.0 | −10.0Penalties applied
| | ||
| 14
▼ 2 | 61.29 | 38.1 | 85.5 | 72.2 | −3.0Penalties applied
| |
Swipe the table sideways to see every column.
Gold, silver and bronze mark the top three scores in each column. Tied scores share a color.
Showing scores in the Overall column ( prompts per humanizer this cycle). A test only earns bypass credit when its output is a real rewrite; that is blended with the category's own meaning and readability. Small categories swing; that's the point. How category scores work
Session recordings
Screen captures of the runs behind the October 2026 board: 14 recordings, one per tool, 3h 20m of testing, 420 runs. Pick a tool to watch it being tested, then jump to any run. Nothing downloads until you press play. Hosted links may expire over time; the published JSON stays the record.
WriteHuman
11:46 · 30 runs
StealthGPT
38:33 · 30 runs
Undetectable.ai
10:13 · 30 runs
Clever AI Humanizer
18:04 · 30 runs
Humanize AI Pro
10:55 · 30 runs
AI Humanize io
9:42 · 30 runs
Super Humanizer
4:03 · 30 runs
Stealth Writer
6:12 · 30 runs
HIX Bypass
11:10 · 30 runs
Walter Writes
14:38 · 30 runs
SupWriter
10:12 · 30 runs
Phrasly
12:41 · 30 runs
Humbot
37:09 · 30 runs
Grammarly
4:56 · 30 runs
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Recordings
How scores are computed
Full methodology- Bypass
- Share of detector tests where the output was classified as human.
- Meaning
- How well the output preserves the original meaning.
- Readability
- How clear, fluent, and natural the writing reads, rated by a language model.
- Consistency
- How evenly the humanizer performs across writing categories.
Penalties
Taken off the total for each flagged test. Each rule stops at −10.0, −50.0 in total, and a score cannot go below 0.
- Identical to input −2.0
- The tool returned the input mostly unchanged. No real humanization happened.
- Refusal −1.0
- The tool refused to humanize the input, often due to a content-policy block.
- Meaning drift −1.0
- The output's meaning drifted significantly from the original input.
- Length inflation −1.0
- The output ran much longer than the input, a common trick that pads text to dilute the AI signal.
- Length deflation −1.0
- The output came back much shorter than the input. The rewrite dropped content instead of paraphrasing it.