AI Humanizer Leaderboard
In August 2026, the top-ranked AI humanizer on HumanizerBench was WriteHuman, scoring 76.69 out of 100. Since July 2026, AI Humanize io climbed 4 places, the biggest gain. Undetectable.ai fell 6 places, the biggest drop.
12 humanizers · 33 prompts each · 5 detectors · Last tested · Methodology v1.2.0
| Rank Position in this cycle by overall score. Learn more → | Humanizer | OverallOverall Weighted blend of bypass rate (42%), meaning preservation (32%), readability (16%), and consistency (10%). Penalties may reduce this score. Learn more → ↓ | BypassBypass Rate Fraction of detector tests where the humanized output was classified as human, across all 5 detectors. Click a score to read the raw outputs and verdicts behind it. Learn more → | Meaning Semantic similarity (embedding cosine) between input and humanized output. Higher = output preserves the input's meaning. Learn more → | Readability Writing quality of the output: clarity, fluency and naturalness, rated by a language model. Higher = better. Learn more → | Penalties Total points deducted from the overall score when output-quality issues were detected. Hover a row's chip to see why. Learn more → | Trend Leaderboard position across recent cycles. The line rises when the tool moves up the ranking, independent of absolute score. Learn more → | Video |
|---|---|---|---|---|---|---|---|---|
| 1
-
| 76.69 | 89.1 | 71.7 | 57.8 | −1.0Penalties applied
| | ||
| 2
▲ 1 | 71.40 | 83.5 | 67.1 | 52.3 | None | | ||
| 3
▲ 1 | 69.83 | 79.7 | 72.0 | 43.7 | −1.0Penalties applied
| | ||
| 4
▲ 2 | 67.93 | 75.2 | 76.7 | 38.3 | −1.0Penalties applied
| | ||
| 5
-
| 66.95 | 71.7 | 75.4 | 43.3 | −1.0Penalties applied
| | ||
| 6
▲ 4 | 64.57 | 70.8 | 64.4 | 54.6 | −2.0Penalties applied
| | ||
| 7
▲ 2 | 63.30 | 76.4 | 62.9 | 67.2 | −8.0×2Penalties applied
| | ||
| 8
▼ 6 | 61.49 | 86.0 | 72.1 | 55.2 | −13.0×2Penalties applied
| | ||
| 9
▼ 2 | 60.83 | 68.0 | 68.5 | 65.3 | −8.0×2Penalties applied
| | ||
| 10
▲ 1 | 55.86 | 41.5 | 73.1 | 58.2 | −1.0Penalties applied
| | ||
| 11
▲ 1 | 52.91 | 0.0 | 94.3 | 79.5 | None | | ||
| 12
▼ 4 | 45.55 | 27.5 | 62.9 | 70.9 | −5.0×3Penalties applied
| |
Swipe the table sideways to see every column.
Gold, silver and bronze mark the top three scores in each column. Tied scores share a color.
Showing scores in the Overall column ( prompts per humanizer this cycle). A test only earns bypass credit when its output is a real rewrite; that is blended with the category's own meaning and readability. Small categories swing; that's the point. How category scores work
Session recordings
Screen captures of the runs behind the August 2026 board: 13 recordings, one per tool, 2h 33m of testing, 429 runs. Pick a tool to watch it being tested, then jump to any run. Nothing downloads until you press play. Hosted links may expire over time; the published JSON stays the record.
WriteHuman
11:56 · 33 runs
Humanize AI Pro
10:58 · 33 runs
Stealth Writer
9:00 · 33 runs
HIX Bypass
9:21 · 33 runs
Humbot
12:34 · 33 runs
AI Humanize io
12:57 · 33 runs
Phrasly
17:08 · 33 runs
Undetectable.ai
14:46 · 33 runs
Walter Writes
15:48 · 33 runs
Super Humanizer
9:42 · 33 runs
Grammarly
10:39 · 33 runs
StealthGPT
16:24 · 33 runs
NoteGPT
1:46 · 33 runs
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Recordings
How scores are computed
Full methodology- Bypass
- Share of detector tests where the output was classified as human.
- Meaning
- How well the output preserves the original meaning.
- Readability
- How clear, fluent, and natural the writing reads, rated by a language model.
- Consistency
- How evenly the humanizer performs across writing categories.
Penalties
Taken off the total for each flagged test. Each rule stops at −10.0, −50.0 in total, and a score cannot go below 0.
- Identical to input −2.0
- The tool returned the input mostly unchanged. No real humanization happened.
- Refusal −1.0
- The tool refused to humanize the input, often due to a content-policy block.
- Meaning drift −1.0
- The output's meaning drifted significantly from the original input.
- Length inflation −1.0
- The output ran much longer than the input, a common trick that pads text to dilute the AI signal.
- Length deflation −1.0
- The output came back much shorter than the input. The rewrite dropped content instead of paraphrasing it.