AI Humanizer Leaderboard
In June 2026, the top-ranked AI humanizer on HumanizerBench was Stealth Writer, scoring 84.47 out of 100.
13 humanizers · 30 prompts each · 5 detectors · Last tested · Methodology v1.0.0
| Rank Position in this cycle by overall score. Learn more → | Humanizer | OverallOverall Weighted blend of bypass rate (42%), meaning preservation (32%), readability (16%), and consistency (10%). Penalties may reduce this score. Learn more → ↓ | BypassBypass Rate Fraction of detector tests where the humanized output was classified as human, across all 5 detectors. Click a score to read the raw outputs and verdicts behind it. Learn more → | Meaning Semantic similarity (embedding cosine) between input and humanized output. Higher = output preserves the input's meaning. Learn more → | Readability Writing quality of the output: clarity, fluency and naturalness, rated by a language model. Higher = better. Learn more → | Penalties Total points deducted from the overall score when output-quality issues were detected. Hover a row's chip to see why. Learn more → | Trend Leaderboard position across recent cycles. The line rises when the tool moves up the ranking, independent of absolute score. Learn more → | Video |
|---|---|---|---|---|---|---|---|---|
| 1
-
| 84.47 | 89.5 | 70.7 | 100.0 | −1.0Penalties applied
| | ||
| 2
-
| 83.59 | 84.5 | 71.8 | 100.0 | None | | ||
| 3
-
| 83.52 | 85.2 | 76.8 | 93.6 | −1.0Penalties applied
| | ||
| 4
-
| 80.28 | 76.9 | 77.6 | 94.3 | None | | ||
| 5
-
| 80.09 | 94.1 | 72.2 | 100.0 | −8.0Penalties applied
| | ||
| 6
-
| 72.95 | 79.2 | 71.6 | 91.0 | −6.0Penalties applied
| | ||
| 7
-
| 71.47 | 82.9 | 67.7 | 100.0 | −9.0×3Penalties applied
| | ||
| 8
-
| 71.10 | 76.9 | 64.1 | 99.2 | −6.0×2Penalties applied
| | ||
| 9
-
| 68.05 | 54.6 | 71.1 | 97.5 | None | | ||
| 10
-
| 67.24 | 73.0 | 70.9 | 85.7 | −8.0Penalties applied
| | ||
| 11
-
| 64.75 | 47.5 | 74.1 | 96.7 | −1.0Penalties applied
| | ||
| 12
-
| 56.50 | 0.0 | 95.3 | 100.0 | None | | ||
| 13
-
| 45.34 | 0.0 | 66.7 | 100.0 | −2.0Penalties applied
| |
Swipe the table sideways to see every column.
Gold, silver and bronze mark the top three scores in each column. Tied scores share a color.
Showing scores in the Overall column ( prompts per humanizer this cycle). A test only earns bypass credit when its output is a real rewrite; that is blended with the category's own meaning and readability. Small categories swing; that's the point. How category scores work
Session recordings
Screen captures of the runs behind the June 2026 board: 14 recordings, one per tool, 2h 37m of testing, 420 runs. Pick a tool to watch it being tested, then jump to any run. Nothing downloads until you press play. Hosted links may expire over time; the published JSON stays the record.
Stealth Writer
6:32 · 30 runs
WriteHuman
7:25 · 30 runs
HIX Bypass
11:26 · 30 runs
Humbot
14:51 · 30 runs
Undetectable.ai
8:49 · 30 runs
AI Humanize io
12:24 · 30 runs
StealthGPT
12:59 · 30 runs
Phrasly
13:22 · 30 runs
Humanize AI Pro
8:49 · 30 runs
Walter Writes
16:17 · 30 runs
Super Humanizer
5:13 · 30 runs
Grammarly
9:37 · 30 runs
NoteGPT
12:41 · 30 runs
BypassGPT
16:46 · 30 runs
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
Runs in this recording
Click a time to jump.
BypassGPT
2026-06-02 00:52 UTC · 16:46 · 30 runs
Runs in this recording
Click a time to jump.
Recordings
How scores are computed
Full methodology- Bypass
- Share of detector tests where the output was classified as human.
- Meaning
- How well the output preserves the original meaning.
- Readability
- How clear, fluent, and natural the writing reads, rated by a language model.
- Consistency
- How evenly the humanizer performs across writing categories.
Penalties
Taken off the total for each flagged test. Each rule stops at −10.0, −50.0 in total, and a score cannot go below 0.
- Identical to input −2.0
- The tool returned the input mostly unchanged. No real humanization happened.
- Refusal −1.0
- The tool refused to humanize the input, often due to a content-policy block.
- Meaning drift −1.0
- The output's meaning drifted significantly from the original input.
- Length inflation −1.0
- The output ran much longer than the input, a common trick that pads text to dilute the AI signal.
- Length deflation −1.0
- The output came back much shorter than the input. The rewrite dropped content instead of paraphrasing it.