Skip to content
The September 2026 results are in: WriteHuman holds #1, and newcomer SupWriter debuts at #5.Read the full analysis →
HumanizerBench

July 2026 prompt evidence

Each of these 33 prompts was written by a frontier model, then handed to all 13 ranked humanizers, and every rewrite was submitted to 5 commercial AI detectors. Every run was screen-recorded. Open a prompt to read the input, every tool's rewrite, each detector's verdict, and to watch the recording of that run.

33
Prompts
13
Tools ranked
429
Humanizer outputs
2143
Detector verdicts
How to read these scores

Each detector returns a human-likelihood on a common 0 to 1 scale, where 1 means it judged the text human-written and 0 means it flagged it as AI. On these pages a verdict counts as passed when that score is at least 0.50, the midpoint of the detector's own scale. That threshold exists only to draw the chips: the bypass rate on the leaderboard is the mean of each test's median score across the 5 detectors, a continuous number, so a tool's pass count and its bypass rate will not be the same figure.

A few runs come back with fewer than 5 verdicts because a detector did not return a usable score for them. A run still counts as passing at least 4 detectors only if that many detectors actually scored it and passed it. A partially scored run is counted as caught if any detector that did score it flagged the output, and is otherwise listed separately as partly scored.

Meaning is the input↔output embedding cosine and readability is a language-model writing-quality rating, both published per test in tests.json. Words is the output's length as a multiple of the input's; the scoring code penalizes ratios above 1.40 or below 0.60. Full definitions live in the methodology.

Prompt Category Written by Input words Passed Failed Open
01 Argumentative Essay Academic Essay GPT-5.5 434 7 7
02 Listicle Blog Blog Post GPT-5.5 398 9 11
03 Cover Letter Application Essay Gemini 3.5 Flash 339 0 13
04 Business Email Business Email Gemini 3.5 Flash 266 1 13
05 Howto Blog Blog Post Gemini 3.5 Flash 384 5 12
06 Lit Review Academic Essay Claude Sonnet 5 384 3 11
07 Product Desc Marketing Copy GPT-5.5 195 5 10
08 Landing Copy Marketing Copy GPT-5.5 273 4 10
09 Personal Statement Application Essay GPT-5.5 377 7 11
10 Business Email Business Email Claude Sonnet 5 276 3 11
11 Product Desc Marketing Copy Claude Sonnet 5 196 9 9
12 Listicle Blog Blog Post Gemini 3.5 Flash 405 8 8
13 Product Desc Marketing Copy Gemini 3.5 Flash 170 8 11
14 Howto Blog Blog Post Claude Sonnet 5 431 4 12
15 Cover Letter Application Essay GPT-5.5 377 2 13
16 Argumentative Essay Academic Essay Gemini 3.5 Flash 387 3 13
17 Discussion Post Discussion Board Gemini 3.5 Flash 255 7 7
18 Lit Review Academic Essay Gemini 3.5 Flash 343 3 12
19 Cover Letter Application Essay Claude Sonnet 5 407 1 12
20 Business Email Business Email GPT-5.5 255 0 13
21 Discussion Post Discussion Board GPT-5.5 255 6 9
22 Personal Statement Application Essay Claude Sonnet 5 369 10 6
23 Landing Copy Marketing Copy Gemini 3.5 Flash 279 4 11
24 News Article News Article Gemini 3.5 Flash 391 7 10
25 Listicle Blog Blog Post Claude Sonnet 5 397 10 5
26 Landing Copy Marketing Copy Claude Sonnet 5 294 8 12
27 Discussion Post Discussion Board Claude Sonnet 5 254 7 9
28 Argumentative Essay Academic Essay Claude Sonnet 5 417 9 8
29 Lit Review Academic Essay GPT-5.5 432 5 9
30 Howto Blog Blog Post GPT-5.5 423 7 8
31 News Article News Article GPT-5.5 463 10 8
32 Personal Statement Application Essay Gemini 3.5 Flash 325 8 7
33 News Article News Article Claude Sonnet 5 424 7 10

"Passed" counts the ranked tools whose rewrite of that prompt at least 4 of the 5 detectors scored as human-written (0.50 or above); "Failed" counts those flagged by at least one detector that scored them; "Hardest" is the single detector that flagged the most rewrites of that prompt.

Recordings

Every tool was run in a recorded sitting. Each recording opens here from the start; the outputs pages jump to the exact moment a prompt was run.

  • WriteHuman 10:15 · 33 prompts · Jul 1, 2026
  • Undetectable.ai 11:36 · 33 prompts · Jul 1, 2026
  • Humanize AI Pro 7:28 · 33 prompts · Jul 1, 2026
  • Stealth Writer 6:33 · 32 prompts · Jul 1, 2026
  • Humbot 15:02 · 33 prompts · Jul 1, 2026
  • HIX Bypass 10:33 · 32 prompts · Jul 1, 2026
  • Walter Writes 18:21 · 33 prompts · Jul 1, 2026
  • StealthGPT 14:40 · 33 prompts · Jul 1, 2026
  • Phrasly 13:11 · 33 prompts · Jul 1, 2026
  • AI Humanize io 17:06 · 33 prompts · Jul 1, 2026
  • Super Humanizer 6:22 · 33 prompts · Jul 1, 2026
  • Grammarly 6:46 · 33 prompts · Jul 1, 2026
  • NoteGPT 12:13 · 33 prompts · Jul 1, 2026

Session recording