Skip to content
The September 2026 results are in: WriteHuman holds #1, and newcomer SupWriter debuts at #5.Read the full analysis →
HumanizerBench

September 2026 prompt evidence

Each of these 33 prompts was written by a frontier model, then handed to all 14 ranked humanizers, and every rewrite was submitted to 5 commercial AI detectors. Every run was screen-recorded. Open a prompt to read the input, every tool's rewrite, each detector's verdict, and to watch the recording of that run.

33
Prompts
14
Tools ranked
462
Humanizer outputs
2268
Detector verdicts
How to read these scores

Each detector returns a human-likelihood on a common 0 to 1 scale, where 1 means it judged the text human-written and 0 means it flagged it as AI. On these pages a verdict counts as passed when that score is at least 0.50, the midpoint of the detector's own scale. That threshold exists only to draw the chips: the bypass rate on the leaderboard is the mean of each test's median score across the 5 detectors, a continuous number, so a tool's pass count and its bypass rate will not be the same figure.

A few runs come back with fewer than 5 verdicts because a detector did not return a usable score for them. A run still counts as passing at least 4 detectors only if that many detectors actually scored it and passed it. A partially scored run is counted as caught if any detector that did score it flagged the output, and is otherwise listed separately as partly scored.

Meaning is the input↔output embedding cosine and readability is a language-model writing-quality rating, both published per test in tests.json. Words is the output's length as a multiple of the input's; the scoring code penalizes ratios above 1.40 or below 0.60. Full definitions live in the methodology.

Prompt Category Written by Input words Passed Failed Open
01 Argumentative Essay Academic Essay GPT-5.5 434 10 9
02 Discussion Post Discussion Board Gemini 3.5 Flash 253 11 10
03 Personal Statement Application Essay GPT-5.5 364 11 9
04 Discussion Post Discussion Board Claude Sonnet 5 246 11 8
05 News Article News Article GPT-5.5 452 11 6
06 Product Desc Marketing Copy Claude Sonnet 5 192 9 10
07 Cover Letter Application Essay Claude Sonnet 5 405 1 14
08 Personal Statement Application Essay Claude Sonnet 5 361 10 8
09 Business Email Business Email Gemini 3.5 Flash 269 1 12
10 Lit Review Academic Essay Gemini 3.5 Flash 401 12 5
11 Discussion Post Discussion Board GPT-5.5 260 11 10
12 Argumentative Essay Academic Essay Gemini 3.5 Flash 374 12 10
13 Landing Copy Marketing Copy Gemini 3.5 Flash 296 7 9
14 Argumentative Essay Academic Essay Claude Sonnet 5 425 8 11
15 Lit Review Academic Essay GPT-5.5 431 11 6
16 Howto Blog Blog Post GPT-5.5 442 10 6
17 Product Desc Marketing Copy Gemini 3.5 Flash 186 9 7
18 Lit Review Academic Essay Claude Sonnet 5 384 12 6
19 Personal Statement Application Essay Gemini 3.5 Flash 316 11 8
20 Landing Copy Marketing Copy GPT-5.5 315 9 11
21 Cover Letter Application Essay Gemini 3.5 Flash 354 4 13
22 Cover Letter Application Essay GPT-5.5 411 1 13
23 Landing Copy Marketing Copy Claude Sonnet 5 313 10 8
24 News Article News Article Gemini 3.5 Flash 366 8 12
25 Listicle Blog Blog Post GPT-5.5 469 7 7
26 Business Email Business Email GPT-5.5 248 1 14
27 Listicle Blog Blog Post Gemini 3.5 Flash 476 13 6
28 Product Desc Marketing Copy GPT-5.5 194 6 12
29 Listicle Blog Blog Post Claude Sonnet 5 404 8 6
30 Howto Blog Blog Post Claude Sonnet 5 415 11 7
31 Business Email Business Email Claude Sonnet 5 278 1 12
32 Howto Blog Blog Post Gemini 3.5 Flash 454 9 8
33 News Article News Article Claude Sonnet 5 405 4 13

"Passed" counts the ranked tools whose rewrite of that prompt at least 4 of the 5 detectors scored as human-written (0.50 or above); "Failed" counts those flagged by at least one detector that scored them; "Hardest" is the single detector that flagged the most rewrites of that prompt.

Recordings

Every tool was run in a recorded sitting. Each recording opens here from the start; the outputs pages jump to the exact moment a prompt was run.

  • WriteHuman 9:39 · 33 prompts · Sep 1, 2026
  • Stealth Writer 7:25 · 32 prompts · Sep 1, 2026
  • HIX Bypass 12:32 · 31 prompts · Sep 1, 2026
  • AI Humanize io 8:59 · 32 prompts · Sep 2, 2026
  • SupWriter 7:51 · 33 prompts · Sep 2, 2026
  • Undetectable.ai 10:31 · 31 prompts · Sep 2, 2026
  • Humanize AI Pro 9:34 · 33 prompts · Sep 2, 2026
  • Phrasly 14:52 · 33 prompts · Sep 2, 2026
  • StealthGPT 46:11 · 32 prompts · Sep 2, 2026
  • Walter Writes 17:35 · 32 prompts · Sep 2, 2026
  • Super Humanizer 6:05 · 33 prompts · Sep 2, 2026
  • Grammarly 8:44 · 33 prompts · Sep 2, 2026
  • StealthBypass 13:24 · 33 prompts · Sep 2, 2026
  • Penlify 20:22 · 33 prompts · Sep 2, 2026
  • Humbot not ranked this cycle 28:26 · 33 prompts · Sep 2, 2026

Session recording