September 2026 prompt evidence
Each of these 33 prompts was written by a frontier model, then handed to all 14 ranked humanizers, and every rewrite was submitted to 5 commercial AI detectors. Every run was screen-recorded. Open a prompt to read the input, every tool's rewrite, each detector's verdict, and to watch the recording of that run.
How to read these scores
Each detector returns a human-likelihood on a common 0 to 1 scale, where 1 means it judged the text human-written and 0 means it flagged it as AI. On these pages a verdict counts as passed when that score is at least 0.50, the midpoint of the detector's own scale. That threshold exists only to draw the chips: the bypass rate on the leaderboard is the mean of each test's median score across the 5 detectors, a continuous number, so a tool's pass count and its bypass rate will not be the same figure.
A few runs come back with fewer than 5 verdicts because a detector did not return a usable score for them. A run still counts as passing at least 4 detectors only if that many detectors actually scored it and passed it. A partially scored run is counted as caught if any detector that did score it flagged the output, and is otherwise listed separately as partly scored.
Meaning is the input↔output embedding cosine and readability is a
language-model writing-quality rating, both published per test in
tests.json. Words is the output's length as a multiple of the input's; the scoring
code penalizes ratios above 1.40 or below 0.60. Full definitions live in the
methodology.
| Prompt | Category | Written by | Input words | Passed | Failed | Hardest | Open |
|---|---|---|---|---|---|---|---|
| 01 Argumentative Essay | Academic Essay | GPT-5.5 | 434 | 10 | 9 | ||
| 02 Discussion Post | Discussion Board | Gemini 3.5 Flash | 253 | 11 | 10 | ||
| 03 Personal Statement | Application Essay | GPT-5.5 | 364 | 11 | 9 | ||
| 04 Discussion Post | Discussion Board | Claude Sonnet 5 | 246 | 11 | 8 | ||
| 05 News Article | News Article | GPT-5.5 | 452 | 11 | 6 | ||
| 06 Product Desc | Marketing Copy | Claude Sonnet 5 | 192 | 9 | 10 | ||
| 07 Cover Letter | Application Essay | Claude Sonnet 5 | 405 | 1 | 14 | ||
| 08 Personal Statement | Application Essay | Claude Sonnet 5 | 361 | 10 | 8 | ||
| 09 Business Email | Business Email | Gemini 3.5 Flash | 269 | 1 | 12 | ||
| 10 Lit Review | Academic Essay | Gemini 3.5 Flash | 401 | 12 | 5 | ||
| 11 Discussion Post | Discussion Board | GPT-5.5 | 260 | 11 | 10 | ||
| 12 Argumentative Essay | Academic Essay | Gemini 3.5 Flash | 374 | 12 | 10 | ||
| 13 Landing Copy | Marketing Copy | Gemini 3.5 Flash | 296 | 7 | 9 | ||
| 14 Argumentative Essay | Academic Essay | Claude Sonnet 5 | 425 | 8 | 11 | ||
| 15 Lit Review | Academic Essay | GPT-5.5 | 431 | 11 | 6 | ||
| 16 Howto Blog | Blog Post | GPT-5.5 | 442 | 10 | 6 | ||
| 17 Product Desc | Marketing Copy | Gemini 3.5 Flash | 186 | 9 | 7 | ||
| 18 Lit Review | Academic Essay | Claude Sonnet 5 | 384 | 12 | 6 | ||
| 19 Personal Statement | Application Essay | Gemini 3.5 Flash | 316 | 11 | 8 | ||
| 20 Landing Copy | Marketing Copy | GPT-5.5 | 315 | 9 | 11 | ||
| 21 Cover Letter | Application Essay | Gemini 3.5 Flash | 354 | 4 | 13 | ||
| 22 Cover Letter | Application Essay | GPT-5.5 | 411 | 1 | 13 | ||
| 23 Landing Copy | Marketing Copy | Claude Sonnet 5 | 313 | 10 | 8 | ||
| 24 News Article | News Article | Gemini 3.5 Flash | 366 | 8 | 12 | ||
| 25 Listicle Blog | Blog Post | GPT-5.5 | 469 | 7 | 7 | ||
| 26 Business Email | Business Email | GPT-5.5 | 248 | 1 | 14 | ||
| 27 Listicle Blog | Blog Post | Gemini 3.5 Flash | 476 | 13 | 6 | ||
| 28 Product Desc | Marketing Copy | GPT-5.5 | 194 | 6 | 12 | ||
| 29 Listicle Blog | Blog Post | Claude Sonnet 5 | 404 | 8 | 6 | ||
| 30 Howto Blog | Blog Post | Claude Sonnet 5 | 415 | 11 | 7 | ||
| 31 Business Email | Business Email | Claude Sonnet 5 | 278 | 1 | 12 | ||
| 32 Howto Blog | Blog Post | Gemini 3.5 Flash | 454 | 9 | 8 | ||
| 33 News Article | News Article | Claude Sonnet 5 | 405 | 4 | 13 |
"Passed" counts the ranked tools whose rewrite of that prompt at least 4 of the 5 detectors scored as human-written (0.50 or above); "Failed" counts those flagged by at least one detector that scored them; "Hardest" is the single detector that flagged the most rewrites of that prompt.
Recordings
Every tool was run in a recorded sitting. Each recording opens here from the start; the outputs pages jump to the exact moment a prompt was run.
-
WriteHuman 9:39 · 33 prompts · Sep 1, 2026
-
Stealth Writer 7:25 · 32 prompts · Sep 1, 2026
-
HIX Bypass 12:32 · 31 prompts · Sep 1, 2026
-
AI Humanize io 8:59 · 32 prompts · Sep 2, 2026
-
SupWriter 7:51 · 33 prompts · Sep 2, 2026
-
Undetectable.ai 10:31 · 31 prompts · Sep 2, 2026
-
Humanize AI Pro 9:34 · 33 prompts · Sep 2, 2026
-
Phrasly 14:52 · 33 prompts · Sep 2, 2026
-
StealthGPT 46:11 · 32 prompts · Sep 2, 2026
-
Walter Writes 17:35 · 32 prompts · Sep 2, 2026
-
Super Humanizer 6:05 · 33 prompts · Sep 2, 2026
-
Grammarly 8:44 · 33 prompts · Sep 2, 2026
-
StealthBypass 13:24 · 33 prompts · Sep 2, 2026
-
Penlify 20:22 · 33 prompts · Sep 2, 2026
-
Humbot not ranked this cycle 28:26 · 33 prompts · Sep 2, 2026