July 2026 prompt evidence
Each of these 33 prompts was written by a frontier model, then handed to all 13 ranked humanizers, and every rewrite was submitted to 5 commercial AI detectors. Every run was screen-recorded. Open a prompt to read the input, every tool's rewrite, each detector's verdict, and to watch the recording of that run.
How to read these scores
Each detector returns a human-likelihood on a common 0 to 1 scale, where 1 means it judged the text human-written and 0 means it flagged it as AI. On these pages a verdict counts as passed when that score is at least 0.50, the midpoint of the detector's own scale. That threshold exists only to draw the chips: the bypass rate on the leaderboard is the mean of each test's median score across the 5 detectors, a continuous number, so a tool's pass count and its bypass rate will not be the same figure.
A few runs come back with fewer than 5 verdicts because a detector did not return a usable score for them. A run still counts as passing at least 4 detectors only if that many detectors actually scored it and passed it. A partially scored run is counted as caught if any detector that did score it flagged the output, and is otherwise listed separately as partly scored.
Meaning is the input↔output embedding cosine and readability is a
language-model writing-quality rating, both published per test in
tests.json. Words is the output's length as a multiple of the input's; the scoring
code penalizes ratios above 1.40 or below 0.60. Full definitions live in the
methodology.
| Prompt | Category | Written by | Input words | Passed | Failed | Hardest | Open |
|---|---|---|---|---|---|---|---|
| 01 Argumentative Essay | Academic Essay | GPT-5.5 | 434 | 7 | 7 | ||
| 02 Listicle Blog | Blog Post | GPT-5.5 | 398 | 9 | 11 | ||
| 03 Cover Letter | Application Essay | Gemini 3.5 Flash | 339 | 0 | 13 | ||
| 04 Business Email | Business Email | Gemini 3.5 Flash | 266 | 1 | 13 | ||
| 05 Howto Blog | Blog Post | Gemini 3.5 Flash | 384 | 5 | 12 | ||
| 06 Lit Review | Academic Essay | Claude Sonnet 5 | 384 | 3 | 11 | ||
| 07 Product Desc | Marketing Copy | GPT-5.5 | 195 | 5 | 10 | ||
| 08 Landing Copy | Marketing Copy | GPT-5.5 | 273 | 4 | 10 | ||
| 09 Personal Statement | Application Essay | GPT-5.5 | 377 | 7 | 11 | ||
| 10 Business Email | Business Email | Claude Sonnet 5 | 276 | 3 | 11 | ||
| 11 Product Desc | Marketing Copy | Claude Sonnet 5 | 196 | 9 | 9 | ||
| 12 Listicle Blog | Blog Post | Gemini 3.5 Flash | 405 | 8 | 8 | ||
| 13 Product Desc | Marketing Copy | Gemini 3.5 Flash | 170 | 8 | 11 | ||
| 14 Howto Blog | Blog Post | Claude Sonnet 5 | 431 | 4 | 12 | ||
| 15 Cover Letter | Application Essay | GPT-5.5 | 377 | 2 | 13 | ||
| 16 Argumentative Essay | Academic Essay | Gemini 3.5 Flash | 387 | 3 | 13 | ||
| 17 Discussion Post | Discussion Board | Gemini 3.5 Flash | 255 | 7 | 7 | ||
| 18 Lit Review | Academic Essay | Gemini 3.5 Flash | 343 | 3 | 12 | ||
| 19 Cover Letter | Application Essay | Claude Sonnet 5 | 407 | 1 | 12 | ||
| 20 Business Email | Business Email | GPT-5.5 | 255 | 0 | 13 | ||
| 21 Discussion Post | Discussion Board | GPT-5.5 | 255 | 6 | 9 | ||
| 22 Personal Statement | Application Essay | Claude Sonnet 5 | 369 | 10 | 6 | ||
| 23 Landing Copy | Marketing Copy | Gemini 3.5 Flash | 279 | 4 | 11 | ||
| 24 News Article | News Article | Gemini 3.5 Flash | 391 | 7 | 10 | ||
| 25 Listicle Blog | Blog Post | Claude Sonnet 5 | 397 | 10 | 5 | ||
| 26 Landing Copy | Marketing Copy | Claude Sonnet 5 | 294 | 8 | 12 | ||
| 27 Discussion Post | Discussion Board | Claude Sonnet 5 | 254 | 7 | 9 | ||
| 28 Argumentative Essay | Academic Essay | Claude Sonnet 5 | 417 | 9 | 8 | ||
| 29 Lit Review | Academic Essay | GPT-5.5 | 432 | 5 | 9 | ||
| 30 Howto Blog | Blog Post | GPT-5.5 | 423 | 7 | 8 | ||
| 31 News Article | News Article | GPT-5.5 | 463 | 10 | 8 | ||
| 32 Personal Statement | Application Essay | Gemini 3.5 Flash | 325 | 8 | 7 | ||
| 33 News Article | News Article | Claude Sonnet 5 | 424 | 7 | 10 |
"Passed" counts the ranked tools whose rewrite of that prompt at least 4 of the 5 detectors scored as human-written (0.50 or above); "Failed" counts those flagged by at least one detector that scored them; "Hardest" is the single detector that flagged the most rewrites of that prompt.
Recordings
Every tool was run in a recorded sitting. Each recording opens here from the start; the outputs pages jump to the exact moment a prompt was run.
-
WriteHuman 10:15 · 33 prompts · Jul 1, 2026
-
Undetectable.ai 11:36 · 33 prompts · Jul 1, 2026
-
Humanize AI Pro 7:28 · 33 prompts · Jul 1, 2026
-
Stealth Writer 6:33 · 32 prompts · Jul 1, 2026
-
Humbot 15:02 · 33 prompts · Jul 1, 2026
-
HIX Bypass 10:33 · 32 prompts · Jul 1, 2026
-
Walter Writes 18:21 · 33 prompts · Jul 1, 2026
-
StealthGPT 14:40 · 33 prompts · Jul 1, 2026
-
Phrasly 13:11 · 33 prompts · Jul 1, 2026
-
AI Humanize io 17:06 · 33 prompts · Jul 1, 2026
-
Super Humanizer 6:22 · 33 prompts · Jul 1, 2026
-
Grammarly 6:46 · 33 prompts · Jul 1, 2026
-
NoteGPT 12:13 · 33 prompts · Jul 1, 2026