Skip to content
The September 2026 results are in: WriteHuman holds #1, and newcomer SupWriter debuts at #5.Read the full analysis →
HumanizerBench

Head to head · September 2026 cycle

Walter Writes vs WriteHuman

WriteHuman finished 16.6 points ahead of Walter Writes in the September 2026 cycle, 78.29 to 61.70, ranking #1 against #10 in a field of 14. WriteHuman scored higher on 4 of 5 detectors and 7 of 7 writing categories. Walter Writes came out ahead on GPTZero (78.3 vs 71.8) and readability (60.4 vs 56.9). The bypass-rate gap, 83.9 to 90.1, sits inside both tools' 95% confidence intervals, so treat bypass as a draw.

61.70 /100
overall score
#10 /14
rank ▼ 1
+16.59
WriteHuman leads →
Detectors
1 4
Categories
0 7
Components
1 4
78.29 /100
overall score
#1 /14
rank -
Last tested
Prompts
33
Methodology
v1.2.0

Score components

What the overall score is made of: bypass rate weighs 42%, meaning preservation 32%, readability 16%, consistency across categories 10%. Output-quality penalties come off the total. How scoring works →

Bypass rate CIs overlap WriteHuman by 6.3
83.9 90.1
Meaning preservation CIs overlap WriteHuman by 5.6
66.6 72.3
Readability Walter Writes by 3.4
60.4 56.9
Consistency WriteHuman by 7.2
74.9 82.1
Penalties fewer is better WriteHuman by 12.0
−12.0 None

CIs overlap marks a gap that sits inside both tools' published 95% confidence intervals for that component; it may not survive another cycle.

Detector by detector

Share of each tool's outputs that a detector classified as human-written. WriteHuman took 4 of 5.

GPTZero Walter Writes by 6.5
78.3 71.8
Originality.ai WriteHuman by 6.9
84.5 91.4
Copyleaks WriteHuman by 27.6
59.1 86.7
Winston AI WriteHuman by 17.1
67.0 84.1
ZeroGPT WriteHuman by 1.9
76.0 77.8

Writing categories

Category score per writing context: bypass credit only for real rewrites, blended with the category's own meaning preservation and readability. Small categories swing; the prompt count is shown on each row. How category scores work →

Academic Essay Application Essay Blog Post Business Email Marketing Copy Discussion Board News Article Academic Essay: Walter Writes 47.9 · WriteHuman 84.1 Application Essay: Walter Writes 57.3 · WriteHuman 71.6 Blog Post: Walter Writes 70.8 · WriteHuman 83.9 Business Email: Walter Writes 54.1 · WriteHuman 56.7 Marketing Copy: Walter Writes 45.3 · WriteHuman 82.1 Discussion Board: Walter Writes 74.4 · WriteHuman 78.2 News Article: Walter Writes 49.8 · WriteHuman 78.3
Walter Writes WriteHuman Outer ring = 100
Academic Essay 6 prompts WriteHuman by 36.2
47.9 84.1
Application Essay 6 prompts WriteHuman by 14.3
57.3 71.6
Blog Post 6 prompts WriteHuman by 13.1
70.8 83.9
Business Email 3 prompts WriteHuman by 2.6
54.1 56.7
Marketing Copy 6 prompts WriteHuman by 36.8
45.3 82.1
Discussion Board 3 prompts WriteHuman by 3.8
74.4 78.2
News Article 3 prompts WriteHuman by 28.5
49.8 78.3

Head-to-head history

Overall score in every cycle both tools were tested. WriteHuman has finished ahead in 4 of 4 shared cycles.

50 60 70 80 90 June 2026 July 2026 August 2026 September 2026 June 2026: WriteHuman 83.59 (#2) June 2026: Walter Writes 67.24 (#10) July 2026: WriteHuman 73.07 (#1) July 2026: Walter Writes 62.84 (#7) August 2026: WriteHuman 76.69 (#1) August 2026: Walter Writes 60.83 (#9) September 2026: WriteHuman 78.29 (#1) September 2026: Walter Writes 61.70 (#10) 61.7 78.3
Walter Writes WriteHuman
Cycle Walter Writes WriteHuman Margin
September 2026 61.70 #10 78.29 #1 → 16.6
August 2026 60.83 #9 76.69 #1 → 15.9
July 2026 62.84 #7 73.07 #1 → 10.2
June 2026 67.24 #10 83.59 #2 → 16.4

Where each one wins

Every measure above that one tool won outright, largest margin first within each group.

Walter Writes

2 wins
  • Readability Score 60.4 vs 56.9
  • GPTZero Detector 78.3 vs 71.8

WriteHuman

15 wins
  • Penalties Score None vs −12.0
  • Consistency Score 82.1 vs 74.9
  • Bypass rate Score 90.1 vs 83.9
  • Meaning preservation Score 72.3 vs 66.6
  • Copyleaks Detector 86.7 vs 59.1
  • Winston AI Detector 84.1 vs 67.0
  • Originality.ai Detector 91.4 vs 84.5
  • ZeroGPT Detector 77.8 vs 76.0
  • Marketing Copy Category 82.1 vs 45.3
  • Academic Essay Category 84.1 vs 47.9
  • and 5 more

Questions people ask

Is Walter Writes better than WriteHuman?

WriteHuman finished 16.6 points ahead of Walter Writes in the September 2026 cycle, 78.29 to 61.70, ranking #1 against #10 in a field of 14. WriteHuman scored higher on 4 of 5 detectors and 7 of 7 writing categories. Both tools ran the same 33 prompts, scored by the same 5 detectors, under methodology v1.2.0; every input, output and verdict is published.

Which bypasses AI detectors better, Walter Writes or WriteHuman?

WriteHuman posted the higher bypass rate, 90.1 vs 83.9, and scored higher on 4 of 5 detectors. Walter Writes led on GPTZero; WriteHuman led on Originality.ai, Copyleaks, Winston AI and ZeroGPT. The overall bypass gap is inside both 95% confidence intervals, so it is not a reliable difference.

Do Walter Writes and WriteHuman keep the original meaning?

WriteHuman preserved meaning better, 72.3 vs 66.6 on our 0–100 similarity scale, though the confidence intervals overlap. On readability, Walter Writes rated higher, 60.4 vs 56.9. Penalties this cycle: Walter Writes −12.0, WriteHuman None.