Benchmarks

We ran Ownhand and 3 frontier models on all 129 inputs from the 4 published HumanizerBench cycles, June 2026 to September 2026, and scored them next to the published outputs of 16 commercial humanizers.

Slop score, lower is better
8.4 / 100 Median of the 16 humanizers: 13.0 / 100. The input, unchanged: 27.4 / 100.
Meaning kept, higher is better
0.98 / 1.00 Median of the humanizers: 0.93 / 1.00. Best frontier model: 0.99 / 1.00.
Facts kept, higher is better
100% of inputs Every number, date, URL and email kept. Median of the humanizers: 72%. Best frontier model: 100%.

Ownhand ran with no Hand and no occasion, like a generic humanizer. Means over 129 inputs.

Results

One row per system. Each value is a mean over the inputs that system completed, with its 95% interval below. The frontier models got one strong generic prompt. The humanizers' outputs are the ones HumanizerBench published.

System Slopout of 100, lower is better Meaning keptout of 1.00 Facts kept% With names% Lengthoutput / input ninputs
OwnhandNo Hand, no occasion 8.46.9 to 10.00.980.98 to 0.98100%100% to 100%100%100% to 100% 0.93 129
Undetectable.aiCommercial humanizer 9.17.4 to 11.00.930.92 to 0.9373%63% to 82%53%43% to 62% 1.51 129
PenlifyCommercial humanizer 9.46.7 to 12.40.880.86 to 0.9061%39% to 83%11%0% to 22% 0.96 32of 33
PhraslyCommercial humanizer 9.57.7 to 11.30.910.91 to 0.9265%54% to 75%20%13% to 27% 1.15 129
Claude Opus 5.5Frontier model, generic prompt 9.78.0 to 11.60.970.96 to 0.97100%100% to 100%88%82% to 94% 1.09 129
NoteGPTCommercial humanizer 9.87.3 to 12.30.900.89 to 0.9149%33% to 64%25%15% to 36% 1.07 63
WriteHumanCommercial humanizer 11.49.6 to 13.30.930.92 to 0.9375%65% to 84%32%23% to 41% 1.09 129
AI Humanize ioCommercial humanizer 11.59.4 to 13.80.920.91 to 0.9472%62% to 82%23%16% to 32% 1.13 129
Humanize AI ProCommercial humanizer 12.010.0 to 13.90.940.93 to 0.9472%62% to 82%26%18% to 34% 1.01 129
Stealth BypassCommercial humanizer 12.67.0 to 19.00.830.78 to 0.8821%5% to 42%15%4% to 30% 0.46 32of 33
Gemini 3.8 FlashFrontier model, generic prompt 12.910.9 to 15.10.970.96 to 0.97100%100% to 100%93%87% to 97% 1.08 129
StealthGPTCommercial humanizer 13.411.0 to 15.80.900.89 to 0.9156%44% to 67%29%21% to 37% 1.12 129
Walter WritesCommercial humanizer 13.711.7 to 15.90.910.89 to 0.9363%52% to 73%14%7% to 20% 1.26 129
SupWriterCommercial humanizer 13.99.8 to 18.40.960.95 to 0.9663%42% to 84%25%11% to 43% 1.06 33
Stealth WriterCommercial humanizer 16.213.8 to 18.70.940.93 to 0.9486%78% to 92%31%23% to 39% 1.17 129
GrammarlyCommercial humanizer 16.413.7 to 19.10.980.97 to 0.9890%82% to 96%87%81% to 93% 1.07 129
GPT-6.1 SolFrontier model, generic prompt 16.814.1 to 19.50.990.98 to 0.99100%100% to 100%85%77% to 91% 1.01 129
HIX BypassCommercial humanizer 17.114.4 to 19.80.940.94 to 0.9577%68% to 86%32%23% to 40% 1.15 129
HumbotCommercial humanizer 18.114.7 to 21.60.940.94 to 0.9577%65% to 87%33%23% to 42% 1.14 96
Super HumanizerCommercial humanizer 20.817.5 to 24.20.930.93 to 0.9487%80% to 94%39%30% to 48% 1.13 129
Input, unchangedReference 27.423.3 to 31.31.001.00 to 1.00100%100% to 100%100%100% to 100% 1.00 129

Facts kept counts inputs whose numbers, dates, URLs and emails all survive. With names also counts multi-word names, which flags reworded headings too.

Charts

All cycles. The bar is the mean and the line is the 95% interval.

Slop score, out of 100

Lower is better.

Ownhand8.4
Undetectable.ai9.1
Penlify9.4
Phrasly9.5
Claude Opus 5.59.7
NoteGPT9.8
WriteHuman11.4
AI Humanize io11.5
Humanize AI Pro12.0
Stealth Bypass12.6
Gemini 3.8 Flash12.9
StealthGPT13.4
Walter Writes13.7
SupWriter13.9
Stealth Writer16.2
Grammarly16.4
GPT-6.1 Sol16.8
HIX Bypass17.1
Humbot18.1
Super Humanizer20.8
Input, unchanged27.4

Meaning kept, out of 1.00

Higher is better.

Input, unchanged1.00
GPT-6.1 Sol0.99
Ownhand0.98
Grammarly0.98
Claude Opus 5.50.97
Gemini 3.8 Flash0.97
SupWriter0.96
Humanize AI Pro0.94
Stealth Writer0.94
HIX Bypass0.94
Humbot0.94
Undetectable.ai0.93
WriteHuman0.93
Super Humanizer0.93
AI Humanize io0.92
Phrasly0.91
Walter Writes0.91
NoteGPT0.90
StealthGPT0.90
Penlify0.88
Stealth Bypass0.83

Facts kept, % of inputs

Higher is better.

Ownhand100%
Claude Opus 5.5100%
Gemini 3.8 Flash100%
GPT-6.1 Sol100%
Input, unchanged100%
Grammarly90%
Super Humanizer87%
Stealth Writer86%
HIX Bypass77%
Humbot77%
WriteHuman75%
Undetectable.ai73%
AI Humanize io72%
Humanize AI Pro72%
Phrasly65%
Walter Writes63%
SupWriter63%
Penlify61%
StealthGPT56%
NoteGPT49%
Stealth Bypass21%

Voice match

The test above checks how a rewrite reads. A Hand is about whose voice it is in. We used 5 fictional writers, from a terse engineer to a formal consultant, and split each one's short samples in two. Ownhand got the first half as a Hand. The frontier models got the same samples and description in their prompt. The second half was kept back for scoring.

Each system rewrote the same 30 drafts for every writer. A nearest-writer test on 13 style features (length, sentence length, punctuation, casing, contractions, line breaks, greetings, sign-offs) then asks which writer each output looks like. Chance is 20%.

System Matched writer% of outputs Same occasion% of outputs Style distancelower is closer Meaning keptout of 1.00 Facts kept% noutputs
Gemini 3.8 Flash with the samples 53%45% to 61%48%29% to 67%0.790.74 to 0.85 0.9197%149
Claude Opus 5.5 with the samples 49%41% to 57%57%38% to 76%0.920.82 to 1.04 0.8997%150
GPT-6.1 Sol with the samples 47%39% to 55%52%33% to 71%0.820.77 to 0.87 0.9398%150
Ownhand with the Hand 41%33% to 49%48%29% to 67%0.790.73 to 0.85 0.8799%150
Ownhand without a Hand 20%14% to 27%33%14% to 52%2.161.85 to 2.47 0.90100%150
The draft, unchanged 20%14% to 26%14%0% to 29%1.071.01 to 1.13 1.00100%150

Same occasion: only the 21 outputs per system whose draft is in the occasion the writer's samples come from, such as a Slack message for a writer whose samples are Slack messages. Style distance is in standard units against the held-back samples.

Matched to the right writer

All drafts. The dashed line is chance, 20%.

Gemini 3.8 Flash with the samples53%
Claude Opus 5.5 with the samples49%
GPT-6.1 Sol with the samples47%
Ownhand with the Hand41%
Ownhand without a Hand20%
The draft, unchanged20%

Across all drafts, Gemini 3.8 Flash with the samples matched the right writer most often, 53% of the time. Ownhand with the Hand reached 41%, and 20% without one. On drafts in the writer's own occasion, Ownhand with the Hand reached 48%, against 57% for Claude Opus 5.5 with the samples. Elsewhere Ownhand follows the occasion, for example a greeting and sign-off on a cold email, even when the writer's samples have none. With 5 writers and 30 short drafts this is a small test, so the intervals are wide.

With an occasion

Three HumanizerBench prompt types match an Ownhand occasion: business email as email_internal, discussion post as chat_channel and how-to blog as docs. We ran those 33 inputs again with the occasion set. Every row covers the same 33 inputs. An occasion changes register and length on purpose, so meaning can move with it.

System Slopout of 100, lower is better Meaning keptout of 1.00 Facts kept% With names% Lengthoutput / input ninputs
Ownhand with occasionOccasion set, no Hand 3.62.1 to 5.50.940.92 to 0.95100%100% to 100%96%89% to 100% 0.67 33
OwnhandNo Hand, no occasion 5.83.7 to 8.00.980.98 to 0.98100%100% to 100%100%100% to 100% 0.93 33
GPT-6.1 SolFrontier model, generic prompt 11.47.5 to 15.70.980.98 to 0.99100%100% to 100%89%78% to 100% 1.01 33
Claude Opus 5.5Frontier model, generic prompt 5.53.6 to 7.70.960.95 to 0.97100%100% to 100%81%67% to 96% 1.11 33
Gemini 3.8 FlashFrontier model, generic prompt 8.75.2 to 12.60.960.95 to 0.97100%100% to 100%89%78% to 100% 1.15 33
Input, unchangedReference 18.711.8 to 26.91.001.00 to 1.00100%100% to 100%100%100% to 100% 1.00 33

How we tested

The full method covers every score, model and setting.