Ownhand (v1, Muse Spark writer) run on 2 October 2026. Other systems run on 2 October 2026. 129 inputs from HumanizerBench.
Benchmarks
We ran Ownhand and 3 frontier models on all 129 inputs from the 4 published HumanizerBench cycles, June 2026 to September 2026, and scored them next to the published outputs of 16 commercial humanizers.
- Slop score, lower is better
- 8.4 / 100 Median of the 16 humanizers: 13.0 / 100. The input, unchanged: 27.4 / 100.
- Meaning kept, higher is better
- 0.98 / 1.00 Median of the humanizers: 0.93 / 1.00. Best frontier model: 0.99 / 1.00.
- Facts kept, higher is better
- 100% of inputs Every number, date, URL and email kept. Median of the humanizers: 72%. Best frontier model: 100%.
Ownhand ran with no Hand and no occasion, like a generic humanizer. Means over 129 inputs.
Results
One row per system. Each value is a mean over the inputs that system completed, with its 95% interval below. The frontier models got one strong generic prompt. The humanizers' outputs are the ones HumanizerBench published.
| System | Slopout of 100, lower is better | Meaning keptout of 1.00 | Facts kept% |
|---|---|---|---|
| OwnhandNo Hand, no occasion | 8.46.9 to 10.0 | 0.980.98 to 0.98 | 100%100% to 100% |
| Undetectable.aiCommercial humanizer | 9.17.4 to 11.0 | 0.930.92 to 0.93 | 73%63% to 82% |
| PenlifyCommercial humanizer | 9.46.7 to 12.4 | 0.880.86 to 0.90 | 61%39% to 83% |
| PhraslyCommercial humanizer | 9.57.7 to 11.3 | 0.910.91 to 0.92 | 65%54% to 75% |
| Claude Opus 5.5Frontier model, generic prompt | 9.78.0 to 11.6 | 0.970.96 to 0.97 | 100%100% to 100% |
| NoteGPTCommercial humanizer | 9.87.3 to 12.3 | 0.900.89 to 0.91 | 49%33% to 64% |
| WriteHumanCommercial humanizer | 11.49.6 to 13.3 | 0.930.92 to 0.93 | 75%65% to 84% |
| AI Humanize ioCommercial humanizer | 11.59.4 to 13.8 | 0.920.91 to 0.94 | 72%62% to 82% |
| Humanize AI ProCommercial humanizer | 12.010.0 to 13.9 | 0.940.93 to 0.94 | 72%62% to 82% |
| Stealth BypassCommercial humanizer | 12.67.0 to 19.0 | 0.830.78 to 0.88 | 21%5% to 42% |
| Gemini 3.8 FlashFrontier model, generic prompt | 12.910.9 to 15.1 | 0.970.96 to 0.97 | 100%100% to 100% |
| StealthGPTCommercial humanizer | 13.411.0 to 15.8 | 0.900.89 to 0.91 | 56%44% to 67% |
| Walter WritesCommercial humanizer | 13.711.7 to 15.9 | 0.910.89 to 0.93 | 63%52% to 73% |
| SupWriterCommercial humanizer | 13.99.8 to 18.4 | 0.960.95 to 0.96 | 63%42% to 84% |
| Stealth WriterCommercial humanizer | 16.213.8 to 18.7 | 0.940.93 to 0.94 | 86%78% to 92% |
| GrammarlyCommercial humanizer | 16.413.7 to 19.1 | 0.980.97 to 0.98 | 90%82% to 96% |
| GPT-6.1 SolFrontier model, generic prompt | 16.814.1 to 19.5 | 0.990.98 to 0.99 | 100%100% to 100% |
| HIX BypassCommercial humanizer | 17.114.4 to 19.8 | 0.940.94 to 0.95 | 77%68% to 86% |
| HumbotCommercial humanizer | 18.114.7 to 21.6 | 0.940.94 to 0.95 | 77%65% to 87% |
| Super HumanizerCommercial humanizer | 20.817.5 to 24.2 | 0.930.93 to 0.94 | 87%80% to 94% |
| Input, unchangedReference | 27.423.3 to 31.3 | 1.001.00 to 1.00 | 100%100% to 100% |
| System | Slopout of 100, lower is better | Meaning keptout of 1.00 | Facts kept% |
|---|---|---|---|
| OwnhandNo Hand, no occasion | 5.23.5 to 7.2 | 0.980.98 to 0.98 | 100%100% to 100% |
| WriteHumanCommercial humanizer | 7.64.9 to 10.7 | 0.930.92 to 0.94 | 89%74% to 100% |
| PhraslyCommercial humanizer | 7.85.4 to 10.5 | 0.920.91 to 0.93 | 68%47% to 89% |
| Undetectable.aiCommercial humanizer | 8.15.2 to 11.5 | 0.950.94 to 0.96 | 84%68% to 100% |
| Claude Opus 5.5Frontier model, generic prompt | 8.25.7 to 11.2 | 0.970.96 to 0.97 | 100%100% to 100% |
| Humanize AI ProCommercial humanizer | 8.55.7 to 12.1 | 0.940.94 to 0.95 | 89%74% to 100% |
| GrammarlyCommercial humanizer | 9.36.0 to 13.6 | 0.970.96 to 0.97 | 79%58% to 95% |
| PenlifyCommercial humanizer | 9.46.7 to 12.4 | 0.880.86 to 0.90 | 61%39% to 83% |
| Walter WritesCommercial humanizer | 9.66.9 to 12.8 | 0.890.83 to 0.93 | 68%47% to 89% |
| AI Humanize ioCommercial humanizer | 11.07.5 to 15.0 | 0.940.93 to 0.95 | 84%68% to 100% |
| Gemini 3.8 FlashFrontier model, generic prompt | 11.57.9 to 15.6 | 0.970.96 to 0.98 | 100%100% to 100% |
| Stealth BypassCommercial humanizer | 12.67.0 to 19.0 | 0.830.78 to 0.88 | 21%5% to 42% |
| HIX BypassCommercial humanizer | 13.49.4 to 17.7 | 0.940.93 to 0.95 | 79%63% to 95% |
| GPT-6.1 SolFrontier model, generic prompt | 13.59.2 to 18.3 | 0.990.98 to 0.99 | 100%100% to 100% |
| SupWriterCommercial humanizer | 13.99.8 to 18.4 | 0.960.95 to 0.96 | 63%42% to 84% |
| Stealth WriterCommercial humanizer | 14.49.7 to 20.0 | 0.950.94 to 0.96 | 79%58% to 95% |
| Super HumanizerCommercial humanizer | 14.410.2 to 19.0 | 0.940.93 to 0.95 | 79%58% to 95% |
| StealthGPTCommercial humanizer | 15.810.8 to 21.3 | 0.890.85 to 0.92 | 63%42% to 84% |
| Input, unchangedReference | 24.217.4 to 31.5 | 1.001.00 to 1.00 | 100%100% to 100% |
Facts kept counts inputs whose numbers, dates, URLs and emails all survive. With names also counts multi-word names, which flags reworded headings too.
Charts
All cycles. The bar is the mean and the line is the 95% interval.
Slop score, out of 100
Lower is better.
Meaning kept, out of 1.00
Higher is better.
Facts kept, % of inputs
Higher is better.
Voice match
The test above checks how a rewrite reads. A Hand is about whose voice it is in. We used 5 fictional writers, from a terse engineer to a formal consultant, and split each one's short samples in two. Ownhand got the first half as a Hand. The frontier models got the same samples and description in their prompt. The second half was kept back for scoring.
Each system rewrote the same 30 drafts for every writer. A nearest-writer test on 13 style features (length, sentence length, punctuation, casing, contractions, line breaks, greetings, sign-offs) then asks which writer each output looks like. Chance is 20%.
| System | Matched writer% of outputs | Same occasion% of outputs | Facts kept% | noutputs |
|---|---|---|---|---|
| Gemini 3.8 Flash with the samples | 53%45% to 61% | 48%29% to 67% | 97% | 149 |
| Claude Opus 5.5 with the samples | 49%41% to 57% | 57%38% to 76% | 97% | 150 |
| GPT-6.1 Sol with the samples | 47%39% to 55% | 52%33% to 71% | 98% | 150 |
| Ownhand with the Hand | 41%33% to 49% | 48%29% to 67% | 99% | 150 |
| Ownhand without a Hand | 20%14% to 27% | 33%14% to 52% | 100% | 150 |
| The draft, unchanged | 20%14% to 26% | 14%0% to 29% | 100% | 150 |
Same occasion: only the 21 outputs per system whose draft is in the occasion the writer's samples come from, such as a Slack message for a writer whose samples are Slack messages. Style distance is in standard units against the held-back samples.
Matched to the right writer
All drafts. The dashed line is chance, 20%.
Across all drafts, Gemini 3.8 Flash with the samples matched the right writer most often, 53% of the time. Ownhand with the Hand reached 41%, and 20% without one. On drafts in the writer's own occasion, Ownhand with the Hand reached 48%, against 57% for Claude Opus 5.5 with the samples. Elsewhere Ownhand follows the occasion, for example a greeting and sign-off on a cold email, even when the writer's samples have none. With 5 writers and 30 short drafts this is a small test, so the intervals are wide.
With an occasion
Three HumanizerBench prompt types match an Ownhand occasion: business email as email_internal, discussion post as chat_channel and how-to blog as docs. We ran those 33 inputs again with the occasion set. Every row covers the same 33 inputs. An occasion changes register and length on purpose, so meaning can move with it.
| System | Slopout of 100, lower is better | Meaning keptout of 1.00 | Facts kept% |
|---|---|---|---|
| Ownhand with occasionOccasion set, no Hand | 3.62.1 to 5.5 | 0.940.92 to 0.95 | 100%100% to 100% |
| OwnhandNo Hand, no occasion | 5.83.7 to 8.0 | 0.980.98 to 0.98 | 100%100% to 100% |
| GPT-6.1 SolFrontier model, generic prompt | 11.47.5 to 15.7 | 0.980.98 to 0.99 | 100%100% to 100% |
| Claude Opus 5.5Frontier model, generic prompt | 5.53.6 to 7.7 | 0.960.95 to 0.97 | 100%100% to 100% |
| Gemini 3.8 FlashFrontier model, generic prompt | 8.75.2 to 12.6 | 0.960.95 to 0.97 | 100%100% to 100% |
| Input, unchangedReference | 18.711.8 to 26.9 | 1.001.00 to 1.00 | 100%100% to 100% |
How we tested
- Inputs: 129 AI-written texts from HumanizerBench (CC BY 4.0), June 2026 to September 2026. Ownhand's rules were tuned on earlier cycles of the same benchmark.
- Scores: Slop Score from slop-score, meaning as embedding similarity on HumanizerBench's scale, and facts kept. Means with 95% bootstrap intervals.
- Ownhand (v1, Muse Spark writer) through its live API on 2 October 2026, and GPT-6.1 Sol, Claude Opus 5.5 and Gemini 3.8 Flash through OpenRouter on 2 October 2026.
The full method covers every score, model and setting.