Ownhand (v1, Muse Spark writer) run on 2 October 2026. Other systems run on 2 October 2026. 129 inputs from HumanizerBench. # Benchmarks We ran Ownhand and 3 frontier models on all 129 inputs from the 4 published HumanizerBench cycles, June 2026 to September 2026, and scored them next to the published outputs of 16 commercial humanizers. Slop score, lower is better 8.4 / 100 (Median of the 16 humanizers: 13.0 / 100. The input, unchanged: 27.4 / 100.) Meaning kept, higher is better 0.98 / 1.00 (Median of the humanizers: 0.93 / 1.00. Best frontier model: 0.99 / 1.00.) Facts kept, higher is better 100% of inputs (Every number, date, URL and email kept. Median of the humanizers: 72%. Best frontier model: 100%.) Ownhand ran with no Hand and no occasion, like a generic humanizer. Means over 129 inputs. ## Results One row per system. Each value is a mean over the inputs that system completed, with its 95% interval below. The frontier models got one strong generic prompt. The humanizers' outputs are the ones HumanizerBench published. ### All cycles, 129 inputs | System | Slopout of 100, lower is better | Meaning keptout of 1.00 | Facts kept% | With names% | Lengthoutput / input | ninputs | | --- | --- | --- | --- | --- | --- | --- | | OwnhandNo Hand, no occasion | 8.46.9 to 10.0 | 0.980.98 to 0.98 | 100%100% to 100% | 100%100% to 100% | 0.93 | 129 | | Undetectable.aiCommercial humanizer | 9.17.4 to 11.0 | 0.930.92 to 0.93 | 73%63% to 82% | 53%43% to 62% | 1.51 | 129 | | PenlifyCommercial humanizer | 9.46.7 to 12.4 | 0.880.86 to 0.90 | 61%39% to 83% | 11%0% to 22% | 0.96 | 32of 33 | | PhraslyCommercial humanizer | 9.57.7 to 11.3 | 0.910.91 to 0.92 | 65%54% to 75% | 20%13% to 27% | 1.15 | 129 | | Claude Opus 5.5Frontier model, generic prompt | 9.78.0 to 11.6 | 0.970.96 to 0.97 | 100%100% to 100% | 88%82% to 94% | 1.09 | 129 | | NoteGPTCommercial humanizer | 9.87.3 to 12.3 | 0.900.89 to 0.91 | 49%33% to 64% | 25%15% to 36% | 1.07 | 63 | | WriteHumanCommercial humanizer | 11.49.6 to 13.3 | 0.930.92 to 0.93 | 75%65% to 84% | 32%23% to 41% | 1.09 | 129 | | AI Humanize ioCommercial humanizer | 11.59.4 to 13.8 | 0.920.91 to 0.94 | 72%62% to 82% | 23%16% to 32% | 1.13 | 129 | | Humanize AI ProCommercial humanizer | 12.010.0 to 13.9 | 0.940.93 to 0.94 | 72%62% to 82% | 26%18% to 34% | 1.01 | 129 | | Stealth BypassCommercial humanizer | 12.67.0 to 19.0 | 0.830.78 to 0.88 | 21%5% to 42% | 15%4% to 30% | 0.46 | 32of 33 | | Gemini 3.8 FlashFrontier model, generic prompt | 12.910.9 to 15.1 | 0.970.96 to 0.97 | 100%100% to 100% | 93%87% to 97% | 1.08 | 129 | | StealthGPTCommercial humanizer | 13.411.0 to 15.8 | 0.900.89 to 0.91 | 56%44% to 67% | 29%21% to 37% | 1.12 | 129 | | Walter WritesCommercial humanizer | 13.711.7 to 15.9 | 0.910.89 to 0.93 | 63%52% to 73% | 14%7% to 20% | 1.26 | 129 | | SupWriterCommercial humanizer | 13.99.8 to 18.4 | 0.960.95 to 0.96 | 63%42% to 84% | 25%11% to 43% | 1.06 | 33 | | Stealth WriterCommercial humanizer | 16.213.8 to 18.7 | 0.940.93 to 0.94 | 86%78% to 92% | 31%23% to 39% | 1.17 | 129 | | GrammarlyCommercial humanizer | 16.413.7 to 19.1 | 0.980.97 to 0.98 | 90%82% to 96% | 87%81% to 93% | 1.07 | 129 | | GPT-6.1 SolFrontier model, generic prompt | 16.814.1 to 19.5 | 0.990.98 to 0.99 | 100%100% to 100% | 85%77% to 91% | 1.01 | 129 | | HIX BypassCommercial humanizer | 17.114.4 to 19.8 | 0.940.94 to 0.95 | 77%68% to 86% | 32%23% to 40% | 1.15 | 129 | | HumbotCommercial humanizer | 18.114.7 to 21.6 | 0.940.94 to 0.95 | 77%65% to 87% | 33%23% to 42% | 1.14 | 96 | | Super HumanizerCommercial humanizer | 20.817.5 to 24.2 | 0.930.93 to 0.94 | 87%80% to 94% | 39%30% to 48% | 1.13 | 129 | | Input, unchangedReference | 27.423.3 to 31.3 | 1.001.00 to 1.00 | 100%100% to 100% | 100%100% to 100% | 1.00 | 129 | ### September 2026 only, 33 inputs | System | Slopout of 100, lower is better | Meaning keptout of 1.00 | Facts kept% | With names% | Lengthoutput / input | ninputs | | --- | --- | --- | --- | --- | --- | --- | | OwnhandNo Hand, no occasion | 5.23.5 to 7.2 | 0.980.98 to 0.98 | 100%100% to 100% | 100%100% to 100% | 0.94 | 33 | | WriteHumanCommercial humanizer | 7.64.9 to 10.7 | 0.930.92 to 0.94 | 89%74% to 100% | 43%25% to 61% | 1.10 | 33 | | PhraslyCommercial humanizer | 7.85.4 to 10.5 | 0.920.91 to 0.93 | 68%47% to 89% | 21%7% to 36% | 1.17 | 33 | | Undetectable.aiCommercial humanizer | 8.15.2 to 11.5 | 0.950.94 to 0.96 | 84%68% to 100% | 68%50% to 86% | 1.22 | 33 | | Claude Opus 5.5Frontier model, generic prompt | 8.25.7 to 11.2 | 0.970.96 to 0.97 | 100%100% to 100% | 89%79% to 100% | 1.10 | 33 | | Humanize AI ProCommercial humanizer | 8.55.7 to 12.1 | 0.940.94 to 0.95 | 89%74% to 100% | 32%18% to 50% | 1.01 | 33 | | GrammarlyCommercial humanizer | 9.36.0 to 13.6 | 0.970.96 to 0.97 | 79%58% to 95% | 79%61% to 93% | 1.27 | 33 | | PenlifyCommercial humanizer | 9.46.7 to 12.4 | 0.880.86 to 0.90 | 61%39% to 83% | 11%0% to 22% | 0.96 | 32of 33 | | Walter WritesCommercial humanizer | 9.66.9 to 12.8 | 0.890.83 to 0.93 | 68%47% to 89% | 14%4% to 29% | 1.30 | 33 | | AI Humanize ioCommercial humanizer | 11.07.5 to 15.0 | 0.940.93 to 0.95 | 84%68% to 100% | 32%18% to 50% | 1.02 | 33 | | Gemini 3.8 FlashFrontier model, generic prompt | 11.57.9 to 15.6 | 0.970.96 to 0.98 | 100%100% to 100% | 96%89% to 100% | 1.08 | 33 | | Stealth BypassCommercial humanizer | 12.67.0 to 19.0 | 0.830.78 to 0.88 | 21%5% to 42% | 15%4% to 30% | 0.46 | 32of 33 | | HIX BypassCommercial humanizer | 13.49.4 to 17.7 | 0.940.93 to 0.95 | 79%63% to 95% | 39%21% to 57% | 1.15 | 33 | | GPT-6.1 SolFrontier model, generic prompt | 13.59.2 to 18.3 | 0.990.98 to 0.99 | 100%100% to 100% | 89%79% to 100% | 1.01 | 33 | | SupWriterCommercial humanizer | 13.99.8 to 18.4 | 0.960.95 to 0.96 | 63%42% to 84% | 25%11% to 43% | 1.06 | 33 | | Stealth WriterCommercial humanizer | 14.49.7 to 20.0 | 0.950.94 to 0.96 | 79%58% to 95% | 36%18% to 54% | 1.14 | 33 | | Super HumanizerCommercial humanizer | 14.410.2 to 19.0 | 0.940.93 to 0.95 | 79%58% to 95% | 39%21% to 57% | 1.16 | 33 | | StealthGPTCommercial humanizer | 15.810.8 to 21.3 | 0.890.85 to 0.92 | 63%42% to 84% | 50%32% to 68% | 1.05 | 33 | | Input, unchangedReference | 24.217.4 to 31.5 | 1.001.00 to 1.00 | 100%100% to 100% | 100%100% to 100% | 1.00 | 33 | Facts kept counts inputs whose numbers, dates, URLs and emails all survive. With names also counts multi-word names, which flags reworded headings too. ## Charts All cycles. The bar is the mean and the line is the 95% interval. ### Slop score, out of 100 Lower is better. Ownhand8.4 Undetectable.ai9.1 Penlify9.4 Phrasly9.5 Claude Opus 5.59.7 NoteGPT9.8 WriteHuman11.4 AI Humanize io11.5 Humanize AI Pro12.0 Stealth Bypass12.6 Gemini 3.8 Flash12.9 StealthGPT13.4 Walter Writes13.7 SupWriter13.9 Stealth Writer16.2 Grammarly16.4 GPT-6.1 Sol16.8 HIX Bypass17.1 Humbot18.1 Super Humanizer20.8 Input, unchanged27.4 ### Meaning kept, out of 1.00 Higher is better. Input, unchanged1.00 GPT-6.1 Sol0.99 Ownhand0.98 Grammarly0.98 Claude Opus 5.50.97 Gemini 3.8 Flash0.97 SupWriter0.96 Humanize AI Pro0.94 Stealth Writer0.94 HIX Bypass0.94 Humbot0.94 Undetectable.ai0.93 WriteHuman0.93 Super Humanizer0.93 AI Humanize io0.92 Phrasly0.91 Walter Writes0.91 NoteGPT0.90 StealthGPT0.90 Penlify0.88 Stealth Bypass0.83 ### Facts kept, % of inputs Higher is better. Ownhand100% Claude Opus 5.5100% Gemini 3.8 Flash100% GPT-6.1 Sol100% Input, unchanged100% Grammarly90% Super Humanizer87% Stealth Writer86% HIX Bypass77% Humbot77% WriteHuman75% Undetectable.ai73% AI Humanize io72% Humanize AI Pro72% Phrasly65% Walter Writes63% SupWriter63% Penlify61% StealthGPT56% NoteGPT49% Stealth Bypass21% ## Voice match The test above checks how a rewrite reads. A Hand is about whose voice it is in. We used 5 fictional writers, from a terse engineer to a formal consultant, and split each one's short samples in two. Ownhand got the first half as a Hand. The frontier models got the same samples and description in their prompt. The second half was kept back for scoring. Each system rewrote the same 30 drafts for every writer. A nearest-writer test on 13 style features (length, sentence length, punctuation, casing, contractions, line breaks, greetings, sign-offs) then asks which writer each output looks like. Chance is 20%. | System | Matched writer% of outputs | Same occasion% of outputs | Style distancelower is closer | Meaning keptout of 1.00 | Facts kept% | noutputs | | --- | --- | --- | --- | --- | --- | --- | | Gemini 3.8 Flash with the samples | 53%45% to 61% | 48%29% to 67% | 0.790.74 to 0.85 | 0.91 | 97% | 149 | | Claude Opus 5.5 with the samples | 49%41% to 57% | 57%38% to 76% | 0.920.82 to 1.04 | 0.89 | 97% | 150 | | GPT-6.1 Sol with the samples | 47%39% to 55% | 52%33% to 71% | 0.820.77 to 0.87 | 0.93 | 98% | 150 | | Ownhand with the Hand | 41%33% to 49% | 48%29% to 67% | 0.790.73 to 0.85 | 0.87 | 99% | 150 | | Ownhand without a Hand | 20%14% to 27% | 33%14% to 52% | 2.161.85 to 2.47 | 0.90 | 100% | 150 | | The draft, unchanged | 20%14% to 26% | 14%0% to 29% | 1.071.01 to 1.13 | 1.00 | 100% | 150 | Same occasion: only the 21 outputs per system whose draft is in the occasion the writer's samples come from, such as a Slack message for a writer whose samples are Slack messages. Style distance is in standard units against the held-back samples. ### Matched to the right writer All drafts. The dashed line is chance, 20%. Gemini 3.8 Flash with the samples53% Claude Opus 5.5 with the samples49% GPT-6.1 Sol with the samples47% Ownhand with the Hand41% Ownhand without a Hand20% The draft, unchanged20% Across all drafts, Gemini 3.8 Flash with the samples matched the right writer most often, 53% of the time. Ownhand with the Hand reached 41%, and 20% without one. On drafts in the writer's own occasion, Ownhand with the Hand reached 48%, against 57% for Claude Opus 5.5 with the samples. Elsewhere Ownhand follows the occasion, for example a greeting and sign-off on a cold email, even when the writer's samples have none. With 5 writers and 30 short drafts this is a small test, so the intervals are wide. ## With an occasion Three HumanizerBench prompt types match an Ownhand occasion: business email as `email_internal`, discussion post as `chat_channel` and how-to blog as `docs`. We ran those 33 inputs again with the occasion set. Every row covers the same 33 inputs. An occasion changes register and length on purpose, so meaning can move with it. | System | Slopout of 100, lower is better | Meaning keptout of 1.00 | Facts kept% | With names% | Lengthoutput / input | ninputs | | --- | --- | --- | --- | --- | --- | --- | | Ownhand with occasionOccasion set, no Hand | 3.62.1 to 5.5 | 0.940.92 to 0.95 | 100%100% to 100% | 96%89% to 100% | 0.67 | 33 | | OwnhandNo Hand, no occasion | 5.83.7 to 8.0 | 0.980.98 to 0.98 | 100%100% to 100% | 100%100% to 100% | 0.93 | 33 | | GPT-6.1 SolFrontier model, generic prompt | 11.47.5 to 15.7 | 0.980.98 to 0.99 | 100%100% to 100% | 89%78% to 100% | 1.01 | 33 | | Claude Opus 5.5Frontier model, generic prompt | 5.53.6 to 7.7 | 0.960.95 to 0.97 | 100%100% to 100% | 81%67% to 96% | 1.11 | 33 | | Gemini 3.8 FlashFrontier model, generic prompt | 8.75.2 to 12.6 | 0.960.95 to 0.97 | 100%100% to 100% | 89%78% to 100% | 1.15 | 33 | | Input, unchangedReference | 18.711.8 to 26.9 | 1.001.00 to 1.00 | 100%100% to 100% | 100%100% to 100% | 1.00 | 33 | ## How we tested - Inputs: 129 AI-written texts from [HumanizerBench](https://github.com/HumanizerBench/humanizerbench) (CC BY 4.0), June 2026 to September 2026. Ownhand's rules were tuned on earlier cycles of the same benchmark. - Scores: Slop Score from slop-score, meaning as embedding similarity on HumanizerBench's scale, and facts kept. Means with 95% bootstrap intervals. - Ownhand (v1, Muse Spark writer) through its live API on 2 October 2026, and GPT-6.1 Sol, Claude Opus 5.5 and Gemini 3.8 Flash through OpenRouter on 2 October 2026. [The full method](https://ownhand.dev/benchmarks/methodology) covers every score, model and setting. --- Page: https://ownhand.dev/benchmarks. Index for agents: https://ownhand.dev/llms.txt