[Benchmarks](https://ownhand.dev/benchmarks) # How we tested The method behind the Ownhand benchmarks. Ownhand (v1, Muse Spark writer) ran on 2 October 2026. The other systems ran on 2 October 2026. ## Test set All 129 inputs from the 4 published cycles of [HumanizerBench](https://github.com/HumanizerBench/humanizerbench): June 2026, July 2026, August 2026 and September 2026. September 2026 is the latest cycle with published outputs, with 33 inputs. The inputs are AI-written texts, 44,835 words in all, in 7 categories: academic essay, application essay, blog post, business email, discussion board, marketing copy and news article. They were written by `claude-opus-4-7`, `claude-sonnet-5`, `gemini-3-5-flash` and `gpt-5-5`. HumanizerBench data is CC BY 4.0. We used it unchanged, including each humanizer's published output. A test HumanizerBench did not mark complete counts as a failure for that humanizer. ## Systems - Ownhand: the live API at `ownhand.dev/v1/humanize`, with the text only. No Hand, no occasion, default options. Writer model `meta/muse-spark-1.3-contributor`. Engine: v1, Muse Spark writer, run on 2 October 2026. - Ownhand with an occasion: the 33 inputs whose prompt type matches an occasion, run again with `email_internal` for business email, `chat_channel` for discussion post and `docs` for how-to blog. - Frontier models through OpenRouter, with provider defaults and one generic prompt: GPT-6.1 Sol (`openai/gpt-6.1-sol`), Claude Opus 5.5 (`anthropic/claude-opus-5.5`) and Gemini 3.8 Flash (`google/gemini-3.8-flash`). The prompt asks for a natural rewrite that keeps every fact, the meaning, the format and the length, and avoids stock AI phrasing. - The 16 commercial humanizers: the outputs HumanizerBench published, named as HumanizerBench names them. - The input, unchanged, as a floor. ## Scores Slop score From [slop-score](https://github.com/sam-paech/slop-score), run from its own code. It counts overused AI words, AI word triples and contrast patterns, and scales them against its public leaderboard. The scale is 0 to 100, and lower is better. We report the mean of the per-text scores. Meaning kept The cosine similarity of `BAAI/bge-small-en-v1.5` embeddings of input and output, mapped onto HumanizerBench's own meaning scale, 0.00 to 1.00. The mapping was fit on 600 HumanizerBench outputs, with Pearson 0.93. On the 1,675 humanizer outputs in this run it agrees with HumanizerBench's score at 0.88. Facts kept Every number, date, URL and email in the input must appear in the output, with numbers as digits. The score is the share of inputs that have such facts where all of them survive. A second column, with names, also checks multi-word capitalized names. It catches reworded headings too, so we keep it apart. Length Output words divided by input words. Intervals 95% percentile bootstrap over inputs, 2,000 resamples. ## Voice match We used 5 fictional writers, each with a short description and a few samples. For each writer, 4 samples went to Ownhand as a Hand, and the same samples and description went into the frontier models' prompt. The other 3 or 4 were kept back for scoring. Every system rewrote the same 30 drafts for every writer, each with its occasion. We measured 13 style features per text: length, sentence length, commas, periods, exclamation and question marks, colons and semicolons, lowercase starts, capital letters, contractions, greetings, sign-offs and line breaks. Each output goes to the writer whose held-back samples it sits closest to. Chance is 20%. The same occasion column uses only the 21 outputs per system whose draft is in the writer's usual occasion. ## One caveat Ownhand's word list and fact check were built with slop-score and earlier HumanizerBench texts, so its slop and facts numbers on these inputs are partly in-sample. --- Page: https://ownhand.dev/benchmarks/methodology. Index for agents: https://ownhand.dev/llms.txt