How we tested

The method behind the Ownhand benchmarks. Ownhand (v1, Muse Spark writer) ran on 2 October 2026. The other systems ran on 2 October 2026.

Test set

All 129 inputs from the 4 published cycles of HumanizerBench: June 2026, July 2026, August 2026 and September 2026. September 2026 is the latest cycle with published outputs, with 33 inputs.

The inputs are AI-written texts, 44,835 words in all, in 7 categories: academic essay, application essay, blog post, business email, discussion board, marketing copy and news article. They were written by claude-opus-4-7, claude-sonnet-5, gemini-3-5-flash and gpt-5-5.

HumanizerBench data is CC BY 4.0. We used it unchanged, including each humanizer's published output. A test HumanizerBench did not mark complete counts as a failure for that humanizer.

Systems

Scores

Slop score
From slop-score, run from its own code. It counts overused AI words, AI word triples and contrast patterns, and scales them against its public leaderboard. The scale is 0 to 100, and lower is better. We report the mean of the per-text scores.
Meaning kept
The cosine similarity of BAAI/bge-small-en-v1.5 embeddings of input and output, mapped onto HumanizerBench's own meaning scale, 0.00 to 1.00. The mapping was fit on 600 HumanizerBench outputs, with Pearson 0.93. On the 1,675 humanizer outputs in this run it agrees with HumanizerBench's score at 0.88.
Facts kept
Every number, date, URL and email in the input must appear in the output, with numbers as digits. The score is the share of inputs that have such facts where all of them survive. A second column, with names, also checks multi-word capitalized names. It catches reworded headings too, so we keep it apart.
Length
Output words divided by input words.
Intervals
95% percentile bootstrap over inputs, 2,000 resamples.

Voice match

We used 5 fictional writers, each with a short description and a few samples. For each writer, 4 samples went to Ownhand as a Hand, and the same samples and description went into the frontier models' prompt. The other 3 or 4 were kept back for scoring.

Every system rewrote the same 30 drafts for every writer, each with its occasion. We measured 13 style features per text: length, sentence length, commas, periods, exclamation and question marks, colons and semicolons, lowercase starts, capital letters, contractions, greetings, sign-offs and line breaks. Each output goes to the writer whose held-back samples it sits closest to. Chance is 20%. The same occasion column uses only the 21 outputs per system whose draft is in the writer's usual occasion.

One caveat

Ownhand's word list and fact check were built with slop-score and earlier HumanizerBench texts, so its slop and facts numbers on these inputs are partly in-sample.