Methodology
Updated July 27, 2026.
Stylome treats an AI-text detector as a measurement system, not an authorship oracle. A public result must be tied to a named model, documented test population, fixed decision threshold, and known limitations. The classifier output is probabilistic evidence and must not be used as proof of who wrote a passage.
Freeze the evidence before tuning
Measurement holdouts are frozen before they are used to judge a model. Their normalized content hashes are audited against training and calibration inputs before a result is reported. If overlap is found, the affected measurement is not treated as independent evidence.
Compare at the same false-positive rate
Raw deployed accuracy numbers are not comparable when detectors use different thresholds. Stylome compares models at matched false-positive rates and reports the threshold, sample counts, and uncertainty needed to interpret the comparison.
Do not hide register shortcuts inside an aggregate
Results are decomposed by source and writing register—including news, reviews, educational writing, creative prose, email, and ESL-focused sets where available. Tail errors matter: an attractive overall score does not excuse concentrated false positives on one kind of writer.
Interpret the three outcomes
- Human-typical
- The passage does not show enough model evidence for AI involvement.
- Mixed or inconclusive
- Some AI-associated signals are present, but the result remains below the decision threshold. Editing, translation, templates, or unusual style may contribute.
- AI generated
- The passage crosses the model's conservative decision threshold. It is still a classifier result—not proof about an author or workflow.
Current deployed contract
The paste box and browser extension use Stylome Large 1.0 for passages from 200 through 10,000 characters. The developer API requires an explicit choice between Stylome Large 1.0, Stylome 1.0, and the lower-latency Stylome Fast 1.0. Each model has its own versioned scoring contract, calibrated threshold, and pinned input/tokenizer contract. Thresholds prioritize a very low false-positive rate. The public report keeps AI-edited and fully AI-generated evaluation categories visible instead of combining them in one headline number.
The methodology and published measurements serve different jobs: this page documents the rules for earning a claim; the transparency report shows the attributable results that currently meet those rules.
View published measurements