For every model
- Overall detection results
- False positives on human writing
- Results for each kind of writing
- Performance at different passage lengths
- Test size and evaluation date
See exactly how each Stylome model performs across public benchmarks, writing types, and source domains. Every result is attributed to a model version, fixed threshold, test population, and overlap audit.
Every released model gets its own report. Choose either model to inspect measurements attributed only to that version.
Stylome 1.0 is the standard model behind the scanner and browser extension. These results come from 24,505 public-benchmark passages scored with the production model and input contract.
6,068 of 6,068 human passages were cleared across all three EditLens test sections.
5,896 of 5,931 fully generated EditLens passages were detected.
2,605 of 2,606 essays were cleared across the frozen ELLIPSE and PERSUADE evaluations.
Every result below uses the deployed Stylome 1.0 threshold, with no threshold fitting on these benchmarks.
The complete labeled validation set, scored at Stylome 1.0’s deployed threshold with no threshold fitting on PAN.
PAN 2025 is an academic shared task for distinguishing human writing from generated text. Its validation set spans essays, fiction, and news, with output from 22 model families and variants.
Every one of its 3,589 passages is long enough for Stylome’s scanner. None has full-text or 1,000-character-prefix overlap with the final Phase 11 training corpus.
1,272 of 1,277 human passages were correctly left unflagged.
2,183 of 2,312 AI-generated passages were detected.
| Writing type | Human cleared | AI generated detected |
|---|---|---|
| Essays | 99.24% (131/132) | 98.50% (525/533) |
| Fiction | 99.57% (924/928) | 92.41% (840/909) |
| News | 100% (217/217) | 94.02% (818/870) |
Evaluation completed July 2026 on PAN’s labeled validation set, not its hidden official test set. Detection on Llama-3.1-8B-Instruct was 70.25%; the aggregate results above span all 22 model families and variants in the validation set.
Only one false positive across a frozen evaluation of 1,713 learner-English essays.
The open-source ELLIPSE corpus contains roughly 6,500 essays written during statewide testing by English learners in grades 8–12.
After removing material inherited by an earlier model lineage, we split the remaining 3,425 eligible essays with a fixed random seed: 1,712 for training and 1,713 for this frozen evaluation.
1,712 of 1,713 human essays were correctly left unflagged.
1 of 1,713 received an incorrect AI-generated result.
| Human essays | Correctly cleared | False positives | Clearance rate | False-positive rate | Training overlap |
|---|---|---|---|---|---|
| 1,713 | 1,712 | 1 | 99.94% | 0.06% | 0 |
Evaluation completed July 2026 at Stylome 1.0’s deployed threshold. The held half has zero full-text or prefix overlap with final model training. This is a Stylome-defined frozen split, not an official detector benchmark split published by the corpus authors.
Zero false positives across a frozen evaluation of 893 student essays.
The open-source PERSUADE 2.0 corpus contains more than 25,000 argumentative essays written by students in grades 6–12 across 15 prompts.
After removing material inherited by an earlier model lineage, we split our 1,785-essay eligible sample with a fixed random seed: 892 for training and 893 for this frozen evaluation.
893 of 893 human essays were correctly left unflagged.
0 of 893 received an incorrect AI-generated result.
| Human essays | Correctly cleared | False positives | Clearance rate | False-positive rate | 95% upper bound | Training overlap |
|---|---|---|---|---|---|---|
| 893 | 893 | 0 | 100% | 0% | 0.34% | 0 |
Evaluation completed July 2026 at Stylome 1.0’s deployed threshold. The held half has zero full-text or prefix overlap with final model training. With no observed errors, 0.34% is the conservative rule-of-three upper bound. This is a Stylome-defined frozen split, not an official detector benchmark split published by the corpus authors.
No human false positives across 6,068 passages, while Stylome 1.0 detected more than 99% of fully generated text in every test section.
EditLens is a dataset from Pangram Labs built to study how much AI has changed a piece of writing. Unlike a simple human-or-AI test, these results use three labels: human written, AI edited, and AI generated.
AI-edited writing remains the harder category. Stylome 1.0’s strongest three-class section was the Llama split: 92.5% overall correct, including 84.7% recall on edited text.
6,068 of 6,068 human passages were correctly left unflagged.
5,896 of 5,931 fully generated passages were detected.
3,812 of 6,220 AI-edited passages were detected at the same conservative threshold.
| Test section | Human writing cleared | False positives | AI edited detected | AI generated detected |
|---|---|---|---|---|
| Standard test | 100% | 0 of 2,038 | 64.11% | 99.36% |
| Enron email | 100% | 0 of 1,992 | 56.37% | 99.13% |
| Llama | 100% | 0 of 2,038 | 64.27% | 99.71% |
The live scanner asks whether the evidence is strong enough to flag a passage. EditLens also asks the model to distinguish AI-edited writing from fully AI-generated writing. A random choice among three balanced labels would score about 33%.
| Test section | Passages | Overall correct | Balanced score | Human found | AI edited found | AI generated found |
|---|---|---|---|---|---|---|
| Standard testReviews, news, educational web text, and creative writing | 6,115 | 89.3% | 89.1% | 92.3% | 77.0% | 98.7% |
| Enron emailA separate business-email domain | 6,147 | 88.1% | 88.2% | 92.5% | 76.2% | 98.2% |
| LlamaAI text produced only with Llama 3.3 | 5,957 | 92.5% | 92.3% | 92.3% | 84.7% | 99.95% |
Overall correct is the share assigned the right label. Balanced score is macro-F1: it gives each of the three labels equal weight. The final three columns show how often Stylome found each class.
The three-way scores complement the deployed-threshold results: together they show Stylome 1.0's binary detection strength and its separate performance on AI-edited writing across each official split.
Evaluation completed July 2026 on all three public EditLens test sections. The final training audit found zero full-text and zero 1,000-character-prefix collisions. Official split rows are shown so the evaluation remains comparable and reproducible.
A deliberately difficult fairness check for false positives on human writing by non-native English speakers.
TOEFL91 contains 91 human-written TOEFL essays used by Liang and colleagues to study detector bias against non-native English writers.
Their 2023 study reported that seven contemporary detectors falsely flagged an average of 61.3% of these essays. Stylome 1.0 cleared 96.70% of the same essays, reducing the false-positive rate to 3.30%.
88 of 91 human essays were correctly left unflagged.
3 of 91 received an incorrect AI-generated result.
| Human essays | Correctly cleared | False positives | False-positive rate | 95% interval | Training overlap |
|---|---|---|---|---|---|
| 91 | 88 | 3 | 3.30% | 0.69–9.33% | 0 |
Evaluation completed July 2026 at Stylome 1.0’s deployed threshold. The essays have zero full-text or prefix overlap with final model training. This small, single-register sample should not be read as a complete measure of performance for multilingual or non-native English writing.
Stylome uses three results instead of forcing every passage into a yes-or-no answer. The “AI generated” result is deliberately reserved for the strongest findings.
| Result | Meaning |
|---|---|
| Human written | No strong signs of AI generation were found. |
| Inconclusive | Some signs of AI generation were found, but not enough for a confident result. |
| AI generated | Strong signs indicate that AI generated or substantially rewrote the passage. |
A model can look excellent overall while struggling with a particular kind of writing. Each report will therefore show separate results for essays, fiction, web articles, social posts, conversations, and writing by people who learned English as an additional language.
We test on writing kept separate from model training and check for accidental duplicates. That gives the published results a fairer chance of reflecting new, unseen text.
The standard and lower-latency models are evaluated on the same core tests, making differences in speed, false positives, and detection performance directly comparable.
| Model | Designed for | Status | Results |
|---|---|---|---|
| Stylome 1.0 | Benchmark record for the previous standard model. | Published benchmark | 5 benchmark reports available |
| Stylome Fast 1.0 | Lower-latency analysis for passages between 200 and 2,000 characters. | Developer API option | 6 benchmark reports available |
Heavily edited or translated passages can be harder to classify, as can templates, highly unusual writing styles, and text near the 200-character minimum. Stylome evaluates the passage itself; it does not identify who wrote it.
Read the evaluation methodology for the rules governing holdouts, overlap audits, calibration, and comparisons between models.