Transparency

How well does Stylome work?

See exactly how each Stylome model performs across public benchmarks, writing types, and source domains. Every result is attributed to a model version, fixed threshold, test population, and overlap audit.

Model reports

Choose a model

Every released model gets its own report. Choose either model to inspect measurements attributed only to that version.

5 benchmark reports

Strong detection with conservative calls on human writing

Stylome 1.0 is the standard model behind the scanner and browser extension. These results come from 24,505 public-benchmark passages scored with the production model and input contract.

6,068 of 6,068 human passages were cleared across all three EditLens test sections.

5,896 of 5,931 fully generated EditLens passages were detected.

2,605 of 2,606 essays were cleared across the frozen ELLIPSE and PERSUADE evaluations.

Every result below uses the deployed Stylome 1.0 threshold, with no threshold fitting on these benchmarks.

Stylome 1.0 · PAN at CLEF 2025PAN 2025Human and AI-generated essays, fiction, and news from an independent detection challenge.Human writing cleared99.61%AI-generated detected94.42%Full results

The complete labeled validation set, scored at Stylome 1.0’s deployed threshold with no threshold fitting on PAN.

PAN 2025 is an academic shared task for distinguishing human writing from generated text. Its validation set spans essays, fiction, and news, with output from 22 model families and variants.

Every one of its 3,589 passages is long enough for Stylome’s scanner. None has full-text or 1,000-character-prefix overlap with the final Phase 11 training corpus.

1,272 of 1,277 human passages were correctly left unflagged.

2,183 of 2,312 AI-generated passages were detected.

Stylome 1.0 on PAN 2025 by writing type
Writing typeHuman clearedAI generated detected
Essays99.24% (131/132)98.50% (525/533)
Fiction99.57% (924/928)92.41% (840/909)
News100% (217/217)94.02% (818/870)

Evaluation completed July 2026 on PAN’s labeled validation set, not its hidden official test set. Detection on Llama-3.1-8B-Instruct was 70.25%; the aggregate results above span all 22 model families and variants in the validation set.

Stylome 1.0 · ELLIPSE CorpusELLIPSEStandardized-test essays written by English learners in grades 8–12.Human essays cleared99.94%Training overlap0Full results

Only one false positive across a frozen evaluation of 1,713 learner-English essays.

The open-source ELLIPSE corpus contains roughly 6,500 essays written during statewide testing by English learners in grades 8–12.

After removing material inherited by an earlier model lineage, we split the remaining 3,425 eligible essays with a fixed random seed: 1,712 for training and 1,713 for this frozen evaluation.

1,712 of 1,713 human essays were correctly left unflagged.

1 of 1,713 received an incorrect AI-generated result.

Stylome 1.0 on the frozen ELLIPSE half
Human essaysCorrectly clearedFalse positivesClearance rateFalse-positive rateTraining overlap
1,7131,712199.94%0.06%0

Evaluation completed July 2026 at Stylome 1.0’s deployed threshold. The held half has zero full-text or prefix overlap with final model training. This is a Stylome-defined frozen split, not an official detector benchmark split published by the corpus authors.

Stylome 1.0 · PERSUADE 2.0PERSUADEArgumentative essays written by United States students in grades 6–12.Human essays cleared100%Training overlap0Full results

Zero false positives across a frozen evaluation of 893 student essays.

The open-source PERSUADE 2.0 corpus contains more than 25,000 argumentative essays written by students in grades 6–12 across 15 prompts.

After removing material inherited by an earlier model lineage, we split our 1,785-essay eligible sample with a fixed random seed: 892 for training and 893 for this frozen evaluation.

893 of 893 human essays were correctly left unflagged.

0 of 893 received an incorrect AI-generated result.

Stylome 1.0 on the frozen PERSUADE half
Human essaysCorrectly clearedFalse positivesClearance rateFalse-positive rate95% upper boundTraining overlap
8938930100%0%0.34%0

Evaluation completed July 2026 at Stylome 1.0’s deployed threshold. The held half has zero full-text or prefix overlap with final model training. With no observed errors, 0.34% is the conservative rule-of-three upper bound. This is a Stylome-defined frozen split, not an official detector benchmark split published by the corpus authors.

Stylome 1.0 · Pangram LabsEditLensHuman-written, AI-edited, and AI-generated text across three test sections.Human writing cleared100%AI-generated detected99.41%Full results

No human false positives across 6,068 passages, while Stylome 1.0 detected more than 99% of fully generated text in every test section.

EditLens is a dataset from Pangram Labs built to study how much AI has changed a piece of writing. Unlike a simple human-or-AI test, these results use three labels: human written, AI edited, and AI generated.

AI-edited writing remains the harder category. Stylome 1.0’s strongest three-class section was the Llama split: 92.5% overall correct, including 84.7% recall on edited text.

6,068 of 6,068 human passages were correctly left unflagged.

5,896 of 5,931 fully generated passages were detected.

3,812 of 6,220 AI-edited passages were detected at the same conservative threshold.

Performance at Stylome 1.0’s deployed threshold
Test sectionHuman writing clearedFalse positivesAI edited detectedAI generated detected
Standard test100%0 of 2,03864.11%99.36%
Enron email100%0 of 1,99256.37%99.13%
Llama100%0 of 2,03864.27%99.71%
Research benchmark

Three labels make this a harder test

The live scanner asks whether the evidence is strong enough to flag a passage. EditLens also asks the model to distinguish AI-edited writing from fully AI-generated writing. A random choice among three balanced labels would score about 33%.

Three-class performance for Stylome 1.0 on EditLens
Test sectionPassagesOverall correctBalanced scoreHuman foundAI edited foundAI generated found
Standard testReviews, news, educational web text, and creative writing6,11589.3%89.1%92.3%77.0%98.7%
Enron emailA separate business-email domain6,14788.1%88.2%92.5%76.2%98.2%
LlamaAI text produced only with Llama 3.35,95792.5%92.3%92.3%84.7%99.95%

Overall correct is the share assigned the right label. Balanced score is macro-F1: it gives each of the three labels equal weight. The final three columns show how often Stylome found each class.

The three-way scores complement the deployed-threshold results: together they show Stylome 1.0's binary detection strength and its separate performance on AI-edited writing across each official split.

Evaluation completed July 2026 on all three public EditLens test sections. The final training audit found zero full-text and zero 1,000-character-prefix collisions. Official split rows are shown so the evaluation remains comparable and reproducible.

Stylome 1.0 · Liang et al.TOEFL91Human essays written by people learning English as an additional language.Human essays cleared96.70%Training overlap0Full results

A deliberately difficult fairness check for false positives on human writing by non-native English speakers.

TOEFL91 contains 91 human-written TOEFL essays used by Liang and colleagues to study detector bias against non-native English writers.

Their 2023 study reported that seven contemporary detectors falsely flagged an average of 61.3% of these essays. Stylome 1.0 cleared 96.70% of the same essays, reducing the false-positive rate to 3.30%.

88 of 91 human essays were correctly left unflagged.

3 of 91 received an incorrect AI-generated result.

Stylome 1.0 on TOEFL91
Human essaysCorrectly clearedFalse positivesFalse-positive rate95% intervalTraining overlap
918833.30%0.69–9.33%0

Evaluation completed July 2026 at Stylome 1.0’s deployed threshold. The essays have zero full-text or prefix overlap with final model training. This small, single-register sample should not be read as a complete measure of performance for multilingual or non-native English writing.

How to read a result

Stylome uses three results instead of forcing every passage into a yes-or-no answer. The “AI generated” result is deliberately reserved for the strongest findings.

What each result means
ResultMeaning
Human writtenNo strong signs of AI generation were found.
InconclusiveSome signs of AI generation were found, but not enough for a confident result.
AI generatedStrong signs indicate that AI generated or substantially rewrote the passage.

What we test

A model can look excellent overall while struggling with a particular kind of writing. Each report will therefore show separate results for essays, fiction, web articles, social posts, conversations, and writing by people who learned English as an additional language.

For every model

  • Overall detection results
  • False positives on human writing
  • Results for each kind of writing
  • Performance at different passage lengths
  • Test size and evaluation date

Before the numbers appear

We test on writing kept separate from model training and check for accidental duplicates. That gives the published results a fairer chance of reflecting new, unseen text.

Comparing model versions

The standard and lower-latency models are evaluated on the same core tests, making differences in speed, false positives, and detection performance directly comparable.

Available Stylome models
ModelDesigned forStatusResults
Stylome 1.0Benchmark record for the previous standard model.Published benchmark5 benchmark reports available
Stylome Fast 1.0Lower-latency analysis for passages between 200 and 2,000 characters.Developer API option6 benchmark reports available

Known limits

Heavily edited or translated passages can be harder to classify, as can templates, highly unusual writing styles, and text near the 200-character minimum. Stylome evaluates the passage itself; it does not identify who wrote it.

Technical reporting details

Included with each report

  • Frozen holdout provenance and SHA-256 overlap audit
  • Matched-FPR model comparison with confidence intervals
  • Per-register and per-source tail metrics
  • TOEFL, Nairaland, and ELLIPSE ESL audits
  • Calibration method, sample counts, model version, and evaluation date

Statistics kept separate

Calibrated false-positive rate
The measured error rate on a fixed evaluation set.
Live bounty findings
Examples found by people actively searching for mistakes. Useful for improving the model, but not an overall error-rate estimate.

Read the evaluation methodology for the rules governing holdouts, overlap audits, calibration, and comparisons between models.