← Back
SiTech Team⏱️ 8 წთ. საკითხავი

AI Text Detectors Struggle When Language Models Mimic an Author's Style — Study

AI Text Detectors Struggle When Language Models Mimic an Author's Style — Study

An Epoch AI study reveals that Pangram, GPTZero, and Originality.ai miss 13% of AI-generated text when it mimics an author's writing style — and in scientific writing, the miss rate climbs to 24-29%.

Epoch AI's Groundbreaking Study: How the Test Worked

A research team at Epoch AI conducted a large-scale evaluation of three widely used AI text detectors: Pangram (version 3.3.2), GPTZero (model 2026-05-11-base), and Originality.ai (Turbo 3.0.2). The goal was to determine how effectively these tools perform when language models deliberately mimic a specific author's writing style — a scenario far closer to real-world usage than simple prompt-and-generate tasks.

The researchers built a corpus of 495 human-written passages from 99 authors, evenly split across three categories: blogging, fiction, and scientific writing. Critically, all texts were written before ChatGPT's release in November 2022, effectively ruling out contamination by language models. This is a key methodological strength — many earlier studies failed to control for this factor, making their results unreliable if control texts could themselves have been AI-generated.

The testing structure had two phases. First, the detectors were evaluated on plain AI-generated text — output created from simple prompts without any style imitation. Second, three frontier AI models — Claude Opus 4.8, GPT-5.5, and Gemini 3.1 Pro — each received five real text passages from a human author and were instructed to produce new content in the same style. This generated 297 style-imitated passages for testing.

📖 Source

Near-Perfect Results on Plain AI Text — But a Critical Caveat

When the detectors were fed plain AI-generated text (without style imitation), all three performed almost flawlessly. The false-negative rate — the percentage of AI texts that go undetected — maxed out at just 0.7%. This means the tools correctly identified over 99% of straightforward AI output, a result that matches the marketing claims of most detection services.

Human-written texts were also classified correctly for the most part. Pangram and GPTZero didn't produce a single false alarm — they correctly identified every human-authored passage as human. This suggests that if a teacher or publisher uses these tools, they will almost never falsely accuse an innocent author of using AI.

However, Originality.ai showed a concerning result: it flagged 19 out of 495 human passages as AI-generated, a false-positive rate of 3.8%. This means roughly one in every 26 human-written texts could be incorrectly labeled as AI-generated — an unacceptably high rate in situations where such an accusation carries serious consequences, such as academic misconduct hearings or journalistic integrity investigations.

📖 Source

Style Imitation Drastically Reduces Detection Accuracy

The picture changes dramatically when AI models receive sample texts from an author and are asked to mimic their style. According to Epoch AI's data, of the 297 style-imitated passages generated for the test, an average of 38 went completely undetected — a false-negative rate of approximately 13%. In other words, more than one in eight AI-generated texts slipped past every detector when the language model had been primed with an author's writing samples.

Performance varied notably between detectors. Pangram missed 10% of style-imitated texts, GPTZero missed 11%, and Originality.ai missed 18% — meaning the latter failed to catch nearly one in five AI-generated passages when style imitation was employed. This represents a dramatic drop from the near-perfect results seen on plain AI text, where the miss rate was below 1%.

Interestingly, the fiction category proved relatively resistant to deception, with false-negative rates across all detectors sitting at just 1–5%. This makes intuitive sense: fiction authors tend to have more distinctive, idiosyncratic voices that are harder for AI to replicate convincingly. The unique turns of phrase, narrative rhythms, and character voices that define literary style apparently remain challenging for even frontier models to simulate perfectly.

📖 Source

Scientific Writing: The Achilles' Heel of AI Detection

If fiction was a relative success story for the detectors, scientific writing was their undoing. This is the category where all three tools performed worst — and it happens to be the genre where AI text detection likely sees the most real-world use, particularly in universities and academic publishing.

Pangram failed to catch 25% of style-imitated academic AI texts, GPTZero missed 24%, and Originality.ai missed 29%. These are not marginal differences — roughly one in every four scientific AI-generated passages slipped through undetected, even with the best-performing detector. For Originality.ai, the figure approached one in three.

The worst individual results emerged from specific model-genre combinations within scientific writing. Pangram missed a staggering 48% of Gemini-generated academic passages — essentially failing to detect every other text. At Originality.ai, 39% of GPT-5.5 academic texts went undetected. These numbers are particularly alarming given that students and researchers increasingly turn to AI for academic writing assistance. The very genre where detection matters most is the one where detectors fail most spectacularly.

The reason likely lies in the nature of scientific writing itself. Academic prose is stylistically more homogeneous — it uses formal language, standardized structures, discipline-specific terminology, and predictable citation patterns. This uniformity makes it easier for AI to mimic convincingly, compared to, say, a blogger's unique voice or a novelist's distinctive narrative style.

📖 Source

Different Methods, Same Blind Spots

Each of the three tested detectors uses a fundamentally different methodology, yet all exhibit the same pattern of vulnerability to style imitation. Pangram relies on a neural network trained on human and machine-generated text — though its founder has candidly described the system as a "black box" since its decisions cannot be traced or explained. This lack of transparency is particularly concerning given the high-stakes environments where Pangram is deployed.

GPTZero measures how predictable word choices are within a text and how much that predictability varies, operating on the premise that language models write more uniformly than humans. Human writing naturally contains more variation in word choice, sentence length, and structural complexity, while AI tends toward statistical averages. In theory, this should make AI text detectable. In practice, when the AI is mimicking a specific author, it appears to replicate some of that natural variation.

Originality.ai searches for statistical patterns learned during training on human and AI-generated text corpora. Despite their different approaches, all three detectors show the same vulnerability. They catch text from simple prompts almost every time but miss imitations far more often. Scientific writing — the genre where AI detection probably sees the most real-world use — remains the hardest to classify correctly regardless of which detector is used.

📖 Source

Why This Study Matters: Completing the Picture

An earlier test by the Authors Guild found that Pangram and Originality.ai reliably classified human texts as human. The Epoch AI study fills in the other half of that picture: a low false-alarm rate on human writing says very little about how many AI-generated texts actually slip through undetected. A detector can be 100% accurate at identifying human text while still missing a significant proportion of AI-generated content — and that is precisely what this study reveals.

This is particularly consequential for the fields where AI detectors see the heaviest use: universities checking student essays, publishers screening submissions, journalism outlets verifying sources, and grant committees evaluating proposals. When a detector reports that a text is "likely human-written," it does not mean the text was actually written by a human — it may simply mean the AI did a good job of mimicking the expected style.

The study raises profound ethical and practical questions about how AI usage should be regulated in academic and professional writing. A 13% average miss rate means roughly one in every seven or eight AI-generated texts escapes detection entirely. In scientific writing, the figure climbs to one in four. These are not negligible error rates — they are systemic blind spots that undermine the reliability of detection as a primary enforcement mechanism.

📖 Source

Looking Ahead: The Future of AI Text Detection

The Epoch AI study serves as an important wake-up call. AI text detectors are not as reliable as previously assumed, and their effectiveness drops sharply when language models deliberately attempt to mimic an author's style. This is particularly problematic in education, where these tools are widely deployed to check student work, often with significant consequences attached to the results.

Style imitation is a readily accessible technique — it does not require sophisticated prompt engineering or specialized tools. Simply providing an AI model with a few samples of an author's writing is enough for it to produce convincingly similar text. As frontier models continue to improve, this capability will only become more refined, making detection progressively harder.

The path forward likely requires new approaches to AI text detection. Pure statistical analysis may no longer be sufficient. Future solutions might combine traditional text analysis with metadata verification, writing history tracking, and multi-modal evidence gathering. Some researchers are exploring watermarking techniques that embed undetectable signals in AI-generated text, while others advocate for a shift away from detection altogether and toward transparency requirements — mandating that AI systems disclose their involvement rather than relying on after-the-fact detection. Whatever approach wins out, one thing is clear from the Epoch AI study: the current generation of AI text detectors cannot be trusted as the sole arbiter of whether content is human-written or AI-generated, especially in scientific and academic contexts.

📖 Source