Back
New paper identifies AI-written web content from structure, not wording
SiTech AI Team2 წთ. საკითხავი

New paper identifies AI-written web content from structure, not wording

A paper on arXiv reports that AI-written commercial blog posts can be told apart from human writing by structure alone: 187 structural features reach 98.0 macro-F1, and the result barely changes when every AI post is reworded.

A new arXiv paper asks whether AI-written commercial web content can be identified from structure rather than wording. “SlopShape” by Jochen Madler of Sitefire reports that an instrument of 214 features, 187 of them structural, separates AI-generated company blog posts from human ones at 98.0 macro-F1.

What was measured

The study transfers StoryScope, earlier work on structural signatures in AI fiction, to commercial content. It pairs 2,250 human blog posts from 268 company domains — pre-ChatGPT Wayback snapshots from 2008 to 2022 — with 11,250 AI mirrors written by five frontier models from the same briefs: GPT-5.4, Claude Sonnet 4.6, Gemini 3 Flash, DeepSeek V3.2 and Kimi K2.5.

Each post was scored against an 11-dimension commercial schema covering purpose, structure, evidence, voice and page format. An LLM applied it; a human gold-annotation session reached 0.928 agreement between annotators and 0.946 between humans and the model.

Study design: the measurement pipeline

Structure beats style and survives rewording

Structural features alone reached 98.0 macro-F1 on companies never seen in training; style features alone managed 88.1. Word-level detectors were perfect on unedited AI text, but that advantage does not survive rewriting: after every AI post was reworded by its own model, the structural model still scored 98.1 and assigned 79.3 percent of posts to the correct one of six sources, against 16.7 percent chance.

Top 20 structural features

The tidy, self-announcing post

Ten features carry most of the signal, and they describe a recognisable shape: the payoff is promised in the title, the thesis and structure are announced before the first section, the voice is that of an editorial explainer, and the close restates the thesis. Human posts lean the other way — no announced flow, no escalation of stakes, no closing restatement — yet a classifier using only those ten features still reaches 93.5 macro-F1.

Human posts also sit in rarer regions of structural space: their mean rarity percentile is 0.838 against 0.435 for AI posts, and among the rarest one percent, 149 are human and 4 are AI.

Release and limits

The pipeline, instrument, prompts, code and aggregate artifacts are public on GitHub. The paper notes a software- and US-heavy corpus, mirrors written from briefs with less context than the original authors had, and a rewording test that covers self-rewriting rather than commercial “humanizer” tools. The author also discloses that he operates Sitefire, a commercial GEO product.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.