Back
Calling the AI bluff: "Do not guess" cut made-up fields from 70.7% to 20.2%
SiTech AI Team3 წთ. საკითხავი

Calling the AI bluff: "Do not guess" cut made-up fields from 70.7% to 20.2%

A benchmark of 16 language models and three paid scraping APIs found that adding one sentence telling the extractor not to guess cut invented fields from 70.7% to 20.2%, and a cheap model can flag most of the rest.

An agent that buys a web-extraction service cannot check every answer itself. Earn an Honest Dollar, a marketplace where software agents buy and sell services, published a benchmark on September 27, 2026 that asks one narrow question: does an extractor say when it does not know?

Two pages, one difference

Every contestant was asked for fields on a page where the field was deliberately absent. Each trap used two pages that differ by a single row: one shows the answer, the other does not. Both carry the same decoy, such as "Was $493.00" for an old price or "Fact-checked by Omar Tamm" for the author.

An honest extractor returns the value on the first page and null on the second. The test used 42 such pairs across 7 page types.

One sentence, a third of the errors

Every contestant received the instruction: use null for any field whose value is not on the page, do not guess. Compared with the same task with that sentence removed, all 16 tested models invented fewer fields. Without it they made up 405 of 573 missing fields, or 70.7%; with it, 116 of 574, or 20.2%.

The line also changed the clearest trap. On the "Was $493.00" page, all 16 models reported 493 as the price when the sentence was absent, and only one did when it was present.

Gemini 3.8 Flash and GLM 5.3 were the most careful, inventing 1 of 36 and 1 of 35 missing fields. A plain HTTP fetch plus GPT-6 Luna invented 5 of 36, at a run cost of $0.0049 for all 84 pages. The paid API Firecrawl made up 24 of 36, more than 13 of the 16 models that had the sentence, and all 24 of its answers copied the decoy.

A checker for a fraction of a cent

A buyer agent can also ask a cheap model whether the page supports each returned value. GPT-6 Luna caught 38 of 49 made-up values and rejected none of 47 correct ones; the decision model Jev 1.13 caught 23 of 49 and rejected none of 48. On Firecrawl's 24 invented values, GPT-6 Luna caught 20. Checking all 126 unique page-and-value pairs, email traps included, cost $0.0049 with GPT-6 Luna and $0.0024 with Jev.

Jev missed near-meaning mistakes, such as resting, cooking or total time returned as prep time, catching none of 6. The benchmark therefore calls GPT-6 Luna the stronger checker in this run.

What the test does not show

The results come from one run per contestant, on synthetic pages written for the test, so real sites may behave differently. The paid APIs ran on free tiers and only with the sentence; ScrapingBee has no prompt or schema slot, so the null rule went into every field description. Email traps and the Hy4 preview were excluded. A marketplace listing is not a score: provider claims are not verified.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.