# The State of Search — what is currently true about AI search

What is currently true about AI search, claim by claim. Every claim traces to published wire items; nothing is asserted that the record has not already reported.

Status values: `settled` (the record agrees), `contested` (credible sources disagree), `moving` (actively changing), `thin` (one source, no corroboration).

Last moved: 2026-08-30

## Can I block AI training without losing search traffic?

**With Google and OpenAI, yes — both state in their own documentation that the training and search crawlers are independent. Anthropic runs separable crawlers but publishes no equivalent guarantee. In every case the block is a request rather than a lock.**

Status: `settled` · verified against the record 2026-08-27 · record last moved 2026-08-27

The three operators the question actually turns on each run a search crawler and a training crawler that can be addressed separately in robots.txt. Two of them say in public what blocking one does to the other; the third does not.

Compliance is a separate question from policy. TollBit measured roughly 15% of the AI page fetchers it tracked in the first half of 2026 reaching European publisher URLs marked disallowed, with OpenAI's ChatGPT-User among the bots most often crossing blocks aimed specifically at it. No operator publishes a compliance rate against its own declared behavior, so a block is worth logging against rather than trusting.

Since August 21, 2026 a fourth party sits in this decision for Cloudflare-proxied sites. Bot Preference Sync conditions continued access for crawlers that both search and train on four public disclosures, which converts what was an operator's private practice into something a publisher can check.

### Where it differs

- **Google** (`settled`) — Documents it. Google-Extended "does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search." — https://anythingengineoptimization.com/item/2026-08-27-no-major-ai-crawler-publicly-meets-all-four-of-cloudflare-s-conditions/
- **OpenAI** (`settled`) — Documents it. Its crawler documentation states a webmaster can allow OAI-SearchBot to appear in search results while disallowing GPTBot. — https://anythingengineoptimization.com/item/2026-08-27-no-major-ai-crawler-publicly-meets-all-four-of-cloudflare-s-conditions/
- **Anthropic** (`thin`) — Runs ClaudeBot and Claude-SearchBot as separate agents, but its published help page does not say that blocking the first preserves the second. — https://anythingengineoptimization.com/item/2026-08-27-no-major-ai-crawler-publicly-meets-all-four-of-cloudflare-s-conditions/

### What changed

- **2026-08-23** — Cloudflare's Bot Preference Sync made continued access for dual-purpose crawlers conditional on four public disclosures — honoring the no-training preference, offering an AI-summary opt-out, reporting URL-level training visibility alongside search metrics, and publicly showing the block does not hurt search. — https://anythingengineoptimization.com/item/2026-08-23-cloudflare-rewrites-robots-txt-gates-access-on-disclosure/
- **2026-08-27** — Previously: Whether any major operator met Cloudflare's conditions was unanswered by the announcement. Checked against each operator's own documentation: none meets all four — Google three, OpenAI two, Anthropic none it documents. The condition all three fail is URL-level reporting of what was made available for training. — https://anythingengineoptimization.com/item/2026-08-27-no-major-ai-crawler-publicly-meets-all-four-of-cloudflare-s-conditions/

### Still unresolved

- Whether Anthropic's separable crawlers carry the same guarantee. Its published pages do not say, and it has not been asked on the record.
- What a robots.txt block is worth per operator. TollBit's ~15% is an aggregate across tracked fetchers; nobody has published a per-crawler compliance rate.
- Which crawlers, if any, meet Cloudflare's four conditions. A press query to Cloudflare was drafted 2026-08-27 and the company has published no compliance list.

### Evidence

- Each operator's documented position on training-versus-search independence, checked the same day against their own pages. — https://anythingengineoptimization.com/item/2026-08-27-no-major-ai-crawler-publicly-meets-all-four-of-cloudflare-s-conditions/ (wire analysis of published documentation — Cloudflare has published no compliance list)
- That a declared block is not reliably observed. — https://anythingengineoptimization.com/item/2026-08-16-chatgpt-ads-expand-to-europe-as-crawler-bypasses-blocks/
- That the access decision is per-crawler rather than blanket, and that a trade body has formalized it into four verdicts. — https://anythingengineoptimization.com/item/2026-08-01-iab-australia-sorts-ai-crawlers-into-four-access-verdicts/

## Why don't AI visibility tools agree with each other?

**Because they measure different engines, query sets and windows. Two trackers reporting opposite trends for the same subject are usually both correct, and the disagreement is methodological rather than factual — which is why the fact of disagreement stays stable while every number under it moves.**

Status: `settled` · verified against the record 2026-08-30 · record last moved 2026-08-30

The clearest worked example the record holds is Reddit in August 2026. BrightEdge reported on August 16 that Reddit's AI citation surge had plateaued at a baseline several times higher than a year earlier, and that Google AI Overviews accounts for roughly 88% of Reddit's AI citation volume. Three days later a tracker reported an 86.4% collapse. Both were right.

Promptwatch's own per-engine breakdown resolved it: ChatGPT Search fell 86.4% (3.83% to 0.52%), while AI Overviews slipped 11.3% and AI Mode 30.5% over the same window. The surface carrying ~88% of the volume barely moved. A decline reported widely as Reddit disappearing from AI search was one engine changing how it assembles background queries.

A separate, single-source measurement of that query-writing step shows why the change moved so much: brands ChatGPT named in its own background search query reached the final answer 68.9% of the time, against 2.1% for brands surfacing only on retrieved pages — a roughly 33x gap — and only 3.1% of 3,554 retrieved pages earned a citation at all. Read against Promptwatch's finding, the two measurements describe the same mechanism nine days apart: the query-writing step, not the page, is what the Aug 8 change actually altered.

Engine-to-engine spread inside a single study is wide enough to swamp most between-study comparisons: Writesonic measured "ghost citation" rates from 52% (Perplexity) down to 19% (Microsoft Copilot) across seven engines and roughly 16 million brand appearances. A figure quoted without its engine is not a figure.

The model underneath moves too. Gemini 3.7 Flash rolled out in Google AI Mode on August 16, 2026, which dates any visibility study measured against the version before it.

### Still unresolved

- Nobody has published a head-to-head of the third-party measurers on the same query set. Until someone does, reconciliation is inference.
- Why ChatGPT's background query assembly changed on August 8. Promptwatch reads the forty-six-fold rise in `site:` operator use as the engine asking named sites directly; OpenAI has not commented.
- Whether query fan-out works the same way outside ChatGPT. Google describes a fan-out technique behind AI Mode, but no comparable measurement of its query-writing step has been published, so nothing here transfers to AI Overviews or AI Mode on evidence.

### Evidence

- That the two headline Reddit measurements were different engines, with per-engine numbers for all three surfaces. — https://anythingengineoptimization.com/item/2026-08-27-reddit-s-citation-drop-is-chatgpt-only-86-there-11-on-ai-overviews/ (Promptwatch study — vendor-published and called provisional by its author)
- That the query-writing step predicts final citation far better than retrieved-page content, connecting the ChatGPT query-fan-out measurement to the Aug 8 change behind the Reddit collapse. — https://anythingengineoptimization.com/item/2026-08-30-query-fan-out-decides-ai-citation-before-a-page-is-ever-read/ (wire analysis connecting two single-source, ChatGPT-only measurements)
- Where Reddit's AI citation volume actually sits (~88% Google AI Overviews), which is what makes the split decisive. — https://anythingengineoptimization.com/item/2026-08-16-brightedge-reddit-s-ai-citation-surge-has-plateaued-at-a-higher/ (BrightEdge study — vendor-published)
- How wide the engine-to-engine spread is inside one methodology. — https://anythingengineoptimization.com/item/2026-07-29-writesonic-study-finds-ai-engines-often-skip-naming-source-brands/ (Writesonic study — vendor-published)
- That the model under a surface changes, dating studies measured on the prior one. — https://anythingengineoptimization.com/item/2026-08-16-gemini-3-7-flash-rolls-out-in-google-ai-mode/
- The standing caveats — citation share is engine-specific, model-sensitive, and methodology-dependent. — /glossary/citation-share/

## What can a publisher actually control about AI answers?

**Appearance and training are separate controls and no single lever covers both. Google now offers the only control that removes a site from AI Overviews and AI Mode without touching ordinary Search — it is rolling out to a subset of accounts, and it does not cover training.**

Status: `moving` · verified against the record 2026-08-27 · record last moved 2026-08-27

Four levers exist and each has a different scope. Choosing wrongly costs more than doing nothing: three of the four take ordinary Search results with them.

The price of the one that does what publishers asked for is stated plainly by Google — opt out and "you won't receive any traffic or impressions from these features" — against AI Overviews appearing on 43% of Google searches in July 2026, up from 15% a year earlier.

A regulator is already pushing on the gap. The UK CMA's Publisher Conduct Requirement, imposed June 3, 2026 and legally binding under the digital markets regime, requires Google to let publishers opt out of training, fine-tuning **and** the grounding behind AI Overviews and AI Mode, with nine months to comply. The Search Console control answers one of those three.

Google's own guidance has not caught up with its own control. "AI features and your website" — the Search Central page telling publishers how to approach inclusion — named `nosnippet` 18 times and the Search generative AI control zero times when checked on August 27, 2026. A publisher following the documentation rather than opening Search Console would not learn the control exists.

### Where it differs

- **noindex** (`settled`) — Removes the page from Search altogether. The bluntest lever, and the only one with no partial mode.
- **nosnippet / data-nosnippet / max-snippet** (`settled`) — Limit what Search displays — and take the ordinary result snippet with them.
- **Google-Extended** (`settled`) — Governs training and grounding for Gemini rather than appearance in Search's own AI features. A robots.txt token with no reporting attached to it.
- **Search generative AI control (Search Console)** (`moving`) — Removes a site from AI Overviews, AI Mode and Discover's generative features in one to two days. Google says it "isn't used as a ranking or inclusion signal affecting other parts of Search." Rolling out to a subset of site owners; does not cover training. — https://anythingengineoptimization.com/item/2026-08-27-what-google-s-ai-overviews-opt-out-does-and-two-things-it-does-not/

### What changed

- **2026-08-27** — Previously: No control separated AI Overviews from ordinary Search: leaving one meant accepting damage somewhere else. Search Console's Search generative AI control was documented, closing the gap for appearance but not for training. Establishing this fact also unblocked the crawler-compliance story, which had stalled on exactly it. — https://anythingengineoptimization.com/item/2026-08-27-what-google-s-ai-overviews-opt-out-does-and-two-things-it-does-not/

### Still unresolved

- Whether opting out costs Top Stories placement. Google places Top Stories carousels inside AI Overviews on roughly 15.5% of US and 17.5% of UK news searches, per NewzDash's John Shehata, who called the link high-confidence; Google has not addressed it.
- Whether the control satisfies the CMA order. It covers appearance; the order covers training, fine-tuning and grounding, with nine months to comply from June 3, 2026.
- Whether the control reaches every account, and when. Google says "a subset of website owners" and has published no schedule.

### Evidence

- What the control does, what it does not cover, and the caveat Google's ring-fence language leaves open. — https://anythingengineoptimization.com/item/2026-08-27-what-google-s-ai-overviews-opt-out-does-and-two-things-it-does-not/ (rolling out to a subset of site owners)
- The Top Stories exposure a news publisher takes on by opting out. — https://anythingengineoptimization.com/item/2026-07-28-google-s-ai-opt-out-setting-may-cost-publishers-top-stories-placement/
- The binding regulatory requirement the control partially answers, and its clock. — https://anythingengineoptimization.com/item/2026-08-01-uk-s-cma-orders-google-to-give-publishers-ai-opt-out-and-attribution/

## Does llms.txt do anything?

**For Google, no — its documentation says outright that no AI text file is needed to appear in its AI features. For every other major engine there is no published answer either way, and the one large-scale measurement found 97% of llms.txt files were never read by an AI crawler.**

Status: `thin` · verified against the record 2026-08-18 · record last moved 2026-08-18

The evidence is one vendor study and one vendor's documentation. That is thin for a tactic this widely recommended, and the thinness is the finding: nobody has published a reason to make the file, and only one company has published a reason not to.

Ahrefs, across roughly 15 million data points, found 97% of sites' llms.txt files were never read by an AI crawler, and that adding schema markup produced no measurable lift in AI citations across 1,885 pages over 30 days. The same research found 88.46% of AI citations still traced to pages in the general search index — the tactic that works is the old one.

The spec itself is still moving. Jeremy Howard published a version 2 update on August 10, 2026, its first revision since the format launched in 2024, adding `rel="alternate" type="text/markdown"` and `rel="describedby"` link relations so agents can find a page's Markdown version.

### Where it differs

- **Google** (`settled`) — Says it is unnecessary: "You don't need to create new machine readable files, AI text files, or markup to appear in these features." — https://anythingengineoptimization.com/item/2026-08-01-google-guidance-machine-readable-ai-files-do-nothing-for-search/
- **OpenAI, Anthropic, Perplexity, Microsoft** (`thin`) — No published statement on whether their systems read the file. Never asked on the record.

### Still unresolved

- Whether any engine other than Google reads llms.txt. Four operators have never said, which is half a story sitting unfinished.
- Whether the v2 link relations change crawler behavior. The spec revision is three weeks old and no measurement post-dates it.

### Evidence

- The only large-scale measurement of whether the files are read at all. — https://anythingengineoptimization.com/item/2026-08-18-ahrefs-study-llms-txt-files-mostly-unread-by-ai-crawlers/ (Ahrefs study — vendor-published)
- Google's own position, stated in its documentation rather than inferred. — https://anythingengineoptimization.com/item/2026-08-01-google-guidance-machine-readable-ai-files-do-nothing-for-search/
- That the spec is under active revision, so measurements have a shelf life. — https://anythingengineoptimization.com/item/2026-08-17-llms-txt-spec-adds-link-relations-for-markdown-discovery/

## Are AI answers costing me traffic?

**It depends on what you publish, and part of the disagreement is instrumentation. Commerce reports AI referrals as growth; news and nonprofit publishers report steep declines; Google's own figure arrives without methodology. Before any of it is compared, 22.4% of AI Overview traffic is misattributed away from Organic Search.**

Status: `contested` · verified against the record 2026-08-29 · record last moved 2026-08-19

The disagreement is real and it sorts by what a site sells. Shopify told its Q2 earnings call that AI-driven traffic and orders tripled year over year, with half of AI-referred sessions landing directly on product pages — 2.5 times the rate of traditional search. Over the same period Blood Cancer UK's leukaemia information page lost 53% of its page views. Both are first-party reports from parties with no incentive to overstate against their own interest.

One widely-cited framing does not survive the record. Zero-click searches hit a record high in June 2026 — 40% of U.S. searches sent a click to the open web — but the same report put Google AI Mode at 0.13% of U.S. search-related visits, usage it described as plateaued or dipping. A surface used by roughly one in a thousand visits cannot be the cause of a decline that broad, whatever else is.

Measurement sits underneath all of it. A nine-month GA4 study of 51,200 clicks found an average 22.4% of AI Overview traffic recorded as Direct rather than Organic Search, and Similarweb found 58.8% of ChatGPT referral traffic lands on publisher homepages even though 65% of cited URLs point two or three folders deep. A citation and the session it produces are not measured in the same place, which is why two honest parties can count the same channel and disagree. See the measurement-disagreement claim.

### Where it differs

- **Google (the engine)** (`thin`) — Says AI features in Search send "billions of clicks" to websites weekly. No methodology accompanied the claim, and it has published no supporting data. — https://anythingengineoptimization.com/item/2026-07-20-google-claims-ai-search-sends-billions-of-clicks-weekly-without-the/
- **Commerce** (`settled`) — Growth. Shopify reported AI-driven traffic and orders tripled year over year in Q2, half of AI-referred sessions landing on product pages — 2.5x traditional search — and framed AI as a complement rather than a substitute. — https://anythingengineoptimization.com/item/2026-08-05-shopify-says-ai-driven-traffic-and-orders-tripled-year-over-year-in-q2/
- **News publishers** (`settled`) — Decline, and not enough leverage to act on it. Google search fell to 21% of People Inc.'s traffic from 25% the prior quarter, with core sessions down 22% year over year and Google search traffic specifically down 40% — and the publisher still declined to block, saying that would "turn off search". — https://anythingengineoptimization.com/item/2026-08-04-people-inc-won-t-block-google-s-ai-crawlers-despite-traffic-slide/
- **Nonprofit and health information** (`settled`) — The sharpest declines in the record. Blood Cancer UK's leukaemia page lost 53% of page views year over year, with 17-45% declines across its other blood cancer pages. Save the Children logged a 303% rise in AI Overview appearances while both clicks and impressions fell. — https://anythingengineoptimization.com/item/2026-08-17-uk-charities-report-steep-traffic-losses-to-google-ai-overviews/
- **The instrument** (`settled`) — Understates the channel before anyone compares it. An average 22.4% of AI Overview traffic was recorded as Direct rather than Organic Search across a nine-month, 51,200-click study — worst at 29.3% in May 2026. — https://anythingengineoptimization.com/item/2026-08-18-study-finds-ai-overview-traffic-misattributed-22-of-time/

### What changed

- **2026-07-23** — Alphabet reported Q2 Search revenue of $63.27 billion, up 17% year over year — the first deceleration after four quarters of accelerating growth. The money says the business has not broken, only slowed. — https://anythingengineoptimization.com/item/2026-07-23-alphabet-q2-search-revenue-up-17-as-growth-eases/
- **2026-08-03** — Previously: The zero-click rise was widely attributed to AI answer surfaces. The same report that recorded the zero-click record put AI Mode at 0.13% of U.S. search-related visits and plateauing, which separates the two trends rather than linking them. — https://anythingengineoptimization.com/item/2026-08-03-zero-click-google-searches-hit-a-record-high-ai-mode-use-stays-flat/

### Still unresolved

- Google holds the only complete dataset and has published no methodology behind "billions of clicks". Until it does, the engine's own position is the least checkable one in the record.
- Whether the sector split is causal or compositional. Nobody has published a study that holds content type constant and varies only AI-surface exposure.
- What share of the zero-click record predates AI answers entirely. The decline is broader and older than AI Mode's 0.13% usage can explain, and no one has decomposed it.
- Whether misattribution is worsening. The 22.4% average spans nine months with a 16.8-29.3% range and no trend line published.

### Evidence

- The engine's own claim, and that it arrived without data. — https://anythingengineoptimization.com/item/2026-07-20-google-claims-ai-search-sends-billions-of-clicks-weekly-without-the/ (vendor statement, no methodology published)
- That AI referrals read as growth in commerce, first-party. — https://anythingengineoptimization.com/item/2026-08-05-shopify-says-ai-driven-traffic-and-orders-tripled-year-over-year-in-q2/ (vendor-published, earnings call)
- That the instrument understates the channel before comparison. — https://anythingengineoptimization.com/item/2026-08-18-study-finds-ai-overview-traffic-misattributed-22-of-time/
- That AI search is layering on top of traditional search rather than replacing it, and that the cited URL and the landing page differ. — https://anythingengineoptimization.com/item/2026-07-30-similarweb-95-of-chatgpt-users-also-search-on-google/
- That crawl volume and referral value are now separately measurable, and diverge sharply. — https://anythingengineoptimization.com/item/2026-08-19-microsoft-clarity-adds-an-ai-scrape-to-referral-ratio/

## What on-page work actually moves AI citation?

**The levers that measure are the ones that were already SEO — being in the search index, ranking well, being reachable in plain HTML, and keeping pages current. The AI-specific file formats measure near zero.**

Status: `settled` · verified against the record 2026-08-29 · record last moved 2026-08-24

The single largest number in the record points back at classic search: 88.46% of AI citations traced to pages already in the general search index, and 76% of the passages Google reused 100 or more times came from pages already ranking first organically. Google's own position is the same — John Mueller said there is "nothing really special you need to do for generative AI responses in search".

That is not the same as saying nothing is specific to AI. Three levers measure, and all three are mechanical rather than semantic: whether a crawler can reach the page at all, how recently it changed, and whether the passage is shaped to be lifted. The AI crawlers are stricter than Googlebot on the first — GPTBot, ClaudeBot and Bingbot found none of a site's JavaScript-injected internal links across 41 days, and recovered 250 pages within 48 hours of the links being converted to HTML.

The tactics named after the category are the ones that do not measure. An analysis spanning roughly 15 million data points found 97% of llms.txt files were never read by an AI crawler, and adding schema markup produced no measurable citation lift across 1,885 pages over 30 days. See the llms-txt-effect claim, which holds that question's own open threads.

### Where it differs

- **Search-index presence** (`settled`) — The dominant lever. 88.46% of AI citations trace to pages in the general search index; 76% of passages reused 100+ times already ranked #1 organically. — https://anythingengineoptimization.com/item/2026-08-18-ahrefs-study-llms-txt-files-mostly-unread-by-ai-crawlers/
- **Crawlability in plain HTML** (`settled`) — Decisive, and stricter than Googlebot. Over 41 days GPTBot, ClaudeBot and Bingbot reached none of a 2,400-page site's JavaScript-linked pages; Googlebot reached 2% against GoogleOther's 48%. — https://anythingengineoptimization.com/item/2026-08-19-41-day-test-finds-ai-crawlers-miss-javascript-only-links-entirely/
- **Recency** (`settled`) — Measures, and varies by engine. Of 47,097 citations, 75% of cited pages had been updated within a year and 88% within two — Gemini 78%, ChatGPT 73%, Perplexity 65%. Last-modified date tracked citation better than original publish date (72% vs 42%). — https://anythingengineoptimization.com/item/2026-08-06-study-ties-ai-citation-likelihood-to-how-recently-a-page-was-updated/
- **Passage shape** (`settled`) — Observable in what gets reused. Across 15.7 million AI Mode citations, 80.9% of passages were cited once, while roughly 2,300 were reused 61+ times and skew short (117-word median), answer-first and self-contained. — https://anythingengineoptimization.com/item/2026-07-29-pillarbase-study-15-7-million-ai-mode-citations-show-what-google/
- **AI-specific files (llms.txt, schema)** (`thin`) — No measurable effect. 97% of llms.txt files were never read by an AI crawler, and schema markup produced no citation lift across 1,885 pages over 30 days. — https://anythingengineoptimization.com/item/2026-08-18-ahrefs-study-llms-txt-files-mostly-unread-by-ai-crawlers/

### What changed

- **2026-08-18** — Schema markup moved from an assumed AI-citation lever to a measured null: no lift across 1,885 pages over 30 days, in the same analysis that found 97% of llms.txt files unread. — https://anythingengineoptimization.com/item/2026-08-18-ahrefs-study-llms-txt-files-mostly-unread-by-ai-crawlers/
- **2026-08-24** — Google stated the position directly — "there’s nothing really special you need to do for generative AI responses in search" — framing AI Overviews and AI Mode as drawing on the standard search index. — https://anythingengineoptimization.com/item/2026-08-24-google-says-geo-needs-no-special-tactics-beyond-seo/

### Still unresolved

- Whether entity phrasing is a content lever. Google's own paper found models fail to recall 26-34% of facts they encoded and tied part of the gap to subject/object order, but the researchers stop short of calling it something a publisher can act on.
- Whether the recency effect is causal or compositional. Actively maintained pages may simply be better pages; no study holds quality constant and varies only update date.
- Whether passage shape can be induced. Pillarbase's finding is observational across published pages — nobody has run the experiment of rewriting a passage to the observed shape and measuring citation change.
- Whether serving a stripped machine-readable page to agents helps or is penalized. BrightEdge shipped Agent Edge on that premise on 2026-08-26; Perplexity blocked Time's agent-facing markdown pages on 2026-08-11. The record contains a product and a punishment and no measurement.

### Evidence

- That index presence dominates, and that the two named AI-specific tactics measure near zero. — https://anythingengineoptimization.com/item/2026-08-18-ahrefs-study-llms-txt-files-mostly-unread-by-ai-crawlers/ (vendor-published (Ahrefs), ~15 million data points)
- That AI crawlers are stricter than Googlebot on JavaScript-rendered links. — https://anythingengineoptimization.com/item/2026-08-19-41-day-test-finds-ai-crawlers-miss-javascript-only-links-entirely/ (single-site experiment, 2,400 pages over 41 days)
- That recency tracks citation likelihood, and by how much per engine. — https://anythingengineoptimization.com/item/2026-08-06-study-ties-ai-citation-likelihood-to-how-recently-a-page-was-updated/ (vendor-published (Seer Interactive), four brands)
- What shape a repeatedly-reused passage has. — https://anythingengineoptimization.com/item/2026-07-29-pillarbase-study-15-7-million-ai-mode-citations-show-what-google/
- Google's own stated position that no separate optimization target exists. — https://anythingengineoptimization.com/item/2026-08-24-google-says-geo-needs-no-special-tactics-beyond-seo/ (spokesperson statement on social, not documentation)
- That the entity-order finding is a model-recall result its own authors do not extend to publisher practice. — https://anythingengineoptimization.com/item/2026-08-17-google-research-entity-order-limits-llm-fact-recall/

## Can I buy my way into AI answers?

**You can buy a slot, but not a citation — they are separate games. Ads run on about a quarter of ChatGPT's commercial prompts and nearly one in three commercial AI Mode queries, and only 3.63% of the advertisers whose ads ran were also cited as a source in the answer above them. The ad stack is scaling faster than the reporting under it.**

Status: `moving` · verified against the record 2026-08-29 · record last moved 2026-08-24

The paid layer arrived quickly and is still moving. ChatGPT Ads launched in the US on February 9, 2026, reached nine markets by August 17 — its first Spanish- and Portuguese-language ones — and roughly 40 countries two days later when 31 European markets went live at once. OpenAI's enterprise CMO put ad revenue growth above 25% since the start of August. Ads remain limited to Free and Go plans.

Paid presence and cited presence are close to unrelated, which is the finding that matters for anyone treating ads as an AEO shortcut. Across more than 50,000 commercial prompts in 20 niches, 3.63% of advertisers running an ad were also cited as a source in the answer above it, and 14.35% of the ads shown were unrelated to the surrounding conversation. The same separation shows on Google: an analysis of AI Mode text ads found advertisers rarely among the answer's cited sources.

Two coverage figures circulate and they are not in conflict — they use different denominators. SE Ranking found ads on 25.94% of commercial prompts; Adthena, measuring all US queries rather than commercial ones, found 4.47% at an average of 1.06 ad items per response, against a 3.53-item average on Google's AI surfaces. Read the first as ad load on the queries advertisers want and the second as ad load across everything.

Reporting lags the levers throughout. Campaign-level platform targeting shipped in August into an Insights view that still groups results as Mobile and Desktop, so a web-only campaign reports entirely as Mobile. View-through conversions appear as their own column but are excluded from CPA, bidding and billing. Automated bidding became the default for new ad groups carrying no guarantee against a CPA, CPC or ROAS target. Third-party tooling is currently the only cross-platform view of what is actually running.

### Where it differs

- **ChatGPT — ad load** (`settled`) — 25.94% of more than 50,000 commercial prompts across 20 niches carried an ad; across all US queries the figure is 4.47%, averaging 1.06 ad items per response. — https://anythingengineoptimization.com/item/2026-08-11-chatgpt-ads-appear-on-a-quarter-of-commercial-prompts/
- **Google AI Mode — ad load** (`settled`) — Ads on nearly one in three commercial-keyword queries, and an average of 3.53 ad items per response across Google's AI surfaces — a heavier load than ChatGPT's. — https://anythingengineoptimization.com/item/2026-07-23-ads-appear-on-a-third-of-commercial-ai-mode-queries-analysis-finds/
- **Overlap with citation** (`settled`) — Near zero. 3.63% of advertisers whose ads ran were also cited as a source in the answer above them, and 14.35% of ads shown were unrelated to the conversation around them. — https://anythingengineoptimization.com/item/2026-08-11-chatgpt-ads-appear-on-a-quarter-of-commercial-prompts/
- **Category concentration** (`settled`) — Retail and fashion drew 39% of observed US ad placements on 24% of query volume. Logistics and home-and-garden carried the highest ad frequency, at 12.41% and 11.99%. — https://anythingengineoptimization.com/item/2026-08-24-retail-takes-39-of-chatgpt-ad-slots-on-24-of-queries/
- **Measurement** (`thin`) — Behind the controls it is meant to measure. Platform targeting reports only Mobile/Desktop; view-through conversions are excluded from CPA, bidding and billing; automated bidding is the default with no performance guarantee. — https://anythingengineoptimization.com/item/2026-08-21-chatgpt-ads-adds-platform-targeting-but-reporting-still-lags/

### What changed

- **2026-08-19** — Previously: ChatGPT Ads ran in nine markets, all English-language until August 17. 31 European markets went live at once, bringing the pilot to roughly 40 countries about six months after the February 9 US launch. — https://anythingengineoptimization.com/item/2026-08-19-chatgpt-ads-go-live-in-31-more-european-markets/
- **2026-08-21** — Automated bidding became the preselected default for new ad groups. The strategy carries no guarantee against a CPA, CPC or ROAS target, which moves spend risk onto advertisers who do not actively opt into a manual bid ceiling. — https://anythingengineoptimization.com/item/2026-08-21-chatgpt-ads-makes-automated-bidding-the-default-for-new-ad-groups/

### Still unresolved

- Whether buying an ad affects the odds of being cited at all, in either direction. The 3.63% overlap is a snapshot of co-occurrence, not a test — nobody has run the same queries with and without a live campaign.
- What ChatGPT ads cost. OpenAI publishes no benchmarks, and making an unguaranteed automated strategy the default puts spend risk on advertisers who have none.
- Whether the Agent campaign type ships. OpenAI has been testing a format that opens a Business Agent conversation instead of sending a click to the advertiser's site, built from a profile it generates by scraping that site — a different unit from an ad entirely.
- Whether Google's AI Mode ad load holds. The 3.53-item average is one measurement window and Google has published nothing of its own.

### Evidence

- Ad load on commercial prompts, the share of ads unrelated to the conversation, and the overlap between advertising and being cited. — https://anythingengineoptimization.com/item/2026-08-11-chatgpt-ads-appear-on-a-quarter-of-commercial-prompts/ (vendor-published (SE Ranking), 50,000+ prompts across 20 niches)
- Ad load across all US queries, items per response, and category concentration. — https://anythingengineoptimization.com/item/2026-08-24-retail-takes-39-of-chatgpt-ad-slots-on-24-of-queries/ (vendor-published (Adthena), ~850,000 US and UK queries, March-May 2026)
- That advertisers are rarely among the cited sources on Google's AI Mode. — https://anythingengineoptimization.com/item/2026-07-23-ads-appear-on-a-third-of-commercial-ai-mode-queries-analysis-finds/
- The pace and shape of geographic rollout, and OpenAI's own revenue-growth figure. — https://anythingengineoptimization.com/item/2026-08-19-chatgpt-ads-go-live-in-31-more-european-markets/ (vendor statement for the revenue figure)
- That the reporting layer lags the targeting controls it is meant to measure. — https://anythingengineoptimization.com/item/2026-08-21-chatgpt-ads-adds-platform-targeting-but-reporting-still-lags/
- That a third-party view of live creatives and landing pages now exists across ChatGPT and Google's AI surfaces. — https://anythingengineoptimization.com/item/2026-08-16-similarweb-launches-ad-tracking-for-chatgpt-and-ai-mode/ (vendor-published (Similarweb))

## Will the courts stop AI engines from using my content?

**Not so far. The one concluded case produced a price rather than a prohibition — roughly $3,000 per work — while 137 AI copyright suits remain pending nationwide, 24 against OpenAI alone. Courts are reaching opposite results on near-identical theories, and a First Amendment defense has appeared from three sets of defendants in six months.**

Status: `moving` · verified against the record 2026-08-29 · record last moved 2026-08-28

One case has produced a number. The Bartz v. Anthropic book-piracy class settlement became effective August 20, 2026, starting a 28-day clock for payouts of roughly $3,000 per work, due September 17. Two appeals filed since challenge only attorneys' fee awards and by the agreement's terms do not affect that date. It is the first figure in the record that prices training on pirated books, and it arrived as a settlement rather than a ruling — so it binds nobody else.

The theories are splitting rather than converging, sometimes within days. A federal judge dismissed most of Google's DMCA anti-circumvention claims against SerpApi on the reasoning that Google's anti-bot system protects ad revenue rather than a copyrighted work; days later a different judge let Reddit's near-identical DMCA claims against Perplexity and SerpApi proceed. Contributory infringement has moved the other way and closed: the New York Times, Daily News and Ziff Davis had those claims dismissed with prejudice against Microsoft after the Supreme Court's Cox Communications decision foreclosed the theory.

Agentic browsing got its first appellate answer, and it favored the agent. The Ninth Circuit reversed an injunction barring Perplexity's Comet browser from Amazon, finding Comet unlikely to violate the Computer Fraud and Abuse Act because it acts on the user's direction rather than accessing servers on its own.

A defense posture is forming in parallel. Musk, Tesla and Warner Bros. Discovery raised a First Amendment defense in an AI-copyright dispute in March 2026; Anthropic added one in the Gilbert case on August 20; Perplexity raised one in Reddit's DMCA suit on August 28. Separately, the pressure that has moved fastest is not copyright but competition — Judge Mehta, who already found Google's search business an illegal monopoly, said in a hearing that Google's use of publisher content in AI Overviews "seems really unfair".

### Where it differs

- **Copyright — training data** (`moving`) — 137 suits pending nationwide, 24 against OpenAI. One settlement concluded, at roughly $3,000 per work. Suits now reach executives personally: Sony Music names Dario Amodei and Benjamin Mann. — https://anythingengineoptimization.com/item/2026-08-22-wikihow-sues-openai-over-ai-training-query-use/
- **DMCA anti-circumvention** (`contested`) — Split. Google's claims against SerpApi were mostly dismissed with prejudice; Reddit's near-identical claims against Perplexity and SerpApi survived dismissal days later before a different judge. — https://anythingengineoptimization.com/item/2026-07-31-reddit-s-dmca-suit-against-perplexity-serpapi-survives-dismissal-bid/
- **Contributory infringement** (`settled`) — Closing. Dismissed with prejudice for the NYT, Daily News and Ziff Davis against Microsoft after the Supreme Court's Cox Communications decision foreclosed the material-contribution theory. — https://anythingengineoptimization.com/item/2026-08-08-nyt-ziff-davis-lose-bid-to-revive-copyright-claims-against-microsoft/
- **CFAA — agentic browsing** (`settled`) — Favors the agent, on the only appellate answer so far. The Ninth Circuit found Perplexity's Comet unlikely to violate the CFAA because it acts on user direction rather than accessing servers on its own. — https://anythingengineoptimization.com/item/2026-08-05-9th-circuit-reverses-perplexity-cfaa-injunction-in-amazon-suit/
- **Antitrust and competition** (`moving`) — The fastest-moving track. Judge Mehta called Google's use of publisher content in AI Overviews "seems really unfair" while hearing Penske Media's suit; nearly 300 French newspapers filed a competition complaint over AI Overviews in France. — https://anythingengineoptimization.com/item/2026-08-26-judge-mehta-calls-google-s-ai-overviews-unfair-to-publishers/
- **First Amendment defense** (`thin`) — An emerging posture rather than a tested one. Three sets of defendants have raised it since March 2026 — Musk/Tesla/Warner Bros. Discovery, then Anthropic, then Perplexity. No court has ruled on it in this context. — https://anythingengineoptimization.com/item/2026-08-28-perplexity-raises-first-amendment-defense-in-reddit-s-dmca-suit/

### What changed

- **2026-08-08** — Previously: Publishers were pursuing contributory-infringement theories against platform intermediaries. The Supreme Court's Cox Communications decision foreclosed the material-contribution theory; the NYT, Daily News and Ziff Davis claims against Microsoft were dismissed with prejudice. — https://anythingengineoptimization.com/item/2026-08-08-nyt-ziff-davis-lose-bid-to-revive-copyright-claims-against-microsoft/
- **2026-08-28** — Previously: No AI training case had produced a price. The Bartz settlement became effective August 20, starting a 28-day payout clock at roughly $3,000 per work, due September 17 — the first concrete cost in the record for training on pirated books. — https://anythingengineoptimization.com/item/2026-08-28-bartz-v-anthropic-settlement-effective-payouts-due-sept-17/

### Still unresolved

- Whether roughly $3,000 per work becomes a benchmark or stays a one-off. Bartz settled rather than ruled, so it binds nobody — and no other case has produced a number.
- Whether training on lawfully-acquired books is fair use. The Mosaic/Databricks summary-judgment hearing is set for October 30, 2026, with publisher trade groups filing against it — the nearest thing in the record to a scheduled answer.
- Whether executives can be held personally liable. Sony Music names Amodei and Mann, and the authors' suit was amended to add Amodei; no court has tested it.
- Whether Judge Mehta applies the monopoly finding to AI Overviews. A hearing remark is not a ruling, and Google's motion to dismiss Penske Media is still pending.
- Why two judges reached opposite results on the same DMCA theory within days. Neither ruling has been reconciled with the other, and the split is unresolved.

### Evidence

- The pending caseload — 137 AI copyright suits nationwide, 24 against OpenAI. — https://anythingengineoptimization.com/item/2026-08-22-wikihow-sues-openai-over-ai-training-query-use/ (single-source count, from one litigation tracker)
- The first concluded price, and the date payouts are due. — https://anythingengineoptimization.com/item/2026-08-28-bartz-v-anthropic-settlement-effective-payouts-due-sept-17/
- That near-identical DMCA theories are producing opposite results. — https://anythingengineoptimization.com/item/2026-07-31-reddit-s-dmca-suit-against-perplexity-serpapi-survives-dismissal-bid/
- That the contributory-infringement route is foreclosed, and why. — https://anythingengineoptimization.com/item/2026-08-08-nyt-ziff-davis-lose-bid-to-revive-copyright-claims-against-microsoft/
- The only appellate answer so far on whether an agentic browser violates the CFAA. — https://anythingengineoptimization.com/item/2026-08-05-9th-circuit-reverses-perplexity-cfaa-injunction-in-amazon-suit/
- That the judge who found Google an illegal monopoly has questioned AI Overviews' fairness to publishers. — https://anythingengineoptimization.com/item/2026-08-26-judge-mehta-calls-google-s-ai-overviews-unfair-to-publishers/ (remark at a hearing, not a ruling)
- That a fair-use summary-judgment hearing is scheduled for October 30, 2026. — https://anythingengineoptimization.com/item/2026-08-28-news-media-alliance-opposes-databricks-summary-judgment-bid/

## I'm getting cited. Am I getting recommended?

**Not necessarily — being cited, being named and being recommended are three different outcomes with different drivers, and the record separates them. Across 1,094 ChatGPT categories the most-cited domain was the most-mentioned brand only 20.8% of the time. Citation measures whose page an engine used; recommendation tracks how well the engine already knows your brand.**

Status: `settled` · verified against the record 2026-08-25 · record last moved 2026-08-25

The three outcomes come apart at every step. An engine can lift a passage from your page and name a competitor in the sentence it supports; it can name you without linking you; it can recommend you having never cited you at all. A dashboard reporting one number cannot tell you which of the three you bought.

What moves recommendation is mostly not on-page work. Models searched for brands they already knew 3.2 times more often than unfamiliar ones — 55.7% versus 17.4% across 66 buyer prompts — and 63% of brand-specific searches surfaced one of each model's five most-familiar brands. Where competing products were otherwise identical, the well-known brand was recommended 100% of the time.

Two findings look like they disagree about reviews and do not. A market census of 4,776 venues found star rating had no effect on whether a venue appeared at all; a controlled study found a rival needed less than a 0.1-star edge to overturn an incumbent's recommendation. Those are different stages of the same funnel — rating does not get you into the candidate set, and decides between candidates once you are in it. Having a business website nearly doubled the odds of the first.

The uncomfortable finding for this discipline sits in the same study: when every brand adopted the same marketing tactics, the incumbent's advantage collapsed from a payoff of +0.802 to +0.007. Tactics that everyone runs stop being an edge and become the price of entry — while brands that opted out received no recommendations at all.

### Where it differs

- **Cited** (`settled`) — Your URL appears as a source under the answer. This is what citation-share tools measure, and it is the only one of the three most of them measure. — https://anythingengineoptimization.com/item/2026-07-23-ahrefs-ranks-the-50-most-cited-domains-across-four-ai-surfaces/
- **Named** (`settled`) — Your brand appears in the answer text, with or without a link. Engines skip naming the source brand between 19% (Microsoft Copilot) and 52% (Perplexity) of the time, so citation and naming diverge by engine before anything else does. — https://anythingengineoptimization.com/item/2026-07-29-writesonic-study-finds-ai-engines-often-skip-naming-source-brands/
- **Recommended** (`settled`) — The engine puts you forward as the answer. Driven mostly by prior brand familiarity, then by reviews at the margin — not by the page-level work that moves citation. — https://anythingengineoptimization.com/item/2026-08-25-llms-favor-well-known-brands-until-reviews-tip-the-scale/

### Still unresolved

- Whether the 20.8% mismatch holds outside ChatGPT and Semrush's citation data. It is one analysis of one engine's estimated demand, and nobody has replicated it elsewhere.
- Whether anything a site controls moves recommendation once familiarity is accounted for. The evidence that tactics equalize to +0.007 when everyone runs them is a single controlled study on one product category.
- Whether recommendation varies by country and prompt language the way practitioners report. No published measurement covers it — the record has nothing.

### Evidence

- That the most-cited domain and the most-mentioned brand are usually not the same, with the size of the gap. — https://anythingengineoptimization.com/item/2026-07-29-analysis-most-ai-search-demand-still-has-no-clear-chatgpt-category/
- That naming diverges from citation, and by how much per engine. — https://anythingengineoptimization.com/item/2026-07-29-writesonic-study-finds-ai-engines-often-skip-naming-source-brands/ (Writesonic study — vendor-published)
- That prior familiarity drives whether a model looks for you at all. — https://anythingengineoptimization.com/item/2026-07-30-study-finds-ai-models-search-more-often-for-brands-they-already-know/ (geoSurge study — vendor-published)
- The incumbent advantage, the review threshold that overturns it, and the collapse when every brand runs the same tactics. — https://anythingengineoptimization.com/item/2026-08-25-llms-favor-well-known-brands-until-reviews-tip-the-scale/
- How much of a real market never gets recommended at all, and that a website moves it where star rating does not. — https://anythingengineoptimization.com/item/2026-08-23-ai-assistants-miss-85-6-of-bali-restaurants-in-market-census/ (Norly Research study — vendor-published, funded and conducted by an AI-visibility vendor)

Canonical: https://anythingengineoptimization.com/state/
From Anything Engine Optimization (AEO Wire) — https://anythingengineoptimization.com/
