Anything Engine Optimization The rolling record of the AI-search industry.
Exterior of bright orange wall of building with door with metal casing and drawing in daylight in city street outside
Photo by rotekirsche 20 on Pexels
arXiv GEO · Aug 30 ResearchGEO

New buyer-persona corpus measures demand for AI citations

A new arXiv preprint describes PersonaGen-1M, a corpus of 1,031,732 synthetic buyer personas spanning 511 industry labels and 4 market contexts, built to study how brands enter or fail to enter the shortlists that generative engines like ChatGPT, Gemini and Perplexity name in answers. Each persona carries a commercial-intent label, a set of search queries, and a preferred-sources field the authors say can be checked against citation-provenance data. The authors built it from roughly 40 million raw persona descriptions across four public datasets, deduplicated with GPU-accelerated MinHash LSH and semantic matching.

Why it matters: It's measurement infrastructure rather than a finding — a corpus sized to let researchers check which sources generative engines cite against what buyers actually say they'd trust.

The record: ChatGPTGeminiPerplexity

Via arXiv GEO ↗

Posted to the wire August 31, 2026.