New buyer-persona corpus measures demand for AI citations
A new arXiv preprint describes PersonaGen-1M, a corpus of 1,031,732 synthetic buyer personas spanning 511 industry labels and 4 market contexts, built to study how brands enter or fail to enter the shortlists that generative engines like ChatGPT, Gemini and Perplexity name in answers. Each persona carries a commercial-intent label, a set of search queries, and a preferred-sources field the authors say can be checked against citation-provenance data. The authors built it from roughly 40 million raw persona descriptions across four public datasets, deduplicated with GPU-accelerated MinHash LSH and semantic matching.
Why it matters: It's measurement infrastructure rather than a finding — a corpus sized to let researchers check which sources generative engines cite against what buyers actually say they'd trust.
The record: ChatGPTGeminiPerplexity
Via arXiv GEO ↗
Posted to the wire August 31, 2026.