# New buyer-persona corpus measures demand for AI citations

Published: 2026-09-01T02:19:51.474Z · Source: arXiv GEO (https://arxiv.org/abs/2608.30023v1)
Source date: 2026-08-30
Entities: chatgpt, gemini, perplexity

A new arXiv preprint describes PersonaGen-1M, a corpus of 1,031,732 synthetic buyer personas spanning 511 industry labels and 4 market contexts, built to study how brands enter or fail to enter the shortlists that generative engines like ChatGPT, Gemini and Perplexity name in answers. Each persona carries a commercial-intent label, a set of search queries, and a preferred-sources field the authors say can be checked against citation-provenance data. The authors built it from roughly 40 million raw persona descriptions across four public datasets, deduplicated with GPU-accelerated MinHash LSH and semantic matching.

Why it matters: It's measurement infrastructure rather than a finding — a corpus sized to let researchers check which sources generative engines cite against what buyers actually say they'd trust.

## What this answers

**What is a demand-side corpus for generative engine optimization?**

PersonaGen-1M is a corpus of 1,031,732 synthetic buyer personas across 511 industries, each carrying a commercial-intent label, a set of search queries, and a preferred-sources field that can be checked against which sources generative engines actually cite.

**Do most AI shopping queries have commercial intent?**

No, per the corpus's intent labels: 78.3% of buyer personas are tagged informational, versus 17.4% commercial and 4.3% transactional.


Canonical: https://anythingengineoptimization.com/item/2026-09-01-new-buyer-persona-corpus-measures-demand-for-ai-citations/
From Anything Engine Optimization (AEO Wire) — https://anythingengineoptimization.com/ · Standards: https://anythingengineoptimization.com/standards/
