A site the AI engines had never cited, not once in 125 measurement runs, got four of five new studies cited within weeks — three at position one. This is the full log: method, controls, costs, and what failed.
The vocabulary in this entry — Citation Vacuum, Retrieval Floor — is defined in Entry 01.
Find the specific questions an AI engine hedges on ("no publicly available data breaks this down…"), then publish the first authoritative, first-party answer as a question-shaped study. On a zero-authority domain, 4 of 5 fills were cited on Perplexity within 2–4 weeks, 3 at position 1, while untouched controls stayed at zero. The same play on an authority site converted to +544% search impressions. Mining a batch of 20 candidate questions cost $0.42; each study cost $0–8 in data plus a day of work.
In Entry 01 I defined a Citation Vacuum as a high-demand question AI engines cannot crisply answer, detectable because the model hedges. This entry is the evidence behind that definition: what happened when I started deliberately finding those hedges and publishing the first real answer into them, on two very different sites. Both sites are mine, which is the point. LumenGEO is a GEO measurement tool I run that had, when this started, essentially zero domain authority. GrantCompass is the grant-discovery site with real traffic behind the case study. One sat below the Retrieval Floor, one comfortably above it.
First, the baseline, stated carefully, because it frames everything. Before July 2026, LumenGEO had been cited in zero of 125 measurement runs across its entire history. On its own that is unremarkable: a young, unknown domain should measure zero, and I would not write an article about it. What makes the zero interesting is what failed to move it. I had spent months running the standard on-page playbook on that site: answer-first rewrites, comparison tables, expert-quote blocks, heading restructures. Those format experiments produced zero citations across 133 runs. Not "disappointing results". Zero. That is what the Retrieval Floor looks like from underneath: the formatting advice is not wrong exactly, it is multiplied by zero, because the engines are not retrieving the site at all. The question this experiment answers is what, if anything, a site in that position can actually do.
Ask an AI engine a specific, stat-shaped question nobody has crisply answered, and it tells on its own sources. The tell is gap-acknowledgment language, and once you know to look for it you see it everywhere:
That hedge is the vacuum signal. It is not a refusal, and it is not a low citation count; an engine with web search on always finds some adjacent source to point at. The detector has to read what the model says about its own sources. Two things made detection reliable in practice. First, ChatGPT is the right detector and Perplexity is the wrong one: Perplexity always synthesizes a confident, citation-dense answer whether or not good data exists, so it cannot reveal a gap. Second, the scoring discriminates: a mining batch of 20 candidate questions cost $0.42, correctly rejected 18 of the 20, and did not false-positive on any of the three deliberately well-answered control questions I mixed in.
Each study cost between nothing and about $8 of data collection, plus a day of build. The fills were real research: original data captures with saved analysis scripts, so every published number is recomputable. That is not virtue signalling, it is the moat — the vacuum exists because nobody has published a real answer, and only a real answer fills it.
Each dot is one Perplexity measurement run at the latest checkpoint; filled dots are runs where the study was cited.
| Study (the filled vacuum) | Live for | Runs cited | Result |
|---|---|---|---|
| Which industries does Perplexity cite Reddit for? S1-001 · deployed June 16 |
4 weeks | 10/15 (was 6/15 at wk 2 — held and improved) |
cited · rising |
| What domain authority do ChatGPT-cited sites have? S1-002 · deployed June 16 |
4 weeks | 4/10 (0 at wk 2 — slow fuse, same destination) |
position 1 |
| What % of Perplexity citations are paywalled? S1-003 · deployed July 3 |
13 days | 10/15 |
position 1 |
| How fresh are the sources AI Overviews cite? S1-004 · deployed July 3 |
13 days | 15/15 — every single run |
position 1 |
| Untouched control pages no changes, measured on the same cadence |
throughout | 0 — never cited |
flat zero |
Four of five fills cited, three at position one, within two to four weeks, on a domain whose all-time citation count had been zero. The controls never moved. As far as I know this is the fastest zero-to-cited path anyone has published with controls attached, and the studies themselves are public: every row in that table links to the actual page Perplexity adopted.
A log that only records wins is marketing. Three findings from the same experiment cut the other way, and they are load-bearing.
The obvious objection: fine on a tiny domain, does it work where it matters? GrantCompass ran the same play in parallel, seeded from questions its audience actually asks. Its first fill went live June 17. By the day-14 read, and again at day 30, four of five target queries were citing GrantCompass in ChatGPT, the fill page was cited on three of three samples of its headline query, and the page's Search Console impressions were up 544%. On an authority property the citation converts to actual search traffic — the thing the zero-authority site could prove mechanically but never collect on.
Two protocol notes from that run, for anyone replicating. GrantCompass was measured on ChatGPT rather than Perplexity and read cleanly anyway; I put that down to it being an established domain (the citation-hiding problem bites hardest at zero authority). And its first fill was an enrichment of an existing page rather than a dedicated question-shaped study; it won regardless, but round two there uses dedicated question-shaped pages, because the shape evidence from LumenGEO is one-directional.
Pull candidate questions from your own Search Console queries and your audience's actual asks. This was my mistake to learn from: mine for any vacuum and you will win low-demand ones that prove the mechanism and pay nothing. Demand-weight from the start.
Web search on, batch the questions, score gap-acknowledgment language. Include two or three well-answered control questions per batch; a good detector rejects most candidates and never flags a control. Budget: under a dollar per batch of 20.
Title the page as the literal question. Original first-party data, named author, honest methodology, limitations included. Only fill vacuums your data can honestly support — a fabricated fill poisons the one query space where an engine was still willing to trust somebody new.
Five runs × three query variants per checkpoint on Perplexity (the engine that still shows its sources reliably). A single check is a coin flip. Keep untouched controls in the loop the whole time.
Decide before deploying: not cited by week 8 = failed, and gets reported. If fills fail broadly, the mechanism does not transfer to your vertical — that is a finding too, and publishing it is worth more than pretending.
If your site is already being retrieved: mine your vertical's hedges, fill the two or three where you genuinely hold first-party data, and let the week-8 line decide. If your site is below the Retrieval Floor: on my data this is the single most effective thing you can do — not because the citations bring traffic yet, but because it is the one lever that measurably worked from zero while everything the checklists recommend measured exactly nothing.
Methodology. LumenGEO figures: 147 Perplexity measurement runs (DataForSEO; 5 runs × 3 query variants per study per checkpoint at baseline / week 2 / week 4) against untouched controls, June 16 – July 16, 2026; pre-experiment site baseline 0 citations in 125 all-time runs; on-page format experiments 0 citations in 133 runs. Mining detector: ChatGPT with web search, gap-language scoring; batch of 20 = $0.42, 18/20 rejected, 0/3 control false-positives. GrantCompass figures: parallel run, 3 ChatGPT samples per query at day 14 and day 30; impressions from Google Search Console. Every published number is recomputable from saved capture and analysis scripts. LumenGEO and GrantCompass are both my properties.
Citation Vacuum, Retrieval Floor, and five more coined terms this entry puts to work — defined first, here.
Learn the terms →The authority property in this experiment, and the first-party data behind it — charts, confounders, honest caveats.
See the data →87 experiments and what the numbers actually show — including the format tactics that measured nothing here.
Read the findings →The nine moves that earn citations across every engine — vacuum-filling is the sharpest of them for new sites.
Read the playbook →Next in the log: the attribution-failure piece — why 76 of my 87 experiments were confounded, and what that means for everyone's tactic claims, mine included. Meanwhile, the tool this experiment was built on is LumenGEO (mine, disclosed), and if you want an operator who publishes his controls, let's talk.
First published August 6, 2026 · All figures from the July 2026 measurement captures; methodology above.