AI SEARCH OPTIMIZATION

Feature Story
The Single Biggest Lever in AI Search. And Why Most People Waste It.
Okay, I need to talk about original research. Specifically, I need to talk about how every piece of AEO advice you've ever read tells you to "publish original data" like it's some kind of magic incantation, and how that advice is simultaneously completely right and almost entirely useless.
Because here's the thing: original research is the single most powerful lever for getting cited by AI systems. And the evidence for that claim is now coming from so many directions that arguing against it is getting genuinely difficult.
The Case Is Basically Closed
Three independent studies, different teams, different methodologies, same conclusion.
The first peer-reviewed academic research on generative engine optimisation — out of Princeton and Georgia Tech, published at ACM KDD 2024 — tested nine content optimisation strategies across 10,000 queries and ten search engines. They tried everything: more words, better structure, more authoritative language. The single most effective technique was adding statistics to content. Specific, citable numbers. That alone improved AI visibility by 41 per cent. Meanwhile, simply increasing word count — the strategy most content teams default to — produced zero measurable improvement. Not "marginal." Zero. The signal is data density, not volume. Which is a beautiful irony given how many "comprehensive guides" are basically 4,000 words of nothing wearing a trench coat.
An information gain study took a different angle entirely — scoring 150 top-ranking Google pages across 50 keywords and 10 verticals, measuring how much each page contributed beyond the rest of its ranking cohort. Original data correlated with that information gain score more strongly than any other page-level trait, including length. Pages with one or fewer unique data points averaged a score of 40.2. Pages with 15 or more averaged 62.1, climbing steadily at every step in between. The relationship was basically linear, which in content research is the equivalent of the data screaming at you.
Then came the citation dataset that really made me sit up. An analysis of 301 live pages that AI systems cited across 316 unique prompts and seven verticals, carrying 1,075 citations between them. Only eight of those 301 pages — 2.7 per cent — qualified as genuine primary research where the data and methodology actually lived on the page itself. Those eight pages earned 90 citations. Primary research averaged 11.3 citations per page. Everything else averaged 3.4. A first-party research page was 3.3 times as citation-dense as a non-primary one.
The logic behind this advantage is structural. AI systems don't browse the web curating interesting facts like some kind of digital librarian with too much free time. They respond to queries. When a model needs a specific data point to answer a question, it needs a source that contains that data point. If the number exists only on your page — because you collected it, ran the test, have the customer base — your page is the one the system has to reach for. There's no alternative. And here's the kicker: top organic results typically carry only four unique data points on average. The bar to outperform the field is sitting on the floor and most people are still stepping over it.
The evidence is settled. But — and this is the bit that keeps nagging at me — most original research still fails to earn a single citation. The gap between "has original data" and "gets cited by ChatGPT" is enormous, and almost nobody is talking about what actually sits in that gap.
So let's talk about it.
It's Not About Having Data. It's About the Shape of the Data.
If the conclusion were simply "publish original data and win," this would be a very short newsletter and I could go do something useful with my evening. But the citation data tells a much more specific — and honestly, slightly annoying — story.
A dataset of 301 pages cited by AI systems found only eight that qualified as genuine primary research. Those eight earned 90 citations between them. Sounds great. Except 75 of those 90 came from a single content format: benchmarks that answer buying comparisons. One page — a cloud data warehouse benchmark published in 2022 — accounted for 44 citations on its own. Nearly half of every primary-research citation in the entire dataset. From one page. Still earning citations in 2026.
(I had to read that three times. I was not hallucinating.)
Strip the benchmark cluster out and first-party research barely registers at all.
AI doesn't reward original data by default. It rewards a particular shape of original data — the benchmark that answers a measurable comparison. Which is fastest. Which is cheapest at scale. Which performs best under a specific workload. In verticals without a clear benchmark format — B2B SaaS, education, professional services — no primary-research page earned a citation at all. Not because original data didn't exist, but because nobody had packaged it into a comparison a model could extract and use.
The retrieval logic makes this obvious once you see it. When someone asks "which project management tool is fastest for enterprise teams," the system needs a source that answers that comparison directly. A benchmark page fits the retrieval slot. An explainer with proprietary numbers buried in paragraph seven does not. (And I say this as someone who has definitely buried good data in paragraph seven. Multiple times. Recently.)
What to Actually Build
That dominant benchmark page is worth studying not because it's flashy — it's not — but because everything that makes it work is craft, not data. And once you see the pattern, it's repeatable.
Here's what citation-earning pages consistently do, and the order matters more than you'd think:
Lead with the comparison result. The headline finding — X is fastest, Y is cheapest at scale — goes in the first third of the page. Broader data shows 44 per cent of ChatGPT citations come from the first 30 per cent of a document, with retrieval probability dropping roughly 2.5 times after that point. If your conclusion sits after two thousand words of context-setting, the page may never surface for the query it was built to answer. Front-load the finding. Save the methodology for the people who want it. Serve both audiences, in the right order.
Make the comparison explicit. A table comparing named options on named specifications is the format AI reaches for on "which is best" prompts. If your data supports a comparison but the page doesn't present it that way, the retrieval system may not recognise it as a benchmark source. Structure is signal. Don't make the model work to find the answer.
Show your methodology. Sample size, time window, what was measured, how. This isn't academic box-ticking — it's what makes a number citable rather than dismissible. A model assessing whether to use a data point is, at some level, assessing credibility. No visible method, no reason to treat it as authoritative. A boxed methodology section near the top takes twenty minutes to write and is the difference between "interesting content" and "citable source."
Keep the URL alive. One canonical page, not migrated or renamed with every redesign cycle. Of 365 cited URLs in the dataset, 64 were dead or broken — taking 203 citations down with them. A fifth of all citations, lost to what is essentially digital negligence. A citation earned this quarter only compounds if the page is still live next quarter. This is not complicated. It is apparently very difficult.
Don't gate it. If your research sits behind a lead-capture form, AI systems can't access it. You've optimised for zero visibility. Congratulations. Put the data on the page. Collect leads somewhere else.
The Mistakes That Kill Most Research Pages
Knowing what to build is half of it. The other half is not sabotaging yourself — which, based on the data, is where most brands quietly fail.
The most common citation killers aren't about data quality. They're about packaging. Brands bury proprietary numbers in narrative prose instead of surfacing them in scannable, extractable formats. They skip the methodology entirely, turning a citable statistic into an unsourced claim. They move the URL during a site redesign and never set up a redirect. They publish the research as a PDF download instead of an indexable page.
Every one of these is fixable. None of them require better data. They require treating the research page as infrastructure — something that needs to stay live, stay accessible, and stay structured — rather than a one-off content play.
The Door That's Standing Wide Open
Here's the bit that actually excites me, and I realise I should probably be more cynical about it but I genuinely can't help myself.
In several categories right now, AI systems are answering comparison queries with listicles, product pages, and generic explainers — not because those are adequate sources, but because nobody has published a proper benchmark. The buyer question exists. The AI prompt exists. The clean, citable answer does not.
If you're a SaaS platform with implementation speed metrics, a law firm with case resolution data, a logistics company with delivery performance benchmarks, a consultancy with client outcome numbers, or frankly any business sitting on proprietary performance data — this is a structural gap in the market. AI systems are answering commercial comparison queries right now with sources that are materially worse than what you could produce, because nobody in your category has built the benchmark page.
The bar is not "publish original data." The bar is "publish a benchmark that answers a buying comparison, with visible methodology, in a structure AI can parse, at a URL that doesn't move." That's more specific, more demanding, and considerably more valuable than the generic instruction to create original content.
The Bottom Line
Original research is the single strongest content-level play in AI search. But the brands that will actually capture the value aren't the ones with the most data. They're the ones who understand that the data is the starting point — and that everything between the number and the citation is where the real work lives.
Which, if you think about it, is kind of how expertise has always worked. Having the information was never the hard part. Knowing how to present it so other people can actually use it? That's the whole game. Always has been. We've just got machines grading the homework now.
Did someone forward you this newsletter? Subscribe here:
Behind The Writing
ABOUT THE WRITER

Jo Lambadjieva is an entrepreneur and AI expert in the e-commerce industry. She is the founder and CEO of Amazing Wave, an agency specializing in AI-driven solutions for e-commerce businesses. With over 13 years of experience in digital marketing, agency work, and e-commerce, Joanna has established herself as a thought leader in integrating AI technologies for business growth.
