Question Headlines Didn’t Win More AI Citations—A Pre-Registered Test Found the Difference Too Small to Matter

Question Headlines Didn’t Win More AI Citations—A Pre-Registered Test Found the Difference Too Small to Matter
Sponsored

Question-style headlines are often recommended for generative engine optimization because they appear to match the way people prompt AI assistants. A small pre-registered experiment from AskedAbout found no meaningful evidence that this formatting choice increased AI citations.

Across the posts that qualified for the experiment’s final read, question headlines were cited by an average of 0.5 AI engines out of four, compared with 0.4 for statement headlines. The 0.1 difference was far below the threshold the researchers had defined in advance for treating the format as a useful winner, and the entire excess came from a single Gemini result.

The most important result is therefore not that questions “won” by 0.1. It is that the pre-registered AskedAbout test concluded the result was inconclusive and stopped on the date specified before the experiment began.

The experiment enrolled 22 posts, but only 11 entered the final comparison

AskedAbout enrolled 22 ordinary daily posts between August 24 and September 11, 2026. Eleven were assigned question-style titles and eleven statement-style titles.

Assignment followed the enrollment sequence rather than an editor choosing whichever format seemed more appropriate after seeing the story. Even sequence numbers received the question treatment and odd sequence numbers received the statement treatment. The newsroom knew the assigned arm before writing the headline and URL.

Not all 22 posts were old enough to enter the pre-registered final read on September 13. The scoring rule used each post’s second eligible Sunday measurement, producing six scored question posts and five scored statement posts.

That leaves an effective final comparison of only 11 articles—an important limitation when interpreting a difference as small as 0.1 citing engines per post.

The test measured citations across four AI engines

Every Sunday at 12:37 UTC, AskedAbout sent each enrolled article’s associated question to ChatGPT, Perplexity, Gemini and Claude through their APIs.

Each engine received three samples for each post. An engine-post cell counted as cited if any of the three generated answers cited askedabout.com.

The primary metric was therefore not total citation count. It was the number of AI engines, out of four, that cited each post, averaged across the posts in each treatment arm.

That design asks a useful practical question: does changing the title shape make a new article more likely to appear across multiple AI systems?

Question titles averaged 0.5 engines; statements averaged 0.4

At the final September 13 read, the six scored question posts averaged 0.5 citing engines out of four. The five scored statement posts averaged 0.4.

Read superficially, the question format finished slightly ahead.

But the experiment had defined its decision rule before any enrolled post existed. For the question style to receive credit, it needed to outperform the statement style by at least 0.5 engine-cells per post, with the advantage appearing on at least two engines.

The observed advantage was only 0.1 and appeared on one engine.

The instrument therefore returned the verdict: “INCONCLUSIVE — reported, not iterated.”

The entire difference came from one Gemini citation

The engine-level results show how fragile the headline difference was.

ChatGPT cited two of the six scored question posts and two of the five scored statement posts. Perplexity cited none of the scored posts in either arm. Claude also cited none.

Gemini cited one question post and zero statement posts.

That single Gemini cell accounts for the question arm’s numerical advantage. There was no broad pattern in which question headlines consistently performed better across the four systems.

This is precisely why the pre-registered rule required an advantage to appear on at least two engines before the newsroom would adopt the style as a citation tactic.

This was not a conventional statistical-significance test

It is tempting to describe the outcome as “not statistically significant,” but that wording would overstate the methodology.

AskedAbout did not report a p-value, confidence interval or conventional null-hypothesis significance test for the headline comparison.

Instead, it used a pre-committed practical decision rule. A question-title advantage of at least 0.5 citing engines per post across at least two engines would count as a useful win. Anything smaller was defined in advance as inconclusive.

The result failed that threshold decisively.

The defensible conclusion is therefore that the experiment found no meaningful advantage under its pre-registered rule—not that it statistically proved question headlines have zero effect everywhere.

Pre-registration prevented the most interesting post from deciding the story

One question-style article was cited by two of four engines at its scored second-Sunday measurement, the best result in the experiment.

Viewed in isolation, that article could easily become a case study claiming question headlines improve AI visibility.

But another statement-style post had also reached two engines on an earlier measurement before settling at one on its scored Sunday.

The pre-registration prevented either anecdote from redefining the experiment after the results were visible. The metric, timing, sample threshold and possible verdicts had been specified on August 23, before the first post was enrolled.

That methodological discipline is arguably more informative than the 0.5-versus-0.4 result itself.

Most posts were never cited at all

The low baseline citation rate overwhelmed the headline-format comparison.

Twenty enrolled posts had received at least one Sunday measurement by the final read. Fourteen of those 20 were never cited by any of the four AI engines on any measured Sunday.

Every measured post scored two or fewer engines out of four on every Sunday.

Across the September 13 run, only eight of 80 engine-post cells cited AskedAbout. At the individual sample level, 12 of 240 generated answers cited the site.

For a small domain publishing new articles, simply getting into an AI engine’s source set appeared to be a much larger problem than whether the H1 ended with a question mark.

Perplexity and Claude produced no scored citations

Neither Perplexity nor Claude cited any of the scored articles in the final comparison.

That matters because a headline optimization cannot demonstrate cross-engine robustness when two of the four systems never produce a positive cell.

The question-versus-statement comparison effectively depended on ChatGPT and Gemini, and ChatGPT showed no directional advantage for question titles.

AskedAbout’s own conclusion reflects that sparse environment: the 0.1 difference is small relative to the much larger challenge of earning citations at all.

The experiment changed both headline shape and URL shape

Another methodological detail limits how narrowly the result can be interpreted.

The treatment was defined as title and URL shape. Question posts used an interrogative H1 with a question mark and a slug beginning with an interrogative token. Statement posts used a number-led declarative sentence and did not carry those question markers.

The server enforced consistency between the assigned arm and the published H1 and slug.

That means the experiment cannot isolate whether any observed effect came from the visible headline, the URL wording or their combination.

For editorial decision-making, that may be acceptable because title and slug often change together. For a mechanistic SEO claim about H1 syntax alone, it is a limitation.

The articles were not matched for topic or quality

The experiment alternated assignments across ordinary daily stories rather than creating matched pairs about the same subject.

AskedAbout explicitly acknowledges that variation between stories may be larger than the tiny difference between treatment arms.

One article may address a topic that AI systems frequently retrieve from the site, while another may answer a question for which the domain has little authority or demand. Editorial quality, source competition, freshness and subject matter can all influence whether a page is cited.

With six question posts and five statement posts in the scored comparison, those uncontrolled differences can easily dominate a 0.1-cell gap.

The test asked each article’s own question

Every post was measured using the verbatim user question under which it had been enrolled.

This creates a clear relationship between article and prompt, but it does not measure citation performance across a demand-weighted set of real user prompts.

A page that receives zero citations might have an ineffective headline. It might also answer a question that few users ask, compete against much stronger sources or simply fail to enter retrieval for reasons unrelated to title syntax.

The experiment cannot separate those possibilities.

It tests whether question-shaped titles helped these particular AskedAbout posts when their associated questions were put to four engines.

Question headlines still have editorial uses

An inconclusive citation experiment does not make question headlines bad.

A direct question can communicate intent clearly, align naturally with a reader’s problem and create an intuitive article structure. For some topics, “How does X work?” is simply a better headline than an awkward declarative alternative.

The result argues against a narrower claim: publishers should not expect question syntax by itself to create a reliable AI citation advantage.

AskedAbout’s newsroom consequently returned title shape to editorial judgment after the pre-registered read rather than continuing to force the question format.

Semantic clarity may matter more than punctuation

The GEO industry often converts plausible retrieval theories into rigid formatting advice. Question headlines are a good example.

The theory is intuitive: users prompt AI systems with questions, so a page whose headline exactly mirrors a question should be easier for a model to match and cite.

But retrieval systems do not evaluate headlines in isolation. They can use body text, entities, passages, links, source reputation, freshness and other signals when deciding which documents to retrieve or cite.

A clear declarative headline can communicate the same topic and intent as a question headline without sacrificing semantic relevance.

The AskedAbout test does not prove those other factors dominate, but it provides no evidence that punctuation and interrogative structure deserve special treatment.

The result challenges checklist-style GEO advice

Many AI optimization recommendations are difficult to test because generative outputs vary between runs and citation rates can be low.

That environment encourages retrospective success stories: change a headline, receive a citation, then attribute the citation to the change.

A pre-registered experiment is valuable because it establishes what will count as success before the outcome is known.

Here, the question arm technically finished ahead, but the experiment did not allow the researchers to call that tiny difference a win after seeing it.

For GEO practitioners, that is a useful standard. Optimization claims should survive a decision rule stronger than “one page improved after we changed it.”

Small sites may face a retrieval problem before a formatting problem

The population numbers point toward a broader hypothesis.

If 14 of 20 measured posts were never cited by any engine, improving citation visibility may require first understanding why the domain or page is not entering the source set.

That could involve authority, indexing, topical relevance, link relationships, source diversity, content uniqueness or simply the engine’s preference for other documents.

Headline shape becomes a second-order question if the page is rarely retrieved in the first place.

The experiment cannot identify the dominant cause, but its low citation floor makes that diagnostic priority difficult to ignore.

Engine-specific behavior complicates universal GEO rules

The result also reinforces a recurring problem in AI search research: different engines behave differently.

ChatGPT generated the majority of positive scored cells in this experiment. Gemini produced the one cell separating the two headline arms. Perplexity and Claude produced none.

A tactic that appears useful on one platform can disappear when evaluated across several systems.

That makes universal recommendations such as “AI prefers question headlines” especially difficult to justify without larger cross-engine samples.

Optimization strategies should be tested against the actual platforms that matter to a business rather than inferred from one model.

The experiment is too small to prove equivalence

It would be equally wrong to reverse the conclusion and say question and statement headlines are definitively identical for AI citations.

Only 11 posts entered the final scored comparison, all from one small website. Topics and article quality were not matched. Citation rates were sparse, and two engines produced no positive scored cells.

A larger experiment across established domains and matched content could find a real effect that this test was unable to detect.

The current evidence simply does not justify treating question formatting as a proven citation tactic.

The strongest finding is the absence of a useful signal

GEO practitioners often want a simple rule: write the title as a question and improve the chance of being cited.

This test did not deliver that rule.

Question posts averaged 0.5 citing engines out of four; statement posts averaged 0.4. The difference was one-tenth of an engine per post, confined to a single Gemini citation and far below the threshold specified before the experiment began.

Meanwhile, 14 of the 20 measured posts were never cited by any engine at all, and neither Perplexity nor Claude cited a scored article in either treatment arm.

The experiment is small and cannot establish a universal null effect. But its practical message is clear: there is currently no good reason to turn every headline into a question merely because an AI citation checklist says it should help.

Write the headline that communicates the story most clearly. Until larger controlled evidence says otherwise, citation visibility appears to have bigger problems to solve than whether the title ends with a question mark.

0%