What the GEO Study Actually Measured
A 2024 research paper tested content edits across 10,000 queries and found which ones increased citation by generative engines. Here is what it found, and what it did not.
Paste content or fetch a URL and get six sub-scores for the factors that measurably increase how often generative engines cite a source. Three come from the peer-reviewed GEO study (Aggarwal et al., KDD 2024) that tested content edits across 10,000 queries. Three come from how retrieval systems select passages. Every single sub-score displays its formula, the numbers it counted and the evidence it counted them from, because a score you can't check is a score you can't act on.
The Content Citability Scorer rates content on six factors: fact density, source citations and quotability (measured by the Aggarwal et al. GEO study at KDD 2024), plus answer-first openings, question headings and self-containment. Every sub-score prints its formula, its counted inputs and the evidence it counted, so the arithmetic is checkable rather than a black-box number. Runs entirely in your browser.
The AI-visibility category has a credibility problem: dozens of tools output a confident number with no stated basis, which makes the number useless the moment you disagree with it. This scorer takes the opposite approach. Fact density is statistics per 100 words, and it lists the statistics it found. Quotability is the share of sentences that are complete, standalone, liftable claims of 8–40 words, and it shows you three of them. Source citations counts outbound reference links per 500 words. Those three factors aren't invented: the GEO study from Princeton, Georgia Tech, IIT Delhi and the Allen Institute measured which content edits actually increased visibility in generative-engine answers, and adding statistics, quotations and cited sources were the top performers. The other three sub-scores are structural, reflecting how retrieval works rather than what the study measured: whether the page opens with a direct answer in 40 words or fewer, whether section headings are phrased the way people ask questions, and whether sections stand alone instead of leaning on 'as mentioned above'. Each sub-score carries a weight, the overall is their weighted mean, and the whole calculation is printed on the page. Paste plain prose, Markdown or full HTML, or fetch a live URL: analysis runs in your browser either way.
Last updated
Plain prose, Markdown or full page HTML all work. Fetching a URL enables the outbound-citation and heading factors, which plain text can't provide.
A weighted mean of six sub-scores, with a one-line summary of what the number means in practice.
Every row shows its formula, its counted inputs and its evidence. Start with the lowest-scoring row that carries the highest weight.
Add real numbers, link the sources you're paraphrasing, rewrite the opening to answer in 40 words, name the subject instead of writing 'it'.
Paste the revision. The sub-scores move as the counts move, so you can see exactly what your edit bought.
Three sub-scores implement factors the KDD 2024 GEO paper measured across 10,000 queries. The paper is linked from every row that uses it, so you can read the source rather than trust the tool.
"Statistics per 100 words ÷ 2 × 100, capped at 100". Every sub-score states its arithmetic, its target and the raw counts it used. Nothing is hidden behind a proprietary model.
It shows the statistics it found, the quotable sentences it identified, the sections that failed self-containment and the exact phrase that failed them.
Below-target rows explain what to change and why that change is the one the evidence supports, not generic 'write better content' filler.
Paste a draft before publishing, or fetch a live URL to score heading structure and outbound citations too. The tool says plainly which factors it can't measure in plain-text mode.
Scoring runs entirely in your browser. Nothing you paste is uploaded, logged or stored, so embargoed and internal content is safe to test.
"Add statistics" is advice. "You have 0.4 statistics per 100 words, target is 2" is a task. The counted numbers convert vague optimization into a finite editing job.
When someone asks why the score is 61, you can show the six formulas and the counted inputs. Black-box scores collapse under that question. This one is designed for it.
Throat-clearing intros, label-style headings and 'as mentioned above' feel like normal prose and are quietly fatal once a page is read as isolated fragments.
The highest-leverage moment is the draft. Paste it, fix the two lowest sub-scores, ship a page that was citable on day one.
From quick one-off fixes to daily workflows, see how people put this tool to use.
The quotability sub-score identifies whether your sentences are liftable claims or subordinate clauses that only work in context: the difference between being quoted and being skimmed.
Traditional SEO can rank thin, hedged content. Generative engines quote specific claims. This scorer measures the gap between the two.
Docs that assume the reader arrived at page one fail hard when a model retrieves section seven alone. Self-containment scoring finds those sections.
Score before and after: the counted deltas (statistics added, quotable sentences created) are a deliverable a client can verify themselves.
Boundaries stated plainly, with the right tool for each neighbouring job.
Accepts URL, TXT, MD and HTML, and produces Report and Share link, all processed locally in your browser.
See how the content citability scorer fits into a step-by-step journey with related tools.
See your page the way a RAG system does: split into retrieval chunks and read one at a time, with no surrounding context. Flags every fragment that collapses alone: pronoun openers, 'as mentioned above', sections that never name their subject.
Check whether every section of your page opens with a self-contained answer in 40 words or fewer. Flags throat-clearing intros, buried answers and openings that depend on missing context, with a rewrite target for each failure.
One URL in, one AI-visibility report out: crawler access with Cloudflare detection, JS rendering, citability scored with published methodology, retrieval chunks, answer snippets and schema. Each finding linked to the free tool that fixes it.
See your page the way AI crawlers see it: without JavaScript. Fetch the raw HTML, count the words that survive, detect client-rendering framework markers, and diff against the rendered DOM to list exactly which content is invisible to GPTBot, ClaudeBot and PerplexityBot.
Paste your description as it appears on your site, GitHub, LinkedIn, Product Hunt, Crunchbase and G2, and see where your own profiles contradict each other: name-spelling drift, conflicting numbers, conflicting years, profiles too thin to corroborate anything.
Check whether GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, Bingbot and every other AI crawler can read your site. Real RFC 9309 robots.txt matching with the deciding rule quoted, plus Cloudflare default-blocking detection. Free, instant, no signup.
Hit a snag? Here are quick fixes for the issues people run into most.
Plain-text mode can't see links, so that factor is scored neutrally at 50 and says so in its row. Fetch the URL or paste full HTML to measure real outbound citations.
They may well be good headings for a human reader. The factor measures a specific retrieval behavior: queries are questions, and question-phrased headings match them literally. It's weighted lowest of the six for exactly this reason: treat it as a suggestion, not a mandate.
Some topics genuinely resist quantification, and the score will reflect that honestly rather than pretend otherwise. Look for the specifics you do have: versions, dates, durations, thresholds, limits, counts. If there are truly none, prioritize the quotability and self-containment factors instead.
Open the sub-scores and compare their counted inputs: the arithmetic is visible on both, so the difference is always locatable. That's the whole reason the formulas are printed.
The two highest-weighted factors are fact density and quotability: a single paragraph of concrete numbers usually moves the overall score more than a full rewrite.
One claim per sentence, 8–40 words, no leading pronoun: that shape is what makes a sentence liftable verbatim by a model.
Link outward to primary sources (specs, standards, papers, benchmark pages). The GEO study found citing sources among the highest-lift edits, and outbound links cost you nothing.
Delete 'as mentioned above' everywhere. Retrieved fragments have no above, and the phrase converts a good section into an unusable one.
Score the page you most want cited, not your homepage. Homepages are rarely retrieved as answers to anything.
Recent updates and improvements to the content citability scorer.
Initial release: six weighted sub-scores (fact density, source citations, quotability, opening answer, question headings, self-containment) with per-row formulas, counted inputs and quoted evidence. Plain-text / Markdown / HTML / URL input. Methodology section citing the KDD 2024 GEO study.
All scoring runs in your browser. Pasted content is never uploaded, logged or stored. URL fetches use a stateless proxy that fetches, returns and discards. Draft and embargoed content is safe here.
Free and instant, Content Citability Scorer processes your file securely and removes it right after.