How Will AI Find It Works
We chose these 11 dimensions from the patterns that recur in the published GEO research — from Princeton's generative-engine benchmark to the 2026 analyses of what ChatGPT actually cites — and from what a page has to contain when the writing is all a model sees. The scoring is built on a simple premise — an assumption, not a measured finding: AI search engines treat content the same way a careful researcher does. They look for specificity, authority, and structure. They skip generic filler the same way you do.
Detection Risk Signals
These dimensions measure patterns that make prose interchangeable — writing that says what thousands of other pages already say, and so gives an AI search engine no reason to cite this one. They are not a test of who wrote it. Lower scores are better — a low score means the content avoids these patterns.
Repetitive Patterns
Count how many paragraphs start with “Furthermore,” “Additionally,” or “Moreover.” Repetitive phrases, parallel list structures, and formulaic transitions are the signature of content built from a template, whoever filled it in. A page built that way says what every other page built from the same template says, and a model assembling an answer has no reason to quote this copy of it. Editorial judgment, not a measured finding.
Generic Phrasing
“Leverage,” “cutting-edge,” “streamline,” “game-changer” — these words add no information. They could describe any product on any page, so there is nothing in a sentence built from them that a model can extract, attribute, and repeat. Writing that makes specific, bounded claims is quotable; corporate filler is not. The fix is almost always the same: replace the adjective with a number.
Passive Voice
Excessive passive voice (“was implemented,” “has been shown,” “can be achieved”) hides the subject: something was implemented, by someone, with some result. Active, direct voice produces definitive statements with an agent and an outcome. Search Engine Land's 2026 ChatGPT citation analysis found cited passages nearly 2x more likely to use definitive language than hedged framing. “We reduced churn by 14%” is quotable. “Churn was reduced” is not.
Content Quality Signals
These dimensions measure whether the prose reads as one person working through an idea rather than a template being filled — rhythm, vocabulary, and the connective tissue between paragraphs. They are editorial judgments about what makes writing quotable, not a test of who wrote it. Higher scores are better.
Sentence Variety
Writing that works through an idea has rhythm. Short sentences land. Longer, more complex constructions develop an idea across clauses. Uniform mid-length sentences — 15 to 25 words, similar construction, paragraph after paragraph — read as filled in rather than thought through, and a passage that reads that way is rarely the one worth lifting into an answer. No study in the Research Foundation measures this directly; it is our editorial judgment about quotable prose.
Lexical Diversity
Vocabulary range relative to text length. If the same key terms appear repeatedly without synonyms or varied phrasing, the writer has one word for the thing because they have one thought about it. A writer with genuine expertise reaches for more precise terms — “correlation” instead of “connection,” “deprioritize” instead of “skip.” Diverse vocabulary reflects depth. Narrow vocabulary reflects thin material. Editorial judgment, not a measured finding.
Natural Flow
Here's the difference between a mechanical transition and a real one. “Furthermore, it is important to consider...” connects nothing. “That's the obvious part. The less obvious part is...” builds an argument, and an argument is what gets quoted. This dimension measures how organically paragraphs and ideas connect — through contextual bridges, callbacks to earlier points, and implicit logical connections rather than connector words. Editorial judgment, not a measured finding.
Authority & Citability Signals
These dimensions measure whether the page has something worth quoting: a voice, a number, a structure, a judgment, a date. Higher scores are better.
Personal Elements
First-person experience, named examples, and original voice. Content with first-hand authority carries information not duplicated across thousands of other pages, which is the only reason to cite this page rather than any other on the topic. Content without original voice is interchangeable — and interchangeable content gives an AI engine no reason to cite yours over anyone else's.
Concrete Data
Specific numbers, statistics, dates, and cited sources. Vague claims are uncitable — an AI search engine cannot reference “many studies show” or “experts agree.” A specific, attributed claim — “Semrush's analysis of 10M+ keywords found AI Overviews on 15.69% of queries in November 2025” — gives the model something it can extract, attribute, and repeat. In the Princeton GEO study, adding statistics and citing sources were among the strongest visibility methods tested.
Structural Clarity
Clear headings, definitive statements, and extractable answer blocks. 44% of ChatGPT citations come from the first 30% of a page (Search Engine Land, 2026). A key point buried deep in the content is competing against the top of the page, which draws nearly half of the citations on its own.
Emotional Intelligence
Content that shows judgment, weighs options, or acknowledges complexity demonstrates human expertise. A blog post that says “this approach works perfectly for every team” is less authoritative than one that says “this approach worked for our 12-person team but broke down when we tried it with 50 — here's why.” A balanced perspective is more useful to the reader asking the question, which is who a model assembling an answer is standing in for. Editorial judgment, not a measured finding.
Cultural Anchoring
Real-world references, temporal markers, and current context. Content that could have been written at any point in any year gives a reader no way to tell whether it is current. Anchored content — referencing specific market conditions, recent research, or current events — shows the writer was working from the present, which is what a reader asking about the present needs. Ahrefs found AI-cited pages measurably fresher than organic results, and Seer found three quarters of cited pages updated within the year. Both measured page dates rather than the writing, so this dimension is consistent with that research rather than measured by it.
Research Foundation
The dimensions are informed by published GEO research where it exists. These are the primary sources we rely on, each linked. A dimension without a listed source rests on the mechanism described in its own entry, not on a measured finding.
GEO methods boosted visibility by up to 40% in generative engine responses. Cite Sources, Quotation Addition, and Statistics Addition were the strongest of the methods tested. Two caveats: the benchmark ranked five sources per query, so the zero-sum setup amplifies relative gains, and the study reflects 2024 model behavior, so absolute percentages do not transfer to current engines.
Informs: Concrete Data
Semrush AI Overviews Study (2025)
Across 10M+ keywords tracked from January to November 2025, 91.3% of the queries that triggered an AI Overview in January were informational. By October that share had fallen to 57.1% as commercial, transactional, and navigational triggers expanded. AI Overviews appeared on 15.69% of tracked queries in November 2025.
Informs: Context only. Shows how much informational search now routes through AI answers; informs no single dimension.
Search Engine Land ChatGPT Citation Analysis (Feb 2026)
44% of ChatGPT citations come from the first 30% of a page. Cited passages are nearly 2x more likely to use definitive language.
Informs: Structural Clarity. Consistent with Passive Voice and Generic Phrasing, which it does not measure directly.
Writesonic AI Crawler Study (March 2026)
Across 62 tests on six AI assistants, ChatGPT, Claude, and Gemini fetched raw HTML and parsed it as text, while Copilot, DeepSeek, and Grok executed JavaScript for a short window. Meta descriptions, Open Graph tags, and JSON-LD reached none of the six. Within that test, several major AI assistants primarily consume transformed HTML body content, so the writing itself carries the signal.
Informs: Structural Clarity
Ahrefs AI Citation Freshness Study (July 2025)
Across 16.975 million cited URLs from ChatGPT, Perplexity, Gemini, Copilot, Google AI Overviews, and organic Google results, AI-cited content averaged 1,064 days old against 1,432 days for organic top-10 results, or 25.7% fresher. ChatGPT showed the strongest preference. The average cited page was still nearly three years old, and the study measured page dates, not the writing.
Informs: Consistent with Cultural Anchoring, which scores temporal references in the text. The study measured publish and update dates, not the text.
Seer Interactive Content Recency Study (July 2026)
Across 7,683 pages carrying 47,097 citations from ChatGPT, Gemini, and Perplexity between March and June 2026, 75% of cited pages had been updated within the last year. Where both dates were readable, 72% had been updated in the past year but only 42% were published in it, so the freshness being rewarded came from maintaining older pages. Four brands, dated by structured signals.
Informs: Consistent with Cultural Anchoring. Like Ahrefs, it measured page dates rather than temporal references in the text.
How the Score Is Computed
Everything about the method is published here except the prompt text itself.
Model and settings
Scoring runs on Anthropic's Claude Sonnet 4.6 (claude-sonnet-4-6) at temperature 0, with a 4,096-token output limit. Each scan is a single pass: one scoring call, one response. If a response fails to parse, the call is retried once and the retry replaces it. Nothing is averaged across runs.
The formula
The model scores each of the 11 dimensions from 1 to 10. The three Detection Risk dimensions are inverted (10 minus the raw score) so that every dimension contributes in the same direction, higher is better. Those eleven display values are the ones on the radar and in the table. The overall score is their sum, normalized:
overall = min(100, round(sum(display values) / 91 * 100))
The maximum possible sum is 107, so any sum of 91 or more scores 100. The divisor was recalibrated from 110 to 91 in April 2026 so that strong, sourced, first-person writing can reach the top band.
Server-side verification, and how often it fires
The model also reports its own overall score. We do not trust it. The server recomputes the score from the 11 dimension values and, when the model's number differs from the recomputed one by more than 2 points, displays the recomputed value.
That happens on 96% of scans. Of 149 scans since the current formula shipped on 2026-04-07, the server overrode the model on 143, and on zero of the 149 did the model's number match the recomputed one exactly. The model's own number ran 10.3 points below the recomputed one on average: it tends to normalize against the theoretical maximum rather than follow the formula. So this check is not an occasional safeguard. The model produces the eleven dimension scores. The server produces the number.
The other 4%, 6 of 149, fell within the 2-point tolerance and kept the model's number. On those scans the displayed score can sit up to 2 points off its own dimension sum. Figures reproduced by a committed query in the project repository (scripts/queries/override-rate.sql), run on 2026-09-04 with our own six verification scans excluded, and the result is committed beside it. The rate will be re-run and this paragraph updated as scans accumulate.
Run-to-run variance, measured
Temperature 0 does not make the model deterministic, so we measured the spread rather than estimating it. On 2026-09-05 we scored five committed test pieces spanning the range, five times each, calling the API directly with the production prompt, model, and temperature. Across the 55 dimension-and-piece pairs, 30 returned the identical value all five times, 54 stayed within 1 point, and the largest spread on any single dimension was 2 points. Mean per-dimension standard deviation was 0.20.
The overall score moves more than any one dimension, because several 1-point shifts add up: across the five runs it ranged 2 to 5 points per piece, 3.8 on average. So a re-score of unchanged text can move the number by a few points while nothing about the writing has changed, and a change of that size is not evidence of anything. Look at which dimensions moved and by how much.
The script, the five pieces, and the raw results are committed in the project repository (scripts/variance-test.ts, scripts/variance-results/), so the measurement can be re-run and will be, when the model or the prompt changes.
What the research supports, and what it does not
The Research Foundation above is real but uneven, and it is fairer to say where it reaches than to imply it covers all 11 dimensions equally.
- Directly measured by a cited study (2): Concrete Data (Princeton) and Structural Clarity (Search Engine Land, Writesonic).
- Consistent with the research, not measured by it (4): Passive Voice and Generic Phrasing (the Search Engine Land study measured definitive versus hedged language, not voice or filler), Personal Elements (Princeton tested adding quotations, not original voice), and Cultural Anchoring (Ahrefs and Seer measured page dates, not temporal references in the text).
- Editorial judgment about what makes prose quotable (5): Repetitive Patterns, Sentence Variety, Lexical Diversity, Natural Flow, and Emotional Intelligence. No cited study measures these. They are in the rubric because we think they are right, and each entry above says so.
None of the 11 has been validated against actual citation outcomes. That study has not been run, by us or by anyone we can cite.
Honest Scoring
Will AI Find It is a diagnostic tool, not a prediction engine. The 11 dimensions are informed by published GEO research where it exists and by editorial judgment where it does not; the section above says which is which. They measure writing patterns, not citation outcomes.
AI search engine behavior is not standardized across providers: Ahrefs found the preference for fresh content differed by engine, and Writesonic found the assistants fetch pages differently. Will AI Find It scores the writing quality patterns that research suggests these systems weight — it does not guarantee citation by any specific engine.
Last updated: September 2026