Users increasingly skip the ten blue links and ask ChatGPT, Perplexity, or Google’s AI Overviews directly. Every answer those systems generate cites sources — and being one of them is the new front page position. This is GEO: Generative Engine Optimization. This checklist is what we test for in our own live GEO auditor — every item below maps to something a machine can check about your site today.
How AI search actually selects sources
Next-generation search runs on retrieval: the system searches, then a language model synthesizes an answer from the results it found. Two consequences follow:
- Ranking still matters. If a page doesn’t surface in retrieval, the model never sees it. GEO builds on top of SEO; it doesn’t replace it.
- Parseability decides citation. Among retrieved pages, the model quotes the one that answers the intent most directly and unambiguously. Clean structure beats beautiful prose.
[user query] ──> [AI search: Perplexity / ChatGPT / AI Overviews]
│
▼
[retrieval: classic search under the hood]
│
▼
[your page: answer capsule + fact density + Schema.org] ──> [cited answer]
1. The answer capsule: answer first, elaborate second
The single highest-leverage edit: a direct, complete answer to the page’s core intent within the first 100 words. Generative engines extract passages — if your answer is buried under three paragraphs of throat-clearing, a competitor’s page gets quoted instead.
Practical rules:
- Every H2 section starts with its conclusion in the first sentence.
- Numbers, dates, and names live in those first sentences, not in a footnote.
- Kill the filler: high fact density is a measurable advantage. Models (and skimmers) favor pages where every paragraph adds information.
2. Schema.org: make meaning unambiguous
Structured data removes guesswork about what your page is. In JSON-LD:
| Schema type | What it gives AI search | Priority |
|---|---|---|
Article / TechArticle | Authorship, dates, expertise signals (E-E-A-T) | High |
FAQPage | Question-answer pairs ready for extraction and voice | Critical |
Organization | Canonical facts about your company (name, contacts, IDs) | High |
LocalBusiness | Coordinates and hours for local intent | High (local) |
Speakable | Passages marked for voice assistants | Medium |
The FAQ type earns its “critical” label honestly: it hands the model pre-shaped question-answer units — exactly the format answer engines assemble.
3. Let the crawlers in (robots.txt + llms.txt)
A surprising number of sites silently block the very bots they want to be cited by. Check your robots.txt:
# AI crawlers you want citing you
User-agent: GPTBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: ClaudeBot
Allow: /
Add an llms.txt file — a Markdown map of your content for LLM crawlers. One honest caveat: Google has stated it does not use llms.txt, and it is not a ranking factor. It costs an hour to maintain and helps other AI services navigate large catalogs, so we ship it — as one detail of the strategy, never as the strategy.
Speed belongs to this layer too: keep TTFB low (under ~200 ms is the comfort zone) because retrieval systems operate on strict budgets and won’t wait for a slow origin.
4. Local and entity signals
If you serve a region, consistency decides what models “know” about you:
- NAP consistency — name, address, phone identical across your site, Google Business Profile, and directories. Mismatched data becomes conflicting facts in the knowledge layer.
- Entity completeness — an About page with real people, an Organization schema with identifiers, pages that stay on-topic. Topical focus beats encyclopedic sprawl for citation.
Common failure modes
Not appearing in AI Overviews. Cause: watery lead, no answer capsule. Fix: rewrite the opening around a direct factual answer.
Fresh content invisible to AI search. Cause: AI crawlers blocked in robots.txt or pages too slow to fetch. Fix: Allow directives + TTFB under control.
The model describes your company incorrectly. Cause: no authoritative, structured facts about the entity. Fix: Organization schema + consistent external sources.
The measurable version of this checklist
Everything above is checkable by machine — that’s the point. We run a live auditor on the automata.sale homepage that scores any URL for exactly these factors: structured data, llms.txt, AI crawler access, meta hygiene, and speed. The homepage you’d test it on scores 94/A, and the article you’re reading ships at 100/100 — the same gate our publishing pipeline enforces on every new page.
Run your site through it (free, no sign-up): automata.sale.
Summary
| Layer | Action | Effect |
|---|---|---|
| Content | Answer capsule in the first 100 words | Extractable, citable passages |
| Semantics | FAQPage + Article/Organization JSON-LD | Unambiguous meaning |
| Access | robots.txt Allow for AI bots, llms.txt, low TTFB | Crawlable and fetchable |
| Entity | NAP consistency, topical focus | Trustworthy knowledge-layer facts |
GEO is not a trick that games the models — it’s removing every reason a system might prefer someone else’s page as its source. Start with the free audit, fix what it flags, and give the retrieval layer 30–45 days to notice.
Want this done for your site end-to-end — audit, markup, content restructure? Get in touch.
📞 +7 (906) 311-77-69 · ✉ hello@automata.sale · 💬 Telegram: @automatasale · 🌐 automata.sale
Sole proprietor Evgeny Uryadov (Automata) · Tax ID 645112058391 · Reg. 312645301900058