# Envoyix: full content
> Envoyix audits any public URL for Generative Engine Optimization: how readily ChatGPT, Perplexity, Claude and Gemini can crawl, understand and cite it.
This file contains the complete Markdown source of every guide on https://envoyix.vercel.app, for language models and other automated readers. The canonical HTML version of each guide is linked in its heading.
Last generated: 2026-09-29
---
# What is GEO (Generative Engine Optimization)?
Source: https://envoyix.vercel.app/learn/what-is-geo
Author: The Envoyix team
Published: 2026-01-15 · Updated: 2026-09-25
> Generative Engine Optimization (GEO) is the practice of structuring web content so that AI answer engines such as ChatGPT, Perplexity, Claude and Google AI Overviews can crawl it, understand it, and cite it as a source.
Generative Engine Optimization (GEO) is the practice of structuring web content so that AI answer engines — ChatGPT, Perplexity, Claude, Gemini and Google AI Overviews — can crawl it, understand it, and cite it as a source. Where search engine optimisation competes for a position in a ranked list of links, GEO competes to be the passage a model quotes inside a single generated answer.
## Why does GEO exist as a separate discipline?
The shape of the result page changed, and with it the shape of the prize.
A classic search result gives a user ten links and lets them choose. An answer engine gives them one synthesised paragraph with a handful of citations attached. There is no second page of results to rank on, and frequently no click at all — the user reads the answer and leaves. Being the source named in that paragraph is a different objective from ranking third for a keyword, and it responds to different work.
Three technical realities drive most of that difference:
- **Most AI crawlers do not run JavaScript.** Googlebot has rendered JavaScript for years. The retrieval bots that fetch pages for answer engines largely do not — they read the raw HTML response and move on. A client-rendered application that ships an empty `
` is, to them, a blank page.
- **Retrieval works on passages, not pages.** A page is split into chunks, each chunk is embedded as a vector, and the chunks nearest the user's question are what the model actually reads. Your page does not compete as a unit; its individual sections do.
- **Attribution needs something to attribute.** A model deciding whether to name you as a source weighs whether there is an identifiable author, a date, a publisher and a canonical address. Unattributed text gets absorbed; attributed text gets cited.
## What does an answer engine actually look for?
Four things, in roughly this order of consequence.
**Can it fetch the page?** This is binary and it gates everything else. Your `robots.txt` either permits the retrieval bots or it does not. A `noindex` meta tag, a `nosnippet` directive, a 404, or content that only exists after a JavaScript bundle runs will each remove you from consideration entirely, no matter how good the writing is.
**Can it find the answer inside the page?** A single `
` stating the subject, `` sections that descend one level at a time, headings phrased as the questions readers actually ask, and paragraphs short enough that each covers one idea. Lists and tables matter here more than they do in SEO, because a model lifts them out intact — the markup already states the relationship between the items, so nothing has to be re-derived.
**Is there a reason to trust it?** JSON-LD structured data telling the engine what type of thing this page is, a named author, a `dateModified` that shows the page is maintained, outbound links to primary sources, and a canonical URL so citation credit lands in one place.
**Is the prose itself quotable?** This is the part most teams miss. A sentence beginning "It does this by…" is useless once extracted — the referent is gone. A sentence beginning "Generative Engine Optimization is…" survives extraction intact. Concrete figures beat vague claims, because a specific number is something a model can attribute where a hedge is something it must soften.
## How do I start optimising for AI answer engines?
Work in the order above, because the categories gate each other. There is no point rewriting your opening paragraph if `robots.txt` blocks PerplexityBot.
1. **Check your `robots.txt` against the retrieval bots.** Confirm `OAI-SearchBot`, `ChatGPT-User`, `PerplexityBot` and `Claude-User` can reach your content pages. Decide separately whether you want to allow the training crawlers.
2. **Confirm the content exists without JavaScript.** Fetch your page with `curl` and read what comes back. If the body is empty, nothing else you do will matter.
3. **Lead with the answer.** Make the first 100 words state what the page is about in a definition-shaped sentence that names its subject. Cut the "In today's fast-paced world" opening.
4. **Add structured data.** At minimum an `Organization` node for the site and the type matching each page — `Article` for a guide, `FAQPage` for a Q&A section, `HowTo` for instructions.
5. **Cite your own sources.** Pages that link to primary sources get cited more readily than pages that assert.
6. **Publish an `llms.txt`.** A Markdown index at your site root pointing models at the clean version of your key pages. Support is still emerging, but the cost is one file.
## What GEO is not
GEO is not a way to trick a model into citing you. The techniques that work are, almost without exception, the techniques that make a page genuinely easier for a person to read and verify: state your claims plainly, back them with figures, say who wrote it and when, and structure the page so the answer is findable. Keyword stuffing, hidden text and prompt-injection attempts in page copy are, at best, ignored.
It is also not a replacement for SEO. The two share a foundation — crawlability, structure, canonical URLs — and diverge at the top of the stack. A page that is invisible to Googlebot is usually invisible to the answer engines too.
## Where to go next
- [GEO vs SEO](/learn/geo-vs-seo) covers what carries over from your existing SEO work and what does not.
- [How to structure content for AI citations](/learn/structure-content-for-ai-citations) is the practical, page-level version of this guide.
- Or [run an audit](/analyze) and get the list for a page you actually own.
## Frequently asked questions
### Is GEO just SEO with a new name?
No. They share a foundation — a page still has to be crawlable and well-structured — but they optimise for different outcomes. SEO competes for a position in a list of links. GEO competes to be the passage an assistant quotes inside a single synthesised answer, where there is no list and often no click.
### Do I need to block AI crawlers to protect my content?
That is an editorial decision, and the two kinds of bot are worth separating. Training crawlers (GPTBot, CCBot, Google-Extended) build model corpora. Retrieval bots (OAI-SearchBot, PerplexityBot, ChatGPT-User, Claude-User) fetch pages to answer a question right now, with attribution. Blocking retrieval bots removes you from cited answers; blocking training bots does not.
### How long does it take to see results from GEO changes?
Crawlability fixes — unblocking a bot in robots.txt, server-rendering content — take effect as soon as the page is re-crawled, often within days. Content and structure changes depend on the engine re-reading and re-embedding the page, which typically takes longer.
### Does GEO work if my site is built with React or Vue?
Only if you server-render or statically generate the content. Most AI crawlers do not execute JavaScript, so a client-rendered page looks empty to them. Next.js, Nuxt, Remix and similar frameworks solve this; a plain client-side SPA does not.
---
# GEO vs SEO: what carries over and what does not
Source: https://envoyix.vercel.app/learn/geo-vs-seo
Author: The Envoyix team
Published: 2026-02-03 · Updated: 2026-09-25
> GEO and SEO share crawlability and structure as a foundation, but SEO optimises for a position in a ranked list of links while GEO optimises to be the passage quoted inside a single generated answer.
GEO and SEO share a technical foundation but optimise for different outcomes. SEO earns a position in a ranked list of ten links that a user chooses from; GEO earns a sentence inside a single synthesised answer where there is no list and often no click. Most of your crawlability and structure work carries over unchanged. What changes is what you do at the top of the stack.
## What is the same?
More than the framing of "GEO vs SEO" suggests. Both disciplines need:
- A page that returns HTTP 200 over HTTPS with a sane content type.
- Content present in the server HTML.
- One `` and a heading hierarchy that descends one level at a time.
- A self-referencing canonical URL.
- An XML sitemap, declared in `robots.txt`.
- Structured data describing what the page is.
- Fast loading and a page that works on a phone.
If your SEO fundamentals are in order, you have already done a large share of the GEO work. The audit categories in Envoyix reflect this: *crawlability* and *structure* are largely shared ground, while *authority* and *content* are where the disciplines diverge.
## Where do they diverge?
| Dimension | Classic SEO | GEO |
| --- | --- | --- |
| The prize | A position in a list of links | Being the cited source inside one answer |
| JavaScript | Googlebot renders it before indexing | Most AI crawlers read raw HTML only |
| Unit of competition | The page | The passage, chunked and embedded |
| Query shape | Short keyword phrases | Full natural-language questions |
| Ranking signals | Links, relevance, page experience | Extractability, specificity, attribution |
| Duplicate handling | Canonical consolidates ranking | Canonical decides who gets named |
| Success metric | Clicks and position | Citation frequency and share of voice |
| Failure mode | Ranking on page three | Being absorbed without attribution |
### JavaScript is the sharpest difference
Googlebot has rendered JavaScript for years, which let a generation of client-rendered applications rank acceptably. The retrieval bots behind answer engines largely do not render. This means a single architectural decision — client-side rendering versus server-side rendering — can leave a site visible in Google and simultaneously invisible in ChatGPT and Perplexity.
The fix is the ordinary one: server-render or statically generate the main content. Next.js, Nuxt, Astro, Remix and plain server-rendered templates all satisfy it.
### Retrieval competes on passages, not pages
An answer engine splits your page into chunks, embeds each one, and retrieves the chunks nearest the user's question. This changes what "optimising a page" means in a concrete way:
- A 300-word paragraph covering four ideas produces one muddy embedding that matches nothing strongly. Two to four focused sentences produce a chunk that matches its question sharply.
- A heading reading "Renewals" carries almost no signal. "How do I renew a passport?" sits close to the user's actual question in vector space.
- A sentence starting "It costs £88" loses its meaning once extracted. "A standard adult passport renewal costs £88" survives.
### Attribution replaces ranking as the scarce resource
In SEO, ten sources can appear on one result page. In an answer, typically three to five are cited, and often only one is named in the sentence that matters. The signals that decide *which* source gets named — a real author, a `dateModified`, a publisher, links to primary sources, a canonical URL — are therefore worth considerably more attention than they get in a typical SEO checklist.
## What should I actually change?
If you have a healthy SEO baseline, the GEO-specific work is roughly this:
1. **Audit `robots.txt` for the AI bots specifically.** A `User-agent: *` group that was fine for Googlebot may be blocking `OAI-SearchBot` by accident. Name the retrieval bots in their own groups so a later wildcard change cannot catch them.
2. **Verify raw-HTML rendering.** `curl` your own page and read the output.
3. **Rewrite openings to answer first.** Delete the throat-clearing introduction. The first sentence should define the subject.
4. **Rephrase headings as questions.** Aim for roughly 40% of your ``/`` headings to be the questions readers ask. Forcing every heading into a question reads badly; ignoring the pattern entirely wastes the signal.
5. **De-pronoun your key sentences.** Search for sentences beginning "It", "This" or "They" and name the subject instead.
6. **Add figures and cite them.** Replace "significantly faster" with a number and a source.
7. **Add an FAQ block with `FAQPage` JSON-LD.** Question-and-answer pairs are the single most retrievable structure on the web, because each one is already a self-contained unit.
8. **Publish `llms.txt` and, optionally, `llms-full.txt`.**
## How do I measure it?
There is no Search Console for answer engines yet, so measurement is more manual:
- **Server logs.** Confirm `OAI-SearchBot`, `PerplexityBot`, `ClaudeBot` and friends are actually fetching your pages. If they are not, nothing downstream matters.
- **Citation tracking.** Ask the engines the questions your pages target and record whether you are named. Do it on a fixed schedule so the results are comparable.
- **Referral traffic.** Assistant referrals show up in analytics from hosts such as `chat.openai.com` and `perplexity.ai`. The volume is usually small relative to search; the intent is usually high.
## Where to go next
- [What is GEO?](/learn/what-is-geo) for the conceptual grounding.
- [How to structure content for AI citations](/learn/structure-content-for-ai-citations) for the page-level mechanics.
- [Run an audit](/analyze) to see where a specific page stands.
## Frequently asked questions
### Should I stop doing SEO and do GEO instead?
No. Around 70% of the technical foundation is shared — crawlability, server-rendered HTML, heading structure, canonical URLs and structured data serve both. Treat GEO as an additional layer on top of a healthy SEO baseline, not a replacement for it.
### Do backlinks matter for GEO?
Less directly than for ranking, but they are not irrelevant. Answer engines lean on established notions of site authority when deciding which of several sources to name, and many of those notions are built from link graphs. The difference is that a single well-structured page on a modest domain can win a citation for a specific question in a way it could rarely win a top-three ranking.
### How do I measure GEO when there is no rank to track?
Measure citations rather than positions. Query the answer engines for the questions your page targets and record whether you are named. Watch referral traffic from chat.openai.com, perplexity.ai and similar hosts in your analytics. Track the AI crawler user agents in your server logs to confirm you are being fetched at all.
### Does keyword research still apply?
The intent research does; the keyword density does not. People type full questions into an assistant rather than two-word queries, so phrase headings as those questions. What no longer helps is repeating a target phrase for density — retrieval matches on meaning, not on term frequency.
---
# How to structure content for AI citations
Source: https://envoyix.vercel.app/learn/structure-content-for-ai-citations
Author: The Envoyix team
Published: 2026-03-11 · Updated: 2026-09-25
> To be cited by an AI answer engine, structure each page so that any passage lifted out of it still makes sense on its own: lead with the answer, phrase headings as questions, keep paragraphs to one idea, name subjects instead of using pronouns, and mark the page up with JSON-LD.
To be cited by an AI answer engine, structure each page so that any passage lifted out of it still makes sense on its own. That single rule explains nearly every recommendation below: retrieval systems extract fragments, not pages, and a fragment that depends on its surroundings becomes useless the moment it is separated from them.
## Why does self-containment matter so much?
An answer engine does not read your page the way a person does. It splits the page into chunks of a few hundred tokens, converts each chunk into a vector, and retrieves the handful whose vectors sit closest to the user's question. The model then writes an answer from those fragments.
Two consequences follow, and everything else is downstream of them:
1. **Your page competes as a set of passages, not as a document.** A brilliant page with one muddy section will lose that section's questions to a mediocre page with a sharp one.
2. **Context does not travel with the passage.** Whatever a fragment needs in order to make sense has to be inside the fragment.
## How should I open a page?
Answer the question in the first sentence, and name the subject while doing it.
Extractive systems weight the opening of a document heavily, which makes the first 100 words the most valuable real estate you have. Most pages spend it on a preamble.
**Before:**
> In today's fast-paced digital landscape, businesses are constantly searching for ways to stay ahead. In this article, we'll explore everything you need to know about our approach.
**After:**
> A content audit is a systematic review of every page on a site, scoring each one against its purpose so you can decide whether to keep, rewrite, merge or delete it. A typical audit of a 500-page site takes two to three weeks.
The second version is quotable as it stands. The first cannot be quoted at all, because it says nothing.
The test: read your first sentence out of context. If someone who has never seen your page could not tell what it is about, rewrite it.
## How should I phrase headings?
As the questions your readers actually ask.
People type full questions into an assistant — "how do I renew a UK passport" rather than "passport renewal". A heading phrased as that question sits close to the user's query in vector space; a heading reading "Renewals" does not.
- Write "How much does a passport renewal cost?" rather than "Pricing".
- Write "What documents do I need?" rather than "Requirements".
- Keep the answer directly beneath the heading, in the first sentence of the first paragraph.
Aim for roughly 40% of your `` and `` headings to be question-shaped. Forcing every heading into a question reads badly, and reference sections legitimately need noun headings. Ignoring the pattern altogether wastes the signal.
Structural rules that matter alongside phrasing:
- Exactly one ``, stating the page's subject.
- Levels descend one at a time. An `` followed by an `` breaks the nesting, and a chunker will attach that section to the wrong parent.
- No empty headings, and no headings used purely for their font size.
## How long should paragraphs and sentences be?
Paragraphs: 40 to 120 words, one idea each. Sentences: under 40 words.
A 300-word paragraph that covers pricing, availability and support produces one embedding that is a blurry average of three topics. It will lose to a focused competitor on all three questions. Splitting it into three paragraphs produces three sharp embeddings.
Sentence length matters for a different reason: a model quoting you has to take the whole sentence. A 60-word sentence with three subordinate clauses is one a model will paraphrase rather than quote — and a paraphrase is much less likely to carry a citation.
Aim for a Flesch Reading Ease score of 50 to 80. Note that higher is not automatically better: a technical page scoring above 90 has usually had its substance removed, and there is nothing left worth citing.
## How do I make individual sentences quotable?
Name the subject instead of pointing at it.
This is the most mechanical improvement available, and the most commonly missed. Scan your key sentences for ones beginning with "It", "This", "That", "They" or "These". Each of those is a sentence that becomes meaningless when extracted.
**Before:** "It typically takes three weeks and costs around £88."
**After:** "A standard adult passport renewal typically takes three weeks and costs £88."
Then add specifics. Concrete figures are what a model can attribute; vague claims are what it must hedge:
- "significantly faster" → "34% faster"
- "many organisations" → "61% of the 1,200 organisations surveyed"
- "recently" → "in March 2026"
Attribute each figure inline — "according to the 2026 Stack Overflow survey" — because an unsourced number is a liability rather than an asset.
## Which structures extract best?
In rough order of usefulness:
1. **FAQ sections.** Each entry is already a self-contained question-and-answer pair, which is exactly the unit retrieval wants. Use `` questions with a 40–60 word answer beneath, or ``/`` — the latter stays readable in raw HTML, unlike a JavaScript accordion.
2. **Definition lists and tables.** The markup states the relationship between items, so a model lifts the whole structure intact rather than re-deriving it. Give tables a `` and `| ` headers.
3. **Ordered and unordered lists.** Any enumeration buried in prose is a list waiting to be extracted.
4. **Short, focused paragraphs.** The default, done well.
## What markup should I add?
JSON-LD, in a ` |