---
name: agent-ready-site
description: Make the user's website readable and usable by AI agents. Audit the site the way an agent sees it and find the customer questions it can't answer, then add what the site needs, scaled to its size: robots.txt content signals, a generated llms.txt, Markdown versions of every page (by .md URL and by Accept header), documentation for the actions a visitor can take, and structured data, and finally verify all of it by using the site as an agent. Works with any framework or host. Use when the user wants their site to work with AI agents, AI search, ChatGPT, Claude, or Perplexity, or asks about llms.txt, AEO, or GEO.
---

# Make a website agent-ready

An AI agent visiting a site is usually doing one of two jobs: answering a question for its person ("who can fix my
roof in Bryant?") or doing something for them ("send this shop an inquiry"). You're making both jobs easy:

1. **Readable:** the agent can find every page and read it as clean text, without layout markup in the way.
2. **Usable:** the agent knows what actions the site offers and exactly how to take them.
3. **Answerable:** the site actually says what customers ask about. A perfectly readable site still loses the lead
   when it never says which towns it serves or what a job costs.

Work in the phases below, but not every site needs all of them; phase 1 decides which ones. Stop for the user's OK at
each **checkpoint**. Nothing ships until phase 7, and you never submit a real form, order, or booking without the
user's explicit permission.

## 0. Learn the site (read first; ask what you can't read)

- **How pages are built.** Find the framework and the host, and work out whether pages are server-rendered, statically
  generated, or rendered in the browser. Find where the content lives: Markdown files, a CMS, or the templates themselves.
- **What already exists.** Look for robots.txt, a sitemap, llms.txt, JSON-LD, and any redirects or middleware. Reuse what
  is there, and generate new files from the same sources the site already uses.
- **What a visitor can do.** List every action: contact forms, booking, quotes, checkout, downloads, and sign-up. For each
  one, find the endpoint it posts to and what it requires.
- **Ask the user:**
  - Whether AI companies may train on their content. This sets `ai-train` in phase 2, and it's their call, not yours.
    Explain the trade-off before asking. Training means the content becomes part of what future models learn, so a
    model can know about the business without searching for it. The cost is giving the content away with no control
    over how it's used later. This is separate from `ai-input`, which covers an agent reading the site live to answer
    a question; most businesses that want leads from agents should allow that either way.
  - Which actions an agent may take on a person's behalf. A contact form is usually fine. Checkout or account
    changes usually are not.

## 1. Audit the site as an agent sees it

Fetch the live site twice for each check: once like a browser (`Accept: text/html`) and once like an agent that prefers
Markdown (`Accept: text/markdown, text/html;q=0.9, */*;q=0.8`). Many agent fetch tools send the second header.

- **Is the content in the HTML?** Fetch a few pages with curl and look for their real text. If the HTML is an empty
  shell that JavaScript fills in, most agents see a blank page. That's the most important problem to fix, and it comes
  before everything else here.
- **Can agents get in?** Check robots.txt for blocked AI crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended)
  and check whether the host's bot protection or a firewall challenges non-browser clients.
- **Is there a CDN in front?** Check the response headers (`server: cloudflare` and `cf-ray` mean Cloudflare) and the
  domain's nameservers. A CDN can block AI crawlers or rewrite robots.txt without touching the site's code, and a
  test from your machine can't see it: a CDN identifies real AI crawlers by where they connect from, so a request
  that only borrows GPTBot's user agent proves nothing. See phase 2.
- **Can agents find everything?** Compare the sitemap with the pages that actually exist.
- **Can agents act?** Note forms that need JavaScript, a captcha, or a third-party widget to submit.
- **Do a real task.** Pick something a real customer would ask an agent ("can they fix my drainage in Bryant, and what
  will it cost?"), then answer it from the live site with your own web fetch tool. Write down every part of the
  answer the site couldn't give: whether they serve a nearby town, prices or price ranges, hours, how fast they
  respond, licensing, or a way to reach them besides calling. These **content gaps** are often the most valuable thing
  this skill finds, because each one is a lead the agent sends somewhere else. No routing or markup fixes them; the
  answers have to be written on the page.

**Size the work to the site.** If the HTML already contains the content and the site is small (a few pages and no
blog or docs), agents can read it fine as it is. Do phase 2, a short llms.txt from phase 3, and phase 6, and skip
phase 4 entirely: Markdown routing on a one-page site adds moving parts for almost no gain. Do the full set of phases
when the site has many pages, a blog or docs, heavy layout markup, or actions an agent can take.

**Checkpoint:** show the user, in plain language:

1. what an agent can and can't do on their site today;
2. the content gaps from the task, as questions the site should answer (for example, "Add the towns you serve" or
   "Add a typical price range"), which the user, not you, must fill in with real facts;
3. which phases you plan to do and which you'll skip, and why.

If the content isn't server-rendered, propose fixing that first (prerendering or server rendering in their framework)
and get approval before touching how their app renders. Never invent prices, service areas, or other facts to fill a
content gap.

## 2. robots.txt and content signals

Allow agents in, unless the user wants something else. Add a
[Content-Signal](https://contentsignals.org) line, which states what the content may be used for:

```
User-Agent: *
Content-Signal: search=yes, ai-input=yes, ai-train=<their answer>
Allow: /

Sitemap: https://example.com/sitemap.xml
```

If the framework generates robots.txt from a typed config with no field for this line, serve robots.txt from a plain
route or a static file instead.

**Check the CDN and host settings too.** You usually can't change these from the code, so give the user a checklist
for their dashboard, and recheck robots.txt on the live site after they're done:

- **Cloudflare** sorts AI traffic into three kinds, each with its own setting: **Search** (indexing to answer questions
  later), **Agent** (fetching a page live for a person), and **Training**. Since September 15, 2026, new domains that
  show ads default to blocking Agent traffic on pages with ads and disallowing Training. Agent is the setting this
  whole skill depends on, so make sure it's allowed. Cloudflare's managed robots.txt can also add its own rules and
  Content-Signal line in front of the site's, which can contradict the line you added. Bot Fight Mode and challenge
  rules can show agents a challenge page instead of the site. Each one should match what the user decided in phase 0.
- **Other hosts and CDNs** (Vercel, Netlify, AWS, Akamai, and others) have similar bot and firewall settings under
  different names. Look for anything that blocks bots, AI crawlers, or automated traffic.

## 3. llms.txt

Serve `/llms.txt` in the [llmstxt.org](https://llmstxt.org) format:

- an H1 with the business name and a one-paragraph summary in a blockquote;
- what the business offers and who it serves, in plain sentences;
- a link to every important page, pointing at its Markdown version from phase 4;
- a section for each action the user approved in phase 0 (see phase 5).

**Generate it from the site's own data** (the page list, the posts, and the services) whenever the framework allows it,
so a new page or post shows up here without anyone editing the file. A hand-written llms.txt goes out of date the first
time someone publishes something.

## 4. Markdown versions of every page

Serve each page as Markdown in two ways:

- **By URL:** `/about.md`, `/blog/my-post.md`, and `/index.md` for the homepage.
- **By content negotiation:** the page's normal URL, requested with `Accept: text/markdown`.

Where the Markdown comes from:

- If the content already is Markdown (a blog, docs, or a CMS that stores Markdown), serve it as written, with the title,
  author, and dates added at the top.
- Otherwise, convert the page's own rendered HTML: take the `<main>` element, drop scripts, styles, forms, and
  navigation, and turn headings, paragraphs, lists, and links into Markdown. Make relative links absolute. Converting the
  real page means the Markdown can't drift from what people see. Never write a second copy of the content by hand.

On every Markdown response, set `Content-Type: text/markdown; charset=utf-8`, a `Link: <html url>; rel="canonical"`
header, and `X-Robots-Tag: noindex`, so search engines keep indexing the HTML page. In each HTML page's head, add
`<link rel="alternate" type="text/markdown" href="...md">` so an agent reading the HTML can find the Markdown.

**Get the routing right. This is where it breaks.**

- **Rewrite only real pages.** Build the list of URLs to rewrite from the sitemap or the site's page list, never from a
  pattern like "every path" or "every path without a dot." Agents that send `Accept: text/markdown` also fetch
  llms.txt, robots.txt, the sitemap, images, generated images with no file extension, downloadable files, and API
  endpoints. If the rewrite catches any of these, those agents get a 404 on llms.txt and a 405 when they try to submit
  the contact form, while every browser test still passes.
- **Use the same list as an allowlist** in whatever serves the Markdown, so it never converts an API route or an
  arbitrary path.
- **Real `.md` files must still be served as themselves.** If the site hosts downloadable `.md` files, the `.md` URL
  rule has to run after the host checks for a real file.
- **Static hosts:** generate the `.md` files at build time. If the host can't vary a response by request header,
  serve the `.md` URLs only and say so in llms.txt.

## 5. Document the actions

For each action the user approved, add a section to llms.txt that an agent can follow on its first try: the method and
full URL, the content type, every field and whether it's required, an example request body, and what success and each
error look like. Read the endpoint's code to get this right; don't guess from the form's labels.

Write the consent rule into the docs: the agent should take the action only when its person asked it to, and should
use that person's real name and contact details so the business can reply to a real human.

- **If the only action is a phone call or an email address**, there's nothing to document beyond putting the number
  or address, and when to use it, in llms.txt and the JSON-LD. Don't build anything else. An agent can't place the
  call; it hands the number to its person. If the business wants agents to send inquiries, tell the user that a
  simple contact form would allow it, and leave the decision to them.
- **Don't add new endpoints** for agents. Document the ones the site already has. If an action exists only inside a
  third-party widget, link to that widget's page instead.
- **Don't remove security.** If a captcha or bot check blocks agents from submitting, tell the user about the trade-off
  and let them decide. A honeypot field that agents leave empty already lets agents through.

## 6. Structured data

Make sure each page carries JSON-LD for what it actually is: `Organization` or `LocalBusiness` (with address, phone,
and area served) on the homepage, `Service` on service pages, `Article` or `BlogPosting` on posts, `FAQPage` where the
page shows questions and answers, and `Person` for the founder or authors. Mark up only facts that appear on the page.
Agents and AI search engines both read this, so a wrong price or hours in JSON-LD spreads quickly.

## 7. Verify by using the site as an agent

Build the site, run it locally in production mode, and test every type of URL with both Accept headers:

| Request | Browser `Accept` | Markdown `Accept` |
| --- | --- | --- |
| Each page in the sitemap | HTML | Markdown |
| `/<page>.md` | Markdown | Markdown |
| llms.txt, robots.txt, sitemap.xml | Unchanged | Unchanged |
| Images, generated images, and downloadable files (including `.md` files) | Unchanged | Unchanged |
| A path that doesn't exist, and `/api/<anything>.md` | 404 | 404 |
| A POST to each documented endpoint with an invalid body | The endpoint's own 400 | The endpoint's own 400 |

If you skipped phase 4, every row should return what it did before your changes, except llms.txt, which now exists.

Then do the same real task from phase 1 again, starting from llms.txt alone, following its links, and working out how
to take the action. Compare the answer with the one from phase 1, and list any content gaps the user hasn't filled
yet.

**Checkpoint:** show the user the results and the diff. Ship only with their OK.

After it deploys, run the same table against the live site, and fetch llms.txt and a page with your own web fetch tool,
since that's what real agents use. Ask the user before sending one valid request to each action. Label it clearly as a
test, and have the user confirm it arrived.

## What to skip

Don't build an MCP server, WebMCP tools, an `ai-catalog.json`, `/.well-known` agent manifests, or an OpenAPI spec for a
site with one or two simple actions. Agents rarely look for them yet, and llms.txt already covers what they would.
Suggest them only when the site offers many actions or an agent needs to complete multi-step flows like booking or
checkout.
