ContraptionSoft solutions

NOTES · OCT 10, 2026

How to make your website usable by AI agents


A free skill for Claude Code or any agent: it makes your site readable and usable by AI agents, finds the questions your site leaves unanswered, and tests it all with a real agent.

By Tyler Malone

Why I bothered

Some of the people looking for a business like mine never see my website. They ask ChatGPT or Claude, and an agent reads the site for them. So I set out to make contraptionsoft.com easy for an agent to read and use, and then I did the part most guides skip: I had an agent actually use it, the same way a customer's agent would.

It worked, mostly. The agent started from one file, figured out what we do, found a past job that matched its task almost exactly, and sent a real inquiry through our contact form that landed in our CRM. It also found a bug I'd shipped that broke the site for exactly the agents I was trying to help, and every normal test I ran had missed it.

I packaged everything I learned as a skill. Give it to your AI agent and it does the same thing for your site, including the testing.

Not into terminals? We'll do it for you. Just mention the agent-ready skill.

What I changed

None of it is new tech. It's four small things that tell an agent where to look and how to read what it finds.

  • An llms.txt file. It's a plain-text page at /llms.txt that summarizes the business and links to every page, written for a model instead of a person.[1] Ours is generated from the same lists the site uses, so a new post shows up without anyone editing it.
  • Markdown versions of every page. An agent can add .md to any URL, or ask for Markdown in its request, and get clean text without the layout and styling mixed in.
  • Instructions for the contact form. The llms.txt file explains exactly how to send us an inquiry, with the fields it needs, and tells the agent to do it only when its person asked and with that person's real name and email.
  • A Content-Signal line in robots.txt. It's a short line that says what our content may be used for: search, answering questions, and training.[2] More on training below.

The agent found a bug I couldn't

I tested all of it the normal way, with a browser and with curl, and everything passed. Then I had Claude use the live site through its own web fetch tool, and the first file it asked for, llms.txt, came back as a 404.

The reason is a little nerdy, but it's the most useful thing in this post. Agents that prefer Markdown say so in every request, and I'd set the site up to send any request asking for Markdown to the Markdown version of the page. "Any request" turned out to include llms.txt itself, the sitemap, images, the files you can download, and the contact form. So for agents that prefer Markdown, which includes Claude's fetch tool, llms.txt didn't exist and the contact form returned an error. Browsers never ask for Markdown, so every test I ran in a browser looked perfect.

The fix took a few lines: only actual pages get the Markdown treatment, and everything else is left alone. The lesson is the part that matters. You can't test an agent-ready site with a browser. You have to test it with an agent, using the same tools a real one uses. The skill does that before it calls anything done.

After the fix, I ran the whole thing again. The agent read llms.txt, followed it to the right pages, answered a made-up customer's questions about automating purchase orders in Zoho, and sent one inquiry clearly labeled as a test. It showed up in our CRM.

The second site needed something different

Before posting this, I ran the skill on a site we built for a client, Malone Excavation. It's a single page with a phone number and no contact form, which is how a lot of small business sites look.

Agents could already read it fine. The content is in the page, and nothing was blocking them. When I asked an agent to help a homeowner in Bryant get a French drain installed, it found that Malone does French drains, that you deal directly with the owner, and that estimates are free. What it couldn't find out was whether Malone works in Bryant (the site says Saline County, but it never names the towns), what a French drain might cost, or how to reach them other than by calling.

None of that is a technical problem. Markdown files and llms.txt can't fix a question the site never answers. For most small business sites, the biggest improvement isn't anything technical. It's writing down the answers people ask on the phone every day: the towns you serve, a rough price range, your hours, and how to reach you. The skill now looks for those gaps first, gives them to you as a list, and skips the heavier setup when a site doesn't need it.

Check Cloudflare, or whatever sits in front of your site

If your site goes through Cloudflare, like mine does, the settings there can override everything above. Cloudflare can block AI traffic before it ever reaches your site, and it can rewrite your robots.txt.

It now splits AI traffic into three kinds you can allow or block separately: search crawlers, agents fetching a page live for a person, and training crawlers.[3] Since September 15, 2026, new sites that show ads default to blocking agents on pages with ads and disallowing training.[4] If your goal is for agents to read your site and send you customers, the agent setting is the one you can't get wrong. Your code can be perfect and an agent will still get a block page.

You can't fix these from your site's code, and an agent testing from its own machine can't see them, so the skill gives you a short checklist for your dashboard. Other hosts and CDNs have similar settings under different names.

Should you let AI train on your site?

This is a real choice, so here's what it means.

When an agent reads your site to answer someone's question, it reads the page live, uses it, and moves on. Training is different. Training means your content goes into the data a company uses to build its next model, so the model learns about your business and remembers it without needing to look anything up. The Content Signals standard keeps the two separate: ai-input covers the live reading, and ai-train covers training.[2]

Some reasons to say no: your content is your product (a publisher, a course, original research), you don't want it used in ways you can't control, or you'd rather be paid for it. Those are fair reasons, and it's your content.

I say yes. When someone asks a model who builds custom software in central Arkansas, I want the model to already know about ContraptionSoft, without having to search for it. The only way to get into the model's memory is to be in what it trained on. For a small local business that wants to be found, giving that content away is a trade I'm happy to make.

Either way, I'd allow the live reading. Blocking it means an agent helping a customer can't read your site at all, and it'll send them somewhere else.

What this won't do

I'll be straight about this: nobody knows yet how much of this pays off. Google says you don't need special files or markup to show up in its AI answers,[5] and plenty of agents never look for llms.txt at all. What's definitely true is that agents read websites today, some prefer Markdown, and a site that answers their questions and lets them reach you gets the lead while a site that doesn't gets skipped. The setup is a small one-time job, so I'd rather be early.

If you sell products online, agents buying things is a separate and much bigger project. I wrote about who should build agent checkout and who should wait. If you sell a service, this note covers the part that applies to you right now.

What your agent will do

  1. Learn how your site is built and what a visitor can do on it. It also asks whether to allow training and which actions an agent may take for someone.
  2. Use your site the way an agent does, then show you what works, what doesn't, and which questions your site can't answer. You decide what to fix.
  3. Update robots.txt and give you a checklist for Cloudflare or your host.
  4. Add an llms.txt file, generated from your site so it doesn't go out of date.
  5. Add Markdown versions of your pages, but only if your site is big enough to need them.
  6. Document what an agent can do on your site, like sending an inquiry.
  7. Check the structured data that tells search engines and agents what your business is.
  8. Test everything with real agent requests before it ships, then again on the live site. It asks before sending any real form, and it labels that request as a test.

It stops for your OK before changing anything that matters, and it never makes up prices, service areas, or other facts to fill a gap.

What you'll need

  • An AI agent that can edit code and run commands, like Claude Code
  • Your website's code, or access to wherever it's built
  • Access to your Cloudflare or hosting dashboard, if you use one

Get the skill

For Claude Code on Mac or Linux:

mkdir -p ~/.claude/skills/agent-ready-site && \
curl -fsSL https://contraptionsoft.com/skills/agent-ready-site/SKILL.md -o ~/.claude/skills/agent-ready-site/SKILL.md

Then open your website's code and say: "make my site agent-ready."

On Windows, or with any other agent, point it at https://contraptionsoft.com/skills/agent-ready-site/SKILL.md and ask it to install the skill and follow the steps.

Here's exactly what you're installing:

View the files1 file
---name: agent-ready-sitedescription: Make the user's website readable and usable by AI agents. Audit the site the way an agent sees it and find the customer questions it can't answer, then add what the site needs, scaled to its size: robots.txt content signals, a generated llms.txt, Markdown versions of every page (by .md URL and by Accept header), documentation for the actions a visitor can take, and structured data, and finally verify all of it by using the site as an agent. Works with any framework or host. Use when the user wants their site to work with AI agents, AI search, ChatGPT, Claude, or Perplexity, or asks about llms.txt, AEO, or GEO.--- # Make a website agent-ready An AI agent visiting a site is usually doing one of two jobs: answering a question for its person ("who can fix myroof in Bryant?") or doing something for them ("send this shop an inquiry"). You're making both jobs easy: 1. **Readable:** the agent can find every page and read it as clean text, without layout markup in the way.2. **Usable:** the agent knows what actions the site offers and exactly how to take them.3. **Answerable:** the site actually says what customers ask about. A perfectly readable site still loses the lead   when it never says which towns it serves or what a job costs. Work in the phases below, but not every site needs all of them; phase 1 decides which ones. Stop for the user's OK ateach **checkpoint**. Nothing ships until phase 7, and you never submit a real form, order, or booking without theuser's explicit permission. ## 0. Learn the site (read first; ask what you can't read) - **How pages are built.** Find the framework and the host, and work out whether pages are server-rendered, statically  generated, or rendered in the browser. Find where the content lives: Markdown files, a CMS, or the templates themselves.- **What already exists.** Look for robots.txt, a sitemap, llms.txt, JSON-LD, and any redirects or middleware. Reuse what  is there, and generate new files from the same sources the site already uses.- **What a visitor can do.** List every action: contact forms, booking, quotes, checkout, downloads, and sign-up. For each  one, find the endpoint it posts to and what it requires.- **Ask the user:**  - Whether AI companies may train on their content. This sets `ai-train` in phase 2, and it's their call, not yours.    Explain the trade-off before asking. Training means the content becomes part of what future models learn, so a    model can know about the business without searching for it. The cost is giving the content away with no control    over how it's used later. This is separate from `ai-input`, which covers an agent reading the site live to answer    a question; most businesses that want leads from agents should allow that either way.  - Which actions an agent may take on a person's behalf. A contact form is usually fine. Checkout or account    changes usually are not.## 1. Audit the site as an agent sees it Fetch the live site twice for each check: once like a browser (`Accept: text/html`) and once like an agent that prefersMarkdown (`Accept: text/markdown, text/html;q=0.9, */*;q=0.8`). Many agent fetch tools send the second header. - **Is the content in the HTML?** Fetch a few pages with curl and look for their real text. If the HTML is an empty  shell that JavaScript fills in, most agents see a blank page. That's the most important problem to fix, and it comes  before everything else here.- **Can agents get in?** Check robots.txt for blocked AI crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended)  and check whether the host's bot protection or a firewall challenges non-browser clients.- **Is there a CDN in front?** Check the response headers (`server: cloudflare` and `cf-ray` mean Cloudflare) and the  domain's nameservers. A CDN can block AI crawlers or rewrite robots.txt without touching the site's code, and a  test from your machine can't see it: a CDN identifies real AI crawlers by where they connect from, so a request  that only borrows GPTBot's user agent proves nothing. See phase 2.- **Can agents find everything?** Compare the sitemap with the pages that actually exist.- **Can agents act?** Note forms that need JavaScript, a captcha, or a third-party widget to submit.- **Do a real task.** Pick something a real customer would ask an agent ("can they fix my drainage in Bryant, and what  will it cost?"), then answer it from the live site with your own web fetch tool. Write down every part of the  answer the site couldn't give: whether they serve a nearby town, prices or price ranges, hours, how fast they  respond, licensing, or a way to reach them besides calling. These **content gaps** are often the most valuable thing  this skill finds, because each one is a lead the agent sends somewhere else. No routing or markup fixes them; the  answers have to be written on the page. **Size the work to the site.** If the HTML already contains the content and the site is small (a few pages and noblog or docs), agents can read it fine as it is. Do phase 2, a short llms.txt from phase 3, and phase 6, and skipphase 4 entirely: Markdown routing on a one-page site adds moving parts for almost no gain. Do the full set of phaseswhen the site has many pages, a blog or docs, heavy layout markup, or actions an agent can take. **Checkpoint:** show the user, in plain language: 1. what an agent can and can't do on their site today;2. the content gaps from the task, as questions the site should answer (for example, "Add the towns you serve" or   "Add a typical price range"), which the user, not you, must fill in with real facts;3. which phases you plan to do and which you'll skip, and why. If the content isn't server-rendered, propose fixing that first (prerendering or server rendering in their framework)and get approval before touching how their app renders. Never invent prices, service areas, or other facts to fill acontent gap. ## 2. robots.txt and content signals Allow agents in, unless the user wants something else. Add a[Content-Signal](https://contentsignals.org) line, which states what the content may be used for: ```User-Agent: *Content-Signal: search=yes, ai-input=yes, ai-train=<their answer>Allow: /Sitemap: https://example.com/sitemap.xml``` If the framework generates robots.txt from a typed config with no field for this line, serve robots.txt from a plainroute or a static file instead. **Check the CDN and host settings too.** You usually can't change these from the code, so give the user a checklistfor their dashboard, and recheck robots.txt on the live site after they're done: - **Cloudflare** sorts AI traffic into three kinds, each with its own setting: **Search** (indexing to answer questions  later), **Agent** (fetching a page live for a person), and **Training**. Since September 15, 2026, new domains that  show ads default to blocking Agent traffic on pages with ads and disallowing Training. Agent is the setting this  whole skill depends on, so make sure it's allowed. Cloudflare's managed robots.txt can also add its own rules and  Content-Signal line in front of the site's, which can contradict the line you added. Bot Fight Mode and challenge  rules can show agents a challenge page instead of the site. Each one should match what the user decided in phase 0.- **Other hosts and CDNs** (Vercel, Netlify, AWS, Akamai, and others) have similar bot and firewall settings under  different names. Look for anything that blocks bots, AI crawlers, or automated traffic. ## 3. llms.txt Serve `/llms.txt` in the [llmstxt.org](https://llmstxt.org) format: - an H1 with the business name and a one-paragraph summary in a blockquote;- what the business offers and who it serves, in plain sentences;- a link to every important page, pointing at its Markdown version from phase 4;- a section for each action the user approved in phase 0 (see phase 5). **Generate it from the site's own data** (the page list, the posts, and the services) whenever the framework allows it,so a new page or post shows up here without anyone editing the file. A hand-written llms.txt goes out of date the firsttime someone publishes something. ## 4. Markdown versions of every page Serve each page as Markdown in two ways: - **By URL:** `/about.md`, `/blog/my-post.md`, and `/index.md` for the homepage.- **By content negotiation:** the page's normal URL, requested with `Accept: text/markdown`. Where the Markdown comes from: - If the content already is Markdown (a blog, docs, or a CMS that stores Markdown), serve it as written, with the title,  author, and dates added at the top.- Otherwise, convert the page's own rendered HTML: take the `<main>` element, drop scripts, styles, forms, and  navigation, and turn headings, paragraphs, lists, and links into Markdown. Make relative links absolute. Converting the  real page means the Markdown can't drift from what people see. Never write a second copy of the content by hand. On every Markdown response, set `Content-Type: text/markdown; charset=utf-8`, a `Link: <html url>; rel="canonical"`header, and `X-Robots-Tag: noindex`, so search engines keep indexing the HTML page. In each HTML page's head, add`<link rel="alternate" type="text/markdown" href="...md">` so an agent reading the HTML can find the Markdown. **Get the routing right. This is where it breaks.** - **Rewrite only real pages.** Build the list of URLs to rewrite from the sitemap or the site's page list, never from a  pattern like "every path" or "every path without a dot." Agents that send `Accept: text/markdown` also fetch  llms.txt, robots.txt, the sitemap, images, generated images with no file extension, downloadable files, and API  endpoints. If the rewrite catches any of these, those agents get a 404 on llms.txt and a 405 when they try to submit  the contact form, while every browser test still passes.- **Use the same list as an allowlist** in whatever serves the Markdown, so it never converts an API route or an  arbitrary path.- **Real `.md` files must still be served as themselves.** If the site hosts downloadable `.md` files, the `.md` URL  rule has to run after the host checks for a real file.- **Static hosts:** generate the `.md` files at build time. If the host can't vary a response by request header,  serve the `.md` URLs only and say so in llms.txt. ## 5. Document the actions For each action the user approved, add a section to llms.txt that an agent can follow on its first try: the method andfull URL, the content type, every field and whether it's required, an example request body, and what success and eacherror look like. Read the endpoint's code to get this right; don't guess from the form's labels. Write the consent rule into the docs: the agent should take the action only when its person asked it to, and shoulduse that person's real name and contact details so the business can reply to a real human. - **If the only action is a phone call or an email address**, there's nothing to document beyond putting the number  or address, and when to use it, in llms.txt and the JSON-LD. Don't build anything else. An agent can't place the  call; it hands the number to its person. If the business wants agents to send inquiries, tell the user that a  simple contact form would allow it, and leave the decision to them.- **Don't add new endpoints** for agents. Document the ones the site already has. If an action exists only inside a  third-party widget, link to that widget's page instead.- **Don't remove security.** If a captcha or bot check blocks agents from submitting, tell the user about the trade-off  and let them decide. A honeypot field that agents leave empty already lets agents through. ## 6. Structured data Make sure each page carries JSON-LD for what it actually is: `Organization` or `LocalBusiness` (with address, phone,and area served) on the homepage, `Service` on service pages, `Article` or `BlogPosting` on posts, `FAQPage` where thepage shows questions and answers, and `Person` for the founder or authors. Mark up only facts that appear on the page.Agents and AI search engines both read this, so a wrong price or hours in JSON-LD spreads quickly. ## 7. Verify by using the site as an agent Build the site, run it locally in production mode, and test every type of URL with both Accept headers: | Request | Browser `Accept` | Markdown `Accept` || --- | --- | --- || Each page in the sitemap | HTML | Markdown || `/<page>.md` | Markdown | Markdown || llms.txt, robots.txt, sitemap.xml | Unchanged | Unchanged || Images, generated images, and downloadable files (including `.md` files) | Unchanged | Unchanged || A path that doesn't exist, and `/api/<anything>.md` | 404 | 404 || A POST to each documented endpoint with an invalid body | The endpoint's own 400 | The endpoint's own 400 | If you skipped phase 4, every row should return what it did before your changes, except llms.txt, which now exists. Then do the same real task from phase 1 again, starting from llms.txt alone, following its links, and working out howto take the action. Compare the answer with the one from phase 1, and list any content gaps the user hasn't filledyet. **Checkpoint:** show the user the results and the diff. Ship only with their OK. After it deploys, run the same table against the live site, and fetch llms.txt and a page with your own web fetch tool,since that's what real agents use. Ask the user before sending one valid request to each action. Label it clearly as atest, and have the user confirm it arrived. ## What to skip Don't build an MCP server, WebMCP tools, an `ai-catalog.json`, `/.well-known` agent manifests, or an OpenAPI spec for asite with one or two simple actions. Agents rarely look for them yet, and llms.txt already covers what they would.Suggest them only when the site offers many actions or an agent needs to complete multi-step flows like booking orcheckout.
SKILL.md · 205 lines · 13.7 KBDownload

Want us to make your site agent-ready, or build a site that answers what your customers actually ask? That's what we do. See web design or get in touch.

Sources

  1. llms-txt, The /llms.txt file
  2. Cloudflare, Giving users choice with Cloudflare's new Content Signals Policy
  3. Cloudflare, New options to manage AI traffic
  4. Cloudflare, Have it both ways: stay discoverable in search while disallowing AI training
  5. Google Search Central, AI Features and Your Website