Find where AI drops off your funnel
SEO gets you found by search engines. This gets you found, understood and acted on by AI: ChatGPT, Claude, Perplexity, Copilot, and the agents your customers use.
Free to use, share, and modify.
---
name: agent-funnel-audit
version: 1.0.0
description: Audit whether AI can find a website, understand it, and act on it for its user. Use when the user says "agent funnel audit," "AI discoverability," "audit my site for AI," "can ChatGPT find my site," "AI search," "GEO," "AEO," "llms.txt," "AI crawlers," "robots.txt for AI bots," or "agent-ready," or gives a URL and asks how it shows up in ChatGPT, Claude, Perplexity, Copilot, or Gemini.
---
# Agent Funnel Audit
SEO gets a site found by search engines. This audit checks the funnel AI goes through before it sends a customer to a site: ChatGPT, Claude, Perplexity, Copilot, Gemini, and the agents people use to get things done. Can AI find the site, understand it, and act on it for its user?
You run it on one URL. It has four parts:
1. **Found: search basics.** AI answers are built on search indexes, so the basics still matter.
2. **Let in and understood: AI access.** Which AI bots can read the site (robots.txt, firewall, JavaScript), and what the site tells them (llms.txt).
3. **Acted on: agent actions.** Could an AI agent sign up, book, or buy for its user?
4. **The report.** A score for each stage, the evidence, a ranked fix list, and ready-to-paste text.
## Rules
- **Read-only.** Use GET and HEAD requests only. Never submit a form, create an account, join a list, book a slot, start a checkout, or pay. Not even with test data. Reading a form is fine. Sending it is not.
- **Cite evidence for every finding:** the URL, the status code, and the exact line or snippet you saw. No evidence, no finding.
- **Never invent facts about the site.** If you could not check something, write "Not checked" and say why. Don't fill gaps from memory or from what sites like this usually do.
- **Plain language.** The owner may not be technical. Explain a term in a few words the first time you use it.
- **Be polite to the server.** Sample pages. Don't crawl everything. About 25 pages is plenty for the in-depth checks. Quick status checks of sitemap URLs (a HEAD request each) don't count toward that.
- **If a file changes while you audit** (for example llms.txt is redeployed), audit the newer version and say what changed.
- If you have no web access at all, say so and stop. Don't audit from memory.
## Before you start
**1. Know your tools.** Work out what you can actually see, and say so at the top of the report:
- Raw HTML: the page source, including the head, meta tags, JSON-LD, and hidden form fields. Some fetch tools return only cleaned-up text. With text only, you can't check meta tags, schema, or honeypots. Mark those "Not checked".
- Status codes and redirects (200, 301, 308, 403).
- Custom user agents, to test what a bot sees.
- A real browser that runs JavaScript.
A shell with curl can do all of this. Whatever you can't do, cover with "Commands for the owner" at the end of the report.
**2. Get the context.** Fetch the homepage and work out, from the site itself:
- What it is, who it's for, and the type of business: software or app, online store, local business, service business, or content.
- The main actions it wants visitors to take (its calls to action).
- The canonical host: https://example.com or https://www.example.com.
If the user told you the business goal or the searches they want to win, use that. Otherwise don't stop to ask. Run the audit and state your assumptions.
**3. Pick the sample.** The homepage, pricing, product or features, about, contact, every signup, waitlist, demo, quote, and download page, docs if any, one or two blog posts, then sitemap URLs up to about 25 pages. You may try common paths like /pricing, /docs, and /blog directly. A 404 is a fine answer. Say how many pages you checked.
---
## Part 1. Found: search basics
ChatGPT search, Copilot, Perplexity, and Google's AI answers all start from a search index. If search engines can't crawl and trust the site, AI won't cite it.
### 1.1 HTTPS and redirects
Request all four versions of the homepage: http://example.com, http://www.example.com, https://example.com, and https://www.example.com.
- Every version except the canonical one must redirect to it with a permanent redirect: 301 or 308. A 302 or 307 is temporary. Flag it.
- One hop is best. Flag chains (http to https to www to the final URL).
- The final URL must match the page's canonical tag (link rel="canonical") and the host used in the sitemap. All three must agree.
- Every page loads over HTTPS with a valid certificate.
### 1.2 Sitemap
- Find it: the "Sitemap:" line in /robots.txt, then /sitemap.xml and /sitemap_index.xml.
- It must be referenced in robots.txt with a full URL.
- It returns 200 and is valid XML on the canonical https host.
- It lists every public page. Compare it with the pages linked from the nav, the footer, and the homepage. List public pages that are missing.
- Every sitemap URL returns 200 directly. Flag URLs that redirect, return 404, or are set to noindex. Status-check them all if there are 50 or fewer (a HEAD request is enough). Otherwise sample 25 across sections.
### 1.3 Every sampled page
- **Title tag:** present, unique, 60 characters or fewer, and contains the search the page should win. Say which search you think it targets.
- **Meta description:** present, unique, about 150 characters. Flag under 70 or over 160.
- **Exactly one H1.**
- **A self-referencing canonical:** an absolute https URL pointing to the page itself.
- **Not noindexed by accident.** Check the robots meta tag and the X-Robots-Tag header. A noindex on a page that is in the sitemap or the nav is almost always a mistake.
### 1.4 Structured data (JSON-LD)
Structured data is how a machine reads the facts on a page without guessing. Check for:
- **Organization** on the homepage: name, url, logo, and sameAs (links to the official social profiles, GitHub, Crunchbase, and Wikipedia or Wikidata where they exist). If the site is a product or brand of a larger company, add parentOrganization.
- **WebSite:** name and url.
- **The product type** that fits the business: SoftwareApplication (apps and SaaS), Product (goods), LocalBusiness (a place people visit, with address, phone, and hours), or Service (a service business). It must include **offers** with a **price** and **priceCurrency**. A free plan is price "0". If prices are "contact us", note it. People ask AI "how much does it cost" all the time. With no price in the markup, AI guesses or skips the site.
- **FAQPage** where the page shows real questions and answers. Only mark up questions that are visible on the page. Google shows few FAQ rich results now, but the markup still hands AI clean question and answer pairs.
- **BreadcrumbList** on inner pages.
- The JSON must parse, and every value must match the visible page. A price in the markup that differs from the pricing page is a finding.
### 1.5 Images and social cards
- Content images have real alt text that describes them. Decorative images use alt="". Flag images with no alt attribute, and alt text like "image" or a file name.
- Each page has its own og:title, og:description, and og:image. The image is an absolute URL that returns 200, ideally 1200 by 630. Flag it if every page shares one generic image.
### 1.6 Content flags
- **Thin pages:** under about 300 words of real content, not counting the nav and footer. Contact and legal pages can be short. Mention them, but rank them low.
- **Orphan pages:** pages in the sitemap that no other page links to.
- **Duplicate titles** or descriptions across pages.
### 1.7 Indexed in Google AND Bing
- You can't see the owner's accounts. Look for signs of verification: a google-site-verification meta tag, an msvalidate.01 meta tag (Bing), or /BingSiteAuth.xml. Their absence proves nothing, because DNS verification leaves no trace on the page. Ask the owner to confirm.
- If you can search, run site:example.com on Google and on Bing. Report roughly what you see. If your search tool ignores "site:", say so and mark it "Not checked".
- **Stress Bing.** ChatGPT search leans on Bing's index (alongside OpenAI's own crawler). Microsoft Copilot is built on Bing. DuckDuckGo's results come largely from Bing. A site missing from Bing is missing from all of them. The fix is quick: add the site to Bing Webmaster Tools (it can import a Google Search Console setup in minutes), submit the sitemap, and turn on IndexNow so new pages reach Bing right away.
---
## Part 2. Let in and understood: AI access
### 2.1 robots.txt
Fetch /robots.txt.
- It should return 200 as plain text. A 404 means everything is allowed. That's fine, but there's no sitemap reference. A 5xx error can make crawlers back off the whole site. An HTML page served at /robots.txt is a finding.
- How to read it: a bot obeys the most specific "User-agent" group that names it. If none names it, it obeys the "User-agent: *" group. A named group replaces the * group completely. It does not add to it. Within a group, the longest matching path wins, and Allow wins a tie.
For each bot below, write Allowed, Blocked, or Partly blocked (and which paths), for the homepage and your sampled pages. If robots.txt only has a "User-agent: *" group, one verdict covers every bot. Say so once.
| Bot | Company | Type | What blocking it does |
|---|---|---|---|
| GPTBot | OpenAI | Training | Keeps content out of future OpenAI models. Does not remove the site from ChatGPT search. |
| OAI-SearchBot | OpenAI | Search and answers | Removes the site from ChatGPT search results and citations. |
| ChatGPT-User | OpenAI | User fetch | ChatGPT can't open a page when a user asks it to. |
| ClaudeBot | Anthropic | Training | Keeps content out of future Claude models. |
| Claude-SearchBot | Anthropic | Search and answers | Hurts how the site shows up in Claude's search answers. |
| Claude-User | Anthropic | User fetch | Claude can't open a page when a user asks it to. |
| PerplexityBot | Perplexity | Search and answers | Removes the site from Perplexity answers. |
| Perplexity-User | Perplexity | User fetch | Perplexity can't open a page when a user asks it to. |
| Google-Extended | Google | Training and Gemini | Opts out of Gemini training and grounding. Does not affect Google Search or AI Overviews, which use Googlebot. |
| Applebot-Extended | Apple | Training | Opts out of Apple's AI training. Does not affect Siri or Spotlight, which use Applebot. |
| Bingbot | Microsoft | Search and answers | Removes the site from Bing, Copilot, and much of ChatGPT search and DuckDuckGo. Critical. |
| CCBot | Common Crawl | Training | Keeps content out of an open dataset that many AI models train on. |
| Meta-ExternalAgent | Meta | Training | Keeps content out of Meta's AI training. |
| Amazonbot | Amazon | Mixed | Feeds Alexa answers and may train Amazon's models. |
Also check Googlebot. Blocking it removes the site from Google Search and Google's AI answers.
Then judge each block:
- **Training bots:** blocking them can be a fair business choice. Report it neutrally.
- **Search and answer bots, and user fetchers:** blocking them makes the site invisible in AI answers, and stops AI from reading a page a user sends it. Treat it as a Fail unless the owner clearly wants it.
- **Does it look intentional?** Signs it's deliberate: comments explaining it, or a clean pattern such as only training bots blocked. Signs it's an accident: "User-agent: *" with "Disallow: /" left over from a staging site, search bots blocked while training bots are allowed (backwards), key pages like /pricing or /signup disallowed, or a block list pasted from a plugin or template. Say which, and why.
### 2.2 Firewall and CDN blocking
robots.txt is a request. A firewall is a wall. Many sites block AI bots at the CDN without knowing it. Cloudflare, for example, has a one-click "Block AI bots" setting, blocks AI crawlers by default on many newer sites, and can add lines to robots.txt for you. So the live robots.txt may not match the one in the code.
- Request the homepage, /robots.txt, and /llms.txt with a normal browser user agent. Then request them with bot user agents: at least GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, and bingbot. Compare the status codes and the content.
- Blocked looks like: 403, 429, or 503 for the bot but 200 for the browser; a challenge page ("Just a moment...", "Attention Required", "Verify you are human"); or an empty page.
- Name the CDN if the headers show it. For example "server: cloudflare" with a "cf-ray" header, or Vercel, Fastly, Akamai, or CloudFront headers.
- Caveat: real bots come from known IP addresses, and some firewalls check that. A user agent test is a strong hint, not proof. Tell the owner where to confirm: the CDN's bot or AI crawler settings, and the server logs.
- In the bot table, write "Not tested" for bots you didn't test.
- If you can't set a user agent, mark this "Not checked" and include the owner command.
### 2.3 Can AI read the page without JavaScript?
Most AI crawlers don't run JavaScript. They read the raw HTML the server sends.
- Compare the raw HTML with what a browser shows. Are the headline, the product description, the prices, and the main call-to-action links in the raw HTML?
- If the raw HTML is an empty shell (for example a lone div with id="root"), AI crawlers see an empty page. The fix is server-side rendering or static pages, at least for the key pages.
### 2.4 llms.txt
llms.txt is a plain text file at the site root. It tells AI what the site is and where the important pages are. It's a young convention and not every AI company reads it yet. But the agents and AI tools that do land on a site look for it, and it takes about an hour to write. Fetch /llms.txt:
- It returns 200 as plain text or markdown ("text/plain" or "text/markdown"). An HTML page (it starts with "<!DOCTYPE" or "<html"), a redirect to the homepage, or a styled 404 page all count as missing.
- It follows the format at llmstxt.org: a "# Name" heading, a one-paragraph summary as a "> " quote, then "## Section" headings with lists of "- [Page title](url): one-line description".
- It says what the product is, who it's for, and what it costs: real prices or "free", with plan names.
- It links the key pages, each with one line on what's there. Every link returns 200 and uses the canonical URL.
- **Every claim is true today.** Check each one against the live site: prices, plans, features, platforms, integrations, availability ("beta", "coming soon", "waitlist"), and company names. AI answers repeat llms.txt almost word for word. A stale price in llms.txt becomes a wrong price in ChatGPT.
- Note whether /llms-full.txt exists. It's optional: a longer version with the full page content.
### 2.5 Not worth doing
If the site has /.well-known/ai-plugin.json (the old ChatGPT plugin manifest, retired in 2024) or an agent.json file, note that no major AI assistant reads them today. They do no harm. Don't spend time on them, and make sure they don't contradict llms.txt. Don't recommend creating them.
---
## Part 3. Acted on: agent actions
More people now ask an AI to do things for them: "sign me up," "book a demo," "get me a quote." This part checks whether an agent could do that on this site without a human looking at the screen. Remember the rules: look, never submit.
### 3.1 Find the actions
Collect every call to action on the homepage, the nav, the pricing page, and the footer: sign up, get started, join a waitlist, start a trial, book a demo or call, get a quote, download, contact, subscribe, and buy. For each, note the button label, where it goes, and what's there: a form on the site, a pop-up, a booking tool like Calendly or Cal.com, an app store, or a checkout page like Stripe. Merge duplicates. Most sites have two to six real actions.
### 3.2 Test each action
For each action, answer these with evidence.
**Reachable by a plain URL?**
- Does the action have its own URL that loads the form directly, without a login? Or does it only exist as a pop-up that opens when a button is clicked?
- Can options be preselected in the URL? For example /signup?plan=pro&billing=annual. Only say yes if you see proof: a link on the site that uses the parameter, or a hidden field that reads it. A parameter that only appears in docs or llms.txt is the owner's claim. Report it as "documented, not verified".
- Is there a way to tag agent traffic, like ?source=agent or a utm_source? If not, recommend one so the owner can see what agents bring in.
**Forms**
- Is it a real form element with a method and an action? Or a JavaScript-only widget, such as clickable divs or a form that only appears after scripts run? An agent without a browser can't use the second kind. For embedded third-party forms (Typeform, Tally, HubSpot, Google Forms, Calendly), find the direct third-party URL.
- Does every field have a label (a label element or an aria-label)? Placeholder text alone doesn't count.
- Are the field names, types (email, tel), and required fields clear?
- Is there a CAPTCHA that needs a human, such as a reCAPTCHA checkbox or image puzzle, hCaptcha, or an interactive Turnstile? Invisible, score-based checks may or may not block agents. Flag those as "may block agents".
- Where does the data go? Note the form's action URL, or an endpoint named in the page, the docs, or llms.txt. Never call it. Reading the site's JavaScript files is fine (they're GET requests), but an endpoint found only inside a JavaScript file is not documented. Mention it in the findings, not in the Actions draft.
**Honeypot fields**
A honeypot is a hidden field that only bots fill in. It's a spam trap. Look for text inputs hidden with CSS (display:none, visibility:hidden, or positioned off-screen), often with tabindex="-1", autocomplete="off", or aria-hidden, a class name containing "honeypot", a label like "Leave this empty", and names like website, url, homepage, fax, hp, honeypot, _gotcha, bot-field, or Mailchimp's "b_" followed by long IDs. If the field is hidden by a CSS class you can't see, rely on these signs and say how sure you are.
- Name each one exactly.
- Why it matters: an agent that fills in every field it finds fills in the trap, and the signup is silently thrown away. Nobody sees an error. So llms.txt must tell agents to leave these fields blank.
- Don't confuse honeypots with type="hidden" fields that carry values, like security tokens, form IDs, or tracking. Agents must send those unchanged.
**Money**
Agents must never pay. If an action involves payment, the right pattern is a hand-off. The agent builds a link with the plan, the quantity, and the user's email already filled in, and the human finishes checkout. Stripe Payment Links, for example, accept prefilled_email and client_reference_id in the URL. Check whether the site offers a link like that. Flag any flow where the only way through is typing in card details.
**Consent**
A signup an agent starts should trigger a confirmation email the person must click (double opt-in). It stops anyone from signing up someone else's address, and it keeps lists clean. You can't test this without submitting, so look for evidence on the page ("check your inbox to confirm"), in the docs, or in the privacy policy. Otherwise mark it "Owner to confirm".
**Downloads**
There should be a stable URL that always serves the latest version, like /download/mac, not a versioned file name that goes stale. Check it with a HEAD request. Don't download big files. It should return 200 and the right file type.
### 3.3 Machine-readable actions
schema.org lets a site describe its actions in its JSON-LD with potentialAction. Few AI tools read these today. But they're standard, cheap, and the most likely hook for agents. Check for any that apply:
- **RegisterAction:** signup or waitlist.
- **SubscribeAction:** newsletter.
- **BuyAction**, or offers with a url: the pricing or checkout hand-off.
- **ReserveAction:** book a demo, a call, or a table.
- **DownloadAction**, or downloadUrl on a SoftwareApplication.
Each one needs a target with the real URL. For example:
~~~json
"potentialAction": {
"@type": "RegisterAction",
"name": "Join the waitlist",
"target": {
"@type": "EntryPoint",
"urlTemplate": "https://example.com/waitlist?source=agent"
}
}
~~~
### 3.4 Watch, not build yet
Mention these two, but don't recommend building for them yet:
- **WebMCP:** a proposed browser standard that lets a page offer tools that an AI agent in the browser can call directly.
- **NLWeb:** an open project from Microsoft that turns a site's structured data into an endpoint AI can ask questions in plain language.
Both are early. Clean pages, good structured data, and a clear llms.txt are what they build on anyway.
---
## Part 4. The report
Write the report in this order. Keep it plain. Use tables where they help.
### Header
- The site, the date, and how many pages you checked. If you don't know today's date, use the Date header from a response.
- What you could and couldn't see: raw HTML, status codes, user agent tests, JavaScript.
- One line on what the site is and who it's for, taken from the site itself.
### Summary
Two or three sentences. Can AI find, understand, and act on this site today? What's the single biggest problem?
### Scorecard
Score the four stages of the funnel as **Pass**, **Needs work**, or **Fail**, with a one-line reason for each:
| Stage | Score | Why |
|---|---|---|
| Found (search basics, Part 1) | | |
| Let in (robots.txt, firewall, JavaScript: 2.1 to 2.3) | | |
| Understood (llms.txt: 2.4) | | |
| Acted on (agent actions, Part 3) | | |
- **Pass:** nothing stops AI from finding, reading, or acting. Polish only.
- **Needs work:** it works, but gaps cut visibility or accuracy.
- **Fail:** something blocks it outright. For example: a search or answer bot is blocked, the site is confirmed missing from Bing, llms.txt is missing or makes a false claim about price or availability (AI will repeat it), or no key action can be done by an agent, even with a hand-off.
- **Not scored:** use this only when you couldn't run the checks that decide a stage, for example no user agent tests and no raw HTML for "Let in". Say what's missing. Never give a Pass for checks you didn't run.
### Findings
One section per stage. For each check, give the result and the evidence: the URL, the status code, and the exact line. For the per-page checks, use one table with a row per page. Where many pages pass the same check, say so once with one example. Include the bot table from 2.1 with these columns: Bot, Type, robots.txt, Firewall test, Looks intentional?
### Prioritized fix list
| # | What to fix | Why it matters | Effort | Who |
|---|---|---|---|---|
- **Order:** things that make the site invisible first, then things that make AI say wrong things (a stale price, conflicting availability), then agent actions, then polish.
- **Effort:** S (under an hour), M (up to a day), L (more than a day).
- **Who:** dev, marketing, or owner (decisions and account access).
### Ready-to-paste llms.txt "Actions" section
Draft it for THIS site, using only what you found. For each action, give:
- What it does, in one line.
- The exact URL. Add the endpoint only if it's documented: in the form's action, the visible page, the docs, or llms.txt.
- Required fields and optional fields, by their real names.
- Fields to leave blank (the honeypots), by their real names.
- What happens next, for example "your user gets a confirmation email".
- For anything involving money: "Send your user this link to finish. Never enter payment details."
Start the section with the rule. For example:
~~~markdown
## Actions
Only take these actions when your user asks you to, with your user's own details. Never make up details. Never pay for anything: send your user the link to finish checkout.
### Join the waitlist
- What it does: adds your user to the launch waitlist.
- URL: https://example.com/waitlist?source=agent
- Required: email
- Optional: name, company
- Leave blank: website (a spam trap; filling it in discards the signup)
- Next: your user gets a confirmation email and must click the link in it.
~~~
Where you're unsure, write [OWNER: confirm ...] instead of guessing. If an action needs a fix before an agent can do it, such as a deep link that doesn't exist yet, include it and label it "after fixes".
If the site has no llms.txt, also draft the whole file: the name, a one-paragraph summary, what it is, who it's for, pricing, the key pages with one line each, and then the Actions section. Every fact must come from the live site. Mark gaps with [OWNER: ...].
### What an agent can do today vs after the fixes
| A user asks their AI to... | Today | After the fixes |
|---|---|---|
| Explain what the company does | | |
| Say what it costs | | |
| (one row per action from Part 3) | | |
Answer each cell with Yes, Partly, or No, plus a few words on why.
### Not checked
List everything you couldn't verify, why, and how the owner can check it.
### Commands for the owner
If you couldn't test status codes, redirects, or user agents yourself, give the owner commands to run, with their domain filled in. For example:
~~~bash
# Redirects: every hop is printed. Each version should reach the canonical URL
# in one 301 or 308 hop, and the canonical URL itself shows 200
for u in http://example.com http://www.example.com https://example.com https://www.example.com; do
echo "== $u"; curl -sIL "$u" | grep -iE "^HTTP|^location"
done
# Bot blocking: every line should show 200
for ua in "Mozilla/5.0" GPTBot OAI-SearchBot ChatGPT-User ClaudeBot Claude-User PerplexityBot bingbot; do
printf "%-16s" "$ua"; curl -s -o /dev/null -w "%{http_code}\n" -A "$ua" https://example.com/
done
# llms.txt: should be 200 with text/plain or text/markdown
curl -sI https://example.com/llms.txt | grep -iE "^HTTP|content-type"
~~~What this skill does
Found
Redirects, sitemap, titles, canonicals, and schema with real prices. Plus Bing, which ChatGPT search, Copilot, and DuckDuckGo lean on.
Understood
Which AI bots your robots.txt and firewall let in, and whether your llms.txt says what you sell, what it costs, and nothing untrue.
Acted on
Could an agent sign up, book a demo, or get a quote for its user? It checks every call to action and never submits a thing.
Action plan
Pass, Needs work, or Fail for each stage, with evidence, and a prioritized fix list with effort and owner.
What you get back
- A scorecard for each stage of the funnel: found, let in, understood, and acted on.
- Evidence for every finding: the URL, the status code, and the exact line.
- A prioritized fix list with effort (S, M, L) and who owns each fix.
- A ready-to-paste llms.txt Actions section written for your site, plus a full llms.txt if you have none.
- A table of what an AI agent can do on your site today, and after the fixes.
How to use it
Copy the skill once, then add it to whichever AI you already use.
Claude
- Click Customize in the sidebar
- Click Skills
- Click Create new skill
- Paste the copied skill content
- Ask: “Run an agent funnel audit on mysite.com”
ChatGPT
- Start a new chat with Search turned on
- Paste the copied skill
- Add your ask on the last line: “Audit mysite.com”
- To reuse it, create a GPT and upload the skill as a Knowledge file
Google Gemini
- Click Gems in the sidebar
- Click Create new Gem
- Paste the skill as the Gem’s instructions
- Name it and save
- Ask: “Run an agent funnel audit on mysite.com”
AI agent (OpenClaw, etc.)
- Save the skill as
SKILL.mdin your agent’s skills folder - For example
skills/agent-funnel-audit/SKILL.md - Agents with a shell get the fullest audit: they can test each bot with
curl - Ask: “Run an agent funnel audit on mysite.com”
Want more AI skills like this?
We build custom AI agents and skills for businesses. Subscribe for free tools every week, or let's talk about what AI can do for you.