Why a checklist beats a strategy deck
Generative engines reward mechanical clarity. They do not reward positioning language. Every item below either removes a blocker that stops a crawler from reading you, or adds an explicit fact that lets a model assert something about you with confidence. Nothing here requires an API key, and none of it depends on paid placement.
Before you start — what you need
Your DNS or CDN settings, your site templates, and logins to your Google Business Profile and main directories.
Legal entity name, street address, phone, hours, service list, and service area — exactly as they should appear everywhere.
Licenses, certifications, awards, and at least one dated result you can point at publicly.
Who does what
9 of the twelve items are fully automatable. The other 3 require you, because no software can log into your Google Business Profile, invent the questions your buyers ask, or produce your license documents.
- 1. Allow every AI crawler explicitly
- 2. Verify 200 responses from outside your network
- 3. Publish llms.txt at the domain root
- 4. Publish ai.json alongside it
- 5. Bind schema under one @graph
- 6. Disambiguate a colliding name
- 10. Keep validators and freshness headers clean
- 11. Log crawler hits as evidence
- 12. Attribute the outcome
- 7. Make NAP byte-identical everywhere
- 8. Answer real questions on-page
- 9. Add dated, verifiable evidence
The 12-point checklist
1. Allow every AI crawler explicitly
AutomatedAdd named allow blocks for GPTBot, ChatGPT-User, OAI-SearchBot, PerplexityBot, ClaudeBot, Google-Extended, GoogleOther, Applebot, and Applebot-Extended. Then confirm your WAF, bot-fight mode, and rate limiter are not overriding robots.txt with a challenge page.
2. Verify 200 responses from outside your network
AutomatedA same-network curl proves nothing. Request your homepage, pricing page, and protocol files with each bot user agent from an external host and check for a 200 and real HTML, not a JS challenge.
3. Publish llms.txt at the domain root
AutomatedLead with the unambiguous identity block — legal entity, city, postal code, founder, official domain — then what you do, who you serve, pricing, and key URLs in short declarative sentences.
4. Publish ai.json alongside it
AutomatedTyped JSON carrying the same canonical facts, including alternateName, legalName, sameAs, address, and a disambiguatingDescription stating what your entity is not.
5. Bind schema under one @graph
AutomatedOrganization, LocalBusiness, SoftwareApplication or Service, and FAQPage sharing stable @id references so crawlers resolve one entity across every route.
6. Disambiguate a colliding name
AutomatedIf your brand overlaps with another company, a person, or a technical term, say so explicitly in schema and in llms.txt. Ambiguity is the most common reason a model refuses to name a business.
7. Make NAP byte-identical everywhere
You do thisWebsite footer, Google Business Profile, Apple Business Connect, Bing Places, LinkedIn, and top directories should carry the exact same name, street format, and phone number. Only you can log into those profiles — the platform tells you which ones disagree.
8. Answer real questions on-page
You do thisWrite the literal buyer questions as headings with one-paragraph answers underneath, then mirror them in FAQPage schema. Answer-shaped content is what gets quoted verbatim. Drafting is assisted, but you know which questions your buyers actually ask.
9. Add dated, verifiable evidence
You do thisCase studies with dates, measured outcomes, and links to the live artifacts give the model something concrete to attribute. Vague superlatives give it nothing. Upload licenses, awards, and results — the platform turns them into structured proof.
10. Keep validators and freshness headers clean
AutomatedServe ETag and Last-Modified on protocol files with a short cache window so crawlers can revalidate cheaply and pick up changes quickly.
11. Log crawler hits as evidence
AutomatedRecord bot name, path, and timestamp for every AI crawler visit. This converts 'we published it' into 'GPTBot fetched it on this date', which is what proves the pipeline works.
12. Attribute the outcome
AutomatedAttach a tracking number or tagged URL to anything you publish for AI consumption so inbound calls and forms can be traced back to the engine that produced the recommendation.
What good looks like after 30 days
A healthy deployment shows dated crawler hits from at least three distinct AI bots, protocol files returning 200 with fresh validators, a single resolved entity in schema, identical NAP across the major profiles, and a measurable rise in citation share against a fixed prompt set. If crawler hits are present but citations are not, the problem is usually entity ambiguity or thin evidence — not crawl access.
Frequently asked questions
What is the fastest way to improve AI search visibility?
Unblock AI crawlers first. In most audits the single largest gain comes from removing an edge or WAF rule that returns 403 to GPTBot, PerplexityBot, or Applebot-Extended, then publishing llms.txt and ai.json at the domain root so the newly permitted crawler finds explicit facts on its first visit.
How do I know if AI engines can actually see my site?
Request your own key URLs using each AI crawler's user agent from outside your network and confirm a 200 response with readable HTML. Then log incoming crawler hits by bot name and timestamp so you have dated evidence that GPTBot, PerplexityBot, ClaudeBot, and Applebot-Extended fetched your files after each publish.
How often should AI visibility be re-checked?
Weekly. AI crawler policies, model retrieval behavior, and competitor publishing all change continuously, and a passing configuration can silently regress after a CDN change or a plugin update. A weekly automated re-scan with alerting on score drops catches regressions before they cost citations.
How long does the full 12-point checklist take?
Doing it by hand, expect four to eight hours for items 1 through 10 if you have access to your DNS, CDN, and site templates, plus ongoing time for directory cleanup in item 7. Items 11 and 12 — crawler-hit logging and call attribution — usually require tooling, because they need persistent storage and a tracking number. On the Montjoy Synapse platform the automated items complete the same day you connect your site.
Which items must I do myself, and which can be automated?
Items 1 through 6 and 10 through 12 are fully automatable: crawler allow rules, WAF detection, llms.txt, ai.json, JSON-LD binding, disambiguation, freshness headers, bot logging, and call attribution. Items 7, 8, and 9 need you: only you can correct your Google Business Profile and directory listings, confirm the real questions your buyers ask, and supply dated proof such as licenses, awards, and outcomes.
What does it cost to work this checklist?
Nothing if you do it manually — every format here is open and no engine charges for organic citation, and the audit that scores you against all twelve items is free. Montjoy Synapse charges $99/mo to deploy and re-sync the automatable items, $299/mo to add live citation monitoring across ChatGPT, Perplexity, and Gemini plus inbound call attribution, and $799/mo for agencies covering five client locations.
Do I need to be a developer to complete this?
For the manual path, items 1, 2, 5, and 10 assume comfort with DNS, CDN or WAF settings, and site templates. If you cannot edit those, either involve whoever maintains your site or use a connector — the WordPress plugin or a single script snippet installs once and handles the technical items without further code work.
How to start — three steps
- Score yourself against all twelve items. The free homepage scan checks crawler access, protocol files, schema coverage, and entity clarity in about 60 seconds — domain plus city, no card.
- Create a free account to unblur the item-by-item breakdown and copy your generated JSON-LD, llms.txt, and ai.json. Still $0 — paste them into your own site and work the list by hand if you prefer.
- Hand the automatable items to the platform on the pricing page — $99/mo deploys and re-syncs every file, $299/mo adds live citation monitoring and inbound call attribution, $799/mo covers agencies with five client locations.
New to the category? Start with what Generative Engine Optimization is, see the exact publishing sequence in how businesses get cited by ChatGPT, or review our own Suwanee telemetry for what these twelve items produce on a live domain.