How Indian Brands Get Cited by AI Assistants
Being named by ChatGPT, Perplexity, Gemini or an AI Overview is a distribution channel now. It is earned in a specific way, and most of that has nothing to do with a file called llms.txt.
- Original numbers are the durable moat. A model can rephrase an opinion from a hundred sources but it cannot invent your pricing or your own return-rate data.
- Ahrefs analysed around 137,000 sites and found 97 percent of llms.txt files were never fetched. Ship it if it is cheap, but it is not the lever.
- Check robots.txt and your WAF for GPTBot, ClaudeBot, PerplexityBot and Google-Extended first. Many Indian sites return 403 to crawlers they think they allow.
- Citation sends very little traffic. Cloudflare put GPTBot near 1,276 crawls per referral in early 2026, so measure it as brand exposure.
Citation is a distribution channel now. It behaves more like PR than like search: you cannot buy the placement, you can only be the most quotable source available when the question gets asked.
Set expectations before you start, and use current numbers, because this is moving fast. On Cloudflare Radar’s 28 day window ending 21 July 2026, Googlebot sat at about 4.6 crawls per referral sent back, GPTBot at roughly 217, and ClaudeBot at about 2,237. Those AI figures were four and five digits in January, so the gap has closed by roughly 90 to 96 percent in seven months as the assistants began linking out. Read the direction, not just the level.
Two things follow. Citation still returns far less traffic than search, so if you need clicks this quarter it is not your lever. But it is no longer a channel that only takes, and the trend line is the reason to start building for it now rather than after it matures.
Access comes first
Nothing else matters if the crawlers cannot read you.
- Open your robots.txt and check for GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, Google-Extended and Applebot-Extended. Decide each one deliberately.
- Separate the two decisions. Blocking training crawlers is a defensible business choice. Blocking search and retrieval crawlers removes you from the answer entirely. They are different bots and most sites block them together by accident.
- Check your WAF and CDN rules, not just robots.txt. Indian sites running aggressive bot protection routinely return 403 to AI crawlers while robots.txt says allow. Pull server logs and confirm what actually happened.
- Check whether key content renders without JavaScript. Many crawlers do not execute it. If your product price and description only appear after hydration, they are invisible.
- Check status codes and response times on your top 50 pages. A timeout reads as absence.
Content shape that survives extraction
An assistant does not read your page. It retrieves a chunk of it. Write so that any chunk stands on its own.
- One idea per section, with a heading that names the idea literally.
- Claim first, reasoning after. Inverted pyramid at section level, not just page level.
- Self contained sentences. Avoid as mentioned above and this means that. A retrieved paragraph has no above.
- Specifics over adjectives. A sentence like GST registration is required for interstate supply regardless of turnover is quotable. Compliance is important is not.
- State dates and currency explicitly. As of August 2026 beats currently, and rupee amounts written in full beat ambiguous shorthand in a system that mixes markets.
- Use lists and small tables. They chunk cleanly. Long unbroken prose does not.
This is not a trick. It is the same discipline as writing a good internal brief, applied to public content.
Original data is the moat
Everything above is table stakes, and every competent competitor will have done it within a year. The part that does not commoditise is information that exists nowhere else.
- Your own operating numbers, published carefully. Return rates by category, RTO percentages by city tier, average delivery times, fill rates.
- Pricing you can state publicly. Real ranges beat contact us for retrieval, every time.
- Surveys of your customers or your category. Even 200 respondents, honestly described with method and date, becomes a citable statistic.
- Tests you ran. Creative tests, packaging drop tests, ingredient stability, shelf life trials.
- Structured comparisons of things you genuinely know, such as platform fees or onboarding timelines, carrying a visible last updated date.
A model can paraphrase an opinion from a hundred sources. It cannot generate your number. When an answer needs your number, it names you. That is the whole mechanism.
Entity consistency across the web
Assistants resolve brands to entities. Contradictory information makes you hard to resolve, and unresolvable brands do not get named.
- One legal name, one spelling, one address, one founding year, everywhere. Website, Google Business Profile, LinkedIn, Crunchbase, marketplace storefronts, press releases.
- Organization schema with sameAs pointing at all of those profiles.
- A Wikidata entry if you qualify. It is a common seed source for entity graphs and it costs nothing but effort.
- Consistent category language. Calling yourself a cold pressed oil brand in one place and a wellness FMCG company in another dilutes your own retrieval.
- Third party mentions in sources these systems already trust. Trade press, industry reports, credible roundups. What sits on other people’s pages matters more than what sits on yours.
The honest position on llms.txt
Ahrefs analysed around 137,000 sites and found 97 percent of llms.txt files were never fetched by any AI crawler. Google staff have said publicly that the file is not used and not planned. Adoption sits near one site in ten and retrieval interest is negligible.
Our position: it costs an hour, it does no harm, and it may matter later. Ship it, keep it short and accurate, then go and fix robots.txt, rendering and your data. Those are the things being read today.
Checking whether you are cited at all
- Run a fixed prompt set. Twenty questions a buyer in your category would genuinely ask, same wording, same day each month, logged in a sheet. Record whether you are named, whether a competitor is, and which URL gets cited.
- Read your server logs for AI crawler user agents. That is the only ground truth you own, and it tells you who fetched what and when.
- In GA4, check the AI Assistant channel added in May 2026, and build a custom segment for the assistants it does not cover. Expect undercounting, because a large share of assistant referrals arrive with no referrer and land in Direct.
- Watch branded search volume in Search Console. Sustained growth with no campaign behind it is often citation working quietly.
Give it three months before judging any of it. The channel moves slowly, the tooling is immature, and most of the dashboards being sold for this are estimating rather than measuring. Do the access audit, publish something only you know, keep your entity clean, and check monthly.