# Local Wire | localwire.ca # # Public content site. Nothing here is disallowed to anyone. Every documented crawler is # named one by one, which is not decoration: robots.txt resolves the MOST SPECIFIC matching # group, so a per-bot Allow survives anything that later adds a Disallow to the * group. # # Each AI company splits its crawlers across three jobs, answered deliberately here: # may this content train models? yes # may an AI answer engine surface and cite it? yes # may a live user's question fetch it? yes # # Generated by the Blue Crane seo-canon robots generator. Every token below was read from # that vendor's own documentation on 2026-09-10, cited inline. Crawlers that publish no # user-agent (Bytespider, xAI's Grok fetcher, DeepSeek) are deliberately absent: they cannot # be addressed in this file, and naming a bot that documents no token would be guesswork. User-agent: * # Cloudflare Content Signals (contentsignals.org). Advisory, not enforcement. It declares # what may happen to the content AFTER a crawler is let in. # # THIS LINE MUST STAY INSIDE THIS GROUP, ABOVE ITS Allow. Content-Signal is a directive # scoped to the User-agent group it sits in, exactly like Allow and Disallow. Moved to the # foot of the file it silently binds to whichever group came last instead of to everyone, # which is a real defect this estate shipped on several hosts before 2026-09-10. # # search building a search index, returning links and short excerpts # ai-input feeding a live model, which is what a cited AI answer actually is # ai-train training or fine-tuning models # use immediate (reuse nothing) | reference (index, excerpt, LINK BACK) | full # (summarize and reproduce). reference is the one that asks for the citation. Content-Signal: search=yes, ai-input=yes, ai-train=yes, use=reference Disallow: /media Disallow: /tiktok Allow: / # ---- OpenAI ----------------- developers.openai.com/api/docs/bots # OAI-SearchBot is the token that decides ChatGPT search visibility, not GPTBot. OpenAI: # "Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers." # ChatGPT-User is user-triggered and OpenAI says robots.txt "may not apply" to it. User-agent: GPTBot User-agent: OAI-SearchBot User-agent: OAI-AdsBot User-agent: ChatGPT-User Disallow: /media Disallow: /tiktok Allow: / # ---- Anthropic -------------- support.claude.com/en/articles/8896518 # The only vendor here that honours robots.txt on its user-triggered fetcher as well. User-agent: ClaudeBot User-agent: Claude-SearchBot User-agent: Claude-User Disallow: /media Disallow: /tiktok Allow: / # ---- Anthropic (legacy tokens) - not in Anthropic's current doc; historical names still seen in server logs # Superseded by the three above. Kept because allowing a retired token costs nothing and # several sites on this estate had already granted them; never silently withdraw a # permission a site already gave. User-agent: anthropic-ai User-agent: Claude-Web Disallow: /media Disallow: /tiktok Allow: / # ---- Google ----------------- developers.google.com/search/docs/crawling-indexing/google-common-crawlers # Google-Extended governs Gemini training only. Google: "Google-Extended does not impact a # site's inclusion in Google Search nor is it used as a ranking signal in Google Search." User-agent: Googlebot User-agent: Googlebot-Image User-agent: Googlebot-News User-agent: Googlebot-Video User-agent: Google-Extended User-agent: Google-InspectionTool User-agent: Google-CloudVertexBot User-agent: GoogleOther User-agent: GoogleOther-Image User-agent: GoogleOther-Video User-agent: Storebot-Google Disallow: /media Disallow: /tiktok Allow: / # ---- Microsoft -------------- bing.com/webmasters/help/webmaster-guidelines-30fba23a # Bing matters twice: its own results, and ChatGPT's web search runs on it. No Bing index # means no ChatGPT citation, whatever the OpenAI tokens above say. User-agent: bingbot User-agent: BingPreview User-agent: msnbot User-agent: msnbot-media Disallow: /media Disallow: /tiktok Allow: / # ---- Perplexity ------------- docs.perplexity.ai/guides/bots # Perplexity states Perplexity-User "generally ignores robots.txt rules" because a user # asked for the fetch. Named anyway so the intent is on the record either way. User-agent: PerplexityBot User-agent: Perplexity-User Disallow: /media Disallow: /tiktok Allow: / # ---- Meta ------------------- developers.facebook.com/docs/sharing/webmasters/web-crawlers/ # facebookexternalhit is the link-preview fetcher: it renders the share card when a page is # posted to Facebook, Instagram or Threads, so it earns its place on any shared site. User-agent: meta-externalagent User-agent: meta-webindexer User-agent: meta-externalads User-agent: meta-externalfetcher User-agent: facebookexternalhit Disallow: /media Disallow: /tiktok Allow: / # ---- Apple ------------------ support.apple.com/en-us/119829 # Applebot-Extended controls Apple Intelligence training only and does not crawl. Apple: # pages that disallow it "can still be included in search results". User-agent: Applebot User-agent: Applebot-Extended Disallow: /media Disallow: /tiktok Allow: / # ---- Mistral ---------------- docs.mistral.ai/robots User-agent: MistralAI-Index User-agent: MistralAI-Training User-agent: MistralAI-User Disallow: /media Disallow: /tiktok Allow: / # ---- Amazon ----------------- developer.amazon.com/amazonbot # Also an IndexNow participant, so a deploy ping reaches it directly. User-agent: Amazonbot Disallow: /media Disallow: /tiktok Allow: / # ---- Common Crawl ----------- commoncrawl.org/ccbot # The open corpus much of the open-model ecosystem is built from. Blocking it has a wider # effect than blocking any single company's crawler. User-agent: CCBot Disallow: /media Disallow: /tiktok Allow: / # ---- Other search engines --- indexnow.org/searchengines.json # Yandex, Seznam and Naver are the remaining IndexNow participants; DuckDuckGo leans on # Bing's index but runs its own fetcher. User-agent: YandexBot User-agent: SeznamBot User-agent: Yeti User-agent: DuckDuckBot Disallow: /media Disallow: /tiktok Allow: / # ---- Also named on this site - carried over from the previous robots.txt so no permission is withdrawn # Not in the canon roster. Kept because this site had already named them. User-agent: Bytespider User-agent: DuckAssistBot User-agent: cohere-ai User-agent: YouBot Disallow: /media Disallow: /tiktok Allow: / # The exclusions above are repeated in every group on purpose. robots.txt resolves the # most specific matching group, so a named crawler with a bare Allow would otherwise be # invited into paths the * group excludes. Sitemap: https://localwire.ca/sitemap-index.xml Sitemap: https://localwire.ca/news-sitemap.xml