# Niche — https://www.getniche.ai/robots.txt # # Permissive crawl policy. Niche wants to be both indexed by # traditional search engines and ingested by LLM training / # retrieval crawlers. The only paths off-limits are server-side # routes that have no business being in an index (POST endpoints, # preview/admin pages, OG image generation routes). User-agent: * Allow: / Disallow: /api/ # Content signals — declare how we want crawled content used. # Spec: https://contentsignals.org/ (draft-romm-aipref-contentsignals) # `search` and `ai-input` (AI search / RAG) are encouraged; we opt # out of `ai-train` (using our content to train foundation models) # while remaining permissive about traditional indexing and # AI-assisted answer surfaces. Content-Signal: search=yes, ai-input=yes, ai-train=no # --- AI training + retrieval crawlers -------------------------------- # Explicitly allow the major LLM providers so changes in their default # behavior won't accidentally block us. Each is its own block so a # single provider can be toggled off later without disturbing the rest. User-agent: GPTBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: OAI-SearchBot Allow: / User-agent: anthropic-ai Allow: / User-agent: ClaudeBot Allow: / User-agent: Claude-Web Allow: / User-agent: Claude-SearchBot Allow: / User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / User-agent: Google-Extended Allow: / User-agent: GoogleOther Allow: / User-agent: Bytespider Allow: / User-agent: Applebot Allow: / User-agent: Applebot-Extended Allow: / User-agent: CCBot Allow: / User-agent: meta-externalagent Allow: / User-agent: FacebookBot Allow: / User-agent: cohere-ai Allow: / User-agent: Diffbot Allow: / User-agent: YouBot Allow: / User-agent: AmazonBot Allow: / User-agent: ImagesiftBot Allow: / User-agent: Omgilibot Allow: / User-agent: DuckAssistBot Allow: / User-agent: MistralAI-User Allow: / # --- Discovery aids -------------------------------------------------- Sitemap: https://www.getniche.ai/sitemap-index.xml # Niche publishes /llms.txt (index) and /llms-full.txt (full content # dump) for LLM crawlers — see https://llmstxt.org for the spec.