# Grewal RE Group — robots.txt # Last updated: 2026-07-30 # # POLICY: "cited everywhere, trained nowhere." # # Every crawler that can put this site into an AI answer or a search result is # allowed. Every crawler whose purpose is harvesting content to TRAIN a model # is disallowed. That is the ai-train=no policy actually enforced, not merely # declared — see the four declarations listed at the foot of this file. # # Content Signals (https://contentsignals.org/): # search = listing this site in a search engine # ai-input = using this content to answer a user's query (AI search / RAG) # ai-train = using this content to train an AI/ML model # Our policy: search=yes, ai-input=yes, ai-train=no. Declared per agent below # (not just implied by crawl access), as a Content-Signal HTTP response header # on every response, and under the W3C TDM Reservation Protocol. # # WHY BLOCKING TRAINING DOES NOT COST CITATIONS # --------------------------------------------- # The major labs run SEPARATE agents for training vs. answering. Blocking the # training agent does not affect the answering one: # # OpenAI GPTBot = training (blocked) OAI-SearchBot + ChatGPT-User = answers (allowed) # Anthropic ClaudeBot = training (blocked) Claude-SearchBot + Claude-User = answers (allowed) # Google Google-Extended = Gemini training (blocked) Googlebot = Search + AI Overviews (allowed) # Apple Applebot-Extended = training (blocked) Applebot = Siri/Spotlight (allowed) # Meta Meta-ExternalAgent = training (blocked) Meta-ExternalFetcher = answers (allowed) # # Do NOT "simplify" this file by allowing everything. The split is the policy. # scripts/verify-ai-visibility.py fails the build if these lists drift. # ── Standard search engines ────────────────────────────────── # These power the search indexes that AI assistants cite from, so they matter # to AI visibility as much as the AI-native agents below. User-agent: Googlebot Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: Googlebot-Image Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: GoogleOther Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: Bingbot Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: Slurp Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: DuckDuckBot Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: Applebot Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: YandexBot Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: Baiduspider Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: PetalBot Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md # ── AI search, assistants, and live-fetch agents ───────────── # The citation pathway. Each of these either builds an AI search index or # fetches a page at answer time on a user's behalf. All allowed. # NOTE: Claude-User and Claude-SearchBot are Anthropic's CURRENT agents; # Claude-Web is the legacy identifier, kept allowed for compatibility. User-agent: ChatGPT-User Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: OAI-SearchBot Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: Claude-User Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: Claude-SearchBot Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: Claude-Web Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: PerplexityBot Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: Perplexity-User Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: MistralAI-User Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: DuckAssistBot Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: YouBot Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: Amazonbot Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: Meta-ExternalFetcher Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: AI2Bot Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: Timpibot Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: Diffbot Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: ImagesiftBot Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md # ── Social / link-preview agents ───────────────────────────── User-agent: FacebookBot Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: facebookexternalhit Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: Twitterbot Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md User-agent: LinkedInBot Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md # ── TRAINING CRAWLERS — DISALLOWED (enforces ai-train=no) ──── # These exist to harvest content into model-training corpora. Allowing any of # them would contradict the ai-train=no policy declared above. None of them # can cite this site — that is what the agents above do. # Training licences are available: shivraj.grewal@compass.com # ── AI TRAINING CRAWLERS ── generated by scripts/build-robots.py ── # Policy unchanged: ai-train=no. These crawlers exist to harvest text for # model training, which this site does not permit. Expressed as explicit # rules rather than a blanket wall, so the block reads as deliberate # policy and the terms themselves stay readable. # # Grouped rather than repeated thirteen times. One User-agent line per # crawler followed by one rule set is standard (RFC 9309 section 2.2.1) # and keeps the rules in one place instead of thirteen copies that can # drift apart. # # DO NOT hand-edit. Run scripts/build-robots.py --apply. --check proves, # with a Google-compatible parser, that this block blocks exactly what # the old `Disallow: /` blocked. User-agent: GPTBot User-agent: ClaudeBot User-agent: anthropic-ai User-agent: Google-Extended User-agent: Applebot-Extended User-agent: CCBot User-agent: Bytespider User-agent: cohere-ai User-agent: cohere-training-data-crawler User-agent: Meta-ExternalAgent User-agent: Ai2Bot-Dolma User-agent: Webzio-Extended User-agent: omgilibot Content-Signal: search=yes, ai-input=yes, ai-train=no # The terms under which this content is available, always readable. Allow: /robots.txt Allow: /ai.txt Allow: /ai-policy.json Allow: /licensing Allow: /licensing.html Allow: /.well-known/tdmrep.json # Everything else. The homepage needs $ because a bare / would be a # blanket rule again, and prefix matching means /about also covers # /about.html and /about.md. Disallow: /$ Disallow: /assets Disallow: /blog Disallow: /calculators Disallow: /communities Disallow: /data Disallow: /listings Disallow: /netlify Disallow: /questions Disallow: /relocation-edition Disallow: /schools-edition Disallow: /westlake-edition Disallow: /404 Disallow: /about Disallow: /accessibility Disallow: /austin-schools-guide Disallow: /buy Disallow: /contact Disallow: /faq Disallow: /index Disallow: /library Disallow: /listing Disallow: /photo-credits Disallow: /privacy Disallow: /relocation-guide Disallow: /reviews Disallow: /search Disallow: /sell Disallow: /terms Disallow: /thank-you Disallow: /westlake-78746-guide Disallow: /*.md Disallow: /*.json Disallow: /*.txt Disallow: /*.xml # ── end generated training-crawler block ── # ── Default policy: allow everything except internal docs ──── # SOURCES.md and AGENTS.md are intentionally crawlable: they are agent # resources. CLAUDE.md, AUTOMATION.md, and the internal editorial/handoff # docs below are internal and must never be indexed. User-agent: * Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /assets/guides/ Disallow: /CLAUDE.md Disallow: /AUTOMATION.md Disallow: /CONTENT-STYLE.md Disallow: /NAP-AUDIT.md Disallow: /WIKIDATA-DRAFT.md # ── Sitemaps + AI entry points ─────────────────────────────── Sitemap: https://grewalregroup.com/sitemap-index.xml Sitemap: https://grewalregroup.com/sitemap.xml Sitemap: https://grewalregroup.com/sitemap-images.xml # AI agents: machine-readable summary at /llms.txt (full detail at # /llms-full.txt), an MCP server at /mcp (discovery card at # /.well-known/mcp-server-card.json), and an A2A endpoint at /a2a (agent card # at /.well-known/agent-card.json). Both protocols expose the same skills # under the same ids. Prefer these over scraping HTML. # # AI usage policy — all four say the same thing, in the four formats # crawlers actually read. Change one, change all four: # https://grewalregroup.com/licensing (human-readable terms) # https://grewalregroup.com/ai-policy.json # https://grewalregroup.com/.well-known/tdmrep.json # https://grewalregroup.com/ai.txt # Content-Signal / tdm-reservation / tdm-policy HTTP response headers # # References: https://llmstxt.org/ https://contentsignals.org/ # https://www.w3.org/community/reports/tdmrep/