# Machine-readable summaries for LLMs: /llms.txt and /llms-full.txt # # The CityChat place database is produced by Applied AI sp. z o.o. (Przemyśl, PL), # its producer under the Polish Act of 27 July 2001 on the protection of databases. # Crawling and citing individual entries: welcome. Bulk extraction or re-utilisation # of a substantial part: not permitted. Licensing: kontakt@citychat.pl # # Text-and-data-mining rights are expressly reserved (art. 8a ust. 2 of that # Act, t.j. Dz.U. 2024 poz. 1769: for a database available online the # reservation is made in a machine-readable format). The reservation is the # Content-Signal line and the training-crawler group below. It lives in THIS # file on purpose: until 2026-09-27 it was left to Cloudflare's managed # robots.txt, which silently stopped being prepended while its zone setting # stayed on. Search and AI answer crawlers stay welcome; AI Crawl Control # blocks the same training crawlers at the edge. # AI training crawlers and training-use tokens (Google-Extended and # Applebot-Extended control training use of what Googlebot/Applebot fetch; # disallowing them costs no search visibility). User-agent: GPTBot User-agent: ClaudeBot User-agent: CCBot User-agent: Google-Extended User-agent: Applebot-Extended User-agent: Bytespider User-agent: Amazonbot User-agent: meta-externalagent Disallow: / # SEO-tool crawlers (2026-09-11): they index this site for other people's # SEO tools and give it nothing back, and once the ~67k rendered place pages # went into the sitemaps they became the heaviest traffic on the site # (AhrefsBot alone: 68,588 requests on 2026-09-10, every one a Pages Function # invocation). Every crawler listed here honours robots.txt. User-agent: AhrefsBot User-agent: SemrushBot User-agent: MJ12bot User-agent: DotBot User-agent: BLEXBot User-agent: DataForSeoBot User-agent: SERankingBacklinksBot # KeenableBot (keenable.ai, "AI visibility" SEO tooling) walked ~18k rendered # place pages 2026-09-14..17 (7,788 on 09-15 alone); it reads robots.txt. User-agent: KeenableBot Disallow: / # GoogleOther is Google's NON-search crawler (research and product use, not # indexing — Googlebot stays welcome). It rendered ~600 place pages 09-18..21, # each executing the page JS and counting an owner-stats view (2026-09-21). User-agent: GoogleOther User-agent: GoogleOther-Image User-agent: GoogleOther-Video Disallow: / # AI *search* index (not training — that is ClaudeBot, disallowed in the # training group above): welcome, because /llms.txt exists for exactly this, but paced. # Anthropic documents that Crawl-delay is honoured; 5 s caps it at ~17k # requests a day (45,908 on 2026-09-10 unpaced). User-agent: Claude-SearchBot Crawl-delay: 5 Allow: / User-agent: * Content-Signal: search=yes, ai-train=no Allow: / Sitemap: https://citychat.pl/sitemap.xml